Disaster recovery method, device and equipment based on large model
By acquiring system operation data and business knowledge, and combining them with large-scale models to simulate defects, disaster models and disaster recovery models are generated. This solves the efficiency and accuracy problems of existing disaster recovery solutions in complex environments, realizes automated disaster recovery support, and ensures business continuity and data security.
Patent Information
- Application Number
- CN202510899911.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-04
AI Technical Summary
Existing disaster recovery solutions struggle to provide efficient and accurate disaster recovery support when facing complex and ever-changing system environments. Traditional IT system disaster recovery solutions rely on manual construction and lack automation and intelligent support, while general large-scale model disaster recovery solutions lack customized IT system knowledge support, resulting in deviations between the generated solutions and the actual IT architecture.
By acquiring system operation data and business knowledge, and combining them with large-scale models to simulate defects, disaster models and corresponding disaster recovery models are generated. When a disaster occurs in real time, these models are quickly matched and executed, realizing an automated process from solution formulation to emergency response, reducing the delay and bias of manual decision-making.
It improves the accuracy of disaster recovery solutions, enabling the generation of disaster and disaster recovery models that better fit the actual IT architecture, adapting to dynamic changes in the system environment, providing more efficient and reliable disaster recovery support, and ensuring business continuity and data security.
Smart Images

Figure CN120896836A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of data processing, and particularly relates to a large model-based disaster recovery method and device, equipment and a storage medium. BACKGROUND
[0002] IT (Information Technology) system disaster recovery technology refers to ensuring the continuity of business and the security of data by establishing a backup system or data backup to quickly switch to a backup system in the event of a main system failure or disaster.
[0003] In the prior art, IT system disaster recovery technology mainly includes a traditional IT system disaster recovery scheme and a general large model-based disaster recovery scheme. The traditional IT system disaster recovery scheme relies on manual construction of disaster recovery tools and development of strategies, and lacks automation and intelligent support. The general large model disaster recovery scheme lacks customized IT system disaster recovery knowledge support, resulting in a deviation between the generated disaster recovery scheme and the actual IT architecture.
[0004] Therefore, the existing disaster recovery scheme is difficult to provide efficient and accurate disaster recovery support when facing complex and variable system environments. SUMMARY
[0005] The purpose of the embodiments of the application is to provide a large model-based disaster recovery method, device, equipment and storage medium, which can solve the problem that the existing disaster recovery scheme is difficult to provide efficient and accurate disaster recovery support when facing complex and variable system environments.
[0006] In a first aspect, the embodiments of the application provide a large model-based disaster recovery method, which comprises:
[0007] obtaining system running data and business knowledge;
[0008] inputting the system running data and the business knowledge into a large model for defect simulation to generate a disaster model and a corresponding disaster recovery model, the disaster recovery model being used to dispose the corresponding disaster model;
[0009] After monitoring a real-time disaster, a target disaster model matched with the real-time disaster is determined, and a target disaster recovery model corresponding to the target disaster model is executed.
[0010] Optionally, the inputting the system running data and the business knowledge into a large model for defect simulation to generate a disaster model and a corresponding disaster recovery model comprises:
[0011] inputting the system running data and the business knowledge into a large model for defect simulation to generate a disaster model;
[0012] The disaster model is added to the execution queue, and the disaster models in the execution queue are executed sequentially to determine whether they cause operational failure.
[0013] If this leads to operational failure, a disaster recovery model corresponding to the disaster model will be generated.
[0014] Optionally, adding the disaster model to the execution queue further includes:
[0015] Generate a snapshot of the current system environment;
[0016] After generating the disaster recovery model corresponding to the disaster model, the process further includes:
[0017] Restore the environment based on the aforementioned runtime snapshot;
[0018] If no operational failure occurs, the environment is restored based on the aforementioned operational snapshot.
[0019] Optionally, the step of inputting the system operation data and the business knowledge into a large model for defect simulation to generate a disaster model includes:
[0020] Determine whether a disaster model exists that matches the system's operational data and the business knowledge;
[0021] If it does not exist, the system operation data and the business knowledge are input into the large model to simulate defects and generate a new disaster model;
[0022] If it exists, then determine whether there is a corresponding response plan for the disaster model. If there is a corresponding response plan, then generate a disaster recovery model based on the response plan.
[0023] Optionally, after generating the disaster recovery model corresponding to the disaster model, the method further includes:
[0024] After executing the disaster recovery model, determine whether the system defects caused by the disaster model have been corrected;
[0025] If not, return to the step of generating the disaster recovery model corresponding to the disaster model.
[0026] Optionally, before obtaining system operation data, the following steps are also included:
[0027] Obtain initial running data;
[0028] Based on the initial operating data, the business operating status is modeled to obtain a feature model, which is used to indicate the correlation between the initial operating data and preset business indicators.
[0029] Verify whether the correlation between the initial operating data and the preset business indicators is valid;
[0030] If invalid, return to the step of modeling the business operation status based on the initial running data to obtain the feature model;
[0031] If valid, the initial running data will be used as the system running data.
[0032] Optionally, if no target disaster model matches the real-time disaster, the method further includes:
[0033] The real-time disaster is analyzed to generate a corresponding real-time disaster recovery model.
[0034] Secondly, embodiments of this application provide a disaster recovery device based on a large model, the device comprising:
[0035] The acquisition module is used to acquire system operation data and business knowledge.
[0036] The analysis module is used to input the system operation data and the business knowledge into the large model to simulate defects, generate a disaster model and a corresponding disaster recovery model, and the disaster recovery model is used to handle the corresponding disaster model.
[0037] The handling module is used to monitor the occurrence of a real-time disaster, determine the target disaster model that matches the real-time disaster, and execute the target disaster recovery model corresponding to the target disaster model.
[0038] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0039] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0040] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0041] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0042] As can be seen from the above, this application firstly overcomes the limitations of traditional solutions that rely on manual intervention and the lack of customized IT disaster recovery knowledge in general large-scale models by acquiring system operation data and business knowledge and combining them with large-scale models for defect simulation. This allows for more accurate generation of disaster models and disaster recovery models that fit the actual IT architecture, improving the accuracy of disaster recovery solutions. Secondly, this method pre-generates disaster models and corresponding disaster recovery models and quickly matches and executes them when a disaster occurs in real time, realizing an automated process from solution formulation to emergency response, reducing the delay and bias of manual decision-making. Compared with traditional static strategies and automated solutions limited by preset logic, this method is more adaptable to the dynamic changes in the system environment, providing more efficient and reliable disaster recovery support, and effectively ensuring business continuity and data security. Attached Figure Description
[0043] Figure 1 This is a flowchart illustrating a disaster recovery method based on a large model according to an exemplary embodiment;
[0044] Figure 2 This is an overall architecture diagram of a disaster recovery method based on a large model, as illustrated in an exemplary embodiment.
[0045] Figure 3 This is a schematic diagram illustrating a disaster simulation scheme generation according to an exemplary embodiment;
[0046] Figure 4 This is a schematic diagram illustrating disaster recovery model generation using a large model, according to an exemplary embodiment.
[0047] Figure 5 This is a schematic diagram illustrating disaster response using a large model, according to an exemplary embodiment.
[0048] Figure 6 This is a block diagram illustrating a disaster recovery device based on a large model, according to an exemplary embodiment.
[0049] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment;
[0050] Figure 8 This is a schematic diagram of the hardware structure of an electronic device according to an exemplary embodiment. Detailed Implementation
[0051] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0052] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0053] First, let me explain the terms mentioned in this application:
[0054] Enhanced Knowledge Retrieval: A methodology for supplementing the static knowledge within a large model by dynamically retrieving external knowledge bases (such as structured databases, unstructured documents, etc.).
[0055] Dynamic context management: Compress historical context through the model itself or external tools (such as text summarization models) and dynamically discard redundant information.
[0056] ReACT (Reasoning & Acting): A cognitive framework that explicitly breaks down the reasoning process into an iterative cycle of "thinking-action" to achieve an interpretable task execution chain.
[0057] PaaS (Platform as a Service) technology: a cloud service paradigm that provides a standardized application development / runtime environment, covering a complete technology stack from middleware to operation and maintenance automation.
[0058] IaaS (Infrastructure as a Service) technology is a service-oriented delivery model for virtualized hardware resources, enabling on-demand provision of computing, storage, and network resources.
[0059] In related technologies, IT system disaster recovery technologies mainly include traditional IT system disaster recovery solutions and disaster recovery solutions based on general large-scale models. Among them, traditional IT system disaster recovery solutions rely on manual construction of disaster recovery tools and formulation of strategies, lacking automation and intelligent support; while general large-scale model disaster recovery solutions lack customized IT system disaster recovery knowledge support, resulting in deviations between the generated disaster recovery solutions and the actual IT architecture.
[0060] Therefore, existing disaster recovery solutions struggle to provide efficient and accurate disaster recovery support when facing complex and ever-changing system environments. Based on this, this application proposes a disaster recovery method based on a large model to address the aforementioned problems.
[0061] The disaster recovery method based on a large model provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0062] Figure 1 This is a flowchart illustrating a large-model-based disaster recovery method according to an exemplary embodiment, which includes the following steps.
[0063] In step S11, system operation data and business knowledge are acquired.
[0064] This step requires the comprehensive and accurate collection of various information related to the IT system's operational status and business logic, including but not limited to system operational data and business knowledge. System operational data encompasses a wide range of content, including not only multi-dimensional characteristic indicators such as network traffic, service response, security operations, and hardware usage, but also potentially system logs, performance monitoring data, and configuration information, reflecting the system's health status and potential risks in real time. Business knowledge focuses on the system's business processes, key services, data flow, user interaction patterns, and business rules.
[0065] By leveraging system operation data and business knowledge from both technical and business perspectives, a solid data foundation and contextual environment can be provided for subsequent in-depth defect simulation and model building of large models, ensuring that the model remains relevant to the actual system's operational scenarios and business objectives.
[0066] In step S12, system operation data and business knowledge are input into the large model to simulate defects, generate a disaster model and a corresponding disaster recovery model, and the disaster recovery model is used to handle the corresponding disaster model.
[0067] During the defect simulation and model generation phase, the acquired system operation data and business knowledge can be used as input to the large model. ReACT technology can be used to deeply explore the intrinsic relationships between the data, and based on pre-designed business defects, the potential defects that the system may have can be simulated. The consequences of these defects and their impact on business continuity can be evaluated by combining business knowledge.
[0068] Based on these analytical results, the large model can be abstracted into a specific "disaster model". A disaster model is a standardized and structured description of a specific type of failure or disaster scenario, which may include failure characteristics, triggering conditions, scope of impact, etc.
[0069] Furthermore, based on the simulation results of the disaster model, historical experience and business needs in the knowledge base can be combined to generate corresponding disaster recovery models. These disaster recovery models contain specific response strategies and operating procedures for the disaster model, which are used to deal with the problems caused by the corresponding disaster model, aiming to quickly and effectively mitigate or eliminate the impact of disasters and restore the normal operation of the system.
[0070] This process transforms data insights into actionable disaster recovery plans, enabling a shift from reactive response to proactive defense and planning.
[0071] In step S13, after a real-time disaster is detected, a target disaster model matching the real-time disaster is determined, and the target disaster recovery model corresponding to the target disaster model is executed.
[0072] During the disaster response and disaster recovery execution phase, the system needs to continuously monitor its operational status. Once a real-time disaster that meets the preset anomaly threshold is detected, subsequent processes will be triggered. At this point, the system will quickly compare and match the current real-time disaster information (such as specific error codes, outlier performance metrics, etc.) with all disaster models previously generated by the large model. Through pattern recognition and similarity calculation, it determines which pre-generated disaster model best represents the current actual situation; this selected model is the "target disaster model."
[0073] Once a match is successful, the system will automatically and quickly execute the "target disaster recovery model" associated with the target disaster model. The target disaster recovery model defines a series of operation sequences, which may include automatically switching to a backup server, isolating faulty components, starting a data recovery process, and adjusting system resource configuration, in order to correct system defects caused by real-time disasters.
[0074] This achieves an automated closed loop from disaster identification to emergency response, greatly shortening fault recovery time, reducing the complexity and potential errors of manual intervention, and ensuring that disaster recovery measures can be implemented in a timely and accurate manner.
[0075] like Figure 2 The diagram shown illustrates the overall architecture of this application, which is divided into two main modules: a large model service and a disaster recovery service. The large model service primarily includes dynamic context management, large model inference, knowledge retrieval enhancement, and capability invocation via ReACT. The disaster recovery service provides the large model service with capabilities for feature acquisition, PaaS resource manipulation, IaaS resource manipulation, and knowledge base querying. The PaaS resource manipulation capabilities mainly target traffic services, application services, and data services; the IaaS resource manipulation capabilities mainly target the domain name access layer and infrastructure.
[0076] In one implementation, in step S12, system operation data and business knowledge are input into a large model for defect simulation to generate a disaster model and a corresponding disaster recovery model, including:
[0077] Input system operation data and business knowledge into a large model to simulate defects and generate a disaster model;
[0078] Add disaster models to the execution queue and execute the disaster models in the queue in sequence to determine whether they cause operational failures.
[0079] If this leads to operational failure, a disaster recovery model corresponding to the disaster model will be generated.
[0080] In this implementation, firstly, the large model uses system operation data and business knowledge as the basis for analysis. With the help of predefined potential business defects, disaster models such as "cache penetration" and "forced instance shutdown" are constructed to simulate various failure scenarios that the system may face.
[0081] Subsequently, the generated disaster models are added to the execution queue in sequence, and the system executes these models one by one according to the queue order for simulation testing. During execution, the system monitors the business operation status in real time to determine whether the disaster model simulation will cause operational failures, such as business system response timeouts or database connection anomalies.
[0082] If the simulated disaster causes system malfunctions, it indicates that the disaster scenario will indeed have a real impact on business operations. In this case, the large model will generate a disaster recovery model based on the specific circumstances of the simulation, combined with historical experience from the knowledge base and business requirements. This disaster recovery model covers specific response strategies and operational procedures for the disaster model, aiming to quickly and effectively handle problems and ensure business continuity and system stability when a real disaster occurs.
[0083] In one implementation, adding the disaster model to the execution queue also includes:
[0084] Generate a snapshot of the current system environment;
[0085] After generating the disaster recovery model corresponding to the disaster model, it also includes:
[0086] Environment restoration based on runtime snapshots;
[0087] If no operational failure occurs, the environment will be restored based on the operational snapshot.
[0088] In other words, during the process of adding disaster models to the execution queue, in order to ensure the safety and stability of the system when simulating disaster scenarios, the system will pre-generate a snapshot of the current operating environment, fully recording the system's various configurations, data status and operating parameters before the simulation, so that it can quickly backtrack in case of unexpected situations.
[0089] After executing a disaster model simulation, the system will take different actions based on the simulation results. If the simulated disaster causes system malfunction, it means that the simulation exposed risks in the system's ability to cope with such disasters. In this case, the large model will generate a corresponding disaster recovery model. After the disaster recovery model is generated, the system will accurately restore the system environment to its pre-simulation state based on a pre-saved operational snapshot. This ensures that the disaster recovery model is based on the results of real simulations and avoids the simulation process from having a lasting impact on the normal operation of the system, providing a clean and stable environmental foundation for the subsequent verification and optimization of disaster recovery solutions.
[0090] If the disaster model simulation does not cause operational failures, environmental restoration will still be performed based on the operational snapshot. This is because even if the simulation does not cause significant failures, the system's state and data have changed during the simulation. By restoring the operational snapshot, the temporary impact of the simulation operation can be eliminated, allowing the system to return to normal operation. This also prepares the system for continuing to execute other disaster model simulations or handle actual business requests, ensuring that the system is always in a reliable and controllable operating environment.
[0091] In one implementation, system operation data and business knowledge are input into a large model for defect simulation to generate a disaster model, including:
[0092] Determine whether a disaster model exists that matches the system's operational data and business knowledge;
[0093] If it does not exist, the system operation data and business knowledge will be input into the large model to simulate defects and generate a new disaster model;
[0094] If it exists, then determine whether there is a corresponding disaster response plan. If there is a corresponding response plan, then generate a disaster recovery model based on the response plan.
[0095] In this implementation, the first step is to determine whether a known disaster model that matches the current system operation data and business knowledge exists. This determination process can be achieved by checking whether the current system state has similar characteristics to a certain disaster model in the historical records, and the specific method is not limited.
[0096] If no matching known disaster model is found, it indicates that the potential risks faced by the system have not yet been recorded and summarized. In this case, the large-scale model will delve into system operational data and business knowledge to conduct comprehensive defect simulations at the business design, architecture design, and hardware levels. Based on these analyses, a new disaster model is constructed to simulate potential risk scenarios the system may encounter. For example, it simulates potential service crashes caused by undiscovered system architecture vulnerabilities, providing a basis for subsequent disaster recovery strategy development.
[0097] If a matching known disaster model is identified, the system will further verify whether a corresponding mitigation plan exists. If a mature mitigation plan already exists, it indicates that past experience has formed an effective response strategy for this type of disaster. In this case, the large model will directly generate a corresponding real-time disaster recovery model based on the mitigation plan and the current actual operating status of the system.
[0098] This approach allows for full utilization of historical experience, avoids duplication of effort, and ensures that the real-time disaster recovery model is practically feasible and effective, enabling a rapid response in the event of a disaster and guaranteeing business continuity.
[0099] In one implementation, after generating the disaster recovery model corresponding to the disaster model, the following steps are also included:
[0100] After executing the disaster recovery model, determine whether the system defects caused by the disaster model have been corrected;
[0101] If not, return to the steps for generating the disaster recovery model corresponding to the disaster model.
[0102] In other words, after successfully generating the disaster recovery model corresponding to the disaster model, the entire disaster recovery process enters the critical stage of practical verification and optimization iteration. The system can execute the generated disaster recovery model and handle the simulated disaster according to the preset strategies and procedures. After execution, the system will conduct a comprehensive evaluation of the handling effect to determine whether the risks simulated by the disaster model have been effectively eliminated, that is, whether the system defects caused by the disaster model have been corrected.
[0103] If the assessment results show that the disaster was not handled properly, it means that the current disaster recovery model has shortcomings and has failed to effectively cope with the simulated disaster scenario. At this time, the system will go back to the stage of generating the disaster recovery model. Based on the feedback results of this execution, the large model will re-examine the problems exposed during the disaster simulation, combine system operation data, business knowledge, and any new information that may be acquired, and optimize and adjust the disaster recovery strategy again to generate a new disaster recovery model.
[0104] In this way, through continuous practice, evaluation, and improvement, we can ensure that the disaster recovery model can accurately and efficiently handle actual disaster problems, gradually improve the system's ability to cope with risks, and ultimately form a mature and reliable disaster recovery solution that can effectively guarantee business continuity.
[0105] In one implementation, before obtaining system operation data in step S11, the method further includes:
[0106] Obtain initial running data;
[0107] Based on the initial operating data, the business operation status is modeled to obtain a feature model, which is used to indicate the correlation between the initial operating data and the preset business indicators.
[0108] Verify the validity of the correlation between the initial operating data and the preset business indicators;
[0109] If invalid, return to the steps of modeling the business operation status based on the initial running data to obtain the feature model;
[0110] If valid, the initial running data will be used as the system running data.
[0111] In other words, before acquiring system operation data, initial operation data can be collected first. Initial operation data covers information from multiple dimensions such as network traffic, business response, security operation, and hardware usage, recording the basic status of the system's daily operation.
[0112] Next, based on this initial operational data, the large model can model the operational status of the business and construct a feature model. This feature model aims to reveal the intrinsic relationship between the initial operational data and preset business indicators, such as analyzing the correlation between server CPU (Central Processing Unit) utilization and business processing speed, or the correlation between network bandwidth utilization and data transmission efficiency, and so on. Through mathematical analysis and algorithmic calculations, the feature model presents these complex relationships in a clear and quantitative form.
[0113] Furthermore, the correlation between the initial operating data and the preset business indicators can be verified. By comparing the actual data performance with the model prediction results, it can be evaluated whether the model accurately reflects the system's operating rules. If the verification finds that the correlation is invalid, it indicates that the feature model has a bias and cannot truly reflect the system's condition. At this point, the system will return to the modeling step, readjust the analysis perspective and algorithm parameters, and rebuild the feature model until a reasonable and effective model is obtained.
[0114] When the verification results show that the correlation is valid, it means that the feature model can accurately describe the system's operating characteristics. At this point, the initial operating data is recognized as a reliable data source and is officially used as system operating data for subsequent core processes such as defect simulation based on large models, disaster model and disaster recovery model generation, thus laying a solid data foundation for the formulation of the entire disaster recovery plan.
[0115] In one implementation, if in step S14 there is no target disaster model that matches the real-time disaster, the method further includes:
[0116] Perform disaster analysis on real-time disasters and generate corresponding real-time disaster recovery models.
[0117] In other words, during the execution of a disaster recovery process based on a large model, when the system detects a real-time disaster, the primary task is to find a matching target disaster model in the existing disaster model library. However, in reality, a special situation may arise: the system fails to find a target disaster model that perfectly matches the characteristics of the real-time disaster. This situation often indicates that the disaster is unique or a new type of risk, exceeding the pre-defined scope of disaster simulation.
[0118] Once it is confirmed that no matching target disaster model exists, the system will quickly activate the emergency response mechanism to conduct a comprehensive and in-depth analysis of the real-time disaster. This process can be implemented using a large-scale model. The large-scale model will comprehensively utilize system operational data, business knowledge, and experience from historical disaster cases to analyze the characteristics and potential hazards of the disaster from multiple dimensions, including the source of the disaster, its scope of impact, and potential chain reactions. Based on the analysis results, the large-scale model, with its powerful algorithms and reasoning capabilities, will quickly generate a real-time disaster recovery model specifically for this real-time disaster.
[0119] This allows for the immediate containment of disasters, minimizing their impact on business continuity and data security, ensuring that disaster recovery solutions can adapt to complex and ever-changing disaster scenarios, and providing comprehensive and seamless security for IT systems.
[0120] Figure 3 This is a schematic diagram illustrating a disaster simulation scheme generation according to an exemplary embodiment, including the following process steps:
[0121] First, it is necessary to acquire system operation data in parallel, including static operation data, such as artifact code, physical architecture, physical location, and operation environment information; and dynamic operation data, such as operation logs, resource snapshots, and service call topology.
[0122] Acquire business knowledge, including business architecture description, business metrics, and first-level business operation status;
[0123] Then, by combining the large model with static and dynamic operational data and business knowledge to analyze defects, defect simulation is achieved.
[0124] Specifically, disaster models can be generated for pre-designed business defects. These business defects include business design defects, architecture design defects, and hardware layer defects, as shown in Table 1. Among them, business design defects involve high-concurrency traffic failures, data consistency failures, third-party service dependency failures, and compliance risks of business logic vulnerabilities; architecture design defects involve unreasonable cache settings, message queue overload, gateway single point of failure, and single instance deployment; and hardware layer defects involve hardware aging, network device configuration errors, etc.
[0125]
[0126]
[0127] Table 1. Simulation Approach by Defect Type Level
[0128] Then, add all disaster models to the execution queue;
[0129] Backup runtime snapshots for use in restoring the runtime environment;
[0130] Retrieve disaster models from the queue to be executed and inject faults;
[0131] The effectiveness of the simulation results was assessed using a large model.
[0132] If the proposed solution can affect business operations, a disaster simulation solution will be generated, and environmental restoration will be performed using operational snapshots; if the proposed solution has no impact on business operations, environmental restoration will be performed directly.
[0133] If the queue to be executed is empty, the disaster simulation scheme generation ends; if it is not empty, return to the step of retrieving the disaster model from the queue to be executed and performing fault injection.
[0134] Figure 4 This is a schematic diagram illustrating disaster recovery model generation using a large model, according to an exemplary embodiment, including the following process steps:
[0135] First, the operational characteristics of the business itself need to be modeled according to the characteristics shown in Table 2;
[0136]
[0137] Table 2 Feature Parameters
[0138] After modeling is completed, the main characteristics of business operation are analyzed through the large model, feature engineering is carried out, and the model is output.
[0139] Data is collected and preprocessed according to the dimensions of the model, and the large model automatically calls relevant mathematical tools through ReACT technology to determine whether the feature construction is accurate.
[0140] If it is inaccurate, reanalyze the main features using a larger model;
[0141] If the features are accurate, they need to be aggregated according to the time dimension, and the existence of disasters corresponding to this type of feature needs to be queried from historical disasters.
[0142] If a disaster has already occurred for this type of business, determine whether there are corresponding manual handling methods.
[0143] If corresponding manual handling methods already exist, they will be summarized and optimized by the large model to form a disaster recovery model;
[0144] Disaster simulation schemes are generated from large models, followed by the simultaneous generation of disaster recovery models and integration of corresponding tools (if a corresponding disaster recovery model already exists, it is executed directly), as well as disaster recovery environment construction and traffic simulation.
[0145] The large model calls upon tools according to the disaster recovery model to implement specific disaster recovery operations;
[0146] The large model determines whether the disaster has been handled correctly. If it has not been handled correctly, the corresponding disaster recovery model is regenerated.
[0147] If the issues have been handled correctly, the large model will determine whether all critical vulnerabilities in the system have been identified.
[0148] If the large model analysis determines that not all vulnerability analyses have been completed, then the large model is combined with business characteristics and operating environment to analyze business disaster risks, and a disaster recovery plan is generated based on the disaster risks. This process is repeated to identify all system vulnerabilities.
[0149] If there is no disaster specific to this type of business, then conduct a disaster risk analysis directly and generate a disaster simulation plan based on the disaster risk.
[0150] Figure 5 This is a schematic diagram illustrating disaster response using a large model according to an exemplary embodiment, including the following process steps:
[0151] First, the large model summarizes the current disaster models and searches for corresponding disaster recovery models;
[0152] The usability of the current disaster recovery model is determined by verifying whether the similarity between the model's preset disaster scenario and the actual disaster situation exceeds 95%.
[0153] If a highly similar model exists, the larger model will call the disaster recovery model to perform disaster recovery operations.
[0154] If no available disaster recovery model exists, a large model is used to analyze the current disaster situation, generate a disaster recovery model, and execute it.
[0155] The next step is to analyze whether the disaster recovery was successful. If it was successful, the disaster recovery is completed. If it was unsuccessful, the disaster recovery model is analyzed, optimized, and executed until the disaster recovery is completed.
[0156] As can be seen from the above, this application firstly overcomes the limitations of traditional solutions that rely on manual intervention and the lack of customized IT disaster recovery knowledge in general large-scale models by acquiring system operation data and business knowledge and combining them with large-scale models for defect simulation. This allows for more accurate generation of disaster models and disaster recovery models that fit the actual IT architecture, improving the accuracy of disaster recovery solutions. Secondly, this method pre-generates disaster models and corresponding disaster recovery models and quickly matches and executes them when a disaster occurs in real time, realizing an automated process from solution formulation to emergency response, reducing the delay and bias of manual decision-making. Compared with traditional static strategies and automated solutions limited by preset logic, this method is more adaptable to the dynamic changes in the system environment, providing more efficient and reliable disaster recovery support, and effectively ensuring business continuity and data security.
[0157] The disaster recovery method based on a large model provided in this application can be executed by a disaster recovery device based on a large model. This application uses the method of data storage executed by a disaster recovery device based on a large model as an example to illustrate the apparatus of the disaster recovery method based on a large model provided in this application.
[0158] Figure 6 This is a block diagram of a disaster recovery device based on a large model, illustrated according to an exemplary embodiment, comprising:
[0159] Module 301 is used to acquire system operation data and business knowledge;
[0160] Analysis module 302 is used to input the system operation data and the business knowledge into the large model to simulate defects, generate a disaster model and a corresponding disaster recovery model, and the disaster recovery model is used to handle the corresponding disaster model;
[0161] The handling module 303 is used to detect a real-time disaster, determine a target disaster model that matches the real-time disaster, and execute the target disaster recovery model corresponding to the target disaster model.
[0162] As can be seen from the above, this application firstly overcomes the limitations of traditional solutions that rely on manual intervention and the lack of customized IT disaster recovery knowledge in general large-scale models by acquiring system operation data and business knowledge and combining them with large-scale models for defect simulation. This allows for more accurate generation of disaster models and disaster recovery models that fit the actual IT architecture, improving the accuracy of disaster recovery solutions. Secondly, this method pre-generates disaster models and corresponding disaster recovery models and quickly matches and executes them when a disaster occurs in real time, realizing an automated process from solution formulation to emergency response, reducing the delay and bias of manual decision-making. Compared with traditional static strategies and automated solutions limited by preset logic, this method is more adaptable to the dynamic changes in the system environment, providing more efficient and reliable disaster recovery support, and effectively ensuring business continuity and data security.
[0163] The disaster recovery method based on a large model provided in this application can be executed by a terminal access terminal. This application uses the terminal access terminal executing the terminal access method as an example to illustrate the apparatus of the disaster recovery method based on a large model provided in this application.
[0164] The disaster recovery device based on a large model in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific implementation.
[0165] The disaster recovery device based on a large model provided in this application embodiment can achieve... Figures 1 to 5 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0166] Optionally, such as Figure 7As shown, this application embodiment also provides an electronic device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they implement the various steps of the above-described disaster recovery method embodiment based on a large model and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0167] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0168] Figure 8 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0169] The electronic device 1000 includes, but is not limited to, components such as: radio frequency unit 1001, network module 1002, audio output unit 1003, input unit 1004, sensor 1005, display unit 1006, user input unit 1007, interface unit 1008, memory 1009, and processor 1010.
[0170] Those skilled in the art will understand that the electronic device 1000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 8 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0171] As can be seen from the above, this application firstly overcomes the limitations of traditional solutions that rely on manual intervention and the lack of customized IT disaster recovery knowledge in general large-scale models by acquiring system operation data and business knowledge and combining them with large-scale models for defect simulation. This allows for more accurate generation of disaster models and disaster recovery models that fit the actual IT architecture, improving the accuracy of disaster recovery solutions. Secondly, this method pre-generates disaster models and corresponding disaster recovery models and quickly matches and executes them when a disaster occurs in real time, realizing an automated process from solution formulation to emergency response, reducing the delay and bias of manual decision-making. Compared with traditional static strategies and automated solutions limited by preset logic, this method is more adaptable to the dynamic changes in the system environment, providing more efficient and reliable disaster recovery support, and effectively ensuring business continuity and data security.
[0172] It should be understood that, in this embodiment, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0173] The memory 1009 can be used to store software programs and various data. The memory 1009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0174] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor 1010.
[0175] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described disaster recovery method embodiment based on a large model and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0176] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0177] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described disaster recovery method embodiment based on a large model, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0178] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0179] This application provides a computer program product stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the disaster recovery method embodiment based on the large model described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0180] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0181] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0182] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A disaster recovery method based on a large model, characterized in that, The method includes: Acquire system operation data and business knowledge; The system operation data and business knowledge are input into the large model to simulate defects, generate a disaster model and a corresponding disaster recovery model, and the disaster recovery model is used to handle the corresponding disaster model. After a real-time disaster is detected, a target disaster model matching the real-time disaster is determined, and the target disaster recovery model corresponding to the target disaster model is executed.
2. The disaster recovery method based on a large model according to claim 1, characterized in that, The step of inputting the system operation data and the business knowledge into a large model for defect simulation to generate a disaster model and a corresponding disaster recovery model includes: The system operation data and business knowledge are input into a large model to simulate defects and generate a disaster model. The disaster model is added to the execution queue, and the disaster models in the execution queue are executed sequentially to determine whether they cause operational failure. If this leads to operational failure, a disaster recovery model corresponding to the disaster model will be generated.
3. The disaster recovery method based on a large model according to claim 2, characterized in that, Adding the disaster model to the execution queue also includes: Generate a snapshot of the current system environment; After generating the disaster recovery model corresponding to the disaster model, the process further includes: Restore the environment based on the aforementioned runtime snapshot; If no operational failure occurs, the environment is restored based on the aforementioned operational snapshot.
4. The disaster recovery method based on a large model according to claim 2, characterized in that, The step of inputting the system operation data and the business knowledge into a large model for defect simulation and generating a disaster model includes: Determine whether a disaster model exists that matches the system's operational data and the business knowledge; If it does not exist, the system operation data and the business knowledge are input into the large model to simulate defects and generate a new disaster model; If it exists, then determine whether there is a corresponding response plan for the disaster model. If there is a corresponding response plan, then generate a disaster recovery model based on the response plan.
5. The disaster recovery method based on a large model according to claim 2, characterized in that, After generating the disaster recovery model corresponding to the disaster model, the process further includes: After executing the disaster recovery model, determine whether the system defects caused by the disaster model have been corrected; If not, return to the step of generating the disaster recovery model corresponding to the disaster model.
6. The disaster recovery method based on a large model according to claim 1, characterized in that, Before obtaining system operation data, the following is also included: Obtain initial running data; Based on the initial operating data, the business operating status is modeled to obtain a feature model, which is used to indicate the correlation between the initial operating data and preset business indicators. Verify whether the correlation between the initial operating data and the preset business indicators is valid; If invalid, return to the step of modeling the business operation status based on the initial running data to obtain the feature model; If valid, the initial running data will be used as the system running data.
7. The disaster recovery method based on a large model according to claim 1, characterized in that, If no target disaster model matches the real-time disaster, the method further includes: The real-time disaster is analyzed to generate a corresponding real-time disaster recovery model.
8. A disaster recovery device based on a large model, characterized in that, The device includes: The acquisition module is used to acquire system operation data and business knowledge. The analysis module is used to input the system operation data and the business knowledge into the large model to simulate defects, generate a disaster model and a corresponding disaster recovery model, and the disaster recovery model is used to handle the corresponding disaster model. The handling module is used to monitor the occurrence of a real-time disaster, determine the target disaster model that matches the real-time disaster, and execute the target disaster recovery model corresponding to the target disaster model.
9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the disaster recovery method based on a large model as described in any one of claims 1 to 7.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the disaster recovery method based on a large model as described in any one of claims 1-7.