Disaster recovery data health monitoring method and device, electronic equipment and storage medium
By extracting AI feature codes and business characteristic information from disaster recovery data in the cloud storage environment and performing consistency calculations, the shortcomings of disaster recovery data health checks are addressed, thereby improving recovery efficiency and system reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA TELECOM CLOUD TECH CO LTD
- Filing Date
- 2025-11-26
- Publication Date
- 2026-04-17
AI Technical Summary
The lack of effective disaster recovery data health check methods in existing cloud storage environments leads to the risk of incomplete or unavailable disaster recovery data when customers attempt to restore business operations in the event of production system failures, resulting in low recovery efficiency.
By acquiring scheduling instructions for disaster recovery data, extracting AI feature codes and obtaining corresponding business feature information, determining the health status of disaster recovery data based on the consistency results, and using a cloud-based AI big data system for multi-dimensional detection, health status information and early warnings are generated.
It improved the health monitoring capabilities of disaster recovery data, reduced blind attempts during business recovery, lowered the risk of business interruption, and enhanced customer acceptance of the cloud system.
Smart Images

Figure CN121880098A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of disaster recovery data health monitoring technology, and in particular to a disaster recovery data health monitoring method, a disaster recovery data health monitoring device, an electronic device, and a readable storage medium. Background Technology
[0002] Disaster recovery refers to the use of scientific and technological means and methods to proactively establish emergency response mechanisms for data security risks such as data loss, corruption, or malware infection. It typically includes a disaster recovery system and a backup system. Backup is the process of saving the data set of a business system from the production array to other storage media. Therefore, disaster recovery data refers to the data set that is saved or transferred to the backup / disaster recovery system within the entire disaster recovery system to ensure data security and business continuity. In various disaster recovery applications and levels, the security and availability of disaster recovery data are paramount.
[0003] Current cloud storage environments offer a wide range of disaster recovery solutions, such as cloud server disaster recovery, local data cloud disaster recovery, and file disaster recovery, which have been widely adopted and developed. However, these technologies lack effective methods for health checks on disaster recovery data. As a result, when customers attempt to restore their businesses using backup data after a production system failure, they may face the risk of incomplete, corrupted, or unusable disaster recovery data, leading to low recovery efficiency. Summary of the Invention
[0004] The present invention provides a method, apparatus, electronic device, and readable storage medium for monitoring the health of disaster recovery data, in order to overcome or at least partially solve the above-mentioned problems.
[0005] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide a method for monitoring the health of disaster recovery data, including: Obtain scheduling instructions for disaster recovery data, extract AI feature codes from the original parameters of the scheduling instructions, and obtain business feature information corresponding to the AI feature codes; The target disaster recovery data is determined, and the health status of the target disaster recovery data is determined based on the consistency between the target disaster recovery data and the business characteristic information.
[0006] Optionally, before the step of extracting AI feature codes from the original parameters of the scheduling instruction, the method further includes: Filter key information from the original parameters of the scheduling instruction to directly drive the execution of subsequent tasks and resource scheduling; The key information is used to generate a verification task.
[0007] Optionally, the step of obtaining business feature information corresponding to the AI feature code includes: The verification task is divided into multiple sub-tasks based on the historical availability of cloud resources. The subtasks are stored in the scheduling pool; Determine the current idle status of cloud resources. When the cloud resources meet the preset conditions based on the current idle status of cloud resources, extract the target subtask from the scheduling pool. When executing the target sub-task, AI feature codes are extracted from the original parameters corresponding to the target sub-task, and business feature information corresponding to the AI feature codes is obtained.
[0008] Optionally, the step of determining the health status of the target disaster recovery data based on the consistency result between the target disaster recovery data and the business characteristic information includes: Based on the business feature information, the matching degree of the target disaster recovery data is calculated to generate a matching degree result that expresses the matching degree between the target disaster recovery data and the business feature information; The final verification result of the overall disaster recovery data is generated by using the consistency results of all target disaster recovery data. The health status of the disaster recovery data is determined using the final verification result.
[0009] Optionally, the step of generating the final verification result of the overall disaster recovery data by using the consistency results of all target disaster recovery data includes: When the matching result is greater than or equal to a preset threshold, the health status of the target disaster recovery data is determined to be good, and first health status information is generated to express that the health status of the target disaster recovery data is good. When the matching result is less than a preset threshold, the health status of the target disaster recovery data is determined to be questionable, and second health status information is generated to express that the health status of the target disaster recovery data is questionable. The final verification result of the overall disaster recovery data is generated using at least the first health status information and the second health status information.
[0010] Optionally, it also includes: Generate an early warning message for the second health status information.
[0011] Optionally, the step of obtaining business feature information corresponding to the AI feature code includes: Access the cloud-based big data center and retrieve business feature information corresponding to the AI feature code from the cloud-based big data center.
[0012] Secondly, embodiments of this application provide a disaster recovery data health monitoring device, comprising: The business feature information acquisition module is used to acquire scheduling instructions for disaster recovery data, extract AI feature codes from the original parameters of the scheduling instructions, and acquire business feature information corresponding to the AI feature codes. The health status generation module is used to determine the target disaster recovery data and, based on the consistency result between the target disaster recovery data and the business characteristic information, determine the health status of the target disaster recovery data.
[0013] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0014] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0015] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0016] The embodiments of the present invention have the following advantages: In this embodiment of the invention, by acquiring scheduling instructions for disaster recovery data, extracting AI feature codes from the original parameters of the scheduling instructions, and obtaining business feature information corresponding to the AI feature codes; determining target disaster recovery data, and determining the health status of the target disaster recovery data based on the consistency result between the target disaster recovery data and the business feature information, the invention enhances the multi-dimensional detection capabilities based on the AI system, solving the shortcoming of traditional backup solutions in providing customer data health monitoring, thereby greatly improving customer acceptance of the cloud system and avoiding increased business interruption time caused by blind attempts during business recovery. Attached Figure Description
[0017] Figure 1 This is a flowchart of the steps of a disaster recovery data health monitoring method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a disaster recovery system data health check system provided in an embodiment of the present invention; Figure 3 This is a flowchart illustrating a disaster recovery data health monitoring method provided in an embodiment of the present invention; Figure 4This is a schematic diagram of the architecture of a data integrity verification layer module provided in an embodiment of the present invention; Figure 5 This is a flowchart illustrating another disaster recovery data health monitoring method provided in this embodiment of the invention; Figure 6 This is a structural block diagram of a disaster recovery data health monitoring device provided in an embodiment of the present invention; Figure 7 This is a hardware structure block diagram of an electronic device provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of a computer-readable medium provided in an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] Reference Figure 1 The diagram illustrates a flowchart of a disaster recovery data health monitoring method provided in an embodiment of the present invention, which may specifically include the following steps: Step 101: Obtain scheduling instructions for disaster recovery data, extract AI feature codes from the original parameters of the scheduling instructions, and obtain business feature information corresponding to the AI feature codes; Step 102: Determine the target disaster recovery data, and determine the health status of the target disaster recovery data based on the consistency result between the target disaster recovery data and the business characteristic information.
[0022] This invention provides an embodiment that can acquire scheduling instructions for disaster recovery data, extract AI feature codes from the original parameters of the scheduling instructions, and obtain business feature information corresponding to the AI feature codes to establish a basis for intelligent detection. This step aims to identify the core technical elements (AI feature codes) required for data health checks from high-level instructions and utilize a cloud-based AI big data system to obtain feature models of business data, preparing a reference standard for subsequent consistency calculations.
[0023] Scheduling instructions: Commands issued by a unified management and control platform (or a multi-strategy scheduling layer) to guide the disaster recovery data health check process.
[0024] Raw parameters: Meta-information carried in the scheduling instruction, used to locate the inspection target, specify the inspection method and constraint execution strategy, such as copy information, AI signature code, etc.
[0025] AI Feature Code: A code used to identify a specific business feature model, which is a key identifier for the system to retrieve corresponding business feature information from the big data center.
[0026] Business characteristic information: numerical values or models extracted from cloud AI big data systems to describe the characteristics of backup data business, serving as a reference standard for judging whether indicators such as data integrity and availability meet the standards.
[0027] In this embodiment of the invention, target disaster recovery data can also be identified, and based on the degree of consistency between the target disaster recovery data and the business feature information, the health status of the target disaster recovery data can be determined to quantitatively assess the quality of the disaster recovery data. This step generates a measurable data health status by efficiently comparing the actual disaster recovery data with the intelligent feature information and calculating the degree of consistency between the two.
[0028] Target disaster recovery data: The actual collection of data stored in a traditional disaster recovery system that requires health and integrity checks.
[0029] Matching result: The numerical result obtained by efficiently calculating the degree of matching between the target disaster recovery data and the business characteristic information (reference standard) is the direct basis for judging whether the data is healthy.
[0030] Health status: A measure of the quality, integrity, and reliability of disaster recovery data, such as "fully passed," "partially passed," or "pending manual inspection," used to guide customer recovery decisions or system optimization.
[0031] In this embodiment of the invention, by acquiring scheduling instructions for disaster recovery data, extracting AI feature codes from the original parameters of the scheduling instructions, and obtaining business feature information corresponding to the AI feature codes; determining target disaster recovery data, and determining the health status of the target disaster recovery data based on the consistency result between the target disaster recovery data and the business feature information, the invention enhances the multi-dimensional detection capabilities based on the AI system, solving the shortcoming of traditional backup solutions in providing customer data health monitoring, thereby greatly improving customer acceptance of the cloud system and avoiding increased business interruption time caused by blind attempts during business recovery.
[0032] Based on the above embodiments, modified embodiments of the above embodiments are proposed. It should be noted that, in order to keep the description brief, only the differences from the above embodiments are described in the modified embodiments.
[0033] In an optional embodiment of the present invention, prior to the step of extracting AI feature codes from the original parameters of the scheduling instruction, the method further includes: Filter key information from the original parameters of the scheduling instruction to directly drive the execution of subsequent tasks and resource scheduling; The key information is used to generate a verification task.
[0034] In this embodiment of the invention, key information that directly drives subsequent task execution and resource scheduling can be filtered from the raw parameters of the scheduling instruction to clarify the instruction intent. This step aims to preprocess the high-level raw parameters, filtering out core instruction elements that can specifically locate target data, specify verification methods, and constrain scheduling strategies. This ensures that subsequent processes (such as task generation and resource allocation) have clear and executable guidance.
[0035] The raw parameters of the scheduling instruction: a set of key information issued by the control layer, including information such as replicas, signatures, and whether to force execution.
[0036] Key information: Elements selected from the original parameters that have target-oriented characteristics (such as the range of data to be verified) and methodological characteristics (such as the required AI model identifier) directly drive the generation and scheduling of subsequent tasks.
[0037] In this embodiment of the invention, the aforementioned key information can be used to generate a verification task, thereby transforming instructions into action plans. This step concretizes the abstract "key information" into an execution entity—the verification task—that the system can recognize and allocate resources to.
[0038] Verification task: An execution work unit generated based on key information to guide subsequent data health checks, and is the basis for the system to allocate resources and control processes.
[0039] In this embodiment of the invention, by transforming high-level instructions into structured verification tasks, a clear and executable task framework and data foundation are provided for subsequent accurate extraction of AI feature codes, efficient multi-resource scheduling, and utilization of idle computing power in the cloud.
[0040] In an optional embodiment of the present invention, the step of obtaining business feature information corresponding to the AI feature code includes: The verification task is divided into multiple sub-tasks based on the historical availability of cloud resources. The subtasks are stored in the scheduling pool; Determine the current idle status of cloud resources. When the cloud resources meet the preset conditions based on the current idle status of cloud resources, extract the target subtask from the scheduling pool. When executing the target sub-task, AI feature codes are extracted from the original parameters corresponding to the target sub-task, and business feature information corresponding to the AI feature codes is obtained.
[0041] In this embodiment of the invention, the verification task can be divided into multiple sub-tasks based on the historical availability of cloud resources, thereby optimizing resource utilization and improving execution efficiency. The decomposition logic can be based on an understanding of historical resources, breaking down large tasks into parallel-processable sub-tasks to maximize the utilization of idle computing power on the cloud, thus shortening the overall verification time.
[0042] Verification task: A high-level, pending disaster recovery data health check work unit generated from key information.
[0043] Historical cloud resource availability: Availability data of cloud system computing resources (such as CPU and memory) over a past period, used as a basis for guiding the optimal granularity of task splitting.
[0044] Subtask: The smallest unit of work into which a verification task is broken down, which can be independently scheduled and executed.
[0045] In this embodiment of the invention, the subtasks can be stored in a scheduling pool for centralized task management. The split subtasks are placed in a central queue for unified resource allocation, priority sorting, and status tracking.
[0046] Scheduling pool: A system component or queue used to receive, store, and manage all subtasks to be executed.
[0047] This invention allows for the determination of the current idle status of cloud resources. When the idle status of cloud resources determines that the resources meet preset conditions, a target subtask is extracted from the scheduling pool to achieve dynamic resource allocation. Based on real-time resource availability, tasks are only extracted and executed when resources are plentiful, thus avoiding the occupation of core resources during peak business periods and ensuring no impact on the operation of customers' main businesses.
[0048] Current cloud resource availability: Real-time monitoring of the availability of cloud servers and computing power.
[0049] Preset conditions: The minimum resource threshold required to start a subtask as set by the system.
[0050] Target subtask: The specific subtask selected by the scheduling pool and about to be allocated to computing resources for execution.
[0051] In this embodiment of the invention, when executing the target sub-task, AI feature codes can be extracted from the original parameters corresponding to the target sub-task, and business feature information corresponding to the AI feature codes can be obtained to provide a technical basis for calculation. This step involves obtaining the AI feature code used to identify the correct verification model from its parameters at the moment the sub-task begins execution, and using this as evidence to obtain business feature information from the big data center, providing a reference standard for subsequent consistency calculations.
[0052] The original parameters corresponding to the target subtask: metadata fragments contained in the subtask, inherited from higher-level scheduling instructions, used to guide its own execution.
[0053] AI Feature Code: A unique identifier used for indexing and retrieving corresponding business feature models in cloud big data centers.
[0054] Business characteristic information: A set of model data or metrics obtained from the big data center to measure the integrity and availability of disaster recovery data.
[0055] In this embodiment of the invention, by utilizing the elastic scaling characteristics and idle resources of the cloud system, tasks are efficiently scheduled in parallel, and the feature data required for AI detection is accurately acquired during execution, ensuring the efficiency, low resource consumption, and high accuracy of the data health check process.
[0056] In an optional embodiment of the present invention, the step of determining the health status of the target disaster recovery data based on the consistency result between the target disaster recovery data and the business characteristic information includes: Based on the business feature information, the matching degree of the target disaster recovery data is calculated to generate a matching degree result that expresses the matching degree between the target disaster recovery data and the business feature information; The final verification result of the overall disaster recovery data is generated by using the consistency results of all target disaster recovery data. The health status of the disaster recovery data is determined using the final verification result.
[0057] In this embodiment of the invention, the target disaster recovery data can be calculated based on the business feature information to generate a consistency result expressing the consistency between the target disaster recovery data and the business feature information, thereby performing actual data health measurement. By using business feature information obtained from an AI big data system as a reference standard, efficient calculations and comparisons are performed on the target disaster recovery data, thereby quantifying the consistency result of the degree of matching between the two.
[0058] Business characteristic information: A set of feature models or indicators extracted from the cloud AI big data center, which serve as criteria for judging data integrity and business availability.
[0059] Target disaster recovery data: The data entities that are actually stored in the disaster recovery system and selected for inspection.
[0060] Matching result: A quantified percentage or score that directly expresses the degree of matching between the target disaster recovery data and its expected business characteristic model (business characteristic information).
[0061] In this embodiment of the invention, the consistency results of all target disaster recovery data can be used to generate the final verification result of the overall disaster recovery data, thereby summarizing and aggregating the scattered inspection results. Since the verification task is broken down into multiple sub-tasks, each sub-task has a consistency result. This step uniformly calculates and integrates the consistency results generated by all sub-tasks to form the final verification result for the entire backup set (overall disaster recovery data).
[0062] The consistency result of all target disaster recovery data: a set of numerical values generated during the verification phase of all sub-tasks to express the data matching degree.
[0063] The final verification result of the overall disaster recovery data: a comprehensive report on the health status of the entire disaster recovery backup data set, such as "pass rate 99.5%" or "number of questionable data blocks", which serves as the basis for qualitative judgment.
[0064] In this embodiment of the invention, the health status of the disaster recovery data can be determined using the final verification result, providing an actionable qualitative conclusion. This step transforms the complex final verification result into a concise and easy-to-understand health status label, providing customers and the system with a clear basis for recovery decisions and risk warnings.
[0065] Final verification result: The aggregated value or report of overall health generated in the previous step.
[0066] Health status of disaster recovery data: The data availability level ultimately determined by the system, usually expressed as "fully passed", "partially passed" or "awaiting manual inspection".
[0067] In this embodiment of the invention, by transforming efficient consistency calculation into quantifiable final verification results and operable health status, customers can quickly select the optimal replica for recovery at critical moments, thereby reducing the risk of business interruption and enhancing trust in the cloud disaster recovery system.
[0068] In an optional embodiment of the present invention, the step of generating the final verification result of the overall disaster recovery data by using the consistency results of all target disaster recovery data includes: When the matching result is greater than or equal to a preset threshold, the health status of the target disaster recovery data is determined to be good, and first health status information is generated to express that the health status of the target disaster recovery data is good. When the matching result is less than a preset threshold, the health status of the target disaster recovery data is determined to be questionable, and second health status information is generated to express that the health status of the target disaster recovery data is questionable. The final verification result of the overall disaster recovery data is generated using at least the first health status information and the second health status information.
[0069] In this embodiment of the invention, when the consistency result is greater than or equal to a preset threshold, the health status of the target disaster recovery data is determined to be good, and a first health status information is generated to express the good health status of the target disaster recovery data, thereby defining the success standard of the health data. This step compares the numerical consistency result calculated by a single subtask with the preset threshold, qualitatively classifies the data units that meet the standard as healthy, and generates a first health status information label for subsequent statistics.
[0070] Matching result: A quantitative score of the degree to which a single data unit matches the business feature information (reference standard).
[0071] Preset threshold: The passing score set in advance by the system or user, which is the basis for judging whether the data meets the expected health standards.
[0072] First health status information: This indicates that the data unit has passed verification and is in good condition, and is used for the final result summary count.
[0073] In this embodiment of the invention, when the consistency result is less than a preset threshold, the health status of the target disaster recovery data is determined to be questionable, and second health status information is generated to express the questionable health status of the target disaster recovery data, thereby identifying potential data risks. This step is used to mark data units whose consistency does not meet the standard, transforming the risk into a qualitative conclusion of questionable health status, and generating second health status information to ensure that risk points are not ignored.
[0074] Health status questionable: This indicates that the data unit may be damaged or incomplete, requiring further manual inspection or system repair.
[0075] Second health status information: This indicates that the data unit has failed verification or its status is questionable, and serves as a direct basis for risk warning.
[0076] In this embodiment of the invention, at least the first health status information and the second health status information can be used to generate the final verification result of the overall disaster recovery data, so as to aggregate and form a final conclusion. This step involves statistically summarizing the good status information and questionable status information of all individual data units, and finally forming a final verification result of the overall disaster recovery data that can represent the health level of the entire backup set.
[0077] Final verification result: A comprehensive report on the health status of the entire disaster recovery backup data set, such as pass rate and percentage of suspicious items, which serves as the main basis for customers or systems to make decisions.
[0078] By setting a matching threshold, this invention transforms complex numerical calculations into clear information on qualified and risky statuses, effectively quantifying and presenting the health of the entire disaster recovery dataset, thereby greatly supporting customers' recovery decisions and risk warnings.
[0079] Optionally, embodiments of the present invention can generate early warning information for the second health status information to achieve timely early warning and remediation of risks. This step aims to immediately convert risk information into early warning information and push it to customers or disaster recovery system administrators when the system determines that the disaster recovery data status is questionable, thereby ensuring timely detection of insecure disaster recovery data and buying time for customers or cloud disaster recovery providers to handle the situation.
[0080] Warning information: Based on the second health status information, the system generates notifications or prompts at the data display and alarm layers to alert customers that their data is at risk and requires timely manual inspection or strategy adjustment.
[0081] In this embodiment of the invention, by immediately generating early warnings for questionable data, it ensures the timely detection and handling of insecure disaster recovery data, preventing customers from using damaged copies at critical moments, thereby effectively guaranteeing the reliability of disaster recovery data.
[0082] Optionally, in this embodiment of the invention, the system can access a cloud-based big data center to retrieve business feature information corresponding to the AI feature code, thereby obtaining a reference standard for data health. Using the AI feature code extracted in the previous step, the system accesses the cloud-based AI big data system and retrieves business feature information matching the data type, providing a comparison benchmark for subsequent consistency calculations.
[0083] Cloud-based Big Data Center: A centralized data system that stores various business data feature models and is the provider of AI capabilities in this solution.
[0084] Business characteristic information: Model data or a set of indicators obtained from the big data center, describing indicators such as backup data integrity and business availability, serving as a reference standard for judging whether the data is healthy.
[0085] This invention, through its embodiment, ensures high accuracy and business adaptability of health checks by dynamically and accurately acquiring the required AI feature models during subtask execution, and is a core component in realizing AI-based big data monitoring capabilities.
[0086] For example, the monitoring process may include the following stages.
[0087] Phase 1: Task initiation and preparation; It all begins with scheduling instructions, which are raw commands issued by the system's higher-level layer (multi-policy scheduling layer) to guide the entire health check process. These instructions do not contain actual backup data, but rather metadata such as: "Please check backup copy A created at a certain point in time," "Please use the database-specific validation model," or "Please execute immediately."
[0088] From instructions to key information: After receiving the scheduling instructions, the system filters out key information (such as the replica ID to be verified and the verification model ID used).
[0089] Task generation and splitting: The system uses this key information to generate a general verification task, and splits the task into multiple sub-tasks based on the historical availability of cloud resources.
[0090] Resource allocation: These subtasks are sent to the scheduling pool. When cloud resources are sufficient, the scheduling pool extracts the target subtasks for execution.
[0091] Phase Two: Obtaining Verification Basis; When the target subtask begins execution, the system begins preparing the "reference points" required for verification.
[0092] AI signature: It is a unique identifier or index key extracted from the original parameters corresponding to the target subtask. Its function is to tell the system: "This backup data is of a certain type (such as a certain version of SQL database backup), and a matching verification model is expected." Business Feature Information: The system uses the AI feature codes extracted in the previous step to access the big data center in the cloud. The big data center stores a large amount of model or feature data that has been trained and extracted by AI, representing the structure, integrity, and availability indicators of health data. Business feature information is the reference standard that the system retrieves from the big data center that matches the data type to be tested.
[0093] Phase 3: Performing calculations and determining health status; With the verification target (target disaster recovery data) and the verification reference (business characteristic information), the system can perform core health assessments.
[0094] The target disaster recovery data is the reality (the actual data backed up), while the business characteristic information is the ideal (the structure and indicators that health data should have). The two are the relationship between the object to be tested and the reference standard.
[0095] The system calculates the fit between the target disaster recovery data (real-world data) and business characteristic information (ideal reference). This calculation is more in-depth than traditional hash verification, verifying multiple levels of indicators such as the structure of data blocks, the integrity of business logic, and the correctness of key metadata.
[0096] If the fit result is high (e.g., above a preset threshold), it indicates that the deviation between the real data and the ideal model is small, the data structure is complete, and the data is in good health.
[0097] If the fit is low, it indicates that the real data deviates significantly from the expectations of the ideal model, and there may be a risk of data corruption, loss or unavailability, raising questions about the health of the data.
[0098] Finally, the system summarizes all the consistency results, generates the final verification result of the overall disaster recovery data, and determines the overall health status based on the result, thereby generating an early warning prompt.
[0099] To enable those skilled in the art to better understand the embodiments of the present invention, a complete example is used below to illustrate the embodiments of the present invention.
[0100] The purpose of this invention is to provide a health check method for disaster recovery data based on a cloud-native architecture, relying on a cloud-based data disaster recovery system, abundant cloud systems and storage space, and massive and flexible cloud computing capabilities. This solution does not require the customer system to provide computing power and space support for health checks; the entire process is completed in the cloud through the rational utilization of idle servers and computing power. Furthermore, it can provide multiple scheduling mechanisms, such as scheduled health checks and triggered health checks (which require waiting a certain period to generate health monitoring results), according to the customer's customized requirements.
[0101] This invention provides cloud storage providers with high-quality and cost-effective services. Addressing customer concerns about the security of cloud systems and disaster recovery systems, this solution not only offers data integrity health checks but also, in conjunction with other cloud security services, provides data security health checks. This solution expands the application scenarios of cloud systems and dispels customer concerns about their security.
[0102] The functions to be implemented in this embodiment of the invention include: a data pre-verification code function layer, a disaster recovery data health check scheduling strategy function layer, a disaster recovery data health check monitoring layer, and a disaster recovery data health check multi-level report display layer.
[0103] The overall solution requires no customer involvement in business configuration changes, involves only a minor increase in cloud storage space, and can be promoted in conjunction with other security features. It does not require complex modifications to the customer's business systems or disaster recovery systems; only the disaster recovery system needs to open a data read interface. Furthermore, the entire process is completed within the cloud system, ensuring the privacy of customer data. Timely data security checks can greatly enhance customer acceptance of the cloud system.
[0104] The embodiments of the present invention are divided into four main functional layers: a data pre-verification layer that realizes the pre-processing function of disaster recovery data inspection, an integrity monitoring layer that realizes the health check and monitoring of disaster recovery data, a multi-strategy scheduling layer that realizes the strategy function of disaster recovery data health check, and a display and alarm function layer that realizes the display function of disaster recovery data health check.
[0105] Overall functional solution reference Figure 2 , Figure 2 Figure 2 This is a schematic diagram of the structure of a disaster recovery system data health check system provided in an embodiment of the present invention; As shown in the diagram, this solution is completed solely within the cloud disaster recovery system. It requires no cooperation from the customer's business operations and does not need to be aware of the customer's business or the cloud disaster recovery system's backup logic. It simply adds health check logic to the traditional disaster recovery functions, and implements customizable health check logic. Through integration with the disaster recovery data system, it provides customers with various customized health check solutions (scheduled, immediate, etc.), and can provide health check reports and more advanced alarm functions according to customer needs.
[0106] Customers / cloud systems using this solution can achieve the following objectives: 1. Timely detection of insecure disaster recovery data, timely early warning and remediation, alerting customers to the risks to their disaster recovery system data, allowing customers or cloud disaster recovery providers to handle the situation promptly; 2. Providing customers with data health alerts during business recovery, preventing customers from wasting valuable business time by attempting recovery multiple times; 3. Providing professional data health check reports for the disaster recovery system, offering optimization directions, and providing suggestions for adjusting disaster recovery strategies based on the results of various health checks.
[0107] The above analysis shows that this solution addresses the shortcomings of traditional backup solutions, which only provide backup and cannot offer customer data health monitoring. Building upon traditional disaster recovery solutions, it provides metrics for customer data health checks. Without affecting basic disaster recovery functions and scheduling, it can offer optimization suggestions for customer disaster recovery strategies based on the data health of different strategies, thereby improving the data availability metrics of the customer's disaster recovery system.
[0108] Overall control process transformation, such as Figure 3 As shown, Figure 3 This is a flowchart illustrating a disaster recovery data health monitoring method provided in this embodiment of the invention. This solution can perform disaster recovery data health checks in existing backup systems via plugins and add-ons. It extracts data based on mature business features provided by a cloud-based AI big data system and implements the disaster recovery data health check function through data integrity verification. The entire solution does not require any business modifications to the customer's system and does not affect the operation of the cloud-based disaster recovery system. After monitoring is completed, the results can be immediately pushed to the display and alerting layer (which can be displayed to the customer or used as a basis for adjusting disaster recovery system strategies).
[0109] refer to Figure 4 , Figure 4 This is a schematic diagram of the architecture of a data integrity verification layer module provided in this embodiment of the invention. The data integrity verification layer consists of two key functions: data detection and AI feature extraction. Relying on the backup data business-specific values extracted by AI big data, it performs multi-level verification on the backup data to confirm whether the integrity, business availability, and other indicators of the backup business data meet the standards. Since the entire system is built on the cloud, relying on the elastic scaling and elastic resource expansion characteristics of the cloud system, it can make full use of idle resources for feature calculation, result analysis, and other processes based on the current business capacity and the actual idleness of cloud resources. When cloud resources are scarce (such as during peak business periods), feature calculation resources can be released immediately to return resources to the business system without affecting the operation of the cloud business system and disaster recovery system, and other major customer businesses.
[0110] The flowchart of the data integrity verification layer functional module is as follows: Figure 5 As shown, Figure 5 This is a flowchart illustrating another disaster recovery data health monitoring method provided in this embodiment of the invention; Step 1: Extract key information based on the parameters (copy, signature, whether to force execution, etc.) issued by the control layer, generate tasks and split them into multiple resource scheduling pools, and then the scheduling pools issue the tasks.
[0111] Step 2: After the task is released, if there are available resources, the task will be executed immediately; otherwise, wait.
[0112] Step 3: During task execution, obtain the corresponding business characteristic information from the big data center based on the task's AI feature code.
[0113] Step 4: Efficiently calculate whether the AI feature information matches the data. If the match exceeds the threshold, the data health check is considered passed; otherwise, it is marked as questionable.
[0114] Step 5: Calculate and publish the results: All passed, percentage of questionable, and completely non-compliant.
[0115] Step Six: After the overall task is completed, calculate whether the overall task has passed the verification based on the release results of each individual task (the result is set as: fully passed, partially passed, pending manual inspection).
[0116] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0117] Reference Figure 6 The diagram illustrates a structural block diagram of a disaster recovery data health monitoring device provided in an embodiment of the present invention, which may specifically include the following modules: The business feature information acquisition module 601 is used to acquire scheduling instructions for disaster recovery data, extract AI feature codes from the original parameters of the scheduling instructions, and acquire business feature information corresponding to the AI feature codes. The health status generation module 602 is used to determine the target disaster recovery data and, based on the consistency result between the target disaster recovery data and the business characteristic information, determine the health status of the target disaster recovery data.
[0118] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0119] In addition, embodiments of the present invention also provide an electronic device, such as... Figure 7 As shown, it includes a processor 701, a communication interface 702, a memory 703, and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704. Memory 703 is used to store computer programs; When the processor 701 executes the program stored in the memory 703, it implements any of the disaster recovery data health monitoring methods described in the above embodiments: The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0120] The communication interface is used for communication between the aforementioned terminal and other devices.
[0121] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0122] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0123] like Figure 8As shown, in another embodiment of the present invention, a computer-readable storage medium 801 is also provided, which stores instructions that, when executed on a computer, cause the computer to perform the disaster recovery data health monitoring method described in the above embodiment.
[0124] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described disaster recovery data health monitoring method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0125] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0126] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0127] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0128] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for monitoring the health of disaster recovery data, characterized in that, include: Obtain scheduling instructions for disaster recovery data, extract AI feature codes from the original parameters of the scheduling instructions, and obtain business feature information corresponding to the AI feature codes; The target disaster recovery data is determined, and the health status of the target disaster recovery data is determined based on the consistency between the target disaster recovery data and the business characteristic information.
2. The method according to claim 1, characterized in that, Prior to the step of extracting AI feature codes from the original parameters of the scheduling instruction, the method further includes: Filter key information from the original parameters of the scheduling instruction to directly drive the execution of subsequent tasks and resource scheduling; The key information is used to generate a verification task.
3. The method according to claim 2, characterized in that, The step of obtaining the business feature information corresponding to the AI feature code includes: The verification task is divided into multiple sub-tasks based on the historical availability of cloud resources. The subtasks are stored in the scheduling pool; Determine the current idle status of cloud resources. When the cloud resources meet the preset conditions based on the current idle status of cloud resources, extract the target subtask from the scheduling pool. When executing the target sub-task, AI feature codes are extracted from the original parameters corresponding to the target sub-task, and business feature information corresponding to the AI feature codes is obtained.
4. The method according to claim 1, characterized in that, The step of determining the health status of the target disaster recovery data based on the consistency result between the target disaster recovery data and the business characteristic information includes: Based on the business feature information, the matching degree of the target disaster recovery data is calculated to generate a matching degree result that expresses the matching degree between the target disaster recovery data and the business feature information; The final verification result of the overall disaster recovery data is generated by using the consistency results of all target disaster recovery data. The health status of the disaster recovery data is determined using the final verification result.
5. The method according to claim 4, characterized in that, The step of generating the final verification result of the overall disaster recovery data by using the consistency results of all target disaster recovery data includes: When the matching result is greater than or equal to a preset threshold, the health status of the target disaster recovery data is determined to be good, and first health status information is generated to express that the health status of the target disaster recovery data is good. When the matching result is less than a preset threshold, the health status of the target disaster recovery data is determined to be questionable, and second health status information is generated to express that the health status of the target disaster recovery data is questionable. The final verification result of the overall disaster recovery data is generated using at least the first health status information and the second health status information.
6. The method according to claim 5, characterized in that, Also includes: Generate an early warning message for the second health status information.
7. The method according to claim 5, characterized in that, The step of obtaining the business feature information corresponding to the AI feature code includes: Access the cloud-based big data center and retrieve business feature information corresponding to the AI feature code from the cloud-based big data center.
8. A disaster recovery data health monitoring device, characterized in that, include: The business feature information acquisition module is used to acquire scheduling instructions for disaster recovery data, extract AI feature codes from the original parameters of the scheduling instructions, and acquire business feature information corresponding to the AI feature codes. The health status generation module is used to determine the target disaster recovery data and, based on the consistency result between the target disaster recovery data and the business characteristic information, determine the health status of the target disaster recovery data.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the method as described in claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the method as described in claims 1-7.