A data asset disaster recovery processing method and system
Patent Information
- Application Number
- CN202611041921.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-14
AI Technical Summary
这些规则往往基于简单的数据分类(如文件类型、访问频率)或孤立的业务影响分析(BIA)报告,缺乏对多维度、动态化风险因素的综合考量
[0063]本申请实施例公开了一种数据资产容灾处理方法及系统。该方法通过采集结构化与非结构化数据的元数据并生成统一的数据资产目录,为后续处理奠定了基础。在此基础上,方案创造性地引入了业务关键度系数B、时间紧迫度因子T、数据丢失容忍度因子P、规范约束系数C及环境风险概率得分V这五个核心量化指标。这些指标分别对应了业务影响、时间容忍度、数据丢失容忍度、合规约束和环境脆弱性,从而构建了一个多维度、全覆盖的评估体系。通过预设加权规则计算出的恢复优先级指数(RPI),能够科学、客观地反映每个数据资产在当前环境下的综合恢复价值与紧迫性。
Smart Images

Figure CN122547613B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data disaster recovery and business continuity management technology, and more specifically, to a data asset disaster recovery processing method and system. Background Technology
[0002] In modern enterprise IT environments, data assets have become a core production factor, and their importance is self-evident. To cope with potential disasters such as natural disasters, hardware failures, and cyberattacks, enterprises have generally deployed data disaster recovery systems in order to quickly restore normal operations after a disaster. Traditional disaster recovery solutions typically rely on predefined static rules to identify critical data and determine its recovery priority. These rules are often based on simple data classifications (such as file type, access frequency) or isolated Business Impact Analysis (BIA) reports, lacking a comprehensive consideration of multi-dimensional and dynamic risk factors.
[0003] However, with increasing business complexity and increasingly stringent external regulatory environments, this static, single-perspective disaster recovery approach has revealed serious shortcomings. On the one hand, it struggles to adapt to rapid changes in business needs or external threats, leading to a disconnect between recovery strategies and actual business objectives. For example, a security policy mandating the recovery of a signature-based database within a very short timeframe might be relegated to the back of non-critical logs due to low access frequency. On the other hand, existing assessment models typically separate business metrics (such as RTO / RPO), recovery constraints, and technical risks, failing to establish a unified, quantifiable decision-making basis and thus hindering effective guidance for automated recovery scheduling.
[0004] In view of this, a data asset disaster recovery processing method and system are proposed. Summary of the Invention
[0005] This application proposes a data asset disaster recovery processing method and system to solve one of the aforementioned technical problems.
[0006] The technical solution adopted in this application is as follows:
[0007] In a first aspect, embodiments of this application provide a data asset disaster recovery processing method, including:
[0008] Collect metadata of structured and unstructured data in the enterprise IT environment, map it to unified common fields, and generate a data asset catalog with unique identifiers;
[0009] The data assets in the data asset catalog are classified according to a preset multi-branch decision logic strategy, and classification result labels are generated.
[0010] Based on the business impact analysis level, recovery time target, recovery point target, classification compliance attributes and environmental risk data, the business criticality coefficient B, time urgency factor T, data loss tolerance factor P, regulatory constraint coefficient C and environmental risk probability score V are generated.
[0011] The Recovery Priority Index (RPI) is calculated according to a preset weighting rule, and a disaster recovery execution queue is generated accordingly.
[0012] In response to external state change events, update the preset weighting rules and rearrange the disaster recovery execution queue;
[0013] According to the rearranged disaster recovery execution queue, control commands are sent to the network controller and virtualization resource pool to perform resource preemption, locking, and disaster recovery data recovery.
[0014] Preferably, it is mapped to a unified public field and a data asset catalog with a unique identifier is generated, including:
[0015] The metadata is collected from at least one data source, including database servers, file servers, cloud storage services, and local disks.
[0016] The table name, field name, number of table records, and update time in the structured data, and the file name, storage path, number of file bytes, and modification time in the unstructured data are respectively mapped to the asset name, asset location, data volume, data format, and timestamp fields in the data asset catalog;
[0017] Generate data feature labels for field relationships in structured data, and generate extended feature values for storage paths, file extensions, and storage media attributes in unstructured data;
[0018] Write the data feature tags and extended feature values into the extended metadata field of the data asset catalog, and configure a unique asset identifier for each data asset.
[0019] Preferably, the data assets in the data asset catalog are classified according to a preset multi-branch decision logic strategy to generate classification result labels, including:
[0020] Extract the content keywords of the data assets. If the content keywords match the preset set of highly sensitive keywords, directly output the classification result tags containing high risk level and core asset markers.
[0021] If the preset set of highly sensitive keywords is not matched, the file type is determined based on the file extension of the data asset, and the volume level is determined based on the data volume parameter of the data asset.
[0022] A combined condition determination is performed based on the content keywords, file type, and size level;
[0023] Specifically, if the file type is code or configuration file, the output will include a classification result label containing the core asset marker; if the file type is a structured data file and the size level reaches the preset large size level, or hits the preset set of sensitive keywords, the output will include a classification result label containing financial data, operational data, or sensitive data markers; if no preset branch is matched, the output will be a classification result label pending manual review.
[0024] Preferably, the business criticality coefficient B, the regulatory constraint coefficient C, and the environmental risk probability score V are generated, including:
[0025] The business impact analysis level associated with the data asset is read from the configuration management database or business continuity management system, and converted into the business criticality coefficient B through a preset business level mapping table; if the business impact analysis level is missing, a preset median business criticality coefficient is assigned and a missing alarm is generated.
[0026] Read the data sensitivity identifier or industry compliance attribute identifier from the classification result label, and convert it into the normative constraint coefficient C through a preset constraint mapping table;
[0027] The environmental risk probability score V is calculated based on at least two of the following: physical risk probability, technological risk probability, and human risk probability.
[0028] The environmental risk probability score V is used to characterize the current threat level of the data asset under the dimensions of physical risk, technical risk, or human risk, and the normative constraint coefficient C is used to characterize the degree of mandatory compliance recovery constraints on the data asset.
[0029] Preferably, generating the time urgency factor T includes:
[0030] Obtain the recovery time target of the data asset and convert the recovery time target into a second-level value RTOsec;
[0031] If the RTOsec is missing or less than or equal to zero, the RTOsec is replaced with a preset median second value; if the RTOsec is greater than a preset maximum truncation threshold, the RTOsec is truncated to the preset maximum truncation threshold; if the RTOsec is less than a preset minimum emergency threshold, the RTOsec is replaced with the preset minimum emergency threshold.
[0032] Based on the boundary-processed RTOsec, a monotonically decreasing normalization calculation is performed to obtain the time urgency factor T, such that the smaller the RTOsec, the larger the time urgency factor T.
[0033] Preferably, generating the data loss tolerance factor P includes:
[0034] Obtain the recovery point target of the data asset and convert the recovery point target into a second-level value RPOsec;
[0035] If RPOsec is missing or less than zero, RPOsec is replaced with a preset median second value; if RPOsec is equal to zero, the data loss tolerance factor P is assigned a preset maximum score; if RPOsec is greater than a preset maximum tolerance threshold, RPOsec is truncated to the preset maximum tolerance threshold.
[0036] Based on the boundary-processed RPOsec, a monotonically decreasing normalization calculation is performed to obtain the data loss tolerance factor P, such that the smaller RPOsec is, the larger the data loss tolerance factor P is.
[0037] Preferably, the recovery priority index (RPI) is calculated according to a preset weighting rule, including:
[0038] Obtain the first basic weight coefficient Wb corresponding to the business criticality coefficient B, the second basic weight coefficient Wt corresponding to the time urgency factor T, the third basic weight coefficient Wp corresponding to the data loss tolerance factor P, and the fourth basic weight coefficient Wv corresponding to the environmental risk probability score V;
[0039] Normalize and verify the first basic weight coefficient Wb, the second basic weight coefficient Wt, the third basic weight coefficient Wp and the fourth basic weight coefficient Wv to make them satisfy Wb+Wt+Wp+Wv=1;
[0040] The Recovery Priority Index (RPI) is calculated using the following formula:
[0041] RPI=B×Wb+T×Wt+P×Wp+V×C×Wv;
[0042] The product term V×C×Wv of the environmental risk probability score V and the normative constraint coefficient C is used to positively weight the recovery priority index RPI to improve the ranking priority of the corresponding data assets in the disaster recovery execution queue.
[0043] The Recovery Priority Index (RPI) is compared with a preset index threshold. If the RPI is greater than the preset index threshold, the corresponding data asset is marked as a first-level priority recovery asset.
[0044] Preferably, updating the preset weighted rules and rearranging the disaster recovery execution queue in response to external state change events includes:
[0045] External state change events are obtained through the event capture interface. These external state change events include at least one of the following: security events, vulnerability discovery events, external rule change events, and protection device configuration change events.
[0046] The external state change event is converted into a formatted structured message by the event converter. The formatted structured message includes at least the event type, event source, event time, event payload, and severity of impact.
[0047] The formatted structure message is routed to the corresponding weight adjustment process according to the event type, and the weight adjustment range is determined according to the severity of the impact.
[0048] Update at least one basic weight coefficient in the preset weighting rule according to the weight adjustment range, and perform weight normalization processing after the update;
[0049] A weight update signal is sent to the recovery priority index calculation module to trigger the recalculation of the recovery priority index RPI, and the disaster recovery execution queue is rearranged according to the recalculated recovery priority index RPI.
[0050] Preferably, according to the rearranged disaster recovery execution queue, control commands are sent to the network controller and virtualization resource pool to perform resource preemption, locking, and disaster recovery data recovery, including:
[0051] Data assets whose Recovery Priority Index (RPI) is greater than the preset index threshold are identified as Level 1 priority recovery assets.
[0052] For the first-priority recovery assets, a bandwidth expansion command and a bandwidth preemption command are sent to the software-defined network controller to pre-lock a preset proportion of network bandwidth resources;
[0053] Send computing resource scheduling instructions and storage resource scheduling instructions to the virtualization resource pool to allocate a specified number of CPU cores, memory resources and solid-state storage resources to the first-priority recovery assets, and perform parallel recovery;
[0054] For data assets not identified as first-priority recovery assets, allocate the remaining shared bandwidth and default resource quotas, and perform serial recovery;
[0055] After the disaster recovery data is restored, a consistency check is performed on the restored data, and the pre-locked network bandwidth resources, computing resources and storage resources are released after the check passes.
[0056] Secondly, embodiments of this application provide a data asset disaster recovery processing system to implement the method described in any of the above-mentioned methods, including:
[0057] The heterogeneous data alignment module is used to collect metadata of structured and unstructured data in the enterprise IT environment, map it to a unified common field, and generate a data asset catalog with a unique identifier.
[0058] The multi-branch rule classification module is used to classify data assets in the data asset catalog according to a preset multi-branch decision logic strategy and generate classification result labels;
[0059] The indicator quantification module is used to generate a business criticality coefficient B, a time urgency factor T, a data loss tolerance factor P, a regulatory constraint coefficient C, and an environmental risk probability score V based on the business impact analysis level, recovery time target, recovery point target, classification compliance attributes, and environmental risk data.
[0060] The priority calculation module is used to calculate the recovery priority index (RPI) according to a preset weighting rule, and generate a disaster recovery execution queue accordingly.
[0061] The dynamic feedback module is used to update the preset weighting rules and rearrange the disaster recovery execution queue in response to external state change events;
[0062] The disaster recovery automated scheduling module is used to send control commands to the network controller and virtualization resource pool according to the rearranged disaster recovery execution queue, and to perform resource preemption, locking and disaster recovery data recovery.
[0063] This application discloses a data asset disaster recovery processing method and system. The method collects metadata from structured and unstructured data and generates a unified data asset catalog, laying the foundation for subsequent processing. Based on this, the solution creatively introduces five core quantitative indicators: business criticality coefficient B, time urgency factor T, data loss tolerance factor P, regulatory constraint coefficient C, and environmental risk probability score V. These indicators correspond to business impact, time tolerance, data loss tolerance, compliance constraints, and environmental vulnerability, respectively, thus constructing a multi-dimensional and comprehensive evaluation system. The Recovery Priority Index (RPI), calculated through preset weighted rules, can scientifically and objectively reflect the comprehensive recovery value and urgency of each data asset in the current environment.
[0064] This solution transforms recovery constraint indicators such as business impact analysis level, recovery time target, and recovery point target, as well as recovery priority requirements in external constraint rules, into quantifiable factors (B, T, P, C) that can be directly used in calculations through specific mathematical transformations and mapping logic. In particular, the design of the normative constraint coefficient C ensures that data assets with high-level constraint labels, even if their environmental risk probability scores are low, can still obtain a higher ranking in the recovery priority index calculation due to the increased constraint strength, thus being prioritized for recovery.
[0065] The solution not only calculates the RPI but also incorporates a dynamic feedback mechanism. When the system detects external state changes such as security events or changes in external rules, it automatically updates the preset weighted rules and triggers the recalculation of the RPI and the reordering of the disaster recovery queue. Ultimately, the system can directly send control commands to the network controller and virtualization resource pool to execute resource preemption, locking, and data recovery operations. This complete closed loop from "intelligent assessment" to "automatic execution" greatly improves the speed and efficiency of disaster recovery response, ensuring that recovery actions remain consistent with the latest business and security posture in complex and ever-changing environments. Attached Figure Description
[0066] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0067] Figure 1 A flowchart illustrating a data asset disaster recovery processing method provided in an embodiment of the present invention;
[0068] Figure 2 This is a schematic diagram of the structure of a data asset disaster recovery system provided in an embodiment of the present invention. Detailed Implementation
[0069] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0070] like Figure 1As shown in some embodiments of this application, this embodiment provides a data asset disaster recovery processing method, specifically, the method includes the following steps:
[0071] Step S101: Collect metadata of structured and unstructured data in the enterprise IT environment, map it to a unified common field, and generate a data asset catalog with a unique identifier.
[0072] The core of this step lies in addressing the issues of heterogeneous data asset formats and fragmented management within enterprise IT environments. Specifically, this step comprises three progressive operational levels:
[0073] Comprehensive Data Collection: The system needs to proactively or passively identify and extract metadata from all corners of the enterprise IT environment, including but not limited to database servers (such as MySQL, Oracle), file servers (such as NAS, SMB shared storage), public or private cloud storage services (such as AWS S3, Alibaba Cloud OSS), and local disks, to identify and extract metadata from all potential data assets. This metadata includes not only basic information such as filenames, paths, and sizes, but also, for structured data, table names, field definitions, number of records, and relationships; for unstructured data, it may include file extensions, content summaries, and hash values.
[0074] Heterogeneous Alignment: Due to the significant differences in dimensionality and structure between the metadata of structured data (such as database tables) and unstructured data (such as Office documents), direct comparison or unified processing is extremely difficult. Therefore, this solution introduces the key operation of "mapping to unified common fields." This is not a simple format conversion, but a process of dimensionality reduction and abstraction. The system predefines a set of common field models covering the core attributes of data assets (e.g., asset name, asset location, data volume, data format, timestamp, etc.). Then, through a set of preset mapping rules, the original metadata from different sources and with different structures is "translated" and populated into this set of unified common fields. For example, although the "number of records" in a database table and the "number of bytes" in a file have different units, they are both mapped to a unified "data volume" field, allowing the subsequent risk assessment module to perform differentiated processing based on this field.
[0075] Catalog Construction: After completing the collection and alignment of metadata, the system generates a globally unique asset ID for each identified data asset. It then integrates all aligned public field information, simplified features of the original metadata (usually stored in extended fields in JSON format), and this unique ID to form a structured "data asset catalog." This catalog serves as the sole data source and operational object for all subsequent processing steps (classification, evaluation, scheduling), ensuring data consistency and traceability throughout the entire process.
[0076] For example, suppose the following two types of data assets exist simultaneously in an enterprise's IT environment:
[0077] Structured data asset A: A MySQL database table named customer_info, located in an internal database instance, contains a preset number of business object records, and the last update time is the first timestamp.
[0078] Unstructured data asset B: An Excel file named Q2_Financial_Report.xlsx, stored in the path \\fileserver\finance\, with a file size of 2.5MB, last modified on May 21, 2024.
[0079] After executing step S101, the system will generate directory entries in the following uniform format:
[0080] The catalog entries for asset A are: unique asset identifier, DB_TBL_001; asset name, customer_info; asset location, internal_db / customer_info; data volume, preset number of records; data format, MySQL; timestamp, first timestamp; extended metadata, {"number of fields":10,"including the first sensitive field"}.
[0081] The catalog entry for Asset B is as follows: Unique Asset Identifier, FILE_002; Asset Name, Q2_Financial_Report.xlsx; Asset Location, “\\fileserver\finance\Q2_Financial_Report.xlsx”; Data Size, 2,621,440 (bytes); Data Format, xlsx; Timestamp, 2024-05-21T09:15:00; Extended Metadata, {"File Hash":"a1b2c3d4...","Storage Media":"Enterprise NAS"}.
[0082] In this way, previously unrelated database tables and Excel files have a comparable and computable unified description in the data asset catalog, paving the way for subsequent automated processing.
[0083] It should be noted that, in specific implementation scenarios, an extended solution based on the above approach can be adopted. That is, the collection of metadata is not limited to automatic scanning tools, but can also include receiving push or subscription messages from the enterprise's existing management systems (such as CMDB configuration management database, DLP data leakage prevention system, SIEM security information and event management platform) to dynamically detect newly created or changed data assets.
[0084] A dynamic mapping rule scheme is adopted, meaning that the mapping rules for unified public fields are not static. The system allows administrators to customize or adjust the mapping logic according to business needs. For example, for constraint rules specific to a particular industry, a new "Industry Data Category" public field can be added, and corresponding mapping rules can be configured to extract this information from the original metadata or external tagging services. All of the above optional solutions fall within the scope of protection of this application.
[0085] Step S102: Classify the data assets in the data asset catalog according to the preset multi-branch decision logic strategy and generate classification result labels.
[0086] This step aims to conduct in-depth semantic analysis and risk assessment on the unified data asset catalog constructed in step S101, in order to achieve refined and intelligent asset classification. Its core lies in the "multi-branch decision-making logic strategy," a hierarchical and conditional judgment mechanism that simulates the expert decision-making process, rather than a simple, single linear rule.
[0087] The specific operating procedure is as follows:
[0088] Content Keyword Extraction and High-Sensitivity Matching: The system first attempts to extract keywords from the content of the data asset (or its metadata summary). For structured data, this may involve scanning sample data within the table; for unstructured data, text parsing techniques may be used. The extracted keywords are immediately compared with a pre-defined set of "high-sensitivity keywords" (such as "identity identifier characteristics," "financial account characteristics," "high-security identifier," "source code," etc.). If a match is found, the highest priority decision branch is triggered, directly assigning the asset a strong label such as "high-risk level" and "core asset," without requiring further judgment, ensuring rapid identification of known high-risk assets.
[0089] Secondary Feature Analysis: For assets not matched by high-sensitivity keywords, the system proceeds to secondary analysis. This stage primarily relies on the unified field information mapped in step S101. For example, the file type is determined by the data format field (e.g., .java for code files, .sql for configuration files, .dbf for structured data files), and the data size field is used to determine its size level (e.g., less than 1GB for small size, 1-10GB for medium size, and more than 10GB for large size).
[0090] Combination Condition Judgment: After obtaining the file type and size level, the system executes more complex combination logic judgments. For example, if the file type is code or configuration file, regardless of its size, it is considered a core asset supporting business operations. If the file type is a structured data file and its size reaches the "large size" level, or if its content keywords match the "medium-sensitive keyword set" (such as "customer," "contract," "financial," etc.), it will be labeled with more business semantics such as "financial data," "operational data," or "sensitive data."
[0091] As a fallback: For assets that cannot be clearly categorized in any of the above branches, the system will generate a "Pending Manual Review" label. This ensures both the efficiency of automated processing and preserves a channel for manual intervention in complex or edge cases, thus guaranteeing the rigor of the classification results.
[0092] Ultimately, all these judgments will be encapsulated into structured "classification result labels," serving as a key attribute of the data asset catalog for use by the subsequent quantitative evaluation module.
[0093] For example, continuing with the previous example of assets A and B:
[0094] Asset A (customer_info table): During content sampling, the system found that the fields contained keywords such as "identity identifier characteristics," which are present in the preset "highly sensitive keyword set." Therefore, the system skipped subsequent judgments and generated a classification result label for it: {"Risk Level":"High","Asset Type":"Core Asset","Data Category":"Identity Feature Data"}.
[0095] Asset B (Q2_Financial_Report.xlsx file): This file's content did not match the highly sensitive keyword set. The system then analyzed its data format fields, identifying it as an Excel file. Further analysis of its content keywords revealed the presence of terms such as "revenue," "net profit," and "assets and liabilities," which belong to the "medium sensitive keyword set." Therefore, based on the combination judgment logic, the system generated the following classification result label: {"Risk Level":"Medium","Asset Type":"Sensitive Data","Data Category":"Financial Data"}.
[0096] New Asset C (system_backup_20240521.tar.gz): This is a large backup compressed file, and no keywords were matched. Its data format is .tar.gz (a common compression format), and its data size is 50GB (large). Since its file type is neither code / configuration nor a typical structured data file, and no sensitive words were matched, the system cannot determine its business value through preset rules. Therefore, the following tags are generated: {"Risk Level":"Pending","Asset Type":"Pending Manual Review"}.
[0097] It should be noted that, in specific implementation scenarios, a dynamic keyword set update scheme can be adopted based on the above solution, meaning that the highly sensitive and moderately sensitive keyword sets are not static. The system can be designed to periodically update keywords from external threat intelligence sources, compliant external rule databases, or the enterprise's internal security policy library, ensuring that the classification logic keeps pace with the times and addresses new data breach risks. A configurable decision logic scheme can be adopted, meaning that the specific conditions and thresholds of multi-branch decision logic strategies (such as large GB volumes, which file types are considered core assets) can be customized by the administrator through the management interface. This allows this solution to flexibly adapt to the specific business scenarios and security policies of enterprises of different industries and sizes. All of the above optional solutions are within the scope of protection of this application.
[0098] Step S103: Based on the business impact analysis level, recovery time target, recovery point target, classification compliance attributes and environmental risk data, generate the business criticality coefficient B, time urgency factor T, data loss tolerance factor P, regulatory constraint coefficient C and environmental risk probability score V.
[0099] The purpose of this step is to transform qualitative and quantitative information from different dimensions and of different natures into standardized numerical factors that can be calculated and compared. These factors together form the basis for the subsequent calculation of the Recovery Priority Index (RPI).
[0100] Business Criticality Coefficient (B): This coefficient directly reflects the strategic importance of data assets to business continuity. It is generated based on the ratings (e.g., "Critical," "Important," "Normal") in the company's formal Business Impact Analysis (BIA) report. The system uses a pre-defined mapping table (e.g., "Critical" → 1.0, "Important" → 0.7, "Normal" → 0.4) to convert it into a continuous value between 0 and 1. If an asset lacks a BIA rating, a median value (e.g., 0.5) is used as the default value, and an alarm is triggered to remind the administrator to complete the information, ensuring that the assessment is not interrupted due to missing data.
[0101] The Time Urgency Factor (T) and Data Loss Tolerance Factor (P) originate from two core metrics in disaster recovery: Recovery Time Objective (RTO) and Recovery Point Objective (RPO). The system first converts RTO / RPO from units like hours and days to seconds (RTOsec / RPOsec). Then, these raw values undergo boundary processing (such as missing value imputation, truncation of excessively large values, and protection against extremely small values), and are mapped to the 0-1 range using a monotonically decreasing normalization function (such as an inverse proportional function or a negative exponential function). The result of this design is that the more stringent the RTO / RPO requirements (i.e., the smaller the values), the larger the corresponding T / P factor, thus gaining a higher weight in the RPI calculation.
[0102] The Compliance Multiplier (C) is a unique design element used to reflect the binding force of external rules or industry standards on data recovery. It is not a simple weighting term, but rather a multiplier. The system extracts "data sensitivity identifiers" (such as "first sensitive field identifiers") or "industry compliance attribute identifiers" (such as "second sensitive field identifiers") from the classification result labels generated in step S102, and converts them into a value greater than or equal to 1 through a constraint mapping table (e.g., C=1.0 for ordinary data, C=1.5 for data constrained by the second sensitive field identifier). This coefficient will be multiplied by the environmental risk probability score V in subsequent calculations, forming a V×C term, thus producing a significant amplification effect in high-risk and highly constrained scenarios.
[0103] Environmental Risk Probability Score (V): This score comprehensively assesses the external threats currently faced by data assets. It is calculated by aggregating at least two independent risk sources, such as: physical risks, the probability of natural disasters (earthquakes, floods) occurring in the area where the data center is located; technical risks, the severity of known unpatched vulnerabilities in the system hosting the asset and their potential for exploitation; and human risks, the success rate of unusual access behavior by internal personnel or social engineering attacks. These risk probabilities from different dimensions are combined through weighted averaging or taking the maximum value, ultimately forming an Environmental Risk Probability Score (V) between 0 and 1, used to characterize the vulnerability of the asset at the current moment.
[0104] For example, suppose we quantify the aforementioned asset A (customer information table):
[0105] Generation of B: According to the BIA report, this table supports core sales operations and is classified as "critical." Through the mapping table, B=1.0. Generation of T: Its RTO is 2 hours (7200 seconds). After boundary processing, calculated using a normalization function, T=0.85. Generation of P: Its RPO is 15 minutes (900 seconds), requiring extremely high precision. After processing and normalization, P=0.95. Generation of C: Its classification label includes a first-class constraint identifier. Through the constraint mapping table, C=1.5. Generation of V: The system detected a high-risk, unpatched vulnerability in the database server (technical risk probability 0.6) and recent suspicious login attempts (human risk probability 0.4). After comprehensive calculation, V=0.5. Finally, asset A obtains a complete set of quantification factors: B=1.0, T=0.85, P=0.95, C=1.5, V=0.5.
[0106] It should be noted that, in specific implementation scenarios, a custom mapping table and function scheme can be adopted based on the above approach. That is, the specific form of the business level mapping table, constraint mapping table, and normalization function used to calculate T and P (such as linear, exponential, or logarithmic) can all be customized by the user according to their own business characteristics and risk preferences. This makes the quantification model more flexible and adaptable.
[0107] An expanded approach to risk data sources can be adopted, meaning the calculation of the environmental risk probability score V can integrate more diverse risk intelligence sources. For example, accessing a real-time cybersecurity threat intelligence platform to obtain the latest attack trends, or obtaining anomaly alert data from physical security systems such as access control and video surveillance, can provide a more comprehensive assessment of environmental risks.
[0108] This approach introduces a coupling relationship between factors. While each factor is generated independently, coupling logic can be designed between factors in certain advanced application scenarios. For example, when the normative constraint coefficient C is very high, the default value of the business criticality coefficient B can be automatically increased, or a stricter threshold can be imposed on the calculation of the time urgency factor T. This coupling relationship deepens the core quantification logic and is an extension of the protection provided by this solution. All of the above optional solutions fall within the protection scope of this application.
[0109] Step S104: Calculate the recovery priority index RPI according to the preset weighting rules, and generate a disaster recovery execution queue accordingly.
[0110] This step involves merging multiple independent quantitative factors (B, T, P, C, V) generated in the previous steps to form a single, sortable comprehensive indicator—the Recovery Priority Index (RPI), and using this as the basis to construct the execution sequence for automated recovery.
[0111] The structure of the pre-set weighted rules: The core of the pre-set weighted rules is a set of basic weight coefficients, namely the first to fourth basic weight coefficients (Wb, Wt, Wp, Wv) corresponding to the business criticality coefficient B, the time urgency factor T, the data loss tolerance factor P, and the environmental risk probability score V, respectively. These weight coefficients reflect the degree of importance that enterprises attach to different dimensions of risk under normal circumstances. For example, financial enterprises may assign a higher weight to Wt (time urgency), while research institutions may place more emphasis on Wb (business criticality).
[0112] Weight normalization verification: To ensure the stability and comparability of RPI calculation results, the system performs a normalization verification on the four basic weight coefficients before calculation, forcing their sum to equal 1 (i.e., Wb+Wt+Wp+Wv=1). This mechanism ensures that regardless of how the weights are allocated, the numerical range of RPI remains within a controllable range, facilitating subsequent threshold comparison and queue sorting.
[0113] The RPI calculation formula: This scheme adopts a carefully designed linear weighted summation formula: RPI = B × Wb + T × Wt + P × Wp + (V × C) × Wv; the (V × C) × Wv term in the formula multiplies the environmental risk probability score V by the regulatory constraint coefficient C and then weights them, meaning that only when a data asset simultaneously satisfies both an environmental risk probability score V greater than a preset risk threshold and a regulatory constraint coefficient C greater than a preset regulatory threshold will its contribution to the RPI be significantly amplified. This design accurately captures the extreme urgency of recovery under the "high risk + strong regulation" scenario.
[0114] Generation of the disaster recovery execution queue: After calculating the RPI for each data asset, the system globally sorts all assets from high to low based on their RPI values, thereby generating an ordered "disaster recovery execution queue". In addition, the system compares the RPI with a preset index threshold, marking assets with higher than this threshold as "first-level priority recovery assets" so that they can receive the highest level of resource protection in subsequent resource scheduling phases.
[0115] For example, continuing with asset A, assume its quantification factors are B=1.0, T=0.85, P=0.95, C=1.5, V=0.5, and the company's preset weights are Wb=0.3, Wt=0.3, Wp=0.2, Wv=0.2 (totaling 1).
[0116] RPI calculation: B×Wb=1.0×0.3=0.3; T×Wt=0.85×0.3=0.255; P×Wp=0.95×0.2=0.19; (V×C)×Wv=(0.5×1.5)×0.2=0.75×0.2=0.15; RPI=0.3+0.255+0.19+0.15=0.895.
[0117] Queue Generation: Assume there is another asset D with an RPI of 0.75. In the generated disaster recovery execution queue, asset A (RPI=0.895) will be ranked before asset D (RPI=0.75). If the preset index threshold is 0.8, asset A will be marked as a "Level 1 Priority Recovery Asset," while asset D will not.
[0118] It should be noted that, in specific implementation scenarios, a nonlinear fusion formula can be applied based on the above scheme. That is, although the embodiment scheme uses linear weighting, the scope of protection can cover the use of more complex nonlinear fusion functions (such as fuzzy logic based on expert systems or simple product terms) to calculate RPI, as long as the core idea is still to fuse the key factors B, T, P, V, and C, and retain the coupling effect of V×C. All of the above optional schemes are within the scope of protection of this application.
[0119] Step S105: In response to external state change events, update the preset weighted rules and rearrange the disaster recovery execution queue.
[0120] This step breaks the rigidity of traditional static disaster recovery solutions, enabling the system to proactively adjust its decision-making logic and re-optimize the recovery sequence based on real-time changes in the internal and external environment.
[0121] Identification of external state change events: The system continuously listens for or subscribes to a series of predefined "external state change events." These events can be divided into two main categories:
[0122] Internal operational events: such as emergency change requests submitted by business departments (e.g., a marketing campaign going live ahead of schedule, resulting in a temporary increase in the importance of related data assets), high-risk vulnerability alerts issued by the IT operations team, or abnormal surges in access to specific assets detected by the monitoring system.
[0123] External environmental events: such as regional natural disaster warnings issued by meteorological departments (e.g., typhoons, floods), APT attack alerts issued by cybersecurity emergency response centers, or notices of the effective date of newly promulgated external rules.
[0124] Dynamic updates to weighted rules: Once a valid state change event is captured, the system triggers a weight adjustment engine. This engine has a built-in mapping relationship between event types and weight adjustment strategies. For example, when a cybersecurity alert of "increased ransomware activity" is received, the engine automatically increases the fourth basic weight coefficient Wv (environmental risk weight); when a notification of "planned downtime maintenance of core business systems" is received, the Wb (business criticality weight) of the relevant assets may be temporarily decreased. This update is not arbitrary, but based on preset, auditable strategies, ensuring the rationality and controllability of the adjustment.
[0125] Incremental reordering of the execution queue: After the weighted rules are updated, the system does not recalculate the RPI for all assets, but instead uses an incremental calculation strategy. It identifies the assets most affected by the weight changes (usually those with high V or B values), recalculates the RPI only for these assets, and reorders the entire disaster recovery execution queue locally or globally based on the new RPI values. This ensures efficient response and avoids performance bottlenecks caused by frequent full calculations in large-scale environments.
[0126] For example, suppose the current time is early Monday morning, and the system is in normal operation.
[0127] Scenario 1: Responding to a security incident
[0128] Event: An external threat intelligence source issued a high-severity alert at the first moment, indicating that a new type of malware was launching a targeted attack on database systems in a specific industry. Response: Upon capturing this event, the system's weighting engine immediately increased the Wv (Environmental Risk Weight) from 0.2 to 0.4, and correspondingly decreased other weights to maintain a sum of 1 (e.g., Wb=0.25, Wt=0.25, Wp=0.1). Reordering: The system recalculated the RPI for all data assets with V>0.3 (i.e., assets considered to be at higher technical risk). Asset A, due to its high V value (0.5) and high C value (1.5), saw its RPI significantly rise from 0.895 to over 0.95, further advancing its position in the queue and ensuring it could be prioritized for backup hardening or isolation before the attack wave arrived.
[0129] Scenario 2: Responding to Business Changes
[0130] Event: The Marketing Department submitted an urgent work order at 5:00 AM that day, announcing that the promotional activity originally scheduled for next month would be launched earlier this Friday due to unforeseen circumstances. The order required all relevant user profiles and transaction data assets to complete disaster recovery preparations by Thursday. Response: The system parsed the work order, identified the relevant asset IDs, and temporarily manually overwrote their B-value (Business Criterion Index) to 1.0 (the highest level). Wt may also be slightly adjusted to emphasize the urgency of the timeframe. Reordering: The RPI of these marked assets was recalculated, quickly rising to the top of the queue. The scheduling system immediately allocated additional network bandwidth and storage resources to them to meet the new business deadline requirements.
[0131] It should be noted that, in specific implementation scenarios, a diversified access scheme for event sources can be adopted on the basis of the above solution. That is, the sources of external state change events can be continuously expanded, such as accessing the enterprise's internal work order system, ITSM platform, Security Orchestration Automation and Response (SOAR) platform, or even connecting to public news and public opinion or supply chain risk monitoring services through API, so as to make the system's perception capabilities more comprehensive.
[0132] A smooth transition scheme for weight adjustment is adopted. To avoid drastic queue jitter caused by sudden weight changes, the system can introduce a "smooth weight transition" mechanism. That is, the weight adjustment is not done all at once, but gradually changes to the target value linearly or exponentially over a time window (such as 10 minutes), thereby making the queue reordering more stable and reducing interference with the ongoing recovery task.
[0133] A historical rollback and version management scheme is adopted, meaning the system can fully record every weight change and queue reordering operation triggered by an event, and supports automatic or manual rollback to the previous weight configuration after the event is resolved. This provides a solid data foundation for disaster recovery drills, post-event reviews, and auditing. All of the above optional solutions fall within the scope of protection of this application.
[0134] Step S106: According to the rearranged disaster recovery execution queue, send control commands to the network controller and virtualization resource pool to perform resource preemption, locking and disaster recovery data recovery.
[0135] This step aims to transform the dynamic priority queue generated in the preceding steps into actual control actions over the underlying infrastructure, ensuring that high-priority data assets can obtain the necessary computing, storage, and network resources in the event of a disaster, thereby enabling efficient recovery.
[0136] The system employs a queue-based instruction scheduling mechanism: using the rearranged disaster recovery execution queue as the sole scheduling basis, it processes data assets in the queue sequentially. For each asset, the system parses its associated metadata (such as the target recovery location, required virtual machine specifications, dependent network policies, etc.) and generates a set of structured control instructions accordingly.
[0137] Instructions for the network controller include: bandwidth guarantee and isolation, allocating dedicated network transmission channels or QoS (Quality of Service) policies to high-priority assets to ensure that their disaster recovery data enjoys the lowest latency and highest throughput during the recovery process; dynamic deployment of security policies, automatically enabling corresponding firewall rules, encrypted tunnels, or access control lists on network paths based on the asset's binding identifier to prevent data leakage during the recovery process; and path optimization, selecting the path with the lowest latency and packet loss rate for data transmission if multiple available recovery paths exist (such as primary and backup links).
[0138] Commands for virtualization resource pools include: Resource Preemption: When high-priority assets lack sufficient resources (such as CPU, memory, and storage IOPS), the system can temporarily reclaim virtual machine resources serving low-priority tasks according to preset policies. Preempted tasks will be suspended or migrated to secondary resource pools instead of being terminated directly, thus ensuring business continuity; Resource Locking: Once a specific virtual machine instance or storage volume is allocated to an asset, a logical lock is applied to that resource to prevent other recovery tasks or routine business operations from occupying it, ensuring the exclusivity and stability of the recovery process; Disaster Recovery Data Execution: Triggers specific recovery operations, including but not limited to: pulling data snapshots from backup media at the corresponding time point, mounting recovery volumes on the target virtual machine, starting database instances and performing log replay to meet RPO requirements, etc.
[0139] The entire execution process emphasizes closed-loop feedback: after each asset recovery operation is completed, the system records the actual time consumed, resource consumption and success status, and sends this information back to the evaluation module for optimization of the weight adjustment strategy in the subsequent step S105.
[0140] For example, continuing with asset A (customer information table), it has jumped to the top of the disaster recovery execution queue after step S105.
[0141] Network control command: The system sends a command to the SDN (Software Defined Networking) controller, requesting the creation of a dedicated VLAN channel with a bandwidth of no less than 1Gbps for the recovery traffic of asset A, and the automatic deployment of an IPSec encrypted tunnel between the source (backup storage node) and the target (production database server) to meet the preset encryption policy requirements for cross-domain data transmission.
[0142] Virtualization resource pool control instructions: The system detects that the host machine where the target database server is located is experiencing resource shortages, but asset D (with a lower RPI) in the queue is running a non-critical reporting service. The system sends a resource preemption instruction to the virtualization management platform (such as vCenter or OpenStack) to suspend the virtual machine corresponding to asset D and migrate it to a standby host machine. Subsequently, the system allocates a high-performance SSD storage volume on the host machine to asset A and sends a resource locking instruction to prevent other tasks from using the volume. Finally, the system calls the backup management interface to initiate a full recovery task for the customer_info table at RPO=15 minutes and monitors its progress until the database service verification is successful.
[0143] Through this series of coordinated operations, Asset A was restored in the shortest possible time with the highest level of security, demonstrating the advantages of the integrated "decision-execution" approach of this solution.
[0144] It should be noted that, in specific implementation scenarios, a multi-level resource pool collaborative scheduling scheme can be adopted based on the above solution. That is, in addition to the local virtualization resource pool, the system can also link the resource pools of public cloud, edge nodes, or off-site disaster recovery centers. When local resources are insufficient to meet high-priority recovery needs, some recovery tasks (such as read-only replica construction) can be automatically offloaded to the cloud to achieve elastic disaster recovery in a hybrid cloud environment.
[0145] The system adopts an atomicity and rollback mechanism for recovery tasks. For complex multi-component assets (such as an overall business system containing applications, databases, and configuration files), the entire recovery process can be encapsulated as a "recovery transaction." If any sub-step fails, an automatic rollback can be triggered, releasing preempted resources and restoring the system to its pre-disaster state, thus avoiding the creation of a partially recovered, dirty data environment.
[0146] A hierarchical authorization and auditing scheme for command execution is adopted, meaning that all control commands sent to the network controller and virtualization platform can be bound to operation permission policies. For example, resource preemption operations require secondary confirmation from the security administrator (in non-emergency mode), and the sending, receiving, and execution results of all commands are recorded in an immutable log, meeting the preset system security audit requirements. All of the above optional schemes fall within the scope of protection of this application.
[0147] In some embodiments of this application, to address the issues of inconsistent metadata formats and semantic fragmentation in multi-source heterogeneous data environments, and thus provide structurally consistent and informationally complete basic inputs for step S102 and subsequent classification, quantification, and scheduling, the metadata is mapped to unified common fields and a data asset catalog with unique identifiers is generated, including:
[0148] The metadata is collected from at least one data source, including database servers, file servers, cloud storage services, and local disks.
[0149] The table name, field name, number of table records, and update time in the structured data, and the file name, storage path, number of file bytes, and modification time in the unstructured data are respectively mapped to the asset name, asset location, data volume, data format, and timestamp fields in the data asset catalog;
[0150] Generate data feature labels for field relationships in structured data, and generate extended feature values for storage paths, file extensions, and storage media attributes in unstructured data;
[0151] Write the data feature tags and extended feature values into the extended metadata field of the data asset catalog, and configure a unique asset identifier for each data asset.
[0152] As described above, the system first collects raw metadata from at least one data source, including but not limited to relational database servers (such as MySQL and Oracle), file servers (such as NAS and SMB shared storage), public or private cloud object storage services (such as AWS S3 and Alibaba Cloud OSS), and local disks on terminal devices. The collection process can be implemented through agents, API interfaces, or log monitoring to ensure coverage of the main data storage locations within the enterprise.
[0153] Subsequently, the system standardizes and maps the collected metadata. For structured data (such as database tables), the table name is mapped to the "asset name" in the data asset catalog, the database instance or connection string containing the table is mapped to the "asset location," the number of records in the table is mapped to the "data volume," the data type distribution of the field set is mapped to the "data format" (e.g., "primarily numeric," "mixed text"), and the last data update time is mapped to the "timestamp." For unstructured data (such as documents, images, and log files), the file name is mapped to the "asset name," the complete storage path (including bucket name and directory level) is mapped to the "asset location," the file size in bytes is mapped to the "data volume," the file extension or MIME type is mapped to the "data format," and the last modification time of the file is mapped to the "timestamp."
[0154] Building upon this foundation, the system further extracts higher-order semantic features. For structured data, it analyzes the primary and foreign key relationships, dependency references, or business logic connections between fields (e.g., the "User ID" field appears in multiple tables), generating "data feature tags" that describe the inherent relationships within the data, such as "core user identifier association" or "transaction flow chain." For unstructured data, it combines the semantics of its storage path (e.g., the path contains " / finance / " or " / HR / "), file extensions (e.g., ".pdf", ".xlsx"), and underlying storage media attributes (e.g., whether it is located on a high-performance SSD, whether WORM write protection is enabled) to generate "extended feature values," such as "financially sensitive PDFs" or "regulated personnel files."
[0155] Finally, the system writes the generated data feature tags and extended feature values into the "Extended Metadata Field" of the data asset catalog, and assigns a globally unique asset identifier (such as a UUID or a namespace-based hash ID) to each data asset included in the catalog. This unique identifier is used throughout the entire disaster recovery assessment and recovery process, ensuring that subsequent steps can accurately and unambiguously reference and operate the corresponding data assets.
[0156] In some embodiments of this application, to avoid misjudgments or omissions caused by a single rule, improve the accuracy and business relevance of classification results, and quickly identify high-risk, high-value assets, thereby reducing manual intervention costs and enhancing the targeting of disaster recovery strategies, the data assets in the data asset catalog are classified according to a preset multi-branch decision logic strategy to generate classification result tags, including:
[0157] Extract the content keywords of the data assets. If the content keywords match the preset set of highly sensitive keywords, directly output the classification result tags containing high risk level and core asset markers.
[0158] If the preset set of highly sensitive keywords is not matched, the file type is determined based on the file extension of the data asset, and the volume level is determined based on the data volume parameter of the data asset.
[0159] A combined condition determination is performed based on the content keywords, file type, and size level;
[0160] Specifically, if the file type is code or configuration file, the output will include a classification result label containing the core asset marker; if the file type is a structured data file and the size level reaches the preset large size level, or hits the preset set of sensitive keywords, the output will include a classification result label containing financial data, operational data, or sensitive data markers; if no preset branch is matched, the output will be a classification result label pending manual review.
[0161] As described above, the system first extracts keywords from the content of the data assets. This process can be implemented based on full-text indexing, regular expression matching, or a lightweight NLP model to extract key terms that reflect the semantics of the data (such as "identity features," "bank card," "salary," "source code," etc.). If any of the extracted keywords matches a preset set of highly sensitive keywords (high-risk words explicitly defined by external rules), no further judgment is needed. The system directly generates a classification result label containing "high-risk level" and "core asset marker" for the asset, ensuring that such assets are prioritized in subsequent processes.
[0162] If the set of highly sensitive keywords is not matched, the secondary judgment process is initiated: the system determines the file type of the data asset based on its file extension (such as ".sql", ".xlsx", ".log", ".py", ".conf", etc.); at the same time, it determines the size level (such as "small size", "medium size", "large size") by comparing its data volume parameter (i.e. the number of bytes or records mapped in the previous steps) with a preset threshold.
[0163] Based on this, the system performs a combined conditional judgment: if the file type is identified as program source code (such as ".java", ".py"), script file (such as ".sh", ".bat"), or system / application configuration file (such as ".xml", ".yaml", ".ini"), even if its content does not explicitly contain sensitive words, it is marked as a "core asset" because it plays a fundamental supporting role in the system's operation; if the file type is a structured data file (such as a database export file ".csv", ".parquet", or table snapshot), and its size reaches the preset large size level (e.g., exceeding 1TB or 1 billion records), or its content keywords match the preset set of medium-sensitive keywords (such as "orders", "customer names", "internal reports", etc.), then the corresponding semantic label is output according to the business context, such as "financial data", "operational data", or "sensitive data"; if none of the above branch conditions are met (e.g., a small log fragment without an extension and without any sensitive keywords), then the system cannot automatically determine its business attributes or risk level, and outputs the classification result label "pending manual review", which is then handed over to security or data governance personnel for further analysis.
[0164] Through the aforementioned multi-branch, progressively advancing judgment mechanism, the system can ensure that high-risk assets are not overlooked while taking into account the business characteristics of different types of data assets, achieving efficient, reliable, and interpretable automated classification.
[0165] In some embodiments of this application, by introducing data from external authoritative systems, semantic classification labels, and multidimensional risk assessment, the quantitative results are ensured to possess business authenticity, external rule constraints, and environmental adaptability. The generation of a business criticality coefficient B, a normative constraint coefficient C, and an environmental risk probability score V includes:
[0166] The business impact analysis level associated with the data asset is read from the configuration management database or business continuity management system, and converted into the business criticality coefficient B through a preset business level mapping table; if the business impact analysis level is missing, a preset median business criticality coefficient is assigned and a missing alarm is generated.
[0167] Read the data sensitivity identifier or industry compliance attribute identifier from the classification result label, and convert it into the normative constraint coefficient C through a preset constraint mapping table;
[0168] The environmental risk probability score V is calculated based on at least two of the following: physical risk probability, technological risk probability, and human risk probability.
[0169] The environmental risk probability score V is used to characterize the current threat level of the data asset under the dimensions of physical risk, technical risk, or human risk, and the normative constraint coefficient C is used to characterize the degree of mandatory compliance recovery constraints on the data asset.
[0170] As mentioned above, regarding the generation of the business criticality coefficient B: The system first attempts to obtain the Business Impact Analysis (BIA) level associated with the data asset from the enterprise's existing Configuration Management Database (CMDB) or Business Continuity Management System (BCMS). This level is typically assessed by the business department during the disaster recovery planning phase, reflecting the degree of impact of the asset's outage on business operations, revenue, or reputation (e.g., "critical," "important," "moderate"). The system converts this level into a standardized business criticality coefficient B using a pre-defined business level mapping table (e.g., "critical" → B=1.0, "important" → B=0.75, "moderate" → B=0.5). If the BIA level is missing due to the asset not being registered or not covered by the system, the system will not skip the asset but will assign a pre-defined median business criticality coefficient (e.g., B=0.6) and generate a "BIA level missing" alarm for subsequent governance improvements, preventing the asset from being incorrectly underestimated due to incomplete information.
[0171] Regarding the generation of the normative constraint coefficient C: The system parses the classification result labels generated in the preceding steps, extracting data sensitivity identifiers (such as "identity feature data" and "health data") or industry compliance attribute identifiers. Based on a preset constraint mapping table, these identifiers are converted into corresponding normative constraint coefficients C. This constraint mapping table is pre-configured by the system administrator according to the enterprise's industry, business region, and constraint rules, and supports dynamic updates. For example, its core rules are as follows:
[0172] If the data asset is data that matches the first preset data protection specification, then it may correspond to C=1.5;
[0173] If the data asset is system data that meets the requirements of the preset third security protection level, then it may correspond to C=1.3;
[0174] If the asset belongs to a critical infrastructure industry such as finance, healthcare, or energy and is subject to specific industry data security regulations (such as pre-defined industry data security classification configuration rules), then C=1.2;
[0175] If multiple conditions are met simultaneously (such as data that matches the first preset data protection specification and system data that meets the requirements of the preset third security protection level), then the maximum value of the corresponding C values is taken, or calculated according to the preset superposition rules (such as C=1+max(0.5,0.3,0.2)=1.5) to avoid repeated amplification;
[0176] For data assets without any compliance labels or that only meet basic internal policies, the default value is C=1.0.
[0177] This mapping table is stored in the policy configuration repository in key-value pairs, supporting CRUD operations via the management interface to ensure that changes to constraint rules can be quickly synchronized to the disaster recovery decision-making logic. This coefficient directly affects the final RPI value, reflecting the mandatory requirements of external rules or industry standards for recovery timeliness, integrity, or audit traceability, ensuring that compliant assets receive the necessary protection during resource scheduling.
[0178] Regarding the generation of the environmental risk probability score V: The system calculates the V value by integrating risk probabilities from at least two dimensions. Physical risk probability can be derived from weather warnings, geographical location risk maps (such as earthquake zones, flood areas), or data center infrastructure monitoring (such as UPS load, abnormal temperature and humidity); technical risk probability can be based on vulnerability scan results, security event logs (such as recent DDoS attacks), system patch status, or backup link health; human risk probability can refer to the complexity of permission allocation, recent personnel change records, and the frequency of violations discovered by internal audits. Each risk dimension can be assigned different weights, and the scores can be merged into a single environmental risk probability score V (typically ranging from 0 to 1) through weighted averaging, maximum value selection, or logical combination (such as "any high risk increases V"). The higher the V value, the more unstable the current environment of the asset, the greater the possibility of damage, and the more priority should be given to restoration to avoid potential losses.
[0179] In summary, this step, by structurally integrating business, compliance, and environmental factors, constructs a multi-dimensional quantitative profile for each data asset, laying a core data foundation for achieving accurate and adaptive disaster recovery scheduling.
[0180] In some embodiments of this application, to address the common issues of missing, abnormal, or inconsistent units in the original RTO value in actual systems, mathematical normalization is used to ensure that the business logic of "the shorter the RTO, the higher the urgency" is accurately reflected in the quantitative model. Generating the time urgency factor T includes:
[0181] Obtain the recovery time target of the data asset and convert the recovery time target into a second-level value RTOsec;
[0182] If the RTOsec is missing or less than or equal to zero, the RTOsec is replaced with a preset median second value; if the RTOsec is greater than a preset maximum truncation threshold, the RTOsec is truncated to the preset maximum truncation threshold; if the RTOsec is less than a preset minimum emergency threshold, the RTOsec is replaced with the preset minimum emergency threshold.
[0183] Based on the boundary-processed RTOsec, a monotonically decreasing normalization calculation is performed to obtain the time urgency factor T, such that the smaller the RTOsec, the larger the time urgency factor T.
[0184] As described above, the system first retrieves the Recovery Time Objective (RTO) set for the data asset from the Configuration Management Database (CMDB), Service Level Agreement (SLA) document, or Business Continuity Plan. The unit is usually hours, minutes, or seconds. The system then converts this value into a value in seconds, denoted as RTOsec, for subsequent calculations and processing.
[0185] Subsequently, the system performs boundary checks and corrections on RTOsec: If RTOsec is missing (i.e., RTO is not configured) or its value is less than or equal to zero (invalid configuration), the system does not consider it to be without urgency, but assigns it a preset median value in seconds (e.g., 43200 seconds, or 12 hours), representing the default recovery tolerance window for typical business systems, and simultaneously records an "RTO missing or invalid" alarm for subsequent governance; if RTOsec exceeds the preset maximum truncation threshold (e.g., 86400 seconds, or 24 hours), it indicates that the asset has extremely low requirements for recovery timeliness. To avoid lowering the overall distribution sensitivity during normalization, the system forcibly truncates it to this maximum threshold; if RTOsec is lower than the preset minimum urgency threshold (e.g., 300 seconds, or 5 minutes), it indicates that the asset belongs to an extremely high timeliness requirement scenario. To highlight its urgency and prevent the value from being too small and causing calculation distortion, the system replaces it with this minimum urgency threshold to ensure that it obtains an urgency factor close to the maximum value after normalization.
[0186] After completing the above boundary processing, the system performs a monotonically decreasing normalization calculation based on the corrected RTOsec to generate the time urgency factor T. Typical implementations include, but are not limited to, using an inverse proportional function (e.g., T = 1 / (1 + α·RTOsec)) or a linear mapping (e.g., T = (RTO_max − RTOsec) / (RTO_max − RTO_min)), where α is an adjustment parameter. Regardless of the form used, it is ensured that the smaller the RTOsec, the larger the T value, thus assigning a higher priority weight to assets with high timeliness requirements in the RPI formula.
[0187] Through this mechanism, the system not only improves the availability and robustness of RTO data, but also ensures the reasonable driving role of the time dimension in recovery scheduling decisions.
[0188] In some embodiments of this application, to address issues such as missing values, extreme values, or inconsistent units in RPO configurations often found in real-world environments, normalization is used to ensure that the business logic of "the smaller the RPO (i.e., the less data loss is allowed), the larger the P value" is accurately reflected. This allows for priority protection of assets with higher data integrity requirements during resource scheduling. Generating the data loss tolerance factor P includes:
[0189] Obtain the recovery point target of the data asset and convert the recovery point target into a second-level value RPOsec;
[0190] If RPOsec is missing or less than zero, RPOsec is replaced with a preset median second value; if RPOsec is equal to zero, the data loss tolerance factor P is assigned a preset maximum score; if RPOsec is greater than a preset maximum tolerance threshold, RPOsec is truncated to the preset maximum tolerance threshold.
[0191] Based on the boundary-processed RPOsec, a monotonically decreasing normalization calculation is performed to obtain the data loss tolerance factor P, such that the smaller RPOsec is, the larger the data loss tolerance factor P is.
[0192] As described above, the system first retrieves the Recovery Point Objective (RPO) set for the data asset from the Configuration Management Database (CMDB), data governance platform, or business continuity plan. This metric represents the maximum acceptable data loss time window in the event of a disaster. The system then converts this value into a uniform value in seconds, denoted as RPOsec, for subsequent numerical processing.
[0193] Subsequently, the system performs boundary checks and corrections on RPOsec: If RPOsec is missing (not configured) or less than zero (invalid value), the system uses a preset median value in seconds (e.g., 3600 seconds, or 1 hour) as a substitute value, representing the default data loss tolerance level under typical business scenarios, and generates an "RPO missing or invalid" alarm for subsequent data governance improvement; if RPOsec equals zero, it indicates that the asset requires zero data loss (such as the core accounting system for financial transactions). In this case, the system directly assigns the data loss tolerance factor P to the preset maximum score (e.g., P=1.0) to highlight its extreme requirements for data integrity; if RPOsec exceeds the preset maximum tolerance threshold (e.g., 86400 seconds, or 24 hours), it indicates that the asset has an extremely high tolerance for data loss (such as archived logs or non-critical reports). To avoid distorting the overall distribution during normalization, the system truncates it to the maximum tolerance threshold.
[0194] After completing the boundary processing described above, the system performs a monotonically decreasing normalization calculation based on the corrected RPOsec to generate the data loss tolerance factor P. Typical methods include inverse proportional mapping (e.g., P = 1 / (1 + β·RPOsec)) or linear compression (e.g., P = (RPO_max - RPOsec) / (RPO_max - RPO_min)), where β is an adjustment parameter. Regardless of the form used, it is ensured that the smaller RPOsec is, the larger the P value, thereby assigning higher weights to low-tolerance assets in the RPI model.
[0195] Through this mechanism, the system not only improves the availability and consistency of RPO data, but also achieves precise quantification of data integrity requirements, providing a reliable basis for the differentiated allocation of disaster recovery resources.
[0196] In some embodiments of this application, in order to achieve the coordinated quantification of multi-dimensional heterogeneous indicators, disaster recovery resources can be dynamically and accurately prioritized based on comprehensive risk and business value. The Recovery Priority Index (RPI) is calculated according to a preset weighting rule, including:
[0197] Obtain the first basic weight coefficient Wb corresponding to the business criticality coefficient B, the second basic weight coefficient Wt corresponding to the time urgency factor T, the third basic weight coefficient Wp corresponding to the data loss tolerance factor P, and the fourth basic weight coefficient Wv corresponding to the environmental risk probability score V;
[0198] Normalize and verify the first basic weight coefficient Wb, the second basic weight coefficient Wt, the third basic weight coefficient Wp and the fourth basic weight coefficient Wv to make them satisfy Wb+Wt+Wp+Wv=1;
[0199] The Recovery Priority Index (RPI) is calculated using the following formula:
[0200] RPI=B×Wb+T×Wt+P×Wp+V×C×Wv;
[0201] The product term V×C×Wv of the environmental risk probability score V and the normative constraint coefficient C is used to positively weight the recovery priority index RPI to improve the ranking priority of the corresponding data assets in the disaster recovery execution queue.
[0202] The Recovery Priority Index (RPI) is compared with a preset index threshold. If the RPI is greater than the preset index threshold, the corresponding data asset is marked as a first-level priority recovery asset.
[0203] As mentioned above, the system first obtains the basic weight coefficients corresponding to four dimensions: the first basic weight coefficient Wb, which is used to adjust the influence of the business criticality coefficient B in RPI; the second basic weight coefficient Wt, which is used to adjust the contribution of the time urgency factor T; the third basic weight coefficient Wp, which is used to reflect the importance of the data loss tolerance factor P; and the fourth basic weight coefficient Wv, which is used to control the weight ratio of environmental risk-related items.
[0204] These weighting coefficients can be preset by the operations team or security governance department based on organizational strategy, industry characteristics, or disaster scenarios (for example, the financial industry might assign higher values to Wb and Wp, while cloud service providers might focus more on Wt). To ensure the comparability and stability of the sum of contributions from each dimension, the system performs normalization checks on Wb, Wt, Wp, and Wv, forcibly satisfying the constraint: Wb + Wt + Wp + Wv = 1. If the initial configuration does not meet this condition, the system will scale proportionally or prompt a configuration error to ensure the mathematical rationality of the calculation results.
[0205] Subsequently, the system calculates the Recovery Priority Index (RPI) according to the following formula: RPI = B × Wb + T × Wt + P × Wp + V × C × Wv. The first three terms are linear weighted terms, reflecting business value, timeliness requirements, and data integrity needs, respectively. The fourth term, V × C × Wv, is a composite weighted term with a specific semantic meaning: when a data asset is both in a high-environment-risk state (high V value, such as being located in an earthquake-prone area or recently subjected to a cyberattack) and has strong compliance constraints (C > 1, such as involving GDPR or financial regulation), this term is significantly amplified, thus generating a "risk-compliance superposition effect" in the RPI, ensuring that such assets receive higher priority in the recovery queue. This coupling mechanism avoids scheduling blind spots caused by relying solely on static attributes or isolated risk assessments.
[0206] Finally, the system compares the calculated RPI value with a preset index threshold (e.g., 0.85). If the RPI is greater than the threshold, the data asset is determined to be an extremely high-priority object, automatically marked as a "Level 1 Priority Recovery Asset," and prioritized in the disaster recovery execution phase by allocating recovery resources such as bandwidth, storage, and computing to ensure that recovery is completed within a limited window, minimizing business interruption losses and rule-related risks.
[0207] In summary, this step, through configurable weights, multi-factor fusion, and a risk-compliance linkage mechanism, constructs a dynamic priority decision-making model that combines flexibility, compliance sensitivity, and risk response capabilities, significantly improving the intelligence level and business adaptability of the disaster recovery system.
[0208] In some embodiments of this application, to overcome the limitations of traditional static weight configuration and enable the recovery strategy to automatically adjust resource allocation priorities based on changes in security posture, constraint rules, or infrastructure status, the following steps are taken: Responding to external status change events by updating the preset weighting rules and rearranging the disaster recovery execution queue, including:
[0209] External state change events are obtained through the event capture interface. These external state change events include at least one of the following: security events, vulnerability discovery events, external rule change events, and protection device configuration change events.
[0210] The external state change event is converted into a formatted structured message by the event converter. The formatted structured message includes at least the event type, event source, event time, event payload, and severity of impact.
[0211] The formatted structure message is routed to the corresponding weight adjustment process according to the event type, and the weight adjustment range is determined according to the severity of the impact.
[0212] Update at least one basic weight coefficient in the preset weighting rule according to the weight adjustment range, and perform weight normalization processing after the update;
[0213] A weight update signal is sent to the recovery priority index calculation module to trigger the recalculation of the recovery priority index RPI, and the disaster recovery execution queue is rearranged according to the recalculated recovery priority index RPI.
[0214] As described above, the system continuously monitors external state change events from Security Information and Event Management (SIEM), vulnerability scanning systems, regulatory bulletin push services, network protection devices (such as firewalls and WAFs), or configuration management databases through a standardized event capture interface. These events include at least one of the following types: security events (such as ransomware attack alerts), vulnerability discovery events (such as a zero-day vulnerability with a CVSS score ≥ 9.0 being exposed in a critical system), external rule change events (such as newly issued data localization requirements), or protection device configuration change events (such as the failure of firewall policies in core areas). These events reflect substantial changes in operational environment risks or compliance constraints.
[0215] After the raw event is captured, a dedicated event converter parses it and converts it into a unified formatted structured message. This message contains at least five fields: event type (used for classification and processing logic), event source (identifying the system that generated the event), event time (used for timeliness determination), event payload (containing specific context, such as a list of affected assets, vulnerability ID, external rule clause number, etc.), and impact severity (usually expressed as a numerical value or a rating, such as "high / medium / low" or a score between 0 and 1), to quantify the intensity of the event's impact on disaster recovery strategies.
[0216] Subsequently, the system routes the formatted structured message to the corresponding preset weight adjustment process based on the event type. For example, a security incident or a high-risk vulnerability incident may trigger a process that increases the environmental risk-related weight Wv; an external rule change incident may activate logic that increases the business criticality weight Wb or the influence of compliance coupling items; and if a change in the configuration of protection equipment leads to an expansion of the exposure surface in a certain area, both Wv and Wt may be increased simultaneously. In the corresponding process, the system determines the specific weight adjustment range based on the severity of the impact carried by the event—the higher the severity, the greater the increase or decrease in the corresponding basic weight coefficient. For example, when an active attack against the core database is detected (severity = 0.9), Wv can be temporarily increased by 0.15.
[0217] After the weight adjustment is completed, the system performs normalization processing on the updated basic weight coefficients (at least one of Wb, Wt, Wp, and Wv) to force Wb+Wt+Wp+Wv=1, so as to maintain the numerical stability and comparability of the RPI calculation model.
[0218] Finally, the system sends a weight update signal to the recovery priority index calculation module. This signal triggers a full or incremental (affected assets only) recalculation of the RPI. Based on the newly generated RPI value, the disaster recovery scheduling engine dynamically rearranges the existing disaster recovery execution queue to ensure that data assets whose risk has increased or whose constraints have become more stringent due to external events can obtain a higher recovery position, thereby completing the deployment of critical resources before the actual threat occurs or the compliance window closes.
[0219] Through the above mechanism, this solution has evolved from "passive response disaster recovery" to "proactive perception-dynamic optimization disaster recovery", which significantly improves the system's risk resistance and compliance assurance level in complex and ever-changing operating environments.
[0220] In some embodiments of this application, to address the delay in critical business recovery caused by "equal recovery of all assets" or "static resource allocation" in traditional disaster recovery systems, and to ensure that first-priority recovery assets have preemptive scheduling capabilities in critical resources such as bandwidth, computing, and storage, thereby significantly shortening their recovery time (RTO) and improving overall business continuity, control commands are sent to the network controller and virtualized resource pool according to the rearranged disaster recovery execution queue to execute resource preemption, locking, and disaster recovery data recovery, including:
[0221] Data assets whose Recovery Priority Index (RPI) is greater than the preset index threshold are identified as Level 1 priority recovery assets.
[0222] For the first-priority recovery assets, a bandwidth expansion command and a bandwidth preemption command are sent to the software-defined network controller to pre-lock a preset proportion of network bandwidth resources;
[0223] Send computing resource scheduling instructions and storage resource scheduling instructions to the virtualization resource pool to allocate a specified number of CPU cores, memory resources and solid-state storage resources to the first-priority recovery assets, and perform parallel recovery;
[0224] For data assets not identified as first-priority recovery assets, allocate the remaining shared bandwidth and default resource quotas, and perform serial recovery;
[0225] After the disaster recovery data is restored, a consistency check is performed on the restored data, and the pre-locked network bandwidth resources, computing resources and storage resources are released after the check passes.
[0226] As described above, the system first identifies data assets with a Recovery Priority Index (RPI) greater than a preset threshold (e.g., 0.85) based on the rearranged disaster recovery execution queue, and explicitly marks them as "Level 1 Priority Recovery Assets". These assets typically have high business criticality, stringent recovery time requirements, zero tolerance for data loss, or are in high-risk compliance scenarios.
[0227] For assets requiring Level 1 priority recovery, the system sends two types of instructions to the Software-Defined Networking (SDN) controller: first, a bandwidth expansion instruction to temporarily increase the total available bandwidth of the target transmission link (e.g., from 1Gbps to 10Gbps); and second, a bandwidth preemption instruction to pre-lock a preset proportion (e.g., 70%) of dedicated bandwidth resources from the current shared bandwidth pool to ensure that the disaster recovery data replication stream is not interfered with by other low-priority traffic. This locking state continues until recovery is complete and verification is successful.
[0228] Simultaneously, the system sends compute and storage resource scheduling instructions to the virtualization resource pool (such as a resource platform based on VMware vSphere, OpenStack, or Kubernetes) to dynamically allocate a specified number of CPU cores, memory capacity (such as 32GB RAM), and high-performance solid-state storage (SSD) space (such as 1TB) to the first-priority recovery assets. These resources are exclusively reserved, do not participate in shared scheduling, and multi-threaded or distributed parallel recovery mechanisms are enabled to maximize recovery throughput efficiency.
[0229] For the remaining data assets that are not marked as first-priority recovery assets, the system only allocates the remaining shared network bandwidth (i.e., the remaining amount of total bandwidth after deducting the locked portion) and the default resource quota (such as 2-core CPU, 4GB memory, and mechanical hard disk storage), and executes the recovery tasks sequentially in a serial manner to avoid resource contention for high-priority recovery processes.
[0230] After all disaster recovery data restoration operations are completed, the system automatically performs consistency checks on the restored data, including but not limited to verification and comparison, transaction log replay integrity checks, and primary / standby data fingerprint matching. Only when the verification results pass, confirming that the data is complete and available, will the system send resource release commands to the SDN controller and virtualization resource pool, releasing the previously pre-locked network bandwidth, CPU, memory, and storage resources, allowing them to return to the shared resource pool for subsequent regular business or other disaster recovery tasks.
[0231] Through the above mechanism, this solution achieves "priority-driven dynamic resource protection", which not only improves the disaster recovery efficiency of critical assets, but also takes into account resource utilization efficiency and overall system stability, effectively supporting the operational needs of modern IT infrastructure with high availability and high constraints.
[0232] Compared with existing technologies, this application discloses a data asset disaster recovery processing method. This method collects metadata from structured and unstructured data and generates a unified data asset catalog, laying the foundation for subsequent processing. Based on this, the solution creatively introduces five core quantitative indicators: business criticality coefficient B, time urgency factor T, data loss tolerance factor P, regulatory constraint coefficient C, and environmental risk probability score V. These indicators correspond to business impact, time tolerance, data loss tolerance, compliance constraints, and environmental vulnerability, respectively, thus constructing a multi-dimensional and comprehensive evaluation system. The Recovery Priority Index (RPI), calculated through preset weighted rules, can scientifically and objectively reflect the comprehensive recovery value and urgency of each data asset in the current environment.
[0233] This solution transforms recovery constraint indicators such as business impact analysis level, recovery time target, and recovery point target, as well as recovery priority requirements in external constraint rules, into quantifiable factors (B, T, P, C) that can be directly used in calculations through specific mathematical transformations and mapping logic. In particular, the design of the normative constraint coefficient C ensures that data assets with high-level constraint labels, even if their environmental risk probability scores are low, can still obtain a higher ranking in the recovery priority index calculation due to the increased constraint strength, thus being prioritized for recovery.
[0234] The solution not only calculates the RPI but also incorporates a dynamic feedback mechanism. When the system detects external state changes such as security events or changes in external rules, it automatically updates the preset weighted rules and triggers the recalculation of the RPI and the reordering of the disaster recovery queue. Ultimately, the system can directly send control commands to the network controller and virtualization resource pool to execute resource preemption, locking, and data recovery operations. This complete closed loop from "intelligent assessment" to "automatic execution" greatly improves the speed and efficiency of disaster recovery response, ensuring that recovery actions remain consistent with the latest business and security posture in complex and ever-changing environments.
[0235] Based on the same inventive concept as the methods described above, this application also proposes a data asset disaster recovery system, such as... Figure 2 The diagram shown is a structural schematic of a data asset disaster recovery system, which includes:
[0236] The heterogeneous data alignment module is used to collect metadata of structured and unstructured data in the enterprise IT environment, map it to a unified common field, and generate a data asset catalog with a unique identifier.
[0237] The multi-branch rule classification module is used to classify data assets in the data asset catalog according to a preset multi-branch decision logic strategy and generate classification result labels;
[0238] The indicator quantification module is used to generate a business criticality coefficient B, a time urgency factor T, a data loss tolerance factor P, a regulatory constraint coefficient C, and an environmental risk probability score V based on the business impact analysis level, recovery time target, recovery point target, classification compliance attributes, and environmental risk data.
[0239] The priority calculation module is used to calculate the recovery priority index (RPI) according to a preset weighting rule, and generate a disaster recovery execution queue accordingly.
[0240] The dynamic feedback module is used to update the preset weighting rules and rearrange the disaster recovery execution queue in response to external state change events;
[0241] The disaster recovery automated scheduling module is used to send control commands to the network controller and virtualization resource pool according to the rearranged disaster recovery execution queue, and to perform resource preemption, locking and disaster recovery data recovery.
[0242] As mentioned above, the heterogeneous data alignment module is responsible for connecting to various data sources within the enterprise, including relational databases, NoSQL systems, file servers, object storage, and log platforms, collecting metadata information for both structured (e.g., table fields, primary keys, foreign keys) and unstructured (e.g., document type, creator, access frequency) data. This module has a built-in mapping rule engine that converts metadata fields from different sources (e.g., "table_name", "file_path", "owner") into predefined unified public fields (e.g., "asset_name", "data_owner", "sensitivity_level"), and generates a globally unique identifier (e.g., UUID or hash ID) for each data asset, ultimately forming a standardized data asset catalog as the basic input for subsequent processing.
[0243] The multi-branch rule classification module automatically classifies each asset in the data asset catalog based on a preset multi-branch decision logic strategy (such as a decision tree or rule chain). This logic can include multiple decision levels, for example: first, determining whether it falls within the scope of external rule constraints; then, determining the business domain to which it belongs; and finally, determining the data format. According to the matching path, the system assigns one or more classification result tags to the asset (such as "Core Business - First Type of Constraint Identifier - High Sensitivity"), which are used to reference the corresponding compliance attributes and processing strategies when quantifying indicators later.
[0244] The indicator quantification module integrates business impact analysis results, service level agreement configuration, constraint rule base, and real-time environmental monitoring data to calculate five core quantitative indicators: Business Criticality Coefficient B, mapped to a value between 0 and 1 based on the business impact analysis level (e.g., high / medium / low); Time Urgency Factor T, converted and normalized from Recovery Time Objective (RTO), with a larger T for shorter RTOs; Data Loss Tolerance Factor P, calculated from Recovery Point Objective (RPO) through boundary processing and monotonically decreasing normalization; Normative Constraint Coefficient C, matched with external rule clauses based on classification labels, where C>1 if pre-set strong security constraints are involved, otherwise C=1; and Environmental Risk Probability Score V, a probabilistic score generated by integrating the security posture of the network area where the asset is located (e.g., vulnerability exposure surface, attack frequency, geographical location risk).
[0245] The priority calculation module receives the aforementioned quantitative indicators and calculates the Recovery Priority Index (RPI) for each data asset based on preset weighting rules (i.e., the basic weights Wb, Wt, Wp, and Wv for each dimension). The formula is: RPI = B × Wb + T × Wt + P × Wp + V × C × Wv. After calculation, the system sorts the assets by RPI value from high to low, generates an initial disaster recovery execution queue, and marks assets with RPI values exceeding the threshold as first-level priority recovery assets.
[0246] The dynamic feedback module continuously monitors external state change events (such as security alarms, new external rule changes, and firewall policy changes). Through event capture, formatting, routing, and severity assessment, it dynamically adjusts one or more basic weight coefficients in the aforementioned weighted rules and triggers the recalculation of RPI after normalization. This, in turn, rearranges the disaster recovery execution queue, enabling the system to have real-time response capabilities to changes in the operating environment.
[0247] The disaster recovery automated scheduling module performs physical recovery operations based on the latest queue: for first-priority recovery assets, it sends bandwidth preemption and expansion commands to the SDN controller, requests dedicated CPU, memory, and SSD resources from the virtualization resource pool, and initiates parallel recovery; for other assets, it allocates remaining shared resources and recovers them serially. After recovery is completed, it performs data consistency verification. Once the verification passes, it releases the occupied resources, completing the entire disaster recovery scheduling closed loop.
[0248] In summary, by organically integrating data alignment, intelligent classification, multi-dimensional quantification, dynamic priority calculation, and automated execution, this system achieves end-to-end disaster recovery capabilities from "asset visibility" to "recovery controllability," significantly improving the business continuity assurance level of enterprises in complex hybrid IT environments.
[0249] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program goods. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program goods embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0250] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program goods according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0251] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0252] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0253] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A data asset disaster recovery processing method, characterized in that, include: Collect metadata of structured and unstructured data in the enterprise IT environment, map it to a unified common field, and generate a data asset catalog with a unique identifier; The data assets in the data asset catalog are classified according to a preset multi-branch decision logic strategy, and classification result labels are generated. The business impact analysis level associated with the data asset is read from the configuration management database or business continuity management system, and converted into a business criticality coefficient B through a preset business level mapping table; if the business impact analysis level is missing, a preset median business criticality coefficient is assigned and a missing alarm is generated. Read the data sensitivity identifier or industry compliance attribute identifier from the classification result label, and convert it into the normative constraint coefficient C through a preset constraint mapping table; Based on at least two of the physical risk probability, technical risk probability, and human risk probability, an environmental risk probability score V is calculated. The environmental risk probability score V is used to characterize the current degree of threat to the data asset under the dimensions of physical risk, technical risk, or human risk. The normative constraint coefficient C is used to characterize the degree of enforcement of compliance recovery constraints on the data asset. The recovery time target of the data asset is obtained and converted into a second-level value RTOsec. If the RTOsec is missing or less than or equal to zero, the RTOsec is replaced with a preset median second-level value. If the RTOsec is greater than a preset maximum truncation threshold, the RTOsec is truncated to the preset maximum truncation threshold. If the RTOsec is less than a preset minimum urgency threshold, the RTOsec is replaced with the preset minimum urgency threshold. Based on the boundary-processed RTOsec, a monotonically decreasing normalization calculation is performed to obtain a time urgency factor T, such that the smaller the RTOsec, the larger the time urgency factor T. The recovery point target of the data asset is obtained and converted into a second-level value RPOsec. If RPOsec is missing or less than zero, it is replaced with a preset median second-level value. If RPOsec is equal to zero, the data loss tolerance factor P is assigned a preset maximum score. If RPOsec is greater than a preset maximum tolerance threshold, it is truncated to the preset maximum tolerance threshold. Based on the boundary-processed RPOsec, a monotonically decreasing normalization calculation is performed to obtain the data loss tolerance factor P, such that the smaller RPOsec is, the larger the data loss tolerance factor P is. Obtain the first basic weight coefficient Wb corresponding to the business criticality coefficient B, the second basic weight coefficient Wt corresponding to the time urgency factor T, the third basic weight coefficient Wp corresponding to the data loss tolerance factor P, and the fourth basic weight coefficient Wv corresponding to the environmental risk probability score V; normalize and verify the first basic weight coefficient Wb, the second basic weight coefficient Wt, the third basic weight coefficient Wp, and the fourth basic weight coefficient Wv to ensure that Wb + Wt + Wp + Wv = 1; calculate the recovery priority index RPI according to RPI = B × Wb + T × Wt + P × Wp + V × C × Wv, and generate a disaster recovery execution queue accordingly, wherein V × C × Wv is used to positively weight the recovery priority index RPI; compare the recovery priority index RPI with a preset index threshold, and if the recovery priority index RPI is greater than the preset index threshold, mark the corresponding data asset as a first-level priority recovery asset. In response to external state change events, the basic weight coefficients are updated and weight normalization is performed. The recovery priority index (RPI) is recalculated and the disaster recovery execution queue is rearranged. According to the rearranged disaster recovery execution queue, control commands are sent to the network controller and virtualization resource pool to perform resource preemption, locking, and disaster recovery data recovery.
2. The method as described in claim 1, characterized in that, The process of mapping it to a unified public field and generating a data asset catalog with a unique identifier includes: The metadata is collected from at least one data source, including database servers, file servers, cloud storage services, and local disks. The table name, field name, number of table records, and update time in the structured data, and the file name, storage path, number of file bytes, and modification time in the unstructured data are respectively mapped to the asset name, asset location, data volume, data format, and timestamp fields in the data asset catalog; Generate data feature labels for field relationships in structured data, and generate extended feature values for storage paths, file extensions, and storage media attributes in unstructured data; The data feature tags and extended feature values are written into the extended metadata field of the data asset catalog, and a unique asset identifier is configured for each data asset.
3. The method as described in claim 1, characterized in that, The process of classifying data assets in the data asset catalog according to a preset multi-branch decision logic strategy and generating classification result tags includes: Extract the content keywords of the data assets. If the content keywords match the preset set of highly sensitive keywords, directly output the classification result tags containing high risk level and core asset markers. If the preset set of highly sensitive keywords is not matched, the file type is determined based on the file extension of the data asset, and the volume level is determined based on the data volume parameter of the data asset. A combined condition determination is performed based on the content keywords, file type, and size level; Specifically, if the file type is code or configuration file, the output will include a classification result label containing the core asset marker; if the file type is a structured data file and the size level reaches the preset large size level, or hits the preset set of sensitive keywords, the output will include a classification result label containing financial data, operational data, or sensitive data markers; if no preset branch is matched, the output will be a classification result label pending manual review.
4. The method as described in claim 1, characterized in that, The process of updating the basic weight coefficients in response to external state change events and performing weight normalization, recalculating the recovery priority index (RPI), and rearranging the disaster recovery execution queue includes: External state change events are obtained through the event capture interface. These external state change events include at least one of the following: security events, vulnerability discovery events, external rule change events, and protection device configuration change events. The external state change event is converted into a formatted structured message by the event converter. The formatted structured message includes at least the event type, event source, event time, event payload, and severity of impact. The formatted structure message is routed to the corresponding weight adjustment process according to the event type, and the weight adjustment range is determined according to the severity of the impact. Update at least one basic weight coefficient according to the weight adjustment range, and perform weight normalization processing after the update; A weight update signal is sent to the recovery priority index calculation module to trigger the recalculation of the recovery priority index RPI, and the disaster recovery execution queue is rearranged according to the recalculated recovery priority index RPI.
5. The method as described in claim 1, characterized in that, The step of sending control commands to the network controller and virtualization resource pool according to the rearranged disaster recovery execution queue to perform resource preemption, locking, and disaster recovery data recovery includes: For the first-priority recovery assets, a bandwidth expansion command and a bandwidth preemption command are sent to the software-defined network controller to pre-lock a preset proportion of network bandwidth resources; Send computing resource scheduling instructions and storage resource scheduling instructions to the virtualization resource pool to allocate a specified number of CPU cores, memory resources and solid-state storage resources to the first-priority recovery assets, and perform parallel recovery; For data assets not identified as first-priority recovery assets, allocate the remaining shared bandwidth and default resource quotas, and perform serial recovery; After the disaster recovery data is restored, a consistency check is performed on the restored data, and the pre-locked network bandwidth resources, computing resources and storage resources are released after the check passes.
6. A data asset disaster recovery system, characterized in that, The system includes: The heterogeneous data alignment module is used to collect metadata of structured and unstructured data in the enterprise IT environment, map it to a unified common field, and generate a data asset catalog with a unique identifier. The multi-branch rule classification module is used to classify data assets in the data asset catalog according to a preset multi-branch decision logic strategy and generate classification result labels; The indicator quantification module is used to read the business impact analysis level associated with the data asset from the configuration management database or business continuity management system, and convert it into a business criticality coefficient B through a preset business level mapping table. When the business impact analysis level is missing, a preset median business criticality coefficient is assigned and a missing alarm is generated. It reads the data sensitivity identifier or industry compliance attribute identifier from the classification result labels and converts it into a normative constraint coefficient C through a preset constraint mapping table. It calculates the environmental risk probability score V based on at least two of the physical risk probability, technical risk probability, and human risk probability. It converts the data asset's recovery time target into a second-level value RTOsec, performs missing value replacement, maximum value truncation, and minimum emergency threshold protection on the RTOsec, and performs monotonically decreasing normalization calculation based on the boundary-processed RTOsec to obtain the time urgency factor T. It converts the data asset's recovery point target into a second-level value RPOsec, performs missing value replacement, zero value processing, and maximum tolerance threshold truncation on the RPOsec, and performs monotonically decreasing normalization calculation based on the boundary-processed RPOsec to obtain the data loss tolerance factor P. The priority calculation module is used to obtain the first basic weight coefficient Wb, the second basic weight coefficient Wt, the third basic weight coefficient Wp, and the fourth basic weight coefficient Wv corresponding to the business criticality coefficient B, the time urgency factor T, the data loss tolerance factor P, and the environmental risk probability score V, respectively. It performs normalization verification on each basic weight coefficient to make Wb+Wt+Wp+Wv=1, and calculates the recovery priority index RPI according to RPI=B×Wb+T×Wt+P×Wp+V×C×Wv. Based on this, it generates a disaster recovery execution queue, and compares the recovery priority index RPI with a preset index threshold. Data assets with an RPI greater than the preset index threshold are identified as first-level priority recovery assets. The dynamic feedback module is used to update at least one basic weight coefficient in response to external state change events and perform weight normalization processing, recalculate the recovery priority index RPI and rearrange the disaster recovery execution queue. The disaster recovery automated scheduling module is used to send control commands to the network controller and virtualization resource pool according to the rearranged disaster recovery execution queue, and to perform resource preemption, locking and disaster recovery data recovery.
Citation Information
Patent Citations
Intelligent data protection method and system
CN118819964A
Task allocation system for disaster recovery and use method thereof
CN119088540A