Intelligent data backup method and system based on AI large model

Through an intelligent data backup method based on AI large model, a holographic data semantic perception model is built to perform real-time risk detection and self-trigger preventive backup decisions, solving the dynamic adjustment and real-time fault detection of data backup in the existing technology, and achieving efficient data backup and recovery capabilities.

CN120560907AActive Publication Date: 2025-08-29ANHUI FEIWEI INFORMATION TECHNOLOGY CO LTD +1

Patent Information

Application Number
CN202511062998.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-08-29
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Existing data backup technologies cannot dynamically adjust according to the importance of data, access frequency or business context, resulting in redundant backups or missing critical data, and lack real-time fault detection and dynamic strategy optimization, making it difficult to achieve efficient scheduling and resource management in large-scale concurrent data writing and high-frequency access scenarios, affecting backup efficiency and recovery reliability.

Method used

Using an intelligent data backup method based on AI large model, we use the holographic data semantic perception model to conduct real-time transient risk detection, self-trigger preventive backup decisions, and combine dynamic scheduling of multi-storage cloud environment resources and intelligent incremental backup decisions to realize adaptive compression coding and multi-task parallel execution, and build an intelligent incremental backup execution engine.

Benefits of technology

It realizes high-dimensional semantic consistency recognition and cross-system consistency of data, has real-time and context-aware capabilities, improves the accuracy and efficiency of backup behavior, optimizes resource utilization and SLA completion rate of backup tasks, and ensures the reliability and traceability of backup results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120560907A_ABST
    Figure CN120560907A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data backup, in particular to an intelligent data backup method and system based on an AI large model. The method comprises the following steps: acquiring an enterprise global data list, performing intelligent data structure deconstruction and dynamic attribute mapping modeling, and constructing a holographic data semantic perception model; performing real-time transient risk mutation detection on the holographic data semantic perception model, and constructing an intelligent backup triggering mechanism; carrying out storage resource demand prediction based on an intelligent backup trigger mechanism, carrying out multi-storage cloud environment resource dynamic scheduling, and constructing an elastic backup storage resource pool; carrying out incremental backup analysis and self-adaptive compression coding to obtain an incremental backup coding packet; and performing dynamic backup sequence adjustment and intelligent incremental backup decision on the incremental backup coding packet based on the elastic backup storage resource pool, and constructing an intelligent incremental backup execution engine. According to the method, the reliability, the accuracy and the traceability of a backup result are improved through self-adaptive intelligent incremental backup.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data backup, and in particular to an intelligent data backup method and system based on an AI big model. Background Art

[0002] Existing data backup technologies primarily rely on periodic task scheduling and incremental difference comparison mechanisms to implement backup operations. These core methods include full backups, incremental backups, and snapshots. While these methods can achieve basic data protection in specific environments, they suffer from several significant limitations: On the one hand, traditional backup solutions cannot dynamically adjust based on data importance, access frequency, or business context, resulting in redundant backups or omissions of critical data. On the other hand, faced with challenges such as heterogeneous data sources, multi-system interactions, and limited backup windows, traditional solutions struggle to achieve efficient scheduling and resource elasticity, easily leading to wasted backup resources or task failures. Furthermore, structural changes, semantic evolution, and cross-source dependencies of data during the backup process are not perceived by existing systems, weakening the intelligence and business-awareness of backups.

[0003] Furthermore, existing backup systems often lack real-time fault detection and dynamic policy optimization mechanisms, relying heavily on manual configuration and passive responses. When anomalies occur, such as backup task interruptions, data consistency check failures, or resource bottlenecks, they are often difficult to detect and address in a timely manner, impacting overall backup efficiency and recovery reliability. This is especially true in scenarios with large-scale concurrent data writes, high-frequency access, and complex business dependencies. Traditional solutions' static policies and single scheduling models are no longer sufficient, necessitating the introduction of intelligent mechanisms with self-learning and adaptive capabilities. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention proposes an intelligent data backup method and system based on an AI big model to solve at least one of the above technical problems.

[0005] To achieve the above objectives, the present invention provides an intelligent data backup method based on an AI large model, comprising the following steps: Obtain an enterprise-wide data inventory, perform intelligent data structure deconstruction and dynamic attribute mapping modeling, and build a holographic data semantic perception model; Conduct real-time transient risk mutation detection on the holographic data semantic perception model, make self-triggered preventive backup decisions, and build an intelligent backup trigger mechanism; Predict storage resource demand based on an intelligent backup trigger mechanism, dynamically schedule resources in multiple storage cloud environments, and build a flexible backup storage resource pool. Obtain the previous version of the system backup data, perform incremental backup analysis and adaptive compression encoding, and obtain an incremental backup encoding package; Based on the elastic backup storage resource pool, the incremental backup code package is dynamically adjusted in backup order and intelligent incremental backup decision is made to build an intelligent incremental backup execution engine; Based on the intelligent incremental backup execution engine, multi-task distributed parallel backup execution is performed, and iterative verification and optimization are carried out to build an iterative optimization backup data model.

[0006] In this specification, an intelligent data backup system based on an AI big model is provided, which is used to execute the intelligent data backup method based on the AI ​​big model as described above, including: The semantic perception module is used to obtain the enterprise's global data list, perform intelligent data structure deconstruction and dynamic attribute mapping modeling, and build a holographic data semantic perception model; The self-triggered backup module is used to perform real-time transient risk mutation detection on the holographic data semantic perception model, make self-triggered preventive backup decisions, and build an intelligent backup trigger mechanism; The storage demand prediction module is used to predict storage resource demand based on the intelligent backup trigger mechanism, dynamically schedule resources in multiple storage cloud environments, and build an elastic backup storage resource pool; The backup encoding module is used to obtain the system backup data of the previous version, perform incremental backup analysis and adaptive compression encoding, and obtain an incremental backup encoding package; The backup sequence adjustment module is used to dynamically adjust the backup sequence of incremental backup code packages and make intelligent incremental backup decisions based on the elastic backup storage resource pool, thus building an intelligent incremental backup execution engine. The distributed parallel backup module is used to perform multi-task distributed parallel backup execution based on the intelligent incremental backup execution engine, and perform iterative verification optimization to build an iterative optimization backup data model.

[0007] The beneficial effects of the present invention include: unified metadata extraction and structural deconstruction of an enterprise's multi-source, heterogeneous data assets (structured, semi-structured, and unstructured), enabling high-dimensional representation and modeling of semantic dependencies between fields. Based on the Transformer architecture, the AI ​​large model vectorizes attributes, constructs a semantic graph between data, and achieves cross-system data semantic consistency. This semantic-aware model supports dynamic labeling of features such as data sensitivity, access frequency, and data lineage, providing a semantically driven, contextual decision-making foundation for subsequent intelligent backup strategies. Time series anomaly detection algorithms (such as the LSTM+Attention mechanism or a GRU-based change point detection model) are used to model data state evolution in real time, detecting sudden data access fluctuations, write anomalies, or structural changes. Backup events are triggered based on anomaly scores and risk thresholds, making backup actions immediate and context-aware. The system's embedded self-triggering decision logic dynamically adjusts thresholds based on data sensitivity labels and mutation factors (such as tampering probability and access surge), achieving a highly accurate preventative backup initiation mechanism.

[0008] Deep time series forecasting models (such as Prophet, DeepAR, or Transformer-based forecasting) are used to accurately model future storage capacity requirements, supporting resource usage predictions down to the hourly level. Integrating with cloud platform APIs, the system automatically mounts and migrates cross-cloud storage resources (such as object storage, block storage, and cold data archives) through containerized scheduling or serverless logic functions. An intelligent resource scheduling engine enables reallocation of resources in real time based on load forecasts and SLA constraints, building a distributed, elastic backup resource pool with scalability. Block-level hash comparisons and change snapshot analysis (such as those based on SHA-256 or content-defined block partitioning) enable rapid identification and location of differing data blocks. The system dynamically selects the optimal compression encoding strategy (such as Zstandard, LZ4, or Huffman+RLE hybrid encoding) based on content characteristics (text, log, or multimedia), improving compression ratios and reducing encoding latency. The compressed incremental encoding package includes version identification and integrity verification mechanisms and can be embedded in a self-describing structure for rapid decompression, verification, and recovery.

[0009] The system dynamically optimizes the execution order of backup tasks using priority queues and reinforcement learning scheduling strategies (such as those based on the DDPG or PPO algorithms), taking into account multiple factors such as storage I / O load, backup latency constraints, and data value density. The intelligent decision-making engine integrates real-time resource pool status with historical performance feedback to implement policy matching based on resource constraints and risk levels, improving resource utilization and backup task SLA achievement. The task scheduling engine supports distributed locking and replication mechanisms to ensure data consistency and high availability during concurrent execution of multiple tasks. Backup tasks are executed in parallel across distributed computing nodes based on a Directed Action Graph (DAG), and a hash-based task sharding mechanism automatically divides backup granularity. The system integrates an iterative optimization module that performs offline training based on historical task logs and backup results to optimize model parameters (such as compression rate prediction accuracy and recovery time estimation), continuously improving the convergence speed and generalization of backup decisions. Automated regression testing and verification scripts provide continuous consistency verification of backup data (such as CRC verification and sample restoration comparison), ensuring the reliability, accuracy, and traceability of backup results. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a flowchart of the steps of an intelligent data backup method based on an AI large model of the present invention; Figure 2 Detailed implementation flow chart of step S1; Figure 3 Detailed implementation flow chart of step S2; Figure 4 Schematic diagram of the detailed implementation steps of step S3. DETAILED DESCRIPTION

[0011] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0012] This application provides an AI-based big model intelligent data backup method and system. The execution entities of the AI-based big model intelligent data backup method and system include, but are not limited to, the following: mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc. equipped with the system, which can be regarded as general computing nodes of this application. The data processing platform includes, but is not limited to, at least one of: an audio and image management system, an information management system, and a cloud data management system.

[0013] See also Figures 1 to 4 The present invention provides an intelligent data backup method based on an AI big model, comprising the following steps: Obtain an enterprise-wide data inventory, perform intelligent data structure deconstruction and dynamic attribute mapping modeling, and build a holographic data semantic perception model; Conduct real-time transient risk mutation detection on the holographic data semantic perception model, make self-triggered preventive backup decisions, and build an intelligent backup trigger mechanism; Predict storage resource demand based on an intelligent backup trigger mechanism, dynamically schedule resources in multiple storage cloud environments, and build a flexible backup storage resource pool. Obtain the previous version of the system backup data, perform incremental backup analysis and adaptive compression encoding, and obtain an incremental backup encoding package; Based on the elastic backup storage resource pool, the incremental backup code package is dynamically adjusted in backup order and intelligent incremental backup decision is made to build an intelligent incremental backup execution engine; Based on the intelligent incremental backup execution engine, multi-task distributed parallel backup execution is performed, and iterative verification and optimization are carried out to build an iterative optimization backup data model.

[0014] In the embodiment of the present invention, see Figure 1 , is a flowchart of the steps of an intelligent data backup method based on an AI large model of the present invention. In this example, the steps of the intelligent data backup method based on the AI ​​large model include: Obtain an enterprise-wide data inventory, perform intelligent data structure deconstruction and dynamic attribute mapping modeling, and build a holographic data semantic perception model; In this embodiment, a distributed data discovery crawler system is deployed to perform a full-domain scan of an enterprise's internal structured databases, unstructured document repositories, multimedia resource repositories, real-time data streams, and other related information. This system creates an enterprise data asset inventory containing basic information such as data source type, storage location, access rights, and data volume. Based on this acquired data asset inventory, the multimodal understanding capabilities of a large language model are utilized to perform deep semantic analysis of each data entity, extracting multidimensional feature vectors such as the data's business semantics, structural features, and content themes. Experimental parameters are set to a semantic vector dimension of 1024 and a semantic similarity threshold of 0.85. Subsequently, intelligent data structure deconstruction is performed, and graph neural network algorithms are used to model the relationships between data. These relationships, such as data dependency links, impact propagation paths, and business correlation strength, are identified, and an enterprise data association graph is constructed. The number of nodes in the graph typically reaches 10^6, and the number of edges reaches 10^7. Based on this constructed data association graph, dynamic attribute mapping modeling is performed. A time series analysis algorithm is used to learn the changing patterns of data attributes and establish a dynamic mapping model for the time-dependent evolution of data attributes. The time window is set to 7 days, with an hourly sampling frequency. Ultimately, the semantic feature vectors, association graphs, and dynamic mapping models are integrated to construct a holographic data semantic perception model that covers data semantic understanding, structural relationships, and evolutionary trends. This model can achieve all-round intelligent perception and understanding of enterprise data.

[0015] Conduct real-time transient risk mutation detection on the holographic data semantic perception model, make self-triggered preventive backup decisions, and build an intelligent backup trigger mechanism; In this embodiment, after establishing a holographic data semantic perception model, real-time risk monitoring and the development of a self-triggered backup strategy are implemented. By monitoring the data flow within the model, mutation behavior between data can be identified. Real-time analysis of data change rates, dynamic changes in data clusters, and potential data emergencies can effectively detect potential risks in the system. These risks can include hardware failures, system crashes, or data loss. A deep learning model based on time series data analysis (such as an LSTM neural network) is used for risk prediction. By combining historical data with real-time data streams, it utilizes methods such as volatility analysis and trend line identification to capture risk fluctuations. Experimental results show that the LSTM network-based risk mutation detection algorithm achieves an accuracy rate exceeding 90%. This model automatically triggers the backup mechanism upon identifying potential risks and makes intelligent decisions about backup timing based on reasoning from historical backup data. This automated backup reduces human intervention and effectively improves the system's disaster recovery capabilities.

[0016] Predict storage resource demand based on an intelligent backup trigger mechanism, dynamically schedule resources in multiple storage cloud environments, and build a flexible backup storage resource pool. In this embodiment, a storage resource demand prediction algorithm is activated based on the backup decision signal output by the intelligent backup trigger mechanism. This algorithm comprehensively considers factors such as backup data volume, backup frequency, storage duration, and redundancy requirements to calculate resource requirements. A time series prediction model, LSTM, is used to predict backup data volume for the next seven days. Model inputs include historical backup data volume, business activity, and seasonal factors, achieving a prediction accuracy of over 95%. Based on the prediction results, specific requirements such as required storage capacity, network bandwidth, and computing resources are calculated. Storage capacity requirements are typically 3-5 times the actual data volume to account for redundancy and compression. Network bandwidth requirements are planned at 120% of peak throughput. Dynamic resource scheduling is then performed across multiple storage cloud environments, establishing a hybrid storage architecture comprising local storage, public cloud storage, and private cloud storage. A cost-benefit analysis algorithm is used to select the optimal storage resource combination, with local storage responsible for 70% of hot data backup, the public cloud for 20% of warm data backup, and the private cloud for 10% of cold data backup. Implement a dynamic resource scheduling strategy to automatically adjust resource allocation across storage nodes based on real-time load. The scheduling algorithm uses a genetic algorithm for multi-objective optimization, targeting minimizing cost, maximizing performance, and optimizing reliability. Build a flexible backup storage resource pool with automatic capacity expansion and contraction capabilities. It automatically expands when storage utilization exceeds 80% and automatically contracts when it falls below 30%. Response times for expansion and contraction are controlled within 5 minutes.

[0017] Obtain the previous version of the system backup data, perform incremental backup analysis and adaptive compression encoding, and obtain an incremental backup encoding package; In this embodiment, the backup data of the previous version is analyzed to determine which data has changed and needs to be backed up. Through incremental backup technology, only the changed data is backed up, thereby significantly reducing the amount of backup data and backup time. The core of incremental backup analysis is to identify which data has changed by comparing the differences between the current backup and the previous version. An adaptive compression coding algorithm is used. The changed data is identified through a differentiation algorithm, and the incremental data is compressed using an efficient compression coding algorithm (such as LZ77, Huffman coding, etc.). After using adaptive compression coding, the compression ratio of the backup data is improved by about 30%, reducing the storage space requirement. The compressed incremental data is packaged into an incremental backup coding package to ensure the efficiency of the backup process and save storage space.

[0018] Based on the elastic backup storage resource pool, the incremental backup code package is dynamically adjusted in backup order and intelligent incremental backup decision is made to build an intelligent incremental backup execution engine; In this embodiment, a dynamic backup order adjustment algorithm is initiated based on the constructed elastic backup storage resource pool and the generated incremental backup code packets. This algorithm comprehensively considers factors such as data priority, dependencies, resource availability, and network conditions, and uses a priority-based scheduling algorithm to sort backup tasks. The priority calculation formula incorporates factors such as data importance weight (40%), risk urgency (30%), dependency complexity (20%), and resource availability (10%). For data with dependencies, a topological sorting algorithm is used to ensure that dependent data is backed up first, avoiding data consistency issues during the backup process. During the backup execution process, an intelligent incremental backup decision-making mechanism is implemented. This mechanism dynamically adjusts the backup strategy based on real-time factors such as storage resource usage, network bandwidth conditions, and business impact. It automatically reduces backup concurrency when network bandwidth utilization exceeds 70%, and automatically initiates data compression optimization when storage resource utilization exceeds 85%. An intelligent incremental backup execution engine is constructed, which utilizes a multi-threaded parallel processing architecture and supports up to 128 concurrent backup tasks. The processing capacity of a single task reaches 100MB / s, and the overall backup throughput reaches 10GB / s. The execution engine has a built-in load balancing mechanism that uses a round-robin algorithm to distribute backup tasks to different storage nodes, ensuring balanced load across all nodes and keeping load deviation within 5%. A backup task monitoring mechanism is also established to monitor task execution status, progress, and error messages in real time. Tasks that fail are automatically retried, with a maximum of three retries. The retry interval uses an exponential backoff algorithm, with an initial interval of 30 seconds.

[0019] Based on the intelligent incremental backup execution engine, multi-task distributed parallel backup execution is performed, and iterative verification and optimization are carried out to build an iterative optimization backup data model.

[0020] In this embodiment, a multi-task distributed parallel backup execution mechanism is implemented based on a constructed intelligent incremental backup execution engine. This mechanism utilizes a master-slave architecture, with the master node responsible for task scheduling and coordination, and the slave nodes responsible for specific backup task execution. Distributed deployments of up to 256 slave nodes are supported, with a single slave node capable of processing up to 500MB / s. During task allocation, an intelligent scheduling algorithm based on load and geographic location is employed to prioritize tasks to nodes with lower loads and minimal network latency. The scheduling algorithm has a time complexity of O(n log n), and scheduling decision time is controlled within 100 milliseconds. A dynamic load balancing mechanism is implemented, automatically migrating tasks to other nodes when a node's load exceeds 90%. This migration process utilizes hot migration technology to ensure uninterrupted backup tasks. Real-time performance monitoring is performed during the backup execution process, collecting key metrics such as backup speed, resource utilization, error rate, and network latency. Monitoring data is sampled once per second and stored in a time series database for subsequent analysis. Iterative validation and optimization are performed based on monitoring data, using a machine learning algorithm to predict and optimize backup performance. The training data consists of nearly three months of backup execution history, encompassing over 10^6 backup task records. The optimization algorithm uses the Q-learning algorithm from reinforcement learning to improve backup policy parameters through trial and error. The learning rate is set to 0.01 and the exploration rate is set to 0.1. After 1,000 rounds of iterative training, backup efficiency increased by over 25%. An iteratively optimized backup data model was constructed, which includes multiple sub-models such as backup performance prediction, resource demand prediction, and fault prediction. The model accuracy exceeds 92%, providing a scientific basis for future backup optimization.

[0021] In this embodiment, refer to Figure 2 , which is a flowchart of the detailed implementation steps for dynamically mapping the enterprise data-related knowledge graph to build a holographic data semantic perception model. In this embodiment, the detailed implementation steps for dynamically mapping the enterprise data-related knowledge graph to build a holographic data semantic perception model include: Obtain an enterprise-wide data list; identify and filter abnormal redundant data in the enterprise-wide data list to generate a redundant filtered data list; Intelligent data structure deconstruction is performed on the redundant filtered data list to generate multi-dimensional enterprise data structure characteristics, wherein the multi-dimensional enterprise data structure characteristics include data structure body, business attributes, data types and data element hierarchical characteristics; Perform deep semantic analysis on the redundant filtered data list to extract multimodal semantic features; Based on the multi-dimensional enterprise data structure characteristics and multi-modal semantic features, we mine the correlation between cross-data sources, identify the business dependency links and impact propagation paths between data, and build an enterprise data association knowledge graph; Perform dynamic attribute mapping modeling on enterprise data-related knowledge graphs and build a holographic data semantic perception model.

[0022] In this embodiment, all enterprise data is uniformly collected and organized to form a complete data asset catalog. Typically, enterprise data is distributed across multiple heterogeneous data sources, including relational databases (such as Oracle and MySQL), non-relational databases (such as MongoDB), distributed storage systems (such as HDFS), logging systems (such as ELK), business middleware APIs, and data lakes. Therefore, standardized access is first implemented through the data source access module. Data connectors or collection agents (such as Sqoop and Kafka Connect) are used to uniformly capture metadata and partial sampled data, and a data source mapping table is established. During the collection process, large AI models can be introduced to assist in data identification and classification, initially classifying the captured data into structured, semi-structured, and unstructured types. Simultaneously, data cleansing and preliminary transformation are performed in conjunction with the ETL (Extract-Transform-Load) process. To ensure the completeness and continuous updateability of the inventory, a scheduled task strategy is established, using differential update technology to dynamically add incremental data. During the experiment, access to five data sources can be configured, with a single source data sample size of no less than 10GB. Model training is used to assess data source coverage and quality integrity, laying a solid data foundation for subsequent intelligent backup strategies. Redundant, anomaly, and duplicate data within the original data list are intelligently identified and cleaned to ensure the effectiveness of subsequent analysis and modeling. First, clustering algorithms (such as K-Means and DBSCAN) are used to analyze data distribution density and identify data segments with high similarity in structure and content. Semantic similarity calculations (such as Cosine Similarity and Siamese Network) are then performed on data field content using large AI models (such as BERT and GPT) to determine whether redundant data exists, where the meaning of the fields is repeated but the names are different. Furthermore, statistical analysis of the data frequency distribution is performed, using histograms to identify abnormal peaks and outliers. The possibility of redundancy is comprehensively assessed based on data dimensions, timestamps, and business tags. During this process, we set an anomaly threshold of α = 0.85 and a redundancy matching accuracy of β = 0.9. In an experimental scenario, we cleaned a data list of 100,000 rows, identifying approximately 27% of redundant fields and eliminating 18% of outliers. The resulting "redundancy-filtered data list" serves as a high-quality data input source, significantly reducing subsequent computing resource consumption and training errors.

[0023] After redundancy filtering, the data content undergoes in-depth structural analysis to extract multidimensional structural features, providing logical support for data association and modeling. First, a structural analysis engine analyzes the physical and logical structure of fields, identifying information such as table structure, primary and foreign key constraints, field naming conventions, nesting hierarchies, and index relationships. An AI-powered model performs semantic analysis of field names and their context, combining named entity recognition (NER) with contextual clustering algorithms to generate "data structure entities" and "business attribute labels" (such as "user ID," "transaction amount," and "time dimension"). Next, a type inference model is used to categorize field types (such as integer, floating point, date, and enumeration) and extract the hierarchical relationships of data elements within the data warehouse (e.g., ODS → DWD → ADS). Using 500 tables and 10,000 fields as experimental subjects, the average accuracy rate for correct field name semantic recognition reached 93%, and the accuracy rate for field hierarchy recognition reached 91%. Ultimately, a multi-dimensional set of enterprise data structure features, encompassing field structure, type, hierarchy, and business meaning, is generated, laying the foundation for the construction of a knowledge graph. Structured data fields and data values ​​are combined to construct a context window. Pre-trained language models (such as RoBERTa and T5) are used to embed field names and sample data, extracting word- and sentence-level semantic features. For semi-structured data such as JSON / XML, tree-structured parsing methods are used to extract hierarchical semantic labels. For unstructured data such as logs and reports, text content is recognized through optical character recognition (OCR) and integrated into a large multimodal model (such as CLIP and GPT-V) for joint modeling of images, text, tables, and text, extracting corresponding concepts, actions, temporal sequences, and causal relationship features. During this process, a multi-head attention mechanism can be introduced to enhance semantic dependency modeling between entities, and a top-K strategy is used to select high-confidence semantic feature vectors. In experiments, using 20 business data samples, the average F1 score for semantic extraction reached 0.89, and the accuracy of extracting business intent from unstructured data was approximately 84%. Ultimately, a semantic embedding representation matrix tailored to the business context is generated, providing critical support for subsequent data graph construction and semantic perception.

[0024] Graph neural networks (such as GCN and GAT) are used to model the potential dependencies between different data entities. Edge prediction mechanisms are used to identify relationships between fields, such as master-slave, reference, and influence. This leads to the construction of a knowledge graph with nodes representing "data fields / tables" and edges representing "associations / dependencies / constraints." Furthermore, AI-powered models are used to determine semantic similarity and contextual relevance between fields, enabling entity alignment and mapping of data objects with different names but consistent semantics. Data propagation paths can also be extracted based on temporal features (such as time field linkage) and event-driven paths (e.g., data changes caused by certain operations), identifying key data influencing sources. In an experiment, graph fusion was performed based on data from a financial system and a CRM system, successfully constructing a knowledge graph with approximately 400,000 nodes and 1.2 million edges. The average node centrality increased to 2.3, achieving 92% coverage. The resulting graph serves as the core infrastructure for intelligent data backup strategy deduction, supporting dynamic dependency tracking and prioritized backup of critical data. This upgrade of static knowledge graphs into dynamic, semantically aware models enables real-time understanding and prediction of enterprise data status, changing trends, and business context. First, a temporal GNN model based on the knowledge graph model models the temporal changes in entity attributes, monitoring attribute dynamics such as field value change frequency, update cycle, and dependency strength. Second, a multi-task learning framework is introduced to integrate multiple sources of features, including structural, semantic, and behavioral, to construct a semantically aware vector space, enabling the system to "understand" the business meaning of the data. On this basis, a large AI model is used to model the mapping relationship between "semantic label → data behavior → backup strategy." For example, for data items identified as "key decision fields," the awareness model can automatically determine whether to perform high-frequency incremental backup or multi-site disaster recovery backup. In experiments, the semantic awareness model, trained with the Transformer framework, achieved an accuracy of 91% in data label prediction tasks and an 87% top-3 hit rate in recommending backup strategies for key fields. Ultimately, the semantic awareness model serves as the "cognitive core" of the intelligent data backup system, enabling holistic intelligent backup strategy scheduling based on business semantics, data influence, and dynamic changes.

[0025] In this embodiment, the specific steps of performing dynamic attribute mapping modeling on the enterprise data-related knowledge graph and constructing a holographic data semantic perception model are as follows: Identify multiple data entity nodes based on the enterprise data association knowledge graph; Calculate the centrality index of the data entity nodes one by one to obtain the graph criticality of each data entity; Perform data access frequency statistics on multiple data entity nodes to obtain data access frequency; quantifying the business importance values ​​of the plurality of data entity nodes; Calculate the comprehensive value weight based on the graph criticality, data access frequency and business importance, and perform intelligent label fitting to generate a priority label for each data entity; Based on multimodal semantic features, personalized semantic feature mining is performed on multiple data entity nodes to obtain multiple data entity semantic fingerprints; Generate a unique semantic identifier based on the data entity semantic fingerprint; Dynamic attribute mapping modeling is performed on the enterprise data association knowledge graph based on the unique semantic identifier and the priority label to build a holographic data semantic perception model.

[0026] In this example, multiple "data entity nodes" representing business-critical content are systematically identified from the enterprise data-related knowledge graph. These nodes typically include database fields, data tables, log entries, interface parameters, and other entities. These nodes are labeled using the graph's business naming, relationship type, and contextual linking. Leveraging the knowledge graph's entity categories (such as "customer information," "order details," and "payment voucher") and edge semantic types (such as "dependency," "reference," and "influence propagation"), graph traversal and filtering algorithms (such as BFS plus attribute filtering) are used to extract high-value data entities that meet the criteria. At this stage, a large model assists in understanding the entity's business context, for example, using BERT embedding to determine that "cust_id" and "client_no" are the same conceptual entity. In an experimental scenario, 12,370 valid data entity nodes were screened based on business criticality labels from a knowledge graph with 500,000 nodes and 1.5 million edges, initially establishing an entity pool for backup decision-making. This step provides the foundational entity units for subsequent multi-dimensional evaluation and priority modeling. Multiple "data entity nodes" representing business-critical content are systematically identified from the enterprise data-linked knowledge graph. These nodes typically include database fields, data tables, log entries, and interface parameters, and are labeled using the graph's business naming, relationship types, and contextual linking. Leveraging the knowledge graph's entity categories (such as "customer information," "order details," and "payment voucher") and edge semantic types (such as "dependency," "reference," and "influence propagation"), graph traversal and filtering algorithms (such as BFS plus attribute filtering) are used to extract high-value data entities that meet the criteria. At this stage, a large model assists in understanding the business context of the entities, for example, using BERT embedding to determine that "cust_id" and "client_no" are the same conceptual entity. In an experimental scenario, 12,370 valid data entity nodes were screened based on business criticality labels from a knowledge graph with 500,000 nodes and 1.5 million edges, initially establishing an entity pool for backup decision-making. This step provides the foundational entity units for subsequent multi-dimensional evaluation and priority modeling.

[0027] Measuring the access intensity of data entities in actual business operations helps determine their activity and usage value. This approach involves extracting entity access events from database audit logs, API call logs, ETL scheduling logs, and user operation records. For each data entity node, key metrics such as access frequency, peak access time, and concurrent call volume within a fixed period are counted. During this phase, the AI ​​model uses behavioral sequence modeling (e.g., the Transformer model models access time series) to detect anomalies and identify peaks. To improve accuracy, access frequency is weighted based on user roles and business scenario labels. For example, nodes accessed by the CEO are given higher priority. In an experimental scenario, using a company's CRM system as an example, nearly 3 million access records were collected over a 7-day period. After clustering, it was found that the top 10% of data entity nodes accounted for over 60% of access frequency. High-frequency nodes, such as "customer credit score" and "order status," are highly weighted in the intelligent backup prioritization model. Ultimately, an "access frequency distribution table" is generated, providing the behavioral characteristics underlying the entity weighting scoring system. Using business rules, institutional documents, and historical decision-making behaviors as input, an AI big model is introduced for semantic parsing and business association modeling. First, natural language processing techniques are used to extract entities and construct dependency graphs from process documents, audit reports, and policy specifications, identifying "critical business data" labels. Next, the big model's reasoning capabilities (such as GPT-4's chain reasoning) are leveraged to determine the impact of a particular entity across different business chains. For example, the "settlement amount" field impacts not only financial statements but also tax filing and reconciliation processes, significantly outweighing auxiliary fields such as "promotional activity identifier." During the experiment, semantic modeling was performed on 17 types of business process documents from a large manufacturing enterprise, leading to the extraction of 184 high-priority business data entities. Furthermore, combined with user feedback and risk assessment reports, a set of "business importance scores" between 0 and 1 was quantified for subsequent value weight modeling. A multi-metric fusion mechanism was used to generate a comprehensive value score and automatically fit a "priority label." First, a normalized weight matrix was constructed, normalizing the three metrics to the [0, 1] range. The weighted combination relationship was learned using weighted averaging or a multi-layer perceptron (MLP). The weight distribution can refer to actual business feedback, such as giving a higher sovereign weight α=0.5, access frequency β=0.3, and graph centrality γ=0.2 to business-critical values. Clustering algorithms (such as K-Means or DBSCAN) are then used to hierarchically divide the comprehensive scores into three categories, generating labels such as "high priority," "medium priority," and "low priority." The AI ​​big model provides label semantic alignment and anomaly labeling correction functions to ensure that the labels are business interpretable. In one experiment, 2,000 core data nodes were modeled, and the label fitting accuracy reached 92.3%. Among them, the response time of "high priority" nodes in data backup and recovery tests increased by approximately 45%.The final output label is used to guide the backup strategy, such as setting high-priority nodes to real-time synchronization or multi-active backup.

[0028] A unique semantic representation, or "semantic fingerprint," is constructed for each data entity node to support fine-grained data identification and semantic backup. First, a multimodal large model (such as FLAN-T5 or GPT-4) is used to embed and semantically align data entities, combining structural information (field names, table structures), contextual descriptions (documentation, interface descriptions), and example values. Feature extraction includes, but is not limited to, naming semantics, business roles, temporal behavior, and access paths. A transformer model is then used to generate high-dimensional semantic vectors (typically 512-1024 dimensions). Principal component analysis (PCA) is then used to reduce the dimensionality and retain the principal features, ultimately forming a compact "semantic fingerprint." Experiments generated semantic fingerprints for 10,000 data entity nodes at a financial institution, demonstrating a top-1 accuracy of 88.6% in similarity recognition tasks. Semantic fingerprints are unique, stable, and portable, enabling unified entity identification across multiple systems and providing the foundational semantic genes for building a highly reliable data backup mapping mechanism. A unique semantic representation, or "semantic fingerprint," is constructed for each data entity node to support fine-grained data identification and semantic backup. First, a multimodal large-scale model (such as FLAN-T5 or GPT-4) is used to embed and semantically align data entities, combining structural information (field names, table structures), contextual descriptions (documentation, interface specifications), and example values. Feature extraction includes, but is not limited to, naming semantics, business roles, temporal behavior, and access paths. A transformer model is then used to generate high-dimensional semantic vectors (typically 512-1024 dimensions). Principal component analysis (PCA) is then used to reduce the dimensionality while retaining the principal feature components, ultimately forming a compact "semantic fingerprint." Experiments generated semantic fingerprints for 10,000 data entity nodes from a financial institution, demonstrating a top-1 accuracy of 88.6% in similarity recognition tasks. Semantic fingerprints are unique, stable, and portable, supporting unified entity recognition across multiple systems and providing the foundational semantic building blocks for building highly reliable data backup mapping mechanisms.

[0029] Combining "unique semantic identifiers" and "priority tags" to enhance the knowledge graph's attributes and dynamically model it, a self-aware holographic semantic model is constructed. This modeling approach utilizes dynamic graph technologies (such as Temporal Graph Network) to record the evolution of data entity states over time. Semantic attribute nodes and backup strategy nodes are introduced to form a multi-layered graph structure. Furthermore, semantic identifiers are used as entity primary keys, and tags are used as scheduling criteria. The model continuously optimizes the semantic weights and association paths of entity nodes. The AI ​​model is responsible for real-time updates to entity states, such as increased access frequency, tag adjustments, and the addition of new dependency chains, and triggers corresponding intelligent backup strategies (such as hot data migration and cold data archiving). In actual deployments, this model can be integrated with the backup engine, reducing average data recovery time (RTO) by 31% and increasing the backup strategy hit rate to 89%. This model achieves a complete semantic closed loop from structure to semantics to strategy, providing a future-oriented, perceptual capability system for enterprise intelligent backup.

[0030] In this embodiment, refer to Figure 3 , which is a flowchart of detailed implementation steps for performing real-time transient risk mutation detection on the holographic data semantic perception model, making self-triggered preventive backup decisions, and building an intelligent backup trigger mechanism. In this embodiment, the detailed implementation steps for performing real-time transient risk mutation detection on the holographic data semantic perception model, making self-triggered preventive backup decisions, and building an intelligent backup trigger mechanism include: Analyze the changes in time series data streams on the holographic data semantic perception model, extract historical data modification trajectories, access patterns, and changes in inter-data dependencies, and fit the multi-dimensional time series data change characteristics; Mining data evolution rules based on the changing characteristics of multi-dimensional time series data to generate data evolution prediction graphs; Calculate the probability of data corruption, business impact, and recovery complexity based on the data evolution prediction graph; Conduct data risk assessment based on data corruption probability, business impact scope, and recovery complexity, and build a data risk assessment matrix; Analyze the risk propagation path based on the data risk assessment matrix, mark key risk nodes and risk propagation nodes, and construct a data risk propagation map; Perform real-time transient risk mutation detection on the data risk propagation graph, make self-triggered preventive backup decisions, and build an intelligent backup trigger mechanism.

[0031] In this example, modification logs for each data entity (field, table, API, etc.) at different time points are collected, including insert, update, and delete operation records. Combined with data access audit logs, these logs capture call frequency, access user category, and access time window, forming an access pattern sequence. Secondly, changes in data dependencies are tracked, including the addition / deletion of field references, changes to data interfaces, and changes in semantic mappings between fields. These changes are extracted through knowledge graph snapshot diff comparisons. Based on this information, a multidimensional time series feature representation method is used to construct a data evolution tensor. Dimensions include data entity ID, time point, operation type, access intensity, dependency, and structural change magnitude. In experimental scenarios, this feature sequence is modeled using an LSTM (Long Short-Term Memory) model or a Transformer time series encoder. Multidimensional time series data change features are fitted for each data entity. These features serve as the basic data input for subsequent data evolution law mining. Through machine learning and sequence modeling, the evolutionary patterns and trends of data resources across the structural, call, dependency, and semantic dimensions are mined. First, using time window partitioning techniques (such as sliding window strategies with window lengths of 7 and 30 days), statistical analysis of data behavior over different periods is performed to extract indicators such as feature evolution rate, periodic fluctuation characteristics, and structural mutation markers. Clustering algorithms (such as DBSCAN and K-means) are then used to classify change patterns and identify common evolution types, such as "stable access + structural stability," "surge in access + field adjustment," and "dependency chain diffusion." For sequence prediction, time series prediction models based on the Transformer architecture (such as Informer and Autoformer) are used to model and predict change trends, outputting estimated values ​​for each data entity in the future (such as access intensity and structural risk level). Finally, using a graph neural network (such as Temporal GCN), all entities and their change trends are mapped into nodes and edges, generating a "data evolution prediction graph." Each node in the graph represents a data object, whose attributes include predicted future change trends, and edges represent dependencies and propagation paths, enabling a visual representation of the data's future evolutionary state.

[0032] A data corruption prediction model is constructed, using indicators such as the structural mutation rate, access anomaly amplitude, and dependency chain instability in the evolution prediction graph as input. A Bayesian network or logistic regression model is used to estimate the probability of "data corruption" or "irrecoverable failure" for each data node within the next T time window. Secondly, combined with the mapping relationship between data and business objects in the knowledge graph, the impact of data node anomalies on downstream business processes is evaluated, including the impact on business modules, the number of process nodes, and the degree of criticality. The impact level is divided into 1 (no impact) to 5 (business interruption). In terms of recovery complexity assessment, the number of redundant backups involved in the data node, the differences between versions, the number of dependent nodes, and other dimensions are counted to construct a complexity scoring function.

[0033] After constructing the data risk vector, the system uses matrix modeling to comprehensively quantify the overall risk level of each data entity, supporting subsequent priority scheduling and backup strategy decisions. This risk assessment matrix represents each data node in a row and three core dimensions in a column: probability of data corruption (P), scope of business impact (I), and recovery complexity (C). Each dimension is scored between [0, 1] or [0, 100]. The system can set a risk level weighting strategy based on actual business circumstances. For example, for fields in real-time trading systems, P is weighted at 0.5, I at 0.3, and C at 0.2; while for log fields, P=0.3, I=0.2, and C=0.5. After comprehensive weighting, a comprehensive risk score R is calculated for each node, with risk levels categorized as low (0-40), medium (41-70), high (71-90), and very high (91-100). The risk matrix also supports visualization, including heat maps and multi-dimensional scatter plots, to assist data governance personnel in identifying risk clusters. For example, during the experiment, we found that the R value of a core field in a financial system reached 94, triggering a high-priority asynchronous backup. This matrix will provide data support for the next step of identifying risk propagation paths.

[0034] Identify potential transmission relationships and pathways between high-risk data nodes and build a system-level risk diffusion model. First, based on the enterprise data association knowledge graph, select data nodes with high risk ratings (R>70) or higher. A directed graph G is constructed by combining their structural dependencies, API calls, and semantic references with other nodes. Then, using a transmission path analysis algorithm, such as PageRank or the Weighted Influence Propagation (WIP) algorithm, the influence of each node on downstream nodes is calculated. Nodes with high influence are designated "Key Risk Nodes," while nodes that may be affected by upstream data but are not inherently high risk are designated "Propagation Nodes." For example, if a field is referenced by seven downstream fields, three of which are financial analysis indicators, its propagation weight increases significantly. In the constructed "Data Risk Propagation Graph," each edge is assigned a risk transmission probability and delay parameter, and each node is assigned its own risk rating. Through graph queries, the system can quickly identify high-risk areas and links requiring isolation and protection. Integrating time series anomaly detection techniques (such as the Z-score method, STL decomposition, and LSTM Autoencoder), the system monitors high-risk nodes in real time for indicators such as risk value R, propagation edge weight W, and dependency chain activity L. If any dimension experiences an abnormal jump within a unit of time (e.g., an R value change of >20% with a sustained upward trend), a "transient mutation" flag is triggered. The system also conducts dynamic contextual assessments based on business cycles (such as month-end settlement and tax filing) to avoid misjudgments. Once the trigger conditions are met, the system automatically generates a "preventive backup plan" based on the node's impact range, backup priority, and current system load. This mechanism prioritizes incremental backups, asynchronous replication, or redundant replica promotion to ensure risk mitigation is completed without interrupting business. With an average response time of less than 300ms, this mechanism significantly reduces latency before data failure and serves as the core execution module for implementing a "predictable, preventative, and self-healing" intelligent backup strategy.

[0035] In this embodiment, reference Figure 4 The above is a flowchart of detailed implementation steps for predicting storage resource demand based on the intelligent backup trigger mechanism, dynamically scheduling resources in a multi-storage cloud environment, and building an elastic backup storage resource pool. In this embodiment, the detailed implementation steps for predicting storage resource demand based on the intelligent backup trigger mechanism, dynamically scheduling resources in a multi-storage cloud environment, and building an elastic backup storage resource pool include: Detecting that the intelligent backup trigger mechanism is in the triggered state, identifying all the data in the current enterprise system; Perform backup parameter analysis based on all current enterprise system data to obtain backup parameter configuration, including backup frequency, backup depth, and storage location; Predict storage resource requirements based on backup parameter configuration and required differential backup data, calculate required storage capacity, network bandwidth, and computing resources, and generate a storage resource demand forecast. Dynamically schedule resources in multiple storage cloud environments based on predicted storage resource demand values ​​to build an elastic backup storage resource pool.

[0036] In this embodiment, the trigger mechanism for intelligent backup can be automatically activated based on a scheduled policy (such as daily incremental backup), a threshold mechanism (such as when the data change rate exceeds a set threshold), a risk-driven mechanism (such as security incident detection), or model reasoning (such as using a semantic-aware model to identify changes in the state of critical data). The AI ​​model analyzes access logs, data update logs, and business flow data in real time to determine whether the trigger conditions are met. Once the trigger conditions are met, the system leverages the data asset catalog and knowledge graph to identify all currently active data. This includes structured data (such as MySQL and PostgreSQL), semi-structured data (such as XML and JSON configuration files), and unstructured data (such as contract PDFs and video footage). In an experimental scenario, 12:00 AM was set as the base trigger point daily, while monitoring business system CPU load changes and abnormal access logs. The overall trigger rate reached 98.6%, and the system identified an average of over 150,000 data nodes. This step ensures real-time and business-relevant backup operations, a prerequisite for achieving comprehensive data protection. After comprehensive data identification, parameterized analysis is performed based on each data type's business importance, access frequency, and dependencies to determine the optimal backup strategy. Backup parameter configuration includes, but is not limited to, backup frequency (e.g., real-time, hourly, daily), backup depth (full, incremental, differential), and storage location (local, remote, cloud). The AI ​​model integrates previously identified data structure features, semantic tags, and business dependency paths from the knowledge graph to perform policy mapping. For example, critical path data such as "settlement center ledger data" requires "real-time + incremental + remote active-active," while auxiliary system logs can be set to "daily + full + local storage." The model also dynamically optimizes key metrics such as historical backup success rates, data recovery time (RTO), and recovery point objective (RPO). In experiments, backup parameter configurations were automatically generated for a group's ERP and financial systems. The model's recommended backup frequency achieved 91.4% accuracy, and the depth selection aligned with the actual policy at 93.7%. This ultimately resulted in a structured backup parameter configuration table, providing a clear basis for resource forecasting and task scheduling.

[0037] By combining backup parameter configuration with the amount of data to be backed up, time series forecasting algorithms (such as ARIMA and LSTM) are used to predict data generation within the next 24-72 hours. Combining historical backup logs, the compression ratio, deduplication rate, and transmission efficiency of each data type are evaluated to calculate the actual storage capacity required. Next, network bandwidth requirements are modeled using the transmission rate per unit time (taking into account the bursty nature of traffic during peak transmission periods) to calculate the required Mbps or Gbps. Computing resource forecasting is dynamically estimated based on the workload of operations such as data preprocessing, encryption, and compression involved in the backup process, combining CPU utilization models with GPU accelerator models. In experiments, the resource demand forecasting model was trained using the last 30 days of backup job data. The overall forecast error was within ±7%, and the resource pre-allocation response time under high-load scenarios did not exceed 1.6 seconds. The final output is a storage resource demand forecast table, which provides a quantitative basis for subsequent multi-cloud resource scheduling, effectively reducing the risk of backup failures due to resource shortages. Through a resource orchestration engine (such as Kubernetes or HashiCorp Nomad), storage nodes on different cloud platforms (such as Alibaba Cloud, AWS, Tencent Cloud, and private clouds) are connected and virtualized, forming a unified resource abstraction interface. At this stage, a large AI model combines current predictions with historical resource utilization to learn the optimal allocation strategy for the resource pool (e.g., reinforcement learning strategy optimization or a multi-objective optimization algorithm), allocating resources based on data type and priority. For example, high-priority financial transaction data is scheduled to a cloud platform with SSD acceleration and cross-region active-active capabilities, while log data is allocated to a low-cost cold storage platform. The system supports horizontal and vertical resource upgrades, ensuring high concurrent scheduling capabilities during peak backup periods. In experiments, a virtual resource pool was established across three public clouds and one private cloud platform, with a maximum concurrent scheduling capacity of 3 TB of data per minute and a resource reconfiguration response time of less than 3 seconds. This elastic backup resource pool ensures high availability, low latency, and high throughput for the backup process, supporting the operation of an intelligent data backup system for large-scale enterprise environments.

[0038] In this embodiment, the specific steps of obtaining the previous version of system backup data, performing incremental backup analysis and adaptive compression coding, and obtaining the incremental backup code package are as follows: Obtaining the previous version of system backup data; performing differential data identification on the system backup data based on all current enterprise system data to obtain incremental backup data; Perform semantic-level fine-grained analysis on incremental backup data to obtain content changes, structural adjustments, and relationship updates of the current backup data, and then fit the semantic incremental change graph. Calculate the similarity between data on the semantic incremental change graph to obtain the data semantic similarity between different data; Performing intelligent data deduplication on the incremental backup data based on the data semantic similarity to obtain deduplication optimized incremental backup data; Adaptively compress and encode the deduplication-optimized incremental backup data to obtain an incremental backup encoding package.

[0039] In this embodiment, the metadata index corresponding to the previous backup task is located, including the backup timestamp, data snapshot version number, backup policy tag, and storage path. A backup indexing service (such as a backup mapping table based on Etcd or Redis) is used to quickly retrieve the target data content, ensuring the consistency and integrity of the comparison object. In real-world scenarios, this data is typically stored in a distributed file system (such as HDFS, MinIO) or object storage (such as S3, OSS) in compressed or encrypted form. The system supports version rollback and structured loading for full, incremental, and multi-version backups. A data decoder is simultaneously invoked to decompress and decrypt the data to ensure structural alignment with the current system data. For example, extracting the previous full backup from a government enterprise cloud system involved 3.4TB of data, and the distributed decoding engine took an average of 6.7 minutes. The core of this step is to establish a stable and traceable backup chain, providing a baseline data source for subsequent difference analysis and semantic recognition. By performing a multi-dimensional comparison of content, structure, and semantics between the current enterprise data and the previous backup, the changed portions, namely the "incremental backup data," are extracted. First, a hash comparison mechanism (such as MD5 and SHA-256) is used to quickly compare data blocks or record rows, identifying discrepant data blocks due to insertions, updates, and deletions. For structured data, such as database tables, precise record-level tracking is achieved through primary key indexing or record ID mapping. For unstructured data, such as contract documents and images, a text fingerprinting algorithm and feature hashing techniques are combined to identify document-level changes. AI large-scale models play a key role in this process, assisting in identifying semantic-level changes. For example, if two fields have similar content but the unit changes (e.g., "yuan" to "10,000 yuan"), the semantic model will identify it as a significant update. In experiments, daily transaction data from an e-commerce platform was analyzed for difference identification. The average incremental rate of difference comparisons over a seven-day window was 6.3%, reducing the data volume from 4.5TB to approximately 285GB. This precise identification allows only essential changes to be retained, saving significant resources for subsequent backup, storage, and transmission.

[0040] After extracting basic differences, further exploration is required to understand the semantic meaning behind the data changes, including the business implications of content updates, the impact of structural adjustments, and the evolution of entity relationships. This process utilizes large language models (such as GPT-4 and T5) for fine-grained semantic analysis. First, each data change unit (field, record, paragraph, etc.) is converted into an embedded vector representation. This is then matched against knowledge nodes in the business semantic space to identify additions, deletions, modifications, and expansions of semantic labels. For example, the change from "customer level" to "customer credit rating" is not only a name change but also an expansion of the semantic dimension. Next, the system analyzes the impact of data structural adjustments—such as new fields, table structure adjustments, and field type changes—on business processes. Using knowledge graph models (such as RDF graphs), a "semantic incremental change graph" is dynamically constructed to describe the content, path, and scope of the changes. In an experiment, the proportion of field-level structural adjustments in a manufacturing company's CRM system reached 12%. Using semantic graph modeling, 19 structural chains with cascading impacts on the settlement system were successfully identified. This graph serves as the core metadata for differential semantic backups, enabling precise recovery and intelligent analysis. Semantic similarity calculation is a key analytical tool for identifying redundant information, semantic overlap, and potential for merging between data. It is particularly suitable for identifying data items with semantic duplication, high similarity, or only minor content changes. The system analyzes each node (i.e., a data change record) in the semantic incremental change graph. Using Transformer-based pre-trained language models (such as RoBERTa and DeBERTa), the system encodes its semantic content into a high-dimensional vector. Similarity is then calculated using methods such as cosine similarity, Euclidean distance, or Mahalanobis distance. It supports pairwise comparison analysis of field pairs, record pairs, and structural change pairs to generate a "semantic similarity matrix." In a test, 300,000 incremental field pairs were analyzed within 19 seconds, with over 90% of highly similar pairs (similarity > 0.95) accurately identified. This mechanism not only identifies "renaming redundancy" but also handles redundant data such as "semantic expansion" or "semantic fusion," significantly improving the intelligence of deduplication analysis.

[0041] After determining semantic similarity, the system intelligently dedupes data based on a predefined similarity threshold. Unlike traditional deduplication methods based on byte comparison or exact field content match, semantic deduplication emphasizes duplicate identification at the semantic level. Using clustering algorithms (such as K-Means and Hierarchical Clustering), the system categorizes data entries with similarity above a predefined threshold into a set of redundant candidates. A large model then assists in determining which primary data items to retain and the appropriate merge strategy. For example, if two fields, "Payment Status" and "Payment Status," have different names but identical semantics, the system automatically merges their change records. Deduplication strategies support modes such as "primary-replica merge," "weight preservation," and "cross-version simplification." In experiments, semantic deduplication was performed on an incremental backup dataset. The original incremental data, approximately 286GB, was compressed to 178GB after deduplication, reducing data redundancy by 37.7% and maintaining semantic accuracy by over 95%. This step significantly reduces redundant transmission, storage, and recovery costs, providing an optimized data foundation for efficient backups. The final step is to efficiently compress and encode the processed incremental data, forming an "incremental backup encoded package" suitable for storage, transmission, and recovery. This compression process utilizes an adaptive algorithm, selecting a compression strategy based on data type (structured / text / multimedia), change density, and semantic tag distribution. For example, columnar compression algorithms (such as Parquet encoding combined with ZSTD compression) are applied to structured data, while word vector re-encoding (such as BPE-based word segmentation combined with contextual semantic vector encoding) is used for text data. Images or audio / video data are compressed using H.265 or WebP encoding combined with semantic summary clipping. The AI ​​model at this stage predicts the potential data fidelity, recovery time, and space usage after compression, assisting in selecting the optimal compression parameter combination. All data is packaged in a containerized archive format (such as .tar.gz / .pb / .avro), forming a recoverable compressed package with version identifiers, semantic tags, and structural indexes. Experimental results show that this method can improve the average compression ratio of incremental backup data from 1.8 to 2.6, and the average recovery and decompression time is kept under 5 seconds, achieving an ideal balance between compression efficiency and restoration accuracy. This encoding package becomes the core data unit in AI-driven data backup.

[0042] In this embodiment, the specific steps of dynamically adjusting the backup order of incremental backup code packets and making intelligent incremental backup decisions based on the elastic backup storage resource pool and building an intelligent incremental backup execution engine are as follows: Intelligently split the incremental backup code package into multiple backup code packages and multiple backup code tasks; Perform data dependency mining on multiple backup code packages to extract topological dependencies between backup data; Perform dynamic resource availability evaluation on the elastic backup storage resource pool to generate a dynamic resource availability evaluation value; Dynamically adjust the backup order of multiple backup encoding tasks based on the dynamic available resource evaluation value and the topological dependency relationship between backup data to generate a dynamic backup adjustment sequence; Make intelligent incremental backup decisions for dynamic backup adjustment sequences and build an intelligent incremental backup execution engine.

[0043] In this embodiment, complete incremental backup code packages are intelligently sharded to improve the parallelization of the backup process, task scheduling flexibility, and cross-node transmission efficiency. After compression and semantic packaging, backup code packages are typically large and contain complex data types. The system first logically shards them based on data type, business module, semantic label, and structural boundaries. Each shard corresponds to a semantically consistent and logically independent data unit. For example, a code package may contain multiple tables, multiple business fields, and several log files. The system slices them based on "table + semantic function" to ensure the semantic integrity of each backup task. Machine learning models (such as the XGBoost classifier) ​​are also used to predict the size, backup time, and resource consumption of each shard, and these predictions are factored into task scheduling weights. In an experimental environment, a 500GB incremental code package was divided into 32 backup task units, with an average shard size of 16.2GB and a time difference between the largest and smallest shards of no more than 8%, ensuring scheduling fairness and balanced resource utilization. This step lays the foundation for subsequent backup task parallelization and dynamic scheduling. After sharding the encoded packets, the logical relationships between the data in the shards need to be identified and a topological dependency graph constructed to prevent issues like order confusion and dependency corruption during the backup process. This process relies on technologies such as semantic graph analysis, field dependency tracking, and business process diagram modeling. The system extracts field primary and foreign key relationships and record references from structured data; semantically identifies trigger and data transmission paths between business modules; and analyzes file access chains and call sequences from operation logs. All this information is integrated to construct a "Backup Task Topological Dependency Graph" (BTDG). Nodes in the graph represent sharded tasks, and edges represent dependencies or data transfer directions. An AI model is used to predict and prioritize dependency edge weights, identifying critical path nodes. For example, a topological graph extracted from a manufacturing enterprise's ERP system contained 73 nodes and 192 edges, with a maximum depth of six layers and an average path weight variance of 0.41. When multiple paths exist, the model further selects the path with the smallest loop to avoid task blocking. This ensures the interpretability, order, and consistency of backup tasks, a prerequisite for intelligent decision-making.

[0044] To ensure efficient backup tasks within controllable resource limits, the system must evaluate the current availability of the elastic resource pool in real time, including storage space, computing load, and network bandwidth. This evaluation method relies on multi-source metric fusion and predictive modeling. First, the system collects metrics such as CPU utilization, I / O rate, remaining storage capacity, and network throughput for each storage node. These metrics are then fed into a time series prediction model (such as Prophet or LSTM) for trend prediction and calculation of the estimated resource availability value (ERAV) for the next backup cycle. Furthermore, an anomaly detection model (such as Isolation Forest) is introduced to identify potentially faulty nodes in the resource pool and adjust their contribution weights. The final evaluation results are output as a multi-dimensional set of metrics, such as "Node A: CPU 70%, storage available space 340GB, bandwidth available 2.3Gbps." Experimental data shows that after deploying this evaluation mechanism, the resource scheduling failure rate decreased by approximately 58%, node load balancing improved by 23%, and resource prediction error remained within ±5%. This step provides an accurate and dynamic resource view, essential for task prioritization and scheduling. After understanding resource availability and task dependencies, the system develops a dynamic task execution sequence to ensure the optimal execution path given limited resources and task dependencies. This process incorporates a multi-constraint scheduling optimization model that comprehensively considers task dependency order (provided by the topology map), resource constraints (provided by ERAV), and task priority (determined by semantic labels and data importance). The scheduling optimization approach uses reinforcement learning strategies (such as DQN) to generate task sequences and simulates node load to evaluate resource utilization and task completion times under different strategies. The system also incorporates a preemption mechanism that allows high-priority tasks to be replaced during periods of resource peaks. In actual testing, the backup sequence for a 40-task node took 185 minutes before optimization, but after scheduling optimization, it decreased to 129 minutes, with a task completion rate of 96%. This sequence not only prioritizes critical tasks but also improves overall resource utilization, forming a core scheduling logic for building an efficient backup engine.

[0045] After understanding resource availability and task dependencies, the system develops a dynamic task execution sequence to ensure the optimal execution path given limited resources and task dependencies. This process incorporates a multi-constraint scheduling optimization model that comprehensively considers task dependency order (provided by the topology map), resource constraints (provided by ERAV), and task priority (determined by semantic labels and data importance). The scheduling optimization approach uses reinforcement learning strategies (such as DQN) to generate task sequences and simulates node load to evaluate resource utilization and task completion times under different strategies. The system also incorporates a preemption mechanism that allows high-priority tasks to be replaced during periods of resource peaks. In actual testing, the backup sequence for a 40-task node took 185 minutes before optimization, but after scheduling optimization, it decreased to 129 minutes, with a task completion rate of 96%. This sequence not only prioritizes critical tasks but also improves overall resource utilization, forming a core scheduling logic for building an efficient backup engine.

[0046] In this embodiment, the specific steps of performing multi-task distributed parallel backup execution based on the intelligent incremental backup execution engine, and performing iterative verification optimization to construct an iteratively optimized backup data model are as follows: Perform multi-task distributed parallel backup execution based on the intelligent incremental backup execution engine, conduct real-time backup monitoring, and collect backup status monitoring information; Calculate real-time backup speed, storage resource utilization, and backup accuracy based on backup status monitoring information; Perform backup quality assessment based on the real-time backup speed, storage resource utilization, and backup accuracy, and generate a backup quality assessment report; Perform backup fault diagnosis based on the backup quality assessment report and mark backup fault data points; Perform fault attribution analysis on backup fault data points to obtain backup fault factors; Based on the backup failure factors, the backup strategy parameters of the intelligent incremental backup execution engine are adjusted, and iterative verification and optimization are performed to build an iterative optimization backup data model.

[0047] In this embodiment, after completing dynamic task sequence scheduling, the intelligent incremental backup execution engine officially enters the operational phase. This engine supports concurrent multi-task execution and, based on containerized deployments (such as Kubernetes and a sidecar architecture), distributes each backup encoding task to multiple nodes, executing them in an orderly manner based on topological dependencies and priority. After the task is launched, the system simultaneously activates a real-time monitoring module to continuously collect key metrics during each backup task execution cycle, including task start time, completion status, I / O read / write rates, network latency, compression ratio, and number of abnormal interruptions. This monitoring data is transmitted to a central monitoring center via lightweight proxy components (such as Prometheus Node Exporter and Fluent Bit), and integrated with OpenTelemetry tracking standards to achieve unified and structured monitoring and tracking. For example, during a concurrent execution of 28 backup tasks, the system collected approximately 1,400 metrics per minute, covering real-time parameters such as task status (Success / Failure / Pending), node load (CPU, memory), and task duration. This step enables visual tracking of the entire backup task process, providing the foundation for subsequent quality assessment and troubleshooting. The real-time backup monitoring information collected by the system requires further structured processing to quantify task execution efficiency and backup accuracy. First, real-time backup speed is calculated based on data throughput. The average speed (V = D / T) is calculated by combining the task duration (T) and the data volume (D). Short-term fluctuations are filtered out using a moving average and outlier removal algorithm. Second, storage resource utilization, including disk I / O, free space percentage, and compressed data ratio, is calculated using node-level sampling (e.g., utilization of storage node A = occupied space / total capacity). Finally, backup accuracy is assessed using verification mechanisms, such as CRC32 checks, file hash comparisons, and schema matching to determine whether the task has fully backed up the target content. Logical comparisons are used for structured data, such as comparing the number of primary key entries and schema hash values ​​before and after the backup; content summary comparisons are used for unstructured content. In the experiment, data collected from an enterprise's private cloud daily incremental backup system showed an average backup speed of 520MB / s, a stable resource utilization rate of 74% after compression, and a backup accuracy rate exceeding 98.7%. The core indicators output from this step constitute a quantitative portrait of task execution quality, providing support for further evaluation.

[0048] The real-time backup monitoring information collected by the system requires further structured processing to quantify task execution efficiency and backup accuracy. First, real-time backup speed is calculated based on data throughput. The average speed (V = D / T) is calculated by combining task duration (T) and data volume (D). Short-term fluctuations are filtered out using a moving average and outlier rejection algorithm. Second, storage resource utilization, including disk I / O, free space percentage, and compressed data ratio, is calculated using node-level sampling (e.g., utilization of storage node A = occupied space / total capacity). Finally, backup accuracy is assessed using verification mechanisms, such as CRC32 checks, file hash comparisons, and schema matching to determine whether the task has fully backed up the target content. Logical comparisons are used for structured data, such as comparing the number of primary key entries and schema hash values ​​before and after the backup; content summary comparisons are used for unstructured content. In experiments, data collected from an enterprise private cloud daily incremental backup system showed an average backup speed of 520 MB / s, a stable resource utilization rate of 74% after compression, and a backup accuracy exceeding 98.7%. The core metrics output from this step constitute a quantitative profile of task execution quality, providing support for further evaluation. If a low-scoring task (e.g., a score <60 or an accuracy rate below 90%) is included in the quality assessment report, the system automatically triggers a fault diagnosis process to conduct fine-grained locating and analysis of task failures or deviations. The diagnostic process consists of three parts: 1) Task-level analysis, identifying the time of failure, abnormal log information, and upstream and downstream task status; 2) Node-level analysis, detecting resource bottlenecks, load anomalies, or I / O conflicts on storage and compute nodes; and 3) Data-level analysis, performing logical consistency checks and content integrity verification on the data segments involved in the task. All data segments identified as "abnormal writes, structural corruption, or verification failures" are marked as "faulty data points." For example, during a backup task, network fluctuations occurred while Task B was writing to a distributed node. The system captured a "connection reset error (ECONNRESET)" in the log and located the write failures of three data blocks in the corresponding encoded segments, marking these as "potential data loss failure points." This step establishes a precise fault location mechanism and provides a complete abnormal sample for the next step of attribution analysis.

[0049] If a low-scoring task (e.g., a score <60 or accuracy below 90%) appears in the quality assessment report, the system automatically triggers a fault diagnosis process, conducting fine-grained localization and analysis of task failures or deviations. This diagnostic process consists of three steps: 1) task-level analysis, identifying the time of failure, abnormal log information, and upstream and downstream task status; 2) node-level analysis, detecting resource bottlenecks, load anomalies, or I / O conflicts on storage and compute nodes; and 3) data-level analysis, performing logical consistency checks and content integrity verification on the data segments involved in the task. All data segments identified as "abnormal writes, structural corruption, or verification failures" are marked as "faulty data points." For example, during a backup task, network fluctuations occurred while Task B was writing to a distributed node. The system captured a "connection reset error (ECONNRESET)" in the log and located the write failures of three data blocks in the corresponding encoded segments, marking them as "potential data loss failure points." This step establishes a precise fault localization mechanism, providing a complete anomaly sample for subsequent attribution analysis. After fault attribution, the system automatically adjusts backup policy parameters to improve the stability and efficiency of future tasks. Policy adjustments include task rescheduling (changing execution nodes or adjusting execution times), reducing compression ratios, reconfiguring task granularity, optimizing bandwidth control, and enhancing caching mechanisms. The system uses a combination of reinforcement learning and policy optimization models (such as PolicyGradient + Greedy Search) to rapidly validate the adjustment results in a simulation environment. After each round of backup tasks, the system feeds the current policy execution results into an AI model, forming an "execution-feedback-adjustment" optimization cycle. This mechanism builds an iteratively optimized backup data model based on historical task data. It then fits and predicts the relationship between different task characteristics (such as data type, compression ratio, and task sequence) and policy parameters, assisting in developing optimized execution plans. For example, in the task model of a large educational platform, the average task failure rate was 6.4% in the initial phase. After five rounds of policy iterative optimization, it dropped to 2.1%. The average task execution time per task was reduced by 18%, and task resource allocation efficiency was improved by 22%. The iterative optimization model finally constructed becomes the self-learning and adaptive intelligent core of the backup system, with cross-task and cross-cycle tuning capabilities.

[0050] In this embodiment, an intelligent data backup system based on an AI big model is provided, which is used to execute the intelligent data backup method based on the AI ​​big model described above, including: The semantic perception module is used to obtain the enterprise's global data list, perform intelligent data structure deconstruction and dynamic attribute mapping modeling, and build a holographic data semantic perception model; The self-triggered backup module is used to perform real-time transient risk mutation detection on the holographic data semantic perception model, make self-triggered preventive backup decisions, and build an intelligent backup trigger mechanism; The storage demand prediction module is used to predict storage resource demand based on the intelligent backup trigger mechanism, dynamically schedule resources in multiple storage cloud environments, and build an elastic backup storage resource pool; The backup encoding module is used to obtain the system backup data of the previous version, perform incremental backup analysis and adaptive compression encoding, and obtain an incremental backup encoding package; The backup sequence adjustment module is used to dynamically adjust the backup sequence of incremental backup code packages and make intelligent incremental backup decisions based on the elastic backup storage resource pool, thus building an intelligent incremental backup execution engine. The distributed parallel backup module is used to perform multi-task distributed parallel backup execution based on the intelligent incremental backup execution engine, and perform iterative verification optimization to build an iterative optimization backup data model.

[0051] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.

[0052] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An intelligent data backup method based on AI big model, characterized in that: The following steps are involved: Obtain an enterprise-wide data inventory, perform intelligent data structure deconstruction and dynamic attribute mapping modeling, and build a holographic data semantic perception model; Conduct real-time transient risk mutation detection on the holographic data semantic perception model, make self-triggered preventive backup decisions, and build an intelligent backup trigger mechanism; Predict storage resource demand based on an intelligent backup trigger mechanism, dynamically schedule resources in multiple storage cloud environments, and build a flexible backup storage resource pool. Obtain the previous version of the system backup data, perform incremental backup analysis and adaptive compression encoding, and obtain an incremental backup encoding package; Based on the elastic backup storage resource pool, the incremental backup code package is dynamically adjusted in backup order and intelligent incremental backup decision is made to build an intelligent incremental backup execution engine; Based on the intelligent incremental backup execution engine, multi-task distributed parallel backup execution is performed, and iterative verification and optimization are carried out to build an iterative optimization backup data model.

2. The intelligent data backup method based on AI big model according to claim 1 is characterized in that: The specific steps of obtaining the enterprise's global data list, performing intelligent data structure deconstruction and dynamic attribute mapping modeling, and building a holographic data semantic perception model are as follows: Obtain a list of enterprise-wide data; Identify and filter abnormal redundant data in the enterprise's global data list to generate a redundant filtered data list; Intelligent data structure deconstruction is performed on the redundant filtered data list to generate multi-dimensional enterprise data structure characteristics, wherein the multi-dimensional enterprise data structure characteristics include data structure body, business attributes, data types and data element hierarchical characteristics; Perform deep semantic analysis on the redundant filtered data list to extract multimodal semantic features; Based on the multi-dimensional enterprise data structure characteristics and multi-modal semantic features, we mine the correlation between cross-data sources, identify the business dependency links and impact propagation paths between data, and build an enterprise data association knowledge graph; Perform dynamic attribute mapping modeling on enterprise data-related knowledge graphs and build a holographic data semantic perception model.

3. The intelligent data backup method based on AI big model according to claim 2 is characterized in that: The specific steps of performing dynamic attribute mapping modeling on the enterprise data-related knowledge graph and building a holographic data semantic perception model are as follows: Identify multiple data entity nodes based on the enterprise data association knowledge graph; Calculate the centrality index of the data entity nodes one by one to obtain the graph criticality of each data entity; Perform data access frequency statistics on multiple data entity nodes to obtain data access frequency; quantifying the business importance values ​​of the plurality of data entity nodes; Calculate the comprehensive value weight based on the graph criticality, data access frequency and business importance, and perform intelligent label fitting to generate a priority label for each data entity; Based on multimodal semantic features, personalized semantic feature mining is performed on multiple data entity nodes to obtain multiple data entity semantic fingerprints; Generate a unique semantic identifier based on the data entity semantic fingerprint; Dynamic attribute mapping modeling is performed on the enterprise data association knowledge graph based on the unique semantic identifier and the priority label to build a holographic data semantic perception model.

4. The intelligent data backup method based on AI big model according to claim 1 is characterized in that: The specific steps of performing real-time transient risk mutation detection on the holographic data semantic perception model and making self-triggered preventive backup decisions to build an intelligent backup trigger mechanism are as follows: Analyze the changes in time series data streams on the holographic data semantic perception model, extract historical data modification trajectories, access patterns, and changes in inter-data dependencies, and fit the multi-dimensional time series data change characteristics; Mining data evolution rules based on the changing characteristics of multi-dimensional time series data to generate data evolution prediction graphs; Calculate the probability of data corruption, business impact, and recovery complexity based on the data evolution prediction graph; Conduct data risk assessment based on data corruption probability, business impact scope, and recovery complexity, and build a data risk assessment matrix; Analyze the risk propagation path based on the data risk assessment matrix, mark key risk nodes and risk propagation nodes, and construct a data risk propagation map; Perform real-time transient risk mutation detection on the data risk propagation graph, make self-triggered preventive backup decisions, and build an intelligent backup trigger mechanism.

5. The intelligent data backup method based on AI big model according to claim 1 is characterized in that: The specific steps for predicting storage resource demand based on the intelligent backup trigger mechanism and dynamically scheduling resources in a multi-storage cloud environment to build an elastic backup storage resource pool are as follows: Detecting that the intelligent backup trigger mechanism is in the triggered state, identifying all the data in the current enterprise system; Perform backup parameter analysis based on all current enterprise system data to obtain backup parameter configuration, including backup frequency, backup depth, and storage location; Predict storage resource requirements based on backup parameter configuration and required differential backup data, calculate required storage capacity, network bandwidth, and computing resources, and generate a storage resource demand forecast. Dynamically schedule resources in multiple storage cloud environments based on predicted storage resource demand values ​​to build an elastic backup storage resource pool.

6. The intelligent data backup method based on AI big model according to claim 1 is characterized in that: The specific steps of obtaining the previous version of system backup data, performing incremental backup analysis and adaptive compression coding, and obtaining the incremental backup code package are as follows: Get the previous version of the system backup data; Performing differential data identification on the system backup data based on all current enterprise system data to obtain incremental backup data; Perform semantic-level fine-grained analysis on incremental backup data to obtain content changes, structural adjustments, and relationship updates of the current backup data, and then fit the semantic incremental change graph. Calculate the similarity between data on the semantic incremental change graph to obtain the data semantic similarity between different data; Performing intelligent data deduplication on the incremental backup data based on the data semantic similarity to obtain deduplication optimized incremental backup data; Adaptively compress and encode the deduplication-optimized incremental backup data to obtain an incremental backup encoding package.

7. The intelligent data backup method based on AI big model according to claim 1 is characterized in that: The specific steps of dynamically adjusting the backup order of incremental backup code packets and making intelligent incremental backup decisions based on the elastic backup storage resource pool to build an intelligent incremental backup execution engine are as follows: Intelligently split the incremental backup code package into multiple backup code packages and multiple backup code tasks; Perform data dependency mining on multiple backup code packages to extract topological dependencies between backup data; Perform dynamic resource availability evaluation on the elastic backup storage resource pool to generate a dynamic resource availability evaluation value; Dynamically adjust the backup order of multiple backup encoding tasks based on the dynamic available resource evaluation value and the topological dependency relationship between backup data to generate a dynamic backup adjustment sequence; Make intelligent incremental backup decisions for dynamic backup adjustment sequences and build an intelligent incremental backup execution engine.

8. The intelligent data backup method based on AI big model according to claim 1 is characterized in that: The specific steps of performing multi-task distributed parallel backup execution based on the intelligent incremental backup execution engine, and performing iterative verification optimization to build an iterative optimization backup data model are as follows: Perform multi-task distributed parallel backup execution based on the intelligent incremental backup execution engine, conduct real-time backup monitoring, and collect backup status monitoring information; Calculate real-time backup speed, storage resource utilization, and backup accuracy based on backup status monitoring information; Perform backup quality assessment based on the real-time backup speed, storage resource utilization, and backup accuracy, and generate a backup quality assessment report; Perform backup fault diagnosis based on the backup quality assessment report and mark backup fault data points; Perform fault attribution analysis on backup fault data points to obtain backup fault factors; Based on the backup failure factors, the backup strategy parameters of the intelligent incremental backup execution engine are adjusted, and iterative verification and optimization are performed to build an iterative optimization backup data model.

9. An intelligent data backup system based on AI big model, characterized by: The method for executing the intelligent data backup method based on the AI ​​big model according to claim 1 comprises: The semantic perception module is used to obtain the enterprise's global data list, perform intelligent data structure deconstruction and dynamic attribute mapping modeling, and build a holographic data semantic perception model; The self-triggered backup module is used to perform real-time transient risk mutation detection on the holographic data semantic perception model, make self-triggered preventive backup decisions, and build an intelligent backup trigger mechanism; The storage demand prediction module is used to predict storage resource demand based on the intelligent backup trigger mechanism, dynamically schedule resources in multiple storage cloud environments, and build an elastic backup storage resource pool; The backup encoding module is used to obtain the system backup data of the previous version, perform incremental backup analysis and adaptive compression encoding, and obtain an incremental backup encoding package; The backup sequence adjustment module is used to dynamically adjust the backup sequence of incremental backup code packages and make intelligent incremental backup decisions based on the elastic backup storage resource pool, thus building an intelligent incremental backup execution engine. The distributed parallel backup module is used to perform multi-task distributed parallel backup execution based on the intelligent incremental backup execution engine, and perform iterative verification optimization to build an iterative optimization backup data model.

Citation Information

Patent Citations

  • Data center data backup disaster recovery intelligent management and control platform and method

    CN118245285A

  • Distributed disaster recovery and AI intelligent calculation fusion method and system, equipment and medium

    CN118394568A

  • Data backup method and device based on large model

    CN119201542A

  • System and method for optimizing memory usage during data backup

    US20080140960A1

  • Disaster recovery in a cell model for an extensibility platform

    US20230315580A1

Cited By

  • Automatic data quality inspection system and method based on large model and data flow arrangement

    CN120893585A

  • Method and system for automatically generating data backup strategy

    CN120909848A

  • A data backup strategy automatic generation method and system

    CN120909848B

  • Server BMC intelligent management method and system based on edge computing

    CN121070737A

  • Intelligent early warning method and system based on multiple Agents

    CN121077946A