A data rescue and archiving system based on cloud platform
Through hybrid encryption technology and digital twin simulation modules, combined with edge computing and reinforcement learning optimization modules, the performance bottlenecks and compatibility issues of SSD data recovery and cloud data rescue and archiving are solved, and efficient data transmission and recovery are achieved.
Patent Information
- Application Number
- CN202510803649.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing technologies cannot be effectively applied to data recovery on solid-state drives (SSDs). Traditional encryption methods have performance bottlenecks and compatibility issues during cloud data rescue and archiving, making it difficult to ensure encryption strength and real-time processing speed. There is a lack of virtual-real interaction testing for data rescue and archiving.
Using hybrid encryption technology, digital twin simulation module and reinforcement learning optimization module, combined with edge computing nodes to collect data in real time, a multi-dimensional data state vector model is generated through a lightweight modeling engine to conduct virtual-reality interaction testing and disaster simulation, optimize data transmission paths, and use deep reinforcement learning models to generate optimal strategies.
It achieves end-to-end encryption, improves data compression rate, reduces transmission bandwidth requirements and cloud storage costs, shortens recovery time, enhances data integrity control, and optimizes resource allocation and transmission paths.
Smart Images

Figure CN120315653B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and in particular to a data rescue and archiving system based on a cloud platform. Background Art
[0002] With the popularization of personal computers, users' demand for data security has increased, prompting the emergence of data recovery software such as Norton Utilities. This type of software can scan hard drives and recover accidentally deleted or damaged files. As a result, solid-state drives (SSDs) have gradually replaced traditional mechanical hard drives as the mainstream storage medium. However, since SSDs use flash memory as storage media, their working principle is completely different from that of traditional magnetic media hard drives, especially the mapping of logical addresses to physical addresses through the FTL (Flash Translation Layer). This change makes traditional magnetic media-based data recovery technology ineffective and cannot be directly applied to SSDs. Therefore, data recovery technology for SSDs has to be innovated, and tools specifically for reverse engineering FTL have been developed to parse the data structure inside the SSD, thereby improving the success rate of data recovery.
[0003] Although traditional encryption methods can ensure data security to a certain extent, they often have performance bottlenecks or compatibility issues. In particular, in the process of cloud data rescue and archiving, it is difficult to ensure encryption strength and real-time processing speed. In the process of building digital models that accurately reflect the behavior of the physical world, existing technologies have difficulty in processing large amounts of heterogeneous data in a timely manner, as well as balancing model simplification and accuracy. At the same time, there is a lack of virtual-reality interaction testing for data rescue and archiving, as well as corresponding data solutions. Summary of the Invention
[0004] In order to solve the problems raised in the above background technology, the present invention is implemented through the following technical solutions:
[0005] A data rescue and archiving system based on a cloud platform, comprising:
[0006] The data perception and collection module collects real-time operational data by deploying edge computing nodes. The operational data includes sensor data, business logs, and historical fault data. Hybrid encryption technology encrypts the transmission link and transmits it to the cloud through a unified data interface.
[0007] The digital twin simulation module generates a multidimensional data state vector model based on a lightweight modeling engine, integrates a spatiotemporal data fusion algorithm to dynamically update the twin state, introduces pre-disaster conditions for virtual-reality interaction testing, and uses the Monte Carlo method to simulate the execution effect of data recovery plans under different disaster conditions, thereby simulating the spread of disasters.
[0008] The reinforcement learning optimization module uses a deep reinforcement learning model, takes as input multi-dimensional evaluation indicators, historical fault data, and business logs, and outputs the optimal strategy.
[0009] The strategy execution and feedback module dynamically allocates emergency resources through a two-layer network based on the optimal strategy generated by reinforcement learning, and uses a path planning algorithm to optimize the data transmission path.
[0010] Furthermore, collect operation data:
[0011] Collect sensor data through IoT devices or sensor networks;
[0012] Convert the records of the application during the execution of its functions into business logs;
[0013] Through the maintenance management system, detailed information on system or equipment failures is collected and recorded as historical failure data.
[0014] Furthermore, compress the data and encrypt the transmission link:
[0015] S101: Preprocessing the operation data; inserting time points to obtain time series operation data;
[0016] S102: Use the Zstandard algorithm to perform preliminary compression on the preprocessed time series running data;
[0017] S103: Based on the initial compression of Zstandard, the machine learning model is used to run the input time series data and output patterns and regularities. Through AI predictive coding, the patterns and regularities in the data are used to predict future data points.
[0018] S104: Combining the result of Zstandard compression with the output generated by AI predictive coding for optimization.
[0019] Furthermore, predict future data points:
[0020] S103-1: Divide the Zstandard compressed data stream into time windows and extract multi-dimensional features, including statistical features, frequency domain features, and pattern markers;
[0021] S103-2: Build a dynamic context-aware model, maintain the state vector in real time, and save the evolution pattern of the first Z data points;
[0022] S103-3: Adopting adaptive quantization residual coding, dynamically adjusting the quantization step size, establishing a parameter update dictionary, and transmitting only non-zero parameter gradient blocks;
[0023] S103-4: Cross-sensor collaborative prediction, spatial correlation matrix maintenance, and calculation of spatial correlation coefficients;
[0024] S103-5: Maintain the candidate model pool, switch models based on online evaluation indicators, perform context-adaptive Range encoding on the prediction residuals, use the context model, predict the confidence interval, and quantize the step size level.
[0025] Furthermore, a lightweight modeling engine is constructed to generate a multi-dimensional data state vector model:
[0026] The QEM algorithm is used to simplify the mesh of the 3D geometric model. A hierarchical optimization strategy is used to iteratively shrink vertex pairs to reduce the number of facets. The simplified model is encoded using Draco compression technology. For dynamic parameters, an LSTM prediction network is integrated to generate a time series state vector. The LBFGS quasi-Newton method is used to optimize vertex positions and minimize surface energy. The model is verified using the Metro mesh algorithm, and the maximum geometric error and root mean square error between the 3D geometric model and the mesh simplified model are calculated.
[0027] By deploying edge sensors and cloud platform API interfaces, physical environment data, business data, and network topology data can be acquired in real time and processed in multiple stages.
[0028] Furthermore, update the twin status:
[0029] The multi-source data after preprocessing and alignment are fused. A spatiotemporal data fusion algorithm is used for data from different sources to obtain the target state. Based on the fusion results, a lightweight modeling engine is used to calculate and update the state vector model of the digital twin. A feedback loop is established to continuously monitor changes in the actual system and feed back new observation results to the digital twin, constantly adjusting its internal state and achieving dynamic updates.
[0030] Furthermore, we simulate the execution effect of data recovery plans under different disaster conditions:
[0031] The Monte Carlo method is used to generate random combination scenarios, simulate the data island effect caused by physical damage to the server, and deduce the feasibility of cross-cabinet data migration paths. By estimating the amount of data loss, quantitative indicators are generated based on the backup cycle, deduplication threshold, and recovery time confidence interval.
[0032] Furthermore, the optimal strategy is output:
[0033] The 10-dimensional evaluation vector, fault scalar value, and LSTM hidden state vector generated by the autoencoder are spliced together to form a comprehensive state vector, which integrates static indicators, dynamic time series, and historical risks to provide multi-dimensional input for the digital twin. The current strategy is loaded in the digital twin environment, and the system operation is simulated based on the comprehensive state vector. During the deduction process, trajectory data is generated in real time, and the state, action, and reward signal of each step are recorded to construct the interaction sequence required for reinforcement learning. The trajectory data is analyzed using generalized advantage estimation to calculate the advantage value of each action step. GAE balances the estimated bias and variance by weighted average of short-term returns and long-term value, and quantifies the direction of strategy improvement.
[0034] Furthermore, the path planning algorithm is used to optimize the data transmission path:
[0035] The network path is dynamically configured through the OpenFlow protocol. The control layer detects network devices through the LLDP protocol and builds a real-time topology map. Independent virtual network slices are created according to the task type, including rescue slices and archive slices. The parameters are mapped to the predefined global policy framework. The policy parameters output by the reinforcement learning module are used to increase the resource allocation of a certain type of task on the node.
[0036] The present invention provides a data rescue and archiving system based on a cloud platform, which has the following beneficial effects:
[0037] (1) The present invention achieves end-to-end encryption through hybrid encryption technology; and the dynamic path planning algorithm avoids high-risk nodes through real-time link status perception; adopts Zstandard+AI predictive coding technology, combined with local preprocessing of edge computing nodes, to increase data compression rate by 40%, significantly reducing transmission bandwidth requirements and cloud storage costs.
[0038] (2) The present invention uses a digital twin simulation module to simulate multi-disaster scenarios, such as fire and cyber attacks, using the Monte Carlo method to generate quantitative recovery indicators, such as RTO and data loss rate; the reinforcement learning module dynamically optimizes the backup strategy based on historical failures and real-time data; automatically adjusts the backup frequency, shortens the recovery time, and controls the loss of critical data integrity.
[0039] (2) The present invention allocates resources according to task priority through a two-layer networking network (control layer + data layer), combines dynamic adjustment of deduplication thresholds to reduce redundant storage, integrates spatiotemporal data fusion algorithms to dynamically update twin states, introduces pre-disaster conditions for virtual-reality interaction testing, inputs pre-conditions for data rescue and archiving path strategies, simulates the execution effects of data recovery plans under different disaster conditions through the Monte Carlo method, and conducts disaster spread simulation. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1Schematic diagram of the system flow of the present invention. DETAILED DESCRIPTION
[0041] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0042] See also Figure 1 This embodiment provides a data rescue and archiving system based on a cloud platform, the system comprising:
[0043] The data perception and collection module deploys edge computing nodes to collect real-time operational data, including sensor data, business logs, and historical fault data. It uses Zstandard+AI predictive coding technology to compress the data, combined with hybrid encryption technology to encrypt the transmission link, and transmits the data to the cloud through a unified data interface.
[0044] Selection and configuration of edge computing nodes:
[0045] Select appropriate edge computing devices based on the application scenario, install the corresponding operating system, and configure network settings and dependencies;
[0046] For example:
[0047] In the smart agriculture project, multi-source sensors are deployed to monitor farmland environmental conditions, such as soil moisture, temperature, and light intensity. Irrigation and farming activities are optimized based on this data, ensuring real-time data processing and reducing the amount of data transmitted to the cloud. Multiple edge computing nodes are deployed. Considering power supply constraints and weather resistance requirements, stable and low-power industrial-grade gateways are selected as edge computing nodes. These gateways support multiple input and output interfaces and have good scalability and adaptability. To meet the requirements of low resource usage, high security, and easy maintenance, a lightweight Linux distribution is installed, and an SSH service is configured for remote management and maintenance. Necessary libraries and tools, such as the Python runtime environment and its related data analysis packages, are installed for local data processing. Cron jobs are set up to regularly check system health and automatically update software. Based on actual needs, static or dynamic IP address acquisition methods are configured, and firewall rules are set to only allow access requests from trusted sources to enhance security.
[0048] Sensor data:
[0049] Information collected by various physical sensors, including but not limited to environmental or equipment operating parameters such as temperature, humidity, pressure, and location. This data is typically acquired through IoT devices or sensor networks. For example, in industrial environments, temperature sensors can be deployed to monitor the operating temperature of machines.
[0050] Business log:
[0051] Log records generated by an application or system during the execution of its functions, including user activities, transaction records, error reports, performance indicators, and other information. Business logs are usually automatically generated by the software's built-in logging mechanism, and specific log levels and content can be configured as needed.
[0052] Historical fault data:
[0053] Record detailed information about system or equipment failures, including the time, type, cause, and resolution. This information can be collected through maintenance management systems such as CMMS, event management tools, or manual recording. For manufacturing companies, historical failure data can assist in analyzing equipment failure modes and provide a basis for preventive maintenance plans.
[0054] Compressed data:
[0055] S101: Data preprocessing: Data collected in real time from sources such as sensors, business logs, and historical fault records undergoes preliminary cleaning and formatting, and time points are inserted to obtain time series operating data to ensure data consistency and accuracy. Duplicate data is removed, missing values are filled, and data types are converted.
[0056] S102: Use the Zstandard algorithm to perform preliminary compression on the preprocessed time series data, providing a high compression ratio while maintaining high compression and decompression speeds. Adjust the algorithm based on actual needs to balance compression efficiency and resource consumption.
[0057] Zstandard - a fast and efficient lossless compression algorithm.
[0058] S103: Based on Zstandard's initial compression, AI predictive coding technology is applied, using machine learning models to run input time series data or repeating pattern data, outputting patterns and regularities. AI predictive coding identifies and utilizes patterns and regularities in the data to predict future data points, thereby reducing the amount of data that needs to be stored or transmitted.
[0059] S103-1: Divide the Zstandard compressed data stream into time windows and extract multidimensional features, including statistical features, frequency domain features, and pattern markers. For example, industrial sensor data is divided into 5-second windows and feature vectors such as mean, variance, slope, and FFT dominant frequency component are extracted. Long-term dependencies are established using LSTM or Transformer architectures.
[0060] S103-2: Build a dynamic context-aware model, maintain the state vector in real time, and save the evolution pattern of the first 100 data points. The pseudo code example is as follows:
[0061] def hybrid_predict(current_data, model):
[0062] # Short-term forecast based on statistical models
[0063] ar_pred = autoregressive(current_data[-10:], order=3)
[0064] # Deep learning model prediction
[0065] dl_pred = model(torch.tensor(current_data).unsqueeze(0))
[0066] # Adaptive weight fusion
[0067] error_stats = calculate_recent_errors()
[0068] final_pred = (0.7*dl_pred + 0.3*ar_pred) * error_adjustment_factor
[0069] return quantize(final_pred, precision=0.001)
[0070] S103-3: Adopt adaptive quantization residual coding and dynamically adjust the quantization step size Q = ƒ, where Q is the data volatility and f is the prediction confidence. Build a parameter update dictionary and only transmit non-zero parameter gradient blocks; for example, transmit a 3KB model fine-tuning parameter package for every 1000 data points.
[0071] S103-4: Cross-sensor collaborative prediction, spatial correlation matrix maintenance, calculation of spatial correlation coefficient, the formula is:
[0072] ;
[0073] in, is the spatial correlation coefficient, specifically the spatial correlation coefficient between sensor k and sensor j, with a value range of [−1, 1], which is used to quantify the degree of linear correlation between the two sensor data sequences; and is the observation value of sensor k and sensor j at the same time point; and is the arithmetic mean of the observation values of sensor k and sensor j in the current time window; and is the standard deviation of sensor k and sensor j in the time window; when > 0.8, activate the joint prediction mode; otherwise, keep the current state;
[0074] S103-5: Maintain the candidate model pool, switch models based on online evaluation indicators, perform context-adaptive range coding on the prediction residuals, use an 8-level context model, predict confidence intervals, and quantize step sizes. The code example is:
[0075] ```
[0076] if rolling_MAPE > threshold:
[0077] load_lightweight_model()
[0078] elif data_distribution_shift_detected:
[0079] activate_retraining_protocol()
[0080] ```
[0081] S104: The results of Zstandard compression are combined with the output generated by AI predictive coding for optimization. The trained AI prediction model and the compressed residual data are stored together. During decompression, both the prediction model and the compressed data are needed to restore the original data. The AI prediction model information and the Zstandard compressed residual data can be included simultaneously. For example, a file header can be used to store the model information, followed by the compressed residual data. This achieves a more efficient compression effect, not only leveraging the advantages of traditional compression algorithms, but also using AI to discover and utilize deep structures in the data, optimizing the compression ratio and efficiency.
[0082] Encrypted transmission link:
[0083] Combining the efficiency of symmetric encryption with the security of asymmetric encryption, end-to-end secure transmission is achieved. The cloud server generates an RSA / ECC asymmetric key pair, namely public key encryption and private key decryption, and publishes the public key to the edge node via a digital certificate. The edge node randomly generates a symmetric key, encrypts the key with the cloud public key, generates an encrypted session key package, and verifies the legitimacy of the public key through a digital certificate to prevent man-in-the-middle attacks. The edge node uses Zstandard to compress the original data and combines it with AI predictive coding to reduce redundancy. The compressed data is encrypted in block mode using AES-GCM mode, and an integrity check tag is generated to prevent tampering. The encrypted data and session key package are encapsulated according to a unified interface standard and attached with metadata such as timestamp and device ID.
[0084] Unified data interface standardization:
[0085] The RESTful API architecture is used to define interface specifications, supporting the HTTP / 2 protocol to improve transmission efficiency. Through standardized URI path design, unified encoding of device identifiers and data types is implemented to ensure interface compatibility across different edge nodes. A standardized data structure based on JSON Schema is defined, which is compatible with heterogeneous data sources such as sensor data and business logs, and multimodal data fusion is achieved through unified metadata fields.
[0086] The transmission channel is established through the HTTPS / TLS 1.3 protocol, with forward secrecy ensuring temporary session security. The transmission path is dynamically selected, switching to a backup link when an anomaly occurs, and random padding is added to the data packet to standardize the ciphertext length to resist traffic analysis attacks. After transmission to the cloud, the private key is used to decrypt the session key package, obtain the AES key, and use the AES key to decrypt the data block. The GCM tag is verified to ensure data integrity. If decryption fails or the verification tag does not match, an alarm is triggered and a retransmission is requested, and the key pool is updated. If successful, the symmetric encryption session key and initialization vector in memory are immediately destroyed to avoid the risk of key leakage caused by resident keys.
[0087] It should be noted that: session keys are automatically updated every hour to prevent the risk of leakage due to long-term use; asymmetric key pairs are rotated every 30 days, and expired certificates are managed in conjunction with the certificate revocation list; and a hardware security module is used to store master keys to isolate keys from business logic.
[0088] The digital twin simulation module generates a multidimensional data state vector model based on a lightweight modeling engine, integrates a spatiotemporal data fusion algorithm to dynamically update the twin state, introduces pre-disaster conditions for virtual-reality interaction testing, and uses the Monte Carlo method to simulate the execution effect of data recovery plans under different disaster conditions, thereby simulating the spread of disasters.
[0089] Lightweight modeling engine construction:
[0090] The 3D geometric model is mesh-simplified using the QEM algorithm, employing a hierarchical optimization strategy to significantly reduce the complexity of the digital twin model. The number of facets is reduced by 70% through iterative shrinkage of vertex pairs. The simplified model is encoded using Draco compression technology, achieving a compression rate exceeding 80%. For dynamic parameters, an LSTM prediction network is integrated to generate time-series state vectors, and the LBFGS quasi-Newton method is used to optimize vertex positions and minimize surface energy to improve smoothness. Model consistency is verified using the Metro mesh algorithm, calculating the maximum geometric error and root mean square error between the 3D geometric model and the mesh-simplified model.
[0091] Maximum geometric error:
[0092] The maximum single-point geometric deviation of corresponding points between the original model and the simplified model is usually measured by Hausdorff distance, and the formula is:
[0093] ;
[0094] Where H(A,B) represents the Hausdorff distance, which is used to measure the maximum single-point geometric deviation between two point sets A and B; A and B are the vertex sets of the original model and the simplified model respectively; is the Euclidean distance between vertices; a and b are points in set A and set B respectively; inf is the lower bound, which refers to the minimum value in the set; sup is the upper bound, which refers to the maximum value in the set; by traversing all vertices and calculating the maximum deviation value, verify whether it satisfies requirements;
[0095] Root mean square error:
[0096] By measuring the average statistical characteristics of the overall geometric deviation of the model, the formula is:
[0097] ;
[0098] Where RMSE is the root mean square error; N is the total number of vertices; and are the corresponding vertices of the original model and the simplified model respectively; N is the total number of vertices; i is the current vertex, i=1, 2, 3..., N; the error reflects the square mean of the deviations of all vertices and must satisfy The relative error range of the parameters is 0.001; parameter lightweighting is achieved by screening key parameters through gradient evaluation and channel gain analysis;
[0099] Generate a multidimensional data state vector model:
[0100] By deploying edge sensors and cloud platform API interfaces, physical environment data, business data, and network topology data can be acquired in real time, and data quality can be improved through multi-stage processing;
[0101] Physical environment data includes computer room location, rack layout, and UPS current fluctuation data;
[0102] Business data includes the timestamp sequence of backup logs and the heat map of data access frequency;
[0103] Network topology data: Dynamically captures storage node IP distribution maps and link delay matrices based on the SDN controller;
[0104] Combine time window sliding alignment with asynchronous data streams, eliminate dimensional differences through Z-score normalization, apply PCA dimensionality reduction to reduce redundant features, and improve the efficiency of subsequent model training. For high-dimensional time series data, differential coding and AI prediction compression technology are used to achieve a 40% compression rate increase. AI prediction compression model selection and training:
[0105] Select an appropriate AI model, such as a neural network-based prediction model, and train it using historical high-dimensional time series data. The model's goal is to predict the data value at the next time step.
[0106] Compression operation: Based on the prediction, only the difference between the predicted value and the actual value is stored (similar to the idea of differential encoding) and the relevant parameters of the model (such as the weights of the neural network). In this way, the compression rate can be improved by 40%;
[0107] Because predicted values and actual values often have a certain correlation, the difference is relatively small, which takes up less storage space. By integrating geographic information with semantic understanding and GIS technology, data compliance and consistency are addressed. The processed data is encapsulated in standardized JSON / Protobuf format and transmitted to the cloud via TLS 1.3 encryption, meeting real-time and security requirements.
[0108] Update the twin status:
[0109] The pre-processed and registered multi-source data is fused. A spatiotemporal data fusion algorithm is used for data from different sources to obtain a target state close to the actual situation. Based on the fusion results, a lightweight modeling engine is used to calculate and update the state vector model of the digital twin. To make the digital twin more accurately reflect the actual situation, a feedback loop is established to continuously monitor changes in the actual system and feed new observations back to the digital twin, thereby continuously adjusting its internal state and achieving dynamic updates.
[0110] Simulate the execution effect of data recovery plans under different disaster conditions:
[0111] Generate random combination scenarios using the Monte Carlo method; simulate the data island effect caused by physical damage to the server and deduce the feasibility of cross-cabinet data migration paths; estimate the amount of data loss, based on the backup cycle, deduplication threshold, and recovery time confidence interval, and consider factors related to the backup cycle in the process of generating quantitative indicators based on the backup cycle, deduplication threshold, and recovery time confidence interval; determine the appropriate backup cycle based on system characteristics and requirements. For example, critical business data may need to be backed up daily or even more frequently, while the cycle can be appropriately extended for relatively static data; within each backup cycle, monitor and record data changes and calculate the amount of new, modified, or deleted data to analyze the impact of the backup cycle on data recovery and understand the data loss risk and resource consumption status under different cycles;
[0112] The deduplication threshold and recovery time confidence interval are incorporated into the analysis. After setting the deduplication threshold, deduplication is performed during backup and the deduplication effect is calculated. The impact on recovery time and data loss is evaluated to find the appropriate balance point for the deduplication threshold. At the same time, the recovery time confidence interval is defined. Time data for multiple simulated data loss and recovery is collected, the mean and standard deviation are calculated, and the confidence interval range is determined. This is used to evaluate system performance and construct a comprehensive quantitative indicator that includes the weights of each factor. The system strategy is adjusted based on the indicator, that is, the Monte Carlo simulation results and quantitative indicators are generated. A feedback loop that continuously monitors the actual system is implemented, and new observation data is continuously input into the digital twin to adjust its internal state in a timely manner.
[0113] For example:
[0114] A data center's main server was damaged by a fire, and a digital twin simulation was needed to evaluate the data recovery plan. The system collected real-time physical parameters of the damaged server, including temperature, smoke concentration, power supply status, and hard drive temperature. The system obtained server operation logs and retrieved past records of similar fire incidents. Using a lightweight modeling engine, the data was integrated into a multidimensional state vector. The sensor's real-time temperature data, such as "the computer room temperature suddenly rose to 80°C, timestamp 2023-10-01 14:00," was time-aligned with the "storage I / O interruption" event in the business log, timestamp 2023-10-01 14:05, and associated with the same spatial location, such as "server cabinet 3 in area C of the computer room."
[0115] Based on historical fire data and real-time sensor data, the server's physical state changes are predicted, such as "the hard drive is expected to fail due to high temperature in 10 minutes." Combined with incremental backup cycles and deduplication strategies, such as "data blocks with a duplication rate ≥ 80% are not stored," the amount of unbacked-up data is calculated. For example, if 20% of the new data added in the last hour is unique and needs to be restored, conditions are injected into the twin to simulate a real disaster environment:
[0116] Scenario 1: Assume that the fire spreads faster and backup servers need to be started across regions;
[0117] Result: Data recovery time was extended to 4 hours, but only 5% of the key business data and database master tables were lost due to deduplication technology;
[0118] Scenario 2: Assume that the network bandwidth is insufficient and the incremental backup is interrupted;
[0119] Results: Data loss increased to 15% when restoring from a full backup, but system availability recovery time was reduced to 2 hours; Data loss: 5%-15%, adjusted according to the backup strategy;
[0120] Recovery time objective: 2-4 hours, depending on network and backup integrity;
[0121] System availability: 90%, with critical services restored first and secondary services restarted later;
[0122] The reinforcement learning optimization module uses a deep reinforcement learning model, takes as input multi-dimensional evaluation indicators, historical fault data, and service logs, and outputs the optimal strategy, including backup frequency, deduplication threshold, and recovery priority.
[0123] The 10-dimensional evaluation vector, fault scalar value, and LSTM hidden state vector generated by the autoencoder are concatenated to form a comprehensive state vector, which integrates static indicators, dynamic time series, and historical risks, providing multi-dimensional input for the digital twin. The current policy, such as equipment scheduling rules and disaster recovery plans, is loaded into the digital twin environment, and the system operation is simulated based on the comprehensive state vector. During the deduction process, trajectory data is generated in real time, recording the state, action, and reward signal of each step, and constructing the interaction sequence required for reinforcement learning. The trajectory data is analyzed using generalized advantage estimation to calculate the advantage value of each action step. GAE balances the estimated bias and variance by weighted averaging short-term returns and long-term value, and quantifies the direction of strategy improvement.
[0124] A clipping objective function is set to limit the magnitude of policy updates to avoid excessive deviation from the current policy. The objective function is maximized through gradient ascent to ensure that the policy is iteratively optimized within a stable range. The network parameters of the value function are updated to more accurately predict the state value, that is, the expectation of long-term cumulative rewards. The update process combines temporal difference error and trajectory data to improve the accuracy of the value assessment of the system state and provide a benchmark for subsequent policy optimization.
[0125] For example:
[0126] Optimize based on simulation results:
[0127] If the evaluation indicator shows "high temperature causes hard drive failure time to be shorter than the backup window," we recommend: shortening the incremental backup cycle from 1 hour to 30 minutes; adding off-site backup nodes to reduce network latency risks; otherwise, no action is taken;
[0128] If the deduplication threshold is too high, resulting in critical data loss, the threshold is adjusted to 70%, balancing storage efficiency and data integrity. The optimized recovery process is quickly initiated based on the digital twin's preview results. For example, the core database is restored first: the most recent full backup (24 hours ago) and the latest incremental backup (1 hour ago) are used, combined with deduplication technology, to control data loss to less than 5%. Resource allocation is also dynamically adjusted, reserving backup transmission channels in advance based on the twin's predicted "network bandwidth fluctuations." Otherwise, no action is taken.
[0129] Load the current system state in the digital twin environment, inject disaster parameters, and perform K independent simulations for each scenario, randomly generating disturbance parameters each time. Example code snippet (pseudocode):
[0130] ```Python
[0131] for strategy in strategy_library:
[0132] for i in range(1000):
[0133] # Generate random perturbations
[0134] network_delay = base_delay * (1 + np.random.uniform(-0.2, 0.2))
[0135] resource_available = np.random.choice([True, False], p=[0.9, 0.1])
[0136] # Execute the simulation and record the results
[0137] result = digital_twin.simulate(strategy, network_delay, resource_available)
[0138] save_to_database(result)
[0139] ```
[0140] Record the detailed event chain for each simulation, such as "failure occurs at 2:00 PM → Plan A is initiated at 2:05 PM → recovery is completed at 2:30 PM." When the deviation between the actual recovery result and the simulated value exceeds a threshold, such as actual RTO minus predicted RTO > 15%, trigger model calibration. Update the device failure parameters in the digital twin, adjusting the hard drive MTBF from 100,000 hours to 80,000 hours.
[0141] The strategy execution and feedback module dynamically allocates emergency resources through a two-layer network based on the strategy generated by reinforcement learning, and optimizes the data transmission path using a path planning algorithm;
[0142] Deployed in the cloud, it is responsible for global policy analysis and resource status monitoring. It receives policy parameters output by the reinforcement learning module in real time, distributes them across edge nodes and core data centers, executes specific operational instructions, supports software-defined networking technology, and dynamically configures network paths through the OpenFlow protocol. The control layer detects the connection status of network devices, switches, and storage nodes through the LLDP protocol, builds a real-time topology map, and creates independent virtual network slices based on task types, including rescue slices and archive slices.
[0143] Policy parameters may include weights for allocating network bandwidth and computing resources to different task types, such as rescue and archiving tasks. These parameters are mapped to a predefined global policy framework. When resource status monitoring indicates that the computing resources of an edge node are nearing saturation, and the reinforcement learning module outputs a policy parameter that increases resource allocation for a certain type of task on that node, the system will adjust accordingly.
[0144] Edge nodes:
[0145] After receiving the policy parameters output by the reinforcement learning module, specific operation instructions will be executed according to these parameters. For example, if the policy parameters indicate that the data read and write priority of a certain type of storage node should be increased, the relevant controller in the edge node will adjust the read and write queue management policy of the storage device based on this parameter;
[0146] For networked devices, if the policy parameters involve adjusting the weight of the network path, the network controller in the edge node will send instructions to network devices such as switches through the OpenFlow protocol based on this weight parameter, dynamically changing the configuration of the network path to meet task requirements;
[0147] Core Data Center Operations:
[0148] In the core data center, policy parameters may be used to adjust large-scale data processing and storage strategies. For example, if a policy parameter indicates that more storage space needs to be reserved for archive slices, the core data center management module will replan the allocation of storage resources based on this parameter, which may involve operations such as reconfiguring the disk array or migrating data.
[0149] To create an independent virtual network slice:
[0150] Policy parameters may dictate that data transmission for a rescue slice must pass through specific low-latency links. For example, when the load on a link exceeds the threshold set by the policy parameters, the controller will redirect some traffic to other links using the OpenFlow protocol. Policy parameters play a key role in creating independent virtual network slices (such as rescue slices and archive slices). For rescue slices, policy parameters may specify requirements such as higher bandwidth and lower latency. Based on these parameters, the SDN controller will configure network devices using the OpenFlow protocol when creating the slice to ensure that the slice meets mission requirements. For different types of slices, policy parameters may also include security policies, quality of service (QoS), and other aspects.
[0151] Rescue slice: dedicated channel for high-priority data recovery;
[0152] Archive slice: a cold data transmission channel that allows bandwidth preemptive scheduling;
[0153] Through the path planning algorithm, the network status is perceived in real time, and the link indicators, bandwidth utilization, and delay are collected. , packet loss rate ; Construct a dynamic weight graph, the formula is:
[0154] ;
[0155] Where u and v are two different nodes used to construct a dynamic weight graph, which reflects the connection strength or cost between two nodes in the network; α, β, and γ are weight coefficients of the influence of bandwidth utilization, delay, and packet loss rate. , and dynamically adjust according to the strategy type; represents the bandwidth utilization of the link; τ is the delay of the link; is the packet loss rate of the link;
[0156] For example:
[0157] Path optimization example:
[0158] Task: Transfer 1TB of disaster recovery data from the G1 site center to the G4 site archive center;
[0159] Initial route: G1 → G2 → G3 → G4, estimated time: 2.5 hours;
[0160] Dynamic adjustment: Node G2 experiences sudden congestion, with bandwidth utilization > 90%. Switching to the path G1, G5, and G4 takes 2.1 hours. Actual RTO is recorded with second-level accuracy. Resource utilization, including CPU, storage, and bandwidth, is also measured. Path switching times are also analyzed.
[0161] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.
[0162] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.
[0163] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A data rescue and archiving system based on a cloud platform, characterized in that: The system includes: The data perception and collection module deploys edge computing nodes to collect real-time operational data, including sensor data, business logs, and historical fault data. It uses a hybrid encryption technology combining symmetric and asymmetric encryption to encrypt the transmission link. It uses the Zstandard algorithm and AI predictive coding to compress the data in two stages, and then transmits it to the cloud through a unified data interface. The digital twin simulation module generates a multidimensional data state vector model based on a lightweight modeling engine, integrates a spatiotemporal data fusion algorithm to dynamically update the twin state, introduces pre-disaster conditions for virtual-reality interaction testing, and uses the Monte Carlo method to simulate the execution effect of data recovery plans under different disaster conditions, thereby simulating the spread of disasters. The reinforcement learning optimization module uses a deep reinforcement learning model. Its input includes multi-dimensional evaluation indicators, historical fault data and business logs, and it outputs the optimal strategy. The output of the optimal strategy includes: splicing the 10-dimensional evaluation vector generated by the autoencoder, the fault scalar value, and the LSTM hidden state vector to form a comprehensive state vector, integrating static indicators, dynamic time series and historical risks to provide multi-dimensional input for the digital twin; loading the current strategy in the digital twin environment and simulating the system operation based on the comprehensive state vector; generating trajectory data in real time during the deduction process, recording the state, action and reward signal of each step, and constructing the interaction sequence required for reinforcement learning; using generalized advantage estimation to analyze the trajectory data and calculate the advantage value of each action step. GAE balances the estimated deviation and variance by weighted average short-term returns and long-term value, and quantifies the direction of strategy improvement; The strategy execution and feedback module dynamically allocates emergency resources through a two-layer network based on the optimal strategy generated by reinforcement learning, and uses a path planning algorithm to optimize the data transmission path.
2. The cloud platform-based data rescue and archiving system according to claim 1, characterized in that: The collecting operation data includes: Collect sensor data through IoT devices or sensor networks; Convert the records of the application during the execution of its functions into business logs; Through the maintenance management system, detailed information on system or equipment failures is collected and recorded as historical failure data.
3. The cloud platform-based data rescue and archiving system according to claim 2, characterized in that: The hybrid encryption technology encrypts the transmission link, including: S101: Preprocessing the operation data; inserting time points to obtain time series operation data; S102: Use the Zstandard algorithm to perform preliminary compression on the preprocessed time series running data; S103: Based on the initial compression of Zstandard, the machine learning model is used to run the input time series data and output patterns and regularities. Through AI predictive coding, the patterns and regularities in the data are used to predict future data points. S104: Combining the result of Zstandard compression with the output generated by AI predictive coding for optimization.
4. The cloud platform-based data rescue and archiving system according to claim 3, characterized in that: The predicted future data points include: S103-1: Divide the Zstandard compressed data stream into time windows and extract multi-dimensional features, including statistical features, frequency domain features, and pattern markers; S103-2: Build a dynamic context-aware model, maintain the state vector in real time, and save the evolution pattern of the first Z data points; S103-3: Adopting adaptive quantization residual coding, dynamically adjusting the quantization step size, establishing a parameter update dictionary, and transmitting only non-zero parameter gradient blocks; S103-4: Cross-sensor collaborative prediction, spatial correlation matrix maintenance, and calculation of spatial correlation coefficients; S103-5: Maintain the candidate model pool, switch models based on online evaluation indicators, perform context-adaptive Range encoding on the prediction residuals, use the context model, predict the confidence interval, and quantize the step size level.
5. The cloud platform-based data rescue and archiving system according to claim 1, characterized in that: The method of generating a multi-dimensional data state vector model based on a lightweight modeling engine includes: The QEM algorithm is used to simplify the mesh of the 3D geometric model. A hierarchical optimization strategy is used to iteratively shrink vertex pairs to reduce the number of facets. The simplified model is encoded using Draco compression technology. For dynamic parameters, an LSTM prediction network is integrated to generate a time series state vector. The LBFGS quasi-Newton method is used to optimize vertex positions and minimize surface energy. The model is verified using the Metro mesh algorithm, and the maximum geometric error and root mean square error between the 3D geometric model and the mesh simplified model are calculated. By deploying edge sensors and cloud platform API interfaces, physical environment data, business data, and network topology data can be acquired in real time and processed in multiple stages.
6. The cloud platform-based data rescue and archiving system according to claim 5, characterized in that: The integrated spatiotemporal data fusion algorithm dynamically updates the twin state, including: The multi-source data after preprocessing and alignment are fused. A spatiotemporal data fusion algorithm is used for data from different sources to obtain the target state. Based on the fusion results, a lightweight modeling engine is used to calculate and update the state vector model of the digital twin. A feedback loop is established to continuously monitor changes in the actual system and feed back new observation results to the digital twin, constantly adjusting its internal state and achieving dynamic updates.
7. The cloud platform-based data rescue and archiving system according to claim 6, characterized in that: The simulation of the execution effect of the data recovery plan under different disaster conditions includes: The Monte Carlo method is used to generate random combination scenarios, simulate the data island effect caused by physical damage to the server, and deduce the feasibility of cross-cabinet data migration paths. By estimating the amount of data loss, quantitative indicators are generated based on the backup cycle, deduplication threshold, and recovery time confidence interval.
8. The cloud platform-based data rescue and archiving system according to claim 1, characterized in that: The optimizing of the data transmission path by using a path planning algorithm includes: The network path is dynamically configured through the OpenFlow protocol. The control layer detects network devices through the LLDP protocol and builds a real-time topology map. Independent virtual network slices are created according to the task type, including rescue slices and archive slices. The parameters are mapped to the predefined global policy framework. The resource allocation of the task is adjusted through the policy parameters output by the reinforcement learning module.
Citation Information
Patent Citations
Operation monitoring method and system based on digital twin technology
CN118690546A
Policy selection method for node backup in multi-domain network
CN119316292A
Terminal data automatic backup and recovery method and system
CN119668939A