Data processing method and system for cloud computing and storage medium

By building a multi-level heterogeneous fusion network and adaptive threshold algorithm in a cloud computing environment, combined with the multi-objective decision-making process, the complex anomaly identification and disaster recovery automation problems in the existing technology are solved, and efficient and intelligent data processing and disaster recovery response are achieved.

CN120256196APending Publication Date: 2025-07-04BEIJING JIATU ZHIFENG TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 17 Cited by

Patent Information

Application Number
CN202510326581.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In existing cloud computing environments, data processing methods are difficult to effectively identify complex cross-system and cross-level exceptions. Static rules and fixed thresholds cannot adapt to dynamic changes, resulting in high false alarm rates or underreporting important abnormalities. Disaster recovery plans lack automated decision-making capabilities and continuous optimization mechanisms, which affect data processing efficiency and reliability.

Method used

By building a multi-level heterogeneous fusion network, mining data association relationships, combining adaptive threshold algorithms for abnormal detection and prediction, using a multi-objective decision-making process to determine the disaster recovery plan, and introducing a continuous learning optimization mechanism to achieve automated and intelligent disaster recovery response.

Benefits of technology

It significantly improves the intelligence level of data processing, timely disaster recovery response and system operation stability in the cloud computing environment, reduces the false alarm rate and missed alarm rate, realizes the scientific and automated disaster recovery decisions, and reduces manual intervention and response delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256196A_ABST
    Figure CN120256196A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a data processing method and system for cloud computing and a storage medium. The method comprises the following steps: acquiring preprocessed data through a distributed acquisition node to obtain a structured data packet; performing classification analysis through mixed feature extraction to obtain a safety disaster recovery data set; analyzing the incidence relation through a heterogeneous fusion network to generate a knowledge base; detecting abnormity through a self-adaptive threshold algorithm to form a risk library; determining a recovery scheme through a multi-objective decision to generate an execution log; and a protection strategy is formed through multi-angle comparison and optimization of parameters. According to the method, the data association relationship is mined and analyzed by constructing the multi-level heterogeneous fusion network, the optimal disaster recovery scheme is automatically determined through the multi-target decision process, and meanwhile, a continuous learning optimization mechanism is introduced; the intelligent level of data processing in the cloud computing environment, the timeliness of disaster recovery response and the stability of system operation are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a data processing method, system, and storage medium for cloud computing. Background Art

[0002] With the rapid development of cloud computing technology, more and more enterprises have migrated their business systems and data to the cloud environment. In the cloud computing environment, data processing faces challenges such as large scale, heterogeneous sources, and complex structures. Traditional data processing methods mainly rely on static rules and fixed thresholds for monitoring and alerting, solve abnormal problems through manual analysis, and perform disaster recovery and restoration operations in an experience-driven manner. Some existing technologies attempt to introduce automated tools for data analysis and anomaly detection, but most are limited to data processing in specific fields or single dimensions, lacking the comprehensive processing ability for multi-source heterogeneous data in the cloud environment. In addition, existing disaster recovery solutions usually adopt preset strategies and are difficult to dynamically adjust according to actual situations, with limited disaster recovery efficiency and accuracy.

[0003] However, these existing technologies have obvious deficiencies. First, in the face of a large amount of multi-source heterogeneous data in the cloud environment, traditional single-dimensional processing methods are difficult to effectively identify complex associated anomalies, especially cross-system and cross-level anomaly patterns. Second, static rules and fixed thresholds cannot adapt to the dynamic change characteristics of the cloud environment, resulting in a high false alarm rate or missed reporting of important anomalies. Third, existing disaster recovery solutions often lack the ability of automated decision-making and rely on manual judgment, causing response delays. Finally, existing technologies generally lack a continuous optimization mechanism and cannot learn and improve from historical experience, making it difficult to improve the system performance and security protection ability. These problems seriously affect the efficiency and reliability of data processing in the cloud computing environment, and may lead to serious business interruptions and data losses, especially in critical business scenarios. Summary of the Invention

[0004] This application provides a data processing method, system, and storage medium for cloud computing, which is used to mine and analyze data association relationships by constructing a multi-level heterogeneous fusion network, realize accurate anomaly detection and prediction in combination with an adaptive threshold algorithm, automatically determine the optimal disaster recovery plan through a multi-objective decision-making process, and introduce a continuous learning and optimization mechanism, significantly improving the intelligent level of data processing, the timeliness of disaster recovery response, and the stability of system operation in the cloud computing environment.

[0005] In a first aspect, the present application provides a data processing method for cloud computing. The data processing method for cloud computing includes: collecting and preprocessing the operation logs and network traffic of a multi-source data center through distributed collection nodes to obtain structured data packets; performing intelligent classification and parsing on the data through hybrid feature extraction according to the structured data packets to obtain security domain and disaster recovery domain data sets; mining and analyzing the data association relationships through a multi-level heterogeneous fusion network based on the security domain and disaster recovery domain data sets to generate a data disaster recovery association knowledge base; performing anomaly detection and prediction on the cloud environment security status through an adaptive threshold algorithm based on the data disaster recovery association knowledge base to form a security risk event library; determining a disaster recovery plan through a multi-objective decision-making process according to the security risk event library, implementing an automatic response mechanism, and generating disaster recovery execution logs; adjusting and optimizing the disaster recovery, system, and storage medium parameters through a multi-angle comparison process based on the disaster recovery execution logs to form a disaster recovery and security protection optimization strategy.

[0006] In a second aspect, the present application provides a data processing system for cloud computing. The data processing system for cloud computing includes:

[0007] A processing module for collecting and preprocessing the operation logs and network traffic of a multi-source data center through distributed collection nodes to obtain structured data packets;

[0008] A classification module for performing intelligent classification and parsing on the data through hybrid feature extraction according to the structured data packets to obtain security domain and disaster recovery domain data sets;

[0009] An analysis module for mining and analyzing the data association relationships through a multi-level heterogeneous fusion network based on the security domain and disaster recovery domain data sets to generate a data disaster recovery association knowledge base;

[0010] A prediction module for performing anomaly detection and prediction on the cloud environment security status through an adaptive threshold algorithm based on the data disaster recovery association knowledge base to form a security risk event library;

[0011] An implementation module for determining a disaster recovery plan through a multi-objective decision-making process according to the security risk event library, implementing an automatic response mechanism, and generating disaster recovery execution logs;

[0012] An adjustment module for adjusting and optimizing the disaster recovery, system, and storage medium parameters through a multi-angle comparison process based on the disaster recovery execution logs to form a disaster recovery and security protection optimization strategy.

[0013] In a third aspect of the present invention, a computer device is provided, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory to cause the computer device to execute the above-mentioned data processing method for cloud computing.

[0014] In a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the above-mentioned data processing method for cloud computing.

[0015] In the technical solution provided by this application, the operation logs and network traffic of multi-source data centers are collected and preprocessed by distributed collection nodes, achieving comprehensive coverage and preliminary purification of data sources, and significantly improving the data quality for subsequent analysis; the data is intelligently classified and parsed through hybrid feature extraction based on structured data packets, effectively identifying and extracting key features, reducing the data dimension and computational complexity, and enhancing the discriminability and expression ability of features; based on the data sets of the security domain and the disaster recovery domain, the mining analysis of data association relationships is carried out through a multi-level heterogeneous fusion network, breaking through the limitations of traditional single-dimensional analysis, capturing complex association patterns across levels and domains, and improving the accuracy and comprehensiveness of anomaly detection; based on the data disaster recovery association knowledge base, the anomaly detection and prediction of the cloud environment security state are carried out through an adaptive threshold algorithm, dynamically adjusting the judgment criteria to adapt to system load changes, while taking into account time sensitivity, significantly reducing the false alarm rate and missed alarm rate; according to the security risk event library, the disaster recovery plan is determined through a multi-objective decision-making process, comprehensively considering multiple objectives such as system availability, performance impact, security risk, and recovery time, realizing the scientific and automated disaster recovery decision-making, significantly reducing manual intervention and response delay; based on the disaster recovery execution logs, the disaster recovery, system, and storage medium parameters are adjusted and optimized through a multi-angle comparison process, forming a closed-loop optimization mechanism, continuously improving the system performance and protection ability. In particular, this solution applies a number of artificial intelligence algorithms in the cloud computing environment. Among them, the multi-level heterogeneous fusion network algorithm plays a key role in data association analysis. Through different levels of network structures and inter-layer mappings, it effectively captures the complex direct and indirect associations between data, providing a comprehensive three-dimensional association perspective for anomaly detection; the adaptive threshold algorithm adjusts the judgment criteria in real time according to the dynamic characteristics of the cloud environment, avoiding the limitations of traditional fixed thresholds; the multi-objective decision-making algorithm can balance multiple decision-making objectives in a complex and changeable cloud environment, find the optimal disaster recovery plan, and adapt to the complexity and multi-objectivity of disaster recovery decision-making in the cloud computing environment; the Bayesian optimization algorithm guides the parameter search through a probability model, efficiently finding the optimal parameter configuration of the system, providing a theoretical support for the continuous optimization of the system. The organic combination of these algorithms significantly improves the intelligent level of data processing, the timeliness of disaster recovery response, and the stability of system operation of this solution in the cloud computing environment. Brief Description of the Drawings

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a schematic diagram of an embodiment of a data processing method for cloud computing in an embodiment of the present application;

[0018] Figure 2 It is a schematic diagram of an embodiment of a data processing system for cloud computing in an embodiment of the present application;

[0019] Figure 3 It is a schematic block diagram of the structure of a computer device in an embodiment of the present invention. Detailed Embodiments

[0020] The embodiments of the present application provide a data processing method, system, and storage medium for cloud computing. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, storage medium, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0021] For ease of understanding, the following describes the specific process of the embodiments of the present application. Please refer to Figure 1 , an embodiment of the data processing method for cloud computing in an embodiment of the present application includes:

[0022] Step S101: Collect and preprocess the operation logs and network traffic of the multi-source data center through distributed collection nodes to obtain structured data packets;

[0023] Step S102: Intelligently classify and analyze the data through hybrid feature extraction according to the structured data packets to obtain the security domain and disaster recovery domain data sets;

[0024] Step S103: Based on the datasets of the security domain and the disaster recovery domain, mining and analyzing the data association relationships through a multi-level heterogeneous fusion network to generate a data disaster recovery association knowledge base;

[0025] Step S104: Based on the data disaster recovery association knowledge base, performing anomaly detection and prediction on the security status of the cloud environment through an adaptive threshold algorithm to form a security risk event library;

[0026] Step S105: Determining a disaster recovery plan through a multi-objective decision-making process based on the security risk event library, implementing an automatic response mechanism, and generating a disaster recovery execution log;

[0027] Step S106: Adjusting and optimizing the disaster recovery, system, and storage medium parameters through a multi-angle comparison process based on the disaster recovery execution log to form a disaster recovery and security protection optimization strategy.

[0028] It can be understood that the execution subject of this application can be a data processing, system, and storage medium for cloud computing, or a terminal or a server. Specifically, it is not limited here. This application embodiment is described by taking the server as the execution subject as an example.

[0029] Specifically, the operation logs and network traffic of multiple source data centers are collected and preprocessed through distributed collection nodes. Distributed collection nodes refer to data collection units distributed at different physical locations. These nodes are respectively deployed at the edge positions of different data centers and adopt an adaptive bandwidth control mechanism to dynamically adjust the data transmission rate according to the network conditions. The types of collected data include server operation logs, network traffic data, user behavior data, and application performance indicators. These raw data are processed through a two-layer filtering mechanism. The first layer uses a rule engine to quickly screen out obvious anomalies and redundant data, and the second layer identifies potential anomaly points through an outlier identification algorithm. Subsequently, the data is calibrated and aligned in terms of timestamps to ensure the consistency of data from different sources in the time dimension. Then, the heterogeneous data is converted into a unified format through a format converter, metadata tags (including data source identification, time stamp, service type, and security level) are added, and finally, an integrity verification code is generated and the data body, metadata tags, and verification code are encapsulated to form a structured data packet.

[0030] Intelligent classification and parsing of data are carried out based on structured data packets through hybrid feature extraction. Hybrid feature extraction is a feature extraction technology that combines statistical methods and deep learning. It extracts time features, space features, service features, and security features from structured data packets to form a feature vector set. These feature vectors are processed by dimensionality reduction to remove redundant features and enhance discrimination, resulting in an optimized feature set. The optimized feature set is input into a multi-layer classifier. First, it is initially divided by a decision tree, and then refined classification is performed by a support vector machine. The data is divided into a computing resource domain, a storage resource domain, a network resource domain, and a security domain. Data items related to disaster recovery are extracted from the security domain data to form disaster recovery domain data. Finally, the security domain and disaster recovery domain data are merged and sorted to form a security domain and disaster recovery domain data set. Based on the security domain and disaster recovery domain data set, the data association relationship is mined and analyzed through a multi-level heterogeneous fusion network. The multi-level heterogeneous fusion network is a technology that represents the relationship of heterogeneous data sources as different-level network structures and captures complex indirect associations through cross-layer connections. The specific operations include constructing a three-layer heterogeneous network structure (entity layer, attribute layer, event layer), establishing connection relationships between nodes in the same layer through intra-layer association analysis for the three-layer heterogeneous network, establishing vertical connections between nodes in different layers through inter-layer mapping, then performing path analysis through an improved random walk algorithm to identify key paths and core nodes, extracting temporal association patterns from important associated subgraphs through temporal projection, and integrating and storing the intra-layer association graph, cross-layer connection graph, and temporal association rule set into a graph data structure to generate a data disaster recovery association knowledge base.

[0031] Based on the data disaster recovery association knowledge base, the abnormal detection and prediction of the cloud environment security state are carried out through an adaptive threshold algorithm. The adaptive threshold algorithm can dynamically adjust the determination threshold according to factors such as time period and business load. First, historical operation data is extracted from the data disaster recovery association knowledge base, and a benchmark behavior sequence is constructed through time window segmentation. The normal fluctuation range is calculated through statistical analysis of the benchmark behavior sequence to form a normal baseline profile. The determination threshold is dynamically adjusted based on the normal baseline profile to obtain an upper and lower threshold boundary table. According to the threshold boundary table, multi-level abnormal screening is performed on the real-time data stream. The abnormal point records are associated with the causal links in the data disaster recovery association knowledge base to trace the abnormal propagation path and root node. Finally, a risk assessment is performed on the abnormal event chain to form a security risk event library.

[0032] According to the security risk event library, determine the disaster recovery plan through a multi-objective decision-making process and implement an automatic response mechanism. The multi-objective decision-making process is a decision-making method that comprehensively considers multiple objectives. First, extract the abnormal event characteristics and impact scope information from the security risk event library to construct a decision knowledge graph. Quantify and assign weights to the four core objectives of system availability, performance impact, security risk, and recovery time through multi-objective weight assignment for the decision knowledge graph. Based on the objective weight matrix, comprehensively score the optional disaster recovery strategies to form a candidate set of disaster recovery plans. Select the non-dominated plans from the candidate set through Pareto optimal selection to obtain the disaster recovery plan. According to the disaster recovery plan, use a hierarchical response controller to classify the risk levels of operations to form an operation sequence, convert the operation sequence into four types of system instructions to record the execution status and effect feedback, and generate a disaster recovery execution log. According to the disaster recovery execution log, adjust and optimize the disaster recovery, system, and storage medium parameters through a multi-angle comparison process. The multi-angle comparison process is a method of comparing the actual effect with the expected goal, historical average level, and industry best practice. Extract the response operation records, execution status data, and effect feedback information from the disaster recovery execution log to construct an effect evaluation data set. Quantify and calculate the four key indicators of problem-solving rate, system recovery time, resource utilization efficiency, and user satisfaction for the effect evaluation data set to generate a performance indicator table. Through horizontal comparative analysis of the performance indicator table, identify the advantages and deficiencies to form a difference analysis report. Based on the difference analysis report, adjust and calculate the key parameters through the Bayesian optimization algorithm to obtain a parameter optimization plan. Supplement and correct the abnormal pattern library to construct optimized knowledge rules, integrate the parameter optimization plan and the optimized knowledge rules into an executable configuration instruction set, and form an optimized strategy for disaster recovery and security protection.

[0033] In the embodiments of the present application, the operation logs and network traffic of the multi-source data center are collected and preprocessed by distributed acquisition nodes, achieving comprehensive coverage and preliminary purification of data sources, and significantly improving the data quality for subsequent analysis; the data is intelligently classified and parsed through hybrid feature extraction based on structured data packets, effectively identifying and extracting key features, reducing the data dimension and computational complexity, and enhancing the discriminability and expression ability of features; based on the data sets of the security domain and the disaster recovery domain, the mining analysis of data association relationships is carried out through a multi-level heterogeneous fusion network, breaking through the limitations of traditional single-dimensional analysis, capturing complex association patterns across levels and domains, and improving the accuracy and comprehensiveness of anomaly detection; according to the data disaster recovery association knowledge base, the anomaly detection and prediction of the cloud environment security status are carried out through an adaptive threshold algorithm, dynamically adjusting the judgment criteria to adapt to system load changes, while taking into account time sensitivity, and significantly reducing the false alarm rate and missed alarm rate; according to the security risk event library, the disaster recovery plan is determined through a multi-objective decision-making process, comprehensively considering multiple objectives such as system availability, performance impact, security risk, and recovery time, realizing the scientific and automated disaster recovery decision-making, and significantly reducing manual intervention and response delay; according to the disaster recovery execution log, the disaster recovery, system, and storage medium parameters are adjusted and optimized through a multi-angle comparison process, forming a closed-loop optimization mechanism, and continuously improving the system performance and protection ability. In particular, this solution applies a number of artificial intelligence algorithms in the cloud computing environment. Among them, the multi-level heterogeneous fusion network algorithm plays a key role in data association analysis. Through different levels of network structures and inter-layer mappings, it effectively captures complex direct and indirect associations between data, providing a comprehensive three-dimensional association perspective for anomaly detection; the adaptive threshold algorithm adjusts the judgment criteria in real time according to the dynamic characteristics of the cloud environment, avoiding the limitations of traditional fixed thresholds; the multi-objective decision-making algorithm can balance multiple decision-making objectives in a complex and changing cloud environment, find the optimal disaster recovery plan, and adapt to the complexity and multi-objectivity of disaster recovery decision-making in the cloud computing environment; the Bayesian optimization algorithm guides parameter search through a probability model, efficiently finds the optimal parameter configuration of the system, and provides a theoretical support for the continuous optimization of the system. The organic combination of these algorithms significantly improves the data processing intelligence level, disaster recovery response timeliness, and system operation stability of this solution in the cloud computing environment.

[0034] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0035] (1) Obtain server operation logs, network traffic data, user behavior data, and application performance indicators from multiple data centers through edge perception collectors, and generate an original data stream;

[0036] (2) Clean the original data stream through a double-layer filtering mechanism. In the first layer, use a rule engine to eliminate obvious anomalies and redundant data. In the second layer, use an outlier identification algorithm to screen potential anomalies to obtain the cleaned data;

[0037] (3) Calibrate and align the cleaned data according to timestamps to construct data units with consistent time series;

[0038] (4) Standardize the data units with consistent time series through a format converter to convert heterogeneous data from different sources into a unified format;

[0039] (5) Add metadata tags to the data in the unified format, including data source identifier, time stamp, business type, and security level, to form tagged data;

[0040] (6) Generate a checksum for the tagged data through a checksum calculator, and encapsulate the data body, metadata tags, and checksum to obtain a structured data packet.

[0041] Specifically, obtain the original data from multiple data centers through an edge-aware collector. The edge-aware collector is a data collection device deployed at the edge of the data center network and has the ability to sense the network state and adaptively adjust the collection strategy. The collector obtains four types of key data from multiple data centers: server operation logs (including system-level metrics such as CPU usage, memory occupancy, and disk I / O), network traffic data (including network behavior metrics such as network throughput, connection count, and packet characteristics), user behavior data (including user interaction information such as access patterns, operation sequences, and authentication records), and application performance metrics (including application-layer metrics such as response time, transaction volume, and error rate). These data are collected in time series form and organized into an original data stream. The original data stream is a time-series, multi-source, heterogeneous data set that contains comprehensive information about the operation state of the cloud environment.

[0042] Data cleaning is performed on the original data stream through a two-layer filtering mechanism. The two-layer filtering mechanism is a hierarchical data purification method, which is divided into a rule filtering layer and a statistical filtering layer. The first-layer rule engine quickly screens the data based on a preset rule set, eliminating obvious anomalies and redundant data. The rule set includes data format verification rules, integrity check rules, value range rules, and known anomaly pattern rules. For example, records with a server CPU usage rate exceeding 100% or below 0% will be directly eliminated. The second layer uses an outlier identification algorithm to screen potential outliers. This algorithm identifies anomalies by calculating the deviation degree of data points from their distribution. Common outlier identification methods include statistics-based methods (such as Z-score, modified Z-score) and density-based methods (such as LOF, DBSCAN). The data after two-layer filtering is called cleaned data, which has higher quality and reliability. The cleaned data is calibrated and aligned according to the timestamp. Since there are time differences during multi-source data collection (such as different server clocks being out of sync, network delays, etc.), it is necessary to perform unified processing on all data in the time dimension. The timestamp calibration process includes three steps: clock deviation detection, standard time reference selection, and time offset compensation. The alignment process maps data from different sources and different sampling rates to a unified time scale, and uses interpolation or downsampling methods to handle the problem of inconsistent sampling rates. The data after time dimension processing forms a time-series consistent data unit, ensuring the correct expression of the time correlation between different data sources.

[0043] The time-series consistent data unit is standardized through a format converter. The format converter is a data processing component that can convert heterogeneous data into a unified format. It defines corresponding parsing rules and mapping relationships for different types of data sources, and converts various proprietary formats (such as JSON, XML, CSV, binary logs, etc.) into a unified internal data structure. The conversion process includes field extraction, data type conversion, missing value processing, and naming normalization. The standardized data adopts a unified structured format, which is convenient for subsequent analysis and processing. Metadata tags are added to the data in the unified format. Metadata tags are additional information that describes the characteristics of the data itself, including four core dimensions: data source identifier (recording the source device, system, or application of the data), time stamp (recording the exact time point when the data is generated and processed), business type (identifying the business area or functional module to which the data belongs), and security level (indicating the sensitivity of the data and the access control requirements). The process of adding metadata tags uses a tag allocator to add corresponding tag information to each piece of data according to predefined tag templates and tag rules, forming tagged data.

[0044] Generate a checksum for the tagged data through a checksum calculator and perform data encapsulation. A checksum is a cryptographic mechanism used to verify data integrity and is typically calculated using a hash function (such as MD5, SHA256, etc.) or a checksum algorithm. The checksum calculator calculates the checksum for the data body and metadata tags for integrity verification during subsequent transmission and storage. The data encapsulation process combines the data body, metadata tags, and checksum into a logical whole to form a structured data packet. This data packet is a self-describing and verifiable data unit that provides a standardized input for subsequent classification and parsing.

[0045] Taking the data processing of a certain cloud service provider as an example, when an anomaly in a critical business application is monitored, the edge perception collector will collect relevant data from multiple data centers. The collector obtains various types of data, including server logs containing anomaly alerts, DNS query traffic, user login records, and application response times. The original data stream is processed by a rule engine to eliminate incorrectly formatted log entries and duplicate network packet records. The outlier identification algorithm discovers that the CPU of some servers suddenly soars to 95%, far higher than the normal benchmark value of 40%, and marks these data points as potential anomalies. After timestamp calibration, it is found that the user login failure event occurred 3 minutes before the CPU anomaly, establishing an event time sequence relationship. The format converter converts logs and monitoring data in different formats into a unified JSON structure, adds metadata tags such as "data source: edge node 03", "time: 2025-01-15T08:23:47", "business type: identity authentication", "security level: high", etc., and finally generates a checksum through the SHA256 algorithm and encapsulates it to form a structured data packet, providing a standardized data basis for subsequent security event analysis.

[0046] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0047] (1) Extract time features, spatial features, business features, and security features from the structured data packet through a feature extraction engine to form a feature matrix;

[0048] (2) Perform dimensionality reduction on the feature matrix through a hybrid feature extraction combining statistics and deep learning to remove information redundancy and enhance feature distinguishability, obtaining a core feature set;

[0049] (3) Input the core feature set into a hierarchical classifier, and perform a coarse-grained division of the features through a decision tree to generate preliminary classification labels;

[0050] (4) Perform fine-grained classification on the preliminary classification labels through vector calculation, and divide the data into a computing resource segment, a storage resource segment, a network resource segment, and a security segment;

[0051] (5) Extract data items related to the disaster recovery function from the security segment through the disaster recovery correlation filter to form the disaster recovery segment data;

[0052] (6) Merge and standardize the data in the security segment and the disaster recovery segment data to generate the security domain and disaster recovery domain data sets.

[0053] Specifically, four types of key features are extracted from the structured data packet through the feature extraction engine. The feature extraction engine is a dedicated data analysis component that can identify and extract valuable features from complex data. The time feature refers to the time attribute of the data, including the occurrence time, duration, periodicity, and time sequence pattern, which are extracted through time window analysis and frequency analysis; the space feature refers to the location and topological relationship attributes of the data, including the physical location, network location, and resource distribution, which are extracted through topological analysis and location mapping; the business feature refers to the business process and functional attributes to which the data belongs, including the business type, transaction characteristics, and service quality indicators, which are extracted through business rule mapping and semantic analysis; the security feature refers to the security-related attributes of the data, including the access mode, abnormal features, and risk indicators, which are extracted through security mode recognition and risk assessment. These four types of features form a feature matrix through mathematical representation. Each row in the matrix represents a data instance, and each column represents a feature dimension.

[0054] Performing dimensionality reduction processing on the feature matrix is a key step in reducing data complexity. Hybrid feature extraction is a feature processing technology that combines statistical methods and deep learning. It identifies the main features through statistical methods such as principal component analysis (PCA) and removes redundant features with high linear correlation; then it captures the non-linear relationships between features through deep learning methods such as autoencoders and extracts more expressive hidden layer features. The hybrid feature extraction process can be expressed as:

[0055]

[0056] Among them, F matrix represents the original feature matrix, represents the feature transformation function of the statistical method, represents the feature transformation function of the deep learning method, F core represents the generated core feature set. In specific implementation, first merge the features with high correlation (correlation coefficient exceeding 0.85) in the feature matrix, then retain the principal components with an explained variance of 95% through principal component analysis, and finally further compress the feature dimensions through a three-layer autoencoder to obtain a core feature set with high discrimination.

[0057] The core feature set is then input into a hierarchical classifier for coarse-grained partitioning. The hierarchical classifier is a hierarchical classification architecture. In the first layer, the decision tree algorithm is used to perform a preliminary classification of the features. The decision tree selects the best splitting feature through information gain or Gini impurity, gradually constructs the tree structure, and divides the data into different categories. The splitting process of the decision tree follows the following rules: for feature x i and threshold t i , if x i <t i then the data is assigned to the left subtree, otherwise it is assigned to the right subtree. The optimal splitting point is selected based on the principle of maximizing information gain:

[0058]

[0059] where IG represents information gain, D represents the data set, H represents information entropy, and D j represents the sub-data set after splitting. By recursively constructing the decision tree, the core feature set is divided into several categories, forming preliminary classification labels.

[0060] Performing a fine classification on the preliminary classification labels is the key to achieving accurate resource classification. Vector calculation is a fine classification method based on the vector space model. It converts the preliminary classification labels into feature vectors and further subdivides them by calculating the similarity or distance between the vectors. Specifically, when implementing, the feature vectors are calculated for each preliminarily classified data

[0061]

[0062] where w c1 , w c2 represents the weight of the computing resource relevance, w s1 , w s2 represents the weight of the storage resource relevance, w n1 , w n2 represents the weight of the network resource relevance, w a1 , w a2 represents the weight of the security relevance. Through vector calculation, the membership degrees S comp of the four resource segments are obtained:

[0063]

[0064] The calculation of other resource segments is similar. The data is assigned to the resource segment with the highest membership degree, forming four data sets: the computing resource segment, the storage resource segment, the network resource segment, and the security segment.

[0065] Extracting disaster recovery related data from the security segment is the key to constructing the disaster recovery domain. The disaster recovery association filter is a dedicated data filtering tool that identifies data items related to disaster recovery functions through predefined disaster recovery feature identifiers and association rules. The filtering process includes keyword matching (such as "backup", "recovery", "disaster tolerance"), attribute association analysis (identifying data related to backup servers and storage devices), and functional relevance assessment (evaluating the degree of relevance of data to disaster recovery functions). Data items identified as highly relevant to disaster recovery are extracted to form the disaster recovery segment data.

[0066] Merge and standardize the security segment data and the disaster recovery segment data to generate a unified security domain and disaster recovery domain dataset. The merge and standardization process first unifies the formats and aligns the fields of the two types of data to ensure data structure consistency; then performs data deduplication to avoid the same information appearing repeatedly in the two data segments; finally, adds domain identifiers to clarify the security domain or disaster recovery domain to which the data belongs, forming a security domain and disaster recovery domain dataset with a standardized structure and clear meaning.

[0067] Taking the cloud platform of a financial institution as an example, the initial structured data packet contains server performance data, network traffic records, storage usage, and security alert logs. The feature extraction engine extracts time features (such as the time pattern of alert occurrence), space features (such as the network topology relationship of alert sources), business features (such as the affected business types), and security features (such as threat levels and attack features) from it to form a feature matrix with 80 dimensions. Through hybrid feature extraction, the matrix dimension is reduced to 25 key features. The decision tree of the hierarchical classifier makes a first split based on the feature of "whether data access is involved", initially dividing the data into data operation classes and non-data operation classes. Subsequently, vector calculation analysis finds that the correlation of a certain type of alert with computing resources is 0.2, with storage resources is 0.1, with network resources is 0.3, and with security is 0.4, so it is classified into the security segment. From the security segment, the disaster recovery association filter identifies alerts containing keywords such as "data backup failure" and "disaster recovery replication delay", and extracts monitoring data related to backup servers to form the disaster recovery segment data. The intrusion detection alerts in the security segment and the backup status data in the disaster recovery segment are merged and standardized to generate a security domain and disaster recovery domain dataset, providing a structured basis for subsequent correlation analysis.

[0068] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0069] (1) Organize the security domain and disaster recovery domain dataset into a network structure of different levels through a data layer constructor, where the first layer is the entity layer containing device nodes and business nodes, the second layer is the attribute layer containing security metrics and disaster recovery metrics, and the third layer is the event layer containing historical security events and disaster recovery events, forming a three-layer heterogeneous network structure;

[0070] (2) Establish the connection relationships between nodes in the same layer of the three-layer heterogeneous network structure through intra-layer correlation analysis. Among them, the entity layer is connected through topological relationships, the attribute layer is connected through correlation thresholds, and the event layer is connected through temporal dependencies to obtain an intra-layer correlation graph;

[0071] (3) Establish vertical connections between nodes at different levels from the intra-layer correlation graph through inter-layer mapping, and map the entity layer nodes to their corresponding attribute layer metrics and event layer records to form a cross-layer connection graph;

[0072] (4) Perform path analysis on the cross-layer connection graph through an improved random walk algorithm to identify critical paths and core nodes, and obtain an important associated subgraph;

[0073] (5) Extract the temporal correlation patterns between data from the important associated subgraph through temporal projection, including periodic patterns, trend patterns, and abnormal fluctuation patterns, to form a temporal correlation rule set;

[0074] (6) Integrate and store the intra-layer correlation graph, cross-layer connection graph, and temporal correlation rule set into a graph data structure to form a data disaster recovery association knowledge base.

[0075] Specifically, organize the security domain and disaster recovery domain data sets into a three-layer heterogeneous network structure through a data layer constructor. The data layer constructor is a processing component specifically used for multi-level data organization. It maps data in a hierarchical structure of entity-attribute-event. The entity layer is the bottom layer structure, including two types of nodes: physical and logical nodes: device nodes (such as physical resources like servers, storage devices, network devices, etc.) and business nodes (such as logical resources like application systems, databases, business processes, etc.). The attribute layer is the middle layer, including metric data describing entity characteristics: security metrics (such as security-related metrics like the number of vulnerabilities, frequency of security events, access control intensity, etc.) and disaster recovery metrics (such as disaster recovery-related metrics like recovery point objective, recovery time objective, data backup integrity rate, etc.). The event layer is the top layer, including historical event records: historical security events (such as security anomaly records like intrusion detection, malicious code infection, privilege abuse, etc.) and disaster recovery events (such as disaster recovery operation records like backup failure, disaster recovery switchover, data recovery, etc.). The data layer constructor assigns each piece of data to the corresponding layer according to its type and attributes, and constructs a heterogeneous network infrastructure containing three layers of nodes.

[0076] Perform intra-layer correlation analysis on the three-layer heterogeneous network structure with the aim of establishing connection relationships between nodes in the same layer. Different connection strategies are adopted in different layers for intra-layer correlation analysis: in the entity layer, connections are established through topological relationships, that is, connections between nodes are established based on physical network connections, deployment dependencies, and functional call relationships. For example, two servers are directly connected through the network or one application system depends on another database; in the attribute layer, connections are established through correlation thresholds. The correlation coefficient between different metrics is calculated, and a connection is established when the correlation exceeds a preset threshold (usually 0.7 or higher). For example, a connection is established when there is a high correlation between the CPU usage rate and the memory occupancy rate; in the event layer, connections are established through temporal dependencies. The time sequence and causal relationship of event occurrences are analyzed, and a connection is established when one event frequently appears before another event and the time interval is relatively fixed. For example, the backup failure event often appears after the storage device alarm. Through these three different connection strategies, an intra-layer correlation graph reflecting the relationships between nodes is formed within each layer. Starting from the intra-layer correlation graph, vertical connections between nodes at different levels are established through inter-layer mapping. Inter-layer mapping is a cross-layer data association technique that establishes connections between nodes at one level and related nodes at other levels. The specific implementation includes three mapping relationships: entity-to-attribute mapping, which connects each entity node to the attribute nodes describing its characteristics, such as connecting a server node to its attribute metrics such as CPU usage rate and memory occupancy rate; entity-to-event mapping, which connects entity nodes to historical events related to them, such as connecting a storage device to the disk failure events that have occurred to it; attribute-to-event mapping, which establishes a connection between attribute metrics and related events, such as connecting the backup success rate metric to the backup failure event. Through these vertical mapping relationships, a complete connection network spanning three levels is established, forming a cross-layer connection graph.

[0077] For the constructed cross-layer connection graph, an improved random walk algorithm is used for path analysis. The improved random walk algorithm is a data mining method based on graph structure, which discovers important nodes and paths by simulating the random walking process on the graph. The algorithm starts from a randomly selected starting node, uses the connection strength between nodes as the transition probability, moves along the connection to the next node, and records the access frequency. Different from the traditional random walk, the improved algorithm introduces hierarchical weights and node importance weights, which increases the probability of moving in the vertical direction and accessing key nodes. During the path analysis process, the access frequency of nodes and the passing frequency of paths are calculated. Nodes with high access frequency are identified as core nodes, and paths with high passing frequency are identified as key paths. Nodes and paths with access frequencies and passing frequencies exceeding the threshold are extracted to form an important association subgraph.

[0078] Extract temporal association patterns from important associated subgraphs through temporal projection. Temporal projection is an analytical method that maps graph-structured data onto the time dimension, revealing the dynamic relationships between data by extracting time-related patterns. Temporal projection first arranges the nodes with time attributes in the graph (mainly event-layer nodes) in chronological order, and then analyzes their occurrence patterns, identifying three main types of patterns: periodic patterns, which refer to combinations of events that occur regularly, such as backup operations at zero o'clock every day; trend patterns, which refer to data changes that show an increasing or decreasing trend over time, such as the gradual increase in system load; and abnormal fluctuation patterns, which refer to abnormal data changes that occur within a short period, such as a sudden surge in network traffic. These patterns are extracted from the time series through statistical analysis, frequency analysis, and rate-of-change calculations, forming a set of temporal association rules that describe the temporal relationships of the data.

[0079] Integrate the intra-layer association graph, cross-layer connection graph, and temporal association rule set and store them in the graph data structure. The graph data structure is a database technology specifically used for storing and processing relational data, suitable for expressing complex connection relationships. The integration process includes: unifying the formats of nodes and edges to ensure that data from different sources have a consistent structure; establishing an index system to improve query efficiency; setting up an attribute system to store the relevant attributes of nodes and edges; and defining relationship types to distinguish different types of connection relationships. The integrated data disaster recovery association knowledge base is a comprehensive knowledge base containing multi-level, multi-type nodes and relationships, capable of comprehensively reflecting the complex association relationships between security and disaster recovery in the cloud computing environment.

[0080] Taking the disaster recovery system in a certain enterprise's cloud environment as an example, when the data layer constructor processes the security domain and disaster recovery domain data sets, the servers and storage arrays in the primary and standby data centers are classified into the entity layer; the indicators such as data replication latency and backup integrity rate are classified into the attribute layer; and the event records such as historical network interruptions and backup failures are classified into the event layer. Intra-layer association analysis finds that the storage arrays in the primary and standby data centers are connected through replication links, establishing a connection in the entity layer; the correlation coefficient between data replication latency and network bandwidth utilization reaches 0.85, exceeding the threshold of 0.7, establishing a connection in the attribute layer; network interruption events often cause backup failure events within 30 minutes, establishing a temporal dependence connection in the event layer. Inter-layer mapping establishes a vertical connection between the storage array and its data replication latency indicator, and at the same time establishes an association with historical backup failure events. The improved random walk algorithm, through 10,000 iterations, finds that the core storage array in the primary data center is the node with the highest access frequency, and the path "storage array - replication latency - backup failure" has the highest access frequency. Temporal projection analysis finds that backup failure events show a periodic pattern with an increasing frequency at 0:00 on Mondays, and a trend pattern of a gradual increase in replication latency during the period from 17:00 to 19:00 on weekdays. Integrate these association relationships and patterns into the graph database.

[0081] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0082] (1) Extract historical operation data from the data disaster recovery association knowledge base, divide the continuous data into multiple time periods through time window segmentation, and construct a benchmark behavior sequence;

[0083] (2) Calculate the normal fluctuation range of each business indicator for the benchmark behavior sequence through statistical analysis, including mean, standard deviation, seasonal variation, and periodic trend, to form a normal baseline profile;

[0084] (3) Dynamically adjust the determination threshold based on the normal baseline profile through an adaptive threshold algorithm, where the threshold boundary is relaxed for high-load periods and tightened for sensitive operation periods to obtain an upper and lower threshold boundary table;

[0085] (4) Conduct multi-level anomaly screening on the real-time data stream according to the upper and lower threshold boundary table, first identify obvious violations through rule determination, then identify numerical anomalies through deviation calculation, and finally identify behavior anomalies through pattern matching to generate anomaly point records;

[0086] (5) Conduct correlation analysis on the anomaly point records and the causal links in the data disaster recovery association knowledge base, trace the anomaly propagation path and root node, and form an anomaly event chain;

[0087] (6) Conduct risk assessment on the anomaly event chain through severity scoring, impact range marking, and urgency grading, and attach solution suggestions to form a security risk event library.

[0088] Specifically, extract historical operation data from the data disaster recovery association knowledge base, divide the continuous data into multiple time periods through time window segmentation, and construct a benchmark behavior sequence; calculate the normal fluctuation range of each business indicator for the benchmark behavior sequence through statistical analysis, including mean, standard deviation, seasonal variation, and periodic trend, to form a normal baseline profile;

[0089] For each business indicator X, calculate its mean μ by adding up the values at all time points and then dividing by the total number. The standard deviation σ is obtained by calculating the sum of the squared deviations of each value from the mean, dividing by the total number, and then taking the square root. These two statistical quantities describe the central tendency and dispersion degree of the data.

[0090]

[0091] where μ X represents the mean of indicator X, N represents the total number of data points, and X t represents the value of indicator at time point t.

[0092]

[0093] where, σ X represents the standard deviation of the index X, measuring the fluctuation range of the data.

[0094] The seasonal variations and cyclic trends are obtained through time series decomposition, considering the original data X t as a combination of trend, seasonal, and random components:

[0095] X t = T X,t + S X,t + R X,t

[0096] where, T X,t represents the long-term trend component, calculated by the moving average method; S X,t represents the seasonal component, extracting the repeating pattern by the grouped average method; R X,t represents the random residual.

[0097] Based on the normal baseline profile, the decision threshold is dynamically adjusted by the adaptive threshold algorithm, where the threshold boundary is relaxed for high-load periods and tightened for sensitive operation periods to obtain the upper and lower threshold boundary tables; for each index X at time point t, the upper and lower threshold boundaries are calculated using dynamic formulas based on the mean, standard deviation, and seasonality:

[0098] Threshold X,upper (t) = μ X + k X,t ·σ X + S X,t

[0099] Threshold X,lower (t) = μ X - k X,t ·σ X + S X,t

[0100] where, Threshold X,upper (t) and Threshold X,lower (t) respectively represent the upper and lower thresholds of the index X at time t, and k X,t is the dynamic adjustment coefficient, adjusted according to the current system state. In high-load periods, k X,t takes a larger value (such as 3.0) to relax the threshold boundary; in sensitive operation periods, k X,t takes a smaller value (such as 2.0) to tighten the threshold boundary.

[0101] Perform multi-level anomaly screening on the real-time data stream according to the upper and lower threshold boundary tables. First, identify obvious violations through rule determination, then identify numerical anomalies through deviation calculation, and finally identify behavioral anomalies through pattern matching to generate anomaly point records. Multi-level anomaly screening first identifies obvious violations through rule determination, such as directly checking whether the data exceeds the system's hard limits or business rules. Subsequently, through deviation calculation, compare the actual value with the threshold boundary, and mark it as an anomaly when the data point exceeds the threshold range. Finally, identify behavioral anomalies through pattern matching, analyze the data change pattern in the recent period, and detect abnormal patterns such as sudden increases, sudden decreases, or abnormal fluctuations. The screening process gradually deepens, and the detection accuracy and comprehensiveness are improved through multi-level determination.

[0102] Perform correlation analysis on the anomaly point records and the causal links in the data disaster recovery associated knowledge base, trace the anomaly propagation path and the root node, and form an anomaly event chain. The correlation analysis first queries the data disaster recovery associated knowledge base to find the nodes and relationships related to the anomaly points. Subsequently, trace the possible roots forward along the causal link and analyze the possible impacts backward. The tracing process adopts a heuristic search strategy, giving priority to strong correlation relationships and high-frequency co-occurrence patterns. Construct the anomaly propagation path through the correlation graph algorithm, identify the core nodes and key paths, and form a complete anomaly event chain to clearly present the source, propagation, and impact scope of the anomaly.

[0103] Conduct risk assessment on the anomaly event chain through severity scoring, impact scope marking, and urgency grading, and attach solution suggestions to form a security risk event library. The risk assessment determines the importance and processing priority of the event through quantitative calculation. The severity scoring considers the degree of deviation from the normal value, the duration, and the business importance of the impact of the anomaly; the impact scope marking identifies the affected system components, functional modules, and user groups; the urgency grading evaluates the sensitivity of the processing time, which is divided into five levels from low to high: normal, concerned, important, urgent, and critical. The evaluation results are combined with historical processing experience to generate targeted solution suggestions, including the proposed processing steps, resource requirements, and expected effects, to form a complete record in the security risk event library.

[0104] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0105] (1) Extract the anomaly event characteristics and impact scope information from the security risk event library, retrieve historical processing experience through association rules, and construct a decision knowledge graph;

[0106] (2) Quantitatively assign weights to the four core objectives of system availability, performance impact, security risk, and recovery time through multi-objective weight assignment for the decision knowledge graph to generate an objective weight matrix;

[0107] (3)Comprehensively score the alternative disaster recovery strategies through multi-dimensional evaluation based on the target weight matrix, considering three key indicators: resource consumption, degree of business interruption, and subsequent risks, to form a candidate set of disaster recovery solutions;

[0108] (4)Screen out the non-dominated solutions from the candidate set of disaster recovery solutions through Pareto optimal selection, that is, solutions that cannot be improved in any objective without sacrificing other objectives, to obtain the disaster recovery plan;

[0109] (5)According to the disaster recovery plan, classify the risks of operations through a hierarchical response controller. Low-risk operations are directly triggered, and high-risk operations generate confirmation requests to form an operation sequence;

[0110] (6)Convert the operation sequence into four types of system instructions: resource scheduling instructions, service migration commands, configuration update instructions, and security policy adjustment commands through the API interface protocol, record the execution status and effect feedback, and generate a disaster recovery execution log.

[0111] Specifically, extract the abnormal event characteristics and impact scope information from the security risk event library. Abnormal event characteristics are descriptive data of abnormal situations, including abnormal types, abnormal values, abnormal durations, and occurrence frequencies; impact scope information describes the system components and business functions affected by the abnormality, including the names of affected components, degrees of influence, and impact diffusion paths. The extraction process uses a structured query method to focus on the most recently occurred abnormal events and high-priority abnormal events. The extracted information is then used to retrieve historical processing experiences through association rules. Association rules are a data association technology based on conditional matching that matches the current abnormal event with similar events in history to find corresponding processing methods and effect evaluations. The retrieved historical processing experiences and the current abnormal situation together construct a decision knowledge graph, which is a special knowledge representation structure that includes abnormal event nodes, solution nodes, effect evaluation nodes, and the association relationships between them, providing knowledge support for subsequent decisions.

[0112] Quantify the decision knowledge graph through multi-objective weight assignment. Multi-objective weight assignment is a technology for balancing multiple decision-making objectives. In a cloud computing environment, system availability, performance impact, security risk, and recovery time are four core decision-making objectives. System availability refers to the degree to which the system can provide services normally, performance impact refers to the impact of the solution on the system performance, security risk refers to the potential security threats introduced during the solution process, and recovery time refers to the time required to recover from an abnormal state to a normal state. The weight assignment process assigns corresponding weight values to each objective according to the current business priority and abnormal characteristics to form a target weight matrix. The weight matrix is a data structure that contains the corresponding relationship between objectives and weights, recording the relative importance of each objective in the current decision-making scenario.

[0113] Based on the target weight matrix, a comprehensive score is given to the alternative disaster recovery strategies through multi-dimensional evaluation. Multi-dimensional evaluation is a method of evaluating alternative solutions from multiple aspects. For each possible disaster recovery strategy, it is evaluated from three key dimensions: resource consumption, degree of business interruption, and subsequent risks. Resource consumption refers to the computing, storage, and network resources required to execute the disaster recovery strategy; the degree of business interruption refers to the degree and scope of the impact on business services during the disaster recovery process; subsequent risks refer to the potential residual risks after the disaster recovery operation is completed. In the evaluation process, each disaster recovery strategy is scored on each dimension, and then the weighted scores are calculated in combination with the target weight matrix to form a candidate set of disaster recovery plans containing multiple alternative solutions and their scores. The best plan is selected from the candidate set of disaster recovery plans through Pareto optimal selection. Pareto optimality is an important concept in multi-objective optimization, which refers to a state where other objectives cannot be further improved without compromising one objective. The screening process first identifies non-dominated solutions, that is, those solutions that are not completely outperformed by other solutions in any objective; then these non-dominated solutions are further screened according to actual business requirements, such as giving priority to recovery time requirements, resource limitations, or risk tolerance; finally, a disaster recovery plan that best suits the current situation is determined. The selected disaster recovery plan contains detailed operation steps, resource requirements, and expected effects, which serve as the basis for subsequent automatic response execution.

[0114] According to the disaster recovery plan, the risk levels of operations are classified through a hierarchical response controller. The hierarchical response controller is an operation management component based on risk assessment, which classifies the operations in the disaster recovery plan according to risk levels. The risk level classification follows a preset risk assessment standard, which usually includes factors such as the scope of operation impact, operation reversibility, and consequences of operation failure. The classification results are generally divided into three levels: low risk, medium risk, and high risk. Low-risk operations such as log cleaning and fine-tuning of configuration parameters can usually be directly executed automatically; medium-risk operations such as restarting non-critical services and resynchronizing data need to be executed automatically under specific conditions; high-risk operations such as primary-secondary switchover and core service migration usually require generating confirmation requests and waiting for manual approval. The operations classified by risk levels are arranged in the order of execution to form an operation sequence.

[0115] Convert the operation sequence into specific system instructions through the API interface protocol. The API interface protocol is the standard interface definition for communication between different system components. Through protocol conversion, the abstract operation description is converted into specific executable instructions. Four types of system instructions are generated for different types of operations during the conversion process: resource scheduling instructions, which are used to allocate, release, or adjust computing, storage, and network resources; service migration commands, which are used to move services from one node to another; configuration update instructions, which are used to modify system parameters, policy settings, or service configurations; and security policy adjustment commands, which are used to update firewall rules, access control lists, or encryption settings, etc. Each instruction contains an execution object, execution content, execution parameters, and a callback method, and is sent to the target system through the corresponding API interface. After the instruction is executed, the execution status (success, failure, or partially completed) and effect feedback (such as performance changes, resource occupancy, business impact, etc.) are recorded to generate a disaster recovery execution log, providing data support for subsequent optimization.

[0116] Taking the core transaction system of a certain bank as an example, when a transaction delay anomaly caused by the exhaustion of the database connection pool is detected from the security risk event library, the anomaly features (the connection pool utilization rate reaches 100%, and the duration is 15 minutes) and the affected scope (the response time of the payment processing module increases by 300%) are first extracted. Through association rule retrieval, it is found that similar anomalies in history mainly stem from improper database parameter configuration or connection leakage, and the successful handling experiences include increasing the connection pool capacity, optimizing SQL queries, and fixing the connection leakage code. Based on this information, a decision-making knowledge graph is constructed, connecting the anomaly phenomenon, possible causes, and potential solutions. Multi-objective weight allocation is performed on the decision-making knowledge graph. Given that it is the business peak period currently, the system availability is given a high weight of 0.4, the performance impact is 0.3, the recovery time is 0.2, and the security risk is 0.1. Based on the target weight matrix, three possible disaster recovery strategies (increasing the connection pool capacity, restarting the database service, enabling the standby database) are evaluated, considering the resource consumption, business interruption degree, and subsequent risks of each solution. Through Pareto optimal selection, it is determined that increasing the connection pool capacity is the best solution currently because it can quickly alleviate the problem without interrupting the business. The hierarchical response controller classifies the operation of increasing the connection pool capacity as a medium-risk operation, which needs to be executed during the non-transaction peak period. The operation sequence is converted into a configuration update instruction of "modifying the value of the database configuration parameter max_connections", executed through the database management API, and the effect that the connection pool utilization rate drops from 100% to 60% before and after the execution is recorded, completing the entire disaster recovery execution process and generating a detailed disaster recovery execution log.

[0117] In a specific embodiment, the process of executing step S106 may specifically include the following steps:

[0118] (1) Extract response operation records, execution status data, and effect feedback information from the disaster recovery execution log, and construct an effect evaluation data set through data aggregation;

[0119] (2) Quantify the four key indicators of problem-solving rate, system recovery time, resource utilization efficiency, and user satisfaction through multi-dimensional metrics for the effect evaluation data set, and generate a performance indicator table;

[0120] (3) Compare the performance indicator table with the historical average level, expected target value, and industry best practice through horizontal comparative analysis, identify the advantageous items and deficiency items, and form a difference analysis report;

[0121] (4) Based on the difference analysis report, adjust and calculate the key parameters through the Bayesian optimization algorithm, including four types of parameters: disaster recovery data collection frequency, security detection threshold, prediction window size, and response priority, to obtain a parameter optimization plan;

[0122] (5) Supplement and correct the abnormal pattern library through the incremental update method with the parameter optimization plan, eliminate the outdated abnormal patterns, supplement the newly added abnormal types, and construct optimized knowledge rules;

[0123] (6) Integrate the parameter optimization plan and the optimized knowledge rules, and convert them into an executable configuration instruction set through the system configuration generator to form an optimization strategy for disaster recovery and security protection.

[0124] Specifically, three types of key data are extracted from the disaster recovery execution log: response operation records, execution status data, and effect feedback information. Response operation records are records of the specific operations performed during the disaster recovery response, including operation type, operation time, operation object, and operation parameters; execution status data reflects the execution results of various operations, including execution success rate, execution time consumption, resource occupancy, and execution exception information; effect feedback information includes content such as changes in the system state, changes in business metrics, and user feedback after the operation. The extracted data is integrated through data aggregation. Data aggregation is a technology that centralizes and organizes dispersed data according to specific dimensions. Through time dimension aggregation, operation type aggregation, and target object aggregation, relevant data is merged into a structured effect evaluation data set, providing a basis for subsequent analysis. Multidimensional index quantification is performed on the effect evaluation data set, and quantitative values are calculated for four key indicators. The problem resolution rate refers to the degree of success in resolving problems through disaster recovery operations, obtained by calculating the proportion of abnormal indicators returning to normal; the system recovery time refers to the time required from the execution of disaster recovery operations to the system returning to the normal state, obtained by recording the time interval from the start of the operation to the return of various indicators to normal; the resource utilization efficiency refers to the rationality of resource use during the disaster recovery process, obtained by calculating the ratio of resource input to problem resolution effect; the user satisfaction refers to the degree of satisfaction of users with the system state after disaster recovery operations, indirectly obtained by collecting user feedback or calculating the improvement degree of service quality. During the quantification process, a standardized calculation method is used to unify different indicators to the same measurement standard, generating a performance index table containing the quantitative values of each indicator.

[0125] The performance indicator table is compared with three benchmarks through horizontal comparative analysis. Horizontal comparative analysis is an analytical method that discovers differences by juxtaposing and comparing different data sources. For each indicator, it is respectively compared with the historical average level (the average performance of the indicator over a period of time in the past), the expected target value (the level that the indicator should reach in the plan), and the industry best practice (the best performance of the indicator in similar systems or services). During the comparison process, the differences between the actual value and the benchmark value are calculated, the statistical significance of the differences is judged, and positive and negative determinations are made. The indicators with significantly higher results than the benchmark are marked as advantageous items, and the indicators significantly lower than the benchmark are marked as deficient items. The results of the comparative analysis are compiled into a difference analysis report, which details the comparison results, the degree of difference, and the causes of the differences for each indicator. Based on the difference analysis report, the key parameters are adjusted and calculated through the Bayesian optimization algorithm. Bayesian optimization is a parameter optimization method based on probability models, especially suitable for optimizing black-box functions with high computational costs and difficult to obtain analytical solutions. The optimization process first establishes a probability model (Gaussian process) between the parameters and the performance indicators, and then, through an iterative approach, comprehensively considers the balance between exploration (searching for unknown regions) and exploitation (optimizing known regions) to gradually find the optimal parameter settings. Four types of key parameters are optimized: the disaster recovery data collection frequency, which controls the time interval for collecting monitoring data from each node; the security detection threshold, which determines the boundary value for triggering security alerts; the prediction window size, which is the length of the historical data window for time series prediction; and the response priority, which determines the processing order for different types of exceptions. The optimization process takes into account the mutual influence between the parameters and, under the premise of meeting the system constraints, searches for the parameter combination that can maximize the overall performance to form a parameter optimization plan. The parameter optimization plan is used to supplement and correct the abnormal pattern library through incremental updates. The abnormal pattern library is a knowledge base that stores known abnormal patterns, including abnormal feature descriptions, triggering conditions, influence scopes, and handling methods. Incremental update is an update mechanism that only modifies the necessary parts without reconstructing the entire library, mainly including three operations: removing obsolete abnormal patterns, that is, those patterns that have not occurred for a long time or are no longer applicable due to changes in the system environment; correcting existing abnormal patterns, adjusting the abnormal feature descriptions, determination thresholds, and handling suggestions according to new observed data; and supplementing new abnormal types, adding newly discovered abnormal patterns to the library, including feature descriptions, detection methods, and recommended responses. The incremental update process maintains the timeliness and accuracy of the abnormal pattern library, forming optimized knowledge rules.

[0126] Integrate the parameter optimization scheme and optimization knowledge rules, and convert them into executable configuration instruction sets through the system configuration generator. The system configuration generator is a tool that converts abstract optimization strategies into specific system configurations. The conversion process includes format conversion (converting optimization strategies into standard configuration formats), dependency checking (ensuring there are no conflicts between configurations), version adaptation (adjusting configuration parameters according to the system version), and verification testing (ensuring the configurations are valid and executable). The converted configuration instruction sets contain various types of configuration instructions, such as monitoring configuration instructions, threshold setting instructions, scheduling rule instructions, and response policy instructions, etc. These configuration instructions together constitute the disaster recovery and security protection optimization strategy, providing the ability for continuous improvement in the cloud computing environment.

[0127] Taking the core transaction system of a financial institution as an example, detailed records were extracted from the execution log of a database high-load disaster recovery process: The response operation record shows that three operations were performed, namely connection pool expansion, query optimization, and cache adjustment; the execution status data records that the connection pool parameter modification was successful but the query optimization part failed; the effect feedback information shows that the system response time recovered from 3 seconds to 300 milliseconds. The effect evaluation data set constructed through data aggregation shows the comparison of the system status before and after the operation. The problem-solving rate calculated by multi-dimensional indicators is 85% (due to incomplete query optimization), the system recovery time is 15 minutes, the resource utilization efficiency is medium (because the expansion of the connection pool capacity consumed additional memory), and the user satisfaction is good (the transaction completion rate has returned to the normal level). Horizontal comparative analysis found that the system recovery time is significantly lower than the historical average level (originally 30 minutes), but the resource utilization efficiency is lower than the industry best practice, marked as an item to be improved. Through 10 rounds of iterative calculations by the Bayesian optimization algorithm, an optimized parameter scheme was obtained: adjust the database monitoring collection frequency from 60 seconds to 30 seconds, adjust the connection pool warning threshold from 90% to 80%, expand the prediction window from 3 hours to 6 hours, and increase the priority of database exception responses by one level. The anomaly pattern library was corrected by incremental update, adding a new pattern of "connection pool exhaustion during peak hours", and updating the corresponding detection rules and handling suggestions. The system configuration generator converts these adjustments into four types of configuration instructions: database parameter settings, monitoring configuration, warning rules, and response processes, constituting the disaster recovery and security protection optimization strategy, and realizing the closed-loop optimization of the data processing method in the cloud computing environment.

[0128] The data processing method for cloud computing in the embodiments of the present application has been described above. Next, the data processing system for cloud computing in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the data processing system for cloud computing in the embodiments of the present application includes:

[0129] A processing module for collecting and preprocessing the operation logs and network traffic of a multi-source data center through distributed acquisition nodes to obtain structured data packets;

[0130] A classification module for intelligently classifying and analyzing data through hybrid feature extraction based on the structured data packets to obtain security domain and disaster recovery domain data sets;

[0131] An analysis module for mining and analyzing data association relationships through a multi-level heterogeneous fusion network based on the security domain and disaster recovery domain data sets to generate a data disaster recovery association knowledge base;

[0132] A prediction module for detecting and predicting anomalies in the security status of the cloud environment through an adaptive threshold algorithm based on the data disaster recovery association knowledge base to form a security risk event library;

[0133] An implementation module for determining a disaster recovery plan through a multi-objective decision-making process based on the security risk event library, implementing an automatic response mechanism, and generating disaster recovery execution logs;

[0134] An adjustment module for adjusting and optimizing disaster recovery, system, and storage medium parameters through a multi-angle comparison process based on the disaster recovery execution logs to form disaster recovery and security protection optimization strategies.

[0135] Through the collaborative cooperation of the above-mentioned various components, the operation logs and network traffic of multi-source data centers are collected and preprocessed by distributed acquisition nodes, achieving comprehensive coverage and preliminary purification of data sources, and significantly improving the data quality for subsequent analysis; through hybrid feature extraction on structured data packets, the data is intelligently classified and parsed, effectively identifying and extracting key features, reducing the data dimension and computational complexity, and enhancing the discriminability and expressiveness of features; based on the data sets of the security domain and disaster recovery domain, the mining and analysis of data association relationships are carried out through a multi-level heterogeneous fusion network, breaking through the limitations of traditional single-dimensional analysis, capturing complex association patterns across levels and domains, and improving the accuracy and comprehensiveness of anomaly detection; according to the data disaster recovery association knowledge base, the anomaly detection and prediction of the cloud environment security status are carried out through an adaptive threshold algorithm, dynamically adjusting the judgment criteria to adapt to system load changes, while taking into account time sensitivity, and significantly reducing the false alarm rate and missed alarm rate; according to the security risk event library, the disaster recovery plan is determined through a multi-objective decision-making process, comprehensively considering multiple objectives such as system availability, performance impact, security risk, and recovery time, realizing the scientific and automated disaster recovery decision-making, and significantly reducing manual intervention and response delay; according to the disaster recovery execution log, the disaster recovery, system, and storage medium parameters are adjusted and optimized through a multi-angle comparison process, forming a closed-loop optimization mechanism, and continuously improving the system performance and protection ability. In particular, this solution applies a number of artificial intelligence algorithms in the cloud computing environment. Among them, the multi-level heterogeneous fusion network algorithm plays a key role in data association analysis. Through different levels of network structures and inter-layer mappings, it effectively captures complex direct and indirect associations between data, providing a comprehensive and three-dimensional association perspective for anomaly detection; the adaptive threshold algorithm adjusts the judgment criteria in real time according to the dynamic characteristics of the cloud environment, avoiding the limitations of traditional fixed thresholds; the multi-objective decision-making algorithm can balance multiple decision-making objectives in a complex and changing cloud environment, find the optimal disaster recovery plan, and adapt to the complexity and multi-objectivity of disaster recovery decision-making in the cloud computing environment; the Bayesian optimization algorithm guides parameter search through a probability model, efficiently finds the optimal parameter configuration of the system, and provides a theoretical support for the continuous optimization of the system. The organic combination of these algorithms significantly improves the data processing intelligence level, disaster recovery response timeliness, and system operation stability of this solution in the cloud computing environment.

[0136] Referring to Figure 3 , in the embodiment of the present invention, a computer device is further provided. This computer device can be a server, and its internal structure can be as Figure 3As shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected via a system and storage medium bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and storage medium, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and storage medium and computer programs in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. The computer program, when executed by the processor, implements the above method.

[0137] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0138] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0139] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0140] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, storage medium, system and unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0141] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0142] The above is the case. The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A data processing method for cloud computing, characterized in that, The data processing method for cloud computing includes: Collecting and preprocessing the operation logs and network traffic of a multi-source data center through distributed collection nodes to obtain structured data packets; Intelligently classifying and analyzing the data through hybrid feature extraction according to the structured data packets to obtain security domain and disaster recovery domain data sets; Mining and analyzing the data association relationships through a multi-level heterogeneous fusion network based on the security domain and disaster recovery domain data sets to generate a data disaster recovery association knowledge base; Performing anomaly detection and prediction on the cloud environment security status through an adaptive threshold algorithm according to the data disaster recovery association knowledge base to form a security risk event library; Determining a disaster recovery plan through a multi-objective decision-making process according to the security risk event library, implementing an automatic response mechanism, and generating disaster recovery execution logs; Adjusting and optimizing the disaster recovery, system, and storage medium parameters through a multi-angle comparison process according to the disaster recovery execution logs to form disaster recovery and security protection optimization strategies.

2. The data processing method for cloud computing according to claim 1, characterized in that, The step of collecting and preprocessing the operation logs and network traffic of a multi-source data center through distributed collection nodes to obtain structured data packets includes: Obtaining server operation logs, network traffic data, user behavior data, and application performance indicators from multiple data centers through edge perception collectors to generate an original data stream; Performing data cleaning on the original data stream through a double-layer filtering mechanism, where the first layer uses a rule engine to eliminate obvious anomalies and redundant data, and the second layer uses an outlier identification algorithm to screen potential anomaly points to obtain cleaned data; Calibrating and aligning the cleaned data according to timestamps to construct data units with consistent time sequences; Performing standardization processing on the data units with consistent time sequences through a format converter to convert heterogeneous data from different sources into a unified format; Adding metadata tags to the data in the unified format, including data source identifiers, time stamps, service types, and security levels, to form tagged data; Generating a checksum for the tagged data through a checksum calculator and encapsulating the data body, metadata tags, and checksum to obtain the structured data packet.

3. The data processing method for cloud computing according to claim 1, characterized in that, The step of intelligently classifying and analyzing the data through hybrid feature extraction according to the structured data packets to obtain security domain and disaster recovery domain data sets includes: Extracting time features, space features, service features, and security features from the structured data packets through a feature extraction engine to form a feature matrix; Performing dimension reduction on the feature matrix through hybrid feature extraction combining statistics and deep learning to remove information redundancy and enhance feature distinguishability to obtain a core feature set; Inputting the core feature set into a hierarchical classifier and coarsely dividing the features through a decision tree to generate preliminary classification labels; Performing fine classification on the preliminary classification labels through vector calculation to divide the data into computing resource segments, storage resource segments, network resource segments, and security segments; Extracting data items related to disaster recovery functions from the security segments through a disaster recovery association filter to form disaster recovery segment data; Merging and standardizing the data in the security segments and the data in the disaster recovery segments to generate the security domain and disaster recovery domain data sets.

4. The data processing method for cloud computing according to claim 1, wherein, The data disaster recovery association knowledge base is generated by mining and analyzing the data association relationship between the security domain and the disaster recovery domain data set through a multi-level heterogeneous fusion network, including: The data sets of the security domain and the disaster recovery domain are organized into network structures at different levels through a data layer constructor, where the first layer is the entity layer including device nodes and service nodes, the second layer is the attribute layer including security indicators and disaster recovery indicators, and the third layer is the event layer including historical security events and disaster recovery events, forming a three-layer heterogeneous network structure; The connection relationships between nodes in the same layer are established for the three-layer heterogeneous network structure through intra-layer association analysis. Among them, the entity layer is connected through a topological relationship, the attribute layer is connected through a correlation threshold, and the event layer is connected through a time series dependence, obtaining an intra-layer association graph; Vertical connections between nodes at different levels are established for the intra-layer association graph through inter-layer mapping, mapping the entity layer nodes to their corresponding attribute layer indicators and event layer records, constituting an inter-layer connection graph; Path analysis is performed on the inter-layer connection graph through an improved random walk algorithm to identify key paths and core nodes, obtaining an important association subgraph; Temporal association patterns between data are extracted from the important association subgraph through temporal projection, including periodic patterns, trend patterns, and abnormal fluctuation patterns, forming a temporal association rule set; The intra-layer association graph, the inter-layer connection graph, and the temporal association rule set are integrated and stored in a graph data structure to constitute the data disaster recovery association knowledge base.

5. The data processing method for cloud computing according to claim 1, wherein The security state of the cloud environment is detected and predicted for anomalies through an adaptive threshold algorithm based on the data disaster recovery association knowledge base, forming a security risk event library, including: Historical operation data is extracted from the data disaster recovery association knowledge base, and continuous data is divided into multiple time periods through time window segmentation to construct a benchmark behavior sequence; The normal fluctuation ranges of each business indicator, including mean, standard deviation, seasonal variation, and periodic trend, are calculated for the benchmark behavior sequence through statistical analysis, forming a normal baseline profile; The decision threshold is dynamically adjusted through an adaptive threshold algorithm based on the normal baseline profile. Among them, the threshold boundary is relaxed for high-load periods and tightened for sensitive operation periods, obtaining an upper and lower threshold boundary table; Multi-level anomaly screening is performed on the real-time data stream according to the upper and lower threshold boundary table. First, obvious violations are identified through rule judgment, then numerical anomalies are identified through deviation calculation, and finally behavior anomalies are identified through pattern matching, generating anomaly point records; The anomaly point records are associated and analyzed with the causal links in the data disaster recovery association knowledge base to trace the anomaly propagation path and root cause node, forming an anomaly event chain; Risk assessment is performed on the anomaly event chain through severity scoring, impact range marking, and urgency grading, and solution suggestions are added to constitute the security risk event library.

6. The data processing method for cloud computing according to claim 1, wherein The disaster recovery plan is determined through a multi-objective decision-making process according to the security risk event library, and an automatic response mechanism is implemented to generate a disaster recovery execution log, including: Anomaly event characteristics and impact range information are extracted from the security risk event library, and historical processing experience is retrieved through association rules to construct a decision-making knowledge graph; Quantify and assign weights to the four core objectives of system availability, performance impact, security risk, and recovery time for the decision-making knowledge graph through multi-objective weight assignment to generate an objective weight matrix; Based on the objective weight matrix, comprehensively score the alternative disaster recovery strategies through multi-dimensional evaluation, considering three key indicators of resource consumption, business interruption degree, and subsequent risks, to form a candidate set of disaster recovery solutions; Select non-dominated solutions from the candidate set of disaster recovery solutions through Pareto optimal selection, that is, solutions that cannot be improved in any objective without sacrificing other objectives, to obtain the disaster recovery plan; According to the disaster recovery plan, divide the risks of operations through a hierarchical response controller. Low-risk operations are directly triggered, and high-risk operations generate confirmation requests to form an operation sequence; Convert the operation sequence into four types of system instructions, namely resource scheduling instructions, service migration commands, configuration update instructions, and security policy adjustment commands, through the API interface protocol, record the execution status and effect feedback, and generate the disaster recovery execution log; 7. The data processing method for cloud computing according to claim 1, wherein Adjust and optimize the disaster recovery, system, and storage medium parameters through a multi-angle comparison process based on the disaster recovery execution log to form a disaster recovery and security protection optimization strategy, including: Extract response operation records, execution status data, and effect feedback information from the disaster recovery execution log, and construct an effect evaluation data set through data aggregation; Quantitatively calculate four key indicators of problem-solving rate, system recovery time, resource utilization efficiency, and user satisfaction for the effect evaluation data set through multi-dimensional index quantification to generate a performance index table; Compare the performance index table with the historical average level, expected target value, and industry best practice through horizontal comparative analysis to identify the advantageous items and deficiencies, and form a difference analysis report; Based on the difference analysis report, adjust and calculate the key parameters through the Bayesian optimization algorithm, including four types of parameters: disaster recovery data collection frequency, security detection threshold, prediction window size, and response priority, to obtain a parameter optimization plan; Supplement and correct the abnormal pattern library through the incremental update method according to the parameter optimization plan, eliminate the outdated abnormal patterns, supplement the new abnormal types, and construct optimized knowledge rules; Integrate the parameter optimization plan and the optimized knowledge rules, and convert them into an executable configuration instruction set through a system configuration generator to constitute the disaster recovery and security protection optimization strategy.

8. A data processing system for cloud computing, which is used to implement the data processing method for cloud computing as described in any one of claims 1-7, characterized in that, The data processing, system, and storage medium for cloud computing include: A processing module for collecting and preprocessing the operation logs and network traffic of multi-source data centers through distributed collection nodes to obtain structured data packets; A classification module for intelligently classifying and parsing data through hybrid feature extraction according to the structured data packets to obtain security domain and disaster recovery domain data sets; An analysis module for mining and analyzing the data association relationships based on the security domain and disaster recovery domain data sets through a multi-level heterogeneous fusion network to generate a data disaster recovery association knowledge base; A prediction module for detecting and predicting anomalies in the cloud environment security status through an adaptive threshold algorithm based on the data disaster recovery association knowledge base to form a security risk event library; An implementation module, configured to determine a disaster recovery plan through a multi-objective decision-making process according to the security risk event library, implement an automatic response mechanism, and generate a disaster recovery execution log; An adjustment module, configured to adjust and optimize disaster recovery, system, and storage medium parameters through a multi-angle comparison process according to the disaster recovery execution log, and form a disaster recovery and security protection optimization strategy.

9. A computer device, characterized in that, It includes a memory and a processor, and the memory stores a computer program that can run on the processor. It is characterized in that when the processor executes the computer program, it implements the data processing method for cloud computing described in any one of claims 1 to 7.

10. A computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the processor is caused to execute the data processing method for cloud computing described in any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent operation and maintenance monitoring method and system for data center

    CN120602308A

  • System optimization method and device, electronic equipment, storage medium and program

    CN120803877A

  • Ecological cycle agricultural data acquisition and analysis method and system based on Internet of Things

    CN120851662A

  • Real-time detection method for water level abnormal data

    CN121030603A

  • Backup disaster recovery optimization method

    CN121037195A