Data security disaster recovery method and system of data center
By analyzing the amount and characteristics of historical data and dynamically adjusting the data synchronization strategy, the problem of the inability to flexibly choose full or incremental synchronization in existing technologies is solved, and the data synchronization effect and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202510722241.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Existing data synchronization solutions usually only offer a single choice between full synchronization and incremental synchronization, and are unable to dynamically adjust between the two modes based on actual needs. This results in poor data synchronization effects and an inability to adapt to the data changes in different business scenarios.
By analyzing the data volume in historical time periods, dividing the time stages, and combining data duplication and growth characteristics to predict future data pressure, a reasonable data synchronization strategy is selected based on data concentration, and the full or incremental synchronization strategy is dynamically adjusted to achieve data security and disaster recovery.
It enables flexible adjustment of data synchronization strategies based on actual needs, improves data synchronization effects, and ensures data consistency and resource utilization efficiency of the data center in different business scenarios.
Smart Images

Figure CN120596869A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a data security disaster recovery method and system for a data center. Background Art
[0002] Data security and disaster recovery strategies play a crucial role in the daily operations and management of data centers. They are key to ensuring data integrity, availability, and business continuity. Traditionally, data security and disaster recovery rely primarily on physical disaster recovery measures, such as configuring backup servers and storage devices and building disaster recovery computer rooms. Their core goal is to quickly restore data and maintain normal business operations in the event of emergencies such as hardware failures, natural disasters, or human errors.
[0003] Currently, to achieve efficient data security and disaster recovery, the industry generally adopts data synchronization technologies such as full synchronization or incremental synchronization. These technologies each have their own advantages and can effectively ensure the consistency and real-time performance of data across different systems or storage locations.
[0004] However, existing data synchronization solutions typically only offer a single choice between full synchronization and incremental synchronization, without the ability to dynamically adjust between the two modes based on actual needs. This lack of flexibility often leads to poor data synchronization performance. Summary of the Invention
[0005] The embodiments of the present invention provide a data security disaster recovery method and system for a data center, which can improve data synchronization effect.
[0006] A first aspect of an embodiment of the present invention provides a data security disaster recovery method for a data center, comprising:
[0007] Obtain the data volume of the data center on the current day and the data volume of each historical day in the historical time period;
[0008] Based on the data volume of each historical day within the historical time period, a day is divided into multiple time periods. The data volume per hour within the time period shows consistent fluctuations.
[0009] Predict future data pressure based on the data duplication and growth characteristics of the target time period. The target time period is the time period of the current day at the current moment.
[0010] When the future data pressure is less than or equal to the preset pressure threshold, the data volume of the current day in the target time period is compared with the data volume of each historical day in the target time period to determine the first data concentration of the current day in the target time period;
[0011] Based on the first data concentration, a data synchronization strategy at a current moment is determined, so that data security disaster recovery is performed on the data center through the data synchronization strategy.
[0012] In some possible implementations, a day is divided into multiple time periods based on the amount of data on each historical day within a historical time period, which may include:
[0013] Determine the concentration of the second data per hour based on the data volume per hour of each historical day within the historical time period;
[0014] Perform curve fitting on the second data concentration every hour to obtain a target curve;
[0015] The time intervals between adjacent maximum value points in the target curve are divided into a time period, thereby obtaining multiple time periods.
[0016] In some possible implementations, determining the hourly second data concentration based on the hourly data volume of each historical day within the historical time period may specifically include:
[0017] The proportion of the data volume of each historical day at the nth hour to the total data volume of each historical day is determined as the data volume prominence of each historical day at the nth hour, where n is a positive integer;
[0018] The data concentration coefficient of the nth hour is determined by using the difference in the prominence of the data volume of adjacent historical days at the nth hour;
[0019] The second data concentration at the nth hour is determined by using the minimum value of the prominence of each data amount and the data concentration coefficient.
[0020] In some possible implementations, based on the data duplication characteristics and data growth characteristics of the target time period, predicting future data pressure may include:
[0021] Determine the target data repetition degree of the target time period based on the data repetition characteristics of the target time period;
[0022] Performing mean processing on the second data concentration every hour in the target time period to obtain the first data growth rate in the target time period;
[0023] Use the target data duplication and the first data growth rate to predict future data pressure.
[0024] In some possible implementations, determining a target data duplication degree for a target time period based on data duplication characteristics for the target time period may specifically include:
[0025] For each historical day, perform the following steps: perform standard deviation processing on the frequency of occurrence of each data type within the target time period of the historical day to obtain the first repetition coefficient of the historical day;
[0026] The first repetition coefficient of each historical day is averaged to obtain the first data repetition degree of the target time period;
[0027] For each data type, perform the following steps: perform standard deviation processing on the frequency of occurrence of the data type on each historical day to obtain the second repetition coefficient of the data type;
[0028] Perform mean processing on the second repetition coefficient of each data type to obtain the second data repetition degree of the target time period;
[0029] A target data repetition rate in a target time period is determined by using the first data repetition rate and the second data repetition rate.
[0030] In some possible implementations, comparing the data volume of the current day at the target time period with the data volume of each historical day at the target time period to determine the first data concentration of the current day at the target time period may specifically include:
[0031] Compare the second data concentration of each hour of the current day within the target time period with the second data concentration of each historical day within the target time period to obtain the data concentration difference;
[0032] The concentration differences of each data set are averaged to obtain the concentration of the first data set at the target time stage of the current day.
[0033] In some possible implementations, after comparing the data volume of the current day at the target time period with the data volume of each historical day at the target time period to determine the first data concentration of the current day at the target time period, the data security disaster recovery method of the data center may further include:
[0034] Predict future packet loss rates based on the number of packet losses on each historical day within the target time period;
[0035] Based on the future packet loss rate, determine the frequency adjustment coefficient for the data synchronization transmission frequency during the target time period of the current day;
[0036] The frequency adjustment coefficient is used to adjust the frequency of data synchronization transmission within the target time period of the current day.
[0037] In some possible implementations, predicting the future packet loss rate based on the number of packet losses on each historical day within the target time period may specifically include:
[0038] Obtain the number of packet losses for each historical day within the target time period, and the second data growth rate for the current day within the target time period;
[0039] The average number of packet losses on each historical day within the target time period is calculated to obtain the average number of packet losses on each historical day.
[0040] Comparing the second data growth rate with the first data growth rate to obtain a growth rate difference;
[0041] The future packet loss rate is determined by using the growth difference and the average number of packet losses.
[0042] In some possible implementations, determining a frequency adjustment coefficient for the frequency of synchronous data transmission within a target time period on the current day based on a future packet loss rate may specifically include:
[0043] Obtaining the bandwidth occupancy at the current moment and a packet loss data sequence for a target time period, wherein the packet loss data sequence includes each packet loss moment within the target time period;
[0044] Using the packet loss data sequence of the target time period, determining the packet loss time interval characteristic value of the target time period;
[0045] The frequency adjustment coefficient is determined using the bandwidth usage, the characteristic value of the packet loss time interval, and the future packet loss rate.
[0046] A second aspect of an embodiment of the present invention provides a data security disaster recovery system for a data center, comprising:
[0047] The data volume acquisition module is used to obtain the data volume of the data center on the current day and the data volume of each historical day in the historical time period;
[0048] The stage division module is used to divide a day into multiple time stages based on the data volume of each historical day within the historical time period. The data volume per hour within the time stage shows consistent fluctuations.
[0049] The pressure prediction module is used to predict future data pressure based on the data duplication characteristics and data growth characteristics of the target time period. The target time period is the time period of the current time of the current day;
[0050] A concentration determination module is configured to compare the data volume of the current day in the target time period with the data volume of each historical day in the target time period when the future data pressure is less than or equal to a preset pressure threshold, and determine the first data concentration of the current day in the target time period;
[0051] The strategy determination module is used to determine the data synchronization strategy at the current moment based on the first data concentration, so as to perform data security disaster recovery for the data center through the data synchronization strategy.
[0052] The present invention has the following beneficial effects:
[0053] In the data security disaster recovery method for a data center provided by an embodiment of the present invention, the time phases are divided according to the amount of data of each historical day within the historical time period. Then, based on the data duplication characteristics and data growth characteristics of the target time phase corresponding to the current moment of the current day, the future data pressure is predicted. In the case that the future data pressure is less than or equal to the preset pressure threshold, the amount of data of the current day in the target time phase is further compared with the amount of data of each historical day in the target time phase to determine the first data concentration of the current day in the target time phase. Thus, according to the first data concentration of the current day in the target time phase, the data synchronization strategy at the current moment is selected, so that data security disaster recovery is performed on the data center through the data synchronization strategy. In this way, the present invention can flexibly and dynamically adjust the data synchronization strategy by predicting the concentration of data in the target time phase of the current day, thereby improving the data synchronization effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0055] Figure 1 A schematic diagram of a flow chart of a first data center data security disaster recovery method provided by one embodiment of the present invention;
[0056] Figure 2 A schematic diagram of the process of S102 provided in one embodiment of the present invention;
[0057] Figure 3 A schematic diagram of time phases provided for one embodiment of the present invention;
[0058] Figure 4 A schematic diagram of the process of S103 provided in one embodiment of the present invention;
[0059] Figure 5 A schematic diagram of the process of S104 provided in one embodiment of the present invention;
[0060] Figure 6 A flow chart of a second data center data security disaster recovery method provided by one embodiment of the present invention;
[0061] Figure 7A schematic diagram of the structure of a data security disaster recovery system for a data center provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0062] To further illustrate the technical means and effectiveness of the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail a data center data security disaster recovery method and system according to the present invention, including its specific implementation, structure, features, and effectiveness. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0063] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0064] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of the present invention comply with the relevant provisions of laws and regulations.
[0065] Within the complex and sophisticated daily operations and management systems of data centers, data security and disaster recovery strategies undoubtedly occupy a core position, serving as a solid foundation for ensuring data integrity, availability, and business continuity. In today's highly digitalized era, data has become one of a company's most valuable assets. Any loss, corruption, or unavailability of data can result in immeasurable losses, ranging from financial loss to reputational damage, and even potentially leading to the complete paralysis of a company's operations. Therefore, establishing a comprehensive and efficient data security and disaster recovery strategy is crucial to a company's survival and development.
[0066] With the rapid development of information technology and the explosive growth of data volumes, traditional physical disaster recovery measures are no longer able to meet the increasingly complex data security needs of enterprises. To achieve more efficient data security and disaster recovery, the industry has actively explored and widely adopted data synchronization technologies such as full synchronization and incremental synchronization. Full synchronization is like a comprehensive "data migration," completely replicating all data from the source data system to the target data system, ensuring that the target data system is completely consistent with the source data system at a given moment. This technology offers advantages such as ease of operation and high data consistency. However, its disadvantages are that when the data volume is large, the synchronization process consumes significant network bandwidth and storage resources, resulting in longer synchronization times. Incremental synchronization, on the other hand, is more cost-effective, synchronizing only the data in the source data system that has changed since the last synchronization. This technology offers advantages such as fast synchronization and low resource consumption, but its disadvantages are the need for additional mechanisms to record data changes and the relatively complex process of applying incremental data sequentially in the order of changes during data recovery. Both data synchronization technologies have their own advantages. Enterprises can choose the appropriate synchronization technology based on their business needs, data characteristics, and resource availability to effectively ensure data consistency and real-time performance across different systems or storage locations.
[0067] However, existing data synchronization solutions have a significant limitation, which is that they can usually only make a single choice between full synchronization and incremental synchronization, and cannot dynamically adjust between the two modes according to actual needs. In actual business scenarios, data changes are complex and changeable. For example, during certain business peaks, the frequency of data changes will increase significantly. At this time, the use of incremental synchronization technology may increase the delay in data synchronization and affect data consistency; while during business slack periods, data changes are relatively small, and the use of full synchronization technology will result in a waste of resources. This lack of flexibility makes the data synchronization solution unable to adapt to the characteristics of data changes in different business scenarios, often resulting in poor data synchronization results and an inability to fully utilize the advantages of data synchronization technology.
[0068] The purpose of the present invention is to provide a data security disaster recovery method and system for a data center. In the data security disaster recovery method for a data center provided by an embodiment of the present invention, the time phase is divided according to the amount of data of each historical day in the historical time period. Then, based on the data repetition characteristics and data growth characteristics of the target time phase corresponding to the current moment of the current day, the future data pressure is predicted. In the case where the future data pressure is less than or equal to the preset pressure threshold, the amount of data of the current day in the target time phase is further compared with the amount of data of each historical day in the target time phase to determine the first data concentration of the current day in the target time phase. Thus, according to the first data concentration of the current day in the target time phase, the data synchronization strategy of the current moment is selected, so that data security disaster recovery of the data center is performed through the data synchronization strategy. In this way, the present invention can flexibly and dynamically adjust the data synchronization strategy by predicting the concentration of data in the target time phase of the current day, thereby improving the data synchronization effect.
[0069] The following describes specific embodiments of the data security disaster recovery method and system for a data center provided by the embodiments of the present invention.
[0070] Figure 1 A flow chart of a data security disaster recovery method for a data center is provided. The data security disaster recovery method for a data center can be applied to a server. The data security disaster recovery method for a data center can include the following steps S101 to S105.
[0071] S101, obtaining the data volume of the data center on the current day and the data volume of each historical day in a historical time period.
[0072] In this embodiment, the historical time period is a preset time period. For example, the historical time period may be the past month.
[0073] As an example, dedicated data collection agents can be deployed in the data center. These agents can be distributed across different nodes in the data center and are responsible for collecting information about the amount of data generated locally. For example, an agent can be installed on a database server to count the number of new data records added to the database every hour, the size change of data files, and so on.
[0074] At the same time, establish a dedicated data repository to store the collected data volume information. This can be done using a relational database (such as MySQL, Oracle) or a non-relational database (such as MongoDB, Redis). Categorize and store the collected data volume information by date, storing the current day's data volume and the data volume for each historical day within a month in separate data tables to facilitate subsequent query and analysis.
[0075] S102 , based on the data volume of each historical day in the historical time period, divide one day into multiple time periods, and the data volume per hour in the time period presents consistent fluctuation.
[0076] In this embodiment, a time period includes multiple hours in which the data volume exhibits consistent fluctuations.
[0077] As an example, the server first cleans the data volume information for each historical day within the historical time period to remove outliers and erroneous information. For example, data volume information that significantly deviates from the normal range can be identified and eliminated through statistical analysis methods (such as the 3σ principle). Then, the average data volume per hour for each historical day within the historical time period is calculated to obtain the average data volume per hour for the historical time period.
[0078] Then, cluster analysis is performed on the average hourly data volume of historical days using a clustering algorithm (such as the K-means clustering algorithm). The hourly average data volume is used as a feature vector, and the cluster centers are iteratively adjusted to divide the 24-hour day into multiple categories, where each category corresponds to a time period.
[0079] Furthermore, you can set volatility assessment indicators, such as standard deviation and coefficient of variation, to calculate the volatility of the average hourly data volume within each time period. Based on the magnitude of the volatility indicator, adjust the time period division to ensure that the average hourly data volume within each time period exhibits consistent volatility.
[0080] S103, predicting future data pressure based on data duplication characteristics and data growth characteristics of a target time period, where the target time period is the time period of the current moment of the current day.
[0081] In this embodiment, the future data pressure is used to represent the flow pressure corresponding to the future data. For example, the smaller the future data pressure, the smaller the flow pressure corresponding to the future data, and the more likely it is to consider full data synchronization.
[0082] As an example, the server analyzes the data volume for each historical day within the target time period and counts recurring data patterns, data blocks, and other information. For example, in some business systems, a large number of similar transaction data records are generated during a specific time period every day. The data can be divided into blocks using a hash algorithm to count the frequency of occurrence of identical data blocks.
[0083] At the same time, calculate the data growth rate for each historical day within the target time period and analyze the data growth trend. You can use methods such as linear regression and exponential smoothing to model the data growth rate and predict future data growth.
[0084] Finally, a prediction model is constructed by combining data duplication and growth features. For example, a neural network model (such as an LSTM neural network) can be used to train the model using the extracted features as input to predict the data pressure of the current day in the future.
[0085] S104, when the future data pressure is less than or equal to the preset pressure threshold, the data volume of the current day in the target time period is compared with the data volume of each historical day in the target time period to determine the first data concentration of the current day in the target time period.
[0086] In this embodiment, the first data concentration is used to represent the degree to which data is concentrated during the target time period of the current day.
[0087] As an example, the server sets a reasonable preset pressure threshold based on the hardware resources of the data center (such as storage capacity, network bandwidth, processing power, etc.) and business needs.
[0088] If the future data pressure is less than or equal to the preset pressure threshold, it is time to consider full data synchronization. At this point, calculate the similarity between the data volume of the current day at the target time period and the data volume of each historical day at the target time period. Specifically, methods such as cosine similarity and Euclidean distance can be used for similarity calculation.
[0089] Then, based on the similarity calculation results, the concentration of the first data for the current day in the target time period is determined. For example, if the average similarity between the current day's data volume and the data volume of each historical day is high, it means that the data distribution of the current day is more similar to the data distribution of the historical days. Therefore, based on the concentrated distribution of the data of the historical days in the target time period, the concentration of the first data for the current day in the target time period is determined.
[0090] As another example, when the future data pressure is greater than the preset pressure threshold, it means that the flow pressure corresponding to the future data is greater. At this time, it can be directly determined to adopt the incremental synchronization strategy for data security disaster recovery.
[0091] S105: Determine a data synchronization strategy at a current moment based on the first data concentration, so as to perform data security disaster recovery for the data center through the data synchronization strategy.
[0092] In this embodiment, the data synchronization strategy includes a full synchronization strategy and an incremental synchronization strategy. The full synchronization strategy completely copies all data, while the incremental synchronization strategy only synchronizes data that has changed since the last synchronization.
[0093] As an example, when the concentration of the first data is less than or equal to the set threshold, it means that the data performance in the target time period at the current moment shows a relatively small concentrated trend after comparison with the historical same time period. Therefore, under the premise of the future data pressure analyzed at the current moment, you can choose to adopt a full synchronization strategy for future data, synchronize all data in the data center to the disaster recovery system, and ensure that the data in the disaster recovery system is completely consistent with the main data center; on the contrary, if the concentration of the first data is greater than the set threshold, only an incremental synchronization strategy is adopted, and only the data that has changed on the current day is synchronized, reducing the time and resource consumption of data synchronization to avoid network delays.
[0094] As an optional embodiment, Figure 2 As shown, S102 may specifically include the following S201 to S203.
[0095] S201, determining the second data concentration per hour based on the data volume per hour of each historical day in the historical time period;
[0096] S202, performing curve fitting on the second data concentration every hour to obtain a target curve;
[0097] S203: Divide adjacent maximum value points in the target curve into a time stage to obtain multiple time stages.
[0098] In this embodiment, the second data concentration is used to represent the average degree of concentrated distribution of data in the corresponding hours of each historical day within the historical time period.
[0099] As an example, the server calculates the proportion of hourly data volume in the total data volume of the corresponding historical day, and then averages the proportions corresponding to the same hours in each historical day to obtain the second data concentration of the corresponding hour.
[0100] Next, select an appropriate fitting function. Common fitting functions include polynomial functions, exponential functions, and logarithmic functions. Specifically, by observing a scatter plot of the second data concentration, you can initially determine the distribution trend of the data and select an appropriate fitting function. For example, if the data exhibits obvious periodic fluctuations, consider using a sine or cosine function for fitting; if the data shows a clear growth or decay trend, consider using an exponential function for fitting.
[0101] Then, the selected fitting function is used to determine the parameters of the fitting function through an optimization algorithm such as the least squares method, so that the error between the fitting curve and the actual data points is minimized.
[0102] Finally, in the fitted target curve, record the hours corresponding to the maximum points found. The range of hours between two adjacent maximum points is divided into a time period. For example, if the maximum points appear at the 3rd hour and the 15th hour respectively, then the period between the 3rd hour and the 15th hour (including the 3rd hour and excluding the 15th hour) is a time period.
[0103] For example, Figure 3 The figure shows a schematic diagram of a target curve. The X-axis represents time, and the Y-axis represents the concentration of the second data. Points B, D, and F are the maximum points in the target curve. The time period between points B and D is one time period, and the time period between points D and F is another time period. The remaining time period between points A and B is also divided into one time period, thereby dividing the target curve into multiple time periods.
[0104] This embodiment determines the second data concentration based on the hourly data volume of each historical day within the historical time period, comprehensively considering the data of the same period on different dates, and avoiding the one-sidedness that may be caused by relying solely on single-day data. In this way, the accuracy of the time period division can be improved.
[0105] As an optional embodiment, S201 may specifically include:
[0106] The proportion of the data volume of each historical day at the nth hour to the total data volume of each historical day is determined as the data volume prominence of each historical day at the nth hour, where n is a positive integer;
[0107] The data concentration coefficient of the nth hour is determined by using the difference in the prominence of the data volume of adjacent historical days at the nth hour;
[0108] The second data concentration at the nth hour is determined by using the minimum value of the prominence of each data amount and the data concentration coefficient.
[0109] In this embodiment, the prominence of the data volume at the nth hour of a historical day is the ratio of the data volume at the nth hour of the historical day to the total data volume of the historical day. For example, if the total data volume of a historical day is Z, and the data volume at the nth hour of the historical day is z, then the prominence of the data volume at the nth hour of the historical day is z divided by Z.
[0110] As an example, the data concentration coefficient can be determined by the following formula 1:
[0111]
[0112] In formula 1, U n It is used to represent the data concentration coefficient of the nth hour, and M is used to represent the total number of historical days in the historical time period.m,n It is used to represent the prominence of the data volume at the nth hour on the mth day, q m+1,n It is used to characterize the prominence of the data volume at the nth hour on the m+1th day, and norm is used to characterize the normalization process. It should be noted that in order to ensure that the calculation results are meaningful, when performing fractional operations in the embodiment of the present invention, when the denominator is 0, a parameter adjustment factor greater than 0 needs to be added to the denominator to prevent the denominator from being 0. The value of the parameter adjustment factor is set by the implementer according to the actual situation, and this application does not impose any special restrictions.
[0113] In scenarios with large-scale data changes, the data volume typically remains constant over a specific period. Therefore, we can analyze the balanced trend of the prominence of data volume at the same hour across each historical day within the historical time period. This means that the more prominent the data volume at the same hour across multiple days, the more likely it is that the data across multiple days and hours exhibits a trend that appears in the dataset at that time.
[0114] Then, the second data concentration can be specifically determined by the following formula 2:
[0115] Q n =(min m∈[1,M] q m,n )*(1+U n ) Formula 2
[0116] In formula 2, Q n The second data concentration used to characterize the nth hour, U n Used to characterize the data concentration coefficient of the nth hour. m,n It is used to represent the prominence of the data volume at the nth hour on the mth day. M is used to represent the total number of historical days in the historical time period. m∈[1,M] q m,n Used to represent the minimum value of the data volume prominence of each historical day.
[0117] This embodiment utilizes the hourly data volume for each historical day within a historical time period to accurately determine the hourly concentration of the second data. This allows for the generation of an accurate target curve, which facilitates subsequent use of the target curve to rationally divide time periods, improving the accuracy of time period division.
[0118] As an optional embodiment, Figure 4 As shown, S103 may specifically include the following S401 to S403.
[0119] S401, determining a target data repetition degree for a target time period based on data repetition characteristics for a target time period;
[0120] S402, performing mean processing on the second data concentration every hour in the target time period to obtain the first data growth rate in the target time period;
[0121] S403: predict future data pressure using the target data duplication rate and the first data growth rate.
[0122] In this embodiment, the target data repetition degree is used to characterize the degree of data repetition in the target time period; the first data growth degree is used to characterize the degree of data growth in the target time period on historical days within the historical time period.
[0123] As an example, the server calculates the frequency and pattern of data repetition for the target time period, such as the number of repetitions per day during the target time period. Specifically, the degree of repetition can be expressed as the ratio of repeated data to total data. This means dividing the number of repeated data on each historical day during the target time period by the total amount of data for that day during the target time period, and then calculating the average to obtain the target data repetition rate for the target time period.
[0124] Then, the second data concentration degree of each hour in the target time period is extracted, and the second data concentration degree of each hour in the target time period is averaged to obtain the first data growth degree in the target time period.
[0125] Next, select an appropriate prediction model, such as a linear regression model or a neural network model. For example, if a linear relationship exists between future data pressure, the target data duplication rate, and the first data growth rate, choose a linear regression model. Use historical data to train the prediction model, using the historical target data duplication rate and the first data growth rate as input and the corresponding actual data pressure as output. Determine the model parameters using methods such as least squares.
[0126] Finally, the currently calculated target data duplication and the first data growth are substituted into the trained prediction model to calculate the predicted value of future data pressure.
[0127] As another example, the future data pressure can be specifically determined by the following formula 3:
[0128] Y=D*(1-C)+(1-D)*Z Formula 3
[0129] In Formula 3, Y is used to represent future data pressure, D is used to represent the bandwidth occupancy at the current moment, C is used to represent the target data duplication, and Z is used to represent the first data growth rate.
[0130] Among them, the greater the bandwidth occupancy D at the current moment and the smaller the target data repetition in the target time period, the greater the future data pressure; the smaller the bandwidth occupancy D at the current moment, but the greater the first data growth in the target time period, the greater the future data pressure.
[0131] Through this embodiment, the target data repetition degree and the first data growth degree are combined to predict future data pressure, which fully considers the repetitive characteristics of the data and the data growth trend within the time period, avoids the deviation that may be caused by single factor prediction, and thus improves the accuracy of future data pressure prediction.
[0132] As an optional embodiment, S401 may specifically include:
[0133] For each historical day, perform the following steps: perform standard deviation processing on the frequency of occurrence of each data type within the target time period of the historical day to obtain the first repetition coefficient of the historical day;
[0134] The first repetition coefficient of each historical day is averaged to obtain the first data repetition degree of the target time period;
[0135] For each data type, perform the following steps: perform standard deviation processing on the frequency of occurrence of the data type on each historical day to obtain the second repetition coefficient of the data type;
[0136] Perform mean processing on the second repetition coefficient of each data type to obtain the second data repetition degree of the target time period;
[0137] A target data repetition rate in a target time period is determined by using the first data repetition rate and the second data repetition rate.
[0138] In this embodiment, the first data repetition coefficient reflects the discrete degree of the frequency of occurrence of the data type on historical days within the target time period, and the second data repetition coefficient reflects the discrete degree of the frequency of occurrence of the data type on each historical day within the target time period.
[0139] The first data repeatability reflects the fluctuation of the data type within each historical day, and the second data repeatability reflects the stability of the data type between different historical days. The combination of the two can more accurately describe the repetitive pattern of the data.
[0140] As an example, the server collects data on each historical day within a target time period and classifies the data into multiple data types according to the differences in the data.
[0141] For each historical day, count the number of occurrences of each data type within the target time period and calculate its frequency of occurrence. Frequency of occurrence = number of occurrences of a data type / total amount of data within the target time period for that historical day. Then, for each historical day, calculate the standard deviation of the frequency of occurrence of each data type within the target time period to obtain the first repetition coefficient for each historical day. Then, average the first repetition coefficients for each historical day to obtain the first data repetition degree for the target time period.
[0142] For each data type, we compiled its frequency of occurrence on each historical day. We then calculated the standard deviation of each data type's frequency of occurrence on each historical day to obtain the second repetition coefficient for each data type. We then averaged the second repetition coefficients for each data type to obtain the second data repetition rate for the target time period.
[0143] Finally, the target data repetition rate for the target time period can be determined by combining the first data repetition rate and the second data repetition rate using a weighted average method. Alternatively, the target data repetition rate for the target time period can be obtained by multiplying the first data repetition rate and the second data repetition rate.
[0144] This embodiment analyzes the data from two dimensions, historical date and data type, and calculates the first data repetition rate and the second data repetition rate, respectively. This allows for a more comprehensive characterization of the repetitive characteristics of the data during the target time period, thereby improving the accuracy of the data repetition rate calculation.
[0145] As an optional embodiment, Figure 5 As shown, S104 may specifically include the following S501 to S502.
[0146] S501, comparing the second data concentration of each hour of the current day within the target time period with the second data concentration of each historical day within the target time period to obtain a data concentration difference;
[0147] S502: Perform mean processing on the concentration differences of each data set to obtain the first data concentration of the current day at the target time stage.
[0148] In this embodiment, the data concentration difference represents the difference between the second data concentration of the current day within the target time period and the second data concentration of the corresponding historical day within the target time period. For example, if the second data concentration of the current day at the nth hour within the target time period is 0.7, and the second data concentration of the historical day at the nth hour within the target time period is 0.6, then the corresponding data concentration difference is 0.1.
[0149] As an example, the server calculates the data concentration difference between the second data concentration corresponding to the current day and each historical day for each hour in the target time period.
[0150] Finally, the concentration differences of each data set are averaged to obtain the concentration of the first data set at the target time stage of the current day.
[0151] This embodiment calculates the data concentration difference, visually reflecting the difference between the data concentration for each hour of the current day within the target time period and the corresponding hour of each historical day. This quantitative approach helps to more clearly understand the uniqueness of the current day's data and provides a basis for selecting subsequent data synchronization strategies.
[0152] As an optional embodiment, Figure 6 As shown, after S105, the data security disaster recovery method of the data center may further include the following S601 to S603.
[0153] S601, predicting the future packet loss rate based on the number of packet losses on each historical day within the target time period;
[0154] S602, determining a frequency adjustment coefficient for the frequency of synchronous data transmission within a target time period on the current day based on the future packet loss rate;
[0155] S603: Using the frequency adjustment coefficient, adjust the frequency of synchronous data transmission within the target time period of the current day.
[0156] In this embodiment, the frequency adjustment coefficient is used to represent a coefficient for adjusting the frequency of synchronous data transmission within a target time period on the current day.
[0157] As an example, the server first collects packet loss data for each historical day within a target time period. For example, in a network communication scenario, the number of packet losses within a specific time period (e.g., 9:00-11:00 AM) is recorded for each historical day. The average number of packet losses within the target time period is then calculated.
[0158] Next, select an appropriate forecasting model based on the data characteristics and forecasting requirements. Common models include time series analysis models (such as ARIMA models), regression models (such as linear regression and polynomial regression), or machine learning models (such as neural networks). For example, if there is a linear relationship between the future packet loss rate and the average number of packet losses per day in history, a linear regression model is selected.
[0159] Then, the selected prediction model is trained using historical data to determine the parameters of the prediction model. Finally, the trained prediction model is applied to the current average number of packet losses to predict the future packet loss rate.
[0160] Then, based on the future packet loss rate, define adjustment rules for the data synchronization frequency. For example, you can set a threshold range. When the future packet loss rate is below a certain threshold, the data synchronization frequency is appropriately increased; when the future packet loss rate is above a certain threshold, the data synchronization frequency is appropriately reduced.
[0161] Finally, the data synchronization transmission frequency of the current day in the target time period is adjusted according to the frequency adjustment coefficient to obtain the adjusted data synchronization transmission frequency. Specifically, the data synchronization transmission frequency of the current day in the target time period can be adjusted by the following formula 4:
[0162] P = p1 + (p2 - p1) * (1 - F) Formula 4
[0163] In Formula 4, P is used to represent the adjusted data synchronization transmission frequency, p1 is used to represent the minimum value of the data synchronization transmission frequency, p2 is used to represent the maximum value of the data synchronization transmission frequency, (p2-p1) is used to represent the adjustable range of the data synchronization transmission frequency, and F is used to represent the frequency adjustment coefficient.
[0164] This embodiment adjusts the frequency of synchronous data transmission based on the predicted future packet loss rate, enabling data transmission to better adapt to changing network conditions. When a high packet loss rate is predicted, reducing the transmission frequency can reduce data packet collisions and loss, thereby increasing the success rate of data transmission.
[0165] As an optional embodiment, S601 may specifically include:
[0166] Obtain the number of packet losses for each historical day within the target time period, and the second data growth rate for the current day within the target time period;
[0167] The number of packet losses on each historical day within the target time period is averaged to obtain the average number of packet losses on the historical day;
[0168] Comparing the second data growth rate with the first data growth rate to obtain a growth rate difference;
[0169] The future packet loss rate is determined by using the growth difference and the average number of packet losses.
[0170] In this embodiment, the second data growth rate is used to represent the degree of data growth during the target time period on the current day. Specifically, the second data concentration for each hour of the current day within the target time period is extracted and averaged to obtain the second data growth rate for the current day during the target time period.
[0171] As an example, the server collects the number of packet losses in the target time period of each historical day and the second data growth rate in the target time period of the current day.
[0172] Then, the number of packet losses on each historical day within the target time period is averaged to calculate the average number of packet losses on the historical day; at the same time, the second data growth rate is compared with the first data growth rate to obtain the growth rate difference.
[0173] Then, multiply the growth difference by the average number of packet losses to obtain the initial packet loss rate for the current day during the target time period. Finally, multiply the initial packet loss rate by the future data pressure to obtain the future packet loss rate.
[0174] This embodiment comprehensively considers both data growth and historical packet loss, enabling a more comprehensive assessment of future packet loss rates. Changes in data growth reflect changes in network load or traffic volume, while historical packet loss provides historical information on network stability. Combining these two factors allows for a more accurate prediction of future packet loss, thereby improving the accuracy of future packet loss rate calculations.
[0175] As an optional embodiment, S602 may specifically include:
[0176] Obtaining the bandwidth occupancy at the current moment and a packet loss data sequence for a target time period, wherein the packet loss data sequence includes each packet loss moment within the target time period;
[0177] Using the packet loss data sequence of the target time period, determining the packet loss time interval characteristic value of the target time period;
[0178] The frequency adjustment coefficient is determined using the bandwidth usage, the characteristic value of the packet loss time interval, and the future packet loss rate.
[0179] In this embodiment, the packet loss data sequence includes each packet loss moment within the target time period, and each packet loss moment is arranged in order from early to late in the packet loss data sequence.
[0180] As an example, the server first obtains the bandwidth usage at the current moment, and constructs a corresponding packet loss data sequence based on the historical time period and the packet loss moments of the current day.
[0181] Next, the packet loss intervals between two consecutive packet losses are obtained from the packet loss data sequence, thereby obtaining several types of packet loss intervals. The occurrence count of each type of packet loss interval is then counted. The occurrence count of each type of packet loss interval is divided by the total number of packet loss intervals in the packet loss data sequence to obtain the representativeness of each type of packet loss interval.
[0182] Then, the packet loss time interval representativeness of each type of packet loss time interval and the interval length of each type of packet loss time interval are combined using the weighted average method and normalized to comprehensively determine the packet loss time interval characteristic value of the target time period.
[0183] Finally, the frequency adjustment coefficient is determined using the following formula 5 using the bandwidth occupancy, the packet loss interval characteristic value, and the future packet loss rate:
[0184] F=T*H+(1-T)*D Formula 5
[0185] In Formula 5, F represents the frequency adjustment coefficient of the data synchronization transmission frequency during the target time period on the current day, T represents the characteristic value of the packet loss time interval during the target time period, H represents the future packet loss rate, and D represents the bandwidth occupancy at the current moment.
[0186] Among them, the larger the packet loss time interval characteristic value T in the target time phase and the greater the future packet loss rate, the larger the frequency adjustment coefficient should be; the smaller the packet loss time interval characteristic value T in the target time phase, but the greater the bandwidth occupancy at the current moment, the larger the frequency adjustment coefficient should be.
[0187] This embodiment utilizes bandwidth utilization, packet loss interval characteristic values, and future packet loss rates to accurately determine the frequency adjustment coefficient for the data synchronization transmission frequency during the target time period of the current day. This facilitates subsequent adjustment of the data synchronization transmission frequency during the target time period based on the frequency adjustment coefficient, thereby enabling data transmission to better adapt to changes in network conditions.
[0188] A data security disaster recovery method based on a data center. Accordingly, the present invention also provides a specific embodiment of a data security disaster recovery system for a data center.
[0189] like Figure 7 , a schematic diagram of a data center data security disaster recovery system is provided. The data center data security disaster recovery system 700 includes a data volume acquisition module 710, a stage division module 720, a pressure prediction module 730, a concentration determination module 740, and a strategy determination module 750.
[0190] The data volume acquisition module 710 is used to obtain the data volume of the data center on the current day and the data volume of each historical day in the historical time period;
[0191] The stage division module 720 is used to divide a day into multiple time stages based on the data volume of each historical day in the historical time period, and the data volume per hour in the time stage shows consistent fluctuation;
[0192] Pressure prediction module 730, for predicting future data pressure based on data duplication characteristics and data growth characteristics of a target time period, where the target time period is the time period of the current moment of the current day;
[0193] The concentration determination module 740 is configured to compare the data volume of the current day in the target time period with the data volume of each historical day in the target time period to determine the first data concentration of the current day in the target time period when the future data pressure is less than or equal to a preset pressure threshold;
[0194] The policy determination module 750 is configured to determine a data synchronization policy at a current moment based on the first data concentration, so as to perform data security disaster recovery for the data center through the data synchronization policy.
[0195] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.
[0196] It should also be noted that the exemplary embodiments described herein describe methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the steps described above. In other words, the steps may be performed in the order described in the embodiments, or in a different order, or several steps may be performed simultaneously.
[0197] The above description is only a specific embodiment of the present invention. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention.
Claims
1. A data security disaster recovery method for a data center, characterized in that: The method comprises: Obtain the data volume of the data center on the current day and the data volume of each historical day in the historical time period; Based on the data volume of each historical day within the historical time period, a day is divided into multiple time periods, and the data volume per hour within the time period shows consistent fluctuation; Predicting future data pressure based on data duplication characteristics and data growth characteristics of a target time period, where the target time period is the time period of the current moment of the current day; When the future data pressure is less than or equal to the preset pressure threshold, the data volume of the current day in the target time period is compared with the data volume of each of the historical days in the target time period to determine the first data concentration of the current day in the target time period; Based on the first data concentration, the data synchronization strategy at the current moment is determined, so that data security disaster recovery is performed on the data center through the data synchronization strategy.
2. The data center data security disaster recovery method according to claim 1, characterized in that: The data volume of each historical day within the historical time period is divided into multiple time periods, including: Determining the second data concentration per hour based on the data volume per hour of each historical day within the historical time period; Performing curve fitting on the concentration of the second data per hour to obtain a target curve; The time intervals between adjacent maximum value points in the target curve are divided into one time period, thereby obtaining a plurality of time periods.
3. The data center data security disaster recovery method according to claim 2, characterized in that: The determining of the second data concentration per hour based on the data volume per hour of each historical day within the historical time period includes: Determine the data volume prominence of each historical day at the nth hour as the proportion of the data volume of each historical day at the nth hour to the total data volume of the historical day, where n is a positive integer; Determine the data concentration coefficient of the nth hour by using the difference in the prominence of the data volumes of adjacent historical days at the nth hour; The second data concentration in the nth hour is determined by using the minimum value of each of the data amount prominences and the data concentration coefficient.
4. The data center data security disaster recovery method according to claim 1, characterized in that: The prediction of future data pressure based on the data duplication characteristics and data growth characteristics of the target time period includes: Determining a target data repetition degree for the target time period based on the data repetition characteristics of the target time period; Performing mean processing on the second data concentration every hour in the target time period to obtain the first data growth rate in the target time period; The future data pressure is predicted using the target data duplication rate and the first data growth rate.
5. The data center data security disaster recovery method according to claim 4, characterized in that: The determining of the target data repetition degree of the target time period based on the data repetition feature of the target time period includes: For each of the historical days, respectively performing: performing standard deviation processing on the occurrence frequency of each data type of the historical day within the target time period to obtain a first repetition coefficient of the historical day; Performing mean processing on the first repetition coefficients of each of the historical days to obtain the first data repetition degree of the target time period; For each of the data types, respectively performing: performing standard deviation processing on the occurrence frequency of the data type on each of the historical days to obtain a second repetition coefficient of the data type; Performing mean processing on the second repetition coefficients of the data types to obtain the second data repetition degree of the target time period; The target data repetition rate for the target time period is determined by using the first data repetition rate and the second data repetition rate.
6. The data center data security disaster recovery method according to claim 1, characterized in that: Comparing the data volume of the current day in the target time period with the data volume of each of the historical days in the target time period to determine the first data concentration of the current day in the target time period includes: Comparing the second data concentration of each hour of the current day within the target time period with the second data concentration of each historical day within the target time period to obtain a data concentration difference; The concentration differences of each data set are averaged to obtain the first data concentration of the current day in the target time period.
7. The data center data security disaster recovery method according to claim 1, characterized in that: After comparing the data volume of the current day in the target time period with the data volume of each of the historical days in the target time period to determine the first data concentration of the current day in the target time period, the method further includes: Predicting a future packet loss rate based on the number of packet losses on each of the historical days within the target time period; Determining, based on the future packet loss rate, a frequency adjustment coefficient for the frequency of synchronous data transmission on the current day within the target time period; The frequency adjustment coefficient is used to adjust the frequency of synchronously sending data within the target time period of the current day.
8. The data center data security disaster recovery method according to claim 7, characterized in that: The predicting of the future packet loss rate based on the number of packet losses on each of the historical days within the target time period includes: Obtaining the number of packet losses for each of the historical days within the target time period, and the second data growth rate of the current day within the target time period; Performing average processing on the number of packet losses on each of the historical days within the target time period to obtain the average number of packet losses on the historical day; Comparing the second data growth rate with the first data growth rate to obtain a growth rate difference; The future packet loss rate is determined by using the growth rate difference and the average number of packet losses.
9. The data center data security disaster recovery method according to claim 7, characterized in that: The determining, based on the future packet loss rate, a frequency adjustment coefficient for the frequency of synchronous data transmission within the target time period on the current day includes: Obtaining the bandwidth occupancy at the current moment and a packet loss data sequence of the target time period, wherein the packet loss data sequence includes each packet loss moment within the target time period; Determining a packet loss time interval characteristic value of the target time phase by using the packet loss data sequence of the target time phase; The frequency adjustment coefficient is determined by using the bandwidth occupancy, the packet loss time interval characteristic value, and the future packet loss rate.
10. A data security disaster recovery system for a data center, characterized in that: The system comprises: The data volume acquisition module is used to obtain the data volume of the data center on the current day and the data volume of each historical day in the historical time period; A stage division module is used to divide a day into multiple time stages based on the data volume of each historical day in the historical time period, and the data volume per hour in the time stage shows consistent fluctuation; A pressure prediction module, configured to predict future data pressure based on data duplication characteristics and data growth characteristics during a target time period, wherein the target time period is the time period at the current moment of the current day; a concentration determination module configured to compare the data volume of the current day in the target time period with the data volume of each of the historical days in the target time period to determine a first data concentration of the current day in the target time period when the future data pressure is less than or equal to a preset pressure threshold; A policy determination module is used to determine the data synchronization policy at the current moment based on the first data concentration, so as to perform data security disaster recovery for the data center through the data synchronization policy.
Citation Information
Patent Citations
Data backup method, data backup system, electronic device and readable storage medium
CN113076224A
Backup data processing method and device and server
CN116521449A
Hierarchical distribution model warehouse synchronization system and method
CN119988503A
Disaster recovery method, apparatus and system
WO2025077634A1