File backup method and system of disaster recovery system
By dynamically adjusting the segmentation threshold and blocking strategy of the disaster recovery system of the power enterprise, combining file access mode and hotspot data transfer, the blocking granularity is optimized, and the efficiency and integrity problems of traditional backup methods when facing dynamic file systems are solved, and efficient incremental backup and resource optimization are achieved.
Patent Information
- Application Number
- CN202510490318.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-18
AI Technical Summary
When facing a dynamically changing file system, traditional disaster recovery systems cannot adapt to the fluctuations in file changes frequency, especially when hot data areas are transferred, backup efficiency decreases or resource wastes, making it difficult to balance efficiency and integrity, and existing solutions are difficult to flexibly adjust segmentation thresholds and blocking strategies.
By obtaining the metadata update rate and modification frequency of the disaster recovery system files of the power enterprise, calculating the frequency distribution of changes, dynamically adjusting the segmentation threshold and blocking strategy, combining file access mode changes and hot data transfer signs, a blocking algorithm based on the data modification frequency is adopted to optimize the blocking granularity, and adaptively adjust the backup strategy to improve efficiency.
It realizes efficient incremental backup in the disaster recovery system of power enterprises, improves data backup efficiency and resource utilization, enhances the system's ability to adapt to market changes, and ensures the flexibility and consistency of backup strategies.
Smart Images

Figure CN120386668A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a file backup method and system for a disaster recovery system. Background Art
[0002] Disaster recovery system research, as a key branch of information technology, directly impacts the core requirements of data security and business continuity, making its critical importance self-evident. With the rapid growth and diversification of data, efficient backup strategies have become essential for ensuring system stability and resilience. Traditional backup methods, often based on fixed cycles or static segmentation, meet basic requirements to a certain extent but often fall short in the face of dynamically changing file systems. This limitation primarily manifests itself in their inability to adapt to fluctuations in the frequency of file changes. In particular, when hotspot data areas shift, existing solutions struggle to flexibly adjust, resulting in reduced backup efficiency and wasted resources. Current solutions often struggle to respond to file change patterns when handling complex file systems. In particular, when the rate of file metadata updates exhibits a nonlinear correlation with actual content changes, traditional static segmentation strategies struggle to accurately capture these patterns. Furthermore, shifting access patterns and varying data block modification granularity further exacerbate the adaptability challenges of backup systems. These shortcomings result in backup processes that are either overly redundant or miss critical data, making it difficult to strike a balance between efficiency and integrity. A key challenge in this research area lies in dynamically sensing and adapting to the changing nature of file systems. Specifically, the correlation between the file metadata update rate and content changes, the migration of hot data areas from structured to unstructured, and the dynamic adjustment of access patterns and modification granularity are technical factors that need to be addressed urgently. Because these factors have not been fully addressed, backup systems are often unable to optimize segmentation thresholds and block strategies in a timely manner when facing hotspot transfers, resulting in fluctuations in the efficiency of incremental backups and even data consistency risks. These difficulties not only increase the complexity of system design, but also place higher demands on the real-time performance and resource utilization of disaster recovery. Therefore, how to design a file classification engine driven by the frequency of file changes that can dynamically adjust the segmentation thresholds and block strategies based on the correlation between metadata updates and content changes, and ensure that it can still maintain efficient incremental backups when hot data areas are transferred, has become a key issue in this study. Summary of the Invention
[0003] The present invention provides a file backup method for a disaster recovery system, which mainly includes:
[0004] Obtain metadata update rates and modification frequencies for power company disaster recovery system files, calculate change frequency distribution, determine initial segmentation threshold ranges, analyze timing fluctuations, and adjust segmentation threshold ranges if they exceed preset thresholds.
[0005] Extract the real-time data of the metadata update rate and the modification frequency from the adjusted segmented threshold range, calculate the Pearson coefficient between the two, and judge the priority of optimizing the chunking strategy;
[0006] Monitor the changes in the enterprise file access pattern and the offset of the modification granularity, determine the signs of hot data transfer. If the offset is greater than the preset threshold, mark the enterprise file as an object for dynamic adjustment to obtain a list of enterprise files to be optimized;
[0007] According to the list of enterprise files to be optimized, use a chunking algorithm based on the data modification frequency to divide data chunks, compare the modification frequency and the metadata update rate of adjacent data chunks, and determine the target chunking granularity;
[0008] Based on the target chunking granularity, calculate the redundancy degree of the data chunks in the incremental backup. If it is lower than the preset redundancy threshold, retain the current chunking strategy. If not, readjust the segmented threshold to obtain an updated chunking scheme;
[0009] Obtain the enterprise file change frequency and the dynamic characteristics of hot data transfer of the updated chunking scheme, and dynamically iteratively optimize the segmented threshold and the chunking strategy through an adaptive adjustment algorithm. Combine the volatility of the power market and the changes in the supply-demand relationship to judge the improvement amplitude of the backup efficiency;
[0010] Obtain the latest data on the changes in the enterprise file access pattern. If it is consistent with the hot data transfer trend, apply the optimized chunking scheme to the subsequent backup process to obtain the final backup strategy;
[0011] Through the final backup strategy, continuously monitor the fluctuations of the enterprise file change frequency and the metadata update rate, use dynamic characteristics to capture the update and modification frequency of power trading-related data, and determine the adaptive adjustment ability of the enterprise disaster recovery system.
[0012] Further, obtain the metadata update rate and modification frequency of the power enterprise disaster recovery system files, calculate the change frequency distribution, determine the initial segmentation threshold range, analyze the time series fluctuations. If the preset threshold is exceeded, adjust the segmentation threshold range, including: obtaining the file metadata update records from the disaster recovery database, calculating the metadata update frequency value according to the data modification timestamp, obtaining the total amount of metadata changes through time series statistics of data updates, and statistically obtaining the file storage usage status based on the storage capacity monitoring value. Use a time series fluctuation detector to perform a sliding window analysis on the metadata update frequency value, calculate the metadata modification frequency distribution map according to the total amount of changes, and obtain the change fluctuation value of the metadata in different time periods. Construct a time series feature vector for the metadata change fluctuation value, extract the frequency feature, time series feature, and fluctuation feature in the vector through a clustering operator to obtain the initial segmentation threshold for metadata changes. Divide the metadata update frequency interval according to the initial segmentation threshold, use a regression predictor to calculate the probability distribution of metadata changes in the interval, and obtain the metadata change warning value. Retrieve the preset threshold range parameters from the warning database, and determine whether the metadata change warning value exceeds the preset range by comparing the metadata change warning value with the preset threshold range parameters. If the metadata change warning value exceeds the preset range, recalculate the segmentation threshold range based on the probability distribution of changes to obtain the dynamic threshold interval for metadata updates.
[0013] Further, real-time data of the metadata update rate and the modification frequency are extracted from the adjusted segmented threshold range, and the Pearson coefficient between the two is calculated to determine the optimization priority of the chunking strategy, including: obtaining the time series records of the metadata update rate and the modification frequency from the segmented threshold database, counting the number of metadata chunks through a data chunk counter, and obtaining the basic statistical value of the real-time data chunks based on the number of chunks. A time series synchronization calculator is used to intercept the time window of the metadata update rate and the modification frequency, and the time series characteristics of the data chunks are calculated according to the basic statistical value of the real-time data chunks to obtain the data chunk association vector. The time series association characteristics between the metadata update rate and the modification frequency are extracted according to the data chunk association vector, the Pearson coefficient is obtained through a correlation coefficient calculator, and the correlation degree of the data chunks between the metadata update rate and the modification frequency is judged. The priority evaluation index is constructed based on the correlation degree of the data chunks, and the initial weight value of the chunking priority is generated through a data chunk priority calculator to obtain the chunking priority reference sequence. The preset priority threshold is retrieved from the priority database, and the optimization priority of the chunking strategy is judged according to the comparison result between the chunking priority reference sequence and the preset priority threshold. If the chunking priority reference sequence exceeds the preset priority threshold, the optimization weight value of the chunking strategy is recalculated based on the correlation degree of the data chunks to obtain the chunking priority adjustment sequence.
[0014] Further, monitor the changes in the enterprise file access pattern and the modification granularity offset, determine the signs of hot data transfer. If the offset is greater than the preset threshold, mark the enterprise file as a dynamic adjustment object to obtain a list of enterprise files to be optimized, including: obtaining the real-time electricity price data, electricity demand prediction data, and transaction contract data from the enterprise file database, calculating the file access heat value according to the access record log, counting the modification granularity offset through the modification record log, and obtaining the data activity index based on the access time series record. A heat tracker is used to perform a sliding window analysis on the file access heat value, calculate the file change rate according to the modification granularity offset, and obtain the file hot spot aggregation degree based on the data activity index. The hot data transfer feature vector is constructed according to the file hot spot aggregation degree, and the transfer feature sequences are extracted for the real-time electricity price data, electricity demand prediction data, and transaction contract data respectively. A gradient boosting decision tree is used to perform time series prediction on the transfer feature sequences, and the hot data transfer trend is judged through the file change rate to obtain the hot migration amplitude value. The preset offset threshold parameter is retrieved from the mark database, and it is judged whether the hot data transfer exceeds the preset offset threshold parameter according to the hot migration amplitude value. If the hot migration amplitude value exceeds the preset offset threshold parameter, mark the corresponding file as the dynamic adjustment object and add the dynamic adjustment object to the list of enterprise files to be optimized.
[0015] Further, according to the list of enterprise files to be optimized, a chunking algorithm based on data modification frequency is used to divide data chunks, compare the modification frequency of adjacent data chunks with the metadata update rate, and determine the target chunking granularity, including: obtaining the list of enterprise files to be optimized from the enterprise file database, counting the modification frequency value according to the data modification record, calculating the metadata update rate through the metadata log, and obtaining the initial size of the file chunks according to the modification time sequence record. An adaptive chunker is used to perform chunking processing on the enterprise files to be optimized according to the modification frequency value, calculate the distance between adjacent chunks for the chunking processing result, and obtain the in-block data association value through a chunk feature extractor. A feature vector of adjacent data chunks is constructed according to the in-block data association value, the modification frequency value and the metadata update rate are respectively extracted for the adjacent data chunks, and the feature difference degree between adjacent chunks is calculated. A time-series correlation calculator is used to analyze the feature difference degree between adjacent chunks, and the chunk boundary position is judged through the distance between adjacent chunks to obtain the chunk interval sequence. The preset granularity range parameter is retrieved from the data chunk database, and the target chunk size is calculated according to the chunk interval sequence. If the target chunk size exceeds the preset granularity range parameter, the chunk boundary position is recalculated until the target chunk size meets the preset granularity range parameter to obtain the target chunking granularity.
[0016] Further, calculate the data modification frequency of the enterprise files corresponding to the enterprise file list to be optimized, divide the enterprise files into multiple data blocks using the data modification frequency, compare the modification frequencies between adjacent data blocks, and at the same time compare the metadata update rates between adjacent blocks. According to the comparison results of the modification frequency and the metadata update rate, determine the optimal target block granularity, including: obtaining the enterprise file list to be optimized from the enterprise file database, calculating the data modification frequency according to the file modification records, statistically calculating the metadata update rate through the metadata records, and obtaining the initial size of the data block based on the time series records. Use a dynamic block splitter to divide the enterprise files to be optimized according to the data modification frequency, calculate the data block size value for the division result, and obtain the adjacent block data sequence through a data feature extractor. Extract the difference in modification frequency between blocks according to the adjacent block data sequence, calculate the difference in metadata update rate between blocks for the adjacent block data sequence, and obtain the change characteristics of the adjacent blocks based on the difference calculator. Retrieve the block division threshold parameter from the block database, calculate the block adjustment value according to the change characteristics of the adjacent blocks, and determine whether the block adjustment value meets the block division threshold parameter through a comparison calculator. If the block adjustment value exceeds the block division threshold parameter, recalculate the data block size value based on the change characteristics of the adjacent blocks. Use an optimization iterator to perform multiple rounds of comparison on the data block size value until the block adjustment value meets the block division threshold parameter to obtain the optimal target block granularity.
[0017] Further, calculate the redundancy degree of the data block in the incremental backup through the target block granularity. If it is lower than the preset redundancy threshold, retain the current block division strategy. If it is not lower, readjust the segmentation threshold to obtain an updated block division scheme, including: marking the data block according to the target block granularity, obtaining the backup data size from the incremental backup database, calculating the incremental backup frequency through the backup record, and obtaining the data block change cycle based on the backup time record. Analyze the backup data size using a data comparer, count the data block increment value for the incremental backup frequency, and obtain the data block redundancy degree through a redundancy detector. Construct the redundancy feature vector according to the data block redundancy degree, calculate the backup time interval for the data block change cycle to obtain the data block repetition feature. Retrieve the preset redundancy threshold parameter from the backup configuration library, calculate the adjacent block repetition rate according to the data block repetition feature, and determine whether the adjacent block repetition rate meets the preset redundancy threshold parameter through a comparison calculator. If the adjacent block repetition rate is lower than the preset redundancy threshold parameter, retain the target block granularity. If the adjacent block repetition rate is not lower than the preset redundancy threshold parameter, recalculate the block division interval value based on the data block change cycle. Generate the new block division scheme according to the block division interval value, and determine whether the adjacent block repetition rate in the new block division scheme meets the preset redundancy threshold parameter through a verification calculator.
[0018] Further, obtain the enterprise file change frequency and the dynamic characteristics of hot data transfer of the updated chunking scheme, dynamically iterate and optimize the segmentation threshold and chunking strategy through an adaptive adjustment algorithm, and judge the improvement amplitude of backup efficiency in combination with the volatility of the power market and the change of supply and demand relationship, including: obtaining the updated chunking scheme from the enterprise file database, calculating the file change frequency according to the file change record, counting the hot data transfer rate through the hot data monitor, and constructing the dynamic feature vector based on the change frequency. Use an adaptive optimizer to calculate the segmentation threshold adjustment value according to the dynamic feature vector, extract the data time series characteristics for the hot data transfer rate, and generate the market volatility index through a data volatility calculator. Calculate the supply and demand change amplitude according to the market volatility index, count the chunk size adjustment value for the data time series characteristics, and obtain the backup time series characteristics through a time series correlation calculator. Use an iterative optimizer to perform multiple rounds of calculations on the segmentation threshold adjustment value, judge the chunking optimization direction according to the supply and demand change amplitude, and obtain the chunking adjustment sequence. Retrieve the preset efficiency parameter from the backup configuration library, calculate the backup efficiency improvement value according to the chunking adjustment sequence, and judge whether the backup efficiency improvement value meets the preset efficiency parameter through a comparison calculator. If the backup efficiency improvement value does not meet the preset efficiency parameter, recalculate the segmentation threshold adjustment value based on the backup time series characteristics. Generate the new chunking scheme according to the segmentation threshold adjustment value, and perform efficiency verification on the new chunking scheme through a verification calculator until the backup efficiency improvement value meets the preset efficiency parameter.
[0019] Further, obtain the latest data on changes in the enterprise file access pattern. If it is consistent with the hot data transfer trend, apply the optimized chunking scheme to the subsequent backup process to obtain the final backup strategy, including: obtaining the latest file access records from the enterprise file database, calculating the file access pattern value based on the access logs, statistically analyzing the data heat index through a hot data monitor, and obtaining the access pattern sequence based on the access time series record. Use a hot spot tracker to calculate the hot data transfer trend based on the data heat index, statistically analyze the file access trend for the access pattern sequence, and determine the trend matching degree through a trend comparison calculator. If the matching degree between the hot data transfer trend and the file access trend exceeds the preset matching threshold, retrieve the optimized chunking scheme from the backup configuration library. Construct the backup execution sequence according to the optimized chunking scheme, extract the data processing parameters for the backup execution sequence, and generate the backup schedule through a data processor. Use a backup scheduler to perform time series allocation on the backup schedule, calculate the backup time window based on the data processing parameters, and obtain the backup execution schedule. Generate the final backup strategy based on the backup execution schedule, and verify the feasibility of the final backup strategy through a verification calculator to determine whether the final backup strategy meets the preset backup requirements.
[0020] Further, through the final backup strategy, continuously monitor the fluctuations in the enterprise file change frequency and metadata update rate, and use dynamic feature capture to obtain the update and modification frequency of power trading-related data to determine the adaptive adjustment ability of the enterprise disaster recovery system, including: setting the data sampling period according to the final backup strategy, obtaining the enterprise file change frequency from the enterprise file database, statistically analyzing the metadata update rate fluctuation value through metadata records, and obtaining the data modification characteristics based on the time series record. Use a feature extractor to perform fluctuation analysis on the enterprise file change frequency, calculate the time series fluctuation characteristics for the metadata update rate fluctuation value, and obtain the update and modification frequency of the power trading data through a fluctuation detector. Construct the dynamic feature vector based on the time series fluctuation characteristics, extract the data change pattern for the update and modification frequency of the power trading data, and obtain the disaster recovery adjustment parameters through a time series predictor. Retrieve the preset adjustment threshold from the disaster recovery database, and determine whether the system adjustment response value meets the preset adjustment threshold based on the disaster recovery adjustment parameters. Use an adaptive optimizer to dynamically calculate the system adjustment response value, and generate the disaster recovery system adjustment plan through a response calculator. If the system adjustment response value exceeds the preset adjustment threshold, recalculate the data sampling period based on the data change pattern. Update the adaptive adjustment parameters according to the disaster recovery system adjustment plan, and verify the rationality of the adaptive adjustment parameters through a verification calculator.
[0021] The present invention provides a file backup system for a disaster recovery system, mainly including:
[0022] A metadata monitoring module, which is used to obtain the metadata update rate and modification frequency of the files in the power enterprise disaster recovery system, calculate the change frequency distribution, determine the initial segmentation threshold range, analyze the time series fluctuations, and adjust the segmentation threshold range if it exceeds the preset threshold;
[0023] A segmentation threshold adjustment module, which is used to extract the real-time data of the metadata update rate and modification frequency from the adjusted segmentation threshold range, calculate the Pearson coefficient between the two, and judge the priority of optimizing the chunking strategy;
[0024] A chunking strategy optimization module, which is used to monitor the changes in the enterprise file access pattern and the offset of the modification granularity, determine the signs of hot data transfer, and if the offset is greater than the preset threshold, mark the enterprise files as objects to be dynamically adjusted to obtain a list of enterprise files to be optimized;
[0025] A dynamic adjustment marking module, which is used to divide data blocks according to the list of enterprise files to be optimized by using a chunking algorithm based on the data modification frequency, compare the modification frequencies and metadata update rates of adjacent data blocks, and determine the target chunking granularity;
[0026] A chunking granularity determination module, which is used to calculate the redundancy degree of data blocks in incremental backup through the target chunking granularity. If it is lower than the preset redundancy threshold, the current chunking strategy is retained. If it is not lower, the segmentation threshold is readjusted to obtain an updated chunking scheme;
[0027] A redundancy evaluation module, which is used to obtain the enterprise file change frequency and the dynamic characteristics of hot data transfer of the updated chunking scheme, dynamically iterate and optimize the segmentation threshold and chunking strategy through an adaptive adjustment algorithm, and judge the improvement amplitude of backup efficiency in combination with the volatility of the power market and the changes in the supply and demand relationship;
[0028] An iterative optimization module, which is used to obtain the latest data on changes in the enterprise file access pattern. If it is consistent with the hot data transfer trend, the optimized chunking scheme is applied to the subsequent backup process to obtain the final backup strategy;
[0029] A strategy application module, which is used to continuously monitor the fluctuations of the enterprise file change frequency and metadata update rate through the final backup strategy, capture the update and modification frequencies of power transaction-related data by using dynamic characteristics, and determine the adaptive adjustment ability of the enterprise disaster recovery system.
[0030] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0031] The present invention discloses a file backup method for a disaster recovery system. The method analyzes the changes in the metadata update rate and modification frequency of enterprise files, calculates the distribution characteristics, and dynamically adjusts the segmentation threshold. Combining the signs of hot data transfer and changes in file access patterns, it determines the list of files to be optimized. Adopting a block-based algorithm based on the modification frequency, it compares the characteristics of adjacent data blocks to determine the target granularity. By calculating the data redundancy in incremental backups, it iteratively optimizes the block scheme. Combining the volatility of the power market and changes in the supply-demand relationship, it evaluates the improvement in backup efficiency. Finally, it applies the optimized block scheme to the backup process and continuously monitors the fluctuations in the file change frequency and metadata update rate to achieve the adaptive adjustment of the disaster recovery system. The present invention can effectively improve the data backup efficiency and resource utilization rate of the disaster recovery system of power enterprises and enhance the system's adaptability to market changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flowchart of a file backup method for a disaster recovery system of the present invention.
[0033] Figure 2 It is a structural diagram of a file backup system for a disaster recovery system of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0034] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0035] As Figure 1 , a file backup method for a disaster recovery system in this embodiment may specifically include:
[0036] S101. After the disaster recovery function is started, obtain the metadata update rate and content modification frequency of files from the disaster recovery system of the power enterprise, calculate the change frequency distribution, and determine the initial segmentation threshold range. At the same time, analyze the time series fluctuations and adjust the segmentation threshold range when exceeding the preset threshold to adapt to dynamic changes.
[0037] In the embodiment of the present invention, as the execution subject of the backup method, first extract the file metadata update records and content modification logs through the disaster recovery database. The metadata update records usually contain information such as file names and modification timestamps, while the content modification logs reflect the changes in the actual data of the files. Taking a power enterprise as an example, the metadata update rate may fluctuate significantly due to power grid dispatching or trading peaks. For example, data from a certain power supply bureau shows that the update frequency during peak hours on weekdays can reach more than 300 times per hour. Calculate the metadata update frequency based on the modification timestamp, and combine the content modification logs to calculate the overall frequency distribution of file changes, forming a frequency distribution diagram reflecting the change rules in different time periods.
[0038] S1011. Extract the metadata update records and content modification logs from the disaster recovery database, calculate the metadata update frequency and content modification frequency using timestamps, generate a frequency distribution graph through sliding window analysis, and extract the time series feature vectors to determine the initial segmentation threshold. Use a time series analyzer to perform sliding window processing on the update frequency and modification frequency. For example, set a 60-minute window to capture the mutation characteristics during the business peak period. Extract frequency features such as the average number of updates, time series features such as periodic patterns, and fluctuation features such as intensity change trends from the frequency distribution graph through a clustering algorithm, and construct a multi-dimensional time series feature vector. According to these features, calculate the initial segmentation threshold. For example, in a substation scenario, the threshold is set to 200 times per hour to divide the metadata update frequency interval.
[0039] S1012. Divide the frequency interval according to the initial segmentation threshold, use a regression analyzer to predict the probability distribution of metadata changes within the interval, generate a change warning value, and compare it with the preset threshold range. If it exceeds the range, readjust the segmentation threshold range based on the probability distribution. The regression analyzer integrates the statistical laws of historical data and calculates the change probability of each frequency interval. For example, during equipment maintenance, the change warning value of a power grid system reaches 2.5 times the normal value. Retrieve the preset threshold range parameters from the warning database, usually set to 1.5 to 3 times the normal value, and determine whether adjustment is needed by comparing the warning value with the preset range. If it exceeds, for example, the warning value of a smart substation exceeds the standard for two consecutive hours due to technical transformation, recalculate the segmentation threshold according to the probability distribution in the past 5 days, and increase the upper limit to 2 times the original level to form a dynamic threshold interval to adapt to the change characteristics.
[0040] In the embodiment of the present invention, in the power dispatching system, the metadata update shows the tidal law of morning and evening peaks, which is highly correlated with the change of the power grid load. The adjustment of the dynamic threshold interval fully considers this periodic characteristic. In addition, the selection of the sliding window size is crucial for the sensitivity of anomaly detection. Taking a distribution system as an example, a 60-minute window can effectively identify the update peak caused by the sudden change of the device state, and the frequency distribution graph shows a bimodal characteristic, ensuring the accuracy of the initial segmentation threshold. Through adaptive adjustment, the effectiveness of the backup strategy can be maintained when the hot data is transferred, laying a foundation for subsequent optimization. It can be understood that the specific window size and feature extraction method can be set by technicians according to the actual scenario and are not overly limited.
[0041] In the embodiment of the present invention, the file backup method of the disaster recovery system optimizes the chunking strategy by dynamically analyzing the file change characteristics to improve the backup efficiency and data integrity. The following steps describe the technical implementation process in detail, and the specific parameters and algorithm details can be flexibly adjusted according to the actual scenario.
[0042] S102. After determining the adjusted segmented threshold range, extract the real-time data of the metadata update rate and content modification frequency from the power enterprise disaster recovery system, calculate the correlation coefficient between the two to evaluate the degree of data block association, and accordingly determine the priority ranking of the block strategy optimization.
[0043] In the embodiment of the present invention, first access the segmented threshold database to obtain the time series records of the metadata update rate and content modification frequency that reflect the dynamic changes of the file. These records usually contain information such as timestamps and update event types, and can reflect the system operation status. Taking a power supply bureau as an example, its database records show that the metadata update rate during the peak period of weekdays is about 160 times per hour, and the content modification frequency is 90 times per hour, while during the non-peak period, they drop to 50 times and 30 times per hour respectively. The real-time data is preliminarily processed by a data block statistic to count the current number of data blocks and generate basic statistical values. For example, the total number of data blocks in the operation data of a substation is 900, and these statistical values serve as the basis for subsequent analysis.
[0044] S1021. Extract the time series records of the metadata update rate and content modification frequency from the segmented threshold database, use the data block statistic to calculate the real-time number of data blocks and generate basic statistical values, and then use a time series analyzer to intercept the two types of frequency data through a time window to extract the time series feature vectors of the data blocks. The time series analyzer uses a fixed time window, such as 15 minutes, to synchronously intercept the metadata update rate and modification frequency to capture short-term fluctuation characteristics. For example, during the operation of a certain power distribution system, the metadata update rate surges to 280 times per hour at 10 am, while the modification frequency remains at 130 times per hour. This difference may be triggered by equipment status adjustment. Calculate the time series feature vectors based on the basic statistical values, including indicators such as the average update frequency and fluctuation amplitude, to provide data support for the correlation analysis.
[0045] S1022. Based on the time series feature vectors, use a correlation analyzer to calculate the Pearson correlation coefficient between the metadata update rate and content modification frequency, evaluate the association strength of the data blocks through this coefficient, and combine the association strength to construct a priority evaluation model to generate the optimization priority sequence of the block strategy. The correlation analyzer calculates the Pearson coefficient through statistical methods. For example, during normal operation, the coefficient of a certain power dispatching system stabilizes at 0.85, indicating a high degree of correlation; while during equipment maintenance, the coefficient drops to 0.45, reflecting a weakened association. Construct a priority evaluation model based on the Pearson coefficient. The model integrates the real-time requirements and business importance of the data blocks. For example, the priority weight of the real-time control instruction data block is set to 0.8, and the historical archive data block is set to 0.5. Obtain the preset threshold from the priority database, such as 0.7. If the coefficient is lower than this value, trigger the optimization mechanism to recalculate the block priority sequence to adjust the strategy.
[0046] During this process, the internal relationship between metadata updates and content modifications is revealed through correlation analysis. For example, in a certain intelligent substation, the change in the line switch status causes the metadata update rate to rise to 300 times per hour within a short period, and the Pearson coefficient reaches 0.90, indicating an enhanced correlation of data blocks. At this time, the chunking priority weight is increased to 0.88 to quickly respond to the changes. Correlation analysis not only relies on the Pearson coefficient but can also incorporate the characteristics of business scenarios. For example, the fluctuation range of peak hours in power trading data is from 0.78 to 0.92, and it stabilizes at around 0.85 during off-peak hours. This dynamic characteristic guides the optimization of chunking granularity.
[0047] S1023. Retrieve the preset priority threshold from the priority database, compare the chunking priority sequence with the threshold. If it exceeds the threshold, recalculate the optimized weight value based on the data block association strength, and finally generate an adjusted chunking priority sequence to guide the subsequent backup strategy optimization. The setting of the preset threshold needs to consider business requirements. For example, the threshold for control data blocks is 0.75. By comparison, it is found that 80% of the sequence values are lower than the threshold during a certain period, then the weight is readjusted according to the association strength. For example, in a regional power grid, the association of data blocks decreases during equipment maintenance, and the optimized weight increases from 0.6 to 0.82 to ensure that the backup strategy adapts to the changes.
[0048] In the embodiment of the present invention, the priority judgment of the chunking strategy is realized through the above steps. For example, in a power dispatching data center, the correlation fluctuation between metadata updates and modification frequencies during peak hours is obvious. A 15-minute window is used to capture features to ensure that the priority sequence reflects real-time requirements. In addition, the moving average method can be introduced into the calculation process of the Pearson coefficient to smooth noise. For example, in the data of a certain substation, the coefficient mean continuously drops from 0.80 to 0.55 within 30 minutes, triggering an optimization adjustment. This method effectively improves the adaptability of the chunking strategy to dynamic scenarios and lays a foundation for subsequent backup optimization. It can be understood that the length of the time window and the weight calculation method can be set by technicians according to actual needs and are not overly limited.
[0049] In the embodiment of the present invention, the file backup method of the disaster recovery system realizes efficient backup through dynamic monitoring and optimization, and is particularly suitable for high-dynamic data scenarios of power enterprises.
[0050] S103. During the backup process, continuously monitor the changes in the access pattern and the offset of the modification granularity in the files of power enterprises. By analyzing the transfer signs of hot data such as real-time electricity prices, electricity demand forecasts, and trading contracts, judge whether the offset exceeds the preset threshold and mark the relevant files as objects for dynamic adjustment, generating a list of enterprise files to be optimized.
[0051] In the embodiments of the present invention, real-time electricity price data, electricity demand forecast data, and transaction contract data are extracted from the enterprise file database, and the dynamic characteristics of the files are calculated using access record logs and modification logs. Taking a regional power grid as an example, the access frequency of real-time electricity price data can reach 300 times per minute during peak trading hours, while it is only 80 times during off-peak hours. The modification granularity offset increases to 3 times the normal value during peak periods, reflecting the significant impact of business load on data access. By analyzing these data, the behavioral characteristics of hot data are obtained, providing a basis for subsequent optimization.
[0052] S1031. Obtain relevant data on real-time electricity price, electricity demand forecast, and transaction contract from the enterprise file database. Calculate the file access heat value according to the access record log and combine the modification log to count the modification granularity offset. Then, generate a data activity index through time series analysis to evaluate the file hot spot aggregation degree. The access heat value reflects the frequency of file calls. For example, the access heat value of the electricity demand forecast data of a certain power supply bureau rises to 350 times per minute during the load adjustment period, and the modification granularity offset increases to 2.5 times the reference value due to forecast updates. Using the time series analysis method, for example, with a 10-minute window, calculate the activity index. The activity index of a certain transaction contract file during the settlement period reaches 0.90, much higher than 0.40 in normal days, indicating a significant increase in the hot spot aggregation degree.
[0053] S1032. Construct a hot data transfer feature vector based on the file hot spot aggregation degree. Extract time series feature sequences for real-time electricity price, electricity demand forecast, and transaction contract respectively, and use the gradient boosting decision tree to predict the transfer trend to calculate the hot spot migration amplitude value. The feature vector includes dimensions such as access frequency change and modification rate. For example, in the data of a certain substation, the time series correlation of the real-time electricity price in the feature sequence shows a decrease from 0.75 to 0.50 during the trading period, indicating a shift in the access pattern. The gradient boosting decision tree optimizes the prediction accuracy through multiple rounds of iteration. For example, the prediction results of a certain distribution system show that the migration amplitude value during equipment maintenance periods rises to 0.85, far exceeding 0.30 during normal operation, accurately capturing the hot spot transfer trend.
[0054] In this step, the change law of hot data is identified through dynamic monitoring. For example, in the power trading center, the access heat value of the transaction contract file surges from 100 times per minute to 400 times per minute before and after the monthly auction, the hot spot aggregation degree rises from 0.35 to 0.87, and the modification rate increases by 4 times. This kind of change is usually related to the business cycle. The gradient boosting decision tree can predict the transfer trend and generate the migration amplitude value by analyzing historical data and real-time features, providing a basis for marking.
[0055] S1033. Retrieve the preset offset threshold parameter from the marking database, compare the hot spot migration amplitude value with the threshold. If it exceeds the threshold, mark the corresponding file as an object for dynamic adjustment and add it to the list of enterprise files to be optimized to guide subsequent backup optimization. The preset threshold is usually set according to business requirements. For example, the threshold for real-time electricity price data is 0.70. When the detected migration amplitude value reaches 0.80 during a certain period, the relevant file is automatically marked. For example, in a certain distribution automation system, the migration amplitude value of the electricity price file before and after the peak-valley electricity price switching rises to 0.78, exceeding the threshold of 0.65. Mark it as a priority adjustment object and add it to the list to ensure that the backup strategy responds in a timely manner to changes in the access mode.
[0056] In the embodiment of the present invention, the above monitoring and marking mechanism effectively identifies the files to be optimized. For example, in a certain power dispatching system, the access popularity of the load forecasting file increases sharply during the peak period. The modified granularity offset increases to 3.2 times the normal value, and the migration amplitude value reaches 0.82, exceeding the preset threshold of 0.60. The relevant files are marked and added to the list of files to be optimized. This method ensures that the system can adapt to hot spot transfer and improves the pertinence and efficiency of backup. It can be understood that the window size and threshold setting can be adjusted by technicians according to the actual scenario and are not overly limited.
[0057] In practical applications, the monitoring process can also be optimized in combination with business characteristics. For example, during the monthly transaction settlement period, the access mode of the transaction contract file fluctuates significantly. The heat tracing shows that its aggregation degree rises from 0.45 to 0.89, and the migration amplitude value predicted by the gradient boosting decision tree reaches 0.79. It is quickly marked as an object for dynamic adjustment. This adaptive mechanism significantly improves the response ability of backup to highly dynamic data.
[0058] In the embodiment of the present invention, the file backup method of the disaster recovery system improves the backup performance through dynamic block optimization, and is particularly suitable for the highly dynamic data scenario of power enterprises.
[0059] S104. After obtaining the list of enterprise files to be optimized, use the block algorithm based on the data modification frequency to divide the files, and determine the target block granularity by comparing the modification frequency of adjacent data blocks with the metadata update rate to ensure backup efficiency and data consistency.
[0060] In the embodiment of the present invention, extract the list of enterprise files to be optimized from the enterprise file database, and use the data modification record and metadata log to respectively count the modification frequency value and metadata update rate of the file. For example, in the real-time load forecasting data of a certain regional power grid company, the modification frequency during the peak period can reach 180 times per hour, and the metadata update rate is 260 times per hour. These data reflect the dynamic characteristics of file changes. Perform block processing based on these statistical values, and optimize the block boundary through feature analysis to ensure that the target granularity meets the business requirements.
[0061] S1041. Obtain the list of enterprise files to be optimized from the enterprise file database, calculate the modification frequency value according to the data modification records, and statistically calculate the metadata update rate through the metadata log. Subsequently, use an adaptive block splitter to preliminarily divide the files according to the modification frequency value, calculate the distance between adjacent blocks to extract the intra-block data association features. The modification frequency value and the metadata update rate are calculated through time series records. For example, the modification frequency of the distribution automation data of a certain power supply bureau during the peak load period is 140 times per hour, and the metadata update rate is 220 times per hour. The adaptive block splitter divides the data blocks according to the frequency fluctuation. For example, the initial block size is set to 10 megabytes, and the distance between adjacent blocks is calculated through the time stamp difference, reflecting the data continuity. Use a feature extractor to analyze the intra-block data correlation. For example, the correlation value reaches 0.85 during normal operation of the equipment and drops to 0.50 during maintenance, providing a basis for subsequent boundary adjustment.
[0062] In this step, capture the file change pattern through block processing. For example, in the management of power trading contracts, the modification frequency difference on the monthly settlement date is significant. The modification frequency value of adjacent blocks rises from 50 times per hour to 200 times per hour, and the metadata update rate increases from 80 times per hour to 150 times per hour. The distance between adjacent blocks reflects the clarity of the business boundary. The core of the adaptive block splitter is to dynamically adjust the block size according to the frequency, avoiding the impact of too large or too small block division on the backup efficiency.
[0063] S1042. Construct the feature vectors of adjacent data blocks according to the intra-block data association features, extract the modification frequency value and the metadata update rate of adjacent blocks, calculate the feature difference degree, and then judge the position of the block division boundary through time series analysis and generate a block division interval sequence to determine the target block size. The feature vector integrates multi-dimensional information such as modification frequency, update rate, and correlation value. For example, the feature difference degree of the telemetry data of a certain substation rises from 0.25 to 0.90 during equipment alarm, indicating significant changes between blocks. Time series analysis calculates the fluctuation trend of the difference degree through a sliding window, such as 5 minutes, to determine the boundary position. For example, in the data of a certain distribution line, when the difference degree exceeds 0.80, the system adjusts the boundary to the mutation point of the modification frequency. The finally generated block division interval sequence contains balanced data block division information.
[0064] S1043. Retrieve the preset granularity range parameters from the data block database, compare the target block size with the parameters. If it exceeds the range, re-adjust the block division boundary based on the feature difference degree between adjacent blocks until the requirements are met and the final target block granularity is obtained. The preset granularity range is set according to the business type. For example, the inspection records are 6 to 14 megabytes, and the telemetry data is 15 to 30 megabytes. If the initial block size of a certain power dispatching data is 35 megabytes, exceeding the range, iteratively adjust according to the difference degree. For example, move the boundary to the place where the frequency changes gently, and finally converge to 20 megabytes. This iterative optimization ensures that the block granularity matches the data characteristics.
[0065] In practical applications, the chunking strategy is optimized by the above method. For example, in the device monitoring system of a smart substation, the modification frequency of device status data during the alarm period surges to 300 times per hour, and the difference in characteristics between adjacent chunks rises to 0.87. The system automatically adjusts the chunk boundary, optimizing the target granularity from the initial 18 megabytes to 12 megabytes, improving the processing efficiency of high-frequency data. It can be understood that the granularity range and the number of iterations can be adjusted according to the scenario requirements and are not strictly limited.
[0066] In the processing of power trading data, for the chunk display of day-ahead market clearing data, when the initial size is 22 megabytes, the difference in modification frequency between adjacent chunks reaches 2.5 times. After 3 rounds of iteration, the target granularity is adjusted to 16 megabytes, and the characteristic difference is stabilized within 0.60. This method effectively balances the balance of data chunks and access efficiency, laying a foundation for subsequent backup optimization.
[0067] In the embodiment of the present invention, the file backup method of the disaster recovery system improves the incremental backup efficiency by dynamically optimizing the chunking strategy and is applicable to complex data scenarios of power enterprises. The following steps elaborate on the implementation process in detail, and specific parameters and algorithms can be adjusted according to actual requirements.
[0068] S105. Compare the redundancy degree of the data chunk in the incremental backup with a preset threshold. If the redundancy degree is lower than the threshold, retain the current chunking strategy; otherwise, readjust the segmentation threshold and generate an updated chunking scheme to optimize the backup performance.
[0069] In the embodiment of the present invention, the data chunks are marked according to the target chunk granularity, and the backup data size is obtained from the incremental backup database and combined with the backup records to analyze the change cycle of the data chunks. For example, in the distribution automation system of a certain regional power grid, when the target chunk granularity is 15 megabytes, the daily backup data size is relatively stable, and the change cycle is about 40 minutes. The change trend of the backup data size is analyzed by a data comparator, and the redundancy degree of the data chunks is calculated by a redundancy detector. For example, in the telemetry data of a certain substation, the incremental backup frequency during peak hours rises to 3 times per hour, and the redundancy degree reaches 0.55, indicating an increase in data repeatability, which may be related to the frequent update of device status.
[0070] S1051. Mark data blocks according to the target chunk granularity, extract the backup data size from the incremental backup database, calculate the data block change period through backup records, and then use a data comparer to analyze the change in data size and combine a redundancy detector to evaluate the redundancy of data blocks to construct a redundancy feature vector. The data comparer identifies incremental changes by comparing the previous and subsequent backup data. For example, the incremental value of the load data of a power supply bureau during the peak period only accounts for 35% of the total, indicating a high degree of repeatability. The redundancy detector statistically calculates the redundancy based on the change period. For example, the redundancy of a certain transaction data during monthly settlement is 0.60. Construct a feature vector based on information such as redundancy and change period to provide a multi-dimensional basis for threshold judgment.
[0071] S1052. Calculate the backup time interval based on the redundancy feature vector and extract the data block repetition feature. Then obtain the preset redundancy threshold parameter from the backup configuration library, and determine whether to retain the target chunk granularity or adjust the chunking strategy by comparing the data block repetition rate with the threshold. The backup time interval reflects the rhythm of data update. For example, the repetition rate of a certain distribution line data reaches 0.45 at a 15-minute interval and drops to 0.30 at a 30-minute interval, showing the impact of time on redundancy. The preset threshold is usually set according to business requirements. For example, if the repetition rate of a certain equipment status data is only 0.25, the current strategy is retained; if it rises to 0.50, an adjustment is triggered. During adjustment, recalculate the chunking interval according to the change period. For example, shorten the interval from 20 minutes to 10 minutes to reduce redundancy.
[0072] During this process, optimize the backup plan through redundancy analysis. For example, in the power trading system, during the monthly settlement period, the data block change period is shortened to 12 minutes, and the repetition rate rises to 0.53, exceeding the threshold of 0.35. Adjust the chunk size to 10 megabytes and shorten the backup interval to 18 minutes. The repetition rate of the new plan drops to 0.28. This kind of dynamic adjustment effectively reduces storage waste and improves backup real-time performance.
[0073] In practical applications, flexibly adjust the strategy according to different data characteristics. For example, in an intelligent substation, the redundancy during equipment alarm periods rises from 0.27 to 0.58. By shortening the interval to 15 minutes and optimizing the chunking to 12 megabytes, the repetition rate drops to 0.32, meeting the threshold requirements. This method ensures that the backup adapts to the needs of high-dynamic scenarios, and the threshold and interval can be adjusted according to business characteristics without fixed limitations.
[0074] In the distribution automation system, it is found that the redundancy of load data is as high as 0.65 during the peak period and only 0.22 during the low period. Through dynamic adjustment, the interval is shortened to 25 minutes during the peak period and extended to 50 minutes during the low period. The new chunking plan controls the redundancy rate below 0.33, effectively balancing resource utilization and backup efficiency.
[0075] In the embodiment of the present invention, the file backup method of the disaster recovery system realizes efficient backup through adaptive optimization, and is particularly suitable for the dynamic data scenario of power enterprises. The following steps describe the technical implementation process in detail, and the specific parameters and algorithms can be adjusted according to actual requirements.
[0076] S106. Analyze the dynamic characteristics of the change frequency of enterprise files and the transfer of hot data, use the adaptive adjustment algorithm to iteratively optimize the segmentation threshold and the chunking strategy, and evaluate the improvement amplitude of backup efficiency in combination with the volatility of the power market and the change of supply and demand relationship.
[0077] In the embodiment of the present invention, extract the updated chunking scheme data from the enterprise file database, calculate the file change frequency through the file change record, and use the hot data monitor to count the hot data transfer rate. For example, in the real-time electricity price data of a regional power grid, the file change frequency during the peak trading hours can reach 90 times per minute, and the hot data transfer rate rises to 0.85 when the electricity price switches. These data reflect the dynamic characteristics driven by business. Construct a dynamic feature vector based on the change frequency to provide basic support for optimization.
[0078] In this step, capture the data change law through dynamic feature analysis. For example, in the power trading center, the hot data transfer rate during the monthly settlement period rises from 0.40 to 0.88, and the file change frequency bursts during the settlement period, accounting for 70%. The adaptive optimizer calculates the adjustment value of the segmentation threshold according to these characteristics. For example, the threshold is increased by 2.5 times from the reference value to adapt to high-frequency changes, and at the same time, extract the data time series characteristics to reflect the market fluctuation trend. This method ensures that the chunking strategy is synchronized with the business rhythm and improves the backup real-time performance.
[0079] S1061. Construct a dynamic feature vector according to the file change frequency and the hot data transfer rate, use the adaptive optimizer to calculate the adjustment value of the segmentation threshold, generate a market volatility index through the data volatility calculator, and then combine the change amplitude of supply and demand to count the adjustment value of the chunk size to optimize the chunking strategy. The dynamic feature vector includes information such as frequency peak and transfer rate fluctuation. For example, the market volatility index of the telemetry data of a substation during maintenance reaches 0.80, which is much higher than 0.35 during normal operation. The adaptive optimizer adjusts the threshold through weighted calculation. For example, the chunk size is reduced from 14 megabytes to 7 megabytes to reduce redundancy. According to the change amplitude of supply and demand, such as a 30% surge in trading volume, further adjust the chunk size to ensure that the backup efficiency matches the data dynamics.
[0080] S1062. Analyze the backup timing characteristics through a timing correlation calculator and generate a block adjustment sequence. Obtain preset efficiency parameters from the backup configuration library and use a comparison calculator to evaluate the backup efficiency improvement value. If the standard is not met, re-optimize the segmentation threshold based on the timing characteristics until the requirements are satisfied. The timing characteristics reflect the backup rhythm. For example, the backup of a certain distribution automation data is concentrated in the daily load forecasting period, accounting for 60%. The iterative optimizer usually requires 3 to 6 rounds of calculation. For example, the efficiency improvement value of a certain smart meter data during the peak data collection period increases from 1.3 times to 1.6 times, meeting the threshold of 1.5 times. Confirm the effect of the new solution through a verification calculator. For example, the processing delay of the fault recording data is reduced by 40%, proving the effectiveness of the optimization.
[0081] In practical applications, optimize the strategy for power market fluctuations. For example, when the fluctuation of the clearing price in the day-ahead market exceeds 25%, the file change frequency rises to 100 times per minute, the block size is adjusted to 6 megabytes, and the backup efficiency improvement value reaches 1.9 times. This adaptive mechanism significantly improves the system's response ability to high-dynamic scenarios, and can adjust parameters according to business characteristics without fixed limitations.
[0082] In the embodiment of the present invention, the file backup method of the disaster recovery system realizes efficient backup through dynamic adjustment, and is particularly suitable for the high-dynamic data scenarios of power enterprises. The following steps elaborate on the technical implementation process in detail, and the parameters and algorithms can be flexibly adjusted according to actual business requirements.
[0083] S107. After completing the block optimization, obtain the latest enterprise file access pattern change data and compare it with the hot data transfer trend. If the two are consistent, apply the optimized block scheme to the subsequent backup process to generate an efficient backup strategy.
[0084] In the embodiment of the present invention, extract the latest file access records from the enterprise file database, and use a hot data monitor to analyze the access logs and calculate the data heat index. For example, in the real-time electricity price data of a certain regional power grid, the access frequency during the peak trading period reaches 280 times per minute, and the data heat index rises to 0.90, reflecting significant access concentration. Judge the matching degree between the access pattern and the hot data transfer based on these data, and provide a basis for the application of the backup strategy.
[0085] S1071. Calculate the file access pattern value based on the latest file access records, generate a data heat index through the hot data monitor, then analyze the hot data transfer trend using the hot spot tracker and evaluate the matching degree with the access pattern using the trend comparison calculator. The hot spot tracker analyzes the heat change through a sliding window. For example, taking 10 minutes as a unit, the transfer trend of the telemetry data of a substation during maintenance drops from 0.85 to 0.60, indicating a shift in the access pattern. The trend comparison calculator evaluates the matching degree by calculating the correlation coefficient. For example, the matching degree of a certain transaction data reaches 0.87 on the settlement day, far exceeding 0.70 in normal days, showing a high degree of consistency.
[0086] In this step, ensure that the backup strategy adapts to dynamic changes through trend analysis. For example, in a distribution automation system, the access pattern of load forecasting data shows a bimodal characteristic during peak periods, the heat index reaches 0.88, the hot spot transfer trend coincides with it, and the matching degree is 0.91. The preset matching threshold is set at 0.80. If it exceeds this value, the system triggers the application of the optimization plan. This mechanism effectively improves the pertinence and real-time nature of backups.
[0087] S1072. Retrieve the optimized block scheme from the backup configuration library and construct a backup execution sequence based on the matching result. Then, generate a backup schedule through the data processor and perform time series allocation using the backup scheduler to determine the backup execution schedule. The backup execution sequence includes parameters such as block size and backup interval. For example, the data of a smart substation will be adjusted to a block size of 12 megabytes and the interval will be shortened to 20 minutes. The data processor generates a schedule based on the processing parameters. For example, the backup window for high-frequency data is set at 25 minutes, and for low-frequency data it is 3 hours. The backup scheduler optimizes the time allocation to ensure resource utilization efficiency. For example, during the peak electricity consumption period, hot data is preferentially processed.
[0088] In practical applications, the strategy deployment is achieved through the above steps. For example, in a power trading center, the matching degree during the monthly settlement period reaches 0.89. The optimization plan adjusts the block size from 20 megabytes to 15 megabytes and the backup interval from 1 hour to 30 minutes. The verification results show that the processing delay is reduced by 50% and the storage efficiency is increased by 35%. This dynamic adjustment significantly enhances the system's adaptability to business fluctuations.
[0089] The execution schedule is also optimized for different scenarios. For example, in the distribution line data of a certain power supply station, the access frequency surges during peak periods, the matching degree reaches 0.93, the backup interval is adjusted to 15 minutes, and the verification calculator confirms that the new strategy reduces the data delay to 40% of the original, meeting the pre-set backup requirements. The threshold and parameters can be adjusted according to business characteristics without strict limitations.
[0090] In the embodiment of the present invention, the file backup method of the disaster recovery system optimizes the system performance through continuous monitoring and dynamic adjustment, and is particularly suitable for the complex data environment of power enterprises. The following steps detail the technical implementation details, and the parameters and algorithms can be flexibly adjusted according to the actual scenario.
[0091] S108. After applying the final backup policy, continuously monitor the fluctuation characteristics of the enterprise file change frequency and the metadata update rate, use dynamic analysis to capture the update and modification frequency of power transaction-related data, and evaluate the adaptive adjustment ability of the disaster recovery system.
[0092] In the embodiment of the present invention, set the data sampling period according to the final backup policy, obtain the enterprise file change frequency from the enterprise file database, and sample the metadata update rate fluctuation value through the metadata record statistic. For example, in the real-time electricity price data of a certain regional power grid, the sampling period is set to 10 minutes. The file change frequency during the peak period reaches 95 times per minute, and the metadata update rate fluctuation value rises to 0.80, showing strong business-driven characteristics. Analyze the dynamic response ability of the system through these data to provide support for the adjustment plan.
[0093] S1081. Sample the enterprise file change frequency through the metadata record statistic and calculate the metadata update rate fluctuation value. Then, use the feature extractor to analyze the fluctuation characteristics and the fluctuation detector to extract the update and modification frequency of the power transaction data to construct a dynamic feature vector. The feature extractor identifies the fluctuation law through time series analysis. For example, the update and modification frequency of the telemetry data of a certain substation during the maintenance period increases from 140 times per hour to 500 times per hour, and the fluctuation value reaches 0.85. Construct a dynamic feature vector according to the frequency and the fluctuation value, such as fusing the peak frequency and the change amplitude, to reflect the data update characteristics and lay a foundation for subsequent prediction.
[0094] In this step, capture the business changes through dynamic features. For example, in the power trading center, the update frequency of the transaction data during the monthly settlement period rises to 120 times per minute, and the fluctuation detector shows that its fluctuation value increases from 0.45 to 0.90. The dynamic feature vector clearly reflects the period of intensive data update. The adaptive optimizer uses this information to adjust the policy to ensure that the system adapts to the high-frequency scenario and improves the backup efficiency and integrity.
[0095] S1082. Perform time series prediction based on dynamic feature vectors and generate disaster recovery adjustment parameters. Subsequently, calculate the system adjustment response value through an adaptive optimizer and compare it with a preset adjustment threshold. If it exceeds the threshold, regenerate the adjustment plan and verify its rationality to optimize the adaptive ability. The time series predictor uses regression analysis to predict the change trend. For example, the adjustment parameter of a certain distribution automation data reaches 0.88 during the peak load period, far exceeding 0.30 during the low valley period. The preset threshold is set at 0.70. If the response value reaches 0.92, the sampling period is readjusted, for example, shortened from 15 minutes to 5 minutes. The new plan is confirmed by a verification calculator. For example, the response delay is reduced by 38%, proving the effectiveness of the adjustment.
[0096] In practical applications, optimize the parameters for different business scenarios. For example, in an intelligent substation, during the equipment alarm period, the update frequency surges to 600 times per hour. The adjustment plan shortens the backup interval to 12 minutes. The verification results show that the data capture rate rises to 99% and the system response time is halved. This dynamic mechanism significantly enhances the adaptability of the disaster recovery system to power market fluctuations and can adjust the threshold and period according to requirements without fixed limitations.
[0097] Such as Figure 2 , the present invention provides a file backup system for a disaster recovery system, mainly including:
[0098] A metadata monitoring module, used to obtain the metadata update rate and modification frequency of the files in the disaster recovery system of the power enterprise, calculate the change frequency distribution, determine the initial segmentation threshold range, analyze the time series fluctuations, and adjust the segmentation threshold range if it exceeds the preset threshold;
[0099] A segmentation threshold adjustment module, used to extract the real-time data of the metadata update rate and modification frequency from the adjusted segmentation threshold range, calculate the Pearson coefficient between the two, and judge the optimization priority of the chunking strategy;
[0100] A chunking strategy optimization module, used to monitor the changes in the enterprise file access pattern and the offset of the modification granularity, determine the signs of hot data transfer, and if the offset is greater than the preset threshold, mark the enterprise file as a dynamically adjustable object to obtain a list of enterprise files to be optimized;
[0101] A dynamic adjustment marking module, used to divide data blocks according to the list of enterprise files to be optimized by using a chunking algorithm based on the data modification frequency, compare the modification frequency and metadata update rate of adjacent data blocks, and determine the target chunking granularity;
[0102] A chunking granularity determination module, used to calculate the redundancy degree of the data blocks in the incremental backup through the target chunking granularity. If it is lower than the preset redundancy threshold, retain the current chunking strategy. If it is not lower, readjust the segmentation threshold to obtain an updated chunking plan;
[0103] A redundancy evaluation module, which is used to obtain the change frequency of enterprise files and the dynamic characteristics of hot data transfer in the updated chunking scheme, dynamically iteratively optimize the segmentation threshold and chunking strategy through an adaptive adjustment algorithm, and judge the improvement amplitude of backup efficiency in combination with the volatility of the power market and the change of supply and demand relationship;
[0104] An iterative optimization module, which is used to obtain the latest data on changes in the enterprise file access pattern. If it is consistent with the hot data transfer trend, the optimized chunking scheme is applied to the subsequent backup process to obtain the final backup strategy;
[0105] A strategy application module, which is used to continuously monitor the change frequency of enterprise files and the fluctuation of metadata update rate through the final backup strategy, capture the update and modification frequency of power trading related data by using dynamic characteristics, and determine the adaptive adjustment ability of the enterprise disaster recovery system.
[0106] Although the present invention has been described in detail with general descriptions and specific embodiments above, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A file backup method for a disaster recovery system, characterized in that The method includes: Obtain the metadata update rate and modification frequency of the power enterprise disaster recovery system files, calculate the change frequency distribution, determine the initial segmentation threshold range, analyze the time series fluctuation, and if it exceeds the preset threshold, adjust the segmentation threshold range; Extract the real-time data of the metadata update rate and modification frequency from the adjusted segmentation threshold range, calculate the Pearson coefficient between the two, and judge the priority of optimizing the block strategy; Monitor the changes in the enterprise file access pattern and the offset of the modification granularity, determine the signs of hot data transfer, and if the offset is greater than the preset threshold, mark the enterprise files as objects for dynamic adjustment to obtain a list of enterprise files to be optimized; According to the list of enterprise files to be optimized, use a block algorithm based on the data modification frequency to divide data blocks, compare the modification frequency and metadata update rate of adjacent data blocks, and determine the target block granularity; Through the target block granularity, calculate the redundancy degree of the data block in the incremental backup. If it is lower than the preset redundancy threshold, retain the current block strategy. If it is not lower, readjust the segmentation threshold to obtain an updated block scheme; Obtain the change frequency of the enterprise files and the dynamic characteristics of hot data transfer in the updated block scheme, dynamically iterate and optimize the segmentation threshold and block strategy through an adaptive adjustment algorithm, and judge the improvement amplitude of the backup efficiency in combination with the volatility of the power market and the change of supply and demand relationship; Obtain the latest data on changes in the enterprise file access pattern. If it is consistent with the trend of hot data transfer, apply the optimized block scheme to the subsequent backup process to obtain the final backup strategy; Through the final backup strategy, continuously monitor the fluctuations of the enterprise file change frequency and metadata update rate, capture the update and modification frequency of power transaction-related data using dynamic characteristics, and determine the adaptive adjustment ability of the enterprise disaster recovery system.
2. The method according to claim 1, wherein The obtaining the metadata update rate and modification frequency of the power enterprise disaster recovery system files, calculating the change frequency distribution, determining the initial segmentation threshold range, analyzing the time series fluctuation, and if it exceeds the preset threshold, adjusting the segmentation threshold range includes: Obtain the file metadata update records from the disaster recovery database, calculate the metadata update frequency value according to the data modification timestamps in the update records to obtain the metadata modification time series data; Use a time series fluctuation detector to perform a sliding window analysis on the metadata modification time series data, and obtain a metadata modification frequency distribution map by calculating the time series data. The frequency distribution map includes the change fluctuation values of the metadata in different time periods; Construct a time series feature vector for the metadata change fluctuation value, and extract the frequency feature, time series feature and fluctuation feature in the time series feature vector through a clustering operator to obtain the initial segmentation threshold of the metadata change.
3. The method according to claim 1, wherein The extracting the real-time data of the metadata update rate and modification frequency from the adjusted segmentation threshold range, calculating the Pearson coefficient between the two, and judging the priority of optimizing the block strategy includes: Obtain the time series records of the metadata update rate and modification frequency from the segmentation threshold database, and count the number of metadata blocks through a data block counter to obtain the basic statistical value of the real-time data blocks; Using a time series synchronization calculator, intercept the time window of the metadata update rate and modification frequency based on the real-time data block basic statistical value to obtain the data block time series characteristics; Using a correlation coefficient calculator, extract the time series correlation characteristics between the metadata update rate and modification frequency according to the data block time series characteristics to obtain the Pearson coefficient; If the Pearson coefficient exceeds the preset priority threshold, recalculate the block strategy optimization weight value based on the data block time series characteristics to obtain the block priority adjustment sequence.
4. The method according to claim 1, wherein Monitoring the change of the enterprise file access mode and the modification granularity offset, determining the sign of hot data transfer. If the offset is greater than the preset threshold, mark the enterprise file as a dynamically adjustable object to obtain a list of enterprise files to be optimized, including: Obtain the real-time electricity price data, electricity demand forecast data and trading contract data from the enterprise file database, and calculate the file access heat value and data activity index according to the access record log; Use a heat tracker to perform a sliding window analysis on the file access heat value, and obtain the file hot spot aggregation degree through the data activity index; Construct a hot data transfer feature vector according to the file hot spot aggregation degree, use a gradient boosting decision tree to perform time series prediction on the transfer feature sequence, and judge the hot data transfer trend through the file change rate; If the hot spot migration amplitude value exceeds the preset offset threshold parameter, mark the corresponding file as a dynamically adjustable object and add the dynamically adjustable object to the list of enterprise files to be optimized.
5. The method according to claim 1, wherein According to the list of enterprise files to be optimized, use a block algorithm based on data modification frequency to divide data blocks, compare the modification frequency and metadata update rate of adjacent data blocks, and determine the target block granularity, including: Obtain the list of enterprise files to be optimized from the enterprise file database, count the modification frequency value according to the data modification record, and obtain the metadata update rate through the metadata log; Use an adaptive block divider to perform block processing on the enterprise files to be optimized according to the modification frequency value, calculate the distance between adjacent blocks for the block processing result, and obtain the intra-block data association value through a block feature extractor; Construct an adjacent data block feature vector according to the intra-block data association value, extract the modification frequency value and metadata update rate for adjacent data blocks, and judge the block boundary position through the distance between adjacent blocks; If the target block size exceeds the preset granularity range parameter, recalculate the block boundary position until the target block size meets the preset granularity range parameter to obtain the target block granularity; It also includes: calculating the data modification frequency of the enterprise files corresponding to the list of enterprise files to be optimized, dividing the enterprise files into multiple data blocks by using the data modification frequency, comparing the modification frequencies between adjacent data blocks, and at the same time comparing the metadata update rates between adjacent blocks. According to the comparison results of the modification frequency and metadata update rate, determine the optimal target block granularity.
6. The method according to claim 1, wherein Calculate the redundancy degree of the data block in the incremental backup through the target block granularity. If it is lower than the preset redundancy threshold, retain the current block strategy. If it is not lower, readjust the segmentation threshold to obtain the updated block scheme, including: Mark the data blocks according to the target chunk granularity, obtain the backup data size from the incremental backup database, and get the data block change period through the backup records; Use a data comparer to analyze the backup data size, and obtain the data block redundancy through a redundancy detector for the data block change period; Construct a redundancy feature vector according to the data block redundancy, calculate the backup time interval for the data block change period, and obtain the data block repetition feature; If the adjacent block repetition rate calculated from the data block repetition feature is lower than the preset redundancy threshold parameter, retain the target chunk granularity; If the adjacent block repetition rate is not lower than the preset redundancy threshold parameter, recalculate the chunk interval value based on the data block change period and generate an updated chunking scheme.
7. The method according to claim 1, characterized in that The enterprise file change frequency and the dynamic characteristics of hot data transfer for obtaining the updated chunking scheme are dynamically iteratively optimized for the segmentation threshold and chunking strategy through an adaptive adjustment algorithm, and the improvement range of backup efficiency is judged by combining the volatility of the power market and the change of supply and demand relationship, including: Calculate the file change frequency according to the file change records, count the hot data transfer rate through a hot data monitor, and construct a dynamic feature vector based on the change frequency; Use an adaptive optimizer to calculate the segmentation threshold adjustment value according to the dynamic feature vector, extract the data time series characteristics for the hot data transfer rate, and generate a market volatility index through a data volatility calculator; Calculate the supply and demand change range according to the market volatility index, count the chunk size adjustment value for the data time series characteristics, and obtain the backup time series characteristics through a time series correlation calculator; Use an iterative optimizer to perform multiple rounds of calculations on the segmentation threshold adjustment value, obtain a chunking adjustment sequence according to the supply and demand change range, and judge whether the backup efficiency improvement value meets the preset efficiency parameter through a comparison calculator; If the backup efficiency improvement value does not meet the preset efficiency parameter, recalculate the segmentation threshold adjustment value based on the backup time series characteristics; Generate a new chunking scheme according to the segmentation threshold adjustment value, and verify the efficiency of the new chunking scheme through a verification calculator.
8. The method according to claim 1, characterized in that, The latest enterprise file access mode change data is obtained. If it is consistent with the hot data transfer trend, apply the optimized chunking scheme to the subsequent backup process to obtain the final backup strategy, including: Obtain the latest file access records from the enterprise file database, analyze the access records through a hot data monitor, and obtain the data heat index; Calculate the hot data transfer trend using a hot trackers according to the data heat index, and analyze the hot data transfer trend through a trend comparison calculator to obtain the trend matching degree; If the trend matching degree exceeds the preset matching threshold, retrieve the optimized chunking scheme from the backup configuration library and generate a backup schedule through a data processor; Perform time series allocation using a backup scheduler according to the backup schedule to obtain the final backup strategy.
9. The method according to claim 1, characterized in that Through the final backup strategy, continuously monitor the fluctuations of the enterprise file change frequency and metadata update rate, and use dynamic features to capture the update and modification frequency of power trading related data to determine the adaptive adjustment ability of the enterprise disaster recovery system, including: Obtain the enterprise file change frequency from the enterprise file database, sample the enterprise file change frequency through a metadata recorder to obtain the metadata update rate fluctuation value; Use a feature extractor to perform feature analysis on the metadata update rate fluctuation value, calculate the time series fluctuation characteristics according to the feature analysis results, and obtain the power transaction data update and modification frequency through a fluctuation detector; Construct a dynamic feature vector based on the power transaction data update and modification frequency, perform time series prediction analysis on the dynamic feature vector, and obtain the disaster recovery adjustment parameter through a time series predictor; Use an adaptive optimizer to calculate the response value of the disaster recovery adjustment parameter. If the response value exceeds the preset adjustment threshold, regenerate the disaster recovery system adjustment plan based on the dynamic feature vector, and verify the rationality of the adjustment plan through a verification calculator.
10. A file backup system for a disaster recovery system, characterized in that, The system includes: A metadata monitoring module, which is used to obtain the metadata update rate and modification frequency of the disaster recovery system files of power enterprises, calculate the change frequency distribution, determine the initial segmentation threshold range, analyze the time series fluctuation, and adjust the segmentation threshold range if it exceeds the preset threshold; A segmentation threshold adjustment module, which is used to extract the real-time data of the metadata update rate and modification frequency from the adjusted segmentation threshold range, calculate the Pearson coefficient between the two, and judge the priority of the block strategy optimization; A block strategy optimization module, which is used to monitor the changes in the enterprise file access mode and the modification granularity offset, determine the signs of hot data transfer. If the offset is greater than the preset threshold, mark the enterprise file as a dynamically adjusted object to obtain a list of enterprise files to be optimized; A dynamic adjustment marking module, which is used to divide data blocks using a block algorithm based on the data modification frequency according to the list of enterprise files to be optimized, compare the modification frequency and metadata update rate of adjacent data blocks, and determine the target block granularity; A block granularity determination module, which is used to calculate the redundancy degree of data blocks in incremental backup through the target block granularity. If it is lower than the preset redundancy threshold, retain the current block strategy. If it is not lower, readjust the segmentation threshold to obtain an updated block plan; A redundancy evaluation module, which is used to obtain the enterprise file change frequency and the dynamic characteristics of hot data transfer of the updated block plan, dynamically iteratively optimize the segmentation threshold and block strategy through an adaptive adjustment algorithm, and judge the improvement amplitude of the backup efficiency in combination with the volatility of the power market and the changes in the supply and demand relationship; An iterative optimization module, which is used to obtain the latest data on the changes in the enterprise file access mode. If it is consistent with the hot data transfer trend, apply the optimized block plan to the subsequent backup process to obtain the final backup strategy; A strategy application module, which is used to continuously monitor the enterprise file change frequency and metadata update rate fluctuation through the final backup strategy, capture the power transaction-related data update and modification frequency using dynamic features, and determine the adaptive adjustment force of the enterprise disaster recovery system.
Citation Information
Patent Citations
Data deduplication method and related system
CN117331487A
Fault-tolerant uploading of data to a distributed storage system
US20220121532A1
Cited By
Data backup method and system for set top box
CN120762973A
File backup method and device, computer equipment and computer readable storage medium
CN121579270A