Intelligent data archiving management system for shutdown system

Through the intelligent data archiving management system for shutdown systems, using technologies such as hybrid firework optimization and dynamic entropy weight clustering, the problems of low data archiving efficiency and low classification accuracy in sudden shutdown events are solved, and efficient identification and priority archiving of key data are achieved.

CN119989033AActive Publication Date: 2025-05-13HANGZHOU YIKANGXIN TECH CO LTD

Patent Information

Application Number
CN202510474997.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing data archiving methods for shutdown systems have problems such as low archiving efficiency, lack of dynamic adaptability, insufficient storage resource optimization, and low data classification accuracy in sudden shutdown events.

Method used

A data intelligent archiving management system for shutdown systems is proposed. Through modules such as data collection and preprocessing, data feature extraction and evaluation, data screening based on hybrid firework optimization, dynamic entropy weight clustering analysis, archive path optimization and storage resource allocation, etc., a rescue archive data collection and optimization are formed and the storage path is optimized.

Benefits of technology

It improves the identification accuracy and archive response speed of key data, ensures that key data can be preferred to enter the archived data collection, and improves the accuracy and response speed of data rescue archiving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989033A_ABST
    Figure CN119989033A_ABST
Patent Text Reader

Abstract

The invention discloses a shutdown system-oriented data intelligent archiving management system, which comprises a data acquisition and preprocessing module, a data storage module, a data management module and a data management module, and is characterized in that the data acquisition and preprocessing module is used for acquiring an original operation data set in a shutdown system; the data feature extraction and evaluation module is used for performing multi-dimensional feature extraction on the preprocessed shutdown system operation data set; the data screening module based on hybrid firework optimization is used for performing global search and parameter optimization on the data feature data set; the dynamic entropy weight clustering analysis module is used for performing dynamic entropy weight clustering analysis on the optimized rescue archiving data set; the archiving path optimization and storage resource allocation module is used for closing a system storage hierarchy model and a real-time system state in combination with an archiving data classification result; and the archiving task execution module is used for executing rescue archiving operation according to the archiving task plan. The invention ensures that the key data can be preferentially selected to enter the archived data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of archive management systems, and in particular to a data intelligent archive management system for shutdown systems. Background Art

[0002] With the development of information technology, data security and storage management of shutdown systems have gradually become an important research direction in the field of informatization. In enterprise data centers, industrial control systems and cloud computing environments, sudden system shutdowns may be caused by equipment failures, network anomalies, software vulnerabilities, malicious attacks and many other factors. In this process, key data in the system often faces high risks. If they are not archived and stored in time, it may lead to data loss, business interruption, and even affect the normal operation of the enterprise.

[0003] At present, the mainstream methods for archiving shutdown system data mainly include preset rule archiving, static priority archiving and regular backup. The existing methods have their own advantages and disadvantages, but the following technical defects are exposed in the emergency shutdown scenario:

[0004] Traditional preset rule archiving relies on pre-established strategies, such as determining the importance of data based on file type, directory structure, or historical access frequency. However, in the event of an unexpected shutdown, the importance of data may change dynamically, and the preset rules cannot respond in real time, resulting in the inability to archive critical data in a timely manner, increasing the risk of data loss.

[0005] The static priority archiving method divides the storage priority of data according to fixed rules, such as setting certain business databases as high priority and ordinary documents as low priority. However, the data status in the shutdown system may change rapidly in a short period of time. For example, some real-time transaction data or log files may become critical at a critical moment. Static priority archiving lacks dynamic adjustment capabilities and is prone to missing high-priority data.

[0006] Existing periodic backup methods usually use preset storage paths for data storage without considering the real-time status of system resources, resulting in high-priority data being stored in low-speed storage media, affecting the rapid recovery of data. At the same time, the allocation of storage resources is not optimized, which may lead to the failure of archiving tasks.

[0007] In summary, the existing shutdown system data archiving methods have great limitations when dealing with sudden shutdown events, mainly reflected in low archiving efficiency, lack of dynamic adaptability, insufficient storage resource optimization, and low data classification accuracy. There is an urgent need for a data intelligent archiving management system for shutdown systems to improve the accuracy and response speed of data rescue archiving. Summary of the invention

[0008] One object of the present invention is to provide a data intelligent archiving management system for a shutdown system, which ensures that key data can be preferentially selected to enter the archived data set.

[0009] According to an embodiment of the present invention, a data intelligent archiving management system for a shutdown system includes the following modules:

[0010] The data collection and preprocessing module is used to collect the original operation data set in the shutdown system and perform preprocessing to form the shutdown system operation data set;

[0011] The data feature extraction and evaluation module is used to extract multi-dimensional features of the shutdown system operation data set, calculate the time sensitivity, access frequency and storage requirements of the data, and calculate the importance weight of the data based on the dynamic entropy weight to form a data feature data set;

[0012] The data screening module based on hybrid fireworks optimization is used to perform global search and parameter optimization on data feature data sets. The optimal data subset is selected based on the hybrid fireworks optimization algorithm combined with shutdown system status, data priority and storage resource constraints to form an optimized rescue archive data set;

[0013] The dynamic entropy weight cluster analysis module is used to perform dynamic entropy weight cluster analysis on the optimized rescue archived data set to generate archived data classification results;

[0014] The archiving path optimization and storage resource allocation module is used to optimize the path and allocate storage resources for the archiving data in the archiving data classification results and generate an archiving task plan by combining the archiving data classification results, the shutdown system storage hierarchy model and the real-time system status;

[0015] The archiving task execution module is used to execute rescue archiving operations according to the archiving task plan, so that the shutdown system operation data set is stored according to the optimal path and the archiving status is recorded.

[0016] A data intelligent archiving management method for a shutdown system is applied to a data intelligent archiving management system for a shutdown system, comprising:

[0017] S1. Collecting the original operating data set in the shutdown system and performing preprocessing operations on the original operating data set to form a preprocessed shutdown system operating data set;

[0018] S2. Perform multi-dimensional feature extraction on the pre-processed shutdown system operation data set, calculate the time sensitivity, access frequency and storage requirements of each data sample in the shutdown system operation data set, and generate a data feature data set;

[0019] S3. Perform global search and parameter optimization on the data feature dataset based on the hybrid fireworks optimization algorithm to form a rescue archive data set that meets the global optimal solution;

[0020] S4. clustering the rescue archived data set that meets the global optimal solution using a dynamic entropy weight clustering method, classifying the data according to the adaptively adjusted weights of each data feature, and generating archived data classification results;

[0021] S5. Optimize the path and allocate storage resources for the archived data in the archived data classification results according to the archived data classification results combined with the shutdown system storage hierarchy model and the real-time system status, and generate an archive task plan;

[0022] S6. Perform rescue archiving operations on the optimized rescue archiving data set in the actual environment of the shutdown system according to the archiving task plan, monitor the archiving process in real time and record the archiving status to form an archiving execution record.

[0023] Optionally, S1 includes the following steps:

[0024] S11. Collecting the original operation data set in the shutdown system, the original operation data set includes different types of data samples, each data sample contains multiple data attributes, constituting the original operation data set : ;

[0025] in, Indicates data samples, Indicates The value of a data sample on the mth data attribute, is the total number of data samples, is the total number of data attributes;

[0026] S12. Run the original dataset Perform format conversion to convert data samples in different formats into standardized representations to form a formatted running data set;

[0027] S13. Deduplication of the formatted operation data set is performed to remove data samples with exactly the same content to form a deduplication operation data set;

[0028] S14. Perform anomaly detection on the deduplicated running data set, identify and remove data outliers, and form an anomaly-detected running data set;

[0029] S15. Perform data integrity check on the operation data set after anomaly detection, remove incomplete data, and form a pre-processed shutdown system operation data set : ;

[0030] in, is the data integrity check function, is the running data set after anomaly detection. hour, Satisfy the completeness requirements, otherwise it will be rejected.

[0031] Optionally, S2 includes the following steps:

[0032] S21. Run the shutdown system data set after preprocessing Perform time sensitivity calculations, analyze the time correlation of data samples, and define time sensitivity parameters ,Time sensitivity is used to measure the real-time requirements of shutting down system operation data. The closer the time sensitivity value is to 1, the more urgent the data needs to be archived: ;

[0033] in, Indicates the last update time of the data sample, Indicates the current system time. Parameters for adjusting the magnitude of time sensitivity changes;

[0034] S22. Run the shutdown system data set after preprocessing Calculate the access frequency, count the number of accesses of each data sample, and define the access frequency parameters : ;

[0035] in, Indicates The number of accesses to a data sample within a specified time window, Indicates The number of accesses to a data sample within a specified time window, The access frequency adjustment factor is used to enhance the recognition of high-access data;

[0036] S23. Run the shutdown system data set after preprocessing Perform storage demand calculations, analyze data storage capacity requirements, and define storage demand parameters : ;

[0037] in, Indicates The storage size of data samples, is the storage demand adjustment index, It represents the weighted storage size of the jth sample in the storage requirement calculation, which is used to adjust the influence of data samples on the storage requirement calculation. >1 to enhance the impact of large data samples, <1 to balance the influence of different data samples. It represents the weighted storage size of the i-th sample in the storage requirement calculation;

[0038] S24. Combining time sensitivity parameters , access frequency parameters and storage requirement parameters , calculate the comprehensive weight , which is used to measure the overall importance of each data sample. The larger the comprehensive weight, the data should be archived first: ;

[0039] in, , , is the weight coefficient;

[0040] S25. Based on the calculated comprehensive weight Sort the data samples in the shutdown system operation data set to generate a data feature data set : ;

[0041] in, The archiving threshold is used to filter out data samples that meet the archiving conditions.

[0042] Optionally, S3 includes the following steps:

[0043] S31. Construct a fireworks optimization search space based on a multi-level explosion strategy, targeting a data feature dataset Construct a multi-level explosion optimization model and initialize the fireworks population of the multi-level explosion optimization model , each fireworks individual in the fireworks individual population Represents a candidate subset of salvage archive data: ;

[0044] in, Indicates A subset of candidate salvage archive data for fireworks individuals, is the number of initial fireworks individual population;

[0045] In the multi-level explosion optimization model, a multi-level explosion radius adjustment mechanism is introduced to make each firework individual according to the comprehensive weight. Adjust explosion radius : ;

[0046] in, is the global maximum search radius, Adjust the parameters for the explosion range, Represents data samples The comprehensive weight of Represents data samples The comprehensive weight of

[0047] S32. Adopting an adaptive perturbation mechanism to optimize the generation of candidate rescue archived data subsets, adaptively perturb data of different priorities during the search process, and defining the optimized candidate rescue archived data subset adjustment rules: ;

[0048] in, To indicate the The candidate rescue archive data subset after the optimization of fireworks individuals, and are the addition and removal thresholds, For data samples The probability of being added to the candidate rescue archive data subset: ; For data samples The probability of being removed from the candidate rescue archive data subset: ;

[0049] in, For data samples The probability of being added to the candidate rescue archive data subset;

[0050] S33. In the optimization process, a hierarchical co-evolution strategy is introduced to divide the fireworks individuals into a key data priority layer, a balanced archiving layer, and a storage restricted layer. The corresponding search strategies are adopted respectively. The key data priority layer adopts a fast convergence search, the balanced archiving layer adopts a multi-objective balance strategy, and the storage restricted layer adopts storage optimization adjustment. The data optimization objectives of the key data priority layer, the balanced archiving layer, and the storage restricted layer are defined as follows: ; ; ;

[0051] in, is the objective function of the key data priority layer, is the objective function of the balanced archive layer, is the storage-restricted layer objective function, Indicates The candidate rescue archive data subset after the optimization of fireworks individuals, To store weight adjustment parameters, is the weighted mean of the data samples, For available storage resources;

[0052] S34. Combine the multi-objective optimization strategy to generate the final optimized data set, and perform final screening to select the rescue archive data set that meets the global optimal solution : ;

[0053] in, , , The weight parameter for the archiving optimization objective, The final population set is made to make the rescue archive data set reach the global optimum among key data priority, archive balance and storage utilization.

[0054] Optionally, S4 includes the following steps:

[0055] S41. Construct a clustering optimization model based on two-layer dynamic entropy weights for the rescue archived data set that meets the global optimal solution Construct an adaptive entropy weight calculation framework to initialize the data clustering center set of the clustering optimization model , each cluster center in the data cluster center set Represents an archive data category: ;

[0056] in, Indicates Cluster centers, is the total number of initial archived data categories;

[0057] The global entropy weight is calculated by using a two-layer dynamic entropy weight calculation method. and local entropy weight , comprehensively evaluate the importance of data features, and the global entropy weight is calculated as follows: ;

[0058] in, Indicates The normalized probability value of the data sample on the corresponding feature;

[0059] The local entropy weight calculation uses the contribution of data samples to the cluster center, which is defined as follows: ;

[0060] in, Represents the cluster center In Features The value on Represents the cluster center In Features The value on For all cluster centers in the feature The mean on ;

[0061] S42. Use variable weight fuzzy clustering method to calculate data sample membership and define adaptive weighted fuzzy distance , calculate the data sample Cluster Center Membership : ;

[0062] Among them, the adaptive weighted fuzzy distance is defined as follows: ;

[0063] in, is the variable weight index, is the total number of features of the data sample;

[0064] S43. Dynamic update strategy based on cluster drift detection, using cluster center moving trend vector to optimize cluster update process, introducing drift metric Determine whether the cluster center needs to be adjusted: ;

[0065] in, Indicates that it belongs to the cluster center The number of data samples is hour, is the threshold, the cluster center needs to be adjusted, and the adjustment strategy is as follows: ;

[0066] in, To dynamically adjust the step size, is the new cluster center location, For data samples Cluster Center The degree of membership;

[0067] S44. Use the fuzzy entropy convergence criterion to optimize the clustering termination condition and calculate the fuzzy entropy of the clustering system Measuring the uncertainty of a clustering: ;

[0068] Set the dynamic entropy convergence threshold ,when And the cluster center changes satisfy When , the clustering is terminated, otherwise, S42-S43 is continued to update the cluster center;

[0069] S45. Based on the optimized cluster center and membership For data samples Assign archiving categories and generate archiving data classification results: ;

[0070] in, Indicates the final classification result of archived data. For data samples The archive category to which it belongs, making the rescue archive data collection are effectively classified, To obtain The value of the variable that reaches its maximum value.

[0071] Optionally, S5 includes the following steps:

[0072] S51. Build a path optimization framework based on the shutdown system storage hierarchy model to classify the archived data Optimize the archive path and initialize the storage hierarchy of the shutdown system , where each storage tier Indicates an archive path to a different storage medium: ;

[0073] in, Indicates storage tiers, is the total number of storage tiers, Represents the collection of available storage resources in the system;

[0074] Introducing data-based urgency Storage tier selection function : ;

[0075] in, For data samples The urgency of For storage tier The read and write delays For storage tier The unit storage cost is set so that data with a higher urgency level than the preset value is stored in high-speed storage media first, and data with a lower urgency level than the preset value is stored in low-cost storage media;

[0076] S52. Calculate the optimal storage allocation plan for the data archiving path, and perform archiving of data samples based on the storage hierarchy model. Calculate its optimal storage location : ;

[0077] in, is the comprehensive weight, is the storage size of the data sample, For storage tier of available storage space, For storage tier The historical visit frequency, To optimize the weight coefficient, the storage allocation scheme takes into account data priority, storage space utilization and data access requirements;

[0078] S53. Generate the final archiving task plan based on the optimal storage location : ;

[0079] in, For data samples The archive timestamp of The archiving trigger threshold is set so that the rescue archive data set is stored in the optimal path.

[0080] The beneficial effects of the present invention are:

[0081] (1) The present invention adopts an improved hybrid fireworks optimization algorithm and a multi-level explosion radius adaptive adjustment mechanism, so that data screening no longer relies on fixed rules, but adaptively adjusts the search range according to the time sensitivity, access frequency, and storage requirement key features of the data, thereby improving the recognition accuracy of key data. The multi-level explosion radius adjustment strategy enables high-priority data to obtain a larger search weight in the global search process, ensuring that key data can be preferentially selected into the archived data set.

[0082] (2) The present invention adopts a two-layer dynamic entropy weight clustering method in the data classification process to calculate the global entropy weight and local entropy weight of data features respectively, so that data classification can dynamically adjust the influence weights of different features according to the system state. The two-layer entropy weight calculation framework enhances the stability and accuracy of data classification through an adaptive weighted fuzzy distance calculation method.

[0083] (3) In terms of data storage path optimization, the present invention proposes a storage level selection function based on data urgency, and combines it with reinforcement learning to optimize the storage path allocation strategy, so that data archiving can be dynamically allocated according to the real-time system status. The storage path selection is continuously adjusted through the reinforcement learning method, so that high-priority data can be allocated to high-speed storage media with lower storage latency, and the data migration plan is dynamically adjusted to ensure efficient use of storage space. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0085] Figure 1 This is a flow chart of a data intelligent archiving management system for shutdown systems proposed by the present invention. DETAILED DESCRIPTION

[0086] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0087] refer to Figure 1 , a data intelligent archiving management system for shutdown systems, including the following modules:

[0088] The data collection and preprocessing module is used to collect the original operation data set in the shutdown system, perform format conversion, data deduplication, anomaly detection and data integrity verification on the original operation data set, and form a preprocessed shutdown system operation data set;

[0089] The data feature extraction and evaluation module is used to extract multi-dimensional features from the pre-processed shutdown system operation data set, calculate the time sensitivity, access frequency and storage requirements of the data, and calculate the importance weight of the data based on the dynamic entropy weight to form a data feature data set;

[0090] The data screening module based on hybrid fireworks optimization is used to perform global search and parameter optimization on data feature data sets. The optimal data subset is selected based on the hybrid fireworks optimization algorithm combined with shutdown system status, data priority and storage resource constraints to form an optimized rescue archive data set;

[0091] The dynamic entropy weight clustering analysis module is used to perform dynamic entropy weight clustering analysis on the optimized rescue archived data set, adaptively adjust the data classification weight according to the global entropy weight and local entropy weight of the data characteristics, perform fuzzy clustering, and calculate the archived data category based on the membership degree to generate the archived data classification result;

[0092] The archiving path optimization and storage resource allocation module is used to optimize the path and allocate storage resources for the archived data in the archived data classification results, calculate the optimal storage level, dynamically adjust the data migration strategy, and generate an archiving task plan based on the archived data classification results, the shutdown system storage level model, and the real-time system status;

[0093] The archiving task execution module is used to execute rescue archiving operations according to the archiving task plan, store key data according to the optimal path, monitor the archiving process in real time, and record the archiving status.

[0094] A data intelligent archiving management method for a shutdown system is applied to a data intelligent archiving management system for a shutdown system, comprising:

[0095] S1. Collecting the original operating data set in the shutdown system and performing preprocessing operations on the original operating data set to form a preprocessed shutdown system operating data set;

[0096] S2. Perform multi-dimensional feature extraction on the pre-processed shutdown system operation data set, calculate the time sensitivity, access frequency and storage requirements of each data sample in the shutdown system operation data set, and generate a data feature data set;

[0097] S3. Perform global search and parameter optimization on the data feature dataset based on the hybrid fireworks optimization algorithm to form a rescue archive data set that meets the global optimal solution;

[0098] S4. clustering the rescue archived data set that meets the global optimal solution using a dynamic entropy weight clustering method, classifying the data according to the adaptively adjusted weights of each data feature, and generating archived data classification results;

[0099] S5. Optimize the path and allocate storage resources for the archived data in the archived data classification results according to the archived data classification results combined with the shutdown system storage hierarchy model and the real-time system status, and generate an archive task plan;

[0100] S6. When the shutdown system fails or is about to be shut down, the rescue archiving operation needs to complete the data priority screening, path allocation and storage writing in a very short time. According to the archiving task plan, the archiving task execution module is started first. The module stores the key data in batches and in order according to the preset storage priority and the optimized storage path. For data sets with different priorities, a layered archiving strategy is adopted to ensure that ultra-high priority data is written on the fastest storage medium, while lower priority data adopts delayed archiving or remote storage strategies to balance the efficiency of storage resource utilization. A real-time monitoring mechanism is used in the archiving process to monitor the data writing rate, storage space occupancy and the execution progress of the archiving task, and based on the anomaly detection mechanism, storage failures, write delays or data integrity issues are identified, and the storage strategy is adjusted in real time.

[0101] In this implementation, S1 includes the following steps:

[0102] S11. Collect the original operation data set in the shutdown system. The original operation data set includes different types of data samples. Each data sample contains multiple data attributes, which constitute the original operation data set. : ;

[0103] in, Indicates data samples, Indicates The value of a data sample on the mth data attribute, is the total number of data samples, is the total number of data attributes;

[0104] S12. Run the original dataset Perform format conversion to convert data samples in different formats into standardized representations to form a formatted running data set;

[0105] S13. Deduplication of the formatted operation data set is performed to remove data samples with exactly the same content to form a deduplication operation data set;

[0106] S14. Perform anomaly detection on the deduplicated running data set, identify and remove data outliers, and form an anomaly-detected running data set;

[0107] S15. Perform data integrity check on the operation data set after anomaly detection, remove incomplete data, and form a pre-processed shutdown system operation data set : ;

[0108] in, is the data integrity check function, is the running data set after anomaly detection. hour, Satisfy the completeness requirements, otherwise it will be rejected.

[0109] This implementation method realizes efficient cleaning and standardization of the original operation data of the shutdown system through data collection and preprocessing methods, ensuring the accuracy and integrity of the archived data. The step-by-step data format conversion, deduplication, anomaly detection and integrity verification strategies are adopted to enable the data to reach a structured, high-quality and standardized state before entering the archive optimization process. In addition, by constructing a data sample set, data from different data sources are uniformly managed, avoiding the problems of data loss and duplicate storage, and improving the efficiency of data sorting. The integrity and consistency of the data can be effectively guaranteed after the shutdown system is restored.

[0110] In this implementation, S2 includes the following steps:

[0111] S21. Run the shutdown system data set after preprocessing Perform time sensitivity calculations, analyze the time correlation of data samples, and define time sensitivity parameters ,Time sensitivity is used to measure the real-time requirements of shutting down system operation data. The closer the time sensitivity value is to 1, the more urgent the data needs to be archived: ;

[0112] in, Indicates the last update time of the data sample, Indicates the current system time. A parameter for adjusting the magnitude of time sensitivity changes;

[0113] S22. Run the shutdown system data set after preprocessing Calculate the access frequency, count the number of accesses of each data sample, and define the access frequency parameters : ;

[0114] in, Indicates The number of accesses to a data sample within a specified time window, Indicates The number of accesses to a data sample within a specified time window, The access frequency adjustment factor is used to enhance the recognition of high-access data;

[0115] S23. Run the shutdown system data set after preprocessing Perform storage demand calculations, analyze data storage capacity requirements, and define storage demand parameters : ;

[0116] in, Indicates The storage size of data samples, is the storage demand adjustment index, It represents the weighted storage size of the jth sample in the storage requirement calculation, which is used to adjust the influence of data samples on the storage requirement calculation. >1 to enhance the impact of large data samples, <1 to balance the influence of different data samples. It represents the weighted storage size of the i-th sample in the storage requirement calculation;

[0117] S24. Combining time sensitivity parameters , access frequency parameters and storage requirement parameters , calculate the comprehensive weight , which is used to measure the overall importance of each data sample. The larger the comprehensive weight, the data should be archived first: ;

[0118] in, , , is the weight coefficient;

[0119] S25. Based on the calculated comprehensive weight Sort the data samples in the shutdown system operation data set to generate a data feature data set : ;

[0120] in, The archiving threshold is used to filter out data samples that meet the archiving conditions.

[0121] This implementation method realizes the automated priority evaluation of shutdown system data through data feature extraction and evaluation methods, thereby improving the accuracy and rationality of data archiving. Compared with the traditional fixed priority archiving strategy, the multi-dimensional features of time sensitivity, access frequency and storage requirements are used to analyze the importance of data, so that data screening no longer relies on static rules, but dynamically adjusts the archiving strategy according to the actual business value of the data. The time sensitivity calculation method can effectively identify data that needs to be archived urgently in the event of an emergency shutdown, thereby improving the priority management capability of rescue archiving. By calculating the comprehensive weight, it ensures that key data can be archived first in a resource-constrained environment, avoiding the unreasonable archiving problem caused by over-reliance on data types or file structures in traditional methods, thereby significantly improving the adaptability and intelligence level of the archiving strategy.

[0122] In this implementation, S3 includes the following steps:

[0123] S31. Construct a fireworks optimization search space based on a multi-level explosion strategy, targeting a data feature dataset Construct a multi-level explosion optimization model and initialize the fireworks population of the multi-level explosion optimization model , each fireworks individual in the fireworks individual population Represents a candidate subset of salvage archive data: ;

[0124] in, Indicates A subset of candidate salvage archive data for fireworks individuals, is the number of initial fireworks individual population;

[0125] In the multi-level explosion optimization model, a multi-level explosion radius adjustment mechanism is introduced to make each firework individual according to the comprehensive weight. Adjust explosion radius : ;

[0126] in, is the global maximum search radius, Adjust the parameters for the explosion range, Represents data samples The comprehensive weight of Represents data samples The comprehensive weight of

[0127] S32. Adopting an adaptive perturbation mechanism to optimize the generation of candidate rescue archived data subsets, adaptively perturb data of different priorities during the search process, and defining the optimized candidate rescue archived data subset adjustment rules: ;

[0128] in, To indicate the The candidate rescue archive data subset after the optimization of fireworks individuals, and are the addition and removal thresholds, For data samples The probability of being added to the candidate rescue archive data subset: ; For data samples The probability of being removed from the candidate rescue archive data subset: ;

[0129] in, For data samples The probability of being added to the candidate rescue archive data subset;

[0130] S33. In the optimization process, a hierarchical co-evolution strategy is introduced to divide the fireworks individuals into a key data priority layer, a balanced archiving layer, and a storage restricted layer. The corresponding search strategies are adopted respectively. The key data priority layer adopts a fast convergence search, the balanced archiving layer adopts a multi-objective balance strategy, and the storage restricted layer adopts storage optimization adjustment. The data optimization objectives of the key data priority layer, the balanced archiving layer, and the storage restricted layer are defined as follows: ; ; ;

[0131] in, is the objective function of the key data priority layer, is the objective function of the balanced archive layer, is the storage-restricted layer objective function, Indicates The candidate rescue archive data subset after the optimization of fireworks individuals, To store weight adjustment parameters, is the weighted mean of the data samples, For available storage resources;

[0132] S34. Combine the multi-objective optimization strategy to generate the final optimized data set, and perform final screening to select the rescue archive data set that meets the global optimal solution : ;

[0133] in, , , The weight parameter for the archiving optimization objective, The final population set is made to make the rescue archive data set reach the global optimum among key data priority, archive balance and storage utilization.

[0134] This implementation method uses an improved hybrid fireworks optimization algorithm to adaptively adjust the search range through a multi-level explosion strategy, so that data screening is no longer limited to static weights, but can evaluate the business importance of data in real time and dynamically optimize the data screening strategy. In addition, combined with reinforcement learning to optimize the search path, the screening process can gradually approach the optimal data subset, ensuring that the most important data can be stored first. At the same time, the hierarchical co-evolution strategy is adopted, so that data screening not only considers storage resource constraints, but also balances data priority and storage space utilization, thereby improving the accuracy of data screening and the scientific nature of archiving decisions.

[0135] In this implementation, S4 includes the following steps:

[0136] S41. Construct a clustering optimization model based on two-layer dynamic entropy weights for the rescue archived data set that meets the global optimal solution Construct an adaptive entropy weight calculation framework to initialize the data clustering center set of the clustering optimization model , each cluster center in the data cluster center set Represents an archive data category: ;

[0137] in, Indicates Cluster centers, is the total number of initial archived data categories;

[0138] The global entropy weight is calculated by using a two-layer dynamic entropy weight calculation method. and local entropy weight , comprehensively evaluate the importance of data features, and the global entropy weight is calculated as follows: ;

[0139] in, Indicates The normalized probability value of the data sample on the corresponding feature;

[0140] The local entropy weight calculation uses the contribution of data samples to the cluster center, which is defined as follows: ;

[0141] in, Represents the cluster center In Features The value on Represents the cluster center In Features The value on For all cluster centers in the feature The mean on ;

[0142] S42. Use variable weight fuzzy clustering method to calculate data sample membership and define adaptive weighted fuzzy distance , calculate the data sample Cluster Center Membership : ;

[0143] Among them, the adaptive weighted fuzzy distance is defined as follows: ;

[0144] in, is the variable weight index, is the total number of features of the data sample;

[0145] S43. Dynamic update strategy based on cluster drift detection, using cluster center moving trend vector to optimize cluster update process, introducing drift metric Determine whether the cluster center needs to be adjusted: ;

[0146] in, Indicates that it belongs to the cluster center The number of data samples is hour, is the threshold, the cluster center needs to be adjusted, and the adjustment strategy is as follows: ;

[0147] in, To dynamically adjust the step size, is the new cluster center location, For data samples Cluster Center The degree of membership;

[0148] S44. Use the fuzzy entropy convergence criterion to optimize the clustering termination condition and calculate the fuzzy entropy of the clustering system Measuring the uncertainty of a clustering: ;

[0149] Set the dynamic entropy convergence threshold ,when And the cluster center changes satisfy When , the clustering is terminated, otherwise, S42-S43 is continued to update the cluster center;

[0150] S45. Based on the optimized cluster center and membership For data samples Assign archiving categories and generate archiving data classification results: ;

[0151] in, Indicates the final classification result of archived data. For data samples The archive category to which it belongs, making the rescue archive data collection are effectively classified, To obtain The value of the variable that reaches its maximum value.

[0152] This implementation method realizes the intelligence and adaptability of data classification through a two-layer dynamic entropy weight clustering method, improves the accuracy and stability of archived data classification, and introduces a two-layer calculation framework of global entropy weight and local entropy weight, so that data clustering can automatically adjust feature weights according to changes in data distribution, thereby improving the flexibility of data classification. In addition, the adaptability of the clustering process is further improved by adopting a variable weight fuzzy clustering method, so that different types of data can be accurately classified according to real-time status, and the cluster drift detection mechanism effectively avoids the problem of misclassification of data categories, improves the stability and consistency of data archiving classification, and ensures that key data can be archived to the most appropriate storage level.

[0153] In this implementation, S5 includes the following steps:

[0154] S51. Build a path optimization framework based on the shutdown system storage hierarchy model to classify the archived data Optimize the archive path and initialize the storage hierarchy of the shutdown system , where each storage tier Indicates an archive path to a different storage medium: ;

[0155] in, Indicates storage tiers, is the total number of storage tiers, Represents the collection of available storage resources in the system;

[0156] Introducing data-based urgency Storage tier selection function : ;

[0157] in, For data samples The urgency of For storage tier The read and write delays For storage tier The unit storage cost is set so that data with a higher urgency level than the preset value is stored in high-speed storage media first, and data with a lower urgency level than the preset value is stored in low-cost storage media;

[0158] S52. Calculate the optimal storage allocation plan for the data archiving path, and perform archiving of data samples based on the storage hierarchy model. Calculate its optimal storage location : ;

[0159] in, is the comprehensive weight, is the storage size of the data sample, For storage tier of available storage space, For storage tier The historical visit frequency, To optimize the weight coefficient, the storage allocation scheme takes into account data priority, storage space utilization and data access requirements;

[0160] S53. Generate the final archiving task plan based on the optimal storage location : ;

[0161] in, For data samples The archive timestamp of The archiving trigger threshold is set so that the rescue archive data set is stored in the optimal path.

[0162] This implementation improves the rationality of data storage and the execution efficiency of archiving tasks through storage path optimization and resource allocation strategies, and introduces a storage level selection method based on data urgency, so that the data storage path can be dynamically optimized according to business needs, ensuring that critical data is allocated to high-speed storage media first, while low-priority data is stored in lower-cost storage media.

[0163] Embodiment 1:

[0164] Example 1 occurred at 3:27 am on June 15, 2024, when a financial institution's data center suddenly shut down. When the accident occurred, the core storage cluster "DC-Storage-003" of the data center unexpectedly lost power due to a power line failure, resulting in IO errors in some storage nodes, and about 14.6TB of transaction data, account change records and system logs failed to be archived normally. The last log timestamp of the database server "DB-Node-05" before the failure was 2024-06-1503:26:58, which means that most of the transaction data had not been written within 1 minute before the system crashed.

[0165] 03:27:10—The accident detection module found that the storage node "DC-Storage-003" had generated 37 abnormal IO failure alarms in the past 15 seconds, far exceeding the threshold (20 times). The system determined that "DC-Storage-003" might be suddenly shut down, triggering the data rescue archiving strategy.

[0166] 03:27:18—The data collection module began to scan the active database transaction logs within 30 minutes before the failure (a total of 32,658 unfinished transactions), involving 1,258,402 user transaction records, with a total data volume of approximately 1.2TB. The system also detected that the transaction log file "trx_log_20240615.log" on "DB-Node-05" was accidentally truncated, with a file size of only 872MB, which was 62% less than the expected size (2.3GB). The system determined that some transaction log data was lost.

[0167] 03:27:35—Format conversion and data cleaning began. After the abnormal transaction log was parsed, the system automatically filtered out low-priority log items that were not related to storage errors, such as query transactions without account changes (7,532 in total), and finally obtained 27,894 high-priority transaction records to be rescued, with a total data volume of 982GB.

[0168] 03:28:12—The hybrid fireworks optimization algorithm starts running to screen and optimize the archived data.

[0169] The system calculates the time sensitivity of the transaction log, with the highest value being 0.98 (i.e., data that is not archived at T+0 minutes is extremely important).

[0170] The results of the access frequency analysis showed that the top 10% of accounts were involved in a total of 28,602 transactions, and these transaction data were automatically marked as high priority.

[0171] The calculation results of the storage requirements of log files show that high-priority transaction data accounts for 73.4% of the total data volume.

[0172] 03:28:28—The hybrid fireworks optimization algorithm completes the calculation, generates a rescue archive data set, and selects 750GB of transaction data as the highest priority archiving target. The account ID range involved is: [A0456723-A0973456], and the transaction order ID range is: [TXN-09873423-TXN-09895687].

[0173] 03:28:50—The dynamic entropy weight clustering method is started to automatically classify the 750GB of filtered data.

[0174] The transaction order data is clustered into the “ultra-high priority group” (412GB in total), which contains the high-frequency trading user group (217,634 transaction records).

[0175] The account change data is clustered into the “high priority group” (184GB in total), and the total number of accounts involved in financial data changes is 5,486.

[0176] The log data is clustered into the "medium priority group" (154 GB in total), which is mainly used for system recovery analysis.

[0177] 03:29:10—The storage optimization module is started to optimize the data archiving path.

[0178] The data of the "ultra-high priority group" is allocated to the NVMeSSD storage node (DC-Storage-001), with an estimated write speed of approximately 2.3GB / s and an estimated completion time of within 5 minutes and 45 seconds.

[0179] The data in the "high priority group" is allocated to the HDD array (DC-Storage-002), with an estimated write speed of approximately 700MB / s and an estimated completion time of 4 minutes and 30 seconds.

[0180] The data of the "medium priority group" is temporarily stored in the remote backup server (DC-Backup-004) and is planned to be synchronized to long-term storage after the system is restored.

[0181] 03:29:45—The data archiving task is officially started. The writing progress is as follows:

[0182] 03:30:00—NVMeSSD storage has completed data archiving of 85GB, with a progress of approximately 21%.

[0183] 03:31:30—HDD storage has completed archiving 94GB, with a progress of about 51%.

[0184] 03:34:10—All data is archived, and the log record "archive_log_20240615.log" is generated, confirming that 750GB of data has been successfully archived, with an archiving accuracy rate of 99.7%.

[0185] Within 30 minutes after the system was restored, all high-priority transaction data was restored, the system started normally, and order settlement reconciliation was completed.

[0186] The rescue archived data for this time totals 750GB. Compared with the traditional regular backup method, the method of the present invention performs well in archiving efficiency, data recovery time and storage optimization. The specific improvements are as follows:

[0187] 1. The key data archiving rate increased by 26.5%, ensuring that important transaction data is stored first and reducing the risk of data loss;

[0188] 2. The archiving execution time is shortened by 17 minutes (about 40.5%), which greatly improves the data recovery speed in the event of an unexpected shutdown;

[0189] 3. Storage space utilization increased by 17.1%, effectively reducing the waste of resources of low-priority data occupying high-performance storage devices;

[0190] 4. The critical data recovery time is shortened to 7 minutes (reduced by 61%), ensuring business continuity and reducing the impact of shutdown accidents on operations;

[0191] 5. The archiving success rate is close to 100%, and business data can be automatically restored after the system is restored without additional manual intervention.

[0192] This embodiment verifies the efficiency and accuracy of the present invention in rescue archiving through an unexpected shutdown of a real financial system. It uses hybrid fireworks optimization, dynamic entropy weight clustering, and storage path optimization algorithms to break through the limitations of traditional data archiving methods, and achieves intelligent archiving and rapid recovery of key data in unexpected shutdown situations, effectively improving the business continuity and data protection capabilities of the shutdown system.

[0193] The present invention adopts an improved hybrid fireworks optimization algorithm and a multi-level explosion radius adaptive adjustment mechanism, so that data screening no longer relies on fixed rules, but adaptively adjusts the search range according to the time sensitivity, access frequency, and storage requirement key features of the data, thereby improving the recognition accuracy of key data. The multi-level explosion radius adjustment strategy enables high-priority data to obtain a greater search weight in the global search process, ensuring that key data can be preferentially selected into the archived data set.

[0194] The present invention adopts a two-layer dynamic entropy weight clustering method in the data classification process to calculate the global entropy weight and local entropy weight of data features respectively, so that data classification can dynamically adjust the influence weights of different features according to the system state. The two-layer entropy weight calculation framework enhances the stability and accuracy of data classification through an adaptive weighted fuzzy distance calculation method.

[0195] In terms of data storage path optimization, the present invention proposes a storage level selection function based on data urgency, and combines it with reinforcement learning to optimize the storage path allocation strategy, so that data archiving can be dynamically allocated according to the real-time system status. The storage path selection is continuously adjusted through the reinforcement learning method, so that high-priority data can be allocated to high-speed storage media with lower storage latency, and the data migration plan is dynamically adjusted to ensure efficient use of storage space.

[0196] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A data intelligent archiving management system for shutdown systems, characterized in that: Includes the following modules: The data collection and preprocessing module is used to collect the original operation data set in the shutdown system and perform preprocessing to form the shutdown system operation data set; The data feature extraction and evaluation module is used to extract multi-dimensional features of the shutdown system operation data set, calculate the time sensitivity, access frequency and storage requirements of the data, and calculate the importance weight of the data based on the dynamic entropy weight to form a data feature data set; The data screening module based on hybrid fireworks optimization is used to perform global search and parameter optimization on data feature data sets. The optimal data subset is selected based on the hybrid fireworks optimization algorithm combined with shutdown system status, data priority and storage resource constraints to form an optimized rescue archive data set; The dynamic entropy weight cluster analysis module is used to perform dynamic entropy weight cluster analysis on the optimized rescue archived data set to generate archived data classification results; The archiving path optimization and storage resource allocation module is used to optimize the path and allocate storage resources for the archiving data in the archiving data classification results and generate an archiving task plan by combining the archiving data classification results, the shutdown system storage hierarchy model and the real-time system status; The archiving task execution module is used to execute rescue archiving operations according to the archiving task plan, so that the shutdown system operation data set is stored according to the optimal path and the archiving status is recorded.

2. A data intelligent archiving management method for a shutdown system, applied to the data intelligent archiving management system for a shutdown system according to claim 1, characterized in that: include: S1. Collecting the original operating data set in the shutdown system and performing preprocessing operations on the original operating data set to form a preprocessed shutdown system operating data set; S2. Perform multi-dimensional feature extraction on the pre-processed shutdown system operation data set, calculate the time sensitivity, access frequency and storage requirements of each data sample in the shutdown system operation data set, and generate a data feature data set; S3. Perform global search and parameter optimization on the data feature dataset based on the hybrid fireworks optimization algorithm to form a rescue archive data set that meets the global optimal solution; S4. clustering the rescue archived data set that meets the global optimal solution using a dynamic entropy weight clustering method, classifying the data according to the adaptively adjusted weights of each data feature, and generating archived data classification results; S5. Optimize the path and allocate storage resources for the archived data in the archived data classification results according to the archived data classification results combined with the shutdown system storage hierarchy model and the real-time system status, and generate an archive task plan; S6. Perform rescue archiving operations on the optimized rescue archiving data set in the actual environment of the shutdown system according to the archiving task plan, monitor the archiving process in real time and record the archiving status to form an archiving execution record.

3. The method for intelligent data archiving management for shutdown systems according to claim 2, characterized in that: The S1 comprises the following steps: S11. Collect the original operating data set in the shutdown system ; S12. Run the original dataset Perform format conversion to form a formatted running data set; S13. Deduplication of the formatted operation data set to form a deduplicated operation data set; S14. Performing anomaly detection on the deduplicated running data set to form an anomaly detected running data set; S15. Perform data integrity check on the operation data set after anomaly detection to form a pre-processed shutdown system operation data set .

4. The method for intelligent data archiving management for shutdown systems according to claim 3, characterized in that: The S2 comprises the following steps: S21. Run the shutdown system data set after preprocessing Perform time sensitivity calculations, analyze the time correlation of data samples, and define time sensitivity parameters ; S22. Run the shutdown system data set after preprocessing Calculate the access frequency, count the number of accesses of each data sample, and define the access frequency parameters ; S23. Run the shutdown system data set after preprocessing Perform storage demand calculations, analyze data storage capacity requirements, and define storage demand parameters ; S24. Combining time sensitivity parameters , access frequency parameters and storage requirement parameters , calculate the comprehensive weight , used to measure the overall importance of each data sample. The larger the comprehensive weight, the data should be archived first; S25. Based on the calculated comprehensive weight Sort the data samples in the shutdown system operation data set to generate a data feature data set .

5. The method for intelligent data archiving management for shutdown systems according to claim 4, characterized in that: The S3 comprises the following steps: S31. Construct a fireworks optimization search space based on a multi-level explosion strategy, targeting a data feature dataset Construct a multi-level explosion optimization model and initialize the fireworks population of the multi-level explosion optimization model , each fireworks individual in the fireworks individual population represents a candidate salvage archive data subset; In the multi-level explosion optimization model, a multi-level explosion radius adjustment mechanism is introduced to make each firework individual according to the comprehensive weight. Adjust explosion radius : ; in, is the global maximum search radius, Adjust the parameters for the explosion range, Represents data samples The comprehensive weight of Represents data samples The comprehensive weight of S32. Adopting an adaptive perturbation mechanism to optimize the generation of candidate rescue archived data subsets, adaptively perturb data of different priorities during the search process, and defining the optimized candidate rescue archived data subset adjustment rules: ; in, To indicate the The candidate rescue archive data subset after the optimization of fireworks individuals, and are the addition and removal thresholds, For data samples The probability of being added to the candidate rescue archive data subset, For data samples The probability of being removed from the candidate salvage archive data subset; S33. In the optimization process, a hierarchical co-evolution strategy is introduced to divide the fireworks individuals into a key data priority layer, a balanced archiving layer, and a storage restricted layer. The corresponding search strategies are adopted respectively. The key data priority layer adopts a fast convergence search, the balanced archiving layer adopts a multi-objective balance strategy, and the storage restricted layer adopts storage optimization adjustment. The data optimization objectives of the key data priority layer, the balanced archiving layer, and the storage restricted layer are defined as follows: ; ; ; in, is the objective function of the key data priority layer, is the objective function of the balanced archive layer, is the storage-restricted layer objective function, Indicates The candidate rescue archive data subset after the optimization of fireworks individuals, To store weight adjustment parameters, is the weighted mean of the data samples, is the available storage resources; S34. Combine the multi-objective optimization strategy to generate the final optimized data set, and perform final screening to select the rescue archive data set that meets the global optimal solution .

6. The method for intelligent data archiving management for shutdown systems according to claim 5, characterized in that: The S4 comprises the following steps: S41. Construct a clustering optimization model based on two-layer dynamic entropy weights for the rescue archived data set that meets the global optimal solution Construct an adaptive entropy weight calculation framework to initialize the data clustering center set of the clustering optimization model , each cluster center in the data cluster center set Represents an archive data category; The global entropy weight is calculated by using a two-layer dynamic entropy weight calculation method. and local entropy weight ,The importance of data features is comprehensively evaluated, and the local entropy weight calculation adopts the contribution of data samples relative to the cluster center; S42. Use variable weight fuzzy clustering method to calculate data sample membership and define adaptive weighted fuzzy distance , calculate the data sample Cluster Center Membership ; S43. Dynamic update strategy based on cluster drift detection, using cluster center moving trend vector to optimize cluster update process, introducing drift metric Determine whether the cluster center needs to be adjusted ,when hour, is the threshold, the cluster center needs to be adjusted; S44. Use the fuzzy entropy convergence criterion to optimize the clustering termination condition and calculate the fuzzy entropy of the clustering system Measuring uncertainty in clustering , set the dynamic entropy convergence threshold ,when And the cluster center changes satisfy When , the clustering is terminated, otherwise, S42-S43 is continued to update the cluster center; S45. Based on the optimized cluster center and membership For data samples Assign archive categories and generate archive data classification results .

7. The method for intelligent data archiving management for shutdown systems according to claim 6, characterized in that: The S5 comprises the following steps: S51. Build a path optimization framework based on the shutdown system storage hierarchy model to classify the archived data Optimize the archive path and initialize the storage hierarchy of the shutdown system , where each storage tier Indicates an archive path to a different storage medium; Introducing data-based urgency Storage tier selection function : ; in, For data samples The urgency of For storage tier The read and write delays For storage tier The unit storage cost is set so that data with a higher urgency level than the preset value is stored in high-speed storage media first, and data with a lower urgency level than the preset value is stored in low-cost storage media; S52. Calculate the optimal storage allocation plan for the data archiving path, and perform archiving of data samples based on the storage hierarchy model. Calculate its optimal storage location ; S53. Generate the final archiving task plan based on the optimal storage location .

Citation Information

Patent Citations

  • Data dump method and system

    CN103106247A

  • Data retention management

    CN103631849A

  • Short-term load forecasting method for microgrid based on independent component analysis and support vector machine

    CN109345027A

  • Automatic filing method and system for personnel archives

    CN117827750A

  • Enterprise digital management system

    CN119067597A

Cited By

  • Data rescue and filing system based on cloud platform

    CN120315653A

  • Financial transaction historical data storage optimization method

    CN121365043A