A data intelligent archiving management system for shutdown systems

Through the intelligent data archiving management system of multi-dimensional feature extraction and dynamic weight calculation, the problems of low data archiving efficiency and low classification accuracy in the shutdown system are solved, and the rapid and accurate archiving and storage optimization of key data is achieved, ensuring the business continuity of the system.

CN119989033BActive Publication Date: 2025-07-22HANGZHOU YIKANGXIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510474997.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-22
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing shutdown system data archiving methods have problems such as low archiving efficiency, lack of dynamic adaptability, insufficient storage resource optimization and low data classification accuracy in sudden shutdown events, resulting in the inability to archive key data in a timely manner, increasing the risk of data loss.

Method used

The data acquisition and preprocessing module, data feature extraction and evaluation module, data screening module based on hybrid firework optimization, dynamic entropy weight clustering analysis module and archive path optimization and storage resource allocation module are used to ensure that key data is preferred to enter the archived data collection through multi-dimensional feature extraction, dynamic weight calculation and adaptive adjustment.

Benefits of technology

It improves the identification accuracy and archiving efficiency of key data, ensures rapid data recovery in sudden shutdowns, reduces the risk of data loss, and improves the rationality and business continuity of data storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989033B_ABST
    Figure CN119989033B_ABST
Patent Text Reader

Abstract

The present invention discloses a data intelligent archiving management system for a shutdown system, and the system includes: a data collection and preprocessing module for collecting the original operation data set in the shutdown system; a data feature extraction and evaluation module for performing multi-dimensional feature extraction on the preprocessed operation data set of the shutdown system; a data screening module based on hybrid fireworks optimization for performing global search and parameter optimization on the data feature data set; a dynamic entropy weight clustering analysis module for performing dynamic entropy weight clustering analysis on the optimized rescue archiving data set; an archiving path optimization and storage resource allocation module for combining the archiving data classification result, the shutdown system storage hierarchy model and the real-time system state; and an archiving task execution module for performing rescue archiving operations according to the archiving task plan. The present invention ensures that key data can be preferentially selected to enter the archiving data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of archiving management systems, and in particular to an intelligent data archiving management system for shutdown systems. Background Art

[0002] With the development of information technology, the data security and storage management of shutdown systems have gradually become an important research direction in the field of informatization. In enterprise data centers, industrial control systems, and cloud computing environments, sudden shutdowns of systems may be caused by various factors such as equipment failures, network anomalies, software vulnerabilities, and malicious attacks. During this process, the critical data in the system often faces high risks. If not archived and stored in a timely manner, it may lead to data loss, business interruption, and even affect the normal operation of enterprises.

[0003] Currently, the mainstream methods for archiving data of shutdown systems mainly include three methods: preset rule archiving, static priority archiving, and regular backup. Each of the existing methods has its own advantages and disadvantages, but the following technical defects are exposed in the case of sudden shutdown scenarios:

[0004] Traditional preset rule archiving relies on pre-established policies, such as determining the importance of data based on file type, directory structure, or historical access frequency. However, in the case of sudden shutdowns, the importance of data may change dynamically, and the preset rules cannot respond in real time, resulting in critical data not being archived in a timely manner and increasing the risk of data loss.

[0005] The static priority archiving method divides the storage priority of data according to fixed rules. For example, certain business databases are set as high priority, and ordinary documents are set as low priority. However, the data status in the shutdown system may change rapidly in a short period of time. For example, some real-time transaction data or log files may become crucial at critical moments, and the static priority archiving lacks the ability to dynamically adjust, easily missing high-priority data.

[0006] Existing regular backup methods usually use preset storage paths for data storage without considering the real-time state of system resources, resulting in high-priority data being possibly stored in low-speed storage media, affecting the rapid recovery of data. At the same time, the allocation of storage resources lacks optimization, which may lead to the failure of archiving tasks.

[0007] In summary, the existing methods for archiving data of shutdown systems have great limitations in dealing with sudden shutdown events, mainly reflected in low archiving efficiency, lack of dynamic adaptability, insufficient optimization of storage resources, and low data classification accuracy. There is an urgent need for an intelligent data archiving management system for shutdown systems to improve the accuracy and response speed of data rescue archiving. Summary of the Invention

[0008] An object of the present invention is to provide a data intelligent archiving management system for shutdown systems, and the present invention ensures that critical data can be preferentially selected to enter the archived data set.

[0009] A data intelligent archiving management system for shutdown systems according to an embodiment of the present invention includes the following modules:

[0010] A data collection and preprocessing module, configured to collect the original operation data set in the shutdown system and perform preprocessing to form a shutdown system operation data set;

[0011] A data feature extraction and evaluation module, configured to perform multi-dimensional feature extraction on the shutdown system operation data set, calculate the time sensitivity, access frequency, and storage requirements of the data, and calculate the importance weight of the data based on dynamic entropy weight to form a data feature data set;

[0012] A data screening module based on hybrid fireworks optimization, configured to perform global search and parameter optimization on the data feature data set, and select the optimal data subset based on the hybrid fireworks optimization algorithm in combination with the shutdown system state, data priority, and storage resource constraints to form an optimized rescue archiving data set;

[0013] A dynamic entropy weight clustering analysis module, configured to perform dynamic entropy weight clustering analysis on the optimized rescue archiving data set to generate an archiving data classification result;

[0014] An archiving path optimization and storage resource allocation module, configured to combine the archiving data classification result, the shutdown system storage hierarchy model, and the real-time system state to perform path optimization and storage resource allocation on the archiving data in the archiving data classification result to generate an archiving task plan;

[0015] An archiving task execution module, configured to perform a rescue archiving operation according to the archiving task plan, store the shutdown system operation data set along the optimal path, and record the archiving status.

[0016] A data intelligent archiving management method for shutdown systems, applied to a data intelligent archiving management system for shutdown systems, includes:

[0017] S1. Collect the original operation data set in the shutdown system, and perform preprocessing operations on the original operation data set to form a preprocessed shutdown system operation data set;

[0018] S2. Perform multi-dimensional feature extraction on the preprocessed shutdown system operation data set, calculate the time sensitivity, access frequency, and storage requirements of each data sample in the shutdown system operation data set, and generate a data feature data set;

[0019] S3. Perform a global search and parameter optimization on the data feature dataset based on the hybrid fireworks optimization algorithm to form a rescue archiving data set that meets the global optimal solution;

[0020] S4. Use the dynamic entropy weight clustering method to perform clustering processing on the rescue archiving data set that meets the global optimal solution, classify the data according to the weights of each data feature adjusted adaptively, and generate the classification result of the archiving data;

[0021] S5. According to the classification result of the archiving data, combined with the shutdown system storage hierarchy model and the real-time system state, perform path optimization and storage resource allocation on the archiving data in the classification result of the archiving data, and generate an archiving task plan;

[0022] S6. Implement the rescue archiving operation on the optimized rescue archiving data set in the actual environment of the shutdown system according to the archiving task plan, monitor the archiving process in real time and record the archiving status to form an archiving execution record.

[0023] Optionally, the S1 includes the following steps:

[0024] S11. Collect the original operation data set in the shutdown system. The original operation data set includes different types of data samples, and each data sample contains multiple data attributes, constituting the original operation data set :

[0025] ;

[0026] Among them, represents the th data sample, represents the value of the th data sample on the mth data attribute, is the total number of data samples, is the total number of data attributes;

[0027] S12. Perform format conversion on the original operation data set to uniformly convert data samples in different formats into a standardized representation form to form a formatted operation data set;

[0028] S13. Remove duplicates from the formatted operation data set, and eliminate data samples with exactly the same content to form a deduplicated operation data set;

[0029] S14. Perform anomaly detection on the deduplicated operation data set, identify and eliminate data outliers to form an operation data set after anomaly detection;

[0030] S15. Perform data integrity verification on the operation dataset after anomaly detection, remove incomplete data, and form the operation dataset of the shutdown system after preprocessing :

[0031] ;

[0032] Among them, is the data integrity verification function, is the operation dataset after anomaly detection. When , meets the integrity requirements, otherwise it is removed.

[0033] Optionally, the above S2 includes the following steps:

[0034] S21. Calculate the time sensitivity of the operation dataset of the shutdown system after preprocessing , analyze the time correlation of data samples, and define the time sensitivity parameter . The time sensitivity is used to measure the real-time requirement of the operation data of the shutdown system. The closer the time sensitivity value is to 1, the more urgently the data needs to be archived:

[0035] ;

[0036] Among them, represents the last update time of the data sample, represents the current system time, is the parameter for adjusting the change range of time sensitivity;

[0037] S22. Calculate the access frequency of the operation dataset of the shutdown system after preprocessing , count the access times of each data sample, and define the access frequency parameter :

[0038] ;

[0039] Among them, represents the access times of the th data sample within the specified time window, represents the access times of the th data sample within the specified time window, is the access frequency adjustment coefficient, which is used to enhance the recognition ability of high-access data;

[0040] S23. Calculate the storage requirement of the operation dataset of the shutdown system after preprocessing , analyze the storage capacity requirements of the data, and define the storage requirement parameter :

[0041] ;

[0042] Among them, represents the storage size of the th data sample, is the storage requirement adjustment index, represents the weighted storage size of the jth sample in the storage requirement calculation, which is used to adjust the influence degree of the data sample on the storage requirement calculation, enhances the influence of large data samples when > 1, balances the influence of different data samples when

[0043] and the storage requirement parameter to calculate the comprehensive weight

[0044]

[0045] ; Among them, , , are weight coefficients;

[0046] S25. Sort the data samples in the shutdown system operation dataset according to the calculated comprehensive weight to generate a data feature dataset :

[0047] ;

[0048] Among them, is the filing threshold, which is used to screen out data samples that meet the filing conditions.

[0049] Optionally, the S3 includes the following steps:

[0050] S31. Construct a fireworks optimization search space based on a multi-level explosion strategy, and construct a multi-level explosion optimization model for the data feature dataset , and initialize the fireworks individual population of the multi-level explosion optimization model. Each fireworks individual in the fireworks individual population represents a candidate rescue filing data subset:

[0051] ;

[0052] Among them, Represents the th candidate rescue archival data subset of a fireworks individual, is the initial population size of fireworks individuals;

[0053] In the multi-stage explosion optimization model, a multi-stage explosion radius adjustment mechanism is introduced, enabling each fireworks individual to adjust the explosion radius according to the comprehensive weight : :

[0054] ;

[0055] Among them, is the global maximum search radius, is the explosion range adjustment parameter, represents the data sample of the comprehensive weight, represents the data sample of the comprehensive weight;

[0056] S32. Optimize the generation of candidate rescue archival data subsets using an adaptive perturbation mechanism. During the search process, perform adaptive perturbation on data with different priorities, and define the adjustment rules for the optimized candidate rescue archival data subsets:

[0057] ;

[0058] Among them, represents the th optimized candidate rescue archival data subset of a fireworks individual, and are the addition and removal thresholds respectively, is the probability that the data sample is added to the candidate rescue archival data subset:

[0059] ;

[0060] is the probability that the data sample is removed from the candidate rescue archival data subset:

[0061] ;

[0062] Among them, is the probability that the data sample is added to the candidate rescue archival data subset;

[0063] S33. Introduce a hierarchical co-evolution strategy during the optimization process. Divide the fireworks individuals into a critical data priority layer, a balanced archiving layer, and a storage-constrained layer, and adopt corresponding search strategies respectively. The critical data priority layer adopts a fast convergence search, the balanced archiving layer adopts a multi-objective balance strategy, and the storage-constrained layer adopts storage optimization adjustment. The data optimization objectives of the critical data priority layer, the balanced archiving layer, and the storage-constrained layer are defined as follows:

[0064] ;

[0065] ;

[0066] ;

[0067] Among them, is the objective function of the critical data priority layer, is the objective function of the balanced archiving layer, is the objective function of the storage-constrained layer, represents the th candidate rescue archiving data subset after optimizing the fireworks individual, is the storage weight adjustment parameter, is the mean value of the data sample weights, is the available storage resource;

[0068] S34. Generate the final optimized data set by combining the multi-objective optimization strategy, and perform the final screening to select the rescue archiving data set that meets the global optimal solution :

[0069] ;

[0070] Among them, , , are the weight parameters of the archiving optimization objective, is the final population set, which makes the rescue archiving data set achieve the global optimal solution among the critical data priority, archiving balance, and storage utilization rate.

[0071] Optionally, the S4 includes the following steps:

[0072] S41. Construct a clustering optimization model based on double-layer dynamic entropy weight. For the rescue archiving data set that meets the global optimal solution construct an adaptive entropy weight calculation framework and initialize the data clustering center set of the clustering optimization model , and each clustering center in the data clustering center set represents an archiving data category:

[0073] ;

[0074] Among them, represents the th cluster center, is the total number of initial archived data categories;

[0075] The double-layer dynamic entropy weight calculation method is used to calculate the global entropy weight and the local entropy weight respectively, to comprehensively evaluate the importance of data features. The calculation of the global entropy weight is as follows:

[0076] ;

[0077] Among them, represents the normalized probability value of the th data sample on the corresponding feature;

[0078] The calculation of the local entropy weight uses the contribution degree of the data sample to the cluster center and is defined as follows:

[0079] ;

[0080] Among them, represents the value of the cluster center on the feature , represents the value of the cluster center on the feature , is the mean value of all cluster centers on the feature ;

[0081] S42. The variable weight fuzzy clustering method is used to calculate the membership degree of the data sample. Define the adaptive weighted fuzzy distance , and calculate the membership degree of the data sample to the cluster center :

[0082] ;

[0083] Among them, the adaptive weighted fuzzy distance is defined as follows:

[0084] ;

[0085] Among them, is the variable weight index, is the total number of features of the data sample;

[0086] S43. Based on the dynamic update strategy of clustering drift detection, the clustering center movement trend vector is used to optimize the clustering update process, and the drift metric is introduced to judge whether the cluster center needs to be adjusted:

[0087] ;

[0088] Among them, represents the number of data samples belonging to the cluster center. When , is the threshold, and the cluster center needs to be adjusted. The adjustment strategy is as follows: is the threshold, and the cluster center needs to be adjusted. The adjustment strategy is as follows:

[0089] ;

[0090] Among them, is the dynamic adjustment step size, is the position of the new cluster center, is the data sample for the cluster center membership degree;

[0091] S44. Optimize the clustering termination condition using the fuzzy entropy convergence criterion, and calculate the fuzzy entropy of the clustering system to measure the uncertainty of clustering:

[0092] ;

[0093] Set the dynamic entropy convergence threshold . When and the change of the cluster center satisfies , terminate the clustering; otherwise, continue to execute S42 - S43 to update the cluster center;

[0094] S45. According to the optimized cluster center and the membership degree for the data sample allocate the archival category, and generate the archival data classification result:

[0095] ;

[0096] Among them, represents the finally generated archival data classification result, is the archival category to which the data sample belongs, so that the rescue archival data set is effectively classified, is to obtain the variable value that makes reach the maximum value.

[0097] Optionally, the above - mentioned S5 includes the following steps:

[0098] S51. Construct a path optimization framework based on the storage - level model of the shutdown system, for the archival data classification result Optimize the archiving path and initialize the storage hierarchy of the shutdown system , where each storage hierarchy represents an archiving path for a different storage medium:

[0099] ;

[0100] Among them, represents the th storage hierarchy, is the total number of storage hierarchies, represents the set of available storage resources of the system;

[0101] Introduce a storage hierarchy selection function based on data urgency :

[0102] ;

[0103] Among them, is the urgency of the data sample , is the read / write latency of the storage hierarchy , is the unit storage cost of the storage hierarchy , so that data with an urgency higher than the preset value is preferentially stored in high-speed storage media, and data with an urgency lower than the preset value is stored in low-cost storage media;

[0104] S52. Calculate the optimal storage allocation plan for the data archiving path, and calculate the optimal storage location of the archived data sample on the basis of the storage hierarchy model :

[0105] ;

[0106] Among them, is the comprehensive weight, is the storage size of the data sample, is the available storage space of the storage hierarchy , is the historical access frequency of the storage hierarchy , is the optimization weight coefficient, so that the storage allocation plan takes into account data priority, storage space utilization rate and data access requirements;

[0107] S53. Generate the final archiving task plan based on the optimal storage location :

[0108] ;

[0109] Among them, is the archiving timestamp of the data sample , and is the archiving trigger threshold, enabling the rescue archiving data set to be stored along the optimal path.

[0110] The beneficial effects of the present invention are as follows:

[0111] (1) Through the improved hybrid fireworks optimization algorithm, the present invention adopts a multi-level explosion radius adaptive adjustment mechanism, enabling data screening to no longer rely on fixed rules but to adaptively adjust the search range according to the key features of data time sensitivity, access frequency, and storage requirements, thereby improving the recognition accuracy of key data. The adjustment strategy of the multi-level explosion radius enables high-priority data to obtain a greater search weight during the global search process, ensuring that key data can be preferentially selected into the archived data set.

[0112] (2) During the data classification process, the present invention adopts a two-layer dynamic entropy weight clustering method to calculate the global entropy weight and local entropy weight of data features respectively, enabling data classification to dynamically adjust the influence weights of different features according to the system state. The two-layer entropy weight calculation framework enhances the stability and accuracy of data classification through an adaptive weighted fuzzy distance calculation method.

[0113] (3) In terms of optimizing the data storage path, the present invention proposes a storage level selection function based on data urgency and combines reinforcement learning to optimize the storage path allocation strategy, enabling data archiving to perform dynamic storage allocation according to the real-time system state. By continuously adjusting the storage path selection through the reinforcement learning method, high-priority data can be allocated to high-speed storage media with lower storage latency, and the data migration plan is dynamically adjusted to ensure the efficient utilization of storage space. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0115] Figure 1 is a flowchart of a data intelligent archiving management system for a shutdown system proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0116] The present invention will now be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0117] Refer to Figure 1 , a data intelligent archiving management system for a shutdown system, including the following modules:

[0118] A data collection and preprocessing module, which is used to collect the original operation data set in the shutdown system, perform format conversion, data deduplication, anomaly detection and data integrity verification on the original operation data set, and form a preprocessed operation data set of the shutdown system;

[0119] A data feature extraction and evaluation module, which is used to perform multi-dimensional feature extraction on the preprocessed operation data set of the shutdown system, calculate the time sensitivity, access frequency and storage requirements of the data, and calculate the importance weight of the data based on dynamic entropy weight to form a data feature data set;

[0120] A data screening module based on hybrid fireworks optimization, which is used to perform global search and parameter optimization on the data feature data set, select the optimal data subset based on the hybrid fireworks optimization algorithm combined with the shutdown system status, data priority and storage resource constraints, and form an optimized rescue archiving data set;

[0121] A dynamic entropy weight clustering analysis module, which is used to perform dynamic entropy weight clustering analysis on the optimized rescue archiving data set, adaptively adjust the data classification weight according to the global entropy weight and local entropy weight of the data features, perform fuzzy clustering, and calculate the archiving data category based on the membership degree to generate an archiving data classification result;

[0122] An archiving path optimization and storage resource allocation module, which is used to combine the archiving data classification result, the shutdown system storage hierarchy model and the real-time system status, perform path optimization and storage resource allocation on the archiving data in the archiving data classification result, calculate the optimal storage hierarchy, and dynamically adjust the data migration strategy to generate an archiving task plan;

[0123] An archiving task execution module, which is used to execute the rescue archiving operation according to the archiving task plan, store the key data according to the optimal path, and perform real-time monitoring on the archiving process and record the archiving status.

[0124] A data intelligent archiving management method for a shutdown system, which is applied to a data intelligent archiving management system for a shutdown system, including:

[0125] S1. Collect the original operation data set in the shutdown system, and perform preprocessing operations on the original operation data set to form a preprocessed operation data set of the shutdown system;

[0126] S2. Perform multi-dimensional feature extraction on the preprocessed operation data set of the shutdown system, calculate the time sensitivity, access frequency and storage requirements of each data sample in the operation data set of the shutdown system, and generate a data feature data set;

[0127] S3. Perform global search and parameter optimization on the data feature data set based on the hybrid fireworks optimization algorithm to form a rescue archiving data set that meets the global optimal solution;

[0128] S4. Use the dynamic entropy weight clustering method to cluster the set of rescue archiving data that meets the global optimal solution, classify the data according to the adaptively adjusted weights of each data feature, and generate the classification result of the archiving data;

[0129] S5. According to the classification result of the archiving data, combined with the storage hierarchy model of the shutdown system and the real-time system state, optimize the path and allocate storage resources for the archiving data in the classification result of the archiving data, and generate the archiving task plan;

[0130] S6. When a failure occurs or the shutdown system is about to shut down, the rescue archiving operation needs to complete the priority screening, path allocation, and storage writing of data within an extremely short time. According to the archiving task plan, first start the archiving task execution module. The module stores the key data in batches and orderly according to the preset storage priority and the optimized storage path. For data sets with different priorities, a hierarchical archiving strategy is adopted to ensure that the ultra-high priority data is written on the fastest storage medium, while the lower priority data adopts a delayed archiving or remote storage strategy to balance the use efficiency of storage resources. During the archiving process, a real-time monitoring mechanism is adopted to monitor the data writing rate, storage space occupancy, and the execution progress of the archiving task, and identify storage failures, writing delays, or data integrity problems based on the anomaly detection mechanism, and adjust the storage strategy in real time.

[0131] In this embodiment, S1 includes the following steps:

[0132] S11. Collect the original operation data set in the shutdown system. The original operation data set includes different types of data samples, and each data sample contains multiple data attributes, forming the original operation data set :

[0133] ;

[0134] Among them, represents the th data sample, represents the value of the th data sample on the mth data attribute, is the total number of data samples, is the total number of data attributes;

[0135] S12. Perform format conversion on the original operation data set to uniformly convert different format data samples into a standardized representation form, forming a formatted operation data set;

[0136] S13. Deduplicate the formatted operation dataset, eliminate data samples with exactly the same content, and form a deduplicated operation dataset;

[0137] S14. Perform anomaly detection on the deduplicated operation dataset, identify and eliminate data outliers, and form an operation dataset after anomaly detection;

[0138] S15. Perform data integrity verification on the operation dataset after anomaly detection, eliminate incomplete data, and form a preprocessed shutdown system operation dataset :

[0139] ;

[0140] Among them, is a data integrity verification function, is the operation dataset after anomaly detection. When holds, meets the integrity requirements; otherwise, it is eliminated.

[0141] In this embodiment, through the data collection and preprocessing method, the efficient cleaning and standardization of the original operation data of the shutdown system are realized, ensuring the accuracy and integrity of the archived data. The step-by-step data format conversion, deduplication, anomaly detection, and integrity verification strategies enable the data to reach a structured, high-quality, and standardized state before entering the archived optimization process. In addition, by constructing a data sample set, the data from different sources is uniformly managed, avoiding data loss and duplicate storage problems, and improving the efficiency of data collation. The integrity and consistency of the data can be effectively guaranteed after the shutdown system is restored.

[0142] In this embodiment, S2 includes the following steps:

[0143] S21. Calculate the time sensitivity of the preprocessed shutdown system operation dataset to analyze the time correlation of data samples and define the time sensitivity parameter . The time sensitivity is used to measure the real-time requirement of the shutdown system operation data. The closer the time sensitivity value is to 1, the more urgently the data needs to be archived:

[0144] ;

[0145] Among them, represents the last update time of the data sample, represents the current system time, is a parameter for adjusting the change range of the time sensitivity;

[0146] S22. For the preprocessed shutdown system operation dataset Perform access frequency calculation, count the access times of each data sample, and define the access frequency parameter :

[0147] ;

[0148] Among them, represents the access times of the th data sample within the specified time window, represents the access times of the th data sample within the specified time window, is the access frequency adjustment coefficient, which is used to enhance the recognition ability of high-volume data;

[0149] S23. Perform storage requirement calculation on the preprocessed shutdown system operation dataset Analyze the storage capacity requirements of the data and define the storage requirement parameter :

[0150] ;

[0151] Among them, represents the storage size of the th data sample, is the storage requirement adjustment index, represents the weighted storage size of the jth sample in the storage requirement calculation, which is used to adjust the influence degree of the data sample on the storage requirement calculation, enhances the influence of large data samples when > 1, balances the influence of different data samples when < 1, represents the weighted storage size of the ith sample in the storage requirement calculation;

[0152] S24. Combine the time sensitivity parameter , the access frequency parameter and the storage requirement parameter to calculate the comprehensive weight , which is used to measure the overall importance of each data sample. The larger the comprehensive weight, the higher the priority for archiving the data:

[0153] ;

[0154] Among them, , , are the weight coefficients;

[0155] S25. Sort the data samples in the shutdown system operation dataset according to the calculated comprehensive weight to generate the data feature dataset :

[0156] ;

[0157] Among them, is the archiving threshold, which is used to screen out data samples that meet the archiving conditions.

[0158] In this embodiment, an automated priority evaluation of the shutdown system data is achieved through the data feature extraction and evaluation method, improving the accuracy and rationality of data archiving. Compared with the traditional fixed-priority archiving strategy, the multi-dimensional features of time sensitivity, access frequency, and storage requirements are used for data importance analysis, enabling data screening to no longer rely on static rules but dynamically adjust the archiving strategy according to the actual business value of the data. The time sensitivity calculation method can effectively identify the data that needs to be urgently archived in the case of an unexpected shutdown, improving the priority management ability of rescue archiving. By calculating the comprehensive weight, it is ensured that critical data can be preferentially archived in a resource-constrained environment, avoiding the unreasonable archiving problems caused by over-reliance on data types or file structures in the traditional method, thus significantly improving the adaptability and intelligence level of the archiving strategy.

[0159] In this embodiment, S3 includes the following steps:

[0160] S31. Construct a fireworks optimization search space based on a multi-level explosion strategy, and construct a multi-level explosion optimization model for the data feature dataset Initialize the fireworks individual population of the multi-level explosion optimization model , and each fireworks individual in the fireworks individual population represents a candidate rescue archiving data subset:

[0161] ;

[0162] Among them, represents the candidate rescue archiving data subset of the th fireworks individual, is the initial number of the fireworks individual population;

[0163] Introduce a multi-level explosion radius adjustment mechanism in the multi-level explosion optimization model, so that each fireworks individual adjusts the explosion radius according to the comprehensive weight :

[0164] ;

[0165] Among them, is the global maximum search radius, is the explosion range adjustment parameter, represents the comprehensive weight of the data sample , Represents the comprehensive weight of data samples ;

[0166] S32. Optimize the generation of candidate rescue archival data subsets by adopting an adaptive perturbation mechanism. During the search process, perform adaptive perturbation on data with different priorities, and define the adjustment rules for the optimized candidate rescue archival data subsets:

[0167] ;

[0168] Among them, is the optimized candidate rescue archival data subset of the th firework individual, and are the addition and removal thresholds respectively, is the probability that the data sample is added to the candidate rescue archival data subset:

[0169] ;

[0170] is the probability that the data sample is removed from the candidate rescue archival data subset:

[0171] ;

[0172] Among them, is the probability that the data sample is added to the candidate rescue archival data subset;

[0173] S33. Introduce a hierarchical co-evolution strategy during the optimization process. Divide the firework individuals into a critical data priority layer, a balanced archival layer, and a storage-constrained layer, and adopt corresponding search strategies respectively. The critical data priority layer adopts a fast convergence search, the balanced archival layer adopts a multi-objective balance strategy, and the storage-constrained layer adopts storage optimization adjustment. The data optimization objectives of the critical data priority layer, the balanced archival layer, and the storage-constrained layer are defined as follows:

[0174] ;

[0175] ;

[0176] ;

[0177] Among them, is the objective function of the critical data priority layer, is the objective function of the balanced archival layer, is the objective function of the storage-constrained layer, represents the The optimized candidate rescue archival data subset for each firework individual For storing the weight adjustment parameter Is the mean weight of the data samples Is the available storage resource;

[0178] S34. Generate the final optimized data set by combining multi-objective optimization strategies, and perform a final screening to select the rescue archival data set that meets the global optimal solution :

[0179] ;

[0180] Among them, 、 、 Are the weight parameters of the archival optimization objectives Is the final population set, so that the rescue archival data set achieves global optimality among the key data priorities, archival balance, and storage utilization rate

[0181] In this embodiment, through the improved hybrid firework optimization algorithm, the search range is adaptively adjusted through a multi-stage explosion strategy, so that data screening is no longer limited to static weights, but can evaluate the business importance of data in real time and dynamically optimize the data screening strategy. In addition, combining reinforcement learning to optimize the search path enables the screening process to gradually approach the optimal data subset, ensuring that the most important data can be stored first. At the same time, a hierarchical co-evolution strategy is adopted, so that data screening not only considers storage resource constraints, but also can balance data priorities and storage space utilization, thereby improving the accuracy of data screening and the scientific nature of archival decision-making

[0182] In this embodiment, S4 includes the following steps:

[0183] S41. Construct a clustering optimization model based on double-layer dynamic entropy weight for the rescue archival data set that meets the global optimal solution Construct an adaptive entropy weight calculation framework and initialize the data clustering center set of the clustering optimization model , Each clustering center in the data clustering center set Represents an archival data category:

[0184] ;

[0185] Among them, Indicates the th clustering center, Is the total number of initial archival data categories;

[0186] Use the double-layer dynamic entropy weight calculation method to calculate the global entropy weight And the local entropy weight , comprehensively evaluate the importance of data features, and the global entropy weight is calculated as follows:

[0187] ;

[0188] Among them, represents the normalized probability value of the th data sample on the corresponding feature;

[0189] The local entropy weight is calculated using the contribution degree of the data sample to the clustering center, and is defined as follows:

[0190] ;

[0191] Among them, represents the value of the clustering center on the feature , represents the value of the clustering center on the feature , is the mean value of all clustering centers on the feature ;

[0192] S42. Use the variable weight fuzzy clustering method to calculate the membership degree of data samples, define the adaptive weighted fuzzy distance , and calculate the membership degree of the data sample to the clustering center :

[0193] ;

[0194] Among them, the adaptive weighted fuzzy distance is defined as follows:

[0195] ;

[0196] Among them, is the variable weight index, is the total number of features of the data sample;

[0197] S43. Based on the dynamic update strategy of clustering drift detection, use the clustering center movement trend vector to optimize the clustering update process, and introduce the drift metric to judge whether the clustering center needs to be adjusted:

[0198] ;

[0199] Among them, represents the number of data samples belonging to the clustering center . When , is the threshold value, and the cluster center needs to be adjusted. The adjustment strategy is as follows:

[0200] ;

[0201] Among them, is the dynamic adjustment step size, is the position of the new cluster center, is the data sample The membership degree of the cluster center ;

[0202] S44. Optimize the clustering termination condition by using the fuzzy entropy convergence criterion, and calculate the fuzzy entropy of the clustering system to measure the uncertainty of clustering:

[0203] ;

[0204] Set the dynamic entropy convergence threshold , when and the change of the cluster center satisfies , terminate the clustering, otherwise continue to execute S42 - S43 to update the cluster center;

[0205] S45. According to the optimized cluster center and the membership degree for the data sample allocate the archival category, and generate the archival data classification result:

[0206] ;

[0207] Among them, represents the finally generated archival data classification result, is the archival category to which the data sample belongs, enabling the rescue archival data set to be effectively classified, is to obtain the variable value that makes reach the maximum value.

[0208] In this embodiment, through the double - layer dynamic entropy - weight clustering method, the intelligence and self - adaptability of data classification are realized, the accuracy and stability of archival data classification are improved. By introducing the double - layer calculation framework of global entropy weight and local entropy weight, the data clustering can automatically adjust the feature weights according to the change of data distribution, thereby improving the flexibility of data classification. In addition, by adopting the variable - weight fuzzy clustering method, the self - adaptability of the clustering process is further enhanced, enabling different types of data to be accurately classified according to the real - time state. The clustering drift detection mechanism effectively avoids the problem of misclassification of data categories, improves the stability and consistency of data archival classification, and thus ensures that key data can be archived to the most appropriate storage level.

[0209] In this embodiment, S5 includes the following steps:

[0210] S51. Construct a path optimization framework based on the storage hierarchy model of the shutdown system, optimize the archival path for the classification result of the archived data and initialize the storage hierarchy structure of the shutdown system , where each storage hierarchy represents an archival path of a different storage medium:

[0211] ;

[0212] Among them, represents the th storage hierarchy, is the total number of storage hierarchies, represents the set of available storage resources of the system;

[0213] Introduce a storage hierarchy selection function based on the data urgency :

[0214] ;

[0215] Among them, is the urgency of the data sample , is the read / write latency of the storage hierarchy , is the unit storage cost of the storage hierarchy , so that data with an urgency higher than the preset value is preferentially stored in a high-speed storage medium, and data with an urgency lower than the preset value is stored in a low-cost storage medium;

[0216] S52. Calculate the optimal storage allocation scheme for the data archival path, and calculate its optimal storage location for the archived data sample based on the storage hierarchy model :

[0217] ;

[0218] Among them, is the comprehensive weight, is the storage size of the data sample, is the available storage space of the storage hierarchy , is the historical access frequency of the storage hierarchy , is the optimization weight coefficient, so that the storage allocation scheme takes into account data priority, storage space utilization and data access requirements;

[0219] S53. Generate the final archiving task plan based on the optimal storage location :

[0220] ;

[0221] wherein, is the archiving timestamp of the data sample , and is the archiving trigger threshold, enabling the rescue archiving data set to be stored according to the optimal path.

[0222] In this embodiment, through the storage path optimization and resource allocation strategy, the rationality of data storage and the execution efficiency of the archiving task are improved. By introducing a storage layer selection method based on data urgency, the data storage path can be dynamically optimized according to business requirements, ensuring that critical data is preferentially allocated to high-speed storage media, while low-priority data is stored in low-cost storage media.

[0223] Example 1:

[0224] Example 1 occurred at 3:27 am on June 15, 2024. A sudden shutdown event occurred in the data center of a financial institution. When the accident occurred, the core storage cluster "DC-Storage-003" in the data center was accidentally powered off due to a power supply line failure, resulting in IO errors on some storage nodes. Approximately 14.6TB of transaction data, account change records, and system logs were not archived normally. The last log timestamp of the database server "DB-Node-05" before the failure was 2024-06-15 03:26:58, indicating that most of the transaction data had not been written within 1 minute before the system crashed.

[0225] 03:27:10 - The accident detection module found that the number of abnormal IO failure alarms generated by the storage node "DC-Storage-003" in the past 15 seconds reached 37 times, far exceeding the threshold (20 times). The system determined that "DC-Storage-003" might have a sudden shutdown, triggering the data rescue archiving strategy.

[0226] 03:27:18 - The data collection module began to scan the active database transaction logs within 30 minutes before the failure (a total of 32,658 uncompleted transactions), which involved 1,258,402 user transaction records with a total data volume of approximately 1.2TB. The system also detected that the transaction log file "trx_log_20240615.log" on "DB-Node-05" was accidentally truncated, and the file size was only 872MB, which was 62% less than the expected size (2.3GB). The system determined that some transaction log data was lost.

[0227] 03:27:35 - Format conversion and data cleaning started. After parsing the exception transaction logs, the system automatically filters out low-priority log entries that are not related to storage errors, such as query transactions without account changes (a total of 7,532 entries). Finally, 27,894 high-priority rescue transaction records are obtained, with a total data volume of 982 GB.

[0228] 03:28:12 - The hybrid fireworks optimization algorithm started running to screen and optimize the data to be archived.

[0229] The system calculates the time sensitivity of the transaction logs, with the highest value being 0.98 (i.e., data not archived within T+0 minutes is extremely important).

[0230] The access frequency analysis results show that the top 10% of accounts are involved in a total of 28,602 transactions, and these transaction data are automatically marked as high-priority.

[0231] The calculation result of the storage requirement of the log files shows that the high-priority transaction data accounts for 73.4% of the total data volume.

[0232] 03:28:28 - The hybrid fireworks optimization algorithm completed the calculation, generating a rescue archiving data set, screening out 750 GB of transaction data as the highest-priority archiving target, with the account ID range being: [A0456723 - A0973456], and the transaction order ID range being: [TXN-09873423 - TXN-09895687].

[0233] 03:28:50 - The dynamic entropy weight clustering method was started to automatically classify the 750 GB of screened data.

[0234] The transaction order data was clustered into the "ultra-high-priority group" (a total of 412 GB), which contains a high-frequency trading user group (217,634 transaction records).

[0235] The account change data was clustered into the "high-priority group" (a total of 184 GB), and the total number of accounts involved in financial data changes was 5,486.

[0236] The log data was clustered into the "medium-priority group" (a total of 154 GB), mainly for system recovery analysis.

[0237] 03:29:10 - The storage optimization module was started to optimize the archiving path of the data.

[0238] The data in the "ultra-high-priority group" was allocated to the NVMe SSD storage node (DC-Storage-001), with an estimated write speed of about 2.3 GB / s, and the completion time is expected to be within 5 minutes and 45 seconds.

[0239] The "high-priority group" data is allocated to the HDD array (DC-Storage-002), with an estimated write speed of approximately 700 MB / s and an expected completion time within 4 minutes and 30 seconds.

[0240] The "medium-priority group" data is temporarily stored on the remote backup server (DC-Backup-004) and is planned to be synchronized to long-term storage after system recovery.

[0241] 03:29:45 - The data archiving task is officially started. The write progress is as follows:

[0242] 03:30:00 - The NVMe SSD storage has completed archiving 85 GB of data, with a progress of approximately 21%.

[0243] 03:31:30 - The HDD storage has completed archiving 94 GB of data, with a progress of approximately 51%.

[0244] 03:34:10 - All data archiving is completed, the log "archive_log_20240615.log" is generated, confirming that 750 GB of data has been successfully archived, and the archiving accuracy rate reaches 99.7%.

[0245] Within 30 minutes after system recovery, all high-priority transaction data has been restored, the system is normally started, and the order settlement reconciliation is completed.

[0246] The total amount of rescue archiving data this time is 750 GB. Compared with the traditional regular backup method, the method of the present invention performs excellently in archiving efficiency, data recovery time, and storage optimization. The specific improvement points are as follows:

[0247] 1. The archiving rate of key data is increased by 26.5%, ensuring the priority storage of important transaction data and reducing the risk of data loss;

[0248] 2. The archiving execution time is shortened by 17 minutes (about 40.5%), greatly improving the data recovery speed in case of sudden shutdown events;

[0249] 3. The storage space utilization rate is increased by 17.1%, effectively reducing the resource waste of low-priority data occupying high-performance storage devices;

[0250] 4. The key data recovery time is shortened to 7 minutes (a reduction of 61%), ensuring business continuity and reducing the impact of shutdown accidents on operations;

[0251] 5. The archiving success rate is close to 100%, and the business data can be automatically restored without additional manual intervention after system recovery.

[0252] In this embodiment, through the sudden shutdown event of a real financial system, the efficiency and accuracy of the present invention in rescue archiving are verified. By adopting a hybrid fireworks optimization, dynamic entropy weight clustering, and storage path optimization algorithm, the limitations of traditional data archiving methods are broken through, intelligent archiving and rapid recovery of key data in the case of sudden shutdown are achieved, and the business continuity and data protection capabilities of the shutdown system are effectively improved.

[0253] The present invention adopts a multi-level explosion radius adaptive adjustment mechanism through an improved hybrid fireworks optimization algorithm, enabling data screening to no longer rely on fixed rules but adaptively adjust the search scope according to the key features of data such as time sensitivity, access frequency, and storage requirements, thereby improving the recognition accuracy of key data. The adjustment strategy of the multi-level explosion radius enables high-priority data to obtain a greater search weight during the global search process, ensuring that key data can be preferentially selected into the archived data set.

[0254] In the process of data classification, the present invention adopts a two-layer dynamic entropy weight clustering method to calculate the global entropy weight and local entropy weight of data features respectively, enabling data classification to dynamically adjust the influence weights of different features according to the system state. The two-layer entropy weight calculation framework enhances the stability and accuracy of data classification through an adaptive weighted fuzzy distance calculation method.

[0255] In terms of optimizing the data storage path, the present invention proposes a storage level selection function based on data urgency and combines reinforcement learning to optimize the storage path allocation strategy, enabling data archiving to perform dynamic storage allocation according to the real-time system state. By continuously adjusting the storage path selection through the reinforcement learning method, high-priority data can be allocated to high-speed storage media with lower storage latency, and the data migration plan is dynamically adjusted to ensure the efficient use of storage space.

[0256] As mentioned above, the above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.

Claims

1. A data intelligent archiving management system for shutdown systems, characterized in that, It includes the following modules: The data acquisition and preprocessing module is used to acquire the original operation data set in the shutdown system and perform preprocessing to form the shutdown system operation data set; The data feature extraction and evaluation module is used to perform multi-dimensional feature extraction on the shutdown system operation data set, calculate the time sensitivity, access frequency and storage requirements of the data, and calculate the importance weight of the data based on dynamic entropy weight to form the data feature data set; The data screening module based on hybrid fireworks optimization is used to perform global search and parameter optimization on the data feature data set, and select the optimal data subset based on the hybrid fireworks optimization algorithm combined with the shutdown system state, data priority and storage resource constraints to form the optimized rescue archiving data set; The dynamic entropy weight clustering analysis module is used to perform dynamic entropy weight clustering analysis on the optimized rescue archiving data set to generate the archiving data classification result; The archiving path optimization and storage resource allocation module is used to combine the archiving data classification result, the shutdown system storage hierarchy model and the real-time system state to perform path optimization and storage resource allocation on the archiving data in the archiving data classification result to generate the archiving task plan; The archiving task execution module is used to perform the rescue archiving operation according to the archiving task plan, store the shutdown system operation data set along the optimal path, and record the archiving status; The archiving task execution module specifically includes constructing a path optimization framework based on the storage hierarchy model of the shutdown system, and classifying the archiving data results to optimize the archiving path and initialize the storage hierarchy structure of the shutdown system , where each storage hierarchy represents an archiving path for a different storage medium: ; Among them, represents the th storage level, is the total number of storage levels, represents the set of available storage resources of the system; Introduce a storage level selection function based on data urgency : ; wherein, is the urgency of the data sample , is the read / write latency of the storage hierarchy , is the unit storage cost of the storage hierarchy , so that data with an urgency higher than the preset value is preferentially stored in a high-speed storage medium, and data with an urgency lower than the preset value is stored in a low-cost storage medium; Calculate the optimal storage allocation scheme for the data archiving path, and calculate its optimal storage location based on the storage hierarchy model for the archived data samples Calculate its optimal storage location : ; Among them, is the comprehensive weight, is the storage size of the data sample, is the storage hierarchy of the available storage space, is the storage hierarchy of the historical access frequency, is the optimization weight coefficient, making the storage allocation scheme take into account data priority, storage space utilization rate and data access requirements; Generate the final archiving task plan based on the optimal storage location : ; Among them, is the archival timestamp of the data sample, is the archival trigger threshold, so that the rescue archival data set is stored along the optimal path.

2. A data intelligent archiving management method for a shutdown system, applied to the data intelligent archiving management system for a shutdown system described in claim 1, characterized in that It includes: S1. Acquire the original operation data set in the shutdown system, and perform preprocessing operations on the original operation data set to form the preprocessed shutdown system operation data set; S2. Perform multi-dimensional feature extraction on the preprocessed shutdown system operation data set, calculate the time sensitivity, access frequency and storage requirements of each data sample in the shutdown system operation data set, and generate the data feature data set; S3. Perform global search and parameter optimization on the data feature data set based on the hybrid fireworks optimization algorithm to form the rescue archiving data set that meets the global optimal solution; S4. Use the dynamic entropy weight clustering method to perform clustering processing on the rescue archiving data set that meets the global optimal solution, classify the data according to the adaptively adjusted weights of each data feature, and generate the archiving data classification result; S5. According to the archiving data classification result, combine the shutdown system storage hierarchy model and the real-time system state to perform path optimization and storage resource allocation on the archiving data in the archiving data classification result to generate the archiving task plan; S6. Implement the rescue archiving operation on the optimized rescue archiving data set in the actual environment of the shutdown system according to the archiving task plan, monitor the archiving process in real time and record the archiving status to form the archiving execution record.

3. The data intelligent archiving management method for a shutdown system according to claim 2, characterized in that, The said S1 includes the following steps: S11. Collect the original operation data set in the shutdown system ; S12. For the original operation data set perform format conversion to form a formatted operation data set; S13. Perform data deduplication on the formatted operation data set to form the deduplicated operation data set; S14. Perform anomaly detection on the deduplicated operation data set to form the anomaly-detected operation data set; S15. Perform data integrity verification on the running data set after anomaly detection to form a preprocessed shutdown system running data set .

4. The data intelligent archiving management method for a shutdown system according to claim 3, characterized in that, The said S2 includes the following steps: S21. Calculate the time sensitivity of the preprocessed shutdown system operation dataset Analyze the time correlation of data samples and define the time sensitivity parameter ; S22. Calculate the access frequency of the preprocessed shutdown system operation dataset to count the access times of each data sample and define the access frequency parameter ; S23. Calculate the storage requirements for the preprocessed shutdown system operation dataset Analyze the storage capacity requirements of the data and define the storage requirement parameters ; S24. Combine the time sensitivity parameter , access frequency parameter and storage requirement parameter to calculate the comprehensive weight for measuring the overall importance of each data sample. The larger the comprehensive weight is, the higher priority should be given to archiving this data; S25. According to the calculated comprehensive weight Sort the data samples in the shutdown system operation dataset to generate a data feature dataset .

5. The data intelligent archiving management method for a shutdown system according to claim 4, characterized in that The said S3 includes the following steps: S31. Construct a fireworks optimization search space based on a multi-level explosion strategy for the data feature dataset Construct a multi-level explosion optimization model and initialize the fireworks individual population of the multi-level explosion optimization model , and each fireworks individual in the fireworks individual population represents a candidate rescue archival data subset; Introduce a multi-level explosion radius adjustment mechanism into the multi-level explosion optimization model, so that each firework individual adjusts the explosion radius according to the comprehensive weight Adjust the explosion radius : ; Among them, is the global maximum search radius, is the explosion range adjustment parameter, represents the data sample 's comprehensive weight, represents the data sample 's comprehensive weight; S32. Optimize the generation of the candidate rescue archiving data subset by using the adaptive perturbation mechanism, perform adaptive perturbation on the data with different priorities during the search process, and define the adjustment rule of the optimized candidate rescue archiving data subset: ; Among them, is the optimized candidate rescue archival data subset for the th individual firework, and are the addition and removal thresholds respectively, is the probability that the data sample is added to the candidate rescue archival data subset, is the probability that the data sample is removed from the candidate rescue archival data subset; S33. Introduce a hierarchical co-evolution strategy during the optimization process. Divide the fireworks individuals into a critical data priority layer, a balanced archiving layer, and a storage-constrained layer, and adopt corresponding search strategies respectively. The critical data priority layer adopts a fast convergence search, the balanced archiving layer adopts a multi-objective balance strategy, and the storage-constrained layer adopts a storage optimization adjustment. The data optimization objectives of the critical data priority layer, the balanced archiving layer, and the storage-constrained layer are defined as follows: ; ; ; Among them, is the objective function of the critical data priority layer, is the objective function of the balanced archiving layer, is the objective function of the storage-constrained layer, represents the th candidate rescue archiving data subset optimized by the firework individual, is the storage weight adjustment parameter, is the mean value of the data sample weights, is the available storage resource; S34. Generate the final optimized data set by combining the multi-objective optimization strategy, and conduct the final screening to select the rescue archival data set that meets the globally optimal solution. 。 6. The data intelligent archiving management method for a shutdown system according to claim 5, characterized in that The said S4 includes the following steps: S41. Construct a clustering optimization model based on double-layer dynamic entropy weight for the set of rescue filing data that meets the global optimal solution Construct an adaptive entropy weight calculation framework and initialize the data clustering center set of the clustering optimization model , and each clustering center in the data clustering center set represents an archiving data category; The double-layer dynamic entropy weight calculation method is used to calculate the global entropy weight and the local entropy weight respectively, and the importance of data features is comprehensively evaluated. The contribution degree of data samples to the relative clustering center is used for the calculation of local entropy weight; S42. Calculate the membership degree of data samples using the variable weight fuzzy clustering method, and define the adaptive weighted fuzzy distance , calculate the data samples For the clustering center Membership degree ; S43. Dynamic update strategy based on clustering drift detection, using the clustering center movement trend vector to optimize the clustering update process and introducing drift metrics Determine whether the clustering center needs to be adjusted When At that time is the threshold, and the clustering center needs to be adjusted; S44. Optimize the clustering termination condition using the fuzzy entropy convergence criterion, and calculate the fuzzy entropy of the clustering system Measure the uncertainty of clustering , and set the dynamic entropy convergence threshold , when and the change of the clustering center satisfies , terminate the clustering; otherwise, continue to execute S42 - S43 to update the clustering center; S45. Assign an archiving category to the data samples according to the optimized cluster centers and membership degrees to generate an archiving data classification result . .

Citation Information

Patent Citations

  • Data retention management

    CN103631849A

  • Synthesizing data for training one or more neural networks

    US20210142177A1