A method for rescue archiving and disposal of important data in a shutdown system

Through machine learning models, data structures, storage and conversion, zero-trust architecture encrypted storage and streaming search mechanisms, the problems of format incompatibility, insufficient security and low query efficiency in shutting down system data archives are solved, and efficient and secure data storage and query are achieved.

CN120011311BActive Publication Date: 2025-07-18HANGZHOU YIKANGXIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510475005.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

There are currently problems with incompatible data formats, insufficient security, low query efficiency, and low archive automation in relation to the data archiving method of shutting down the system.

Method used

The machine learning model is used to automatically analyze the data structure, combine the storage and conversion mechanism for format conversion, encrypted storage and dynamic key management based on zero-trust architecture, combined with distributed key storage technology, and use the streaming retrieval mechanism and virtual database interface for data query and access.

Benefits of technology

It realizes data format compatibility between different systems, improves the security of data storage and query efficiency, ensures the security of data during storage, transmission and access, and supports efficient query and cross-data sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011311B_ABST
    Figure CN120011311B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data storage and management, and discloses a method for rescue archiving and disposal of important data in a shutdown system, including the following steps: S101, automatically parsing the data structure of stored data by using a machine learning model and performing scanning and extraction; S102, storing and converting the scanned and extracted data by using a mechanism of storing while converting; S103, protecting the archived data by using an encryption storage mechanism based on a zero-trust architecture, and adopting a dynamic key management strategy to dynamically allocate, recycle, and destroy keys according to access scenarios, and combining distributed key storage technology to protect the data; S104, providing query and access to the archived data by using a streaming retrieval mechanism, and remotely calling and virtually accessing the archived data based on a virtual database interface. The present invention improves the format compatibility of the archived data and enhances the query efficiency and cross-data source access ability of the archived data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data storage and management, and particularly relates to a method for rescue archiving and disposal of important data in a shutdown system. Background Art

[0002] With the rapid development of information technology, the information systems of enterprises, governments, and research institutions are constantly upgraded and iterated. Some old systems are facing shutdown due to reasons such as aging of software and hardware, backward technical architecture, and high operation and maintenance costs. However, these systems often store key data such as important business data, user information, financial records, and historical logs. Direct deletion or improper migration will result in data loss, compliance risks, and business continuity problems.

[0003] The existing data archiving methods for shutdown systems mainly rely on traditional methods such as database backup, log storage, and manual data migration, but there are the following technical limitations: incompatible data formats, insufficient storage security, low query and access efficiency, and low degree of archiving automation. First, the data formats, database structures, and storage methods of different systems are diverse, and simple data export is difficult to meet the cross-system compatibility requirements. Second, traditional data storage methods lack dynamic encryption and key management, and data is vulnerable to tampering or leakage during the archiving and storage process. Third, existing archiving solutions mostly adopt full recovery queries, resulting in data access delays and unable to meet the high-efficiency query requirements. Summary of the Invention

[0004] The present invention provides a method for rescue archiving and disposal of important data in a shutdown system, which solves the technical problems of incompatible data formats, insufficient security, low query efficiency, and low level of archiving automation in the related art.

[0005] The present invention provides a method for rescue archiving and disposal of important data in a shutdown system, including the following steps:

[0006] S101, before archiving the data of the shutdown system, use a machine learning model to automatically analyze the data structure of the stored data and perform scanning and extraction;

[0007] S102, adopt a mechanism of storing while converting to store and convert the scanned and extracted data, and based on content feature analysis, automatically match the corresponding archiving format for archiving;

[0008] S103, adopt an encryption storage mechanism based on the zero-trust architecture to protect the archived data, and adopt a dynamic key management strategy to dynamically allocate, recycle, and destroy keys according to the access scenario, and combine distributed key storage technology to protect the data;

[0009] S104. Provide query and access to the archived data using a streaming retrieval mechanism, and perform remote calls and virtualized access to the archived data based on a virtual database interface.

[0010] Furthermore, for the stored data, use a machine learning model to automatically parse the data structure and perform scanning and extraction. The specific steps include:

[0011] S201. Use an automatic storage type recognition algorithm to identify the data storage type of the shutdown system;

[0012] S202. Use a machine learning model to parse the internal structure of the data stored in the shutdown system;

[0013] S203. Extract time series features, text pattern features, and numerical distribution features from the parsed data structure;

[0014] S204. Scan and extract the stored data according to the extracted features.

[0015] Furthermore, the specific steps of S102 include:

[0016] S301. Convert the storage format of the scanned and extracted data;

[0017] S302. Generate metadata based on the data after format conversion, and store the association relationship between the data;

[0018] S303. Adjust the storage method of the data after format conversion;

[0019] S304. Manage the storage structure of the data after format conversion.

[0020] Furthermore, use a symmetric encryption algorithm to encrypt the archived data; during the data storage and access process, adopt an on-demand allocation strategy, dynamically allocate a unique temporary key for each access, and perform real-time auditing of the access behavior;

[0021] Adopt a key automatic expiration strategy, where the key is only valid within a preset time window and is automatically recycled when it times out;

[0022] Adopt an irrecoverable destruction strategy. After the end of the data life cycle, use a multiple erasure algorithm to destroy the key, and record the key destruction log through a blockchain evidence storage mechanism.

[0023] Furthermore, store the key in a distributed key storage manner. The distributed key storage specifically includes: multi-node key sharding storage and geographical distributed storage.

[0024] Furthermore, use a first security policy to access the archived data, where the first security policy includes: fine-grained access control and zero-knowledge access.

[0025] Furthermore, the streaming retrieval mechanism uses a distributed stream processing framework to perform real-time parsing on the archived data. The specific steps of the streaming retrieval mechanism are as follows:

[0026] S401, Based on streaming processing technology, extract data as needed and use block loading to limit the data reading volume;

[0027] S402, Optimize the query of archived data using vector indexing technology;

[0028] S403, Parse the query request through a streaming SQL engine.

[0029] Furthermore, when it is detected that the system reaches a preset trigger condition, the data archiving operation is automatically executed. The preset trigger conditions include:

[0030] The disk space is lower than the first preset threshold;

[0031] The CPU load is higher than the second preset threshold and the duration exceeds the first preset time interval;

[0032] The system detects an abnormal database connection.

[0033] Furthermore, optimize the query of archived data using vector indexing technology. The specific steps are as follows:

[0034] S501, Extract the text content, metadata, and statistical information for each data item and convert them into feature vectors;

[0035] S502, Set multiple levels. Each level uses a different hash function family. Allocate data items to hash buckets according to the hash values and recursively form a tree-like index structure. Each node in the structure represents a hash bucket, and the leaf nodes contain pointers to the original data items;

[0036] S503, After standardizing the query data, traverse each layer of hash buckets in turn using the same hash function as when constructing the index, and use the stored representative vectors to preliminarily screen the candidate data. Finally, calculate the cosine similarity of the candidate data items at the leaf nodes and return the query results sorted by similarity;

[0037] S504, Automatically update the hash index when new data is added.

[0038] The beneficial effects of the present invention are as follows: The present invention uses a machine learning model to automatically analyze the data structure and combines an automatic storage type recognition algorithm to accurately identify different data storage types such as relational databases, non-relational databases, log files, and image files. Based on the mechanism of storing while converting, format conversion is synchronously completed during the archival storage process. This method can automatically adapt to a variety of archival storage formats, ensure data compatibility between different systems, reduce manual intervention in data format conversion, and improve the archival efficiency and storage standardization level;

[0039] The present invention adopts an encryption storage mechanism with a zero-trust architecture to ensure that data is in an encrypted protection state throughout the processes of storage, transmission, and access. Combining with a dynamic key management strategy, keys can be dynamically allocated, recycled, and destroyed according to access scenarios to prevent long-term exposure of keys. At the same time, distributed key storage technology is adopted to ensure the anti-attack ability and recoverability of keys, prevent single-point key leakage, and improve the long-term security of data after archiving;

[0040] The present invention adopts a streaming retrieval mechanism and combines a distributed stream processing framework to support extracting archival data on demand and improve query efficiency; Based on a virtual database interface, archival data can be directly jointly queried across data sources without physically restoring the data, improving the flexibility and real-time performance of data retrieval. Brief Description of the Drawings

[0041] Figure 1 is a flowchart of a method for rescue archival disposal of important data in a shutdown system of the present invention. Detailed Embodiments

[0042] Now, the subject matter described herein will be discussed with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described relative to some examples can also be combined in other examples.

[0043] As Figure 1 shown, a method for rescue archival disposal of important data in a shutdown system includes:

[0044] S101, before archiving the data of the shutdown system, use a machine learning model to automatically analyze the data structure of the stored data and perform scanning and extraction;

[0045] S102, use the mechanism of storing while converting to store and convert the scanned and extracted data, and based on content feature analysis, automatically match the corresponding archival format for archiving;

[0046] S103. Use an encryption storage mechanism based on the zero-trust architecture to protect the archived data, and adopt a dynamic key management strategy to dynamically allocate, recycle, and destroy keys according to the access scenario, and combine distributed key storage technology to protect the data;

[0047] S104. Provide query and access to the archived data using a streaming retrieval mechanism, and perform remote calls and virtualized access to the archived data based on a virtual database interface.

[0048] In an embodiment of the present invention, a machine learning model is used to automatically parse the data structure of the stored data and perform scanning and extraction. The specific steps include:

[0049] S201. Use an automatic storage type recognition algorithm to identify the data storage type of the shutdown system;

[0050] Specifically, the data storage types include: relational databases, non-relational databases, log files, and image files. The specific steps of the automatic storage type recognition algorithm include:

[0051] S2011. Detect the database meta-information through the database connection driver to identify the database type;

[0052] S2012. Detect and identify log files and image files through file header signatures;

[0053] S2013. Use regular expression matching to analyze the log format and automatically adapt to the log structure. For example, the format is timestamp + event + parameter;

[0054] S202. Use a machine learning model to parse the internal structure of the data stored in the shutdown system;

[0055] Specifically, use a clustering algorithm to analyze the data distribution in the database and automatically group similar data structures. For example, the order table, invoice table, and transaction table are grouped into the financial data category, and the user table, customer table, and employee table are grouped into the personnel information category;

[0056] Use a pattern matching algorithm to perform standardized mapping on the data patterns of different systems to ensure data compatibility of heterogeneous databases. Specifically, use a Schema Mapping algorithm based on template matching to convert different database schemas into a unified data structure. For example: convert VARCHAR2 stored in an Oracle database and VARCHAR stored in a MySQL database into the STRING format;

[0057] Use a deep neural network or decision tree to automatically predict the data type of a field based on field value distribution, length, and format features;

[0058] Model the field correlation between database tables using a Bayesian network to automatically identify the primary and foreign key relationships, so as to preserve the integrity of the database structure;

[0059] S203. Extract time series features, text pattern features, and numerical distribution features from the parsed data structure. Specifically, identify the time fields in the stored data, parse the time format, and determine the time granularity, such as seconds, minutes, hours, and days; use term frequency-inverse document frequency to analyze the important keywords in the text fields, and use regular expressions to identify specific format texts such as addresses, names, and amounts; calculate the mean, standard deviation, and extreme values of the numerical fields to judge the value range and outliers of the data;

[0060] S204. Scan and extract the stored data according to the extracted features. Specifically, adopt a streaming processing mechanism to read the database data row by row to avoid excessive consumption of system resources caused by one-time loading; use regular expressions to parse the log content and extract the key fields; combine optical character recognition and automatic speech recognition to extract the content of images, PDF documents, and audio data to improve the usability of unstructured data.

[0061] In an embodiment of the present invention, the specific steps of S102 include:

[0062] S301. Convert the scanned and extracted data into an archived storage format to ensure data structuring and improve query and recovery capabilities;

[0063] Among them, the specific steps of S301 include:

[0064] S3011. Determine the optimal storage format based on the field type, data hierarchical structure, and distribution characteristics;

[0065] S3012. Adopt a format matching algorithm to convert different types of data. Specifically, convert relational database tables into JSON format, convert non-relational database tables into structured tables for storage, extract the key fields from the log files and store them in the index system, and combine OCR to parse the image files to generate retrievable text;

[0066] S3013. Adopt a streaming processing mechanism to synchronously execute format conversion during the data storage process.

[0067] S302. Generate metadata based on the data after format conversion and store the association relationships between the data to ensure that the archived data can still retain the business logic. Specifically, metadata represents the information attributes of the data, and metadata includes: data source, storage time, data format, field definition, and hash value; the association relationship between the data refers to the primary key-foreign key relationship existing between the data tables, and the association relationship between the data is stored through a relationship table;

[0068] S303. Adjust the storage method of the data after format conversion to adapt to different data storage media. Specifically, classify the data for hierarchical storage according to the data access frequency. Store the data with high access frequency in the cloud database and the data with low access frequency on the disk.

[0069] S304. Manage the storage structure of the data after format conversion to ensure the long-term availability and security of the archived data. Specifically, adopt a partition storage strategy, partition the archived data according to time, business category, and data source to improve the retrieval efficiency.

[0070] In an embodiment of the present invention, the zero-trust model adopts a security policy of never trusting and always verifying. Any data access requires identity authentication, device verification, and behavior analysis. It does not rely on traditional network boundary security mechanisms but is based on fine-grained access control and dynamic trust assessment to ensure that even in a trusted environment, data access still requires strict verification.

[0071] In an embodiment of the present invention, use a symmetric encryption algorithm to encrypt the archived data. During the data storage and access process, adopt an on-demand allocation strategy. Dynamically allocate a unique temporary key for each access and conduct real-time auditing of the access behavior. Based on factors such as data sensitivity, user identity, access time, and device location, perform multi-factor decision-making to determine whether to allow key allocation.

[0072] Adopt a key automatic expiration policy. The key is only valid within a preset time window and is automatically recycled when it times out to prevent the abuse of keys stored for a long time. Combine it with a revocation mechanism. When an abnormal access is detected, immediately revoke the relevant key and terminate the access permission. Abnormal access includes, but is not limited to, behaviors such as logging in from a high-risk IP address and abnormal data downloading.

[0073] Adopt an irrecoverable destruction policy. After the end of the data life cycle, use a multiple erasure algorithm to destroy the key and record the key destruction log through a blockchain evidence storage mechanism to ensure that the key management process is auditable and tamper-proof. The multiple erasure algorithm adopts the NIST800-88 standard.

[0074] In an embodiment of the present invention, store the key in a distributed key storage manner, specifically including: multi-node key sharding storage and geographical distributed storage.

[0075] Among them, multi-node key sharding storage adopts the Shamir encryption sharing algorithm to split the encryption key into multiple segments and store them in different independent security modules to ensure that a single storage node cannot recover the complete key. Only when the preset key segment threshold is met can the complete key be recovered to improve the anti-attack ability of key storage.

[0076] Geographically distributed storage adopts a cross-region key storage strategy, storing keys on secure servers at different physical locations to prevent key loss or attacks caused by a single data center failure. During key recovery, a consensus mechanism is used to ensure that keys can only be recombined in a trusted environment, preventing malicious attackers from obtaining key fragments for synthesis.

[0077] In one embodiment of the present invention, a first security policy is used to access archived data, where the first security policy includes: fine-grained access control and zero-knowledge access. Specifically, fine-grained access control is used to dynamically adjust access permissions based on user identity, access device, time, geographical location, and access behavior. Fine-grained access control includes: attribute-based access control and role-based access control. Attribute-based access control means setting access policies according to user attributes to ensure that data access complies with predefined security rules. Role-based access control means defining different permission levels according to the roles of users in the system to ensure the permission boundaries of different user groups. The permission levels include but are not limited to: administrator permissions, auditor permissions, and ordinary user permissions; zero-knowledge access adopts a zero-knowledge proof mechanism, enabling users to verify their access permissions through encrypted mathematical proofs without directly obtaining the data content.

[0078] In one embodiment of the present invention, the streaming retrieval mechanism uses the distributed stream processing framework Apache Kafka to perform real-time parsing on archived data. The specific steps of the streaming retrieval mechanism include:

[0079] S401, Based on stream processing technology, extract data as needed and use chunked loading to limit the data reading volume, reducing the data reading volume;

[0080] S402, Use vector index technology to optimize the query of archived data, improving the retrieval speed. Specifically, use B+ tree index for SQL data, Elasticsearch index for text and logs to support fuzzy query and full-text retrieval, and OCR + feature vector index for image files to support content-based retrieval;

[0081] S403, Parse the query request through the streaming SQL engine Presto to improve the query performance.

[0082] Traditional database queries usually require local data storage, but archived data may be distributed in different storage systems, such as local storage, cloud storage, distributed databases, etc. The virtual database interface allows cross-data source queries. Users can query data in different archived storages through standard SQL syntax without data migration.

[0083] In an embodiment of the present invention, a virtual database interface is used to perform remote calls and virtualized access to archived data. Specifically, the remote call is used to execute queries when the archived data is located at different physical locations or in distributed storage, including: supporting cross-storage system queries through the database adaptation layer JDBC and automatically discovering data sources; adopting a distributed query optimizer to split SQL queries into sub-queries, locally execute them at the data source end, and then return the results; combining a remote query security control mechanism to perform user authentication before the query.

[0084] The virtualized access is used to provide transparent access in different data storage systems, enabling users to be unaware of the data storage location, including: adopting a virtual data view (Virtual Data View) to dynamically map archived data to a standardized data structure when users query; combining an on-demand loading strategy to only extract the required data during the query process, avoiding full restoration of archived data; adopting a cross-data source federated query mechanism to execute SQL queries on different data storage systems and merge the query results to return to users.

[0085] In an embodiment of the present invention, when the system detects that it reaches a preset trigger condition, a data archiving operation is automatically performed. The preset trigger conditions include:

[0086] When the disk space is lower than 10%, data archiving is automatically triggered;

[0087] When the CPU load is higher than 85% and lasts for more than 1 hour, data archiving is automatically triggered;

[0088] When the system detects an abnormal database connection or multiple query failures, data archiving is automatically triggered.

[0089] In an embodiment of the present invention, a vector index technology is adopted to optimize the query of archived data, specifically including:

[0090] S501, extracting text content, metadata, and statistical information for each data item and converting them into feature vectors;

[0091] S502, setting multiple levels, each level adopting a different hash function family, allocating data items to hash buckets according to hash values, and recursively constructing a tree-like index structure. Each node in the structure represents a hash bucket, and the leaf nodes contain pointers to the original data items;

[0092] S503, after standardizing the query data, sequentially traversing each layer of hash buckets using the same hash function as when constructing the index, and initially screening candidate data using the stored representative vectors. Finally, calculate the cosine similarity of the candidate data items at the leaf nodes and return the query results sorted by similarity.

[0093] S504, when new data is added, the hash index is automatically updated.

[0094] The embodiments of the present invention have been described above. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make many forms, all of which fall within the protection scope of this embodiment.

Claims

1. A method for salvage archiving and disposal of important data in a shutdown system, characterized in that, It includes the following steps: S101, before shutting down the system data archiving, automatically parse the data structure of the stored data using a machine learning model and perform scanning and extraction; S102, adopt an on-the-fly storage and conversion mechanism to store and convert the scanned and extracted data, and based on content feature analysis, automatically match the corresponding archiving format for archiving; S103, adopt an encryption storage mechanism based on the zero-trust architecture to protect the archived data, and adopt a dynamic key management strategy to dynamically allocate, recycle, and destroy keys according to the access scenario, and combine distributed key storage technology to protect the data; S104, adopt a streaming retrieval mechanism to provide query and access to the archived data, and based on the virtual database interface, perform remote call and virtualized access to the archived data; Among them, the streaming retrieval mechanism uses a distributed stream processing framework to perform real-time parsing on the archived data. The specific steps of the streaming retrieval mechanism include: S401, based on the streaming processing technology, extract data as needed, and adopt block loading to limit the data reading volume; S402, adopt vector index technology to optimize the query of the archived data, including: S501, extract text content, metadata, and statistical information for each data item and convert them into feature vectors; S502, set multiple levels, each level uses a different hash function family, allocate data items to hash buckets according to the hash value, and recursively form a tree-like index structure. Each node in the structure represents a hash bucket, and the leaf nodes contain pointers to the original data items; S503, after normalizing the query data, use the same hash function as when constructing the index to traverse each layer of hash buckets in turn, and use the stored representative vectors to preliminarily screen candidate data. Finally, calculate the cosine similarity of the candidate data items at the leaf nodes and return the query results sorted by similarity; S504, when new data is added, automatically update the hash index.

2. The rescue archiving and disposal method for important data of a shutdown system according to claim 1, characterized in that Automatically parse the data structure of the stored data using a machine learning model and perform scanning and extraction. The specific steps include: S201, use an automatic storage type recognition algorithm to identify the data storage type of the shutdown system; S202, adopt a machine learning model to parse the internal structure of the stored data of the shutdown system; S203, extract time series features, text pattern features, and numerical distribution features from the parsed data structure; S204, scan and extract the stored data according to the extracted features.

3. A method for rescue archiving and disposal of important data in a shutdown system according to claim 1, characterized in that, The specific steps of S102 include: S301, perform conversion of the archiving storage format of the scanned and extracted data; S302, generate metadata based on the data after format conversion and store the association relationship between the data; S303, adjust the storage method of the data after format conversion; S304, manage the storage structure of the data after format conversion.

4. A method for rescue archiving and disposal of important data in a shutdown system according to claim 1, characterized in that, Use a symmetric encryption algorithm to encrypt the archived data; during the data storage and access process, adopt an on-demand allocation strategy, dynamically allocate a unique temporary key for each access, and perform real-time auditing of the access behavior; Adopt a key automatic expiration strategy, and the key is only valid within a preset time window and is automatically recycled when it times out; Adopt an irrecoverable destruction strategy. After the end of the data life cycle, use a multiple erasure algorithm to destroy the key, and record the key destruction log through the blockchain evidence storage mechanism.

5. A method for rescue archiving and disposal of important data in a shutdown system according to claim 1, characterized in that, Store the key in a distributed key storage manner. The distributed key storage specifically includes: multi-node key sharding storage and geographical distributed storage.

6. A method for rescue and archival disposal of important data in a shutdown system according to claim 1, characterized in that, Use the first security policy to access the archived data, where the first security policy includes: fine-grained access control and zero-knowledge access.

7. A method for rescue and archival disposal of important data in a shutdown system according to claim 1, characterized in that, The specific steps of the streaming retrieval mechanism further include: S403, Parse the query request through the streaming SQL engine.

8. A method for rescue archiving and disposal of important data in a shutdown system according to claim 1, characterized in that, When it is detected that the system reaches a preset trigger condition, automatically execute the data archiving operation. The preset trigger conditions include: The disk space is lower than the first preset threshold; The CPU load is higher than the second preset threshold and lasts for more than the first preset time interval; The system detects an abnormal database connection.

Citation Information

Patent Citations

  • Data archiving method and device

    CN113779137A

  • Block chain private data sharing method and system

    CN118869243A