Emergency filing processing method for important data of shutdown system

Through the machine learning model, the data structure is analyzed and combined with the storage and conversion mechanism, the problem of data format incompatibility in shutting down the system data archive is solved. Through the zero-trust architecture and streaming retrieval mechanism, data security and query efficiency are improved, and efficient and secure data archiving and query are achieved.

CN120011311AActive Publication Date: 2025-05-16HANGZHOU YIKANGXIN TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510475005.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-16
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

There are currently problems with incompatible data formats, insufficient security, low query access efficiency, and low archive automation in relation to the data archiving method of shutting down the system.

Method used

The machine learning model is used to automatically analyze the data structure, combine the storage and conversion mechanism to convert the data format, and ensure data security and query efficiency through the zero-trust architecture encrypted storage mechanism and streaming retrieval mechanism.

Benefits of technology

It realizes automatic adaptation of data formats, improves archive efficiency and storage standardization; ensures full encryption of data during storage, transmission and access, and improves data security; through the streaming retrieval mechanism, query efficiency and flexibility are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011311A_ABST
    Figure CN120011311A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data storage and management, and discloses a shutdown system important data rescue filing processing method, which comprises the following steps: S101, automatically analyzing a data structure of stored data by adopting a machine learning model, and scanning and extracting; s102, performing storage and format conversion on the scanned and extracted data by adopting a storage-while-conversion mechanism; s103, performing security protection on the archived data by adopting an encryption storage mechanism based on a zero-trust architecture, dynamically distributing, recovering and destroying keys according to an access scene by adopting a dynamic key management strategy, and protecting the data in combination with a distributed key storage technology; and S104, querying and accessing the archived data by adopting a streaming retrieval mechanism, and carrying out remote calling and virtualized access on the archived data based on a virtual database interface. According to the method, the format compatibility of the archived data is improved, and the query efficiency and the cross-data-source access capability of the archived data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of data storage and management, and in particular relates to a method for rescuing and archiving important data of a shutdown system. Background Art

[0002] With the rapid development of information technology, the information systems of enterprises, governments and scientific research institutions are constantly upgraded and iterated. Some old systems are facing closure due to aging software and hardware, backward technical architecture, high operation and maintenance costs, etc. However, these systems often store important business data, user information, financial records, historical logs and other key data. Direct deletion or improper migration will lead to data loss, compliance risks and business continuity issues.

[0003] The existing shutdown system data archiving methods mainly rely on traditional methods such as database backup, log storage, and manual data migration, but there are the following technical limitations: incompatible data formats, insufficient storage security, low query access efficiency, and low archiving automation. First, the data formats, database structures, and storage methods of different systems are different, and simple data exports are difficult to meet cross-system compatibility requirements. Secondly, traditional data storage methods lack dynamic encryption and key management, and data is susceptible to tampering or leakage during the archiving and storage process. Furthermore, existing archiving solutions mostly use full recovery queries, which leads to data access delays and cannot meet efficient query requirements. Summary of the invention

[0004] The present invention provides a method for rescuing and archiving important data of a shutdown system, which solves the technical problems of incompatible data formats, insufficient security, low query efficiency and low archiving automation level in related technologies.

[0005] The present invention provides a method for rescuing and archiving important data of a shutdown system, comprising the following steps: S101, before shutting down the system data archiving, automatically parse the data structure of the stored data using a machine learning model and perform scanning and extraction; S102, using a storage-while-conversion mechanism to store and convert the format of the scanned and extracted data, and automatically matching the corresponding archiving format for archiving based on content feature analysis; S103, adopt an encrypted storage mechanism based on zero-trust architecture to protect the archived data, and adopt a dynamic key management strategy to dynamically allocate, recycle and destroy keys according to access scenarios, and protect data in combination with distributed key storage technology; S104, using a streaming retrieval mechanism to provide query and access to the archived data, and based on a virtual database interface, performing remote calls and virtualized access to the archived data.

[0006] Furthermore, the stored data is automatically parsed using a machine learning model and scanned and extracted. The specific steps include: S201, using an automatic storage type identification algorithm to identify the data storage type of the shutdown system; S202, using a machine learning model to analyze the internal structure of the shutdown system storage data; S203, extracting time series features, text pattern features, and numerical distribution features from the parsed data structure; S204: Scan and extract the stored data according to the extracted features.

[0007] Furthermore, the specific steps of S102 include: S301, converting the scanned and extracted data into an archiving storage format; S302, generating metadata according to the format-converted data, and storing the association relationship between the data; S303, adjusting the storage method of the data after the format conversion; S304, managing the storage structure of the data after format conversion.

[0008] Furthermore, a symmetric encryption algorithm is used to encrypt the archived data; during the data storage and access process, an on-demand allocation strategy is adopted, a unique temporary key is dynamically allocated for each access, and access behavior is audited in real time; Adopting the automatic key expiration strategy, the key is only valid within the preset time window and will be automatically recycled after the timeout; Adopting an irreversible destruction strategy, after the data life cycle ends, the key is destroyed using a multiple erasure algorithm, and the key destruction log is recorded through the blockchain evidence storage mechanism.

[0009] Furthermore, the keys are stored in a distributed key storage manner, and the distributed key storage specifically includes: multi-node key shard storage and geographically distributed storage.

[0010] Further, the archived data is accessed using a first security policy, wherein the first security policy includes: fine-grained access control and zero-knowledge access.

[0011] Furthermore, the streaming retrieval mechanism uses a distributed stream processing framework to perform real-time analysis on archived data. The specific steps of the streaming retrieval mechanism include: S401, based on streaming technology, extract data on demand and use block loading to limit the amount of data read; S402, optimizing the query of archived data by using vector index technology; S403, parsing the query request through the streaming SQL engine.

[0012] Furthermore, when it is detected that the system reaches a preset trigger condition, the data archiving operation is automatically performed, and the preset trigger condition includes: The disk space is below a first preset threshold; The CPU load is higher than the second preset threshold and lasts longer than the first preset time interval; The system detected a database connection abnormality.

[0013] Furthermore, vector indexing technology is used to optimize the query of archived data. The specific steps include: S501, extracting text content, metadata and statistical information for each data item and converting them into feature vectors; S502, setting multiple levels, each level using a different hash function family, assigning data items to hash buckets according to hash values, and recursively forming a tree index structure, in which each node represents a hash bucket, and a leaf node contains a pointer to the original data item; S503, after the query data is standardized, the same hash function as that used in index construction is used to traverse the hash buckets of each layer in turn, and the stored representative vectors are used to preliminarily screen the candidate data, and finally the cosine similarity of the candidate data items is calculated at the leaf node, and the query results are returned in order of similarity; S504, when new data is added, the hash index is automatically updated.

[0014] The beneficial effects of the present invention are as follows: the present invention adopts a machine learning model to automatically parse the data structure, and combines with an automatic storage type recognition algorithm to accurately identify different data storage types such as relational databases, non-relational databases, log files, and image files, and based on a storage-while-conversion mechanism, synchronously completes format conversion during the archiving storage process. The method can automatically adapt to a variety of archiving storage formats, ensure data compatibility between different systems, reduce manual intervention in data format conversion, and improve archiving efficiency and storage standardization. The present invention adopts an encrypted storage mechanism of zero-trust architecture to ensure that data is in an encrypted protection state during the entire process of storage, transmission and access. Combined with dynamic key management strategies, keys can be dynamically allocated, recovered and destroyed according to access scenarios to prevent long-term exposure of keys. At the same time, distributed key storage technology is adopted to ensure the key's anti-attack capability and recoverability, prevent single-point key leakage, and improve the long-term security of data after archiving; The present invention adopts a streaming retrieval mechanism, combined with a distributed stream processing framework, to support on-demand extraction of archived data and improve query efficiency; based on a virtual database interface, archived data can be directly queried across data sources without physically restoring the data, thereby improving the flexibility and real-time performance of data retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 The present invention is a flowchart of a method for rescuing and archiving important data of a shutdown system. DETAILED DESCRIPTION

[0016] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that the discussion of these embodiments is only to enable those skilled in the art to better understand and implement the subject matter described herein, and the functions and arrangements of the elements discussed may be changed without departing from the scope of protection of the contents of this specification. Each example may omit, replace or add various processes or components as needed. In addition, the features described relative to some examples may also be combined in other examples.

[0017] like Figure 1 As shown, a method for rescuing and archiving important data of a shutdown system includes: S101, before shutting down the system data archiving, automatically parse the data structure of the stored data using a machine learning model and perform scanning and extraction; S102, using a storage-while-conversion mechanism to store and convert the format of the scanned and extracted data, and automatically matching the corresponding archiving format for archiving based on content feature analysis; S103, adopt an encrypted storage mechanism based on zero-trust architecture to protect the archived data, and adopt a dynamic key management strategy to dynamically allocate, recycle and destroy keys according to access scenarios, and protect data in combination with distributed key storage technology; S104, using a streaming retrieval mechanism to provide query and access to the archived data, and based on a virtual database interface, performing remote calls and virtualized access to the archived data.

[0018] In one embodiment of the present invention, a machine learning model is used to automatically parse the data structure of the stored data and perform scanning and extraction, and the specific steps include: S201, using an automatic storage type identification algorithm to identify the data storage type of the shutdown system; Specifically, data storage types include: relational databases, non-relational databases, log files, and image files. The specific steps of the automatic storage type recognition algorithm include: S2011, detecting database metadata through database connection driver and identifying database type; S2012, identifying log files and image files through file header signature monitoring; S2013 uses regular expression matching to analyze log formats and automatically adapts log structures, for example, timestamp + event + parameter formats; S202, using a machine learning model to analyze the internal structure of the shutdown system storage data; Specifically, clustering algorithms are used to analyze the data distribution in the database and automatically group similar data structures. For example, order tables, invoice tables, and transaction tables are grouped into financial data categories, and user tables, customer tables, and employee tables are grouped into personnel information categories. Use pattern matching algorithms to standardize and map data patterns of different systems to ensure data compatibility of heterogeneous databases. Specifically, use a Schema Mapping algorithm based on template matching to convert different database patterns into a unified data structure. For example, convert VARCHAR2 stored in Oracle databases and VARCHAR stored in MySQL databases into STRING format. Use deep neural networks or decision trees to automatically predict the data type of a field based on field value distribution, length, and format characteristics; Use Bayesian networks to model the field associations between database tables and automatically identify primary and foreign key relationships to preserve the integrity of the database structure; S203, extracting time series features, text pattern features and numerical distribution features from the parsed data structure, specifically, identifying the time field in the stored data, parsing the time format, and determining the time granularity, such as seconds, minutes, hours and days; using word frequency-inverse document frequency to analyze important keywords in the text field, and using regular expressions to identify text in specific formats such as addresses, names, amounts, etc.; calculating the mean, standard deviation and extreme value of the numerical field, and determining the value range and abnormal points of the data; S204, scan and extract the stored data according to the extracted features. Specifically, adopt a streaming processing mechanism to read the database data line by line to avoid excessive consumption of system resources caused by one-time loading; use regular expressions to parse the log content and extract key fields; combine optical character recognition and automatic speech recognition to extract content from images, PDF documents, and audio data to improve the availability of unstructured data.

[0019] In one embodiment of the present invention, the specific steps of S102 include: S301, converting the scanned and extracted data into an archive storage format to ensure data structuring and improve query and recovery capabilities; The specific steps of S301 include: S3011, determine the optimal storage format based on field type, data hierarchy structure and distribution characteristics; S3012, using format matching algorithms to convert different types of data, specifically, converting relational database tables into JSON format, converting non-relational database tables into structured table storage, extracting key fields from log files and storing them in the index system, and generating searchable texts from image files in combination with OCR analysis; S3013, adopts streaming processing mechanism to synchronously perform format conversion during data storage.

[0020] S302, metadata is generated according to the format-converted data, and the association relationship between the data is stored to ensure that the archived data can still retain the business logic. Specifically, the metadata represents the information attributes of the data, and the metadata includes: data source, storage time, data format, field definition and hash value; the association relationship between the data refers to the primary key-foreign key relationship between the data tables, and the association relationship between the data is stored through the relationship table; S303, adjusting the storage method of the format-converted data to adapt to different data storage media. Specifically, the data is stored in a hierarchical manner according to the data access frequency, with high-access frequency data being stored in a cloud database and low-access frequency data being stored in a disk; S304, managing the storage structure of the format-converted data to ensure the long-term availability and security of the archived data. Specifically, a partition storage strategy is adopted to partition the archived data according to time, business category and data source to improve retrieval efficiency.

[0021] In one embodiment of the present invention, the zero-trust model adopts a security strategy of never trusting and always verifying, and any data access requires identity authentication, device verification, and behavior analysis; it does not rely on traditional network boundary security mechanisms, but is based on fine-grained access control and dynamic trust evaluation to ensure that data access is still subject to strict verification even in a trusted environment; In one embodiment of the present invention, a symmetric encryption algorithm is used to encrypt the archived data; during the data storage and access process, an on-demand allocation strategy is adopted, a unique temporary key is dynamically allocated for each access, and the access behavior is audited in real time. Based on factors such as data sensitivity, user identity, access time, and device location, a multi-factor decision is made to determine whether key allocation is allowed; The key automatic expiration strategy is adopted. The key is valid only within the preset time window and automatically recovered after the timeout to prevent the long-term stored keys from being abused. In addition, the revocation mechanism is combined to revoke the relevant keys immediately and terminate the access rights when abnormal access is detected. Abnormal access includes but is not limited to: high-risk IP address login, abnormal data download and other behaviors; An irreversible destruction strategy is adopted. After the data life cycle ends, the key is destroyed using a multiple erasure algorithm, and the key destruction log is recorded through the blockchain evidence mechanism to ensure that the key management process is auditable and cannot be tampered with. The multiple erasure algorithm adopts the NIST800-88 standard.

[0022] In one embodiment of the present invention, keys are stored in a distributed key storage manner, specifically including: multi-node key sharding storage and geographically distributed storage; Among them, multi-node key shard storage adopts Shamir encryption sharing algorithm to divide the encryption key into multiple fragments and store them in different independent security modules to ensure that a single storage node cannot recover the complete key. The complete key can only be recovered when the preset key fragment threshold is met, so as to improve the anti-attack capability of key storage; Geographically distributed storage adopts a cross-regional key storage strategy, storing keys on secure servers in different physical locations to prevent key loss or attacks due to failure of a single data center. When the key is recovered, a consensus mechanism is used to ensure that the key can only be reassembled in a trusted environment to prevent malicious attackers from obtaining key fragments for synthesis.

[0023] In one embodiment of the present invention, a first security policy is used to access archived data, wherein the first security policy includes: fine-grained access control and zero-knowledge access. Specifically, fine-grained access control is used to dynamically adjust access rights based on user identity, access device, time, geographic location and access behavior. Fine-grained access control includes: attribute-based access control and role-based access control. Attribute-based access control means setting access policies based on user attributes to ensure that data access complies with predefined security rules. Role-based access control means defining different permission levels based on the user's role in the system to ensure permission boundaries for different user groups. Permission levels include but are not limited to: administrator permissions, auditor permissions and ordinary user permissions. Zero-knowledge access uses a zero-knowledge proof mechanism to enable users to verify their access rights through encrypted mathematical proofs without directly obtaining data content.

[0024] In one embodiment of the present invention, the streaming retrieval mechanism uses the distributed stream processing framework Apache Kafka to perform real-time analysis on archived data. The specific steps of the streaming retrieval mechanism include: S401, based on the streaming technology, extract data on demand, and use block loading to limit the amount of data read, thereby reducing the amount of data read; S402, using vector indexing technology to optimize the query of archived data and improve the retrieval speed. Specifically, using B+ tree indexing for SQL data, using Elasticsearch indexing for text and logs, supporting fuzzy query and full-text retrieval, and using OCR+ feature vector indexing for image files, supporting content-based retrieval; S403, parsing the query request through the streaming SQL engine Presto to improve query performance.

[0025] Traditional database queries usually require local data storage, but archived data may be distributed in different storage systems, such as local storage, cloud storage, distributed databases, etc. The virtual database interface allows cross-data source queries. Users can use standard SQL syntax to query data in different archive storages without data migration.

[0026] In one embodiment of the present invention, a virtual database interface is used to perform remote calls and virtualized access to archived data. Specifically, remote calls are used to execute queries on archived data stored in different physical locations or distributed storage, including: supporting cross-storage system queries through the database adaptation layer JDBC and automatically discovering data sources; using a distributed query optimizer to split SQL queries into subqueries, and returning results after local execution at the data source end; combining a remote query security control mechanism to perform user identity authentication before querying; Virtualized access is used to provide transparent access in different data storage systems, so that users do not need to be aware of the data storage location, including: using virtual data view Virtual Data View to dynamically map archived data to standardized data structure when users query; combining on-demand loading strategy to extract only required data during the query process to avoid complete recovery of archived data; using cross-data source joint query mechanism to execute SQL queries on different data storage systems, and merge query results and return them to users.

[0027] In one embodiment of the present invention, when it is detected that the system reaches a preset trigger condition, the data archiving operation is automatically performed, and the preset trigger condition includes: When the disk space is less than 10%, data archiving is automatically triggered; When the CPU load is higher than 85% and lasts for more than 1 hour, data archiving is automatically triggered; When the system detects an abnormal database connection or multiple query failures, data archiving is automatically triggered.

[0028] In one embodiment of the present invention, vector indexing technology is used to optimize the query of archived data, specifically including: S501, extracting text content, metadata and statistical information for each data item and converting them into feature vectors; S502, setting multiple levels, each level using a different hash function family, assigning data items to hash buckets according to hash values, and recursively forming a tree index structure, in which each node represents a hash bucket, and a leaf node contains a pointer to the original data item; S503, after the query data is standardized, the same hash function as that used in index construction is used to traverse the hash buckets of each layer in turn, and the stored representative vectors are used to preliminarily screen the candidate data, and finally the cosine similarity of the candidate data items is calculated at the leaf node, and the query results are returned in order of similarity; S504, when new data is added, the hash index is automatically updated.

[0029] The embodiments of the present invention are described above, but the present invention is not limited to the above-mentioned specific implementation modes. The above-mentioned specific implementation modes are merely illustrative and not restrictive. Under the guidance of the present embodiment, ordinary technicians in this field can also make many forms, which are all within the protection of the present embodiment.

Claims

1. A method for rescuing and archiving important data of a shutdown system, characterized in that: The following steps are involved: S101, before shutting down the system data archiving, automatically parse the data structure of the stored data using a machine learning model and perform scanning and extraction; S102, using a storage-while-conversion mechanism to store and convert the format of the scanned and extracted data, and automatically matching the corresponding archiving format for archiving based on content feature analysis; S103, adopt an encrypted storage mechanism based on zero-trust architecture to protect the archived data, and adopt a dynamic key management strategy to dynamically allocate, recycle and destroy keys according to access scenarios, and protect data in combination with distributed key storage technology; S104, using a streaming retrieval mechanism to provide query and access to the archived data, and based on a virtual database interface, performing remote calls and virtualized access to the archived data.

2. A method for rescuing and archiving important data of a shutdown system according to claim 1, characterized in that: The machine learning model is used to automatically parse the data structure and scan and extract the stored data. The specific steps include: S201, using an automatic storage type identification algorithm to identify the data storage type of the shutdown system; S202, using a machine learning model to analyze the internal structure of the shutdown system storage data; S203, extracting time series features, text pattern features, and numerical distribution features from the parsed data structure; S204: Scan and extract the stored data according to the extracted features.

3. The method for rescuing and archiving important data of a shutdown system according to claim 1, characterized in that: The specific steps of S102 include: S301, converting the scanned and extracted data into an archiving storage format; S302, generating metadata according to the format-converted data, and storing the association relationship between the data; S303, adjusting the storage method of the data after the format conversion; S304, managing the storage structure of the data after format conversion.

4. The method for rescuing and archiving important data of a shutdown system according to claim 1, characterized in that: Use symmetric encryption algorithms to encrypt archived data; during data storage and access, adopt an on-demand allocation strategy, dynamically allocate a unique temporary key for each access, and conduct real-time audits on access behaviors; Adopting the automatic key expiration strategy, the key is only valid within the preset time window and will be automatically recycled after the timeout; Adopting an irreversible destruction strategy, after the data life cycle ends, the key is destroyed using a multiple erasure algorithm, and the key destruction log is recorded through the blockchain evidence storage mechanism.

5. The method for rescuing and archiving important data of a shutdown system according to claim 1, characterized in that: The keys are stored in a distributed key storage manner, which specifically includes: multi-node key sharding storage and geographically distributed storage.

6. The method for rescuing and archiving important data of a shutdown system according to claim 1, characterized in that: The archived data is accessed using a first security policy, wherein the first security policy includes: fine-grained access control and zero-knowledge access.

7. The method for rescuing and archiving important data of a shutdown system according to claim 1, characterized in that: The streaming retrieval mechanism uses a distributed stream processing framework to perform real-time analysis on archived data. The specific steps of the streaming retrieval mechanism include: S401, based on streaming technology, extract data on demand and use block loading to limit the amount of data read; S402, optimizing the query of archived data by using vector index technology; S403, parsing the query request through the streaming SQL engine.

8. The method for rescuing and archiving important data of a shutdown system according to claim 1, characterized in that: When the system detects that the preset trigger conditions have been met, the data archiving operation is automatically performed. The preset trigger conditions include: The disk space is below a first preset threshold; The CPU load is higher than the second preset threshold and lasts longer than the first preset time interval; The system detected a database connection abnormality.

9. The method for rescuing and archiving important data of a shutdown system according to claim 7, characterized in that: Vector index technology is used to optimize the query of archived data. The specific steps include: S501, extracting text content, metadata and statistical information for each data item and converting them into feature vectors; S502, setting multiple levels, each level using a different hash function family, assigning data items to hash buckets according to hash values, and recursively forming a tree index structure, in which each node represents a hash bucket, and a leaf node contains a pointer to the original data item; S503, after the query data is standardized, the same hash function as that used in index construction is used to traverse the hash buckets of each layer in turn, and the stored representative vectors are used to preliminarily screen the candidate data, and finally the cosine similarity of the candidate data items is calculated at the leaf node, and the query results are returned in order of similarity; S504, when new data is added, the hash index is automatically updated.

Citation Information

Patent Citations

  • Data archiving method and device

    CN113779137A

  • Streaming vector search method, device and system based on Flink

    CN117689451A

  • Large-scale network public opinion-oriented Elasticsearch retrieval optimization system

    CN118503512A

  • Block chain private data sharing method and system

    CN118869243A

  • Shutdown system-based important data rescue filing method and system

    CN119415319A