Electronic archive distributed storage method based on content sensitivity classification encryption and fragmentation redundancy

By adopting content sensitivity-driven classification encryption, dynamic sharding redundancy and multi-node collaborative recovery methods in distributed storage systems, single point of failure and security risks in traditional storage methods are solved, and efficient, secure and reliable distributed storage of electronic files is achieved.

CN120067212APending Publication Date: 2025-05-30SHANXI CHENGCHENG ARCHIVES MANAGEMENT SERVICE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510216860.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional centralized storage methods have single point of failure, high-risk security risks and insufficient scalability, while distributed storage technology still has challenges in data security, privacy protection and dynamic resource scheduling.

Method used

A distributed storage method is designed using a classification encryption strategy based on content sensitivity, combining dynamic sharding and redundancy mechanisms, as well as a multi-node collaborative fault tolerance and recovery scheme. This method ensures high security and high availability of data in a distributed environment through dynamic classification encryption, adaptive sharding redundancy and multi-node collaborative recovery.

Benefits of technology

It realizes dynamic adjustment of encryption strength based on the sensitivity of electronic archive content, optimizes storage efficiency and security; through adaptive sharding redundancy strategy, the system reliability and fault tolerance are enhanced; the multi-node collaborative recovery mechanism significantly improves data recovery speed and high availability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067212A_ABST
    Figure CN120067212A_ABST
Patent Text Reader

Abstract

The invention provides an electronic file distributed storage method based on content sensitivity classification encryption and fragmentation redundancy, and aims to improve the safety and reliability of an electronic file storage system. The method comprises the following steps: firstly, carrying out content classification on electronic archives, and carrying out hierarchical encryption according to sensitivity to ensure that different types of archives have adaptive security policies; then, the encrypted archive data is subjected to fragmentation processing and redundancy distribution, so that the fault-tolerant capability and the access efficiency of the data are improved. Fragmented data is stored to a plurality of nodes through a distributed network, and rapid positioning and retrieval are carried out by using a distributed hash table. In a data access process, a user authority management and key distribution mechanism is combined to ensure that the user can only access the archive information within the authority range of the user. The method effectively solves the problems of data security, privacy protection and system robustness in electronic archive distributed storage, and is suitable for various archive management scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of archives and distributed storage, and relates to the encrypted storage, data sharding, and fault tolerance mechanism of electronic archives, and is specifically applied to the distributed storage system of electronic archives with high security and high reliability requirements. Background Art

[0002] With the rapid advancement of digitalization, electronic archives have become key resources for various institutions, enterprises, and government departments, and are widely used to record sensitive information such as important documents, transaction records, patient medical records, and financial data. However, in a centralized storage system, these electronic archives often become high-risk targets for cyberattacks, virus infections, and malicious tampering due to the centralized data location. In addition, centralized storage also faces the problem of single-point failure. Once the core node fails, the entire system may not be able to operate normally, resulting in data unavailability or permanent loss. Therefore, in order to effectively protect this key information, there is an urgent need for an electronic archive storage solution with a distributed architecture, high security, and fault tolerance capabilities to ensure the continuous availability and security of data in the event of cyberattacks, hardware failures, or other accidents.

[0003] Development of classification encryption technology: In traditional electronic archive encryption methods, a unified encryption standard is usually adopted for all data. However, electronic archives contain different levels of sensitive information. For example, the security requirements for financial records and ordinary documents are different. Using unified encryption not only leads to unnecessary consumption of computing resources but also may not fully meet the security requirements of sensitive data. Classification encryption technology provides a solution for this: First, the sensitivity level of the data content is classified, and then an appropriate encryption strategy is selected for different archive categories according to the security requirements. This differential encryption not only improves the security of the data but also effectively reduces the access delay and computing cost. However, classification encryption poses practical challenges in a distributed environment, including how to efficiently classify massive amounts of data automatically, how to select the optimal encryption strategy to meet security and performance requirements, and how to ensure the encryption effect while the data is distributed across nodes.

[0004] Distributed storage systems distribute data across multiple storage nodes, improving resource utilization and data access efficiency by dispersing the storage load and featuring fault tolerance. However, in practical applications, distributed systems inevitably encounter problems such as node failures, network latency, and data desynchronization, which may all affect the system's reliability. For example, under high-concurrency access, excessive load on some nodes can lead to a decline in data access speed and even make the system temporarily unavailable. At the same time, how to quickly recover data and maintain data consistency when nodes fail has become the focus of fault-tolerant design. To this end, distributed systems usually need to adopt various redundancy strategies, such as replica redundancy or erasure coding technology, to be able to rely on backups or redundant data to reconstruct data content in case of node failures or data loss, ensuring the high availability of the system.

[0005] In distributed storage, data sharding and redundancy strategies are commonly used fault-tolerant measures. Data sharding divides large chunks of data into multiple small pieces and disperses them across different nodes, so that even if a certain node fails, the data will not be completely lost. At the same time, redundancy strategies generate backups for data or use erasure coding to distribute some redundant information across multiple nodes to ensure that data can still be reconstructed in case of partial node failures. For example, erasure coding can provide high data recovery capabilities with less storage overhead and is widely used in distributed systems. However, in practical applications, sharding and redundancy strategies need to balance data storage efficiency and security. Especially for sensitive electronic files, how to improve the system's fault-tolerant performance and access efficiency using sharding and redundancy while ensuring data privacy and encryption security is an important technical challenge.

[0006] In distributed storage systems, user privilege management directly affects data security and privacy protection, especially in scenarios where sensitive electronic files are stored. Different users may have different access privileges. Traditional centralized privilege management schemes are difficult to adapt to the distributed environment. Especially when data is distributed across multiple nodes, the complexity of privilege control increases significantly. For example, when synchronizing data between nodes, it may lead to inconsistent privilege information, posing risks of unauthorized data access or privilege leakage. Therefore, a rigorous privilege control scheme needs to be designed in distributed storage systems, combined with key management and distributed authentication, to ensure that users can only access data within their authorized scope. In addition, the privilege synchronization and update mechanism between different nodes also needs to be considered to ensure the real-time performance and security of the system and effectively prevent potential risks caused by unauthorized access to electronic file data. Summary of the Invention

[0007] With the widespread application of electronic files in fields such as government affairs, healthcare, and finance, the amount of data has grown exponentially, while posing higher requirements for the security, reliability, and privacy protection of storage systems. However, traditional centralized storage methods have problems such as single-point failures, high-risk security hazards, and insufficient scalability. Although distributed storage technology has strong fault tolerance, there are still challenges in data security, privacy protection, and dynamic resource scheduling. Therefore, how to design a distributed storage method that can combine intelligent encryption strategies based on content sensitivity, dynamic fragmentation and redundancy mechanisms, and multi-node collaborative fault tolerance and recovery solutions has become a key research direction for solving the problems of electronic file storage and management.

[0008] To solve the above technical problems, the technical solution adopted in the present invention is as follows:

[0009] A distributed storage method for electronic files based on content sensitivity classification encryption and fragmentation redundancy, comprising the following steps:

[0010] S1. Dynamically classify electronic files based on content sensitivity analysis and hierarchically define security levels; introduce a sensitivity-driven encryption algorithm strategy, and dynamically select an appropriate AES (Advanced Encryption Standard) symmetric encryption algorithm to ensure that sensitive files are preferentially encrypted with high strength, improving the security and performance compatibility of the system;

[0011] S2. An adaptive fragmentation and redundancy generation strategy that dynamically adjusts the fragmentation scale and redundancy level according to the access frequency and importance of file data; can reduce storage redundancy while ensuring data security, optimize storage space utilization, and update fragmentation allocation in real time according to the health status of nodes;

[0012] S3. In a distributed network, introduce a classification-aware storage mechanism, and allocate fragmented data and redundant blocks to different nodes considering data categories, node performance, and network topology; different from traditional random or uniform allocation methods, improve the access speed and overall reliability of the system by optimizing the data distribution strategy;

[0013] S4. Combine a scenario-based permission control system to dynamically adjust the permission verification and key distribution strategies according to user permissions, data sensitivity levels, and access scenarios; through the key dynamic update mechanism, overcome the problems of inconsistent permission synchronization and control in traditional distributed systems, and enhance the controllability of user access behavior and the privacy protection ability of data;

[0014] S5. A multi-node collaborative fault tolerance and recovery solution that combines redundant fragmentation and a distributed consistency protocol to dynamically optimize the data recovery path during the data recovery process of a faulty node; adjust the recovery strategy according to the real-time status of nodes to ensure that the system has fault tolerance in scenarios of load or node failure.

[0015] The method for dynamically classifying electronic files and hierarchically defining security levels in S1 is as follows:

[0016] S11. Extracting the content features of the file. For each electronic file D i perform content analysis and extract its feature vector V i = [v i1 , v i2 ,..., v in , where v ij represents the value of the i-th file on the j-th feature, n is the number of features, and the features include keyword frequency, topic distribution, and the number of occurrences of sensitive words;

[0017] S12. Calculating the content sensitivity score of each file where w j is the weight of the j-th feature, reflecting the influence degree of this feature on sensitivity; the weight w j is determined by expert experience and machine learning models; according to the sensitivity score S i the files are divided into different security levels L i , set multiple sensitivity thresholds T i , and the security level L i is defined as:

[0018]

[0019] For each security level L i allocate the corresponding encryption algorithm E i and key length K i , form a mapping relationship E: L i → (E i , K i ), L 1 uses AES-128-bit encryption, L 2 uses AES-192-bit encryption, L 3 uses AES-256-bit encryption.

[0020] The method for an adaptive fragmentation and redundancy generation strategy in S2 is as follows:

[0021] S21. According to the size i of the electronic file D and the preset upper limit of the fragmentation size S max , determine the number of fragments where represents rounding up to ensure that the size of each fragment does not exceed S max ; the file D i is divided into N i data blocks

[0022] For the sensitivity level L of each file i , dynamically adjust the redundancy ratio R i , and generate the number of additional redundant blocks Among them, R i 's value is set according to the sensitivity level L i , R 1 = 0.5, R 2 = 1.0, R 3 = 2.0; Use erasure codes to generate redundant blocks Among them, f(·) represents the erasure code function;

[0023] S22, the sharding and redundancy allocation strategy distributes shards and redundant blocks to different storage nodes N k Among them, uniform distribution evenly distributes the original shards and redundant blocks to K nodes, and defines the shard allocation function g(D ij , R im ):

[0024]

[0025] Among them, Load(N k ) represents the node load, Risk(N k ) represents the node security risk, and α is a trade-off coefficient;

[0026] S23, after completing the sharding and redundancy allocation, the storage node performs integrity verification on the allocated data blocks D ij and redundant blocks R im to generate hash values H ij = Hash(D ij ), H im = Hash(R im ); When a certain node N k fails, use the remaining shards and redundant blocks to recover the lost data; Assume the lost data is {D ij , R im}, and the recovery formula is:

[0027] {D ij} = f -1 (R im , D ij′ )

[0028] Among them, f -1 is the inverse operation of the erasure code, and D ij′ is the remaining complete shard data; It realizes dynamically adjusting the sharding scale and redundancy ratio according to the sensitivity level of the file content, and optimizes the efficiency and security of distributed storage.

[0029] The classification-aware storage mechanism introduced in S3 is as follows:

[0030] S31. For each storage node N k The performance metrics include storage capacity C k , network bandwidth B k , current load L k ; and node security level S k which is scored based on the physical security of the node, the network environment, and the historical data leakage records; setting S k ∈[0,1], the comprehensive priority of each node is scored as:

[0031]

[0032] where w 1 , w 2 , w 3 , w 4 are weight parameters satisfying ∑w i = 1;

[0033] S32. For each shard i and redundant block of each electronic file D , based on the security level L i of the file and the node comprehensive priority P k , a classification-aware allocation is performed; the objective function is to maximize the storage security and performance of the sharded data, and the formula is where W ij represents the sensitivity weight of shard j:

[0034]

[0035] The distributed data storage strategy uses a classification-aware allocation function to select the best node N k for each shard and redundant block.

[0036] The scenario-based permission control system combined in S4 is as follows:

[0037] S41. Define the user permission P u as the set of permissions for the accessing user u, that is, P u = {(R u , A u , C u )}, where the resource scope R u is the set of electronic files that the user can access, the operation permission A u is the set of operations that the user can perform on the files, and the scenario constraint C u is the access restriction condition under a specific scenario;

[0038] Each file D i According to its security level L i and encryption algorithm Generate a unique key K i , The key distribution needs to meet the following conditions: The key K i Is dynamically distributed through the user request after permission verification; After the user obtains the key, decrypt the file Among them, C i Is the encrypted file, Is the corresponding decryption algorithm;

[0039] S42, According to the user's current access scenario C current , Dynamically adjust the permission constraint condition C u ; All file access requests are logged L ={(u, D i , A, t, e)}, Containing the user ID u, the accessed file D i , The operation type A, the timestamp t, and the device information e; Perform anomaly detection on the log, and identify potential abnormal access behaviors by analyzing indicators such as access frequency, geographical location, and device changes. The anomaly score calculation formula:

[0040] AnomalyScore = α·f(F freq ) + β·g(G u , G prev ) + γ·h(e u , e prev )

[0041] Among them, α, β, γ are weights, F freq Is the change in access frequency, and g and h are the change detection functions for geography and devices respectively.

[0042] The key dynamic update mechanism in S4 is as follows:

[0043] S43, Each electronic file D i According to its security level L i and encryption algorithm Generate an initial key The key generation formula is:

[0044]

[0045] Among them, H(·) is a secure hash function, t 0 Is the key generation timestamp, and the initial key is assigned to authorized users. The key distribution rule refers to the user permission verification mechanism;

[0046] S44, In the key update mechanism, the time interval exceeds the preset time T uWhen triggered, a periodic update occurs; user permission P u When it changes, immediately update the key of the relevant file; when a security threat is detected, force the key update; when the key update is triggered, generate a new key where n represents the number of updates; the key update formula is where is the previous key, t n is the current key update timestamp, L i is the security level of the file, ensuring that the key update is bound to the security attributes of the file;

[0047] After updating the key distribute the new key to all authorized users, and at the same time revoke the usage permission of the old key; to support the version control of the key, record the historical versions of the key updates and attach the timestamp t v and the update reason R v to each version, and the key version is

[0048]

[0049] Support rollback to a specific key version in case of data recovery or security events To improve the anti-cracking ability of the key, introduce a random factor r n when updating the key, so that the key update formula becomes:

[0050] A multi-node collaborative fault tolerance and recovery scheme in S5 is as follows:

[0051] S51, each storage node N k periodically reports its health status, including storage availability S k the remaining storage capacity of the node; network connection quality Q k network latency and bandwidth stability, failure rate F k the historical failure records of the node; data integrity check I k ensuring that the stored data has not been tampered with through hash verification, and the node health index H k through weighted calculation:

[0052]

[0053] When H k <H min the node is marked as abnormal;

[0054] S52, when monitoring discovers that node N k is abnormal or fails, trigger the fault tolerance mechanism, and mark the data shards and redundant blocks stored on the failed node {D ij ,Rim} is unavailable, notify other nodes to enter the collaborative recovery mode; for shard data recovery, use the distributed redundancy recovery algorithm to reconstruct the failed data, and the required number of complete shards for recovery is N r ; combine the remaining shards {D′ ij} and the redundant blocks {R′ im} to perform recovery through erasure coding:

[0055] D ij = f -1 ({D ij ′},{R im ′})

[0056] where f -1 is the inverse operation of erasure coding; the recovered shards are reallocated to healthy nodes N k′ , and when selecting the target node, give priority to considering its health index H k′ and the current load L k′ ,

[0057]

[0058] where α is the load balancing coefficient and Healthy Nodes are healthy nodes.

[0059] Compared with the prior art, the beneficial effects of the present invention are:

[0060] By classifying and encrypting according to the content sensitivity of electronic files, the present invention optimizes the storage efficiency while ensuring data security. Compared with the traditional unified encryption method, the present invention can select different encryption strengths according to the security requirements of different files, effectively avoiding the problems of over-encryption and insufficient encryption strength, and improving the flexibility and security of storage.

[0061] The present invention proposes an adaptive sharding and redundancy generation strategy, which dynamically adjusts the number of shards and the redundancy degree according to the security level of the file and the state of the storage node, enhancing the reliability and fault tolerance of the storage system. Different from the fixed sharding or redundancy strategy in the prior art, the present invention can optimize resource allocation according to the actual situation, avoiding the problems of resource waste or uneven storage.

[0062] The present invention introduces a multi-node collaborative fault tolerance and recovery scheme, which can quickly recover data through collaborative work when a node fails or malfunctions, ensuring the high availability of electronic files. Different from the traditional single-node or simple redundancy recovery mechanism, the present invention dynamically adjusts the recovery path and data allocation through the cooperation between nodes, significantly improving the fault tolerance of the system and the data recovery speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary. For those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained by extending the provided drawings.

[0064] The structures, ratios, sizes, etc. illustrated in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have technical substantive significance. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0065] Figure 1 It is a flowchart of an embodiment of the present invention. Specific Embodiments

[0066] To make the purposes, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. These descriptions are only to further illustrate the features and advantages of the present invention, rather than a limitation on the claims of the present invention; based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.

[0067] The following will further describe in detail the specific embodiments of the present invention in combination with the drawings and embodiments. The following embodiments are used to illustrate the present invention, but not to limit the scope of the present invention.

[0068] As Figure 1 shown, the framework of the present invention is mainly divided into the following five steps, which are connected layer by layer and finally integrated. The process mainly includes the following steps:

[0069] S1. Dynamically classify electronic files based on content sensitivity analysis and hierarchically define security levels; introduce a sensitivity-driven encryption algorithm strategy, dynamically select an appropriate AES (Advanced Encryption Standard) symmetric encryption algorithm, ensure that sensitive files are preferentially encrypted with high strength, and improve the security and performance compatibility of the system;

[0070] The method for dynamically classifying electronic files and hierarchically defining security levels in S1 is as follows:

[0071] S11. Feature extraction of file content: For each electronic file D i perform content analysis to extract its feature vector V i =[v i1 , v i2 ,..., v in , where v ij represents the value of the i-th file on the j-th feature, n is the number of features, and the features include keyword frequency, theme distribution, and the number of occurrences of sensitive words;

[0072] S12. Calculate the content sensitivity score of each file where w j is the weight of the j-th feature, reflecting the influence degree of this feature on sensitivity; the weight w j is determined by expert experience and machine learning models; according to the sensitivity score S i divide the files into different security levels L i , set multiple sensitivity thresholds T i , and the security level L i is defined as:

[0073]

[0074] For each security level L i allocate the corresponding encryption algorithm E i and key length K i , to form a mapping relationship E: L i →(E i , K i ), L 1 uses AES-128-bit encryption, L 2 uses AES-192-bit encryption, L 3 uses AES-256-bit encryption.

[0075] S2. An adaptive sharding and redundancy generation strategy that dynamically adjusts the sharding scale and redundancy level according to the access frequency and importance of file data; can reduce storage redundancy, optimize storage space utilization while ensuring data security, and update sharding allocation in real time according to node health status;

[0076] The method of an adaptive sharding and redundancy generation strategy in the above S2 is as follows:

[0077] S21. According to the size i of the electronic file D and the preset upper limit of sharding size S max , determine the number of shards where, represents rounding up to ensure that the size of each shard does not exceed S max; Archive D i is split into N i data blocks

[0078] For the sensitivity level L of each archive i , dynamically adjust the redundancy ratio R i , and generate the number of additional redundant blocks where R i is set according to the sensitivity level L i , R 1 = 0.5, R 2 = 1.0, R 3 = 2.0; Use erasure codes to generate redundant blocks where f(·) represents the erasure code function;

[0079] S22, The sharding and redundancy allocation strategy distributes shards and redundant blocks to different storage nodes N k Among them, the uniform distribution evenly distributes the original shards and redundant blocks to K nodes, and defines the shard allocation function g(D ij ,R im ):

[0080]

[0081] where Load(N k ) represents the node load, Risk(N k ) represents the node security risk, and α is the trade-off coefficient;

[0082] S23, After completing the sharding and redundancy allocation, the storage nodes perform integrity verification on the allocated data blocks D ij and redundant blocks R im to generate the hash value H ij = Hash(D ij ), H im = Hash(R im ); When a certain node N k fails, use the remaining shards and redundant blocks to recover the lost data; Assume the lost data is {D ij ,R im}, and the recovery formula is:

[0083] {D ij} = f -1 (R im ,D ij′ )

[0084] where f -1 is the inverse operation of the erasure code, D ij′For the remaining complete shard data; it realizes dynamically adjusting the shard scale and redundancy ratio according to the sensitivity level of the file content, and optimizes the efficiency and security of distributed storage.

[0085] S3. In a distributed network, introduce a classification-aware storage mechanism. Considering data categories, node performance, and network topology, allocate shard data and redundant blocks to different nodes; different from traditional random or uniform allocation methods, by optimizing the data distribution strategy, improve the system's access speed and overall reliability;

[0086] The classification-aware storage mechanism introduced in the above S3 is as follows:

[0087] S31. Each storage node N k The performance metrics include storage capacity C k , network bandwidth B k , current load L k ; node security level S k Score by evaluating the physical security of the node, network environment, and historical data leakage records; set S k ∈[0,1], and score the comprehensive priority of each node:

[0088]

[0089] where w 1 , w 2 , w 3 , w 4 are weight parameters, satisfying ∑w i =1;

[0090] S32. For each shard i and redundant block of each electronic file D based on the security level L i of the file and the node comprehensive priority P k , perform classification-aware allocation; the objective function is to maximize the storage security and performance of shard data, and the formula is where, W ij represents the sensitivity weight of shard j:

[0091]

[0092] The distributed data storage strategy uses a classification-aware allocation function to select the best node N k for each shard and redundant block.

[0093] S4. The scenario-based permission control system dynamically adjusts the permission verification and key distribution policies according to user permissions, data sensitivity levels, and access scenarios. Through the key dynamic update mechanism, it overcomes the problems of inconsistent permission synchronization and control in traditional distributed systems, enhancing the controllability of user access behavior and the privacy protection ability of data.

[0094] The scenario-based permission control system in S4 is as follows:

[0095] S41. Define user permissions P u as the permission set for accessing user u, that is, P u = {(R u , A u , C u )}, where the resource scope R u is the set of electronic files accessible to the user, the operation permission A u is the set of operations that the user can perform on the files, and the scenario constraint C u is the access restriction conditions under specific scenarios.

[0096] For each file D i , generate a unique key K i according to its security level L and the encryption algorithm i . The key distribution needs to meet the following conditions: The key K i is dynamically distributed through the user request after permission verification. After the user obtains the key, decrypt the file . Among them, C i is the encrypted file, is the corresponding decryption algorithm.

[0097] S42. Dynamically adjust the permission constraint condition C current according to the user's current access scenario C u . All file access requests are logged in L = {(u, D i , A, t, e)}, including the user ID u, the accessed file D i , the operation type A, the timestamp t, and the device information e. Perform anomaly detection on the log. By analyzing indicators such as access frequency, geographical location, and device changes, identify potential abnormal access behaviors. The anomaly score calculation formula is:

[0098] AnomalyScore = α·f(F freq ) + β·g(G u , G prev ) + γ·h(e u , e prev )

[0099] Among them, α, β, γ are weights, F freqFor the change in access frequency, g and h are the change detection functions for geography and device respectively.

[0100] The key dynamic update mechanism in S4 is as follows:

[0101] S43, for each electronic file D i According to its security level L i and the encryption algorithm generate the initial key The key generation formula is:

[0102]

[0103] where H(·) is a secure hash function, t 0 is the key generation timestamp, and the initial key is assigned to authorized users. The key distribution rule refers to the user permission verification mechanism;

[0104] S44, in the key update mechanism, when the time interval exceeds the preset time T u it triggers regular update; when the user permission P u changes, the keys of relevant files are updated immediately; when a security threat is detected, the keys are forced to be updated; when key update is triggered, new keys are generated where n represents the number of updates; the key update formula is where is the previous key, t n is the current key update timestamp, L i is the security level of the file, ensuring that key update is bound to the security attributes of the file;

[0105] After updating the key the new key is distributed to all authorized users, and at the same time the usage permission of the old key is revoked; to support version control of keys, the historical versions of key updates are recorded and a timestamp t v and the update reason R v are attached to each version. The key version is

[0106]

[0107] When data recovery or security events occur, it supports rolling back to a specific key version To improve the anti - cracking ability of keys, a random factor r n is introduced when updating keys, making the key update formula become:

[0108] S5. A fault tolerance and recovery scheme for multi-node collaboration, which combines redundant sharding and distributed consistency protocols to dynamically optimize the data recovery path during the data recovery process of faulty nodes; adjusts the recovery strategy according to the real-time status of nodes to ensure that the system has fault tolerance in scenarios of load or node failure.

[0109] The fault tolerance and recovery scheme for multi-node collaboration in S5 is as follows:

[0110] S51. Each storage node N k periodically reports its health status, including storage availability S k the remaining storage capacity of the node; network connection quality Q k network latency and bandwidth stability, failure rate F k the historical fault records of the node; data integrity check I k ensures that the stored data has not been tampered with through hash verification, and the node health index H k is calculated through weighted calculation:

[0111]

[0112] When H k <H min the node is marked as abnormal;

[0113] S52. When monitoring discovers that node N k is abnormal or fails, it triggers the fault tolerance mechanism, marks the data shards and redundant blocks {D ij , R im} stored on the failed node as unavailable, and notifies other nodes to enter the collaborative recovery mode; the sharded data recovery uses the distributed redundant recovery algorithm to reconstruct the failed data, and the number of complete shards N r required for recovery; combines the remaining shards {D′ ij} and redundant blocks {R′ im} to perform recovery through erasure coding:

[0114] D ij = f -1 ({D ij ′}, {R im ′})

[0115] where f -1 is the inverse operation of erasure coding; the recovered shards are reallocated to healthy nodes N k′ , and when selecting the target node, the health index H k′ and the current load L k′ are given priority,

[0116]

[0117] where α is the load balancing coefficient and Healthy Nodes are healthy nodes.

[0118] This invention patent mainly focuses on the innovation of the distributed storage method of content-sensitivity classification encryption and sharding redundancy for electronic files. Therefore, a series of experiments are needed to verify its effect in practical applications. The following is a description of the datasets used, comparison methods, and experimental results for this invention:

[0119] This invention uses the following two types of datasets for experimental verification: (1) Electronic file datasets: including medical records, financial documents, government documents, etc. These datasets have different content sensitivities and are suitable for testing the effect of content-sensitivity classification encryption algorithms. The content sensitivity of each file is marked according to its privacy level (such as personal information, financial data, etc.) to facilitate the implementation of classification encryption. (2) Distributed storage environment datasets: simulate the storage node status in a distributed storage environment, including node storage capacity, bandwidth, latency, failure rate, etc. This dataset is used to evaluate the effect of the dynamic sharding and redundancy generation strategy and the multi-node collaborative fault tolerance and recovery mechanism of this invention.

[0120] To verify the innovation and superiority of this invention, we selected the following traditional methods for comparative experiments: (1) Traditional centralized encryption method: uses symmetric encryption or asymmetric encryption methods for unified encryption processing, lacking optimization for content sensitivity. (2) Fixed sharding redundancy strategy: uses a fixed number of shards and redundant blocks for storage and cannot be dynamically adjusted. (3) Traditional distributed storage system: uses traditional fault tolerance mechanisms (such as replica replication or simple erasure codes) for fault tolerance recovery but does not involve dynamic node collaborative recovery. The comparison metrics include: storage efficiency, encryption strength, recovery speed, data integrity, system fault tolerance, and security. The specific experimental results are shown in Table 1.

[0121] Table 1 Comparison of experimental results of different methods

[0122]

[0123] In the comparison of experimental results, the storage efficiency of this invention is significantly better than traditional methods. The traditional centralized encryption method and the fixed sharding redundancy strategy do not consider the sensitivity of data content, resulting in too high encryption strength and more waste of storage space. While this invention classifies and encrypts according to the content sensitivity of the files, effectively avoiding over-encryption and saving about 30% of the storage space. This flexible encryption strategy not only improves storage efficiency but also ensures data security.

[0124] In terms of data recovery speed, the multi-node collaborative recovery mechanism of the present invention shows strong advantages. In the case of multiple node failures, the recovery time of traditional methods is relatively long, while the present invention reduces the data recovery time by approximately 40% through dynamic sharding redundancy and node collaborative work. This optimization not only improves the system's recovery speed but also significantly reduces the time for fault recovery, enhancing the availability of data.

[0125] Finally, in the comparison of fault tolerance and data loss rate, the present invention performs even better. Traditional distributed storage systems may have a risk of data loss when nodes fail, while the present invention achieves the goal of zero data loss through an adaptive sharding redundancy strategy and a multi-node collaborative mechanism. In the case of multiple node failures, the present invention can ensure 100% data recovery, while the recovery ability of traditional methods is weak, and the data loss rate can reach 2%. This advantage makes the present invention superior to the prior art in terms of high availability and fault tolerance. Experimental results show that the present invention has significant advantages over traditional methods in terms of storage efficiency, data recovery speed, and fault tolerance. By combining a content-sensitivity encryption strategy, a dynamic sharding and redundancy strategy, and a multi-node collaborative fault tolerance and recovery mechanism, the present invention can optimize storage resources while improving data security and privacy protection, enhancing the high availability and recovery ability of the system, and having broad practical application value.

[0126] The above only elaborates in detail on the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention, and all such changes should be included within the protection scope of the present invention.

Claims

1. A distributed storage method for electronic archives based on content sensitivity classification encryption and sharding redundancy, characterized in that: The following steps are involved: S1. Dynamically classify electronic files based on content sensitivity analysis and define security levels in a hierarchical manner; introduce a sensitivity-driven encryption algorithm strategy, dynamically select an adaptive AES (Advanced Encryption Standard) symmetric encryption algorithm, ensure that sensitive files are given priority to high-intensity encryption, and improve the security and performance compatibility of the system; S2, an adaptive sharding and redundancy generation strategy, dynamically adjusts the sharding scale and redundancy level according to the access frequency and importance of archive data; it can reduce storage redundancy, optimize storage space utilization, and update sharding allocation in real time according to node health status while ensuring data security; S3. In distributed networks, a classification-aware storage mechanism is introduced to allocate shard data and redundant blocks to different nodes, taking into account data categories, node performance, and network topology. Different from traditional random or uniform allocation methods, this method improves the access speed and overall reliability of the system by optimizing data distribution strategies. S4. Combined with the scenario-based permission control system, dynamically adjust the permission verification and key distribution strategy according to user permissions, data sensitivity level and access scenario; Through the dynamic key update mechanism, the problem of inconsistent permission synchronization and control in traditional distributed systems is overcome, and the controllability of user access behavior and the privacy protection of data are enhanced; S5. A multi-node collaborative fault tolerance and recovery solution that combines redundant sharding and distributed consistency protocols to dynamically optimize the data recovery path during the data recovery process of failed nodes; adjust the recovery strategy according to the real-time status of the node to ensure that the system has fault tolerance under load or node failure scenarios.

2. According to claim 1, a distributed storage method for electronic archives based on content sensitivity classification encryption and sharding redundancy is characterized in that: The method for dynamically classifying electronic archives and defining security levels in a hierarchical manner in S1 is as follows: S11, archive content feature extraction for each electronic archive D i Perform content analysis and extract its feature vector V i =[v i1 ,v i2 ,...,v in ], where v ij It represents the value of the i-th file on the j-th feature, n is the number of features, and the features include keyword frequency, topic distribution, and the number of occurrences of sensitive words; Calculate the content sensitivity score for each archive Among them, w j is the weight of the jth feature, reflecting the influence of the feature on sensitivity; weight w j Determined by expert experience and machine learning models; based on sensitivity score S i Divide files into different security levels i , set multiple sensitivity thresholds T i , safety level L i Defined as: S12, for each safety level L i Assign the corresponding encryption algorithm E i and key length K i , forming a mapping relationship E:L i →(E i ,K i ), L1 uses AES-128 bit encryption, L2 uses AES-192 bit encryption, and L3 uses AES-256 bit encryption.

3. According to claim 1, a distributed storage method for electronic archives based on content sensitivity classification encryption and sharding redundancy is characterized in that: An adaptive sharding and redundancy generation strategy method in S2 is: S21, according to electronic file D i Size and the preset upper limit of the shard size S max , determine the number of shards in, Indicates rounding up to ensure that the size of each shard does not exceed S max ; File D i Divided into N i Data blocks The sensitivity level L for each file i , dynamically adjust the redundancy ratio R i , generating additional redundant blocks Among them, R i The value of L depends on the sensitivity level i Set, R1 = 0.5, R2 = 1.0, R3 = 2.0; use erasure code to generate redundant blocks Where f(·) represents the erasure code function; S22, the sharding and redundancy allocation strategy allocates shards and redundant blocks to different storage nodes N k In the uniform distribution, the original shards and redundant blocks are evenly distributed to K nodes, and the shard allocation function g(D ij ,R im ): Among them, Load(N k ) represents the node load, Risk(N k ) represents the node security risk, and α is the trade-off coefficient; S23, after completing the sharding and redundancy allocation, the storage node allocates the data block D ij and redundant block R im Perform integrity verification and generate a hash value H ij =Hash(D ij ), H im =Hash(R im ); When a node N k When failure occurs, the remaining shards and redundant blocks are used to recover the lost data; assuming that the lost data is {D ij ,R im }, the recovery formula is: {D ij }=f -1 (R im ,D ij′ ) Among them, f -1 is the inverse operation of erasure coding, D ij′ For the remaining complete shard data; it realizes the dynamic adjustment of shard size and redundancy ratio according to the sensitivity level of the archive content, and optimizes the efficiency and security of distributed storage.

4. According to claim 1, a method for distributed storage of electronic archives based on content sensitivity classification encryption and sharding redundancy is characterized in that: The classification-aware storage mechanism introduced in S3 is: S31, each storage node N k Performance indicators include storage capacity C k 、Network bandwidth B k 、Current load L k ;Node security level S k By scoring the node's physical security, network environment, and historical data leakage records; setting S k ∈[0,1], the comprehensive priority of each node is scored: Among them w1,w2,w3,w4 are weight parameters, satisfying ∑w i =1; S32, for each electronic file D i Shard and redundant blocks File-based security level L i and node comprehensive priority P k , perform classification-aware allocation; the objective function is to maximize the storage security and performance of sharded data, the formula is Among them, W ij Represents the sensitivity weight of shard j: Distributed data storage strategy using classification-aware allocation function Select the best nodes N for each shard and redundant chunk k .

5. According to claim 1, a distributed storage method for electronic archives based on content sensitivity classification encryption and sharding redundancy is characterized in that: The scenario-based permission control system in S4 is: S41, define user rights P u is the permission set for accessing user u, that is, P u ={(R u ,A u ,C u )}, where the resource range R u A collection of electronic archives accessible to the user, operation authority A u The set of operations that users can perform on archives, scenario constraints C u Access restrictions in specific scenarios; Each file D i According to its safety level L i and encryption algorithm Generate a unique key K i , key distribution must meet the following conditions: key K i Dynamic distribution of user requests after permission verification; after the user obtains the key, the file is decrypted Among them, C i It is an encrypted file. is the corresponding decryption algorithm; S42, based on the user's current access scenario C current , dynamically adjust permission constraints C u ; All file access requests are logged L = {(u,D i ,A,t,e)}, including user IDu, access profile D i , operation type A, timestamp t, device information e; perform anomaly detection on the logs, identify potential abnormal access behaviors by analyzing indicators such as access frequency, geographic location, and device changes, and calculate the anomaly score: AnomalyScore=α·f(F freq )+β·g(G u ,G prev )+γ·h(e u ,e prev ) Among them, α, β, γ are weights, F freq is the access frequency change, g and h are the change detection functions for geography and device, respectively.

6. According to claim 1, a method for distributed storage of electronic archives based on content sensitivity classification encryption and sharding redundancy is characterized in that: The key dynamic update mechanism in S4 is: S43, each electronic file D i According to its safety level L i and encryption algorithm Generate initial key The key generation formula is: Where H(·) is a secure hash function, t0 is the key generation timestamp, the initial key is assigned to the authorized user, and the key distribution rule refers to the user authority verification mechanism; S44, in the key update mechanism, the time interval exceeds the preset time T u When, trigger regular updates; User Permissions u When changes occur, update the keys of related archives immediately; When a security threat is detected, force a key update; when a key update is triggered, generate a new key Where n represents the number of updates; the key update formula is in is the last key, t n Update timestamp for current key, L i For the security level of the archive, ensure that the key update is bound to the archive security attributes; Updating the key After that, the new key is distributed to all authorized users, and the use rights of the old key are revoked; to support key version control, the historical version of the key update is recorded And add a timestamp t to each version v and update reason R v , the key version is Supports rollback to a specific key version in case of data recovery or security incidents In order to improve the key's anti-cracking ability, a random factor r is introduced when updating the key. n , so that the key update formula becomes:

7. According to claim 1, a method for distributed storage of electronic archives based on content sensitivity classification encryption and sharding redundancy is characterized in that: A multi-node coordinated fault tolerance and recovery solution in S5 is: S51, each storage node N k Periodically reports its health status, including storage availability S k Node remaining storage capacity; network connection quality Q k Network latency and bandwidth stability, failure rate F k Node historical fault records; data integrity check I k Hash verification is used to ensure that the stored data has not been tampered with. The node health index H k By weighted calculation: When H k <H min When , the node is marked as abnormal; S52, when the monitoring finds that node N k When an exception or failure occurs, the fault tolerance mechanism is triggered to mark the data fragments and redundant blocks stored on the failed node. ij ,R im } is unavailable, notifying other nodes to enter collaborative recovery mode; shard data recovery uses a distributed redundant recovery algorithm to rebuild the failed data and restore the required number of complete shards N r ; Combine the remaining fragments {D′ ij } and redundant blocks {R′ im }Recovery through erasure coding: D ij =f -1 ({D ij ′},{R im ′}) Among them, f -1 It is the inverse operation of erasure coding; the recovered shards are redistributed to healthy nodes N k′ , when selecting the target node, give priority to its health index H k′ and the current load L k′ , Where α is the load balancing factor and Healthy Nodes are healthy nodes.

Citation Information

Cited By

  • Data security storage protection device and system

    CN120744958A

  • Data sharing method and system for intelligent network connection road

    CN121585399A