Data security processing method and system based on distributed storage

Through dynamic sharding strategies and intelligent threat perception mechanisms, the problems of static key management and cross-table association risks in distributed storage solutions are solved, efficient data security protection and real-time attack blocking are achieved, and the security and performance of the system are improved.

CN120654250AActive Publication Date: 2025-09-16WUHAN ANYU INFORMATION SECURITY TECH CO LTD

Patent Information

Application Number
CN202510744880.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-16
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Existing distributed storage solutions have security issues such as static key management, poor shard security isolation, high cross-table association risks, and delayed threat response.

Method used

Dynamic sharding strategy, multi-dimensional verification mechanism, cross-group key cross-binding and intelligent threat perception are adopted to achieve security protection and real-time attack blocking of data throughout its entire life cycle. Data encryption keys are generated through dynamic sharding strategy, dynamic threshold values ​​are generated by combining multi-dimensional indicators, and an intelligent hazard perception engine is used to identify attack nodes and conduct joint defense response.

Benefits of technology

It effectively reduces the risk of single point failure, improves the attack blocking rate, reduces decryption delay and improves throughput, thereby improving resource utilization and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654250A_ABST
    Figure CN120654250A_ABST
Patent Text Reader

Abstract

The invention relates to a data security processing method and system based on distributed storage, and relates to the technical field of computer information processing. The method comprises the following steps: cutting data into encryption fragments with a configurable number by adopting a dynamic fragmentation strategy, and generating a physically isolated dynamic check block in combination with a timestamp to realize tampering prevention; a dynamic threshold value is dynamically calculated based on the data sensitivity index and the node load, and the node is optimized through the reliability score for cooperative decryption; a database table is divided into independent marshalling storage according to main foreign key association, foreign key fields are encrypted by adopting cross keys, and cross-marshalling access needs to meet a multi-key threshold condition; an intelligent threat perception engine is constructed, access logs and threat intelligence are analyzed in real time, and key rotation, fragment replacement and joint defense response mechanisms are dynamically triggered. According to the invention, full life cycle protection of data is realized, and the problems of key leakage risk and cross-table association attack are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer information processing, and in particular to a data security processing method and system based on distributed storage. Background Art

[0002] With the rapid development of information technologies such as cloud computing, the Internet of Things, and artificial intelligence, as well as the digital transformation of traditional industries, global data volumes are growing exponentially. Traditional relational databases, limited by their centralized architecture and fixed storage capacity, are no longer able to meet the demands of high-concurrency, high-throughput, and massive data storage. Against this backdrop, big data storage technologies based on distributed architectures have emerged. These technologies, which achieve data sharding through horizontally scaling node clusters, significantly improve system scalability and disaster recovery capabilities. However, existing distributed storage solutions suffer from serious security deficiencies.

[0003] The Chinese invention patent with publication number CN115422570B provides a distributed storage data processing method and system, which receives a decryption request for a target ciphertext data encryption key sent by a distributed storage client; Obtain the first encryption zone key of the target ciphertext data encryption key from the relational key library; decrypt the target ciphertext data encryption key according to the first encryption zone key to obtain a decrypted data encryption key, and feed the decrypted data encryption key back to the client, so that the encryption and decryption module of the client calls the target encryption algorithm through the encryption engine according to the data encryption key to encrypt the data file to obtain an encrypted data file, and stores the encrypted data file in the data node to improve security.

[0004] However, the over-reliance on a fixed encryption zone key and data encryption key layering mechanism lacks dynamic response capabilities. Attackers can steal shards over a long period of time or exploit node vulnerabilities to crack the key, leading to the spread of single point failure risks. Summary of the Invention

[0005] This invention addresses existing distributed storage solutions, including static key management, poor shard security isolation, high cross-table association risks, and delayed threat response. By leveraging dynamic sharding strategies, multi-dimensional validation mechanisms, cross-group key cross-binding, and intelligent threat sensing, it achieves full data lifecycle security protection and real-time attack blocking.

[0006] This application provides a data security processing method and system based on distributed storage, the method comprising: S1. Shard the data using a dynamic sharding strategy based on the original data. Each shard independently generates a data encryption key and generates a dynamic check block based on the timestamp and shard content. S2. Obtain the data sensitivity index and node load index based on multi-dimensional indicators and generate dynamic threshold values; The multi-dimensional indicators include the privacy level of the original data, compliance requirements, data value, and the server's CPU usage, memory usage, and network latency; a node refers to an independent storage unit in a distributed system; S3. Divide the database table into independent groups based on primary and foreign key relationships. Each independent group is stored in a physically isolated node cluster. The foreign key fields between groups are encrypted using cross-key encryption technology. Nodes include but are not limited to regional data center nodes of cloud service providers, physical server clusters of local enterprise data centers, etc. S4. The group access log, external threat intelligence, and real-time traffic characteristics are used as data sources for training to obtain an intelligent hazard perception engine. The intelligent hazard perception engine identifies the attacked nodes in the group and sets the nodes associated with the data to automatic defense. A joint defense response is performed based on the intelligent hazard perception engine combined with the group joint defense strategy.

[0007] Furthermore, the number of shards and redundancy coefficients are dynamically configured based on the data sensitivity level; the dynamic check block uses HMAC-SM3 hash chain technology, the master key MEK is generated by the hardware security module, and the shard key is dynamically derived; the dynamic check block storage retains 30 days of historical versions, and mandatory operations are based on the latest version of the dynamic check block verification.

[0008] Furthermore, based on this architecture, the master key (MEK) is generated by the hardware security module (HSM), and the shard key (DEK) is generated by Dynamic derivation ensures that a single shard key leak does not affect other shards. Combining edge pre-decryption with parallel verification reduces decryption latency by 40% and increases throughput by 120%.

[0009] Furthermore, under the support of the dynamic key system, a dynamic threshold value formula is constructed based on the data sensitivity index (DSI) and the node load index (NLI): , Where t is the dynamic threshold value, Nodes is the total number of nodes in the storage shard; ,DSI is the data sensitivity index, 、 、 are all pre-set weight coefficients, and 、 、 The sum of is 1, P is the privacy level, C is the compliance level, and D is the network protocol; , NLI is the node load index, U and M are CPU occupancy and memory occupancy respectively, both in the range of [0,100], N is the real-time network delay in ms, and R is the benchmark delay; node reliability score , represents the reliability score of the i-th node, O is the online rate, Time is the total duration, and Num is the number of historical failures.

[0010] For example, when the node load index rises by 20%, the dynamic threshold automatically increases by 1-2 levels. During this process, nodes are sorted in descending order by their reliability scores, with high RScore nodes prioritized for threshold decryption. Combined with threshold signature verification of shard consistency, this reduces the success rate of tamper injection from 20% to 0.1%.

[0011] Furthermore, in order to block foreign key association attacks, the database table is divided into independent groups according to the primary and foreign key association relationships. Each group is stored in a physically isolated node cluster. The foreign key fields are Cross encryption and decryption must simultaneously meet multiple group key threshold conditions.

[0012] Furthermore, by defining the group safety weight formula: , a is the current grouping weight, b is the associated grouping weight, the sum of a+b is equal to 1, The privacy level of the current group can be set from 1 to 5 from low to high. is the sum of the j-associated group safety weights. If there is no association, the value is 0. The dynamic threshold value is adjusted based on the group safety weight. The adjustment formula is: , is the node load index of the current i-th node, and TN is the total number of nodes, which prevents the risk of cross-table association caused by the leakage of foreign key plaintext.

[0013] Furthermore, based on the above protection, by analyzing group access logs, external threat intelligence and traffic characteristics, it can identify brute force cracking, node hijacking and other attack behaviors in real time, with a threat identification accuracy rate of 99.9%. Automated joint defense response to achieve dynamic key rotation: cross-group key update is automatically triggered in the event of a leak (such as ) and foreign keys to ensure the consistency of ciphertext migration transactions and improve the protection level of associated grouping.

[0014] A data security processing system based on distributed storage, comprising: Multi-dimensional verification storage module, used to obtain data fragments and generate dynamic verification blocks; The intelligent hazard perception engine module is trained based on group access logs, external threat intelligence, and real-time traffic characteristics as data sources. When the corresponding threat type is detected, the engine automatically increases the dynamic threshold value of the relevant node; The group joint defense control module is used for node clusters. When a group is attacked, the other groups will conduct dynamic joint defense. When the group joint defense control module detects an attack, it generates corresponding threat intelligence and synchronizes it to the other nodes in the group through an encrypted channel, automatically updating their firewall rules. The geofencing module dynamically builds geofencing policies based on the geographic compliance requirements of the data.

[0015] One or more technical solutions provided in this application have at least the following technical effects or advantages: By combining dynamic sharding strategy with physical isolation verification mechanism, the problem of fixed sharding keys in existing technologies that easily lead to single point failure is solved; through dynamic threshold adjustment and cross-group cross-key encryption, the defect of traditional solutions that cannot block foreign key association attacks is overcome; based on the real-time joint defense response of the intelligent threat perception engine, the attack blocking rate is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a flow chart of a data security processing method based on distributed storage in an embodiment of the present invention; Figure 2 This is an example diagram of the e-commerce database of the present invention; Figure 3 This is a key distribution rule diagram of the present invention; Figure 4 This is an architecture diagram of a data security processing system based on distributed storage in an embodiment of the present invention; Figure 5 This is a flow chart of a data security processing system based on distributed storage in an embodiment of the present invention; DETAILED DESCRIPTION The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0017] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the described features. In the description of this application, "plurality" means two or more, unless otherwise specifically specified.

[0018] In the description of this application, the term "for example" is used to mean "used as an example, illustration or explanation". Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art will recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in this application.

[0019] Example 1: The present application provides a distributed storage data processing method, the method comprising: Based on the security requirements of different business scenarios, the system is divided into top secret: number of shards n=10, redundancy number r=2; sensitive: number of shards n=6, redundancy number r=3; and ordinary: number of shards n=3, redundancy number r=1 according to the data sensitivity level. When a storage node failure is detected, r is automatically increased, and the system copies the shards to the new node to ensure availability.

[0020] The original data is sharded and divided into n shards according to the data sensitivity level, where n ranges from [2,10]. An independent data encryption key (DEK) is generated for each shard and encrypted using the national secret SM4 or AES algorithm. The encrypted shards are distributed and stored in different nodes to obtain an encrypted shard set. ,in represents the i-th shard, and n is the total number of shards.

[0021] After completing the encrypted shard distributed storage, , calculate the dynamic check block of HMAC-SM3 based on the timestamp, the calculation formula is: , in, To obtain the dynamic verification hash value, i represents the i-th fragment. HMAC-SM3 is a hash message authentication code based on the national secret SM3 hash algorithm, which is used to generate a keyed hash value. Represents the i-th original data content, || represents the data splicing operator, which splices adjacent contents in sequence, and Timestamp is the precise timestamp when the shard is generated.

[0022] After obtaining the dynamic verification hash value, a dynamic verification block is generated and stored in an independent secure storage pool. The calculation formula is: ,in, is the dynamic check block of the i-th shard, is the dynamic check hash value of the i-th shard, and Version is the version number of the dynamic check block, thus obtaining the dynamic check block set ,in represents the i-th dynamic check block, and n is the total number of dynamic check blocks.

[0023] To meet data auditing and rollback protection requirements, a new dynamic checksum is generated for each shard update, and older versions are retained for 30 days. Audits can trace historical data status using version numbers. Each operation must be based on the latest version of the dynamic checksum; older versions are only used for auditing and cannot be used for data recovery. Version control and rollback protection are achieved through the storage of multiple versions of dynamic checksums.

[0024] The key system adopts a hierarchical protection design. The master key (MEK) is generated by the key management service and stored in the hardware security module; the sharded DEK is derived hierarchically: ; in, is the key of the i-th shard, To store the unique identifier of the node of the i-th shard, ensure that the key is bound to the node, the DEK of each shard is independent, and the leakage of a single shard key does not affect other shards.

[0025] When a client requests access to data, it obtains the shard from the storage node. , pull the latest dynamic check block from the security pool , recalculate the hash , Dynamic check block The timestamp when it was generated, if The shard is determined to have been tampered with, an alarm is triggered, and the node is isolated.

[0026] In an anti-attack environment, when an attacker tampers with the content of a shard, a dynamic check block is generated based on the timestamp and shard content. The hash value does not match after tampering, triggering a real-time alarm. The tampering detection rate is: 99.9%, and the response delay is <50ms. When the DEK of a shard is leaked, the attacker attempts to decrypt other shards. Because the shard keys are derived independently, the leaked DEK only affects a single shard. Therefore, combined with the dynamic check block, even if a single shard is decrypted, the tampering will still be detected, reducing the scope of impact by 80%.

[0027] The dynamic sharding strategy is composed of dynamic adjustment of the number of shards according to the data sensitivity level, key hierarchy derivation, dynamic check block generation and version number traceability mechanism.

[0028] For example, if an attacker obtains the subkey of a shard but still cannot decrypt the complete data, the dynamic check block can verify the integrity of the shard. If the shard is tampered with, the recovery process will automatically detect and reject it.

[0029] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages: Dynamic checksums completely neutralize data tampering and rollback attacks. Shard key isolation reduces the impact of key leaks by 80%. Dynamic sharding and redundancy strategies reduce storage costs by 35% and increase resource utilization by 40%. Existing technologies, however, rely on a fixed encryption zone key and data encryption key layering mechanism, lacking dynamic response capabilities. Attackers can exploit long-term sharding or node vulnerabilities to crack keys, leading to the spread of single point failure risks.

[0030] Example 2: In the above embodiment, data security is further improved through sharded encrypted storage and dynamic check blocks. However, although the risk of single point leakage is reduced through sharded key isolation, it cannot prevent attackers from launching coordinated attacks by forging nodes or tampering with transmission content. To address this problem, this embodiment further improves the above embodiment.

[0031] The privacy level, compliance requirements, and data value are used as the Data Sensitivity Index (DSI). Privacy levels are divided into 1 to 5 levels. Compliance requirements are divided into 3 levels based on the General Data Protection Regulation, Cybersecurity Level Protection 2.0, and no requirements. The values ​​are General Data Protection Regulation = 3, Cybersecurity Level Protection 2.0 = 2, and no requirements = 1. Data value is a quantified economic value, such as medical images per MB = 100 yuan. The formula is obtained:

[0032] Among them, DSI is the data sensitivity index, 、 、 are all pre-set weight coefficients, and 、 、 The sum of is 1, P is the privacy level, C is the compliance level, and D is the network protocol; The server's CPU usage, memory usage, and network latency are used as the node load index (NLI). The formula is:

[0033] Among them, NLI is the node load index, U and M are CPU usage and memory usage, both in the range of [0,100]. N is the real-time network delay in milliseconds. R is the benchmark delay, which uses different data according to different scenarios, such as 10ms for local and 100ms for cross-border.

[0034] Generate dynamic threshold values ​​based on data sensitivity index and node load index; The dynamic threshold calculation formula is:

[0035] Where t is the dynamic threshold value, and Nodes is the total number of nodes in the storage shard.

[0036] In view of the stability and credibility of different nodes, a node reliability score (RScore) is added, and the calculation formula is:

[0037] in, represents the reliability score of the i-th node, O is the online rate, Time is the total duration, and Num is the number of historical failures. Based on the descending reliability score, the first t nodes are selected to participate in decryption, and the rest are used as redundant backups.

[0038] Encrypting sharded collections by input , dynamic threshold value t, select t shards for decryption collaboration according to the priority based on the reliability score, each node uses an independent key to decrypt the shard and generate an intermediate result , nodes are based on Generate a partial signature , aggregate t Generate a full signature , verify whether it matches the original data hash. If the verification passes, the decryption is successful.

[0039] When an attacker forges multiple fake nodes to participate in collaborative decryption, the reliability score is dynamically downgraded. The forged node has no historical data or a high threat score, which makes the reliability score infinitely close to 0 and cannot be selected. Therefore, the fake shard cannot generate a valid , resulting in aggregation failure, thus effectively resisting attacks; when the attacker tampers with the content or results, according to the threshold signature scheme verification, any tampering of the fragment will result Verification failed. If a node fails verification three times in a row, it will be marked as high-risk and isolated, reducing the success rate of tampered data injection from 20% to 0.1%.

[0040] In some embodiments, a geo-fence strategy is dynamically constructed based on the geographic compliance requirements of the data to restrict the storage and processing nodes of the shards to be located within a specified geographic area; when a cross-border shard processing request is detected or the node location exceeds the geo-fence range, the shard replacement mechanism is automatically triggered to migrate the original shard to a compliant node cluster, and isolate the illegal node. At the same time, the dynamic check block and cross-key binding relationship are updated to ensure data integrity and regional compliance. For example, EU data is only processed by European nodes, and the dynamic threshold value is adjusted according to regional compliance requirements; the audit log records the entire process log of threshold adjustment, shard selection and decryption verification, and retains it for 180 days.

[0041] For example, a stock exchange processes 50,000 trade orders per second, requiring real-time decryption and risk control analysis to protect against attacks and tampering. The privacy level is 5, the compliance requirement is 2, and the data value is 1KB × 500 yuan / KB = 500 yuan. Substituting this into the Data Sensitivity Index (DSI) formula yields: , the node load index (NLI) is: , thus obtaining the dynamic threshold value: By dynamically reducing the threshold from 5 to 3 and combining it with GPU parallel computing, millisecond-level decryption is achieved. Verified by the threshold signature scheme, even if an attacker controls one node, they still cannot tamper with the transaction data.

[0042] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages: Through dynamic threshold adjustment, shard priority scheduling, and threshold signature verification, we address security concerns caused by coordinated attacks by attackers forging nodes or tampering with transmitted content. Dynamic thresholds are adjusted in real time based on data sensitivity and node load, achieving an optimal balance between security and efficiency.

[0043] Example 3: In the above embodiment, dynamic threshold sharding and collaborative decryption solve the security problem of attackers launching collaborative attacks by forging nodes or tampering with transmission content. However, only single-table data shards are encrypted, and the risk of cross-table foreign key associations is not solved. To address this problem, this embodiment further improves the above embodiment.

[0044] Group the primary and foreign keys, and divide the database tables into multiple groups according to business logic. Each group contains the primary table and the secondary tables directly related by foreign keys. Each group data is independently stored in a physically isolated node cluster, such as Figure 2 As shown, G1 is stored in node cluster A, and G2 is stored in cluster B; the foreign key fields between groups are stored encrypted.

[0045] On the basis of grouping the primary and foreign keys, the grouping keys are divided into shards and bound according to the key distribution rules. Figure 3, cross-key encryption is used for foreign keys between groups, and the UserID field of G2 is encrypted as follows: , K is the encrypted foreign key, It represents the composite key generated by calculating the key K1 of group G1 and the key K2 of group G2. Decryption requires both K1 and K2 to restore the foreign key plaintext.

[0046] When querying the user's order and payment record data, the client requests to decrypt the UserID of G1 and provides the administrator shard and auditor shard, thereby recovering K1 and decrypting the UserID plaintext:

[0047] Among them, SM4_Decrypt is the decryption function of the SM4 algorithm, K1 is the key of group G1, KNC_K1(UserID) represents the ciphertext after the primary key field UserID of the user table G1 is encrypted using the key K1. The decrypted plaintext is obtained. The user data is obtained; a one-time password token is generated based on the client, which is valid for 30 seconds. K2 is generated in conjunction with the administrator shard, and the G2 foreign key is decrypted. , ensure that G1 With G2 If a match is found, the order data is obtained; the auditor activates the K3 shard, verifies the access authorization records of G1 and G2, and dynamically generates a global key: , decrypt G3's OrderID association: , obtain the payment record data; use threshold signatures in the entire access process to verify the consistency of cross-group decryption results to prevent tampering by middlemen.

[0048] The group security weight is set for the group based on the data sensitivity. The formula is:

[0049] Among them, a is the current grouping weight, b is the associated grouping weight, and the sum of a+b is equal to 1. The privacy level of the current group i is set from 1 to 5 from low to high. It is the sum of the safety weights of associated groups. If there is no association, the value is 0.

[0050] Add the group to the dynamic threshold formula to obtain a new dynamic threshold formula:

[0051] Among them, NLI is the node load index of the current i-th node, and TN is the total number of nodes.

[0052] For example, after stealing the order table information, the attacker wants to obtain the user information, but because the foreign key UserID is used If the payment record table is encrypted and the G1 key K1 is not obtained, identity association cannot be achieved. If an insider obtains the payment record table and attempts to unauthorizedly access other data, the administrator and the client must generate a one-time password token to generate the global key, thus preventing unauthorized access. This provides a technological breakthrough for data correlation security protection, moving from single-point protection to global joint defense.

[0053] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages: Dynamic threshold control of primary and foreign key grouping prevents cross-table data relationships from being leaked through foreign key plaintext; access to cross-group data requires meeting threshold conditions of multiple groups, increasing the cost of cracking by attackers; data within a group is stored locally to reduce cross-node query overhead.

[0054] Example 4: In the above embodiment, dynamic threshold control of primary and foreign key grouping is used to prevent cross-table data relationships from being leaked through foreign key plain text; access to cross-group data requires meeting the threshold conditions of multiple groups, which increases the cost of cracking for attackers; data within the group is stored locally to reduce cross-node query overhead; however, crisis handling is often passive and retrospective. To address this problem, this embodiment further improves the above embodiment.

[0055] When the system detects a corresponding threat type, the engine automatically raises the dynamic threshold for the relevant group, blocks access to suspicious IP addresses, and triggers key rotation. The intelligent threat perception engine is trained using group access logs, external threat intelligence, and real-time traffic characteristics as input.

[0056] Based on the engine combined with the group joint defense strategy, it receives brute force information. When the G1 node is subjected to brute force, the dynamic threshold value of the associated group is temporarily increased and the foreign key encryption algorithm is upgraded. The dynamic threshold value is increased and the foreign key encryption is dynamically updated to SM4-256. When the G2 node is hijacked, the G2 high-risk node is isolated, and the associated shards in G1 and G3 are cleaned up synchronously. The blockchain smart contract is used to trigger the deletion of cross-group shards and generate new shards to store in the secure node. When the G3 key is leaked, K3 is automatically rotated and all associated foreign keys are updated, and regenerated. , traverse the G2 table to re-encrypt the foreign key. Because the foreign key is dynamically re-encrypted, the key is also rotated to generate a new key .

[0057] Based on the new key, the ciphertext is migrated and the integrity of the encryption switch is ensured through database transactions to avoid data inconsistency. The new ciphertext formula is: , Text is the new ciphertext, and FK is the corresponding foreign key plaintext.

[0058] Share threat intelligence across groups. When G1 detects a brute force attack, it generates corresponding threat intelligence. This report is synchronized to G2 and G3 through an encrypted channel, automatically updating their firewall rules. Assuming the probability of a single group independently intercepting an attack is P = 90%, the overall interception probability after joint defense is increased to: , greatly improving safety.

[0059] For example, in a cross-border e-commerce data protection scenario, user data is placed in the EU node with GSW=5, order data is placed in the Asian node with GSW=4, and payment data is placed in the North American node with GSW=5. When an attacker hijacks two nodes of G2 and attempts to crack the user identity associated with the order foreign key, the attacker first identifies the abnormal access pattern of the G2 node and determines that the threat level is assessed as 4. The G2 dynamic threshold value is upgraded from 4 to 5, requiring 5 shards to decrypt 0. The G1 foreign key encryption algorithm is upgraded from SM4-128 to SM4-256 to block association inference. G3 automatically rotates the K3 key and updates the payment record ciphertext from G2 to G3. Based on the above joint defense response, the attacker cannot obtain enough shards within 1 hour, and the attack fails. The system automatically generates new shards to replace the hijacked node, and the business recovery time is less than 5 minutes.

[0060] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages: Threat identification has changed from "post-hoc tracing" to "real-time blocking", and the attack success rate has been reduced from 0.1% to 0.001%; ​​when a single group is attacked, the protection level of the associated groups is automatically upgraded; key rotation and ciphertext migration are fully automated, and the recovery time is reduced from hours to minutes.

[0061] Example 5: Figure 4 As shown, the present application also provides a data security processing system based on distributed storage, comprising: The multi-dimensional checksum storage module is used to obtain data fragments and generate dynamic checksum blocks, including: Dynamic verification hash value generation unit generates dynamic verification hash value according to the national secret SM3 and the current timestamp. The master key (MEK) is generated by the HSM hardware security module, and the sharding key DEKi is generated by Dynamic derivation ensures that the key is strongly bound to the node and time; The dynamic check block management unit supports 30-day historical version tracing based on the version number. The dynamic check block is composed of a combination of version number, timestamp, and hash value. The historical version can be retrieved by the timestamp and the version number can be used to confirm the version.

[0062] The intelligent hazard perception engine module is trained based on group access logs, external threat intelligence, and real-time traffic characteristics as data sources. When a corresponding threat type is detected, the engine automatically increases the dynamic threshold value of the relevant node, including: Automatic detection unit: When the system detects the corresponding threat type, it automatically increases the dynamic threshold of the relevant group based on the engine, blocks the access rights of suspicious IP addresses, and triggers key rotation.

[0063] The group joint defense control module is used for physically isolated node clusters. When a group is attacked, the remaining groups will dynamically join forces to defend against it. When the group joint defense control module detects an attack, it generates corresponding threat intelligence and synchronizes it to the remaining nodes in the group through an encrypted channel, automatically updating their firewall rules, including: A data isolation storage unit, used to independently store grouping data in a physically isolated node cluster; The cross-key management unit is used to bind the grouping key fragments based on the grouping division of the primary and foreign keys, and the foreign keys between groups are encrypted with cross-keys; Adaptive threshold adjustment unit for real-time response to changes in threat levels; Transaction consistency assurance unit, used to ensure data integrity during key rotation or ciphertext migration.

[0064] The geofencing module dynamically builds geofencing policies based on the geographic compliance requirements of the data, including: Geographic fence units restrict the storage and processing nodes of shards to be located within a specified geographical area. When a cross-border shard processing request is detected or the node location is outside the geographic fence, the shard replacement mechanism is automatically triggered to migrate the original shard to a compliant node cluster, isolate the illegal node, and update the dynamic check block and cross-key binding relationship to ensure data integrity and regional compliance.

[0065] like Figure 5 As shown in the figure, the system workflow diagram is executed in the following order: Dynamic sharding strategy: Shard the original data file and cut it into n shards according to the dynamic sharding strategy.

[0066] Dynamic check block: Based on sharding, dynamic check blocks are calculated and stored in an independent secure storage pool.

[0067] Hierarchical protection design of the key system: the master key is generated by the key management service and stored in the hardware security module; the sharded DEK is derived hierarchically.

[0068] Data sensitivity index: a combination of privacy level, compliance requirements and data value.

[0069] Node load index: It is composed of privacy level, compliance level, and network protocol.

[0070] Dynamic threshold: Generates a dynamic threshold based on the data sensitivity index and node load index.

[0071] Reliability score: The stability and credibility of different nodes, used to sort nodes.

[0072] Grouping by primary and foreign keys: Divide the database table into multiple groups according to business logic. Each group contains the primary table and the secondary tables directly linked by foreign keys. The data of each group is independently stored in a physically isolated node cluster.

[0073] Group security weight: A combination of data sensitivity and group settings.

[0074] Intelligent Danger Perception Engine: An engine trained based on grouped access logs, external threat intelligence, and real-time traffic features as data sources.

[0075] Cross-group threat intelligence sharing: When one of the nodes in a group is attacked, the associated corresponding nodes will automatically defend themselves.

[0076] Geofencing strategies dynamically limit cross-border sharding: A strategy of adjusting dynamic threshold values ​​is adopted for different environments.

[0077] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0078] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0079] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0080] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0082] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0083] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A data security processing method based on distributed storage, characterized in that: include: S1. Shard the original data based on the dynamic sharding strategy, independently generate a data encryption key for each shard, and generate a dynamic check block based on the timestamp and shard content; S2. Obtain the data sensitivity index and node load index based on multi-dimensional indicators and generate dynamic threshold values; The multi-dimensional indicators include the privacy level of the original data, compliance requirements, data value, and the server's CPU usage, memory usage, and network latency; a node refers to an independent storage unit in a distributed system; S3. Divide the database table into independent groups based on primary and foreign key relationships. Each independent group is stored in a physically isolated node cluster. The foreign key fields between groups are encrypted using cross-key encryption technology. S4. The group access log, external threat intelligence, and real-time traffic characteristics are used as data sources for training to obtain an intelligent hazard perception engine. The intelligent hazard perception engine identifies the attacked nodes in the group and sets the nodes associated with the data to automatic defense. A joint defense response is performed based on the intelligent hazard perception engine combined with the group joint defense strategy.

2. A data security processing method based on distributed storage according to claim 1, characterized in that: The dynamic sharding strategy includes: dynamically configuring the number of shards and the redundancy coefficient based on the data sensitivity level; the dynamic check block adopts HMAC-SM3 hash chain technology, the master key MEK is generated by the hardware security module, and the sharding key is dynamically derived; the dynamic check block storage retains 30 days of historical versions, and mandatory operations are based on the latest version of the dynamic check block verification.

3. The data security processing method based on distributed storage according to claim 1, characterized in that: The dynamic threshold value calculation formula is: , Where t is the dynamic threshold value, Nodes is the total number of nodes in the storage shard; , DSI is the data sensitivity index. 、 、 are all pre-set weight coefficients, and 、 、 The sum of is 1, P is the privacy level, C is the compliance level, and D is the network protocol; , NLI is the node load index, U and M are CPU usage and memory usage respectively, both in the range of [0,100], N is the real-time network delay in milliseconds, and R is the baseline delay; Node reliability score, , represents the reliability score of the i-th node, O is the online rate, Time is the total duration, and Num is the number of historical failures.

4. The data security processing method based on distributed storage according to claim 1, characterized in that: The division of the database tables into independent groups according to the primary and foreign key association relationships includes: The primary and foreign key fields are encrypted using cross-key encryption, and the encryption formula is: , where K is the encrypted foreign key, It is a composite key generated by the operation of key K1 and key K2; the group security weight is calculated using , a is the current grouping weight, b is the associated grouping weight, the sum of a+b is equal to 1, The privacy level of the current group can be set from 1 to 5 from low to high. is the sum of the j-associated group safety weights. If there is no association, the value is 0. The dynamic threshold value is adjusted based on the group safety weight. The adjustment formula is: , is the node load index of the current i-th node, and TN is the total number of nodes.

5. The data security processing method based on distributed storage according to claim 1, characterized in that: The joint defense response includes: The dynamic key rotation mechanism automatically triggers cross-group key updates and foreign key re-encryption in the event of a leak, ensuring the consistency of ciphertext migration transactions and improving the protection level of associated groups.

6. A data security processing system based on distributed storage, implementing a data security processing method based on distributed storage according to any one of claims 1 to 5, characterized in that: include: Multi-dimensional verification storage module, used to obtain data fragments and generate dynamic verification blocks; The intelligent hazard perception engine module is trained based on group access logs, external threat intelligence, and real-time traffic characteristics as data sources. When the corresponding threat type is detected, the dynamic threshold value of the relevant node is automatically increased; The group joint defense control module is used to manage physically isolated node clusters. When a group is attacked, the remaining groups will dynamically conduct joint defense. When the group joint defense control module detects an attack, it generates corresponding threat intelligence and synchronizes it to the remaining nodes in the group through an encrypted channel, automatically updating the node firewall rules. The geofencing module dynamically builds geofencing policies based on the geographic compliance requirements of the data.

7. A data security processing system based on distributed storage according to claim 6, characterized in that: The multi-dimensional verification storage module includes: Dynamic verification hash value generation unit, generates dynamic verification hash value according to national secret SM3 and current timestamp; Dynamic check block management unit supports 30-day historical version tracing based on version number.

8. The data security processing system based on distributed storage according to claim 6, characterized in that: The intelligent hazard perception engine module includes: Automatic detection unit: When the system detects the corresponding threat type, it automatically increases the dynamic threshold of the relevant group based on the engine, blocks the access rights of suspicious IP addresses, and triggers key rotation.

9. The data security processing system based on distributed storage according to claim 6, characterized in that: The group joint defense control module includes: A data isolation storage unit, used to independently store grouping data in a physically isolated node cluster; The cross-key management unit is used to bind the grouping key fragments based on the grouping division of the primary and foreign keys, and the foreign keys between groups are encrypted with cross-keys; Adaptive threshold adjustment unit for real-time response to changes in threat levels; Transaction consistency assurance unit, used to ensure data integrity during key rotation or ciphertext migration.

10. The data security processing system based on distributed storage according to claim 6, characterized in that: Also includes: Geographic fence units restrict the storage and processing nodes of shards to be located within a specified geographical area. When a cross-border shard processing request is detected or the node location is outside the geographic fence, the shard replacement mechanism is automatically triggered to migrate the original shard to a compliant node cluster, isolate the illegal node, and update the dynamic check block and cross-key binding relationship to ensure data integrity and regional compliance.

Citation Information

Patent Citations

  • A distributed storage data processing method and system

    CN115422570B

  • Key value distributed balanced storage method based on programmable data plane

    CN115168346A

  • Data security management method and system based on cloud computing

    CN119475369A

  • Cloud storage protection system based on big data network

    CN119544247A

  • Method and system for placing a workload on one of a plurality of hosts

    US20180219899A1

Cited By

  • Data security protection method and system for civil aviation data center

    CN120951402A

  • A data security protection method and system for civil aviation data middleware

    CN120951402B

  • Battlefield information processing method based on multi-modal data fusion

    CN121389000A

  • A battlefield information processing method based on multi-modal data fusion

    CN121389000B

  • Knowledge vector library-oriented secure storage and efficient query method

    CN121412361A