A method and system for secure data interaction based on a distributed database
By introducing an identity authentication and authorization module, topology awareness and dynamic sharding unit, and a threshold authentication encryption algorithm based on a distributed pseudo-random function into the distributed database, the problem of insufficient response to node security status in traditional data security interaction is solved. Real-time sharding of sensitive data and dynamic adjustment of encryption strategies are realized, thereby improving the system's security and adaptive recovery capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI MOXIU INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-03-10
- Publication Date
- 2026-06-02
AI Technical Summary
In existing data security interaction methods and systems based on distributed databases, traditional data sharding strategies lack a response to the real-time security status of nodes, resulting in high-risk nodes storing highly sensitive data fragments and creating security risks. The system's encryption mechanism relies on centralized key management, which poses a single point of failure and key leakage risk, and it is difficult to support cross-node encrypted collaborative computing, causing data to be frequently decrypted during the interaction process, increasing security risks.
The system employs modules for identity authentication and authorization, distributed database management, data security interaction control, node communication and collaboration, and trusted proof of interaction behavior to achieve unified identity verification and real-time dynamic authorization of fine-grained data access permissions across distributed nodes. It performs adaptive data sharding through topology awareness and dynamic sharding units, utilizes a threshold authentication encryption algorithm based on distributed pseudo-random functions to achieve collaborative encryption and decryption with keys that never need to be reconstructed, and constructs a confused multi-hop communication route between heterogeneous nodes to enhance anti-interference capabilities.
It enables real-time response to node security status, dynamically adjusts data sharding and encryption strategies, enhances the confidentiality and integrity of data during interaction, reduces the risk of single point of failure and key leakage, and improves the system's adaptive recovery capability and security.
Smart Images

Figure CN122137639A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of security situation awareness, adaptive data sharding, and distributed data encryption technology, specifically a data security interaction method and system based on a distributed database. Background Technology
[0002] Security situation awareness technology is a dynamic monitoring mechanism that collects and analyzes security indicators of distributed nodes in real time. It aims to solve the problem of the disconnect between data sharding strategies and security requirements in traditional distributed databases. It enables the system to dynamically trigger shard reorganization based on dual-dimensional indicators of node security level and data sensitivity. It automatically splits highly sensitive data into logically related but physically isolated fragments and distributes them across high-security domain nodes, thereby transforming the physical distribution of data itself into a proactive security barrier rather than a source of risk. Adaptive data sharding technology is a dynamic resource scheduling method based on workload awareness and security situation feedback. It aims to solve the problems of uneven resource utilization and rigid security protection caused by static sharding configuration. By setting splitting and merging thresholds, it achieves elastic scaling of shard size and minimizes cross-shard interactions by combining account relocation algorithms. This allows shard boundaries to dynamically drift with the security situation rather than remain fixed, thereby enabling secure data interaction in a distributed database environment.
[0003] Distributed data encryption technology is an encryption method based on threshold cryptography and distributed collaborative computing. It aims to solve the problems of single point of failure and key reconstruction risk in traditional centralized key management in distributed environments. Through distributed pseudo-random functions and threshold authentication encryption mechanisms, multiple nodes can collaboratively complete encryption and decryption operations without reconstructing the complete key. This ensures that sensitive data remains encrypted throughout its entire lifecycle of cross-node transmission, computation, and storage, effectively resisting man-in-the-middle attacks and node collusion threats.
[0004] Existing data security interaction methods and systems based on distributed databases suffer from several drawbacks. Traditional data sharding strategies lack responsiveness to the real-time security status of nodes, leading to high-risk nodes storing highly sensitive data fragments and creating security vulnerabilities. Furthermore, the system's encryption mechanism relies on centralized key management and independent node encryption, which poses risks of single-point failures and key leakage. It also struggles to support cross-node encrypted collaborative computing, resulting in frequent data decryption during the interaction process and increasing security risks. Summary of the Invention
[0005] The purpose of this invention is to provide a data security interaction method and system based on a distributed database, in order to solve the problems mentioned in the background art. These problems include: traditional data sharding strategies lack response to the real-time security status of nodes, leading to high-risk nodes storing highly sensitive data fragments and causing security risks; the system's encryption mechanism relies on centralized key management and independent node encryption, which poses single-point failure and key leakage risks, and it is difficult to support cross-node encrypted collaborative computing, resulting in frequent data decryption during interaction and increasing security risks.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a data security interaction method and system based on a distributed database, comprising an identity authentication and authorization module, a distributed database management module, a data security interaction control module, a node communication and collaboration module, and an interaction behavior trusted proof module, characterized in that: the identity authentication and authorization module is used for unified identity verification across distributed nodes and real-time dynamic authorization of fine-grained data access permissions; the distributed database management module includes a topology awareness and dynamic sharding unit and a query processing and optimization unit, wherein the topology awareness and dynamic sharding unit proposes an adaptive data sharding algorithm based on security situation awareness, used to monitor the security situation of nodes in real time and automatically split the data into logically related but physically isolated fragments, which are distributed and stored in different security domain nodes; the query processing and optimization unit is used to parse query semantics and dynamically generate differentiated secure execution paths, and automatically adjust the data visibility granularity and desensitization intensity according to the interaction intent; the data security interaction control module includes an identity authentication and authorization module, a distributed database management module, a data security interaction control module, a node communication and collaboration module, and an interaction behavior trusted proof module, characterized in that: the identity authentication and authorization module is used for unified identity verification across distributed nodes and real-time dynamic authorization of fine-grained data access permissions; the distributed database management module includes a topology awareness and dynamic sharding unit and a query processing and optimization unit, wherein the topology awareness and dynamic sharding unit proposes an adaptive data sharding algorithm based on security situation awareness, used for real-time monitoring of node security situation and automatic splitting of data into logically related but physically isolated fragments, which are distributed and stored in different security domain nodes; the query processing and optimization unit is used to parse query semantics and dynamically generate differentiated secure execution paths, and automatically adjust the data visibility granularity and desensitization intensity according to the interaction intent; the data security interaction control module includes an identity authentication and authorization module, a distributed database management module, a data security interaction control module, a node communication and collaboration module, and an interaction behavior trusted proof module, characterized in that: the identity authentication and authorization module is used for unified identity verification across distributed nodes and real The secure interaction control module includes a distributed data encryption unit and a cross-domain consistency security arbitration unit. The distributed data encryption unit proposes a threshold authentication encryption algorithm based on a distributed pseudo-random function to achieve collaborative encryption and decryption processing between distributed database nodes, ensuring that sensitive data remains encrypted throughout the cross-node interaction process. The cross-domain consistency security arbitration unit coordinates minimum common security standards among distributed communication nodes and automatically triggers key rotation and resharding operations when a node's security status degrades. The node communication and coordination module constructs obfuscated multi-hop communication routes, allowing data interaction paths to change dynamically with requests and enabling lightweight negotiation of security parameters between heterogeneous nodes. This transforms communication path diversity into an active obfuscation dimension to enhance anti-interference capabilities. The trusted proof module for interaction behavior generates zero-knowledge proofs containing data tags and policy execution trajectories for each data interaction, supporting post-event verification and auditing without leaking the original data content.
[0007] Preferably, the identity authentication and authorization module employs a dual mechanism of multi-factor dynamic credential generation and zero-knowledge proof verification. When a user initiates a cross-node query, the system generates a one-time dynamic token based on device fingerprints and behavioral characteristics. Each node verifies that the token was generated by a legitimate user without leaking the original credentials through zero-knowledge proof. If an abnormal geographical location is detected, the permissions are automatically downgraded from data export to aggregate statistics. Distributed consensus on identity credibility is achieved before cross-node interaction, enabling real-time dynamic adjustment of fine-grained permissions.
[0008] Preferably, the distributed database management module includes a topology awareness and dynamic sharding unit. The topology awareness and dynamic sharding unit proposes an adaptive data sharding algorithm based on security situation awareness. By monitoring the security level of each communication node in real time, including network isolation strength, hardware trusted execution environment status, geographical location trustworthiness, and data sensitivity in a dual-dimensional indicator, it adaptively performs data sharding operations. At the same time, it adopts a workload awareness rebalancing strategy based on data access patterns to distribute logically related data objects to different security domain nodes, thereby realizing the physical isolation of sensitive data and the proactive construction of security barriers.
[0009] The distributed database management module includes a topology awareness and dynamic sharding unit. This unit employs an adaptive data sharding algorithm based on security situation awareness. It collects multi-dimensional security attributes of each communication node in real time, such as network isolation strength, hardware trusted execution environment status, and geographical location trustworthiness, and constructs a node security situation score function through weighted fusion. The node security situation scores are aggregated to the sharding level and coupled with data sensitivity to build a coupled evaluation model. Based on the coupled evaluation value, shard splitting, merging, and data migration operations are dynamically triggered. Simultaneously, combined with a workload-aware rebalancing strategy based on data access patterns, logically related data objects are distributed across heterogeneous security domain nodes with significant differences in security situation, achieving physical isolation of sensitive data and proactive construction of security barriers. Finally, the weighted entropy value of the security situation differences between shards is calculated to quantify the overall security barrier strength, forming a closed loop of awareness-decision-execution-evaluation.
[0010] Preferably, the adaptive data sharding algorithm based on security situation awareness is as follows: First, the multi-dimensional security attributes of each communication node in the distributed database are transformed into calculable numerical indicators to provide a quantitative basis for subsequent sharding decisions. Three types of security attributes—network isolation strength, hardware trusted execution environment status, and geographical location trustworthiness—are collected in real time for each communication node in the distributed database. A node security situation scoring function is constructed through weighted fusion. This function maps heterogeneous security attributes to a unified numerical space, enabling precise measurement and comparison of security situation, providing a mathematical basis for security situation awareness. The specific calculation formula is expressed as follows:
[0011]
[0012] in, Represented as a communication node The security situation score, Represented as a communication node index. Represented as nodes The quantification value of network isolation strength, Represented as nodes The trustworthiness of the hardware trusted execution environment state. Represented as nodes The credibility of geographical location This is represented by a weighting coefficient indicating the strength of network isolation. This represents the weighting coefficients of the hardware trusted execution environment state. Let be the weighting coefficient for geographic location credibility, and satisfy . Secondly, a coupling model between shard-level security posture and the sensitivity of the data it carries is established to achieve a two-dimensional collaborative assessment of security posture awareness and data sensitivity. The calculated node security posture scores are aggregated to the shard level, and a data sensitivity penalty term is constructed. This ensures that highly sensitive data distributed in low-security-post shards generates a significant negative evaluation value. The constructed coupling model allows shard reorganization decisions to consider both topological security attributes and data asset value, providing a precise decision-making basis for adaptive sharding. The specific calculation formula is expressed as follows:
[0013]
[0014] in, Represented as fragments The security data coupling evaluation value, Represented as a data sharding index, Represented as fragments The set of communication nodes included Represented as a set The number of elements, Represented as fragments Average sensitivity of internal storage data, This is expressed as a data sensitivity penalty coefficient. Then, based on the calculated coupling evaluation value, the sharding splitting and merging operations are dynamically triggered, realizing the core mechanism of adaptive data sharding. The specific calculation formula is expressed as follows:
[0015]
[0016] in, Represented as fragments The results of the restructuring decision, This is represented as the safe split threshold. This is represented as the safe merging threshold. This represents the number of consecutive stable periods required for the merger decision. This indicates the current evaluation cycle number. Represented as periodic Time slice The coupling evaluation value, Represented as a historical evaluation period index variable, used to iterate from... arrive All historical cycles within the time window, This is a split operation command; once triggered, it will split the data into fragments. It is split into multiple logically related but physically isolated sub-shards. This is a merge operation command; once triggered, it will split the data into fragments. Merge resources with low-load shards that have similar interaction patterns. This is represented as a stable state instruction, which is then fragmented upon triggering. Maintain the current configuration and do not reorganize, when sharding of Below The split operation is triggered at certain times, splitting highly sensitive data into logically related but physically isolated fragments. Higher than And continuous When the evaluation cycle remains stable, a merging operation is triggered to optimize resource utilization. The decision-making mechanism enables the data distribution topology to dynamically evolve with the security situation, transforming distribution into a proactive security barrier. Secondly, after fragmentation and reorganization, data migration is performed, distributing logically related data objects across different security domain nodes to achieve physical isolation of sensitive data. The data access pattern characteristics output by the identity authentication and authorization module identify sets of data objects that are frequently accessed jointly. And construct a security domain mapping function to assign objects in the set to domains with security posture differences greater than 100%. The strategy for constructing heterogeneous nodes, while ensuring query performance, utilizes distributed topology to build a defense-in-depth system, directly strengthening the foundation of secure interactions within the system. The specific formula is as follows:
[0017]
[0018] in, Data objects represented as security domain-aware Mapping function to storage nodes, This represents the target data object to be allocated. Represented as the target data object Preceding data objects that have logical relationships Represented as a set of candidate nodes, This is represented as the threshold for the difference in security situation. This represents selecting the node index that maximizes the objective function. Represented as a preceding data object The security posture score of the allocated storage nodes is calculated. Finally, the security barrier strength of the reorganized sharding architecture is quantified to provide closed-loop feedback for the system security posture. The weighted entropy value of the security posture difference between shards is calculated. The larger the entropy value, the more fully sensitive data is distributed in the heterogeneous security domain and the more significant the physical isolation effect. The entropy value metric directly reflects the degree of realization of the distribution, i.e., security paradigm, so that the adaptive sharding process forms a complete closed loop from perception to decision-making to execution to evaluation, continuously optimizing the security interaction environment of the distributed database. The specific formula is expressed as follows:
[0019]
[0020] in, This represents the overall security barrier strength of the system. This represents the current total number of shards. This is expressed as the arithmetic mean of all fragmented coupling evaluation values, serving as a benchmark for assessing the dispersion of fragmented security posture. Represented as fragments The security data coupling evaluation value, This is represented as a numerical stability parameter, used to adjust the denominator to always be positive. This represents the index variable for the inner summation operation, and the traversal range is from 1 to... The sum of the sensitivities of all data fragments is used as the normalized denominator. Represented as fragments Average sensitivity of internal storage data, Represented as The index variable in the function iterates from 1 to... Used to find all fragments The maximum value, It is represented as the natural logarithm function.
[0021] Preferably, the distributed database management module includes a query processing and optimization unit. The query processing and optimization unit predicts the cross-shard data flow path by parsing query semantics and migrates frequently interacting data objects to the same shard, minimizing the amount of cross-node data transmission. At the same time, it uses query intent as a triggering factor for shard reorganization, so that security policies and query semantics are deeply coupled, achieving a dynamic balance between security protection strength and query performance.
[0022] Preferably, the data security interaction control module includes a distributed data encryption unit, which proposes a threshold authentication encryption algorithm based on a distributed pseudo-random function. By keeping sensitive data in a encrypted state throughout the cross-node interaction process, the distributed nodes collaboratively generate key derivation values but never reconstruct the complete key, thereby achieving distributed data protection throughout the key lifecycle.
[0023] The data security interaction control module includes a distributed data encryption unit. This unit employs a threshold authentication encryption algorithm based on a distributed pseudo-random function. The master key is divided into multiple shares and distributed to data slice nodes through a threshold secret sharing mechanism. Upon receiving the message to be encrypted and its semantic commitment, a one-time random number is generated and broadcast to selected nodes. Each node uses its local key share to call the AES pseudo-random function to generate a partial calculation result. A key derivation value is generated through weighted aggregation of Lagrange interpolation coefficients, ensuring the key is never reconstructed. Finally, the key derivation value drives a pseudo-random number generator to produce a key stream, which is concatenated with the message and then XORed to generate ciphertext, ensuring that sensitive data remains encrypted throughout the cross-node interaction process.
[0024] Preferably, the threshold authentication encryption algorithm based on a distributed pseudo-random function is as follows: First, a key distribution operation is performed, distributing the master key... pass Threshold secret sharing mechanism divided into Each share is distributed to the distributed database data slice nodes, so that any number of shares less than 1 share are distributed to the distributed database data slice nodes. The inability of individual communication nodes to reconstruct the master key lays a secure foundation for subsequent threshold-based authentication encryption. The specific calculation formula is as follows:
[0025]
[0026] in, Indicated as distributed to the first The key share of a distributed database data slice node; Represented as the system master key; Represented as the secret shared polynomial in the th One random coefficient; This represents the index of a data slice node in a distributed database. Represents the XOR operation over a finite field; Indicates the index of the number of communication nodes, and then receives the message to be encrypted. And the semantic commitment generated by the identity authentication and authorization module, to generate a one-time random number. Calculate message commitments and Broadcast to selected Each communication node enables a strong binding between encryption operations and data semantics. The specific calculation formula is as follows:
[0027]
[0028] in, Indicates the semantic commitment value of the message; Represents a cryptographically secure hash function; This represents a semantic commitment to the data stored in a distributed database, generated by the identity authentication and authorization module. Represents a one-time random number; This represents a bit string concatenation operation; This indicates the identity of the requesting node. Then, after each participating node verifies the legitimacy of the request source, it uses the key share. Used as a key to call the AES pseudo-random function, for input The evaluation is performed and a mask result is returned, ensuring that the key share is never exposed. The specific calculation formula is as follows:
[0029]
[0030] in, Indicates the first The computation results returned by the distributed database data slice nodes; Indicated by key Perform AES encryption; Represents a node The share of the key held; This indicates the identity of the requester; secondly, all [resources / information] are collected. Partial calculation results returned by each participating node The final key derived value is generated through finite field aggregation operations weighted by Lagrange interpolation coefficients. The aggregation mechanism follows... The mathematical reconstruction principle of threshold secret sharing allows any number of thresholds to be shared. The subset of partial results cannot derive any valid bits of the final key derivation value in an information-theoretic sense. This achieves the key-never-reconstruction characteristic based on a distributed pseudo-random function among data slice nodes in a distributed database. This provides dual security for threshold-authenticated encryption, resisting both single-point key leakage and collusion attacks within the threshold, ensuring that sensitive data remains encrypted throughout cross-node interactions. The specific formula is as follows:
[0031]
[0032] in, Represented as the key-derived value generated by aggregation; Represented as in a finite field Addition operations on top, Indicated as containing A finite field of elements The safety strength factor is set to 256; Represented as the first Each participating node is in the current participating set. The corresponding Lagrange interpolation coefficients; Represented as in a finite field Scalar multiplication operations; This is represented by the set of nodes currently participating in the encryption operation. Finally, the derived key value is calculated. The pseudo-random number generator is driven to generate a key stream, which in turn generates the original message. With the generated one-time random number After concatenation, an XOR operation is performed with the key stream to generate the ciphertext body. At the same time, identify the requester. Semantic commitment value With the ciphertext body Combined to form a complete ciphertext The encrypted data remains encrypted throughout the transmission between distributed database nodes, ensuring both confidentiality and integrity of sensitive data exchanged across nodes. This transforms the node collaboration characteristics of the distributed architecture into a proactive security barrier rather than a source of risk. The specific calculation formula is as follows:
[0033]
[0034] in, Represented as the ciphertext body; This is represented as a pseudo-random number generator built on AES; This represents the original message to be encrypted; This is represented as a bit string concatenation operation, where a pseudo-random number generator derives the key value. Expanding the keystream to the same length as the message allows each bit position to be encrypted using independent key material, eliminating the risk of key reuse and generating a one-time random number. The introduction of this feature allows the same message to generate different ciphertexts in different encrypted sessions, resisting replay attacks and frequency analysis. The ciphertext body... With identity markers Semantic commitment The triplet structure enables the receiver to verify The binding relationship confirms the legitimacy of the encrypted source and restores the data by re-executing the distributed computing process. Decryption is completed without any node requiring reconstruction of the complete master key. This achieves a security paradigm where the key is never reconstructed.
[0035] Preferably, the data security interaction control module includes a cross-domain consistency security arbitration unit. The cross-domain consistency security arbitration unit coordinates the minimum common security standard among distributed database communication nodes through the Merkle tree state synchronization mechanism. When a security state degradation of a corresponding node is detected, it automatically triggers key rotation and data resharding, and conducts decentralized adjudication of disputed transactions through a weighted voting mechanism. This prevents malicious nodes from undermining the overall security level of the system through policy degradation, thereby ensuring the atomicity and policy consistency of cross-domain transactions.
[0036] Preferably, the node communication and collaboration module constructs a confusing multi-hop communication routing strategy, which enables the data interaction path to be dynamically reconstructed according to the node security level for each request. Highly sensitive data is always forwarded by highly trusted nodes, while low-sensitivity data can choose a low-latency path, thereby achieving dynamic confusing of the communication topology. At the same time, a path authentication mechanism based on threshold signature is embedded in cross-node transmission, which prevents intermediate nodes from tampering with the routing information, so that the data interaction path changes dynamically with the request and is strongly correlated with the node security level.
[0037] Preferably, the interactive behavior trust proof module generates a zero-knowledge proof containing data tags and policy execution trajectories for each data interaction, binding the data tag path, security policy execution trajectory, and node identity to an immutable cryptographic evidence chain, enabling the auditor to complete the interaction compliance verification without accessing the original sensitive data.
[0038] Compared with the prior art, the beneficial effects of the present invention are:
[0039] 1. The topology-aware and dynamic sharding unit proposes an adaptive data sharding algorithm based on security situation awareness. First, this algorithm quantifies the multi-dimensional heterogeneous security attributes of each communication node in the network—including network isolation strength, hardware trusted execution environment status, and geographical location trustworthiness—in real time into a unified and comparable security situation score. This provides an accurate and objective security quantification benchmark for constructing distributed data storage topologies, enabling the system to break through the traditional security configuration relying on static policies and human experience. It achieves continuous and accurate awareness of the security status of the underlying infrastructure, laying the foundation for dynamic evolution of all upper-layer security interactions. Second, based on this real-time awareness capability, the algorithm constructs a dynamic relationship between sharding security situation and the sensitivity of the data it carries. The coupled assessment model intelligently matches the storage location of data assets with their intrinsic value and security requirements. High-sensitivity data is no longer statically stored on a fixed node. Instead, through algorithmic calculation, when the overall security posture assessment value of the current shard is found to be low, a split operation is automatically triggered. This operation intelligently splits logically closely related data into multiple physically isolated fragments. Based on the real-time security scores of each node, these fragments are distributed and stored in heterogeneous security domain nodes with higher security posture scores and significant differences between them. This process achieves both logical association and physical isolation, ensuring that even if a data attack breaches a single node, it can only obtain data fragments that cannot reconstruct the complete semantics, thus increasing the risk of stealing complete high-value data. The difficulty and cost of assets, thus constructing a deep security defense system at the physical level of data storage, while the adaptive nature of the algorithm is a continuously running closed-loop optimization process. When a node's security posture score degrades due to network attacks, environmental changes, and policy adjustments, the algorithm can detect and reassess the security of the relevant shards in real time. Once a security threshold is reached, it automatically initiates a data migration and shard reassessment process, dynamically and smoothly migrating sensitive data fragments from the degraded nodes to more secure nodes. When the overall system security posture is good and stable, the algorithm can intelligently trigger a security merging operation, optimizing the utilization efficiency of storage and computing resources while ensuring security redundancy. This elastic architecture that adjusts with changes in the security environment enables... The entire distributed database system possesses powerful adaptive recovery capabilities. Finally, by quantifying the overall security barrier strength of the system, the algorithm provides a measurable and verifiable security gain closed loop for the entire data security interaction system. This indicator profoundly reflects the sufficiency and dispersion of sensitive data distribution in heterogeneous security domains. The larger the entropy value, the more significant the physical isolation effect, and the more robust the security barrier built by the system in a distributed manner. In summary, the algorithm's end-to-end innovation from perception, decision-making, execution to evaluation transforms the topology of data distribution itself into a powerful, dynamic, and proactive security asset, fundamentally strengthening the core capability of distributed databases to achieve reliable and trustworthy data security interaction in complex threat environments.
[0040] 2. The distributed data encryption unit proposes a threshold authentication encryption algorithm based on a distributed pseudo-random function. First, this algorithm, through a threshold secret sharing mechanism, securely splits the master key into multiple shares during system initialization and distributes them across various data slice nodes. This design eliminates the single key storage point at its source, preventing attackers from obtaining the complete key material by compromising any one or a few nodes. This fundamentally solves the single point of failure and fatal leakage risks inherent in traditional key management centers, laying a solid foundation for secure data interaction throughout the system without the need for a single trusted entity. Second, this algorithm achieves [something] throughout the entire encryption and decryption process. The availability of key shares is invisible. When data encryption and decryption are required, the request must be completed collaboratively by at least a number of nodes. Each participating node uses only its own key share, which never leaves the security boundary, to perform localized computation on the input bound to the requester's identity and specific data semantics, generating a partial result. The partial result is then securely aggregated to generate the final encryption key. The complete original master key and valid key derivations are never reconstructed or displayed at any time or place. This mechanism ensures that the key material is always in a distributed stealth state during system operation, compressing the attack surface exposed by the key and achieving information theory-level stealth. Security is enhanced. Furthermore, by strongly binding encryption operations to semantic commitments generated by the identity authentication module, the algorithm delivers verifiable integrity and non-repudiation. Each encryption operation is uniquely associated with a specific requester's identity, data content, and a one-time random number, making the generated ciphertext itself proof of a trusted and secure interaction. This effectively resists ciphertext replay attacks, substitution attacks, and unauthorized decryption attempts, as any invalid or tampered request cannot pass the collaborative verification process of distributed nodes. Simultaneously, since the encryption process mandates the participation of multiple nodes, it introduces a natural distributed witness for all critical data interactions, increasing security. The difficulty and cost for attackers to secretly tamper with and steal data have been reduced, ensuring the auditability and trustworthiness of operations. In summary, this algorithm, through its distributed encrypted computing design, can improve the confidentiality and integrity of sensitive data during cross-node transmission, storage, and even processing. It transforms the traditional model, which relies on static keys and boundary protection, into a dynamic, embedded, proactive security capability based on cryptographic primitives and distributed consensus. This makes maintaining data in a encrypted state throughout the interaction process a default state guaranteed by the architecture, rather than a security objective. Thus, it provides solid encrypted support for building data security interaction methods and systems based on distributed databases.
[0041] 3. In summary, compared with existing distributed databases that employ static sharding strategies and centralized key management, this invention achieves real-time adaptation of data physical distribution and security posture through a security situation awareness-driven multi-dimensional coupled evaluation model and a dynamic splitting and merging mechanism, transforming the topology into an active defense barrier. Simultaneously, by utilizing the threshold authentication encryption and semantic commitment binding mechanism of distributed pseudo-random functions, it achieves encrypted collaborative computation with keys that are never reconstructed, solving the problems of single-point key leakage and encrypted computation in cross-node interactions. This constructs a complete security closed loop from awareness sharding to encrypted interaction and then to trusted proof, realizing a substantial leap from passive protection to dynamic adaptive protection in distributed database data security interaction. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the structure of the present invention; Detailed Implementation
[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] Please see Figure 1This invention provides a data security interaction method and system based on a distributed database, including an identity authentication and authorization module, a distributed database management module, a data security interaction control module, a node communication and collaboration module, and an interaction behavior trusted proof module. The identity authentication and authorization module is used for unified identity verification across distributed nodes and real-time dynamic authorization of fine-grained data access permissions. The distributed database management module includes a topology awareness and dynamic sharding unit and a query processing and optimization unit. The topology awareness and dynamic sharding unit proposes an adaptive data sharding algorithm based on security situation awareness, used to monitor the security situation of nodes in real time and automatically split data into logically related but physically isolated fragments, distributed and stored in different security domain nodes. The query processing and optimization unit is used to parse query semantics and dynamically generate differentiated secure execution paths, automatically adjusting the data visibility granularity and desensitization strength according to the interaction intent. The data security interaction control module... The module includes a distributed data encryption unit and a cross-domain consistency security arbitration unit. The distributed data encryption unit proposes a threshold authentication encryption algorithm based on a distributed pseudo-random function to achieve collaborative encryption and decryption processing between distributed database nodes with keys that are never reconstructed, ensuring that sensitive data remains encrypted throughout cross-node interactions. The cross-domain consistency security arbitration unit coordinates minimum common security standards among distributed communication nodes and automatically triggers key rotation and resharding operations when a node's security status degrades. The node communication and collaboration module constructs obfuscated multi-hop communication routes, enabling data interaction paths to change dynamically with requests and achieving lightweight negotiation of security parameters between heterogeneous nodes, transforming communication path diversity into an active obfuscation dimension to enhance anti-interference capabilities. The interaction behavior trusted proof module generates zero-knowledge proofs containing data tags and policy execution trajectories for each data interaction, supporting post-event verification and auditing without leaking the original data content.
[0045] See Figure 1 Furthermore, the identity authentication and authorization module employs a dual mechanism of multi-factor dynamic credential generation and zero-knowledge proof verification. When a user initiates a cross-node query, the system generates a one-time dynamic token based on device fingerprints and behavioral characteristics. Each node verifies that the token was generated by a legitimate user without disclosing the original credentials through zero-knowledge proof. If an abnormal geographical location is detected, the permissions are automatically downgraded from data export to aggregate statistics. Distributed consensus on identity credibility is achieved before cross-node interaction, enabling real-time dynamic adjustment of fine-grained permissions.
[0046] See Figure 1Furthermore, the distributed database management module includes a topology awareness and dynamic sharding unit. This unit proposes an adaptive data sharding algorithm based on security situation awareness. By monitoring the security level of each communication node in real time, including network isolation strength, hardware trusted execution environment status, geographical location trustworthiness, and data sensitivity, the algorithm adaptively performs data sharding operations. At the same time, it adopts a workload-aware rebalancing strategy based on data access patterns to distribute logically related data objects across different security domain nodes, thereby achieving physical isolation of sensitive data and proactive construction of security barriers.
[0047] See Figure 1 Furthermore, the adaptive data sharding algorithm based on security situation awareness is as follows: First, the multi-dimensional security attributes of each communication node in the distributed database are transformed into calculable numerical indicators to provide a quantitative basis for subsequent sharding decisions. In real time, three types of security attributes—network isolation strength, hardware trusted execution environment status, and geographical location trustworthiness—are collected from each communication node in the distributed database. A node security situation scoring function is constructed through weighted fusion. This function maps heterogeneous security attributes to a unified numerical space, enabling precise measurement and comparison of security situation, providing a mathematical basis for security situation awareness. The specific calculation formula is expressed as follows:
[0048]
[0049] in, Represented as a communication node The security situation score, Represented as a communication node index. Represented as nodes The quantification value of network isolation strength, Represented as nodes The trustworthiness of the hardware trusted execution environment state. Represented as nodes The credibility of geographical location This is represented by a weighting coefficient indicating the strength of network isolation. This represents the weighting coefficients of the hardware trusted execution environment state. Let be the weighting coefficient for geographic location credibility, and satisfy . Secondly, a coupling model between shard-level security posture and the sensitivity of the data it carries is established to achieve a two-dimensional collaborative assessment of security posture awareness and data sensitivity. The calculated node security posture scores are aggregated to the shard level, and a data sensitivity penalty term is constructed. This ensures that highly sensitive data distributed in low-security-post shards generates a significant negative evaluation value. The constructed coupling model allows shard reorganization decisions to consider both topological security attributes and data asset value, providing a precise decision-making basis for adaptive sharding. The specific calculation formula is expressed as follows:
[0050]
[0051] in, Represented as fragments The security data coupling evaluation value, Represented as a data sharding index, Represented as fragments The set of communication nodes included Represented as a set The number of elements, Represented as fragments Average sensitivity of internal storage data, This is expressed as a data sensitivity penalty coefficient. Then, based on the calculated coupling evaluation value, the sharding splitting and merging operations are dynamically triggered, realizing the core mechanism of adaptive data sharding. The specific calculation formula is expressed as follows:
[0052]
[0053] in, Represented as fragments The results of the restructuring decision, This is represented as the safe split threshold. This is represented as the safe merging threshold. This represents the number of consecutive stable periods required for the merger decision. This indicates the current evaluation cycle number. Represented as periodic Time slice The coupling evaluation value, Represented as a historical evaluation period index variable, used to iterate from... arrive All historical cycles within the time window, This is a split operation command; once triggered, it will split the data into fragments. It is split into multiple logically related but physically isolated sub-shards. This indicates a merge operation command, which will split the data upon triggering. Merge resources with low-load shards that have similar interaction patterns. This is represented as a stable state instruction, which is then fragmented upon triggering. Maintain the current configuration and do not reorganize, when sharding of Below The split operation is triggered at certain times, splitting highly sensitive data into logically related but physically isolated fragments. Higher than And continuous When the evaluation cycle remains stable, a merging operation is triggered to optimize resource utilization. The decision-making mechanism enables the data distribution topology to dynamically evolve with the security situation, transforming distribution into a proactive security barrier. Secondly, after fragmentation and reorganization, data migration is performed, distributing logically related data objects across different security domain nodes to achieve physical isolation of sensitive data. The data access pattern characteristics output by the identity authentication and authorization module identify sets of data objects that are frequently accessed jointly. And construct a security domain mapping function to assign objects in the set to domains with security posture differences greater than 100%. The strategy for constructing heterogeneous nodes, while ensuring query performance, utilizes distributed topology to build a defense-in-depth system, directly strengthening the foundation of secure interactions within the system. The specific formula is as follows:
[0054]
[0055] in, Data objects represented as security domain-aware Mapping function to storage nodes, This represents the target data object to be allocated. Represented as the target data object Preceding data objects that have logical relationships Represented as a set of candidate nodes, This is represented as the threshold for the difference in security situation. This represents selecting the node index that maximizes the objective function. Represented as a preceding data object The security posture score of the allocated storage nodes is calculated. Finally, the security barrier strength of the reorganized sharding architecture is quantified to provide closed-loop feedback for the system security posture. The weighted entropy value of the security posture difference between shards is calculated. The larger the entropy value, the more fully sensitive data is distributed in the heterogeneous security domain and the more significant the physical isolation effect. The entropy value metric directly reflects the degree of realization of the distribution, i.e., security paradigm, so that the adaptive sharding process forms a complete closed loop from perception to decision-making to execution to evaluation, continuously optimizing the security interaction environment of the distributed database. The specific formula is expressed as follows:
[0056]
[0057] in, This represents the overall security barrier strength of the system. This represents the current total number of shards. This is expressed as the arithmetic mean of all fragmented coupling evaluation values, serving as a benchmark for assessing the dispersion of fragmented security posture. Represented as fragments The security data coupling evaluation value, This is represented as a numerical stability parameter, used to adjust the denominator to always be positive. This represents the index variable for the inner summation operation, and the traversal range is from 1 to... The sum of the sensitivities of all data fragments is used as the normalized denominator. Represented as fragments Average sensitivity of internal storage data, Represented as The index variable in the function iterates from 1 to... Used to find all fragments The maximum value, It is represented as the natural logarithm function.
[0058] See Figure 1 Furthermore, the distributed database management module includes a query processing and optimization unit. This unit predicts the cross-shard data flow path by parsing query semantics and migrates frequently interacting data objects to the same shard, minimizing the amount of cross-node data transmission. At the same time, it uses query intent as a triggering factor for shard reorganization, thereby deeply coupling security policies with query semantics and achieving a dynamic balance between security protection strength and query performance.
[0059] See Figure 1 Furthermore, the data security interaction control module includes a distributed data encryption unit, which proposes a threshold authentication encryption algorithm based on a distributed pseudo-random function. By keeping sensitive data in a encrypted state throughout the cross-node interaction process, the distributed nodes collaboratively generate key derivation values but never reconstruct the complete key, thereby achieving distributed data protection throughout the key lifecycle.
[0060] See Figure 1 Furthermore, the threshold authentication encryption algorithm based on a distributed pseudo-random function is specifically as follows: First, a key distribution operation is performed, distributing the master key... pass Threshold secret sharing mechanism divided into Each share is distributed to the distributed database data slice nodes, so that any number of shares less than 1 share are distributed to the distributed database data slice nodes. The inability of individual communication nodes to reconstruct the master key lays a secure foundation for subsequent threshold-based authentication encryption. The specific calculation formula is as follows:
[0061]
[0062] in, Indicated as distributed to the first The key share of a distributed database data slice node; Represented as the system master key; Represented as the secret shared polynomial in the th One random coefficient; This represents the index of a data slice node in a distributed database. Represents the XOR operation over a finite field; Indicates the index of the number of communication nodes, and then receives the message to be encrypted. And the semantic commitment generated by the identity authentication and authorization module, to generate a one-time random number. Calculate message commitments and Broadcast to selected Each communication node enables a strong binding between encryption operations and data semantics. The specific calculation formula is as follows:
[0063]
[0064] in, Indicates the semantic commitment value of the message; Represents a cryptographically secure hash function; This represents a semantic commitment to the data stored in a distributed database, generated by the identity authentication and authorization module. Represents a one-time random number; This represents a bit string concatenation operation; This indicates the identity of the requesting node. Then, after each participating node verifies the legitimacy of the request source, it uses the key share. Used as a key to call the AES pseudo-random function, for input The evaluation is performed and a mask result is returned, ensuring that the key share is never exposed. The specific calculation formula is as follows:
[0065]
[0066] in, Indicates the first The computation results returned by the distributed database data slice nodes; Indicated by key Perform AES encryption; Represents a node The share of the key held; This indicates the identity of the requester; secondly, all [resources / information] are collected. Partial calculation results returned by each participating node The final key derived value is generated through finite field aggregation operations weighted by Lagrange interpolation coefficients. The aggregation mechanism follows... The mathematical reconstruction principle of threshold secret sharing allows any number of thresholds to be shared. The subset of partial results cannot derive any valid bits of the final key derivation value in an information-theoretic sense. This achieves the key-never-reconstruction characteristic based on a distributed pseudo-random function among data slice nodes in a distributed database. This provides dual security for threshold-authenticated encryption, resisting both single-point key leakage and collusion attacks within the threshold, ensuring that sensitive data remains encrypted throughout cross-node interactions. The specific formula is as follows:
[0067]
[0068] in, Represented as the key-derived value generated by aggregation; Represented as in a finite field Addition operations on top, Indicated as containing A finite field of elements The safety strength factor is set to 256; Represented as the first Each participating node is in the current participating set. The corresponding Lagrange interpolation coefficients; Represented as in a finite field Scalar multiplication operations; This is represented by the set of nodes currently participating in the encryption operation. Finally, the derived key value is calculated. The pseudo-random number generator is driven to generate a key stream, which in turn generates the original message. With the generated one-time random number After concatenation, an XOR operation is performed with the key stream to generate the ciphertext body. At the same time, identify the requester. Semantic commitment value With the ciphertext body Combined to form a complete ciphertext The encrypted data remains encrypted throughout the transmission between distributed database nodes, ensuring both confidentiality and integrity of sensitive data exchanged across nodes. This transforms the node collaboration characteristics of the distributed architecture into a proactive security barrier rather than a source of risk. The specific calculation formula is as follows:
[0069]
[0070] in, Represented as the ciphertext body; This is represented as a pseudo-random number generator built on AES; This represents the original message to be encrypted; This is represented as a bit string concatenation operation, where a pseudo-random number generator derives the key value. Expanding the keystream to the same length as the message allows each bit position to be encrypted using independent key material, eliminating the risk of key reuse and generating a one-time random number. The introduction of this feature allows the same message to generate different ciphertexts in different encrypted sessions, resisting replay attacks and frequency analysis. The ciphertext body... With identity markers Semantic commitment The triplet structure enables the receiver to verify The binding relationship confirms the legitimacy of the encrypted source and restores the data by re-executing the distributed computing process. Decryption is completed without any node requiring reconstruction of the complete master key. This achieves a security paradigm where the key is never reconstructed.
[0071] See Figure 1 Furthermore, the data security interaction control module includes a cross-domain consistency security arbitration unit. This unit coordinates the minimum common security standard among distributed database communication nodes through a Merkle tree state synchronization mechanism. When a security state degradation of a corresponding node is detected, it automatically triggers key rotation and data resharding. It also conducts decentralized adjudication of disputed transactions through a weighted voting mechanism, preventing malicious nodes from undermining the overall security level of the system through policy degradation, thus ensuring the atomicity and policy consistency of cross-domain transactions.
[0072] See Figure 1 Furthermore, the node communication and collaboration module constructs a confusing multi-hop communication routing strategy, which enables the data interaction path to be dynamically reconstructed based on the node security level for each request. Highly sensitive data is always forwarded by highly trusted nodes, while low-sensitivity data can choose a low-latency path, thus achieving dynamic confusing of the communication topology. At the same time, a path authentication mechanism based on threshold signature is embedded in cross-node transmission, which prevents intermediate nodes from tampering with the routing information, so that the data interaction path changes dynamically with the request and is strongly correlated with the node security level.
[0073] See Figure 1 Furthermore, the interactive behavior trust proof module generates a zero-knowledge proof containing data tags and policy execution trajectories for each data interaction, binding the data tag path, security policy execution trajectory, and node identity to an immutable cryptographic evidence chain, enabling the auditor to complete the interaction compliance verification without accessing the original sensitive data.
[0074] In practical use, firstly, the identity authentication and authorization module is used for unified identity verification across distributed nodes and real-time dynamic authorization of fine-grained data access permissions; secondly, the distributed database management module includes a topology awareness and dynamic sharding unit and a query processing and optimization unit. The topology awareness and dynamic sharding unit proposes an adaptive data sharding algorithm based on security situation awareness, used to monitor the security situation of nodes in real time and automatically split data into logically related but physically isolated fragments, which are distributed and stored on nodes in different security domains. The query processing and optimization unit is used to parse query semantics and dynamically generate differentiated secure execution paths, automatically adjusting the granularity of data visibility and the strength of de-identification according to the interaction intent; then, the data security interaction control module includes a distributed data encryption unit and a cross-domain consistency security arbitration unit. The distributed data encryption unit proposes a distributed... A threshold-based encryption algorithm using pseudo-random functions is used to achieve collaborative encryption and decryption processing between distributed database nodes, ensuring that the key is never reconstructed. This keeps sensitive data encrypted throughout cross-node interactions. A cross-domain consistency security arbitration unit coordinates minimum common security standards among distributed communication nodes and automatically triggers key rotation and resharding operations when a node's security status degrades. Secondly, a node communication and collaboration module is used to construct obfuscated multi-hop communication routes, enabling data interaction paths to change dynamically with requests and achieving lightweight negotiation of security parameters between heterogeneous nodes. This transforms the diversity of communication paths into an active obfuscation dimension to enhance anti-interference capabilities. Finally, an interaction behavior trusted proof module generates zero-knowledge proofs containing data tags and policy execution traces for each data interaction, supporting post-event verification and auditing without leaking the original data content.
[0075] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments and make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A data security interaction method and system based on a distributed database, comprising an identity authentication and authorization module, a distributed database management module, a data security interaction control module, a node communication and collaboration module, and an interaction behavior trusted proof module, characterized in that: The identity authentication and authorization module is used for unified identity authentication across distributed nodes and real-time dynamic authorization of fine-grained data access permissions; the distributed database management module includes a topology awareness and dynamic sharding unit and a query processing and optimization unit. The topology awareness and dynamic sharding unit proposes an adaptive data sharding algorithm based on security situation awareness, which is used to monitor the security situation of nodes in real time and automatically split the data into logically related but physically isolated fragments, which are distributed and stored in different security domain nodes. The query processing and optimization unit is used to parse query semantics and dynamically generate differentiated secure execution paths, and automatically adjust the data visibility granularity and desensitization intensity according to the interaction intent. The data security interaction control module includes a distributed data encryption unit and a cross-domain consistency security arbitration unit. The distributed data encryption unit proposes a threshold authentication encryption algorithm based on a distributed pseudo-random function to achieve collaborative encryption and decryption processing between distributed database nodes, ensuring that sensitive data remains encrypted throughout the cross-node interaction process. The cross-domain consistency security arbitration unit coordinates the minimum common security standards among distributed communication nodes and automatically triggers key rotation and resharding operations when the node's security status degrades. The node communication and collaboration module constructs obfuscated multi-hop communication routes, allowing data interaction paths to change dynamically with requests and enabling lightweight negotiation of security parameters between heterogeneous nodes. This transforms communication path diversity into an active obfuscation dimension to enhance anti-interference capabilities. The interaction behavior trusted proof module generates zero-knowledge proofs containing data tags and policy execution trajectories for each data interaction, supporting post-event verification and auditing without leaking the original data content.
2. The data security interaction method and system based on a distributed database according to claim 1, characterized in that: The identity authentication and authorization module achieves distributed consensus on identity credibility before cross-node interaction through a dual mechanism of multi-factor dynamic credential generation and zero-knowledge proof verification, enabling real-time dynamic adjustment of fine-grained permissions.
3. The data security interaction method and system based on a distributed database according to claim 1, characterized in that: The distributed database management module includes a topology awareness and dynamic sharding unit. This unit employs an adaptive data sharding algorithm based on security situation awareness. It collects multi-dimensional security attributes of each communication node in real time, such as network isolation strength, hardware trusted execution environment status, and geographical location trustworthiness, and constructs a node security situation score function through weighted fusion. The node security situation scores are aggregated to the sharding level and coupled with data sensitivity to build a coupled evaluation model. Based on the coupled evaluation value, shard splitting, merging, and data migration operations are dynamically triggered. Simultaneously, combined with a workload-aware rebalancing strategy based on data access patterns, logically related data objects are distributed across heterogeneous security domain nodes with significant differences in security situation, achieving physical isolation of sensitive data and proactive construction of security barriers. Finally, the weighted entropy value of the security situation differences between shards is calculated to quantify the overall security barrier strength, forming a closed loop of awareness-decision-execution-evaluation.
4. The data security interaction method and system based on a distributed database according to claim 1, characterized in that: The distributed database management module includes a query processing and optimization unit. This unit predicts the cross-shard data flow path by parsing query semantics and migrates frequently interacting data objects to the same shard, minimizing the amount of cross-node data transmission. At the same time, it uses query intent as a triggering factor for shard reorganization, thereby deeply coupling security policies with query semantics and achieving a dynamic balance between security protection strength and query performance.
5. The data security interaction method and system based on a distributed database according to claim 1, characterized in that: The data security interaction control module includes a distributed data encryption unit. This unit employs a threshold authentication encryption algorithm based on a distributed pseudo-random function. The master key is divided into multiple shares and distributed to data slice nodes through a threshold secret sharing mechanism. Upon receiving the message to be encrypted and its semantic commitment, a one-time random number is generated and broadcast to selected nodes. Each node uses its local key share to call the AES pseudo-random function to generate a partial calculation result. A key derivation value is generated through weighted aggregation of Lagrange interpolation coefficients, ensuring the key is never reconstructed. Finally, the key derivation value drives a pseudo-random number generator to produce a key stream, which is concatenated with the message and then XORed to generate ciphertext, ensuring that sensitive data remains encrypted throughout the cross-node interaction process.
6. The data security interaction method and system based on a distributed database according to claim 1, characterized in that: The data security interaction control module includes a cross-domain consistency security arbitration unit. This unit coordinates the minimum common security standard among distributed database communication nodes through a Merkle tree state synchronization mechanism. When a security state degradation of a corresponding node is detected, it automatically triggers key rotation and data resharding. It also conducts decentralized adjudication of disputed transactions through a weighted voting mechanism, preventing malicious nodes from undermining the overall security level of the system through policy degradation, thus ensuring the atomicity and policy consistency of cross-domain transactions.
7. The data security interaction method and system based on a distributed database according to claim 1, characterized in that: The node communication and collaboration module constructs a confusing multi-hop communication routing strategy, which enables the data interaction path to be dynamically reconstructed based on the node security level for each request. Highly sensitive data is always forwarded by highly trusted nodes, while low-sensitivity data can choose a low-latency path, thus achieving dynamic confusing of the communication topology. At the same time, a path authentication mechanism based on threshold signature is embedded in cross-node transmission, which prevents intermediate nodes from tampering with the routing information, so that the data interaction path changes dynamically with the request and is strongly correlated with the node security level.
8. The data security interaction method and system based on a distributed database according to claim 1, characterized in that: The interactive behavior trust proof module generates a zero-knowledge proof containing data tags and policy execution trajectories for each data interaction, binding the data tag path, security policy execution trajectory, and node identity to an immutable cryptographic evidence chain, enabling the auditor to complete the interaction compliance verification without accessing the original sensitive data.