Medical examination data sharing method and system based on cloud authentication
Through dynamic desensitization processing, zero-knowledge proof and anti-quantum attribute-based encryption technology, the problems of insufficient privacy protection and inefficient cross-domain authentication in the existing medical test data sharing system are solved, and the privacy-utility dynamic balance and adaptive defense are achieved in multiple scenarios, improving the security and efficiency of the data sharing system.
Patent Information
- Application Number
- CN202510591423.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the existing medical test data sharing system responds to the dynamic sharing needs, static desensitization rules are difficult to adapt to the differentiated privacy needs of multiple types of test data. Traditional identity authentication has the risk of credential impersonation and privacy leakage. A single encryption strategy cannot balance the real-time nature of edge computing nodes with the long-term security of central storage. Mechanical access control lacks dynamic response capabilities to the context environment and new attack methods.
Dynamic desensitization processing, zero-knowledge proof, anti-quantum attribute-based encryption and blockchain audit technology are adopted to generate privacy-utility balanced desensitization data through the optimization model through the rate distortion theory, combining differential privacy noise injection and context-aware decryption to achieve fine-grained access control and adaptive defense.
It realizes data sharing with dynamic balance of privacy-utility in multiple scenarios, resists quantum computing attacks, provides trustless identity authentication and adaptive defense against new attacks, and improves cross-domain authentication efficiency and data security.
Smart Images

Figure CN120281560A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical test data, and in particular to a medical test data sharing method and system based on cloud authentication. Background Art
[0002] The current medical test data sharing system generally adopts role-based access control (RBAC) combined with traditional encryption technology to achieve cloud authentication, mainly involving key technologies such as static data desensitization, certificate-based identity authentication, and homomorphic encryption operations. Related solutions have been actually deployed in scenarios such as regional medical alliances and cross-institutional test results mutual recognition, relying on two-factor authentication, attribute-based encryption (ABE) and other methods to ensure data flow security. When responding to dynamic sharing needs, existing technical systems usually combine audit logs and access control lists (ACLs) for permission management, and some systems try to introduce blockchain technology to achieve operation evidence.
[0003] However, static desensitization rules are difficult to adapt to the differentiated privacy needs of multiple types of inspection data. Traditional identity authentication has the dual risks of credential misuse and privacy leakage. A single encryption strategy cannot balance the real-time performance of edge computing nodes and the long-term security of central storage. Mechanical access control lacks the ability to dynamically respond to contextual environments and new attack methods, resulting in systemic defects such as insufficient privacy protection, inefficient cross-domain authentication, weak resistance to quantum attacks, and delayed policy iteration in the process of medical data sharing. Summary of the invention
[0004] In view of the shortcomings of the prior art, the present invention provides a medical test data sharing method and system based on cloud authentication, which solves the problems in the prior art that static desensitization rules are difficult to adapt to the differentiated privacy requirements of multiple types of test data, traditional identity authentication has the dual risks of credential misuse and privacy leakage, a single encryption strategy cannot balance the real-time performance of edge computing nodes and the long-term security of central storage, and mechanical access control lacks the ability to dynamically respond to contextual environments and new attack methods.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A medical test data sharing method based on cloud authentication comprises the following steps: Dynamically desensitize the original medical test data at the edge node to generate desensitized data that meets the privacy-utility balance; Perform dynamic desensitization processing on the original medical test data at the edge node to generate desensitized data that meets the privacy-utility balance. Optimize the model through rate-distortion theory, quantitatively analyze the balance relationship between data availability and privacy leakage risk, and adjust the desensitization strategy in real time. The edge node is built with a feature parsing engine to identify sensitive fields and label attribute tags, combine the sliding window to statistically calculate the joint probability distribution, and dynamically generate the optimal desensitization rules. The differential privacy noise injection module generates noise based on a hardware random source, while protecting sensitive information, retaining the statistical characteristics of the data.
[0006] The rate-distortion theory establishes a mathematical optimization framework to constrain the amount of sensitive information leakage at the information theory level; edge computing realizes local data processing to avoid the transmission of original data; differential privacy noise enhances the defense ability against background knowledge attacks.
[0007] Verify the identity and permissions of the requesting end user through the zero-knowledge proof protocol, and generate an authentication result bound to the privacy policy; Verify the identity and permissions of the requesting end user through the zero-knowledge proof protocol, and generate an authentication result bound to the privacy policy. The permission statement submitted by the user is compiled into arithmetic circuit logic constraints, and a verifiable zero-knowledge proof is generated based on elliptic curve bilinear pairing. During the verification process, the cloud policy management module only verifies the validity of the proof without obtaining the user's true identity or certificate content.
[0008] Zero-knowledge proof realizes "proof is authorization" through a cryptographic protocol to ensure that the verification process does not disclose privacy information; bilinear pairing provides mathematical non-forgeability to prevent the forgery or tampering of proofs; dynamic policy binding enables permissions to be associated with access scenarios in real time.
[0009] Perform hierarchical encryption on the desensitized data based on the quantum-resistant attribute-based encryption algorithm to generate ciphertext dynamically associated with access attributes; Perform hierarchical encryption on the desensitized data based on the quantum-resistant attribute-based encryption algorithm to generate ciphertext dynamically associated with access attributes. Use lattice cryptography algorithms to construct a hierarchical encryption system. The edge cloud layer adopts a lightweight LWE encryption scheme, and the central cloud layer adopts the fully homomorphic BGV algorithm to support long-term secure storage. The attribute policy parser encodes the access conditions into Boolean logic expressions, and the ciphertext decryption key takes effect only when the attributes of the requesting end match the policy.
[0010] Lattice cryptography is based on the RLWE mathematical problem to resist quantum computing attacks; hierarchical encryption takes into account both the efficiency of edge computing and the security of central storage; dynamic attribute binding realizes fine-grained access control.
[0011] Dynamically decrypt the data according to real-time context parameters, and perform secondary desensitization processing in combination with the access scenario; Dynamically decrypt data according to real-time context parameters, and perform secondary desensitization processing in combination with the access scenario. The decryption gateway integrates a context-aware engine to collect device fingerprint, geographical location, and timestamp parameters, and trigger dynamic decryption policies. When an unconventional access scenario is detected, a bitmask operation is applied to the decrypted data to hide unnecessary fields and embed invisible watermarks.
[0012] Context parameters construct a multi-dimensional access scenario portrait to achieve dynamic permission control; secondary desensitization reduces the risk of privacy leakage through data degradation; watermark technology provides the ability to trace back leaks.
[0013] Write data operation records into the blockchain for audit and traceability, and iteratively update the desensitization policy according to the privacy attack model; Write data operation records into the blockchain for audit and traceability, and iteratively update the desensitization policy according to the privacy attack model. The blockchain nodes adopt the PBFT consensus mechanism to ensure the immutability of operation logs. The attack model training platform simulates privacy attacks through a generative adversarial network, evaluates the protection ability of the current policy, and drives the dynamic adjustment of rate-distortion optimization parameters.
[0014] The blockchain distributed ledger ensures the credibility and integrity of audit data; adversarial learning simulates real attack behaviors to achieve the active evolution of defense strategies; the hot update mechanism ensures that policy iteration has no service interruption.
[0015] Preferably, the dynamic desensitization processing includes: Construct a rate-distortion theory optimization model, aiming to minimize the mutual information between the desensitized data and the original data, and constrain the leakage amount of sensitive attributes not to exceed a preset threshold; Overlay differential privacy noise on the desensitized data, and the noise scale parameter is determined by the ratio of the global sensitivity of the utility attribute to the privacy budget.
[0016] Construct a rate-distortion theory optimization model, aiming to minimize the mutual information between the desensitized data and the original data, and constrain the leakage amount of sensitive attributes not to exceed a preset threshold. By dynamically adjusting the conditional probability distribution, quantitatively analyze the balance relationship between data availability loss and privacy leakage risk. The optimization engine continuously updates the desensitization rules based on iterative algorithms to ensure that potential inference paths of sensitive information are suppressed while meeting the needs of clinical analysis.
[0017] The rate-distortion theory framework transforms privacy protection into an optimization problem under information-theoretic constraints, reduces data correlation through mutual information minimization, and simultaneously uses the mutual information threshold of sensitive attributes to control the upper limit of privacy leakage. The dynamic adjustment mechanism enables the desensitization policy to adapt to different data types and usage scenarios, avoiding over-protection or under-protection problems caused by static rules.
[0018] Add differential privacy noise to the de-identified data. The noise scale parameter is determined by the ratio of the global sensitivity of the utility attribute to the privacy budget. The noise injection module differentially configures the noise intensity according to the statistical characteristics of the data fields and business requirements. For highly sensitive continuous data such as blood glucose values, a non-linear noise scaling strategy is adopted; for discrete categorical data, a perturbation mechanism based on probability distribution is applied.
[0019] Differential privacy eliminates the impact of individual data on the overall statistical results through a mathematically provable noise mechanism. The global sensitivity quantifies the maximum impact of a single data change on the output. Combining with the privacy budget to control the noise intensity, it achieves an adjustable balance between the privacy protection strength and data availability. The noise generator is based on a hardware true random source to ensure that attackers cannot reverse-infer the original value through a probability model.
[0020] The rate-distortion optimization engine collaborates with the differential privacy module: The preliminary de-identified data generated by the rate-distortion model is used as the input for differential privacy processing, forming a double privacy protection layer. The former eliminates structural privacy leakage through information-theoretic constraints, and the latter resists inference attacks based on background knowledge through random noise.
[0021] The dynamic adjustment unit interacts with the noise parameter library: According to the real-time monitored privacy attack situation (model reconstruction accuracy), dynamically adjust the privacy threshold of the rate-distortion model and the budget allocation of differential privacy to form a closed-loop feedback.
[0022] The sensitivity calculator is linked with the noise injection unit: For different medical test items (blood routine, genetic testing), pre-compute the global sensitivity of each field and store it in the parameter library to guide the dynamic adaptation of the noise scale.
[0023] Preferably, the rate-distortion theory optimization model is solved through the following objective function: ; where is the mutual information between the original data and the de-identified data ; is the data utility loss, is the original utility attribute, is the attribute after de-identification; is the utility loss weight factor, with a value range of 1.0 to 5.0; is the privacy budget, with a value range of 0.1 to 0.5 bits.
[0024] Construct an optimization goal centered on minimizing the mutual information between original data and desensitized data, while constraining the amount of sensitive attribute leakage. By quantitatively analyzing the dynamic balance between data relevance and privacy risk, a composite objective function including utility loss weights is designed. The optimization engine dynamically adjusts the conditional probability distribution based on an iterative algorithm to ensure that sensitive information leakage is suppressed to the greatest extent possible while meeting the clinical data availability requirements.
[0025] Mutual information minimization eliminates the statistical correlation between data through the principle of information theory and cuts off the path of sensitive attribute inference; the utility loss function introduces Manhattan distance to quantify the impact of data degradation on downstream tasks, and dynamically adjusts the balance between privacy protection strength and data availability through weight factors. The privacy budget threshold is used as a hard constraint to ensure that sensitive information leakage is always within a controllable range.
[0026] The utility loss weight factor and privacy budget parameters are dynamically configured according to the data usage scenario. The weight factor adjustment module combines the data field type (continuous test value, discrete classification label) and business requirements (emergency retrieval requires high data fidelity) to adjust the optimization direction in real time. The privacy budget management module monitors real-time attack risks and automatically tightens the privacy leakage threshold when a new attack mode is detected.
[0027] The weight factor acts as an optimization direction regulator. By increasing the weight, data utility can be prioritized (maintaining the accuracy of blood sugar values), and vice versa, privacy protection can be strengthened. The privacy budget threshold acts as a safety valve, combined with attack situation awareness to dynamically shrink or relax the constraint boundary to form an adaptive defense system.
[0028] Preferably, the zero-knowledge proof protocol includes: Compile user permission declarations into arithmetic circuit constraints to generate zero-knowledge proofs based on elliptic curve bilinear pairings; The verification process does not transmit the user's real identity information, and is bound to the device fingerprint and temporary access token.
[0029] The user permission declaration is converted into an arithmetic circuit constraint, and the natural language policy ("attending physician of a tertiary hospital") is mapped into a Boolean logic operation sequence through a logic expression parsing engine. The compiler performs structured encoding of the permission declaration based on a preset rule base and generates a circuit topology containing multiple layers of logic gates to ensure that the permission verification conditions can be expressed in a mathematical formal way.
[0030] Arithmetic circuits transform complex permission logic into verifiable mathematical constraints, eliminating natural language ambiguity; multi-layer logic gate design supports nested conditional judgments to achieve fine-grained permission control; circuit topology optimization reduces verification calculations and improves the efficiency of zero-knowledge proof generation.
[0031] Generate zero - knowledge proofs based on elliptic - curve bilinear pairings, and utilize the discrete logarithm problem in the elliptic - curve group to ensure that the proofs cannot be forged. The proof generator generates a temporary key pair through a random - number seed, and constructs proof parameters in combination with a common reference string, enabling the verifier to only confirm the validity of the claim and preventing the reverse - inference of the user's private credentials.
[0032] The bilinear pairing provides non - interactive proof capabilities, and verification can be completed with a single communication. The mathematical properties of the elliptic - curve group guarantee the strong security of the proof. Even if an attacker obtains the public parameters, they cannot forge a valid proof. The temporary key mechanism implements "one - time - one - key" to prevent privacy leakage caused by the reuse of proofs.
[0033] Generate a temporary access token after verification passes, and cryptographically bind it to the fingerprint of the requesting device. The fingerprint collection module extracts the hardware features of the device (MAC address, processor serial number), generates a unique identifier through a hash function, and encrypts it in combination with the token expiration field and stores it in the verification log.
[0034] Device fingerprints provide hardware - level identity identification, enhancing the anti - tampering ability of authentication. The dynamic binding of tokens ensures that even if the credentials are leaked, unauthorized devices cannot be used. Hash encryption protects fingerprint privacy and prevents the acquisition of original hardware information through reverse engineering.
[0035] The logic compiler and the circuit optimizer work together: the original circuit generated by the compiler is compressed in the logic level by the optimizer, reducing the computational complexity of subsequent proof generation.
[0036] The proof generator and the verification gateway cooperate: the proof parameters output by the generator are adapted to the bilinear - pairing verification algorithm of the verification gateway to ensure protocol consistency.
[0037] The token manager and the audit module interact: the information of the dynamically - bound token is synchronized to the audit log in real - time, providing data support for tracing abnormal access.
[0038] Preferably, the elliptic - curve parameters satisfy: ; And the verification process performs bilinear - pairing operations: ; In the formula, , , are proof parameters, , are the generators of the elliptic - curve group, is a hash function.
[0039] Construct a cryptographic group using a specific elliptic curve equation and modulus parameters. The modulus is selected as a standardized large prime number to ensure that group operations satisfy the mathematical assumptions of the discrete logarithm problem. The curve parameters are verified by international cryptographic standards to eliminate the risk of backdoor implantation and provide compatibility support for bilinear pairing operations.
[0040] The discrete logarithm problem of the elliptic curve group is the basis of security, ensuring that attackers cannot deduce private information through known public parameters; the large prime number modulus enhances the randomness of the distribution of group elements and resists attack methods based on small group substructures; the standardized parameters ensure cross-platform compatibility and adapt to the interoperability requirements between heterogeneous medical systems.
[0041] The verification process is based on the mathematical properties of bilinear mapping, and the validity of the proof is confirmed through pairing operations between group elements. A random perturbation factor is introduced during the generation of proof parameters, making the intermediate variables generated in each verification unpredictable, preventing replay attacks or man-in-the-middle tampering.
[0042] Bilinear pairing maps group elements to an extension field, and uses its non-degeneracy and computability to achieve the mathematical binding of the proof; the random perturbation factor destroys the linear relationship model constructed by the attacker, ensuring the independence and security of each verification process; the uniqueness of the pairing result guarantees that the proof cannot be forged or tampered with.
[0043] The public reference string and the group generator are dynamically updated according to a preset period, and the update strategy combines the abnormal access records in the blockchain audit log. When a potential key leakage risk is detected, the group parameter rotation mechanism is triggered to regenerate the elliptic curve base point and discard the historical parameters.
[0044] The dynamic parameter update cuts off the relevance between historical data and the current system, achieving forward security; the blockchain audit provides a reliable basis for triggering abnormal events, ensuring the objectivity of the parameter rotation decision; the list of discarded parameters is synchronized to all verification nodes to prevent malicious reuse of expired proofs.
[0045] Preferably, the quantum-resistant attribute-based encryption includes: Adopt a lattice cryptography algorithm based on the RLWE problem and configure the dimension of the polynomial ring , modulus ; Generate hierarchical ciphertexts: the edge cloud ciphertext is encrypted using LWE, and the central cloud ciphertext is encrypted using BGV fully homomorphic encryption.
[0046] Construct an encryption system using a lattice cryptography algorithm based on the RLWE (Learning with Errors over Rings) problem, and achieve quantum-resistant attack capabilities through the configuration of polynomial ring structures and modulus parameters. The selection of the polynomial ring dimension and modulus balances computational efficiency and security, ensuring that encryption operations can be quickly completed in an edge computing environment while meeting the long-term confidentiality requirements of medical data.
[0047] The mathematical complexity of the RLWE problem is based on the shortest vector problem in lattice theory, which cannot be cracked by quantum computers even in polynomial time; the high-dimensional polynomial ring expands the key space to resist brute-force attacks; the hierarchical modulus design optimizes the growth of ciphertext noise to ensure the feasibility of fully homomorphic operations.
[0048] The edge cloud nodes adopt a lightweight LWE encryption scheme to quickly encrypt data with high timeliness requirements; the BGV fully homomorphic encryption algorithm is deployed in the persistent layer of the central cloud to support statistical analysis operations in the ciphertext state. The encryption policy generator dynamically matches the encryption level according to the data classification labels (inspection reports, image files) to ensure that high-value data is protected at a higher level.
[0049] LWE encryption achieves high efficiency through linear operations and adapts to the limited computing resources of edge nodes; the BGV algorithm supports ciphertext addition and multiplication operations to meet the data mining requirements of the central cloud; the hierarchical architecture realizes the scenario adaptation of security and availability, avoiding the performance bottleneck of a single encryption policy.
[0050] The attribute policy parser encodes the access control conditions ("department: cardiovascular medicine") as mathematical constraints on the polynomial ring and binds them to the ciphertext. When generating the decryption key, it is necessary to verify whether the attribute set of the request side satisfies the policy expression bound to the ciphertext, and the key parameters match the current encryption level.
[0051] Attribute binding realizes policy embedding through polynomial multiplication on the ring to ensure the inseparability of the ciphertext and the access conditions; hierarchical matching verification prevents low-security-level keys from decrypting high-protection-level data; the dynamic association mechanism realizes fine-grained access control, avoiding the static permission defects of the traditional RBAC model.
[0052] The parameter configurator and the key generator are linked: according to the hardware performance differences between the edge and central nodes, the dimensions and modulus parameters of the polynomial ring are dynamically adjusted to generate encryption keys suitable for different levels.
[0053] The attribute policy engine interacts with the hierarchical encryption unit: the policy conditions are compiled into mathematical constraints in real time to drive the edge cloud and the central cloud to select the corresponding encryption algorithms and parameter sets.
[0054] The noise controller cooperates with the fully homomorphic operation module: monitors the BGV ciphertext noise level, triggers the ciphertext refresh operation when the noise approaches the security threshold, and maintains the feasibility of the fully homomorphic operation.
[0055] Preferably, the LWE encryption process is expressed as: ; Where: is the public matrix, is the random vector; , is a discrete Gaussian distribution error vector with a standard deviation ; is a common parameter matrix.
[0056] Based on lattice cryptography theory, a high-dimensional integer matrix is constructed as the common parameter of the encryption system. The design of the matrix dimension and modulus follows the anti-quantum attack standard. The matrix elements are filled by a provably secure random number generator to ensure that their linear independence satisfies the mathematical assumptions of the learning with errors problem. The matrix update period is associated with the blockchain audit log, and when an abnormal access behavior is detected, matrix reconstruction is triggered to cut off the correlation between the historical ciphertext and the current system.
[0057] As the basic carrier of encryption operations, the security of the common matrix depends on the computational complexity of the shortest vector problem in lattice theory. Even if an attacker obtains the matrix, they cannot deduce the private key; the dynamic reconstruction mechanism achieves forward security and prevents the risk of deciphering historical ciphertexts caused by long-term key leakage; the matrix dimension and modulus parameters balance security strength and computational overhead to adapt to the resource constraints of edge nodes.
[0058] During the encryption process, a random vector and a discrete Gaussian distribution error vector are introduced to confuse the original data through linear combination and modular arithmetic. The random vector generator is based on a physical entropy source (hardware noise) to ensure uniqueness and unpredictability. The standard deviation parameter of the error vector is configured according to the data sensitivity level, and a larger perturbation intensity is used for high-classified data.
[0059] The random vector destroys the determinism of the encryption process and prevents chosen-plaintext attacks; the discrete Gaussian error masks the statistical characteristics of the original data through the probability distribution characteristics, making the ciphertext still indistinguishable under the quantum computing model; the hierarchical error strategy realizes the scenario adaptation of security and computational accuracy and avoids data inaccuracy caused by excessive noise.
[0060] The common matrix and error parameters are configured differently according to the data hierarchy (edge layer, central layer). The edge layer uses a low-dimensional matrix and a smaller modulus to improve the encryption speed, while the central layer uses high-dimensional parameters to ensure long-term security. The parameter switching engine monitors the ciphertext decryption failure rate in real time, and when the error rate exceeds the threshold, it triggers the parameter upgrade process.
[0061] The hierarchical parameters achieve the optimal balance between edge computing efficiency and central storage security; the dynamic monitoring mechanism adaptively adjusts the encryption intensity to avoid local vulnerabilities caused by fixed parameters; the parameter switching process is seamlessly connected to ensure that the continuity of the encryption service is not affected.
[0062] Preferably, the dynamic decryption includes: , geographical location and timestamp ; When the access scenario is a non-essential permission, apply a binary masking operation to the decrypted data: ; Among them, the decrypted data, the mask matrix is dynamically generated according to the role permissions.
[0063] Real-time collect the device fingerprint, geographical location and timestamp parameters of the request side, and construct a multi-dimensional access scenario portrait. The device fingerprint generates a unique identity identifier by fusing the hardware identifiers (MAC address, IMEI number) through a hash function. The geographical location obtains the longitude and latitude coordinates through the GPS / Beidou positioning interface. The timestamp is synchronized with the system clock to ensure timeliness. When an abnormal access scenario (non-working hours, off-site login) is detected, trigger dynamic decryption policy adjustment.
[0064] The device fingerprint provides hardware-level identity non-repudiation to prevent account theft; the geofencing technology restricts the geographical boundary of data access and blocks cross-regional unauthorized access; the timestamp is linked with the preset access policy to achieve hierarchical control of data availability based on time periods.
[0065] Dynamically generate a binary mask matrix according to the role permissions, and selectively hide the decrypted data. The mask generator is based on the principle of minimum necessity, and only retains the fields necessary for the current business scenario (display blood type and allergy history during emergency review, hide medical history details). The hidden field ratio is dynamically adjusted by the policy engine, and the upper limit does not exceed 30%. After the masking operation, an invisible digital watermark is embedded to record the operator's identity and time information.
[0066] The binary mask realizes precise control at the granularity of data fields through bit operations, avoiding an all-or-nothing access mode; the principle of minimum necessity ensures the availability of the core information of clinical diagnosis; the digital watermark provides the ability to trace the source of leakage and deter internal malicious behaviors.
[0067] When the role permissions change, the policy management center automatically issues new mask rules to the edge nodes. The mask matrix version number is bound to the user permission version, and the decryption gateway can load the latest rules only after verifying the version consistency. When it is detected that the masked data is accessed frequently, trigger a risk assessment and tighten the hidden ratio threshold.
[0068] The permission-mask version binding prevents data overexposure caused by lagging policy updates; the access frequency monitoring identifies potential data scraping behaviors and dynamically adjusts the protection intensity; the pre-loading of policies by the edge nodes reduces the decryption delay and ensures the smoothness of high-timeliness scenarios such as emergencies.
[0069] Preferably, the policy update includes: Evaluating the Success Rate of De-identification Data Attacks via Generative Adversarial Networks , dynamically adjusting the privacy budget : ; Among them, is the learning rate, is the success rate threshold of the attack; The BSDiff algorithm is used to generate a policy difference file and hot-update it to the edge nodes.
[0070] A privacy attack simulation environment is constructed through a generative adversarial network (GAN). The generator attempts to reconstruct the original sensitive information from the de-identified data, and the discriminator evaluates the accuracy of the reconstruction result. The success rate of the attack is calculated by counting the reconstruction accuracy of the generator on a preset test set. When the accuracy exceeds the threshold, a privacy budget tightening strategy is triggered. During the adversarial training process, a data distribution drift compensation mechanism is introduced to dynamically adjust the attack strategy of the generator to approximate the real attack behavior.
[0071] The generative adversarial network simulates the worst-case privacy attack through adversarial learning, exposing potential vulnerabilities in the de-identification strategy; the reconstruction accuracy quantifies the actual threat level of the attack, providing an objective basis for dynamic adjustment; the distribution drift compensation ensures that the attack model continuously tracks changes in data characteristics and avoids evaluation bias.
[0072] According to the difference between the attack success rate and the threshold, the privacy budget parameter is adjusted through a negative feedback control algorithm. When the attack success rate is continuously higher than the threshold, the privacy budget is exponentially reduced to enhance the protection intensity; when the attack success rate is lower than the security baseline, the budget is linearly relaxed to improve data utility. The adjustment process introduces smoothing filtering to avoid service oscillations caused by parameter mutations.
[0073] The negative feedback control establishes a dynamic balance relationship between the attack situation and the protection intensity, realizing adaptive defense; the exponential adjustment strengthens the ability to quickly respond to high-risk attacks, and the linear relaxation ensures the stability of data availability recovery; the smoothing filtering eliminates noise interference and ensures the robustness of the policy adjustment.
[0074] The BSDiff algorithm is used to generate a policy difference file and only transmit the changed part of the policy parameters to the edge nodes. The hot deployment engine loads the new policy in memory and maintains the old version for backup, and realizes the update without service interruption through traffic switching. The version rollback module monitors the running metrics (de-identification failure rate) of the new policy and automatically switches to the historical stable version in case of anomalies.
[0075] The differential update reduces the network transmission load and improves the update efficiency of the edge nodes; the in-memory hot loading avoids disk I / O latency and ensures real-time performance in high-concurrency scenarios; the version rollback mechanism provides update fault tolerance and ensures the continuous availability of the system.
[0076] The adversarial training platform is linked with the privacy budget manager: the attack success rate index is synchronously sent to the budget manager in real time to drive the closed-loop adjustment of privacy parameters.
[0077] The BSDiff engine interacts with the edge node policy library: the differential files are written through the incremental update interface of the edge computing node, reducing the bandwidth occupancy and storage overhead.
[0078] The attack log analyzer collaborates with the blockchain audit module: the abnormal samples used in the attack model training are synchronously stored on the blockchain for providing a verifiable data source for policy optimization.
[0079] A system based on the above method, comprising: A dynamic desensitization module, deployed on the edge node, with a rate-distortion optimization engine and a differential privacy noise injection unit built therein; A zero-knowledge authentication gateway, integrating an arithmetic circuit compiler and a bilinear pairing validator, supporting the zk-SNARK protocol; A hierarchical encryption unit, configured with a lattice cryptography parameter generator to implement hierarchical encryption of the LWE and BGV algorithms; A policy management center, connecting to the blockchain audit node and the attack model training platform to drive the closed-loop policy update.
[0080] The present invention provides a medical test data sharing method and system based on cloud authentication. It has the following beneficial effects: 1. The present invention adopts a technical solution of dynamic desensitization and rate-distortion theory optimization, achieving the technical effect of dynamic balance between privacy and utility. Compared with the problems of fixed static desensitization rules, overprotection or underprotection in the prior art, it solves the core deficiency of being unable to adapt to the multi-scenario data sharing requirements.
[0081] 2. The present invention realizes the identity authentication effect without trust dependence through the technical solution of zero-knowledge proof and dynamic binding of device fingerprints. Compared with the defect of easy leakage of privacy information in certificate or password verification in the prior art, it solves the key shortcoming of being unable to resist internal personnel's overstepping authority or misusing identities.
[0082] 3. Based on the technical solution of anti-quantum hierarchical encryption and dynamic attribute association, the present invention forms a quantum-safe fine-grained access control effect. Compared with the problem that a single encryption policy in the prior art is difficult to balance efficiency and long-term security, it solves the significant weakness of being unable to cope with quantum computing attacks and dynamic permission management.
[0083] 4. The present invention combines the technical solutions of context awareness and adversarial learning strategy iteration to construct the technical effect of an adaptive privacy protection system. Compared with the deficiencies of mechanical access control and static defense strategies in the prior art, it solves the inherent limitation of being unable to cope with new attacks and complex scenario evolution. Description of the Drawings
[0084] Figure 1 It is a schematic flowchart of the method of the present invention; Figure 2 It is a system architecture diagram of the present invention. Detailed Embodiments
[0085] Next, in combination with the drawings of the specification of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0086] Embodiment: Please refer to the attached Figure 1 - attached Figure 2 , the embodiment of the present invention provides a method for sharing medical test data based on cloud authentication, including the following steps: Step S1: Dynamic desensitization processing of edge nodes In the method for sharing medical test data based on cloud authentication, step S1, as the core link of data preprocessing, forms a technical closed-loop with subsequent authentication and encryption steps. Specifically, after the original test data is privacy-desensitized at the edge node, it will be uploaded to the cloud through an encrypted channel to provide input data that complies with privacy specifications for the zero-knowledge authentication in step S2.
[0087] In some embodiments, after the edge node receives the original medical test data, it first splits the data fields through a structured parsing engine. Specifically, for HL7 FHIR format data, the patient identifier, test item code, and numerical result fields are extracted. The sensitive attribute annotation module identifies the sensitive field set S based on a preset keyword library ("HIV positive", "tumor marker"). The utility attribute set U is dynamically selected according to the downstream task requirements, such as continuous numerical values of blood glucose level and white blood cell count.
[0088] Generally, the joint probability distribution P(X, S, U) is estimated by a sliding window statistical method: ; wherein, is an indicator function, is the historical data sample size.
[0089] In a possible implementation manner, the objective function is designed to minimize the mutual information between the desensitized data and the original data, and to constrain the leakage amount of sensitive attributes. Specifically, the optimization problem is expressed as: ; In the formula, represents the mutual information; is the utility loss function, select the Manhattan distance ; is an adjustable weight factor (typical value 1.0 - 5.0); is the privacy budget (typical value 0.1 - 0.5 bits).
[0090] In some embodiments, the optimal conditional distribution is iteratively calculated by an improved Blahut - Arimoto algorithm. Specifically: Initialize the conditional probability matrix as a uniform distribution; The iterative update rule is: ; where is the Lagrange multiplier, updated by an adaptive adjustment strategy: ; is the learning rate (typical value 0.01 - 0.1).
[0091] The termination condition is set such that the change in mutual information for 10 consecutive iterations is less than .
[0092] As an option, after obtaining the optimal de - sensitized data , Laplace noise is injected into its utility attribute to further enhance privacy protection. Specifically: ; The noise scale parameter is determined by the global sensitivity and the privacy budget : ; In the formula, , the global sensitivity of the blood glucose value is 200 mg / dL; is the differential privacy budget (typical value 0.1 - 1.0).
[0093] Step S2: Zero - knowledge identity authentication process In the method for sharing medical test data based on cloud authentication, step S2, as the core link of access control, forms a technical linkage with the desensitization process in step S1 and the subsequent encryption and decryption processes. Specifically, the desensitized data generated in step S1 needs to pass the zero-knowledge authentication in step S2 to ensure that only authorized users can access it, while avoiding the leakage of users' real identity information, providing a reliable basis for permission determination for attribute-based encryption in step S3.
[0094] In some embodiments, the identity statements submitted by doctor terminals adopt the JSON-LD structured format. Specifically, the statement content includes fields such as role identifiers, institutional codes, and permission time limits. The statement compiler converts natural language policies (such as "having the permissions of an emergency department doctor in Hospital A") into boolean logic expressions.
[0095] In a possible implementation manner, the arithmetic circuit generator constructs constraint conditions based on the R1CS (Rank-1 Constraint System) model. When verifying whether the digital certificate hash value exists in the pre-compiled legal list, the circuit includes the following constraints: ; In the formula, is the hash value of the certificate to be verified, is the legal hash list.
[0096] Generally, the zk-SNARK protocol uses the BN128 elliptic curve, and its parameters are defined as: ; In the trusted setup phase, the common reference string (CRS) is generated through secure multi-party computation. Specifically, the number of participants is not less than 3, and the random number seeds of all parties are mixed through the threshold signature protocol.
[0097] As another option, when it is necessary to be compatible with the national cryptography standard, the elliptic curve can be replaced by the SM2 curve, and its equation is: ; In some embodiments, the Groth16 algorithm is called in the proof generation phase to generate zero-knowledge proofs. Specifically, for the statement and the private credential , the calculation process of the proof satisfies: ; In the formula, is the public parameter in the CRS, , is a random number, is the output of the hash function, is the polynomial coefficient.
[0098] In the verification stage, a bilinear pairing operation is performed to determine whether the following equation holds: ; where, is a hash function, is a bilinear mapping, , are the generators of the elliptic curve groups , respectively.
[0099] In a possible implementation, a temporary access token is generated after authentication. Specifically, the token is valid for 15 minutes and is bound to the fingerprint information of the requesting device. The device fingerprint is generated through the following feature combination: ; In the formula, is the SHA3-256 hash function, represents the string concatenation operation.
[0100] The above implementation realizes high-security authentication without a trusted third party through zero-knowledge proof. In some extended embodiments: The hash function can be replaced with the BLAKE2b algorithm to improve the calculation efficiency; The declared policy can be extended to support time-lock conditions ("valid only from 9:00 to 18:00"), and in this case, a timestamp comparison constraint needs to be added to the arithmetic circuit; The bilinear pairing operation can adopt the Ate pairing to optimize the calculation speed.
[0101] Step S3: Hierarchical attribute-based encryption storage In the medical test data sharing method based on cloud authentication, step S3, as the core link of data security storage, forms a technical collaboration with the desensitization process in step S1 and the authentication result in step S2. Specifically, the desensitized data generated in step S1 needs to be hierarchically stored through the quantum-resistant encryption algorithm in step S3, and at the same time, the permission authentication result in step S2 will be dynamically bound to the ciphertext attribute policy, providing a policy matching basis for the context-aware decryption in the subsequent step S4.
[0102] In some embodiments, the quantum-resistant attribute-based encryption scheme is constructed based on the RLWE (Ring Learning With Errors) problem. Specifically, the following lattice cipher parameters are selected: Polynomial ring dimension: ; Modulus: (satisfying ); Error distribution: Discrete Gaussian distribution (standard deviation ).
[0103] In the master key generation phase, the central cloud platform generates a public matrix and a trapdoor basis , where . Specifically, the matrix is constructed by the NIST DRBG random number generator and satisfies: ; In the formula, is the publicly available Gadget matrix for implementing efficient lattice operations.
[0104] In a possible implementation, the access policy is expressed through monotone Boolean logic. "Emergency department doctor AND Hospital A" can be encoded as the attribute vector , where is the total number of system preset attributes. The user private key generation algorithm applies the GPV (Gentry - Peikert - Vaikuntanathan) sampling technique: ; Specifically, when the attribute vector satisfies the policy, the trapdoor basis can derive the corresponding decryption key.
[0105] In some embodiments, the desensitized data is encrypted in layers according to timeliness: 1. Edge cloud cache layer: Encrypts recent data (within 7 days) using lightweight LWE encryption: ; In the formula, is the public matrix, is the random vector, , is the discrete Gaussian distribution error vector with standard deviation , is the public parameter matrix.
[0106] The ciphertext - associated attribute label ("Department: Emergency") is stored in the regional edge node.
[0107] 2. Central cloud persistent layer: Encrypts long - term data using the fully homomorphic BGV scheme: ; In the formula, is the random vector, is the error term.
[0108] The ciphertext policy is extended to a time constraint expression ("validity period ≤ 5 years") and stored in the central cloud cold storage.
[0109] As an option, when the access policy changes, an automatic re-encryption process is triggered. Specifically, the differences between the old and new policies Update the ciphertext through the following operations: ; In the formula, is a newly added random vector to ensure forward security.
[0110] The above implementation achieves a balance between storage efficiency and security through hierarchical encryption. In some extended embodiments: The edge cloud cache layer can adopt the LRU replacement algorithm to manage the storage space; The central cloud cold storage can be configured in a 6 + 3 erasure code redundancy mode to improve the disaster tolerance ability; The attribute vector encoding can be replaced with a secret sharing scheme based on Shamir threshold.
[0111] Step S4: Context-aware dynamic decryption In the method for sharing medical test data based on cloud authentication, step S4, as the core link of data access control, forms a technical linkage with the authentication result of step S2 and the encryption storage policy of step S3. Specifically, after verifying the user's permission through zero-knowledge authentication in step S2, step S4 dynamically decrypts the ciphertext generated in step S3 according to the real-time access scenario, and performs secondary desensitization in combination with the context-aware policy to ensure that the data meets clinical requirements under the premise of minimizing the risk of privacy leakage.
[0112] In some embodiments, after the cloud receives the decryption request, it first parses the attribute policy bound to the ciphertext. Specifically, for the edge cloud ciphertext generated in step S3 , , the attribute matcher performs a fast Bloom filter operation to verify whether the request-side attribute set meets the policy tree .
[0113] In a possible implementation, the dynamic decryption key generation adopts the lattice rounding algorithm: ; In the formula, is the user's private key derived from step S3, is the rounding function by the power of 2, which is used to eliminate noise interference.
[0114] Specifically, the context-aware engine collects the following parameters in real time: Device fingerprint: Calculate the unique device identifier through a hash function: ; In the formula, the SM3 national cryptographic hash algorithm is selected, indicating the string concatenation operation.
[0115] Geographical location: Call the GeoIP2 database to parse the IP address and obtain the longitude and latitude coordinates .
[0116] Timestamp: Synchronize the NTP server to obtain the atomic clock time , and verify whether the time window is within the range allowed by the policy.
[0117] When an unconventional access scenario (non-working hours or unfamiliar device) is detected, trigger the secondary desensitization rule library.
[0118] As an option, the decrypted data is field-hidden according to preset rules. Specifically, for access requests from non-attending physicians, bitmask operations are used to hide sensitive fields: ; In the formula, is the binary mask matrix, and its element is defined as follows: ; At the same time, an invisible digital watermark is embedded in the data stream, and the watermark strength parameter is: ; Ensure that the impact of the watermark on downstream analysis tasks does not exceed 5%.
[0119] In a possible implementation, the system monitors the response delay of the decryption request . When , start the edge node bandwidth expansion process: ; In the formula, is the current bandwidth, is the adjusted bandwidth. If the request times out continuously for 5 times, trigger the fuse mechanism and suspend the user's access permission for 30 minutes.
[0120] Step S5: Blockchain auditing and policy update In the method for sharing medical test data based on cloud authentication, step S5, as the feedback optimization link of the technical closed-loop, forms a dynamic association with the decryption access log in step S4 and the rate-distortion optimization parameter in step S1. Specifically, the operation records generated in step S4 are subject to non-tamperable auditing through the blockchain evidence storage in step S5, and the rate-distortion optimization parameter in step S1 is updated according to the latest privacy attack model, forming an adaptive enhanced privacy protection mechanism.
[0121] In some embodiments, the audit logs are stored in the form of key-value pairs, including the operation time, user fingerprint, and data hash field. Specifically, the log entries are structured and encoded in the ProtocolBuffers format, and the hash value is calculated through the SM3 national cryptography algorithm: ; wherein, is a random number (128-bit length) used to prevent rainbow table attacks. The block hash calculation adopts Merkle tree aggregation: ; wherein, is the Merkle root of the hashes of all log entries in this block, is the proof-of-work random number.
[0122] In a possible implementation, the blockchain network adopts the PBFT (Practical Byzantine Fault Tolerance) consensus algorithm. Specifically, for each new block, at least nodes need to reach an agreement, where is the total number of nodes (typical value ).
[0123] The smart contract is deployed with the following functions: Policy update trigger: When a new type of privacy attack pattern is detected, a policy optimization request is automatically generated; Version rollback: If the accuracy of downstream tasks decreases by more than the threshold , roll back to the previous generation of policies.
[0124] In some embodiments, the privacy attack model is constructed using a generative adversarial network (GAN). The generator attempts to reconstruct the original sensitive attribute from the de-sensitized data , and the discriminator distinguishes between real and generated data. The objective function is: ; After the training converges, the attack success rate is used to adjust the privacy budget in step S1: ; wherein, is the learning rate (typical value 0.1), is the threshold (20%).
[0125] As an option, the policy update file is distributed through an incremental transfer protocol. Specifically, the differential file is calculated as follows: ; In the formula, is a binary differential algorithm, and the compression ratio can reach more than 90%. After the edge node receives , it loads the new policy in real time through the memory patching technology: ; Version management retains the snapshots of the latest three generations of policies, and the snapshot metadata is stored in the IPFS distributed network.
[0126] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for sharing medical test data based on cloud authentication, characterized in that, It includes the following steps: Dynamically desensitize the original medical test data at the edge node to generate desensitized data that meets the privacy-utility balance; Verify the identity and permissions of the requesting user through the zero-knowledge proof protocol to generate an authentication result bound to the privacy policy; Hierarchically encrypt the desensitized data based on the quantum-resistant attribute-based encryption algorithm to generate ciphertext dynamically associated with access attributes; Dynamically decrypt the data according to real-time context parameters and perform secondary desensitization processing in combination with the access scenario; Write the data operation record into the blockchain for audit traceability, and iteratively update the desensitization policy according to the privacy attack model.
2. The method for sharing medical test data based on cloud authentication according to claim 1, wherein The dynamic desensitization processing includes: Construct a rate-distortion theory optimization model with the goal of minimizing the mutual information between the desensitized data and the original data, and restricting the leakage amount of sensitive attributes not to exceed a preset threshold; Overlay differential privacy noise on the desensitized data, and the noise scale parameter is determined by the ratio of the global sensitivity of the utility attribute to the privacy budget.
3. The method for sharing medical test data based on cloud authentication according to claim 2, wherein The rate-distortion theory optimization model is solved through the following objective function: ; Where, is the original data and the de-identified data of the mutual information; is the data utility loss, is the original utility attribute, is the attribute after desensitization; is the utility loss weight factor, and its value range is 1.0 to 5.0; is the privacy budget, with a value range of 0.1 to 0.5 bits.
4. The method for sharing medical test data based on cloud authentication according to claim 1, wherein The zero-knowledge proof protocol includes: Compile the user permission statement into arithmetic circuit constraints to generate a zero-knowledge proof based on elliptic curve bilinear pairing; The verification process does not transmit the user's real identity information, and binds the device fingerprint to the temporary access token.
5. The method for sharing medical test data based on cloud authentication according to claim 4, wherein The elliptic curve parameters satisfy: ; And the verification process performs a bilinear pairing operation: ; Wherein, , , are proof parameters, , is the generator of the elliptic curve group, is the hash function.
6. The method for sharing medical test data based on cloud authentication according to claim 1, wherein, The quantum-resistant attribute-based encryption includes: Adopt a lattice cryptography algorithm based on the RLWE problem and configure the dimension of the polynomial ring , modulus ; Generate hierarchical ciphertext: the edge cloud ciphertext is encrypted using LWE, and the central cloud ciphertext is encrypted using BGV fully homomorphic encryption.
7. The method for sharing medical test data based on cloud authentication according to claim 6, wherein The LWE encryption process is expressed as: ; Where: is a common matrix, is a random vector; , is a discrete Gaussian distribution error vector with a standard deviation of ; is a common parameter matrix.
8. The method for sharing medical test data based on cloud authentication according to claim 1, wherein The dynamic decryption includes: , geographical location and time stamp ; When the access scenario is a non-essential permission, apply a binary mask operation to the decrypted data; ; Among them, the decrypted data, the mask matrix is dynamically generated according to the role permissions.
9. The method for sharing medical test data based on cloud authentication according to claim 1, wherein The policy update includes: Evaluating the success rate of desensitized data attacks through generative adversarial networks , dynamically adjusting the privacy budget : ; Among them, is the learning rate, is the attack success rate threshold; Use the BSDiff algorithm to generate a policy difference file and hot-update it to the edge node.
10. A system based on the method according to claim 1, characterized in that, It includes: A dynamic desensitization module deployed at the edge node, with a built-in rate-distortion optimization engine and a differential privacy noise injection unit; A zero-knowledge authentication gateway integrated with an arithmetic circuit compiler and a bilinear pairing validator, supporting the zk-SNARK protocol; A hierarchical encryption unit configured with a lattice cipher parameter generator to implement hierarchical encryption of the LWE and BGV algorithms; A policy management center that connects to the blockchain audit node and the attack model training platform to drive closed-loop policy updates.
Citation Information
Cited By
Data desensitization and integrity verification method and system based on differential privacy algorithm
CN120579227A
Data integrity verification method and system based on label embedding and homomorphic encryption
CN120602238A
A Data Integrity Verification Method and System Based on Tag Embedding and Homomorphic Encryption
CN120602238B
Cloud mobile phone equipment fingerprint disguising method and related equipment
CN120768541A
Low-voltage generator car non-inductive grid-connected control system and oscillation suppression method
CN120879672A