Secure backup and recovery method and system for medical data
By dynamic hierarchical and differentiated encryption of medical data, combined with cross-platform backup topology and multi-node collaborative decryption, the resource waste and security vulnerabilities in traditional backup methods are solved, and efficient and reliable data recovery and secure backup are achieved.
Patent Information
- Application Number
- CN202510753314.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing medical data backup methods, static classification strategies lead to resource waste and security vulnerabilities, which cannot effectively resist targeted attacks. The data recovery process lacks cross-platform verification and multi-node collaborative decryption mechanism, which can easily cause data tampering risks or recovery results are untrusted.
By obtaining multi-dimensional attribute information of medical data, dynamic hierarchy, encrypting and building a cross-platform backup topology, monitoring abnormal node behavior in real time, tracing the source verification and performing multi-node collaborative decryption and reconstruction, a trusted recovery chain is generated to ensure the reliability and security of data recovery.
It realizes the precise matching of hierarchical storage policies and encryption strength, enhances the disaster recovery capabilities of the backup system, prevents data tampering and leakage, ensures the integrity and auditability of data recovery, and reduces redundant storage costs.
Smart Images

Figure CN120256208A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical data, and particularly to a method and system for secure backup and recovery of medical data. Background Art
[0002] As the core carrier of patient diagnosis and treatment records, the secure storage and reliable recovery of medical data are key links to ensure the stable operation of the medical system and the privacy security of patients. With the exponential growth of the scale of medical data and the increasing demand for cross-institutional sharing, traditional backup methods face severe challenges in protecting data integrity, real-time anomaly detection, and trusted recovery mechanisms in complex network environments. In the prior art, medical data backup mostly adopts static classification strategies and unified encryption standards, and fails to implement hierarchical storage according to dynamic attributes such as data sensitivity and access frequency, resulting in waste of storage resources and difficulty in resisting targeted attacks. At the same time, the anomaly monitoring mechanism of the backup system mostly relies on fixed threshold rules, which cannot accurately identify collaborative attack behaviors between distributed nodes, and the data recovery process lacks cross-platform topology verification and multi-node collaborative decryption mechanisms, easily leading to problems such as data tampering risks or untrusted recovery results. Summary of the Invention
[0003] The main object of the present invention is to provide a method and system for secure backup and recovery of medical data, which can reduce the redundant storage cost while ensuring the reliability of data recovery.
[0004] To achieve the above object, the present invention provides a method for secure backup and recovery of medical data, including: Obtaining multi-dimensional attribute information of medical data, dynamically classifying the multi-dimensional attribute information based on a preset classification rule to obtain an attribute information label; Differentially encrypting the medical data according to the attribute information label to obtain encrypted medical data, writing the encrypted medical data into multi-level storage nodes, and performing multi-level backup construction to obtain a cross-platform backup topology; Performing anomaly monitoring on the multi-level storage nodes according to the cross-platform backup topology to obtain abnormal node data access behaviors, extracting replica node information from the multi-level storage nodes according to the abnormal node data access behaviors, and updating the topology of the cross-platform backup topology to obtain a backup topology configuration; Extracting traceable node information from the multi-level storage nodes according to the abnormal node data access behaviors, and performing data traceability verification on the traceable node information according to the backup topology configuration to obtain a trusted recovery chain; Selecting recovery nodes from the multi-level storage nodes according to the trusted recovery chain, and performing multi-node collaborative decryption and reconstruction to obtain a recovery verification report.
[0005] Further, the method for obtaining multi-dimensional attribute information of medical data and dynamically grading the multi-dimensional attribute information based on a preset classification rule to obtain an attribute information label includes: Separating the structural features of the multi-dimensional attribute information to obtain a timestamp sequence, a data format identifier, a device code, and a patient biometric index; Performing offset normalization processing on the timestamp sequence according to the classification rule to obtain a cross-platform time synchronization sequence; Performing format compatibility parsing on the data format identifier based on the classification rule to obtain a metadata format descriptor; Constructing a position association between the device code and the patient biometric to obtain a joint dimension map; Performing multi-modal feature fusion on the time synchronization sequence, the metadata format descriptor, and the joint dimension map to obtain an initial grading cluster; Allocating data retention periods and access permission threshold weights to the initial grading cluster according to preset privacy compliance constraint conditions to obtain a grading coefficient matrix; Performing incremental iterative adjustment on the grading coefficient matrix to obtain the attribute information label.
[0006] Further, the method for differentially encrypting the medical data according to the attribute information label to obtain encrypted medical data, writing the encrypted medical data into multi-level storage nodes, and performing multi-level backup construction to obtain a cross-platform backup topology includes: Constructing an encryption policy based on the attribute information label and the medical data to obtain an encryption algorithm priority list and a key segmentation threshold matrix; Performing hybrid encryption on the medical data based on the encryption algorithm priority list to obtain an encrypted dynamic key group; Performing stream cipher block processing on the medical data according to the key segmentation threshold matrix to obtain encrypted data packets; Mixing the encrypted dynamic key group and the encrypted data packets to obtain a comprehensive encrypted node packet; Identifying nodes of the multi-level storage nodes to obtain local storage nodes, off-site cold storage nodes, and cloud storage nodes; Storing the comprehensive encrypted node packet into the local storage nodes, the off-site cold storage nodes, and the cloud storage nodes respectively, and marking the local storage nodes as hot backup units; Performing erasure code redundant sharding processing on the hot backup units to obtain a shard set, and performing check identification on the off-site cold storage nodes and the cloud storage nodes to obtain a check shard set; Perform backup topology construction on the shard set and the verification shard set to obtain the cross-platform backup topology.
[0007] Further, perform anomaly monitoring on the multi-level storage nodes according to the cross-platform backup topology to obtain abnormal node data access behaviors. Extract replica node information from the multi-level storage nodes based on the abnormal node data access behaviors, and perform topology update on the cross-platform backup topology to obtain the backup topology configuration, including: Collect node monitoring logs from the multi-level storage nodes according to the cross-platform backup topology to obtain a set of node behavior baseline parameters; Perform analysis on the deviation of access behavior thresholds for the set of node behavior baseline parameters to obtain the abnormal node data access behaviors; Perform clustering and grouping processing on the request source address and operation type based on the abnormal node data access behaviors to obtain a list of abnormal node addresses; Perform reverse link node tracing extraction on the multi-level storage nodes according to the list of abnormal node addresses to obtain abnormal replica node codes and abnormal node storage information; Perform hash chain backtracking verification processing on the replica node codes to obtain data integrity verification information; Perform path isolation processing on the cross-platform backup topology based on the data integrity verification information, and perform topology update in combination with the abnormal node storage information to obtain the backup topology configuration.
[0008] Further, perform reverse link node tracing extraction on the multi-level storage nodes according to the list of abnormal node addresses to obtain abnormal replica node codes and abnormal node storage information, including: Perform request feature filtering on the list of abnormal node addresses to obtain a set of reverse tracing parameters; Analyze the link path dependency relationship for the set of reverse tracing parameters according to the backup topology configuration to obtain a reverse tracing path map; Perform replica node verification requests on the multi-level storage nodes based on the reverse tracing path map, and perform consistency consensus verification to obtain a list of candidate nodes; Perform hash chain tampering backtracking on the list of candidate nodes, and extract interrupted node codes to obtain a set of abnormal replica node codes; Perform key binding and storage path association processing on the set of abnormal replica node codes according to the backup topology configuration to obtain the abnormal node storage information.
[0009] Further, extracting traceable nodes from the multi-level storage nodes according to the abnormal node data access behavior to obtain traceable node information, and performing data traceability verification on the traceable node information according to the backup topology configuration to obtain a trusted recovery chain, including: Performing request log correlation extraction on the abnormal node data access behavior to obtain abnormal node log data; Based on the abnormal node log data, extracting traceable nodes from the multi-level storage nodes to obtain a set of traceable node metadata; Performing shard path dependency parsing on the set of traceable node metadata according to the backup topology configuration to obtain a shard verification path graph; Based on the shard verification path graph, sending a shard hash value synchronization request to the multi-level storage nodes and performing multi-node consistency verification to obtain a node trusted verification set; Sorting the node trusted verification set by shards and performing shard index level combination to obtain a node signature recovery chain; Based on a preset non-interactive zero-knowledge proof protocol, comparing the integrity of the initial records of the blockchain for the node signature recovery chain to obtain the trusted recovery chain.
[0010] Further, selecting recovery nodes from the multi-level storage nodes according to the trusted recovery chain and performing multi-node collaborative decryption and reconstruction to obtain a recovery verification report, including: Evaluating the node signature weights of the trusted recovery chain to obtain a list of recovery candidate nodes; Calculating a redundant shard distribution path coverage set for the multi-level storage nodes according to the list of recovery candidate nodes to obtain a set of recovery nodes; Based on the attribute information tags, performing dynamic key matching on the set of recovery nodes to obtain a node key group and a decryption order sequence; Performing initial collaborative decryption processing on the set of recovery nodes through a preset distributed key sharding and the node key group, and decrypting the shard data layer by layer to obtain a set of original data shards; Performing redundancy check and sequential alignment on the set of original data shards according to the decryption order sequence, and performing shard recombination according to a preset reverse erasure code algorithm to obtain a preliminary recovery result; Performing data integrity recovery on the preliminary recovery result based on the trusted recovery chain to obtain the recovery verification report.
[0011] The present invention also provides a medical data security backup and recovery system, which is applied to the medical data security backup and recovery method described in any one of the above, including: A collection module, which is used to obtain multi-dimensional attribute information of medical data, dynamically classify the multi-dimensional attribute information based on a preset classification rule, and obtain an attribute information label; An analysis module, which is used to differentially encrypt the medical data according to the attribute information label to obtain encrypted medical data, write the encrypted medical data into a multi-level storage node, and perform multi-level backup construction to obtain a cross-platform backup topology; An association module, which is used to perform anomaly monitoring on the multi-level storage node according to the cross-platform backup topology to obtain abnormal node data access behavior, extract replica node information from the multi-level storage node according to the abnormal node data access behavior, and update the topology of the cross-platform backup topology to obtain a backup topology configuration; A processing module, which is used to extract a traceability node from the multi-level storage node according to the abnormal node data access behavior to obtain traceability node information, and perform data traceability verification on the traceability node information according to the backup topology configuration to obtain a trusted recovery chain; A control module, which is used to select a recovery node from the multi-level storage node according to the trusted recovery chain, and perform multi-node collaborative decryption and reconstruction to obtain a recovery verification report.
[0012] A method and system for secure backup and recovery of medical data provided by the present invention has the following beneficial effects: Through a dynamic classification mechanism based on multi-dimensional attribute information, it is possible to dynamically divide the data security level according to the privacy level, access frequency, and storage period of medical data, realize the precise matching of the hierarchical storage strategy and encryption intensity, and solve the problems of resource waste and security vulnerabilities caused by traditional static classification. Through the construction of differential encryption and cross-platform backup topology, combined with the redundant sharding distribution of multi-level storage nodes, it can effectively resist the risk of data tampering or leakage of a single storage node and enhance the disaster tolerance of the backup system; by means of real-time monitoring of abnormal node data access behavior and dynamic extraction of replica node information, it is possible to quickly locate and isolate abnormal operation nodes in a distributed network, preventing the spread of malicious attacks in the backup link. Based on the traceability verification of the trusted recovery chain and multi-node collaborative decryption and reconstruction, the integrity and auditability of the data recovery process are ensured through tree-shaped hash aggregation and zero-knowledge proof technology, avoiding the problem of untrusted data caused by single-point failure in traditional centralized recovery mechanisms. Through the dynamic update of the backup topology configuration and the calculation of the recovery node coverage set, the resource scheduling efficiency of multi-level storage nodes is optimized, reducing the redundant storage cost while ensuring the reliability of data recovery. Description of the Drawings
[0013] Figure 1It is a flowchart of a method for secure backup and recovery of medical data provided by the present invention; Figure 2 It is a structural diagram of a system for secure backup and recovery of medical data provided by the present invention.
[0014] The realization of the object, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments
[0015] In order to make the object, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0016] Next, with reference to the accompanying drawings and specific embodiments, the present invention will be further described.
[0017] Referring to Figure 1 As shown, the present invention provides a method for secure backup and recovery of medical data, including: Step S1: Obtain multi-dimensional attribute information of medical data, and dynamically classify the multi-dimensional attribute information based on a preset classification rule to obtain attribute information tags; Step S2: Differentially encrypt the medical data according to the attribute information tags to obtain encrypted medical data, write the encrypted medical data into multi-level storage nodes, and perform multi-level backup construction to obtain a cross-platform backup topology; Step S3: Perform anomaly monitoring on the multi-level storage nodes according to the cross-platform backup topology to obtain abnormal node data access behaviors, extract replica node information from the multi-level storage nodes according to the abnormal node data access behaviors, and update the topology of the cross-platform backup topology to obtain a backup topology configuration; Step S4: Extract traceable nodes from the multi-level storage nodes according to the abnormal node data access behaviors to obtain traceable node information, and perform data traceability verification on the traceable node information according to the backup topology configuration to obtain a trusted recovery chain; Step S5: Select recovery nodes from the multi-level storage nodes according to the trusted recovery chain, and perform multi-node collaborative decryption and reconstruction to obtain a recovery verification report.
[0018] Based on the above steps, the detailed step process is as follows: Step S1: The multi-dimensional attribute information of medical data includes structured and unstructured features such as data sensitivity (e.g., patient privacy level), data type (image, text, gene sequence), access frequency, storage period, and compliance requirements. The dynamic grading process is based on preset classification rules. Through the data feature parsing engine, feature extraction and semantic annotation are performed on the original attribute information. For example, through natural language processing technology, sensitive fields in text data (such as names, diagnosis results) are identified, or the storage period of image data is determined through metadata analysis. The classification rules are dynamically adjusted according to the data governance strategy. For example, the data sensitivity weight is updated by combining real-time access logs, or the compliance label is redefined based on the latest privacy regulations (such as GDPR). The dynamic grading engine maps the multi-dimensional attributes to a grading matrix, generating an attribute information label containing parameters such as data encryption priority, storage node type, and backup redundancy. This label provides a decision basis for subsequent differential encryption and storage strategies, ensuring that high-sensitive data is preferentially assigned high-strength encryption algorithms and multi-copy storage, while low-frequency access data uses lightweight encryption to reduce system load.
[0019] Step S2: Differential encryption selects an appropriate encryption algorithm and key management strategy based on the security level parameter in the attribute information label. For example, AES-256 encryption combined with attribute-based encryption (ABE) technology is used for high-sensitive medical image data to ensure that only authorized roles can decrypt it; the national cipher SM4 algorithm is used for medium- and low-sensitive text data to balance security and processing efficiency. After encryption, the data shards are written to multi-level storage nodes according to the storage node type parameter in the label, including local edge nodes (low-latency access), private cloud nodes (medium security), and public cloud nodes (high availability). The multi-level backup is constructed through a redundant shard distribution algorithm. For example, each data shard and its replicas are dispersed and stored on different platform nodes according to a preset redundancy (such as 3 replicas), and the topological dependency relationship between the shards is recorded (such as the replicas of shard A are stored on nodes X, Y, and Z). The cross-platform backup topology describes the data links and redundant paths between nodes through a graph structure, forming a distributed storage network with balanced disaster tolerance and access efficiency, providing a baseline structure for subsequent anomaly monitoring.
[0020] Step S3: The anomaly monitoring module monitors the access behaviors of multi-level storage nodes in real time and conducts pattern analysis based on the cross-platform backup topology. By collecting node logs (such as access frequency, data traffic, response time), a node behavior baseline model is constructed. For example, the average access interval and data throughput of each node in historical data are statistically analyzed. The threshold deviation analysis engine compares the differences between real-time behaviors and the baseline model. For example, it detects that a certain node has an abnormally high frequency of access in a short period (possibly data scraping) or a sudden increase in data traffic (possibly malicious tampering). After identifying the abnormal node's data access behavior, the replica node information extraction module locates the redundant shard storage paths associated with this node through a reverse tracing algorithm. For example, it retrieves all adjacent nodes containing shard replicas of the abnormal node from the backup topology. The topology update engine dynamically adjusts the backup topology according to the abnormal behavior characteristics. For example, it isolates the abnormal node and reallocates its shard replicas to healthy nodes, or increases the redundancy of critical shards to enhance the anti-attack ability. The updated backup topology configuration is synchronized to all storage nodes to ensure that subsequent data recovery processes are executed based on the latest reliable paths.
[0021] Step S4: The source tracing node extraction locates the potential source of data tampering based on the mapping relationship between the operation logs in the abnormal node's data access behavior and the redundant paths in the backup topology configuration. Through the reverse link tracing algorithm, starting from the abnormal node, it traces back layer by layer along the shard dependency paths of the backup topology to identify all associated nodes participating in the storage or transmission of this data shard, generating a set of source tracing node information containing node addresses, shard version numbers, and operation timestamps. Data source tracing verification is achieved through a distributed consensus mechanism. For example, a shard hash value query request is sent to each node in the source tracing node set, and the returned results are consistency-checked based on the Byzantine fault tolerance protocol. If the hash values returned by the majority of nodes are consistent with the initial hash recorded in the backup topology, then this node is determined to be a trusted node; otherwise, it is marked as a suspicious node and the tampering evidence is recorded. The construction of the trusted recovery chain aggregates the shard hash values and associated metadata (such as shard indexes, node signatures) of all trusted nodes, and calculates the hash digest layer by layer using the Merkle tree structure. Finally, an immutable recovery chain containing the root hash value, shard hierarchical relationship, and verification timestamp is generated, providing an integrity verification benchmark for subsequent data recovery.
[0022] Step S5: The node selection for recovery is based on the shard-level relationship and node credibility scores in the trusted recovery chain. The minimum node set that can completely cover all shards of the data is filtered out through the minimum covering set algorithm. For example, nodes that store the latest version of the shards and have low network latency are preferentially selected, while ensuring that at least one trusted copy of each shard participates in the recovery. The collaborative decryption and reconstruction of multiple nodes rely on the encryption policy parameters in the attribute information tag, extract the dynamic key group (such as the shard key of the master key) bound to the shards from the backup topology configuration, and use the secure multi-party computation protocol (MPC) to achieve the secure reconstruction of the key. Specifically, each node participates in the calculation based on the locally stored key shards, jointly recovers the master key without exposing the plaintext key, and decrypts the encrypted shards layer by layer. The decrypted original data shards are redundantly verified and sequentially reorganized through the reverse erasure code algorithm. For example, the data block boundaries are aligned according to the shard index and timestamp, duplicate or damaged shards are removed, and finally, a complete data file is reorganized. The recovery verification report is generated by comparing the hash value of the reorganized data with the root hash value in the trusted recovery chain. If the consistency check passes, the recovery success status and time consumption are recorded; if there is a deviation, the location of the tampered shards is marked and an alarm is triggered, providing a basis for subsequent auditing.
[0023] A method for secure backup and recovery of medical data provided by the present invention can dynamically divide the data security level according to the privacy level, access frequency, and storage period of medical data through a dynamic grading mechanism based on multi-dimensional attribute information, realizing an accurate match between the hierarchical storage strategy and the encryption intensity, and solving the problems of resource waste and security vulnerabilities caused by traditional static classification. By constructing a differential encryption and cross-platform backup topology, combined with the redundant shard distribution of multi-level storage nodes, it effectively resists the risk of data tampering or leakage targeting a single storage node, enhancing the disaster tolerance of the backup system; by means of real-time monitoring of the abnormal data access behavior of nodes and dynamic extraction of replica node information, it can quickly locate and isolate abnormal operation nodes in the distributed network, preventing the spread of malicious attacks in the backup link. Based on the traceability verification and multi-node collaborative decryption and reconstruction of the trusted recovery chain, the integrity and auditability of the data recovery process are ensured through tree-shaped hash aggregation and zero-knowledge proof technology, avoiding the problem of untrusted data caused by single-point failure in the traditional centralized recovery mechanism. Through the dynamic update of the backup topology configuration and the calculation of the recovery node covering set, the resource scheduling efficiency of multi-level storage nodes is optimized, reducing the redundant storage cost while ensuring the reliability of data recovery.
[0024] In one embodiment, multi-dimensional attribute information of medical data is obtained, and the multi-dimensional attribute information is dynamically graded based on a preset classification rule to obtain an attribute information tag, including: The multi-dimensional attribute information of medical data includes heterogeneous data such as data generation time, storage format, acquisition device identifier, and patient biometric index. Through the data parsing engine, the original attribute information is split into a timestamp sequence, a data format identifier, a device code, and a patient biometric index. The timestamp sequence records a set of time points of data operations with millisecond-level precision, such as the creation and modification times of a certain medical image; the data format identifier is used to distinguish storage formats of different data structures such as DICOM and HL7; the device code uniquely identifies the data acquisition device; the patient biometric index converts biometric features into irreversible unique identifiers through a hash algorithm. The preset time window division rule divides the timestamp sequence in a 24-hour cycle and combines with the format compatibility rule to parse non-standard encoded data, ensuring that subsequent processing is compatible with multi-modal data.
[0025] The generation of a cross-platform time synchronization sequence needs to eliminate clock deviations and time zone differences of different devices. The offset normalization processing engine converts the original timestamp into Coordinated Universal Time (UTC) according to the geographical location database associated with the device code. For example, the timestamp "2023-10-05 14:30:21" in the Beijing time zone is converted into the UTC time of "2023-10-05 06:30:21". The clock error compensation mechanism calibrates the millisecond-level deviation between the local clock and UTC through the device heartbeat signal. For example, a certain device sends a status packet once a minute to correct clock drift. The preset time zone offset calculation rule and clock synchronization error threshold ensure the precise alignment of the time sequence, forming a cross-platform time synchronization sequence.
[0026] The parsing of the data format identifier depends on the format compatibility rule, which converts the header information of a DICOM file into JSON key-value pairs, or splits the segments of an HL7 message into structured fields. Non-standard encoded data is unified into a readable format through UTF-8 transcoding. For example, a custom text field "PatID: 023A" in a certain hospital is converted into a standard patient identifier format. The metadata format descriptor records the data format type, keyword field list, and encoding specification. For example, the metadata of a DICOM image contains parameters such as scan layer thickness and compression algorithm. The preset format conversion priority rule ensures the consistency of metadata parsing for different formats.
[0027] The construction of the location association between device codes and patient biometrics is achieved through a joint dimensionality map. The device location mapping rule queries the registration information database according to the device code to obtain the three-dimensional coordinates of the device deployment location. For example, the device "GE-MR-007" is mapped to the coordinates of "Radiology Department on the 3rd Floor - Area B". The patient biometric spatio-temporal binding rule associates the biometric index with the device coordinates and timestamp to form a spatio-temporal trajectory map. For example, the patient iris hash "H-8a3f…" is associated with the device coordinates within a specific time period to generate the edge relationship of "device node - patient node". The topological structure of the map stores the spatio-temporal relationships of nodes and edges for subsequent feature fusion.
[0028] Multimodal feature fusion aggregates the time-synchronized sequence, metadata format descriptor, and features of the joint dimensionality map into an initial hierarchical cluster. The time series features are transformed into a frequency domain amplitude spectrum through Fourier transform to capture the data access cycle pattern; the metadata semantic embedding uses a natural language processing model to extract the semantic vectors of text fields; the map topological features calculate the node centrality and edge weights, such as the data generation frequency of device nodes. The preset weighted fusion coefficient rule assigns weights (40%, 30%, 30%) to the time series, metadata, and map features, and generates an initial hierarchical cluster through the DBSCAN clustering algorithm. For example, highly sensitive image data is clustered into cluster 1, medium-sensitive text data into cluster 2, and low-sensitive log data into cluster 3.
[0029] Privacy compliance constraints drive the dynamic weight assignment of the hierarchical coefficient matrix. The data retention period rule sets that cluster 1 data is retained for 10 years, cluster 2 for 5 years, and cluster 3 for 1 year according to regulations such as GDPR and HIPAA. The access permission threshold rule requires double authentication for cluster 1 data, single-factor authentication for cluster 2, and only role permission verification for cluster 3. Real-time audit logs trigger dynamic weight adjustment. For example, a data breach event causes the access weight of cluster 1 to increase from 0.7 to 0.9. The hierarchical coefficient matrix stores the retention period, access weight, and encryption level of each cluster in matrix form. For example, the encryption level of cluster 1 is marked as AES-256.
[0030] The attribute information label is generated by incrementally iteratively adjusting the hierarchical coefficient matrix. The incremental data injection rule receives new data in real time and executes the feature separation to fusion process to generate new hierarchical clusters. The sliding time window rule recalculates the cluster weights based on the data in the most recent 30 days. For example, an increase in the proportion of highly accessed data triggers an adjustment of the encryption level. The convergence condition rule determines the stability of the matrix. When the cluster weight change rate is less than 1% in three consecutive iterations, the final attribute information label is output. The label describes the data classification result in JSON format. For example, the label of a certain image data includes encryption level, storage node, and retention period parameters.
[0031] In this embodiment, by performing structural feature separation and offset normalization processing on multi-dimensional attribute information, it is possible to disassemble attributes such as timestamps and format identifiers of heterogeneous medical data into a standardized structure, eliminate clock deviations and time zone differences between devices, and ensure the temporal consistency of cross-platform data operations. Based on the format compatibility parsing rules, unified transcoding and metadata extraction are performed on the data format identifiers to solve the problem of cross-platform parsing of multi-modal medical data and improve the compatibility and semantic coherence of data processing. Through the construction of the position association between device coding and patient biometric features, combined with the spatio-temporal trajectory mapping of the joint dimension map, the anonymous and accurate association of data sources and patient identities is realized, enhancing the data traceability ability while avoiding the risk of exposure of biometric plaintext; a multi-modal feature fusion mechanism is used to aggregate time series, metadata semantics, and map topological features, and initial hierarchical clusters are generated through dynamic weight allocation and clustering algorithms to overcome the problem of mismatch between security policies and data sensitivity caused by traditional single-dimensional classification. According to the privacy compliance constraints, the data retention period and access permission threshold are dynamically allocated to reduce the compliance risk caused by static classification of medical data.
[0032] In one embodiment, the medical data is differentially encrypted according to the attribute information label to obtain encrypted medical data, and the encrypted medical data is written into a multi-level storage node, and a multi-level backup is constructed to obtain a cross-platform backup topology, including: The attribute information label includes parameters such as data security level, access frequency, and storage period. The encryption policy construction engine generates a list of encryption algorithm priorities based on the security level parameters in the label. For example, highly sensitive patient gene data is mapped to the AES-256 algorithm, moderately sensitive diagnostic reports are mapped to the SM4 algorithm, and low-sensitive log data is mapped to the ChaCha20 algorithm. The key segmentation threshold matrix dynamically sets the block threshold according to the data size and access frequency. For example, the block threshold for image data exceeding 1GB is 128MB, and the block threshold for frequently accessed text data is 64MB. The preset security level-algorithm mapping rules define the encryption priorities corresponding to different security levels, and the block threshold rules balance the encryption efficiency and data throughput requirements.
[0033] The hybrid encryption process is based on the encryption algorithm priority list and achieves a balance among the security, efficiency, and dynamic key management of medical data through a hierarchical encryption mechanism. For highly sensitive medical data (such as patient gene sequences), when the encryption process is initiated, a 256-bit random key required for AES-256 symmetric encryption is first generated. This key is generated by a cryptographically secure pseudorandom number generator (CSPRNG) to ensure the unpredictability of the key. The generated symmetric key is directly used to encrypt the original data. The AES-GCM (Galois / Counter Mode) mode is adopted, combined with the initialization vector (IV) and additional authentication data (AAD), to achieve both data confidentiality and integrity protection simultaneously. The IV is a 12-byte random number, independently generated for each data block and transmitted in plaintext in the ciphertext header to ensure that the decryption end can reconstruct it correctly.
[0034] The encrypted symmetric key needs to be protected by an asymmetric encryption algorithm. The system calls the RSA-2048 public key encryption process and uses a pre-generated RSA public key (such as the root certificate public key of the hospital data center) to encrypt the AES key. The RSA encryption process follows the PKCS#1 v1.5 padding standard to ensure that the encrypted key ciphertext meets the interoperability requirements. The dynamic key group consists of the encrypted AES key ciphertext, the RSA public key identifier (such as the public key fingerprint "PK-007"), the key validity period (such as 90 days), and the key version number. The key version number is used to support the coexistence of multiple versions in the key rotation scenario. For example, when the key is updated, the old version key can still decrypt historical data until it is automatically destroyed after expiration.
[0035] The key lifecycle management rules require the dynamic key group to be rotated regularly. The key validity period parameter is dynamically set according to the storage period in the attribute information tag. For example, data that needs to be stored long-term (>5 years) uses a shorter key validity period (such as 30 days) to reduce the risk of key leakage through frequent rotation. When the key rotation is triggered, a new AES key is automatically generated and the data is re-encrypted. The old key ciphertext is retained until the data decryption is completed and then archived and deleted. The key storage is protected by a hardware security module (HSM) to ensure that the plaintext key only exists in the transient memory of the encryption process, avoiding the leakage risk caused by persistent storage.
[0036] For medium- and low-sensitivity data (such as diagnostic texts or operation logs), the encryption process is downgraded according to the algorithm priority list. For medium-sensitivity data, the national cryptography SM4 algorithm is used, with a key length of 128 bits and a CTR (Counter Mode) encryption mode. The counter is derived from the data block index through a hash function to avoid IV repetition. For low-sensitivity data, the ChaCha20 stream cipher is used for encryption, and the key and IV are derived from the master key through HKDF (HMAC-based Key Derivation Function) to ensure the key independence of each data block. The algorithm type, derivation parameters, and key version are recorded in the dynamic key group, supporting unified management in a multi-algorithm environment.
[0037] The encrypted medical data is bound to the dynamic key group through the node mixing rule. Each integrated encryption node package contains 1 dynamic key group and 10 associated encrypted data packets. The root value of the hash chain of the data packet (such as the root node of the SHA-256 hash tree) is embedded in the key group header. During node mixing, the local storage node saves the complete data packet and key group, while the off-site and cloud nodes only store fragmented or verification data, and verify the data integrity through the hash chain. For example, if a data packet is tampered with during transmission, its hash value will not match the root hash recorded in the key group, triggering the automatic isolation and recovery process.
[0038] The stream cipher block processing divides the medical data into continuous data blocks according to the key segmentation threshold matrix and applies the stream cipher algorithm to each block for encryption. For example, a 1.5GB CT image is divided into 12 data blocks of 128MB each, and each block is encrypted using an independent initialization vector (IV). The encrypted data packet contains the ciphertext block, IV, and block index, such as "encrypted data packet: {ciphertext: a3d9…, IV: 5f2b…, index: 3}". The preset stream cipher synchronization rule requires each data block to use a unique IV to prevent security vulnerabilities caused by repeated encryption modes.
[0039] The node mixing process binds the encrypted dynamic key group and the encrypted data packets according to a preset ratio to generate an integrated encryption node package. For example, every 10 data packets are bound to 1 dynamic key group, and the key group is associated with the first data packet through the hash chain to ensure data integrity. The preset node mixing rule defines the binding ratio of the key and the data packet and the hash verification mechanism. For example, the SHA-256 is used to calculate the data packet hash and embed it in the key group metadata.
[0040] The multi-level storage nodes are divided into local storage nodes (hospital internal servers), off-site cold storage nodes (offline tape libraries), and cloud storage nodes (public cloud object storage). The node recognition engine automatically classifies through node registration information (such as IP address, storage medium type). For example, local nodes are marked as "hot backup units", off-site nodes are marked as "cold storage", and cloud nodes are marked as "redundant storage". The preset node classification rules divide node types based on the access latency and durability parameters of the storage medium.
[0041] The comprehensive encryption node package is allocated to different nodes according to the storage strategy: local nodes store high-frequency access data packets and are marked as hot backup units, off-site nodes store low-frequency archived data packets, and cloud nodes store redundant copies. The hot backup marking rule is automatically triggered based on the access frequency threshold (such as more than 100 daily accesses). For example, a certain imaging data packet is marked as a hot backup unit due to high-frequency access and is preferentially stored in the local node.
[0042] The hot backup unit performs redundant sharding processing through the erasure code algorithm, splitting the original data packet into data shards and parity shards. For example, 10 data packets generate 12 shards (8 data shards + 4 parity shards), satisfying the shard quantity formula n = k + m, where k is the number of original data shards and m is the number of redundant parity shards. Off-site nodes and cloud nodes store the parity shards and verify the shard integrity through the hash verification rule. For example, each shard is appended with a SHA-256 hash value. The preset redundant sharding rule ensures that data can be restored through the remaining shards in case of a single point of failure. For example, at least k shards are required to reconstruct the original data.
[0043] The cross-platform backup topology describes the mapping relationship between shards and nodes through a graph structure. For example, data shards are stored in local nodes, and parity shards are stored in off-site and cloud nodes. The topology edges represent the redundant dependency paths between shards. The topology construction rule requires that each data shard has at least 3 redundant copies (1 copy each in local, off-site, and cloud), and the parity shards are stored dispersedly according to the geographical distribution. The preset redundant path rule combines the node geographical location and network latency parameters to optimize the data recovery efficiency and disaster tolerance ability.
[0044] In this embodiment, an encryption policy is dynamically constructed through attribute information tags, and an encryption algorithm is matched according to the data security level and access characteristics, so as to solve the problem of imbalance between security and efficiency in the traditional encryption mode. The hybrid encryption mechanism combines symmetric and asymmetric algorithms, and the dynamic key group is hierarchically protected and rotated regularly to reduce the risk of key leakage due to long-term use. The stream cipher processes data in blocks and dynamically adjusts the threshold according to the data scale. The unique initialization vector encrypts each block to prevent the exposure of ciphertext regularity and improve the anti-cracking ability. The node hybrid rule binds the key and the data packet, and the hash chain verification ensures the transmission integrity and quickly identifies tampering behavior. The multi-level storage nodes are classified and hot backup marked to optimize the rapid recovery of local high-frequency data and the disaster tolerance of remote cloud persistence; the erasure code is sliced and the redundant topology is constructed to disperse the stored data and the parity slices, and the remaining slices are used to efficiently recover in case of a single point of failure. The backup topology enforces geographical dispersion of multiple copies, and the dependency mapping reduces cross-platform latency, improving the efficiency of the recovery path and the fault tolerance of the system.
[0045] In one embodiment, the multi-level storage nodes are monitored for anomalies according to the cross-platform backup topology to obtain the abnormal node data access behavior. According to the abnormal node data access behavior, the replica node information of the multi-level storage nodes is extracted, and the cross-platform backup topology is updated to obtain the backup topology configuration, including: The cross-platform backup topology defines the data slice dependency path and redundancy distribution relationship among multi-level storage nodes (such as local nodes, remote nodes, and cloud nodes). The node monitoring log collection engine collects operation logs such as the access frequency, data throughput, and response time of each node in real time according to the node connection relationship in the topology, and generates a set of node behavior baseline parameters. For example, the baseline parameters of the local node include the average daily access times (such as 1200 times), the average response time (such as 50 ms), and the data traffic threshold (such as 1 GB / s). The preset log aggregation rule requires summarizing the log data according to a time window (such as 5 minutes) and dynamically updating the baseline parameters through a sliding window algorithm to adapt to the temporal changes of the node access pattern.
[0046] The identification of abnormal node data access behavior is achieved through threshold deviation analysis. The threshold deviation analysis engine compares the real-time access log with the set of baseline parameters to detect abnormal behaviors that exceed the dynamic threshold. For example, the access times of a remote node suddenly increase to 3 times the baseline value (such as from 100 times to 300 times) within 10 minutes, or the data traffic exceeds 200% of the baseline threshold (such as from 500 MB / s to 1.5 GB / s), triggering an anomaly mark. The preset dynamic threshold adjustment rule automatically calculates the tolerance range according to the historical volatility (such as the standard deviation). For example, the abnormal threshold of the access frequency is set to the baseline value ±2σ.
[0047] The generation of the abnormal node address list depends on clustering and grouping processing. The clustering algorithm for the request source address and operation type groups abnormal access behaviors by source IP address (such as 192.168.1.10) and operation type (such as data writing, shard deletion). For example, continuous shard deletion operations from the same IP are clustered into the "malicious tampering group". The preset clustering distance rules define the geographical location differences of the source addresses (such as IP segment differences) and the operation type similarities (such as the weights of write / delete operations), and generate the abnormal node address list and associated data tampering features (such as shard version number conflicts, metadata verification failures).
[0048] The replica reverse link node tracing starts from the abnormal node address list, traces back along the shard dependency path of the backup topology in reverse, and extracts the associated replica node codes and storage information. For example, for shard A stored on an abnormal node, its redundant replicas are distributed on nodes X, Y, and Z. The reverse tracing algorithm retrieves all node paths related to shard A in the topology and generates a set of abnormal replica node codes (such as codes N-023, N-045). The preset path tracing rules require at least three layers of redundant paths to be covered to ensure the completeness of the tracing results.
[0049] The hash chain backtracking verification process confirms data integrity by verifying the hash continuity of the replica node codes layer by layer. For example, the hash value H1 of shard A is calculated from its original data, and the hash value H2 of its replica shard A' should be the same as H1; if they are inconsistent, it is determined that shard A' has been tampered with. The preset hash continuity rules require the hash value of each shard to be associated with the hash of its parent node to form an unbreakable hash chain. The verification results generate data integrity verification information, recording the location and type of the tampered shard (such as hash mismatch of shard N-023).
[0050] The path isolation and update of the cross-platform backup topology are performed based on the data integrity verification information. The path isolation engine removes the abnormal replica nodes from the topology. For example, node N-023 is marked as "isolated state" and prohibited from participating in data recovery. The topology update engine recalculates the redundant shard distribution paths according to the storage information of the abnormal nodes. For example, the shards originally stored on N-023 are redistributed to nodes N-087 and N-102, and the topology dependency relationship is updated. The preset redundant path reconstruction rules require at least three replicas to be retained for each shard, and the new nodes need to meet the geographical dispersion requirements (such as different data centers). The updated backup topology configuration records the latest node connections and shard distribution relationships, such as "shard A: nodes X, Y, Z → nodes X, Y, W".
[0051] In this embodiment, through the collection of node monitoring logs of the cross-platform backup topology and the analysis of dynamic threshold deviation, it is possible to capture the abnormal access behavior of multi-level storage nodes in real time. Based on the historical access pattern adaptive threshold adjustment mechanism, the detection sensitivity to low-frequency covert attacks (such as slow data tampering) is improved; based on the clustering and grouping processing of the request source address and operation type, through the multi-dimensional feature analysis of the source IP geographical location difference and operation type weight, malicious node clusters in distributed collaborative attacks are accurately identified, solving the problem of high false alarm rate caused by traditional single-dimensional detection; the replica reverse link node tracking combined with the hierarchical backtracking of redundant paths locates the abnormal replica node encoding and its associated storage location, ensuring the complete traceability of the attack path and the accurate isolation of the tampered shards; the hash chain backtracking verification quickly identifies the tampered or damaged data replicas through the continuity verification of the shard hash values, and combines the irreversible characteristics of the hash chain to ensure the non-repudiation of the verification results; the path isolation and topology update mechanism dynamically adjusts the shard storage nodes and enforces geographical dispersion distribution according to the redundant path reconstruction rules, while blocking the spread of abnormal nodes, optimizing the disaster tolerance ability of the backup topology, and reducing the latency of cross-platform data recovery and the risk of single point of failure.
[0052] In one embodiment, the replica reverse link node tracking extraction of multi-level storage nodes is performed according to the abnormal node address list, and the abnormal replica node encoding and abnormal node storage information are obtained, including: The replica reverse link node tracking and abnormal storage information extraction process is implemented based on the path dependence relationship and consistency verification mechanism of the cross-platform backup topology. The abnormal node address list contains the node IP addresses marked as suspicious and their associated operation types (such as shard deletion, metadata tampering). The request feature filtering engine filters out high-threat request records from the abnormal node address list according to the operation type weight (such as the deletion operation risk weight is 0.8, and the read operation is 0.2) and the time window threshold (such as more than 5 consecutive operations within 10 minutes), and generates a reverse tracking parameter set. For example, an abnormal node initiates 8 shard deletion operations within 5 minutes, and the superposition of its operation type weights exceeds the preset threshold (such as 0.75), triggering the parameter filtering condition. The preset filtering rules require a comprehensive evaluation of the threat level by combining the operation type weight, frequency, and source IP geographical location (such as cross-border IP segments).
[0053] The construction of the reverse tracing path graph depends on the redundant shard distribution paths in the backup topology configuration. The link path dependency parsing engine reversely deduces the upstream and downstream node paths that may be affected by abnormal shards based on the node connection relationships recorded in the topology (such as shard A being stored at nodes X→Y→Z). For example, if an abnormal node stores shard B and its redundant copy distribution path is "node P→Q→R", the reverse path graph "R→Q→P" is generated through reverse parsing. The preset path depth rule requires covering at least three layers of redundant paths to ensure the integrity of the tracing scope. The path graph stores the dependency relationships between nodes in a graph structure, such as "node R—[shard B]—>node Q—[shard B]—>node P".
[0054] The replica node verification request is initiated based on the reverse tracing path graph, and a shard hash value query instruction is sent to each node in the graph. The consistency consensus check adopts the Byzantine fault tolerance protocol, which requires more than half of the nodes to return consistent hash values to determine the validity of the shard. For example, if the path graph contains 7 nodes and 4 of them return the same hash value H1, then H1 is determined to be the valid value, and the nodes that do not reach an agreement are marked as suspicious candidates. The candidate node list records all the responding nodes and their hash states, such as "node Q: hash H1 (valid), node R: hash H2 (invalid)". The preset majority consensus rule defines the minimum valid response ratio (such as 51%) to ensure the credibility of the verification result.
[0055] The hash chain tampering backtracking process starts from the candidate node list and verifies the continuity of the shard hash chain layer by layer. The hash chain consists of the shard data and its parent shard hash value. For example, the hash H1 of shard B is generated from its data, and the hash H2 of its sub-shard B1 should be calculated based on H1. If the hash H2’ returned by a certain node ≠ H2, it is determined that the shard B1 stored by this node has been tampered with. The interrupted node encoding extraction algorithm locates the break point of the hash chain and generates a set of abnormal replica node encodings, such as "node R: shard B1 hash mismatch". The preset hash continuity rule requires that the hash value of each shard strictly depends on the predecessor node, forming an irreversible chain verification structure.
[0056] The storage information of abnormal nodes is generated through the associated processing of key binding and storage paths. According to the key mapping table in the backup topology configuration, the abnormal replica node encoding is associated with the storage location (such as "off-site node - storage cabinet 3") and the bound encryption key (such as key ID: K-045). The storage path association rule requires that the newly allocated nodes meet geographical dispersion (such as different data centers) and minimum redundancy (such as at least 3 replicas for each shard). For example, the storage path of abnormal node R is updated to "cloud node S" and "off-site node T", and the key K-102 is rebound. The preset geographical dispersion rule forces the shard replicas to be stored across regions to avoid data unavailability caused by single-region failures.
[0057] In this embodiment, by filtering the request features of the abnormal node address list and dynamically weighting multi-dimensional parameters (operation type weight, time window, geographical location), high-threat access behaviors can be accurately screened, and the detection efficiency for covert attacks (such as low-frequency data tampering) can be improved; based on the reverse path tracing and link dependency relationship parsing of the backup topology configuration, a path map is reversely constructed along the redundant shard storage path to ensure the complete traceability of abnormal replica nodes and avoid the path omission problem caused by traditional one-way tracing; the replica node verification request combines the consistency consensus verification of the Byzantine fault tolerance protocol, and the validity of the shard is determined by the consistency of the responses of the majority of nodes, solving the problem that malicious nodes in a distributed environment forge responses to interfere with verification; the hash chain tampering traceback quickly locates the node where the hash chain breaks by layer-by-layer verifying the continuity of the shard hashes, and combines the irreversible characteristics of the chain structure to ensure the non-repudiation of tampering behaviors; the key binding and storage path association processing dynamically allocates shard replicas according to the geographical dispersion rule, forcing the new nodes after abnormal node replacement to meet the cross-regional redundant distribution, isolating the tampered nodes while optimizing the disaster tolerance ability of the backup topology and reducing the risk of data recovery delay caused by local failures.
[0058] In one embodiment, source nodes are extracted from multi-level storage nodes according to the abnormal node data access behaviors to obtain source node information, and data source tracing verification is performed on the source node information according to the backup topology configuration to obtain a trusted recovery chain, including: The abnormal node data access behaviors include high-frequency unconventional operations (such as abnormal shard deletion, metadata overwrite) and cross-node collaborative attack characteristics. The request log correlation extraction engine filters out abnormal node log data from the access logs of multi-level storage nodes according to the operation time window rule (such as more than 5 consecutive operations within 10 minutes) and the operation type association rule (such as the temporal correlation between deletion operations and metadata modification). For example, a certain node initiates 8 shard deletion operations within 5 minutes and is accompanied by abnormal modification of metadata fields, triggering the correlation extraction condition. The preset log filtering rules require multi-dimensional matching of the operation timestamp, source address, and operation type. For example, the operation weight of the superposition of a cross-country IP segment and an unconventional time (such as 3 am) is increased.
[0059] The source node metadata set is generated through path tracing and dependency analysis of multi-level storage nodes. The source node extraction engine reversely traces and associates nodes along the redundant path of the backup topology according to the shard identifier in the abnormal node log data. For example, the storage path of shard A is "local node → remote node → cloud node", and all node addresses, shard version numbers, and operation timestamps involved in the storage or transmission of shard A are extracted by reverse tracing. The preset path depth rule requires at least covering 3 layers of redundant paths to ensure the integrity of the traced metadata. The metadata set is stored in a structured format, such as "node X: shard A_v3, timestamp 2023-10-05 14:30:21".
[0060] The shard verification path graph is generated by resolving the shard path dependencies in the backup topology configuration. The path dependency resolution engine constructs a complete dependency graph of the shard storage paths based on the shard identifiers in the metadata set. For example, the dependency path of shard A includes its parent shard (such as shard A0) and child shards (such as shard A1). The graph nodes represent the shard storage locations, and the edges represent the shard version inheritance relationship. The preset version inheritance rule requires that the child shard hash value must be calculated based on the parent shard hash, forming an immutable chain structure. The shard verification path graph is stored in the form of a graph database, such as "shard A_v3 —[hash H1]—> shard A1_v1".
[0061] The node trust verification set is generated by synchronizing shard hash value requests and multi-node consistency checks. Based on the shard verification path graph, a shard hash value query request is sent to the associated nodes, and the Byzantine fault tolerance protocol is used to perform a majority consensus check on the response results. For example, if 5 out of 7 nodes return the same hash value H1, then H1 is determined to be a valid value, and the nodes that do not reach a consensus are marked as low-trust nodes. The preset majority consensus rule requires that the number of valid responses satisfies the formula Nvalid≥⌈2Ntotal+1⌉, where Ntotal is the total number of nodes. The node trust verification set records the node address, hash value, and trust status, such as "Node Y: hash H1 (trusted), Node Z: hash H2' (untrusted)".
[0062] The node signature recovery chain is generated by sorting and aggregating the shards in the trust verification set according to the timestamp. The shard sorting rule is arranged in strict ascending order of the operation timestamp. For example, shard A_v3 (timestamp T1) takes precedence over shard A_v4 (timestamp T2). The shard index hierarchy combination adopts a Merkle tree structure, and the hash digest of the sorted shard hash values is calculated layer by layer, finally generating the root hash value and the node signature (such as the digital signature of node X). The preset hierarchical aggregation rule requires that each shard index contains a version number and a parent hash reference, such as "shard A_v4: {parent hash: H1, signature: SigX}".
[0063] The trusted recovery chain is generated by verifying the integrity of the initial records of the blockchain through the non-interactive zero-knowledge proof protocol (zk-SNARKs). The integrity comparison engine performs a zero-knowledge proof verification on the root hash value in the recovery chain and the initial hash value stored on the blockchain to ensure that the recovery chain has not been tampered with and is consistent with the initial records. For example, if the recovery chain root hash H_root matches the H_block of the blockchain record, a verification passed flag is generated; if not, the differential shard is located and an alarm is triggered. The preset verification protocol rule requires that the proof generation process does not disclose the details of the shard content, and only verifies the hash consistency through mathematical proofs.
[0064] In this embodiment, by correlating and extracting request logs and dynamically screening abnormal node log data with multi-dimensional weight rules, it is possible to accurately identify high-frequency unconventional operations and cross-node collaborative attack behaviors, improving the detection sensitivity of covert attacks; based on the reverse tracing and metadata extraction of the redundant path depth rules, all nodes participating in sharded storage are completely traced along the storage link of the backup topology, solving the problem of path omission caused by traditional one-way tracing; constructing a shard verification path graph and constraining the version inheritance rules to ensure that the chain dependency relationship of shard hash values cannot be tampered with, enhancing the reliability of data version control; the majority consensus check of the node trusted verification set is implemented through the Byzantine fault tolerance protocol, effectively resisting the interference of false responses from malicious nodes in a distributed environment and ensuring the credibility of shard hash verification results. The hierarchical aggregation of the Merkle tree structure and the shard timestamp sorting mechanism generate a recovery chain containing node signatures, ensuring the irreversibility of the data recovery order and the audibility of the operation history; the integrity comparison of the non-interactive zero-knowledge proof protocol (zk-SNARKs) verifies the consistency between the root hash of the recovery chain and the initial record of the blockchain without revealing the details of the shard data, meeting the privacy compliance requirements of medical data (such as GDPR) and preventing tampering behaviors from bypassing the audit mechanism.
[0065] In one embodiment, recovery node selection is performed on multi-level storage nodes according to the trusted recovery chain, and multi-node collaborative decryption and reconstruction are carried out to obtain a recovery verification report, including: The trusted recovery chain contains the verified Merkle root hash, node signatures, and shard hierarchical relationships. The node signature weight evaluation engine generates a list of recovery candidate nodes based on the node credibility scores in the recovery chain (such as signature validity, historical response success rate). For example, node X is included in the candidate list because its historical response success rate reaches 99% and the signature is valid, and the weight score is higher than the threshold. The preset weight rules require that the node needs to meet the minimum credibility threshold (such as score ≥ 0.8) and have a dispersed geographical distribution (such as across at least 3 data centers).
[0066] The recovery node set is generated by calculating the redundant shard distribution path coverage set. The coverage set calculation engine filters out the smallest node set that can cover all data shards based on the candidate node list and the shard storage paths in the backup topology configuration. For example, if the redundant copies of shard A are distributed on nodes X, Y, and Z, and nodes X and Y already meet the shard coverage requirements, then node Z is excluded. The preset minimum coverage rule requires that each shard be covered by at least 2 nodes, and the network latency between nodes does not exceed 50ms.
[0067] Dynamic key matching is performed based on the encryption policy parameters in the attribute information tag. The attribute information tag records the security level of the data and the key binding relationship. For example, highly sensitive data is bound to the AES-256 key group. The key matching engine extracts the key group that matches the tag from the set of recovery nodes and sorts them according to the decryption priority. For example, the key group of shard A includes the main key shards K1, K2, and K3, which need to be jointly decrypted in the order K1 → K2 → K3. The preset key binding rule requires that the main key shards be stored separately in nodes in different geographical regions.
[0068] The distributed key sharding and collaborative decryption process is implemented through the secure multi-party computation protocol (MPC). Each node in the set of recovery nodes participates in the decryption calculation based on the key shards stored locally. For example, node X provides shard K1 and node Y provides shard K2, and the main key is jointly recovered through the MPC protocol and the sharded data is decrypted layer by layer. The preset decryption order rule requires that it be executed according to the priority sequence of the key shards. For example, decrypt the metadata shards first and then the data block shards. The set of decrypted original data shards contains shard content, index, and timestamp.
[0069] Among them, the calculation formula for participating in the decryption calculation is: ; : The minimum number of nodes that meet the conditions. argmin: Find the minimum parameter N that makes the set coverage condition hold. All shards: The union of the shard sets of the selected N nodes contains all shards. N: The number of nodes. : The shard set stored in the i-th node.
[0070] Redundancy check and sequential alignment perform deduplication and sorting on the set of original data shards. The redundancy check rule requires removing duplicate shards (such as shards with the same index but older timestamps) and sorting them strictly according to the shard index. For example, shard A1 (index 1, timestamp T1) takes precedence over shard A2 (index 2, timestamp T2). The reverse erasure code algorithm reconstructs the complete data block based on the shard index and the redundancy check shards. For example, any 8 out of 12 shards (8 data shards + 4 parity shards) can recover the original data, and the missing shards are filled by calculating the parity shards.
[0071] Data integrity recovery verifies the preliminary recovery result based on the root hash value of the trusted recovery chain. The integrity verification engine calculates the hash value of the reconstructed data block and compares it with the root hash in the recovery chain. If they are consistent, a recovery verification report containing the recovery time, the list of participating nodes, and the hash consistency mark is generated; if they are inconsistent, the differential shards are located and an alarm is triggered. The preset integrity rule requires that the reconstructed data hash must exactly match the root hash of the recovery chain.
[0072] In this embodiment, a node signature weight evaluation mechanism and a geographical dispersion rule are used to dynamically screen recovery candidate nodes, which can select a set of highly reliable nodes while ensuring data redundancy, avoid low-reputation nodes from participating in the recovery process, and improve the success rate and efficiency of data recovery; the calculation rule for the minimum number of nodes in the redundant shard coverage set combines the shard distribution path and network latency constraints to optimize the scale and response speed of the recovery node set, reduce cross-platform communication overhead while meeting full shard coverage; the dynamic key matching accurately binds key groups based on the security level of the attribute information tags, and through the cross-region distributed storage and hierarchical decryption order of the main key shards, prevents the risk of data exposure caused by the leakage of a single node's key; the collaborative decryption process driven by the secure multi-party computation protocol (MPC) ensures that the key shards jointly recover the main key without decrypting, taking into account both the security and efficiency of the decryption process; the reverse erasure code algorithm reorders the redundant check shards and data shards, which can efficiently fill in missing or damaged shards and reduce the probability of recovery failure caused by the failure of some nodes; the data integrity recovery mechanism forces the recombined data to be consistent with the initial record through the root hash comparison of the trusted recovery chain, and combines zero-knowledge proof verification to achieve the dual goals of privacy protection and tampering detection, meeting the compliance audit requirements of medical data.
[0073] Referring to Figure 2 As shown, the present invention also provides a secure backup and recovery system for medical data, which is applied to the secure backup and recovery method for medical data in any one of the above, and includes: An acquisition module, which is used to acquire multi-dimensional attribute information of medical data, and dynamically classify the multi-dimensional attribute information based on a preset classification rule to obtain attribute information tags; An analysis module, which is used to differentially encrypt medical data according to the attribute information tags to obtain encrypted medical data, write the encrypted medical data into multi-level storage nodes, and perform multi-level backup construction to obtain a cross-platform backup topology; An association module, which is used to perform anomaly monitoring on the multi-level storage nodes according to the cross-platform backup topology to obtain abnormal node data access behaviors, extract replica node information from the multi-level storage nodes based on the abnormal node data access behaviors, and update the topology of the cross-platform backup topology to obtain a backup topology configuration; A processing module, which is used to extract traceable node information from the multi-level storage nodes according to the abnormal node data access behaviors, perform data traceability verification on the traceable node information according to the backup topology configuration to obtain a trusted recovery chain; A control module, which is used to select recovery nodes for the multi-level storage nodes according to the trusted recovery chain, and perform multi-node collaborative decryption and reconstruction to obtain a recovery verification report.
[0074] A secure backup and recovery system for medical data provided by the present invention can, through a dynamic grading mechanism based on multi-dimensional attribute information, dynamically divide data security levels according to the privacy level, access frequency, and storage period of medical data, achieve an accurate match between the hierarchical storage strategy and the encryption intensity, and solve the problems of resource waste and security vulnerabilities caused by traditional static classification. By constructing differential encryption and cross-platform backup topologies, combined with the redundant sharding distribution of multi-level storage nodes, it can effectively resist the risk of data tampering or leakage of a single storage node and enhance the disaster tolerance of the backup system; by means of real-time monitoring of abnormal node data access behaviors and dynamic extraction of replica node information, it can quickly locate and isolate abnormal operation nodes in the distributed network and prevent the spread of malicious attacks in the backup link. Based on the traceability verification of the trusted recovery chain and multi-node collaborative decryption and reconstruction, the integrity and auditability of the data recovery process are ensured through tree-shaped hash aggregation and zero-knowledge proof technologies, avoiding the problem of untrusted data caused by single-point failure in the traditional centralized recovery mechanism. Through the dynamic update of the backup topology configuration and the calculation of the recovery node coverage set, the resource scheduling efficiency of multi-level storage nodes is optimized, reducing the redundant storage cost while ensuring the reliability of data recovery.
[0075] It should be noted that those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working processes of the above-described system and each module can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0076] The above are only the preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for secure backup and recovery of medical data, characterized in that, Including: Obtain the multi-dimensional attribute information of medical data, dynamically classify the multi-dimensional attribute information based on a preset classification rule, and obtain an attribute information label; Differentially encrypt the medical data according to the attribute information label to obtain encrypted medical data, write the encrypted medical data into multi-level storage nodes, and perform multi-level backup construction to obtain a cross-platform backup topology; Perform anomaly monitoring on the multi-level storage nodes according to the cross-platform backup topology to obtain abnormal node data access behaviors, extract replica node information from the multi-level storage nodes based on the abnormal node data access behaviors, and update the topology of the cross-platform backup topology to obtain a backup topology configuration; Extract traceable node information from the multi-level storage nodes based on the abnormal node data access behaviors, and perform data traceability verification on the traceable node information according to the backup topology configuration to obtain a trusted recovery chain; Select recovery nodes from the multi-level storage nodes according to the trusted recovery chain, and perform multi-node collaborative decryption and reconstruction to obtain a recovery verification report.
2. The security backup and recovery method of medical data according to claim 1, characterized in that The obtaining of the multi-dimensional attribute information of medical data, dynamically classifying the multi-dimensional attribute information based on a preset classification rule, and obtaining an attribute information label includes: Separate the structural features of the multi-dimensional attribute information to obtain a timestamp sequence, a data format identifier, a device code, and a patient biometric index; Perform offset normalization processing on the timestamp sequence according to the classification rule to obtain a cross-platform time synchronization sequence; Perform format compatibility parsing on the data format identifier based on the classification rule to obtain a metadata format descriptor; Construct a location association between the device code and the patient biometric to obtain a joint dimension map; Perform multi-modal feature fusion on the time synchronization sequence, the metadata format descriptor, and the joint dimension map to obtain an initial classification cluster; Allocate data retention periods and access permission threshold weights to the initial classification cluster according to preset privacy compliance constraint conditions to obtain a classification coefficient matrix; Perform incremental iterative adjustment on the classification coefficient matrix to obtain the attribute information label.
3. The method for secure backup and recovery of medical data according to claim 1, wherein The differentially encrypting the medical data according to the attribute information label to obtain encrypted medical data, writing the encrypted medical data into multi-level storage nodes, and performing multi-level backup construction to obtain a cross-platform backup topology includes: Construct an encryption policy based on the attribute information label and the medical data to obtain an encryption algorithm priority list and a key segmentation threshold matrix; Perform hybrid encryption on the medical data based on the encryption algorithm priority list to obtain an encrypted dynamic key group; Perform stream cipher block processing on the medical data according to the key segmentation threshold matrix to obtain encrypted data packets; Mix the encrypted dynamic key group and the encrypted data packets at the node to obtain a comprehensive encrypted node packet; Perform node identification on the multi-level storage nodes to obtain local storage nodes, off-site cold storage nodes, and cloud storage nodes; Store the comprehensive encryption node packages into the local storage node, the off-site cold storage node, and the cloud storage node respectively, and mark the local storage node as a hot backup unit; Perform erasure code redundant sharding processing on the hot backup unit to obtain a shard set, and perform verification and identification on the off-site cold storage node and the cloud storage node to obtain a verification shard set; Construct a backup topology for the shard set and the verification shard set to obtain the cross-platform backup topology.
4. The method for secure backup and recovery of medical data according to claim 1, characterized in that, Perform anomaly monitoring on the multi-level storage nodes according to the cross-platform backup topology to obtain abnormal node data access behaviors, extract replica node information from the multi-level storage nodes based on the abnormal node data access behaviors, and update the topology of the cross-platform backup topology to obtain a backup topology configuration, including: Collect node monitoring logs from the multi-level storage nodes according to the cross-platform backup topology to obtain a set of node behavior baseline parameters; Perform an analysis of the deviation of the access behavior threshold on the set of node behavior baseline parameters to obtain the abnormal node data access behaviors; Perform clustering and grouping processing on the request source address and operation type based on the abnormal node data access behaviors to obtain a list of abnormal node addresses; Perform replica reverse link node tracking extraction on the multi-level storage nodes according to the list of abnormal node addresses to obtain abnormal replica node codes and abnormal node storage information; Perform hash chain backtracking verification processing on the replica node codes to obtain data integrity verification information; Perform path isolation processing on the cross-platform backup topology based on the data integrity verification information, and perform topology update in combination with the abnormal node storage information to obtain the backup topology configuration.
5. The method for secure backup and recovery of medical data according to claim 4, characterized in that Perform replica reverse link node tracking extraction on the multi-level storage nodes according to the list of abnormal node addresses to obtain abnormal replica node codes and abnormal node storage information, including: Filter the request features of the list of abnormal node addresses to obtain a set of reverse tracking parameters; Analyze the link path dependency relationship of the set of reverse tracking parameters according to the backup topology configuration to obtain a reverse tracking path map; Perform replica node verification requests on the multi-level storage nodes based on the reverse tracking path map, and perform consistency consensus verification to obtain a list of candidate nodes; Perform hash chain tampering backtracking on the list of candidate nodes, and extract the interrupted node codes to obtain a set of abnormal replica node codes; Perform key binding and storage path association processing on the set of abnormal replica node codes according to the backup topology configuration to obtain the abnormal node storage information.
6. The method for secure backup and recovery of medical data according to claim 1, characterized in that, Extract traceable nodes from the multi-level storage nodes based on the abnormal node data access behaviors to obtain traceable node information, and perform data traceability verification on the traceable node information according to the backup topology configuration to obtain a trusted recovery chain, including: Perform request log correlation extraction on the abnormal node data access behaviors to obtain abnormal node log data; Extract traceable nodes from the multi-level storage nodes based on the abnormal node log data to obtain a set of traceable node metadata; Perform sharding path dependency resolution on the traceability node metadata set according to the backup topology configuration to obtain a sharding verification path graph; Based on the sharding verification path graph, send sharding hash value synchronization requests to the multi-level storage nodes and perform multi-node consistency verification to obtain a node trusted verification set; Sort the node trusted verification set by sharding and perform sharding index level combination to obtain a node signature recovery chain; Based on a preset non-interactive zero-knowledge proof protocol, perform blockchain initial record integrity comparison on the node signature recovery chain to obtain the trusted recovery chain.
7. The method for secure backup and recovery of medical data according to claim 1, characterized in that, Select recovery nodes for the multi-level storage nodes according to the trusted recovery chain and perform multi-node collaborative decryption and reconstruction to obtain a recovery verification report, including: Evaluate the node signature weights of the trusted recovery chain to obtain a list of recovery candidate nodes; Calculate a redundant sharding distribution path coverage set for the multi-level storage nodes according to the list of recovery candidate nodes to obtain a set of recovery nodes; Perform dynamic key matching on the set of recovery nodes based on the attribute information tag to obtain a node key group and a decryption order sequence; Perform initial collaborative decryption processing on the set of recovery nodes through preset distributed key sharding and the node key group, and decrypt the sharded data layer by layer to obtain a set of original data shards; Perform redundancy check and sequential alignment on the set of original data shards according to the decryption order sequence, and perform shard recombination according to a preset reverse erasure code algorithm to obtain a preliminary recovery result; Perform data integrity recovery on the preliminary recovery result based on the trusted recovery chain to obtain the recovery verification report.
8. A secure backup and recovery system for medical data, characterized in that, Applied to the method for secure backup and recovery of medical data described in any one of claims 1-7 above, including: A collection module for obtaining multi-dimensional attribute information of medical data, dynamically grading the multi-dimensional attribute information based on a preset classification rule to obtain an attribute information tag; An analysis module for differentially encrypting the medical data according to the attribute information tag to obtain encrypted medical data, writing the encrypted medical data into a multi-level storage node, and performing multi-level backup construction to obtain a cross-platform backup topology; An association module for monitoring anomalies of the multi-level storage nodes according to the cross-platform backup topology to obtain abnormal node data access behaviors, extracting replica node information from the multi-level storage nodes according to the abnormal node data access behaviors, and updating the topology of the cross-platform backup topology to obtain a backup topology configuration; A processing module for extracting traceability node information from the multi-level storage nodes according to the abnormal node data access behaviors, performing data traceability verification on the traceability node information according to the backup topology configuration to obtain a trusted recovery chain; A control module for selecting recovery nodes for the multi-level storage nodes according to the trusted recovery chain and performing multi-node collaborative decryption and reconstruction to obtain a recovery verification report.
Citation Information
Cited By
Logistics traceability system and method based on block chain technology
CN120525433A
A logistics traceability system and method based on blockchain technology
CN120525433B
Mobile phone file backup method and system based on intelligent hardware
CN120540905A
Traceable file watermark generation method and device, computer equipment and medium
CN120541040A
Medical data traceability and privacy protection auditing method and system based on block chain
CN120541883A