An AI-based intelligent private data sharding and reorganization method and system
By constructing a medical knowledge graph and using AI-driven intelligent fragmentation and reorganization methods, the problems of information fragmentation and privacy protection in medical data sharing are solved, achieving data security, availability and computability, and supporting cross-institutional joint analysis.
Patent Information
- Application Number
- CN202511127121.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-13
AI Technical Summary
Existing data sharding methods struggle to maintain the integrity and privacy of critical information during medical data sharing, leading to fragmented or over-encrypted information during reassembly, which fails to support joint analysis across institutions.
By constructing a medical knowledge graph, using AI for intelligent fragmentation and reorganization, employing an association-preserving fragmentation strategy to process high-clinical-value data, and differential privacy fragmentation to process highly sensitive data, and using distributed storage and dynamic rotation mechanisms for differentiated reorganization, the security and availability of data are ensured.
It achieves the protection of data privacy while maintaining the integrity of critical medical information, supporting cross-institutional collaborative analysis, reducing the risk of privacy leaks, and ensuring the semantic integrity and usability of reconstructed data.
Smart Images

Figure CN120632943B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data protection technology, specifically to an AI-based intelligent private data fragmentation and reassembly method and system. Background Technology
[0002] Data sharding, a traditional data protection technique, is widely used in distributed systems. Its basic idea is to divide the original data into multiple fragments and store them on different nodes to reduce the risk of single points of failure or data leakage. However, traditional data sharding methods are usually based on fixed rules (such as hash sharding or range sharding), lacking the ability to intelligently identify and differentiate data content, making it difficult to adapt to complex and ever-changing data types and application scenarios. Furthermore, static sharding strategies are easily exploited by attackers through reverse engineering and other methods to reconstruct the original data, thus weakening its security effectiveness.
[0003] In recent years, the rapid development of artificial intelligence (AI) technology, especially deep learning and reinforcement learning, has brought new ideas to the field of data security. AI possesses powerful pattern recognition and decision-making capabilities, enabling it to dynamically formulate optimal data segmentation strategies based on the content characteristics, sensitivity, and application scenarios of the data. Simultaneously, AI can play a crucial role in the data reconstruction stage, achieving efficient and accurate restoration of original data by learning the semantic relationships and structural connections between data points.
[0004] Chinese invention patent CN109033873A discloses a data anonymization method to prevent privacy leaks. Specifically, it includes the following process: removing explicit associations based on the same index fields between different data tables in a database; defining cryptographic functions for the index fields between data tables to process the association ID; calculating the association ID value based on the cryptographic function, writing the association ID value, and then accessing the data. This invention primarily employs a cryptographic approach, using algorithms to process the association fields between data tables, removing strong couplings between different tables and data in the database and user information. This ensures that even with superuser privileges on the user's database, the relationships between different data and information cannot be known, and the obtained data cannot be verified against the user, thus achieving data privacy protection. This method can effectively prevent privacy leaks caused by direct database access due to platform attacks or insider attacks.
[0005] However, in the process of sharing medical data, patients' electronic medical records (including text descriptions, examination reports, etc.) need to be fragmented and stored on different nodes to protect patient privacy. However, this fragmentation process may result in critical information (such as "diabetes + insulin treatment") being broken down, making it difficult to reconstruct the complete medical history. While direct encryption can protect data privacy, it sacrifices computability, hindering cross-institutional collaborative analysis. Furthermore, semantically preserving fragmentation may expose sensitive patient data, allowing for the inference of patient identity or medical history. Summary of the Invention
[0006] The purpose of this invention is to provide an AI-based intelligent private data fragmentation and reassembly method and system to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: an AI-based intelligent private data fragmentation and reassembly method, comprising:
[0008] S1: Constructing a medical knowledge graph: Using the identified entities as knowledge graph nodes, construct a medical knowledge graph and label them with clinical importance levels and sensitive data types;
[0009] S2: Intelligent Sharding: Determines the sharding method based on the partitioning results of the labeled data, including:
[0010] S2.1: Constructing the assessment matrix: Construct a two-dimensional assessment matrix table based on the privacy risk dimension and the clinical value dimension;
[0011] S2.2: Segmentation Decision: Data is segmented using a two-dimensional evaluation matrix table to determine the hash identifiers and metadata of high clinical value data, and the privacy-protected data for external release of data with privacy risks;
[0012] S2.3: Segmented storage: Based on the number of relations retained in each segment and the total number of relations in the original graph, determine the segmentation integrity index, and determine the final segmentation result based on the segmentation integrity index and the preset index threshold;
[0013] S3: Data Reorganization: Based on the sharding results, the shards are stored in heterogeneous nodes, and the data is reorganized according to the access permission level, including:
[0014] S3.1: Sharded storage: Sharded data with verifiable semantic tags are stored in heterogeneous nodes, and the storage location is updated through a dynamic rotation mechanism;
[0015] S3.2: Perceptual Reassembly: Based on access permission levels, reassemble fragmented data, obtain semantic integrity index and semantic deviation, and determine the final reassembled data, including:
[0016] S3.2.1: Access control: Based on the number of valid clinical relations in the reconstructed data and the clinical relation tree of the original knowledge graph, obtain the semantic integrity index, and determine the final reconstruction method based on the semantic integrity index and the preset semantic threshold;
[0017] S3.2.2: Secure Reassembly: Based on the node-specific key pair and the final reassembly method, the split fragments are aggregated, and the semantic deviation is obtained based on the semantic integrity index of the reassembly result and the semantic integrity index of the original data to determine the final reassembly result.
[0018] Furthermore, the clinical importance level and sensitive data types are labeled, including:
[0019] S1.1: Data processing: Obtain patient information, key paragraphs in medical records and DICOM files through the HIS system, EMR text and medical image metadata, and perform data cleaning and normalization on the patient information, key paragraphs in medical records and DICOM files.
[0020] S1.2: Semantic parsing: Using the BioClinicalBERT model, entity units in the diagnostic record text are identified. At the same time, the entity units are used as input to the graph attention network, and the output obtains the weight between different entity units and constructs relation triples.
[0021] S1.3: Privacy Labeling: Based on the life support index, treatment complexity, and update frequency, clinical importance and sensitivity are graded and associated with the aforementioned relational triples to construct a medical knowledge graph.
[0022] Furthermore, an entity co-occurrence matrix is constructed using the entity units, and initial weights for entity unit combinations are set using the entity co-occurrence matrix. Simultaneously, the weights of the graph attention network are adjusted based on these initial weights, including:
[0023] S1.2.1: Determine initial weights: Divide the diagnostic record text into segments according to the set window size, obtain the frequency of occurrence of different entity unit combinations, construct an entity co-occurrence matrix, and determine the initial weights for each entity unit combination, specifically:
[0024]
[0025] in: Let r be the initial weight of the relationship between entity r and entity o. Let r be the number of times entity r and o co-occur in the case. Total number of cases;
[0026] S1.2.2: Weight Adjustment: Based on the initial weights, determine the adjusted weights, specifically as follows:
[0027]
[0028] in: The adjusted weights between entity r and entity o. The weights represent the original relationship between entity r and entity o obtained through the graph attention network. Let r be the initial weight of the relationship between entity r and entity o. It is a normalized exponential function.
[0029] Furthermore, a medical knowledge graph is constructed, including:
[0030] S1.3.1: Clinical Importance Grading: Based on the set life support index score, treatment complexity score, and update frequency score, a grading score is obtained. Simultaneously, the grading score is compared with a preset grading threshold range, and based on the comparison result, the clinical importance level is determined, specifically as follows:
[0031] When the grading score is less than the lower limit of the preset grading threshold range, the clinical importance level is L3; when the grading score is within the preset grading threshold range, the clinical importance level is L2; when the grading score is greater than the upper limit of the preset grading threshold range, the clinical importance level is L1.
[0032] S1.3.2: Sensitivity Marker: Set the sensitivity level based on the text content of the diagnostic record;
[0033] S1.3.3: Knowledge Graph Generation: Based on the entity recognition results, clinical importance level, and sensitivity level, node attributes in the knowledge graph are set. At the same time, based on different entity units and weights, relation attributes in the knowledge graph are set. Through the node attributes and relation attributes, Neo4j graph data is constructed.
[0034] Furthermore, based on PubMed's literature citation data and the clinical research registry platform, keyword combinations in the treatment plan are extracted, and a research value score is determined according to the number of extracted literatures. At the same time, a recommendation level mapping table is established according to the recommendation level of clinical guidelines to determine the treatment effectiveness score, and a clinical value dimension is set based on the treatment effectiveness score and the research value score.
[0035] Furthermore, based on the HIPAA classification mapping table, a basic risk value for rare diseases is set, and the basic risk value is adjusted according to the prevalence of the disease. At the same time, a privacy risk dimension is set based on the adjusted risk value.
[0036] Furthermore, updating the storage location includes:
[0037] S3.1.1: Distributed storage: Each data shard is split into multiple fragments, and these fragments are stored in nodes at different geographical locations.
[0038] S3.1.2: Dynamic Protection: Through the RAFT protocol, the node data that is stored repeatedly is identified, and the validity of the rotation is obtained based on the total number of nodes. At the same time, the hash value of the shard, the timestamp, and the node signature are used as inputs to the zero-knowledge proof model, and the proof result is output.
[0039] Furthermore, the fragmentation integrity index is compared with a preset index threshold, and the final fragmentation result is determined based on the comparison result, specifically as follows:
[0040] When the fragment integrity index is less than the preset index threshold, steps S2.1-S2.3 are repeated to reassemble the fragments until the fragment integrity index is not less than the preset index threshold; otherwise, the corresponding fragmentation result is the final fragmentation result.
[0041] The semantic integrity index is compared with a preset semantic threshold, and the final recombination method is determined based on the comparison result, specifically as follows:
[0042] When the semantic integrity index is less than a preset semantic threshold, missing data retrieval is performed; otherwise, recombinant data is obtained through the clinical decision subgraph.
[0043] The semantic deviation is compared with a preset deviation threshold range, and the final reorganization result is determined based on the comparison result, specifically as follows:
[0044] When the semantic deviation is less than the lower limit of the preset deviation threshold range, the corresponding reconstructed semantics is the final reconstructed result; when the semantic deviation is within the preset deviation threshold range, a warning signal is triggered, and the corresponding reconstructed semantics is marked with low confidence; when the semantic deviation is greater than the upper limit of the preset deviation threshold range, step S3.1 is returned, fragmentation and storage are performed again, and steps S3.2.1 and S3.2.2 are repeated.
[0045] An AI-based intelligent private data fragmentation and reconstruction system uses any one of the AI-based intelligent private data fragmentation and reconstruction methods described above.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] Firstly, this invention, through the construction of a medical knowledge graph and semantic analysis, can accurately identify high clinical value data and highly sensitive data, and process them respectively using association-preserving sharding and differential privacy sharding strategies. This not only ensures the integrity of key medical information, but also effectively reduces the risk of privacy leakage.
[0048] Secondly, this invention ensures the security and traceability of data sharding through distributed storage and dynamic rotation mechanisms. At the same time, it performs differentiated reorganization based on permission levels, so that users with different permissions can obtain the corresponding adapted data views, thereby not only improving data security but also ensuring data availability.
[0049] Thirdly, this invention achieves intelligent decision-making for data sharding through a two-dimensional assessment matrix of clinical value and privacy risk, avoiding the problems of fragmentation or over-encryption of key medical information caused by traditional sharding methods, and improving the feasibility of data sharing and joint analysis.
[0050] Fourthly, this invention utilizes the correlation of knowledge graphs and detects errors during the reorganization stage using semantic integrity index and deviation, thereby ensuring the usability of the reorganized data in clinical decision-making and research, and reducing information distortion. Attached Figure Description
[0051] Figure 1 This is a flowchart illustrating the intelligent private data fragmentation and reassembly method of the present invention;
[0052] Figure 2 This is the medical knowledge graph in this invention;
[0053] Figure 3 This is a clinical value-privacy risk assessment matrix diagram in this invention;
[0054] Figure 4 This is a security verification diagram used in this invention. Detailed Implementation
[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0056] In the process of sharing medical data, patients' electronic medical records (including text descriptions, examination reports, etc.) need to be fragmented and stored on different nodes to protect patient privacy. However, fragmenting patients' electronic medical records may result in the fragmentation of critical information (such as "diabetes + insulin treatment"), making it difficult to reconstruct the complete medical history during reconstruction. While direct encryption can protect data privacy, it sacrifices computability and cannot support cross-institutional collaborative analysis. Furthermore, semantically preserved fragmentation may expose sensitive patient data, allowing for the inference of patient identity or medical history. This application addresses this issue by performing entity recognition and semantic parsing on electronic medical records, medical images, and other data to construct a medical knowledge graph and determine corresponding indicators to label clinical importance and sensitivity levels. Based on a two-dimensional assessment matrix of clinical value and privacy risk, association-preserving fragmentation is used for high-clinical-value data to maintain key medical connections, while differential privacy fragmentation is implemented for highly sensitive data to protect privacy. A dynamic rotation mechanism stores fragments on heterogeneous nodes and performs differentiated reconstruction based on access permissions. This ensures both data security and the semantic integrity of the reconstructed data, achieving an intelligent balance between privacy protection and clinical value of medical data.
[0057] Example 1
[0058] refer to Figures 1-4 This embodiment provides an AI-based intelligent private data fragmentation and reassembly method, which specifically includes the following steps:
[0059] Step S1: Construct a medical knowledge graph. This involves using the BioClinical BERT model to identify electronic medical record data, medical imaging reports, and laboratory test results. The identified entities are then used as nodes in the knowledge graph to construct the medical knowledge graph. Simultaneously, the medical knowledge graph is used to label data based on clinical importance and sensitive data types. Details are as follows:
[0060] Step S1.1: Data Processing. This involves retrieving patient information from the HIS system, extracting key paragraphs from the medical record in the EMR text, and retrieving the DICOM file from the medical image metadata. Specifically, in the HIS system data, core fields, including the patient's unique identifier (e.g., MED_ID:2023P123456), gender, and date of birth, are retrieved through the inpatient master index. In the EMR text data, patient diagnostic text information is retrieved through key paragraphs. In the medical image metadata, the corresponding examination type, equipment model, and scan parameters are extracted from the CT image header file.
[0061] Furthermore, the patient diagnosis text information obtained from the EMR text data is segmented, stop words are removed, and medical terminology is standardized to use unified medical standard terminology.
[0062] Furthermore, the core fields and medical image metadata obtained from the HIS system are matched with the patient ID to ensure that each test result can be associated with the correct patient ID. In this embodiment, the patient ID in the test report is matched with the master index in the HIS system, and the hospitalization number on the medical order is matched with the nursing record.
[0063] In this embodiment, based on the patient ID and corresponding diagnostic record (including main diagnostic content, date of diagnosis, and diagnosing physician) obtained from the HIS system data, the diagnostic content is converted using ICO-10 codes, and the date of diagnosis is converted to ISO 8601 format. Simultaneously, based on the drug record (including internal drug code, drug name, and daily dosage) obtained from the EMR text, both the internal drug code and drug name are converted to standard codes using a drug mapping table, and the units corresponding to the daily dosage are converted to the unit size corresponding to the drug specification.
[0064] In this embodiment, the admission time obtained from the HIS system data is used as the baseline time point. At the same time, the collection time of the laboratory is obtained through the LIS laboratory information system, and the collection time of the image acquisition metadata is obtained through the PACS medical image archiving system. The collection time of the laboratory and the collection time of the image acquisition metadata are set based on the baseline time point. That is, the relative time corresponding to the collection time of the laboratory and the collection time of the image acquisition metadata is determined by the baseline time point.
[0065] Furthermore, based on the completeness of the patient's medical records, specifically the completeness of the fields corresponding to patient ID, diagnosis code, allergy history, and contact information, a corresponding completeness index is determined, as follows:
[0066]
[0067] in: As an integrity index, The weight of the i-th field, For the complete state of the i-th field, To evaluate the total number of fields, To evaluate the field index.
[0068] Specifically, in this embodiment, based on the field completion status, the completeness status corresponding to the completed field is set to 1, and the completeness status corresponding to the missing field is set to 0. Simultaneously, this embodiment sets the weight of key fields (such as patient ID and diagnosis code) to 1, the weight of important fields (such as allergy history) to 0.5, and the weight of general fields (such as contact information) to 0.3.
[0069] Step S1.2: Semantic parsing. This involves using the BioClinical BERT model to identify the content of the diagnostic record text, converting it into multiple sub-word units, and then identifying the corresponding entity units from these sub-word units, such as disease, drug, and surgery sub-word units. Based on the identified entity units, the start and end character positions of each entity unit in the diagnostic record text are determined.
[0070] Furthermore, the identified entity units are used as input to the graph attention network, and the output obtains the weights between different entity units. In other words, based on different entity units and their corresponding weights and relationships, corresponding relational triples are obtained, such as (acute anterior wall myocardial infarction, emergency treatment, aspirin 300mg, weight 0.91) and (acute anterior wall myocardial infarction, typical signs, ST segment elevation, weight 0.87).
[0071] Step S1.3: Privacy Labeling. This involves classifying the clinical importance and sensitivity of patients based on their life support index, treatment complexity, and update frequency in their medical records. The classified clinical importance and sensitivity are then associated with the relational triples obtained in Step S1.2 to construct a corresponding medical knowledge graph. Details are as follows:
[0072] Step S1.3.1: Clinical Importance Grading. This involves setting scores for the Life Support Index, Treatment Complexity, and Update Frequency based on the text content of diagnostic and medication records. Specifically, when setting the Life Support Index score, a score of 0 is set when the data is unrelated to life maintenance. A score of 1 is set when the data indirectly affects vital signs but is not an immediate risk. A score of 2 is set when active intervention is required to maintain life but is not continuously dependent on the data. A score of 3 is set when life is entirely dependent on medical devices / medications.
[0073] Furthermore, in setting the treatment complexity score, the corresponding treatment complexity score is 1 for a single oral medication; 2 for combination therapy / requiring simple monitoring; 3 for invasive procedures / intensive treatment; 4 for multidisciplinary collaborative treatment; and 5 for multi-organ support therapy.
[0074] To elaborate further, in setting the update frequency score, the score is as follows: 1 for real-time updates (i.e., updates every minute); 0.8 for high-frequency updates (every 1-15 minutes); 0.5 for regular updates (every 1-24 hours); 0.3 for low-frequency updates (every 1-30 days); and 0.1 for annual updates (at least one year).
[0075] In this embodiment, based on the obtained life support index score, treatment complexity score, and update frequency score, the corresponding grading score is determined, specifically as follows:
[0076]
[0077] in: To assign scores based on grading, Rate the life support index. To score the complexity of treatment, Rate the frequency of updates.
[0078] Furthermore, the obtained grading score is compared with a preset grading threshold range (which is set according to specific needs, and therefore not specifically described in this embodiment, such as 1.5-2.5), and the corresponding clinical importance level is determined based on the comparison result, specifically as follows:
[0079] When the obtained grading score is less than the lower limit of the preset grading threshold range (i.e., 1.5), the corresponding clinical importance level is L3. When the obtained grading score is within the preset grading threshold range (i.e., 1.5-2.5), the corresponding clinical importance level is L2. When the obtained grading score is greater than the upper limit of the preset grading threshold range (i.e., 2.5), the corresponding clinical importance level is L1.
[0080] Step S1.3.2: Sensitivity Labeling. Based on the patient's genetic data, original records of major infectious diseases and psychological treatment, the sensitivity level corresponding to their characteristics is set to high risk. Simultaneously, based on rare disease diagnoses, history of mental illness, and occupational exposure records, the sensitivity level corresponding to their characteristics is set to medium risk. Based on demographic data, aggregated statistics, and anonymized specimen data, the sensitivity level corresponding to their characteristics is set to low risk.
[0081] Step S1.3.3: Knowledge Graph Generation. Based on the entity recognition results in Step S1.2, the clinical importance level determined in Step S1.3.1, and the sensitivity level determined in Step S1.3.2, the node attributes in the knowledge graph are set. Simultaneously, based on the different entity units and their corresponding weights obtained in Step S1.2, the relational attributes in the knowledge graph are set, thereby constructing the Neo4j graph data.
[0082] Step S2: Intelligent Sharding. This involves using a preservative sharding method to shard high-clinical-value data based on a pre-defined clinical value-privacy-risk two-dimensional assessment matrix, and differential privacy sharding for highly sensitive data, while simultaneously setting corresponding sharding levels. Details are as follows:
[0083] Step S2.1: Construct the assessment matrix. This involves creating a recommendation level mapping table based on the recommendation levels of clinical guidelines (such as NCCN and AHA), as shown in Table 1 below.
[0084] Table 1: Recommendation Level Mapping Table
[0085] Guide Recommendation Level Score Example treatment Class I 1.0 PCI for acute myocardial infarction Class IIa 0.7 New oral anticoagulants for atrial fibrillation patients Class IIb 0.3 Experimental immunotherapy for advanced tumors
[0086] Furthermore, based on PubMed citation data and the clinical research registry platform, keyword combinations in the treatment plan are extracted, and the corresponding research value score is determined according to the number of extracted documents, specifically:
[0087]
[0088] in: To score the research value, This represents the number of times the document has been cited.
[0089] In this embodiment, the corresponding treatment effectiveness score is determined based on the established recommendation level mapping table, and the corresponding clinical value dimension is set based on the determined treatment effectiveness score and research value score.
[0090] Furthermore, based on the established HIPAA classification mapping table, the basic risk values for rare diseases are set, as shown in Table 2 below.
[0091] Table 2: HIPAA Classification Mapping Table
[0092] Sensitivity Level Data Type Risk Value High risk Gene sequence, HIV status 1.0 Medium and high risk Rare disease + occupational combination 0.8 Medium risk Surgical records + medication history 0.6 Low risk Age + Gender 0.3
[0093] Furthermore, based on the prevalence of the disease, the baseline risk values in the HIPAA classification mapping table are adjusted to obtain the adjusted risk values, specifically:
[0094]
[0095] in: This is the adjusted risk value. Basic risk value, For the prevalence of disease.
[0096] In other words, based on the adjusted risk value, a corresponding privacy risk dimension is set. Specifically, based on the set privacy risk dimension and clinical value dimension, a corresponding two-dimensional assessment matrix table is constructed, as shown in Table 3 below.
[0097] Table 3: Two-Dimensional Evaluation Matrix
[0098] Treatment effectiveness score\Privacy risk value 0.3 (Low risk) 0.6 (Medium risk) 1.0 (High Risk) 1.0 (Class I) Regular fragmentation Relationship-based sharding Differential privacy fragmentation 0.7 (Class IIa) Regular fragmentation Dynamic decision-making* Differential privacy fragmentation 0.3 (Class IIb) Low priority storage Desensitization treatment Strict encryption
[0099] Step S2.2: Segmentation Decision. Based on the two-dimensional evaluation matrix table constructed in Step S2.1, high clinical value data and high privacy risk data are divided, and the corresponding segmentation type is set according to the data types of the segmentation.
[0100] Furthermore, the maximum connected subgraph algorithm is used to extract disease types, treatment methods, and testing methods from high clinical value data, and corresponding hash identifiers and metadata are generated based on the extracted disease types, treatment methods, and testing methods.
[0101] Furthermore, based on the sensitive data requiring protection within the high-privacy-risk data and the set privacy protection level, the corresponding privacy-protected data to be released externally is determined, specifically as follows:
[0102]
[0103] in: The output value after perturbation. The original true value, To improve query sensitivity, For privacy budget parameters, Let be a Laplace-distributed random variable.
[0104] In other words, based on the perturbed output value, privacy budget, and data utility corresponding to high privacy risk data, the corresponding differential privacy sharding attributes are set.
[0105] Step S2.3: Fragmented Storage. Based on the fragmentation type set in Step S2.2, a unique semantic identifier is generated for each fragment. Simultaneously, based on the number of relations retained in each fragment and the total number of relations in the original graph, the corresponding fragmentation integrity index is determined. Specifically:
[0106]
[0107] in: This is the fragment integrity index. The number of valid relations retained in the partition. This represents the total number of relations in the original knowledge graph.
[0108] In this embodiment, the obtained fragmentation integrity index is compared with a preset index threshold (which is set according to specific needs, so it is not specifically described in this embodiment, for example, 85%), and the final fragmentation result is determined based on the comparison result. Specifically:
[0109] If the obtained fragment integrity index is less than the preset index threshold (i.e., 85%), then repeat steps S2.1-S2.3 to reassemble the fragments until the obtained fragment integrity index is not less than the preset index threshold. Conversely, if the obtained fragment integrity index is not less than the preset index threshold, then the corresponding fragmentation result is the final fragmentation result.
[0110] Step S3: Data Reassembly. Based on the sharding results determined in Step S2.3, after storing them on heterogeneous nodes, the shards are reassembled in a differentiated manner according to access permission levels. Specifically:
[0111] Step S3.1: Sharded Storage. This involves marking sharded data using verifiable semantics, storing the marked sharded data across heterogeneous nodes, and updating the corresponding storage locations through a dynamic rotation mechanism. Specifically:
[0112] Step S3.1.1: Distributed storage. This involves splitting each data shard into multiple fragments and storing these fragments on nodes in different geographical locations. In other words, each node stores a different combination of shards.
[0113] The specific storage method used in the implementation process is shown in Table 4 below.
[0114] Table 4: Sharding Node Storage Table
[0115] Node position Storage sharding combination Data Center of Top-Tier Hospitals in Shanghai Fragment A + Fragment E Beijing Medical Private Cloud Fragment B + Fragment D Guangzhou edge computing node Fragment C Chengdu Disaster Recovery Center Fragment D + Fragment A Wuhan Scientific Research Cloud Fragment E + Fragment B
[0116] Step S3.1.2: Dynamic Protection. This involves using the RAFT protocol to obtain the nodes corresponding to each shard within a preset time period, identifying duplicate data stored on these nodes, and determining the corresponding round-robin validity based on the total number of nodes. Specifically:
[0117]
[0118] in: For rotational effectiveness, The number of nodes that are repeatedly stored. This represents the total number of nodes.
[0119] Furthermore, the hash value, timestamp, and node signature corresponding to the shard are used as inputs to the zero-knowledge proof model, and the output is the corresponding proof result.
[0120] Step S3.2: Perceptual Reassembly. This involves reassembling the fragmented data according to the access permission level, obtaining the semantic integrity index and semantic deviation of the reassembled data, and determining the final reassembled data based on these indicators. Specifically:
[0121] Step S3.2.1: Access Control. This involves obtaining the corresponding semantic integrity index based on the number of valid clinical relations in the reconstructed data and the clinical relation tree in the original knowledge graph. Specifically:
[0122]
[0123] in: It is a semantic integrity index. The number of valid clinical relationships in the recombinant data. This is the clinical relationship tree in the original knowledge graph.
[0124] In this embodiment, the obtained semantic integrity index is compared with a preset semantic threshold (which is set according to specific needs, so it is not specifically described in this embodiment, for example, 0.9), and the final recombination method is determined based on the comparison result, specifically as follows:
[0125] If the obtained semantic integrity index is less than the preset semantic threshold (i.e., 0.9), missing data retrieval is performed. Conversely, if the obtained semantic integrity index is not less than the preset semantic threshold (i.e., 0.9), complete recombinant data is obtained through the clinical decision subgraph.
[0126] Step S3.2.2: Secure Reassembly. This involves splitting the fragments in each node into multiple segments using the dedicated key pair for each node. Simultaneously, based on the reassembly method determined in step S3.2.1, the split segments from different nodes are aggregated to obtain the corresponding aggregation result.
[0127] Furthermore, based on the semantic integrity index of the reconstructed data corresponding to the aggregation result and the semantic integrity index of the original data, the corresponding semantic deviation is determined, specifically as follows:
[0128]
[0129] in: For semantic deviation, To reorganize the semantic integrity index, This is the semantic integrity index of the original data.
[0130] In this embodiment, the obtained semantic deviation is compared with a preset deviation threshold range (which is set according to specific needs, so it is not specifically described in this embodiment, for example, 10%-15%), and the final recombination result is determined based on the comparison result. Specifically:
[0131] When the obtained semantic deviation is less than the lower limit of the preset deviation threshold range (i.e., 10%), the corresponding reconstructed semantics is the final reconstructed result. When the obtained semantic deviation is within the preset deviation threshold range (i.e., 10%-15%), a warning signal is triggered, and the corresponding reconstructed semantics is marked with low confidence. When the obtained semantic deviation is greater than the upper limit of the preset deviation threshold range (i.e., 15%), step S3.1 is returned, fragmentation and storage are performed again, and steps S3.2.1 and S3.2.2 are repeated.
[0132] This embodiment also provides an AI-based intelligent private data fragmentation and reconstruction system, which uses the aforementioned AI-based intelligent private data fragmentation and reconstruction method.
[0133] Example 2
[0134] This embodiment provides an AI-based intelligent private data fragmentation and reassembly method, which is implemented in the same way as in Embodiment 1. The difference is that an entity co-occurrence matrix is constructed using different entity units, and the initial weights between entity unit combinations are obtained through the entity co-occurrence matrix. The weights obtained by the graph attention network are then adjusted according to the initial weights. The invention will be illustrated below with specific examples of the implementation of this embodiment.
[0135] In this embodiment, the weights obtained by the graph attention network are adjusted according to the initial weights, as follows:
[0136] Step S1.2.1: Determine the initial weights. This involves dividing the diagnostic record text into segments based on the entity units identified by the BioClinicalBERT model, setting the window size on a sentence-by-sentence basis. Simultaneously, the frequency of different entity unit combinations is obtained within the segmented diagnostic record text, and an entity co-occurrence matrix is constructed based on the set entity unit combinations and their corresponding frequencies.
[0137] Furthermore, based on the occurrence frequency and total frequency of each entity unit combination in the entity co-occurrence matrix, the initial weights corresponding to each entity unit combination are set as follows:
[0138]
[0139] in: Let r be the initial weight of the relationship between entity r and entity o. Let r be the number of times entity r and o co-occur in the case. This represents the total number of cases.
[0140] Step S1.2.2: Weight Adjustment. This involves adjusting the weights between different entity units obtained from the graph attention network output based on the initial weights obtained in step S1.2.1. Specifically:
[0141]
[0142] in: The adjusted weights between entity r and entity o. The weights represent the original relationship between entity r and entity o obtained through the graph attention network. Let r be the initial weight of the relationship between entity r and entity o. It is a normalized exponential function.
[0143] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended embodiments and their equivalents.
Claims
1. An AI-based intelligent private data fragmentation and reassembly method, characterized in that, Including: S1: Constructing a medical knowledge graph: Using the identified entities as knowledge graph nodes, construct a medical knowledge graph and label them with clinical importance levels and sensitive data types; S2: Intelligent Sharding: Determines the sharding method based on the partitioning results of the labeled data, including: S2.1: Constructing the assessment matrix: Construct a two-dimensional assessment matrix table based on the privacy risk dimension and the clinical value dimension; S2.2: Segmentation Decision: Data is segmented using a two-dimensional evaluation matrix table to determine the hash identifiers and metadata of high clinical value data, and the privacy-protected data for external release of data with privacy risks; S2.3: Segmented storage: Based on the number of relations retained in each segment and the total number of relations in the original graph, determine the segmentation integrity index, and determine the final segmentation result based on the segmentation integrity index and the preset index threshold; S3: Data Reorganization: Based on the sharding results, the shards are stored in heterogeneous nodes, and the data is reorganized according to the access permission level, including: S3.1: Sharded storage: Sharded data with verifiable semantic tags are stored in heterogeneous nodes, and the storage location is updated through a dynamic rotation mechanism; S3.2: Perceptual Reassembly: Based on access permission levels, reassemble fragmented data, obtain semantic integrity index and semantic deviation, and determine the final reassembled data, including: S3.2.1: Access control: Based on the number of valid clinical relations in the reconstructed data and the clinical relation tree of the original knowledge graph, obtain the semantic integrity index, and determine the final reconstruction method based on the semantic integrity index and the preset semantic threshold; S3.2.2: Secure Reassembly: Based on the node-specific key pair and the final reassembly method, the split fragments are aggregated, and the semantic deviation is obtained based on the semantic integrity index of the reassembly result and the semantic integrity index of the original data to determine the final reassembly result.
2. The AI-based intelligent private data fragmentation and reassembly method according to claim 1, characterized in that, The clinical importance level and sensitive data types are marked, including: S1.1: Data processing: Obtain patient information, key paragraphs in medical records and DICOM files through the HIS system, EMR text and medical image metadata, and perform data cleaning and normalization on the patient information, key paragraphs in medical records and DICOM files. S1.2: Semantic parsing: Using the BioClinicalBERT model, entity units in the diagnostic record text are identified. At the same time, the entity units are used as input to the graph attention network, and the output obtains the weight between different entity units and constructs relation triples. S1.3: Privacy Labeling: Based on the life support index, treatment complexity, and update frequency, clinical importance and sensitivity are graded and associated with the aforementioned relational triples to construct a medical knowledge graph.
3. The AI-based intelligent private data fragmentation and reassembly method according to claim 2, characterized in that, An entity co-occurrence matrix is constructed using the entity units, and initial weights for entity unit combinations are set using the entity co-occurrence matrix. Simultaneously, the weights of the graph attention network are adjusted based on these initial weights, including: S1.2.1: Determine initial weights: Divide the diagnostic record text into segments according to the set window size, obtain the frequency of occurrence of different entity unit combinations, construct an entity co-occurrence matrix, and determine the initial weights for each entity unit combination, specifically: , in: Let r be the initial weight of the relationship between entity r and entity o. Let r be the number of times entity r and o co-occur in the case. Total number of cases; S1.2.2: Weight Adjustment: Based on the initial weights, determine the adjusted weights, specifically as follows: , in: The adjusted weights between entity r and entity o. The weights represent the original relationship between entity r and entity o obtained through the graph attention network. Let r be the initial weight of the relationship between entity r and entity o. It is a normalized exponential function.
4. The AI-based intelligent private data fragmentation and reassembly method according to claim 2, characterized in that, The constructed medical knowledge graph includes: S1.3.1: Clinical Importance Grading: Based on the set life support index score, treatment complexity score, and update frequency score, a grading score is obtained. Simultaneously, the grading score is compared with a preset grading threshold range, and based on the comparison result, the clinical importance level is determined, specifically as follows: When the grading score is less than the lower limit of the preset grading threshold range, the clinical importance level is L3; when the grading score is within the preset grading threshold range, the clinical importance level is L2; when the grading score is greater than the upper limit of the preset grading threshold range, the clinical importance level is L1. S1.3.2: Sensitivity Marker: Set the sensitivity level based on the text content of the diagnostic record; S1.3.3: Knowledge Graph Generation: Based on the entity recognition results, clinical importance level, and sensitivity level, node attributes in the knowledge graph are set. At the same time, based on different entity units and weights, relation attributes in the knowledge graph are set. Through the node attributes and relation attributes, Neo4j graph data is constructed.
5. The AI-based intelligent private data fragmentation and reassembly method according to claim 1, characterized in that, Based on PubMed citation data and the clinical research registry platform, keyword combinations in treatment protocols were extracted, and a research value score was determined according to the number of extracted references. Simultaneously, a recommendation level mapping table was established based on the recommendation levels of clinical guidelines to determine the treatment effectiveness score. Based on the treatment effectiveness score and the research value score, a clinical value dimension was set.
6. The AI-based intelligent private data fragmentation and reassembly method according to claim 1, characterized in that, Based on the HIPAA classification mapping table, a basic risk value for rare diseases is set, and the basic risk value is adjusted according to the prevalence of the disease. At the same time, a privacy risk dimension is set based on the adjusted risk value.
7. The AI-based intelligent private data fragmentation and reassembly method according to claim 1, characterized in that, Updating the storage location includes: S3.1.1: Distributed storage: Each data shard is split into multiple fragments, and these fragments are stored in nodes at different geographical locations. S3.1.2: Dynamic Protection: Through the RAFT protocol, the node data that is stored repeatedly is identified, and the validity of the rotation is obtained based on the total number of nodes. At the same time, the hash value of the shard, the timestamp, and the node signature are used as inputs to the zero-knowledge proof model, and the proof result is output.
8. The AI-based intelligent private data fragmentation and reassembly method according to claim 1, characterized in that, The fragmentation integrity index is compared with a preset index threshold, and the final fragmentation result is determined based on the comparison result, specifically as follows: When the fragment integrity index is less than the preset index threshold, steps S2.1-S2.3 are repeated to reassemble the fragments until the fragment integrity index is not less than the preset index threshold; otherwise, the corresponding fragmentation result is the final fragmentation result. The semantic integrity index is compared with a preset semantic threshold, and the final recombination method is determined based on the comparison result, specifically as follows: When the semantic integrity index is less than a preset semantic threshold, missing data retrieval is performed; otherwise, recombinant data is obtained through the clinical decision subgraph. The semantic deviation is compared with a preset deviation threshold range, and the final reorganization result is determined based on the comparison result, specifically as follows: When the semantic deviation is less than the lower limit of the preset deviation threshold range, the corresponding reconstructed semantics is the final reconstructed result; when the semantic deviation is within the preset deviation threshold range, a warning signal is triggered, and the corresponding reconstructed semantics is marked with low confidence; when the semantic deviation is greater than the upper limit of the preset deviation threshold range, step S3.1 is returned, fragmentation and storage are performed again, and steps S3.2.1 and S3.2.2 are repeated.
9. An AI-based intelligent private data fragmentation and reassembly system, characterized in that, The method for intelligent private data fragmentation and reconstruction based on AI, as described in any one of claims 1-8, was used.
Citation Information
Patent Citations
A data desensitization method for preventing privacy leakage
CN109033873A
Method of using cloud to safely and dynamically transmit data
CN111740951A
HICH intelligent rehabilitation large health system based on image and text recognition multi-round complementation
CN119418850A