Method and System for Constructing a Tumor Early Screening Data Sharing Platform Based on Cloud Computing

Through the organic integration of knowledge graph, encryption technology, distributed storage and blockchain, the problem of insufficient data silos and privacy protection in traditional tumor early screening data management is solved, efficient integration and secure sharing of multi-source heterogeneous data is achieved, and a safe and efficient tumor early screening data sharing platform is built.

CN119920488BActive Publication Date: 2025-07-08SHENZHEN RAPHA BIOTECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510403230.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-08
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The traditional tumor early screening data management model has data silos, low sharing efficiency, and insufficient privacy protection, which is difficult to meet the needs of modern medical care for efficient utilization and secure sharing of large-scale, multi-source heterogeneous data.

Method used

The knowledge graph technology is used to model the multi-dimensional early screening data association, generate a metadata index library, and encrypt sensitive fields; the metadata index library is partitioned by tumor type in a distributed storage system on the cloud platform, and the erasure coding technology is used to perform cross-region redundant backup; access control smart contracts are built based on blockchain technology, and the requesting party qualifications are verified through the digital identity authentication module, dynamically decrypt and transmit data copies that meet the permission level.

Benefits of technology

It realizes efficient integration and sharing of multi-source heterogeneous data, ensures data privacy and security, improves data disaster recovery capabilities and availability, ensures data sharing security and compliance, and builds a safe, efficient and reliable tumor premature screening data sharing platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920488B_ABST
    Figure CN119920488B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of big data technology, in particular to a method and system for constructing a tumor early screening data sharing platform based on cloud computing. The method for constructing the sharing platform includes: using knowledge graph technology to perform associated modeling on multi-dimensional early screening data; encrypting sensitive fields in the metadata index library, and synchronously generating irreversible anonymous identifiers to replace the original identifiers; arranging the metadata index library in partitions according to tumor types in the corresponding distributed storage system of the cloud platform, and using erasure code technology for cross-regional redundant backup; constructing an access control intelligent contract based on blockchain technology, verifying the qualifications of the requesting party through a digital identity authentication module, and after verifying the validity of the digital certificate of the requesting party, dynamically decrypting and transmitting data copies that meet the permission level. The present invention constructs a safe, efficient and reliable tumor early screening data sharing platform through the organic integration of knowledge graph, encryption technology, distributed storage and blockchain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and particularly to a method and system for constructing a tumor early screening data sharing platform based on cloud computing. Background Art

[0002] In recent years, with the continuous rise of the global tumor incidence rate, the importance of tumor early screening in disease prevention and control has become increasingly prominent. The traditional tumor early screening data management mode has problems such as data islands, low sharing efficiency, and insufficient privacy protection, making it difficult to meet the requirements of modern medicine for the efficient utilization and secure sharing of large-scale, multi-source heterogeneous data. In this context, cloud computing technology, with its advantages of elastic expansion, high availability, and on-demand services, provides a new technical architecture for tumor early screening data sharing. Therefore, constructing a tumor early screening data sharing platform based on cloud computing can not only break data islands, achieve the efficient integration and sharing of multi-source heterogeneous data, but also ensure the privacy security and compliance of data during the sharing process through advanced encryption and access control mechanisms, providing strong support for the precise and intelligent development of tumor early screening. Summary of the Invention

[0003] The present invention overcomes the deficiencies of the prior art and provides a method and system for constructing a tumor early screening data sharing platform based on cloud computing.

[0004] The technical solution adopted by the present invention to achieve the above object is as follows:

[0005] In the first aspect of the present invention, a method for constructing a tumor early screening data sharing platform based on cloud computing is disclosed, including:

[0006] Obtain target multi-dimensional early screening data, and use knowledge graph technology to perform association modeling on the multi-dimensional early screening data to generate a metadata index library;

[0007] Encrypt the sensitive fields of the metadata index library, and synchronously generate irreversible anonymous identifiers to replace the original identifiers;

[0008] Arrange the metadata index library in the corresponding distributed storage system of the cloud platform according to tumor types, and use erasure code technology for cross-regional redundant backup;

[0009] Construct an access control smart contract based on blockchain technology, verify the qualifications of the requester through a digital identity authentication module, and after verifying the validity of the requester's digital certificate, dynamically decrypt and transmit data copies that meet the permission level.

[0010] Preferably, using knowledge graph technology to perform association modeling on multi-dimensional early screening data to generate a metadata index library is specifically:

[0011] Perform feature analysis and processing on each multi-dimensional early screening data to obtain the key entities of each multi-dimensional early screening data;

[0012] Iteratively learn the semantic association strength relationship between key entities through a graph neural network, and construct an initial knowledge graph according to the semantic association strength relationship;

[0013] Based on a preset rule library in the field of early cancer screening, logically verify the management strength relationship between key entities in the initial knowledge graph, and identify and delete redundant relationships that do not conform to domain knowledge;

[0014] Calculate the cosine similarity between each key entity in the initial knowledge graph; if the cosine similarity between two key entities is greater than a preset threshold, perform a pruning operation on any one of them;

[0015] Use a rule inference engine to verify the missing relationships in the pruned knowledge graph, and supplement the missing key entities through entity attribute matching and semantic similarity calculation to generate an early screening data knowledge graph;

[0016] Initialize several data indexes, generate a unique semantic identifier for each data index, and map each data index to the key entities in the early screening data knowledge graph to obtain a metadata index library.

[0017] Preferably, encrypt the sensitive fields of the metadata index library, and synchronously generate an irreversible anonymous identifier to replace the original identifier. Specifically:

[0018] According to a preset sensitive information recognition rule library, match the fields in the metadata index library item by item. If the field name or content conforms to the sensitive information characteristics, mark it as a sensitive field;

[0019] Perform format verification on the marked sensitive fields. If the field format conforms to the preset specification, further extract its key feature values;

[0020] Use a hash algorithm to perform irreversible encryption processing on the extracted key feature values to generate a unique anonymous identifier;

[0021] Map the generated anonymous identifier to the original identifier and store it in an encryption mapping table, and at the same time delete the original identifier field;

[0022] If the field format does not conform to the preset specification, use a homomorphic encryption algorithm to encrypt it to ensure that the data can still be calculated in the encrypted state;

[0023] Control the access rights to the encryption mapping table, only authorize specific modules to perform decryption operations when meeting preset conditions, and record all access logs for auditing.

[0024] Preferably, an irreversible encryption process is performed on the extracted key feature values using a hash algorithm to generate a unique anonymous identifier, specifically as follows:

[0025] Preprocess the extracted key feature values. If the feature values contain special characters or spaces, remove the invalid characters and uniformly convert them to a standardized string format;

[0026] According to the preset hash algorithm configuration parameters, perform a hash operation on the standardized string to generate a hash value of a fixed length;

[0027] Perform a uniqueness check on the generated hash value. If the hash value already exists in the anonymous identifier library, use the salt addition technique to add a random salt value to the original string and then perform the hash operation again until a unique hash value is generated;

[0028] Map the generated unique hash value to the original feature value and store it in the encryption mapping table, and at the same time generate the corresponding anonymous identifier;

[0029] If the original feature value is empty or invalid, skip the hash processing and mark it as an invalid record;

[0030] Perform format standardization processing on the generated anonymous identifier to ensure that it conforms to the system identifier specification, and write it into the metadata index library to replace the original identification field.

[0031] Preferably, the metadata index library is partitioned by tumor type and arranged in the corresponding distributed storage system of the cloud platform, and the erasure code technology is used for cross-region redundant backup, specifically as follows:

[0032] Based on the preset tumor type classification rule library, perform type matching on the data entities in the metadata index library. If the data contains specific tumor markers or pathological features, it is automatically assigned to the corresponding tumor type partition;

[0033] Dynamically adjust the storage location of the data partition according to the real-time load status of the distributed storage nodes. If it is detected that the storage pressure of the target load node exceeds the threshold, migrate the data to the load node whose storage pressure does not exceed the threshold;

[0034] Perform sharding processing on the data blocks of each tumor type partition, and use the erasure code algorithm to encode the data shards and parity blocks according to the preset ratio. If the sharded data volume does not meet the encoding requirements, supplement virtual padding data to complete the sharding;

[0035] According to the cross-region backup strategy, store the data shards and parity blocks in the storage nodes of different geographical regions respectively, ensuring that the data shards and parity blocks of the same partition are completely isolated geographically.

[0036] Preferably, an access control smart contract is constructed based on blockchain technology, and the qualification of the requester is verified through a digital identity authentication module. After verifying the validity of the digital certificate of the requester, a copy of the data that meets the permission level is dynamically decrypted and transmitted, specifically:

[0037] Based on the preset access rights policy template, create a smart contract instance in the blockchain network, define data access rights and decryption rules, and deploy the contract to distributed nodes;

[0038] When the requester initiates a data access request, the smart contract automatically triggers the digital identity authentication module. If the digital certificate provided by the requester fails to pass the consensus verification of the blockchain node, the access process is immediately terminated and the abnormal log is recorded;

[0039] For the verified requester, the permission label embedded in its certificate is matched with the corresponding permission level in the metadata index library. If the permission label does not match the minimum access level of the target data, a prompt indicating insufficient permissions is returned.

[0040] For data requests that meet the permissions, the homomorphic encryption module is called to generate a dynamic decryption key. If the data needs to be shared across institutions, the target data shard is jointly decrypted through a multi-party secure computing protocol;

[0041] The decrypted data copy is desensitized again according to the scope of authority and pushed to the requester through the encrypted transmission channel in the blockchain network. At the same time, the complete operation record is stored on the chain.

[0042] Among them, the multidimensional early screening data includes medical images, pathological sections, gene sequencing data and electronic medical record texts.

[0043] The second aspect of the present invention discloses a cloud computing-based tumor early screening data sharing platform construction system, the tumor early screening data sharing platform construction system includes a memory and a processor, the memory stores a tumor early screening data sharing platform construction method program, when the tumor early screening data sharing platform construction method program is executed by the processor, any step of the tumor early screening data sharing platform construction method is implemented.

[0044] The third aspect of the present invention discloses a computer-readable storage medium, which includes a method program for constructing a tumor early screening data sharing platform. When the method program for constructing a tumor early screening data sharing platform is executed by a processor, the steps of any one of the methods for constructing a tumor early screening data sharing platform are implemented.

[0045] The present invention solves the technical defects existing in the background art and has the following beneficial effects: First, obtain target multi-dimensional early screening data, use knowledge graph technology to perform correlation modeling on the data, generate a metadata index library, and realize the semanticization and structuring of the data; Second, encrypt sensitive fields in the metadata index library and generate irreversible anonymous identifiers to replace the original identifiers to ensure data privacy and security; Then, arrange the metadata index library in a distributed storage system of the cloud platform according to tumor types and use erasure coding technology for cross-regional redundant backup to improve the disaster tolerance and availability of the data; Finally, construct an access control smart contract based on blockchain technology, verify the qualifications of the requesting party through a digital identity authentication module, and dynamically decrypt and transmit data copies that meet the permission levels after verification to ensure the security and compliance of data sharing. Through the organic integration of knowledge graph, encryption technology, distributed storage, and blockchain, the present invention constructs a safe, efficient, and reliable tumor early screening data sharing platform, providing technical support for the sharing and application of tumor early screening data. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0047] Figure 1 It is an overall method flowchart of a method for constructing a tumor early screening data sharing platform based on cloud computing;

[0048] Figure 2 It is a partial method flowchart of a method for constructing a tumor early screening data sharing platform based on cloud computing;

[0049] Figure 3 It is a system block diagram of a system for constructing a tumor early screening data sharing platform based on cloud computing. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] In order to more clearly understand the above-mentioned objects, features, and advantages of the present invention, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.

[0051] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0052] As Figure 1 shown, the first aspect of the present invention discloses a method for constructing a tumor early screening data sharing platform based on cloud computing, including:

[0053] S102. Obtain target multi-dimensional early screening data, perform association modeling on the multi-dimensional early screening data by using knowledge graph technology, and generate a metadata index library;

[0054] S104. Encrypt sensitive fields in the metadata index library, and synchronously generate irreversible anonymous identifiers to replace the original identifiers;

[0055] S106. Arrange the metadata index library in the corresponding distributed storage system of the cloud platform according to tumor types, and perform cross-regional redundant backup by using erasure code technology;

[0056] S108. Construct an access control smart contract based on blockchain technology, verify the qualifications of the requester through a digital identity authentication module, and after verifying the validity of the requester's digital certificate, dynamically decrypt and transmit data copies that meet the permission level.

[0057] Among them, the multi-dimensional early screening data includes medical images, pathological sections, gene sequencing data, and electronic medical record texts.

[0058] It should be noted that, first, obtain the target multi-dimensional early screening data, perform association modeling on the data by using knowledge graph technology, and generate a metadata index library to realize the semanticization and structuring of the data; second, encrypt sensitive fields in the metadata index library and generate irreversible anonymous identifiers to replace the original identifiers to ensure data privacy and security; then, arrange the metadata index library in the distributed storage system of the cloud platform according to tumor types and perform cross-regional redundant backup by using erasure code technology to improve the disaster tolerance and availability of the data; finally, construct an access control smart contract based on blockchain technology, verify the qualifications of the requester through a digital identity authentication module, and after passing the verification, dynamically decrypt and transmit data copies that meet the permission level to ensure the security and compliance of data sharing. The present invention constructs a safe, efficient, and reliable tumor early screening data sharing platform through the organic integration of knowledge graph, encryption technology, distributed storage, and blockchain, providing technical support for the sharing and application of tumor early screening data.

[0059] Preferably, performing association modeling on the multi-dimensional early screening data by using knowledge graph technology and generating a metadata index library specifically includes:

[0060] Perform feature analysis and processing on each multi-dimensional early screening data to obtain key entities of each multi-dimensional early screening data;

[0061] Among them, the key entities include but are not limited to patient names, key gene sequence features, key pathological text description features, and lesion area imaging features.

[0062] It should be noted that for format verification of multi-dimensional early screening data, if the data is unstructured text (such as a pathological report), natural language processing technology is used for entity recognition to extract key information such as patient names and pathological features; if the data is structured data (such as gene sequences), the feature fields are directly extracted. Secondly, for medical image data, lesion area detection is performed through a deep learning model (such as a convolutional neural network) to extract imaging features.

[0063] Iteratively learn the semantic association strength relationship between key entities through a graph neural network, and construct an initial knowledge graph according to the semantic association strength relationship.

[0064] It should be noted that the key entities and their initial relationships are represented as a graph structure, where the entities are nodes and the initial relationships are edges, and an initial weight is assigned to each edge. Secondly, a graph neural network model (such as GraphSAGE or GAT) is used to perform multiple rounds of iterative learning on the graph structure. In each round of iteration, the entity representation is updated by aggregating the feature information of neighboring nodes, and the semantic association strength between entities is calculated. Then, the weight of the edge is adjusted according to the updated semantic association strength. If the weight is lower than the preset threshold, the edge is removed. Then, the graph neural network parameters are optimized through the backpropagation algorithm to ensure that the model can more accurately capture the semantic relationship between entities. Finally, the optimized graph structure is output as an initial knowledge graph, where the nodes represent key entities and the edges represent the semantic association strength relationship, providing a basis for subsequent knowledge graph optimization.

[0065] Based on a preset rule library in the field of early tumor screening, logically verify the management strength relationship between key entities in the initial knowledge graph, and identify and delete redundant relationships that do not conform to domain knowledge.

[0066] Among them, the rule library in the field of early tumor screening is a set of rules specifically constructed by relevant technical personnel for the field of early tumor screening, which includes logical constraints on the relationships between entities, semantic association rules, and standardized definitions of domain knowledge. These rules are based on medical literature, clinical guidelines, and expert experience, and are used to describe the reasonable associations and constraint conditions between key entities (such as gene mutations, pathological features, imaging features, etc.) in early tumor screening data. For example, the rule library contains logical rules such as "a specific gene mutation has a strong association with a specific pathological feature" or "a benign tumor should not be directly associated with a malignant gene mutation". Through the rule library, the entity relationships in the knowledge graph can be logically verified to ensure that they conform to the professional knowledge in the medical field, thereby improving the accuracy and reliability of the knowledge graph.

[0067] It should be noted that the rule library in the field of early tumor screening is loaded, which includes logical constraints on the relationships between entities (such as the association rules between "gene mutation type" and "pathological features"); traverse each edge in the initial knowledge graph, if the semantic relationship of the edge conflicts with the logical constraints in the rule library (such as the unreasonable association between "benign tumor" and "malignant gene mutation"), then mark it as a relationship to be deleted; then, conduct a secondary verification on the marked redundant relationships, if it is confirmed that they do not conform to the domain knowledge, then remove them from the knowledge graph; then, check the connectivity of the knowledge graph after removing the redundant relationships, if isolated nodes are found, supplement the necessary relationship edges according to the rule library; finally, output the optimized knowledge graph to ensure that it conforms to the professional knowledge logic in the field of early tumor screening.

[0068] Calculate the cosine similarity between each pair of key entities in the initial knowledge graph; if the cosine similarity between a certain pair of key entities is greater than the preset threshold, then perform a pruning operation on any one of them.

[0069] It should be noted that the cosine similarity between each pair of key entities in the initial knowledge graph is calculated by the cosine similarity algorithm.

[0070] Use the rule inference engine to verify the missing relationships in the pruned knowledge graph, supplement the missing key entities through entity attribute matching and semantic similarity calculation, and generate the early screening data knowledge graph.

[0071] It should be noted that based on the rule library in the field of early tumor screening, define the inference rules for the missing relationships between entities (such as "specific gene mutations may lead to specific pathological features"); traverse the pruned knowledge graph, if it is found that there are potential relationships between entities but they are not directly connected, then trigger the rule inference engine for relationship prediction; then, calculate the attribute similarity between entities through the entity attribute matching algorithm, if the similarity is higher than the preset threshold, then supplement the missing relationship edge; then, use the semantic similarity calculation model (such as BERT-based semantic embedding) to further verify the rationality of the supplemented relationship, if the semantic similarity conforms to the domain knowledge logic, then officially add the relationship to the knowledge graph; finally, output the supplemented early screening data knowledge graph to ensure its integrity and accuracy.

[0072] Initialize several data indexes, generate a unique semantic identifier for each data index, and map each data index to the key entities in the early screening data knowledge graph to obtain the metadata index library.

[0073] It should be noted that based on the key entities in the early screening data knowledge graph, the core features of each piece of data (such as patient ID, gene sequence, pathological features, etc.) are extracted as the initial content of the data index; the hash algorithm is used to perform irreversible encryption on the core features to generate a unique semantic identifier; the generated semantic identifier is bound to the corresponding data record and stored in the index mapping table; according to the semantic relationships of the key entities in the knowledge graph, the data index is mapped to the entity nodes to establish a two-way association between the index and the entity; the mapped index information is integrated into the metadata index library to ensure that each piece of data can be quickly retrieved through the semantic identifier and accurately associated with the entities in the knowledge graph.

[0074] In summary, the present invention realizes the semantic association of multi-dimensional data through knowledge graph technology, solves the standardization problem of multi-source heterogeneous data; based on the collaborative optimization of graph neural network and rule reasoning, ensures the accuracy and integrity of the knowledge graph, and provides reliable knowledge support for subsequent data mining and analysis; by generating a unique semantic identifier and metadata index library, realizes the efficient retrieval and cross-institutional sharing of data, improves the data utilization rate; at the same time, through logical verification and redundant relationship pruning, reduces the impact of data noise on the analysis results, and provides a high-quality data foundation for tumor early screening research.

[0075] Preferably, the sensitive fields in the metadata index library are encrypted, and an irreversible anonymous identifier is synchronously generated to replace the original identifier, specifically:

[0076] According to the preset sensitive information recognition rule library, the fields in the metadata index library are matched item by item. If the field name or content conforms to the sensitive information characteristics (such as patient ID, gene locus, contact information, etc.), it is marked as a sensitive field;

[0077] Among them, the sensitive information recognition rule library refers to a set of rules specifically used to identify and mark sensitive data, which includes the feature definitions, format specifications, and matching rules of sensitive fields. It is used to accurately identify the sensitive information in the metadata index library (such as patient ID, gene locus, contact information, etc.). For example, the rule library defines rules such as "the ID card number consists of 18 digits" or "the gene locus needs to conform to a specific naming specification". Through the rule library, the system can automatically detect and mark sensitive fields, providing a basis for subsequent encryption and anonymization processing to ensure that data privacy protection complies with compliance requirements.

[0078] Perform format verification on the marked sensitive fields. If the field format conforms to the preset specifications (such as ID card number, telephone number, etc.), further extract its key feature values;

[0079] Use the hash algorithm to perform irreversible encryption on the extracted key feature values to generate a unique anonymous identifier;

[0080] Map the generated anonymous identifier to the original identifier, store it in the encrypted mapping table, and delete the original identifier field;

[0081] If the field format does not conform to the preset specification, use the homomorphic encryption algorithm to encrypt it to ensure that calculations can still be performed on the encrypted data;

[0082] It should be noted that data cleaning is performed on fields that do not conform to the format specification, invalid characters are removed, and the encoding format is unified to ensure that the data can be processed by the encryption algorithm; Secondly, select a suitable homomorphic encryption algorithm (such as Paillier or CKKS), and generate an encryption key pair according to the preset key generation rules; Then, perform block processing on the cleaned field data. If the data length exceeds the maximum value supported by the algorithm, it is split into multiple sub-blocks; Then, use the public key to encrypt each sub-block to generate encrypted data blocks, and splice the encrypted data blocks in order into a complete encrypted field; Finally, store the encrypted field in the metadata index library to ensure that it still supports calculations such as addition or multiplication in the encrypted state, and securely store the decryption key for use by the authorization module.

[0083] Control the access rights to the encrypted mapping table, only authorize specific modules to perform decryption operations when meeting the preset conditions, and record all access logs for auditing.

[0084] In summary, through sensitive information identification and encryption processing, it is ensured that sensitive fields such as patient ID and gene loci cannot be reverse-derived during storage and transmission, reducing the risk of data leakage; use irreversible anonymous identifiers to replace the original identifier, maintaining the data's associability while protecting privacy; through homomorphic encryption technology, support data calculations in the encrypted state, taking into account both data security and availability; strict access control and audit log recording further enhance the transparency and controllability of data access.

[0085] Preferably, use the hash algorithm to perform irreversible encryption processing on the extracted key feature values to generate a unique anonymous identifier, specifically:

[0086] Preprocess the extracted key feature values. If the feature values contain special characters or spaces, remove the invalid characters and uniformly convert them to the standardized string format;

[0087] According to the preset hash algorithm configuration parameters, perform a hash operation on the standardized string to generate a hash value of a fixed length;

[0088] Perform uniqueness verification on the generated hash value. If the hash value already exists in the anonymous identifier library, use the salting technique to add a random salt value to the original string and then perform a hash operation again until a unique hash value is generated;

[0089] Among them, the salting technique refers to adding a randomly generated string (referred to as "salt value") to the original data (such as a string) during the hashing operation to enhance the uniqueness and security of the hash result. Its core purpose is to prevent hash collisions (i.e., different inputs generate the same hash value) and resist rainbow table attacks (reverse-deriving the original data through pre-computed hash values).

[0090] It should be noted that the generated hash value is compared with the existing hash values in the anonymous identifier library. If a duplicate is found, the salting technique is triggered for processing; secondly, a random salt value is generated and appended to the end of the original string to ensure that the length and randomness of the salt value meet the preset security requirements; then, the same hashing algorithm is used to re-hash the salted string to generate a new hash value; then, the new hash value is compared with the anonymous identifier library again. If there are still duplicates, the above salting and hashing operation process is repeated until a unique hash value is generated; finally, the unique hash value is stored in the anonymous identifier library, and the corresponding mapping relationship between the original string and the salt value is recorded to ensure the accuracy of subsequent retrieval and verification.

[0091] Map the generated unique hash value to the original feature value and store it in the encrypted mapping table, and at the same time generate the corresponding anonymous identifier;

[0092] If the original feature value is empty or invalid, skip the hashing process and mark it as an invalid record;

[0093] Standardize the format of the generated anonymous identifier to ensure that it conforms to the system identifier specification, and write it into the metadata index library to replace the original identification field.

[0094] In summary, an irreversible anonymous identifier is generated through the hashing algorithm to ensure that the original feature value cannot be reverse-derived, reducing the risk of data leakage; the salting technique is adopted to solve the hash collision problem and ensure the uniqueness of the anonymous identifier; the mapping relationship between the hash value and the original feature value is stored through the encrypted mapping table to support efficient retrieval and verification when necessary.

[0095] Preferably, the metadata index library is partitioned by tumor type and arranged in the corresponding distributed storage system of the cloud platform, and the erasure code technology is used for cross-regional redundant backup, as Figure 2 shown, specifically:

[0096] S202. Based on the preset tumor type classification rule library, perform type matching on the data entities in the metadata index library. If the data contains specific tumor markers or pathological features, it is automatically assigned to the corresponding tumor type partition;

[0097] Among them, the tumor type classification rule library refers to a set of rules specifically used to identify and classify tumor types, which includes matching rules and classification criteria for key indicators such as tumor markers, pathological features, and gene mutations. These rules are based on medical literature, clinical guidelines, and expert consensus, and are used to accurately identify tumor types (such as lung cancer, breast cancer, gastric cancer, etc.) in the metadata index library. For example, the rule library defines rules such as "EGFR gene mutation is related to lung cancer" or "HER2 protein overexpression is related to breast cancer".

[0098] S204. Dynamically adjust the storage location of data partitions according to the real-time load status of distributed storage nodes. If it is detected that the storage pressure of the target load node exceeds the threshold, migrate the data to a load node whose storage pressure does not exceed the threshold.

[0099] S206. Perform sharding processing on the data blocks in each tumor type partition. Use the erasure code algorithm to encode the data shards and parity blocks according to a preset ratio (such as the 4 + 2 mode). If the amount of sharded data does not meet the encoding requirements, supplement virtual padding data to make a complete shard.

[0100] It should be noted that the data blocks in each tumor type partition are sharded according to a preset size (such as 1MB). If the amount of data in the last shard is insufficient, supplement virtual padding data (such as all-zero bytes) to make a complete shard. According to the preset erasure code mode, encode the data shards and parity blocks in proportion. Among them, 4 data shards generate 2 parity blocks to ensure that the original data can still be restored when any 2 shards are lost. Mark and store the encoded data shards and parity blocks in distributed storage nodes respectively, and record the mapping relationship between the shards and parity blocks at the same time. Then, mark the virtual padding data to ensure that the padding part can be accurately identified and removed during data recovery to restore the original data content.

[0101] S208. According to the cross-region backup strategy, store the data shards and parity blocks in storage nodes in different geographical regions respectively, ensuring that the data shards and parity blocks of the same partition are completely geographically isolated.

[0102] It should be noted that according to the preset cross-region backup strategy, select storage nodes in multiple geographical regions to ensure that the geographical distance between the nodes meets the isolation requirements (such as at least 500 kilometers apart). Allocate the data shards and parity blocks of each tumor type partition to storage nodes in different regions in sequence, record the storage location of each shard and parity block to ensure that it is not in the same region as other shards or parity blocks of the same partition. Then, regularly check the availability and data integrity of the storage nodes. If it is found that a node in a certain region is unavailable, trigger the data recovery mechanism to reconstruct the lost data using the shards and parity blocks in other regions; finally, update the storage location record and ensure that the new data shards and parity blocks continue to meet the geographical isolation requirements.

[0103] In summary, through the zonal arrangement of tumor types, the classified management and efficient retrieval of data are realized, facilitating subsequent analysis and application; based on the dynamic adjustment mechanism of real-time load status, the reasonable allocation and efficient utilization of storage resources are ensured, avoiding node overload; the erasure code technology is adopted to fragment and encode data, which not only improves the storage efficiency but also enhances the fault tolerance of data; the cross-regional backup strategy ensures the geographical isolation and redundancy of data, effectively coping with the risk of data loss caused by natural disasters or hardware failures.

[0104] Preferably, an access control smart contract is constructed based on blockchain technology. After verifying the qualifications of the requester through the digital identity authentication module and validating the digital certificate of the requester, a data copy that meets the permission level is dynamically decrypted and transmitted. Specifically:

[0105] Based on a preset access permission policy template, a smart contract instance is created in the blockchain network, defining data access permissions and decryption rules, and deploying the contract to distributed nodes;

[0106] Among them, the access permission policy template is a pre-defined set of rules used to regulate the scope of data access permissions and operating conditions. Based on multi-dimensional factors such as roles, institutions, and data sensitivity, the template defines the access levels (such as read-only, read-write, delete, etc.) of different users or organizations to data and access restriction conditions (such as time range, data category, etc.). For example, the template may stipulate that "research institutions can only access de-identified data" or "medical institutions can access encrypted complete data but cannot download it". Through the access permission policy template, the system can automatically match user permissions with data access requests, ensuring the compliance and security of data sharing. Distributed nodes refer to computers or servers that operate independently in a distributed system. Each node has storage, computing, and communication capabilities and can cooperate with other nodes to complete common tasks.

[0107] When the requester initiates a data access request, the smart contract automatically triggers the digital identity authentication module. If the digital certificate provided by the requester fails the consensus verification of the blockchain node (such as the certificate expiration or invalid signature), the access process is immediately terminated and an exception log is recorded;

[0108] For the requester who passes the verification, according to the permission label embedded in its certificate, the corresponding permission level in the metadata index library is matched. If the permission label does not match the lowest access level of the target data, a permission insufficient prompt is returned;

[0109] For data requests that meet the permissions, the homomorphic encryption module is called to generate a dynamic decryption key. If the data needs to be shared across institutions, the target data shards are jointly decrypted through a multi-party secure computation protocol;

[0110] The decrypted data copy is desensitized again according to the scope of authority and pushed to the requester through the encrypted transmission channel in the blockchain network. At the same time, the complete operation record is stored on the chain.

[0111] It should be noted that, according to the permission level of the requester, the homomorphic encryption module is called to generate a dynamic decryption key to ensure that the key is only valid for the current request and is time-effective. If the data needs to be shared across institutions, the multi-party secure computing protocol is triggered to distribute the encrypted data shards to each participating institution. Each institution uses the local private key to partially decrypt the shards and generate an intermediate decryption result; then, the intermediate decryption result is aggregated to the coordination node through a secure communication channel, and the coordination node completes the final decryption; finally, the decrypted data shards are secondary desensitized according to the permission scope, and pushed to the requester through an encrypted transmission channel, while recording the decryption operation log for auditing.

[0112] In summary, the present invention automatically performs identity authentication, permission matching and data decryption through smart contracts, thereby achieving full-process security and controllability of data sharing.

[0113] In this embodiment, the method for constructing a tumor early screening data sharing platform may further include the following steps:

[0114] Based on the fuzzy clustering algorithm, the tumor early screening data to be released in the tumor early screening data sharing platform are grouped and aggregated, and the statistical indicators of the grouped data (such as mean, proportion, frequency distribution, etc.) are calculated. If the statistical results contain sensitive information (such as extremely low frequency events), they are marked as data that needs to be protected;

[0115] According to the preset differential privacy budget parameters, the utility function value of the data to be protected under the exponential mechanism is calculated, and the perturbation probability of each possible output result is calculated based on the utility function value to generate the perturbation probability distribution of the data to be protected;

[0116] Based on the perturbation probability distribution, the data to be protected is randomly sampled to generate the perturbation result;

[0117] Calculate the degree of overlap between the disturbed result and the original data. If the degree of overlap is greater than a preset overlap threshold, recalculate the disturbance probability distribution and perform sampling again until the degree of overlap is no greater than the preset overlap threshold.

[0118] The statistical results after disturbance are published, and the noise addition parameters and operation logs are recorded to ensure subsequent traceability and auditing.

[0119] It should be noted that the fuzzy clustering algorithm allows data points to belong to multiple categories with a certain probability, and can better handle the uncertainties in early tumor screening data (such as the fuzziness of pathological features), thus generating more reasonable grouping results. This grouping aggregation not only reduces the data dimension, facilitating subsequent statistical calculations and privacy protection processing, but also reveals potential patterns and laws in the data (such as the association between specific gene mutations and tumor types), providing more accurate data support for early tumor screening research. At the same time, grouping aggregation helps to protect individual privacy when data is released, avoiding the reverse derivation of extremely low-frequency events.

[0120] It should be noted that the differential privacy budget parameter refers to the key value used to control the privacy protection intensity in differential privacy technology. Its size directly determines the degree of data perturbation and the privacy protection level. A smaller value indicates a larger added noise and a higher privacy protection intensity, but the data availability will decrease; a larger value indicates a smaller added noise, an increase in data availability, but a weakened privacy protection intensity. This parameter quantifies the privacy leakage risk to ensure that the contribution of individual data cannot be reverse-derived during data release or sharing, while maintaining the accuracy of statistical results as much as possible. The differential privacy budget parameter is the core tool for balancing privacy protection and data availability and can be applied to fields such as data release and machine learning.

[0121] It should be noted that this method realizes the balance between privacy protection and availability of statistical results through the organic connection of data preprocessing, feature extraction, statistical calculation, noise addition, and result verification.

[0122] In this embodiment, the method for constructing the early tumor screening data sharing platform may further include the following steps:

[0123] Deploy a traffic collection module in the early tumor screening data sharing platform to capture access traffic data in real time;

[0124] Perform multi-dimensional feature extraction on the collected access traffic data (such as access frequency, data volume, time distribution, user behavior patterns, etc.). If the deviation of the feature value from the data generated by the model exceeds the preset threshold, it is determined as a suspicious abnormal access behavior;

[0125] Use the request source IP address of the suspicious abnormal access behavior and the target data node as nodes in the graph structure, and the access behavior as an edge to construct an abnormal behavior graph;

[0126] Use a graph neural network to perform multiple rounds of iterative learning on the abnormal behavior graph. In each round, update the node representation by aggregating the feature information of neighbor nodes and calculate the association strength between nodes;

[0127] Construct a dense subgraph based on the updated association strength. If the association strength between an IP address node and its neighbor nodes exceeds a preset threshold, mark it as the core node of the dense subgraph;

[0128] Integrate the core node with its associated neighbor nodes, access behavior edges, and abnormal behavior characteristics (such as access frequency and time distribution) to construct an abnormal access verification topology graph for suspicious abnormal access behaviors;

[0129] Obtain the similarity between the abnormal access verification topology graph and a preset verification topology graph; if the similarity is greater than the preset similarity threshold, determine the suspicious abnormal access behavior as an abnormal access behavior;

[0130] Encrypt and store the detailed information of the abnormal access behavior in the distributed audit log, and notify the administrator in real time through the intelligent alarm module.

[0131] It should be noted that through real-time traffic collection and multi-dimensional feature extraction, early detection and accurate determination of abnormal access behaviors are ensured. Using graph neural networks to construct abnormal behavior graphs and identify dense subgraphs can efficiently capture abnormal patterns in complex access behaviors; through the similarity comparison of the abnormal access verification topology graph, the accuracy and reliability of abnormal behavior determination are further improved. Encrypting and storing abnormal behavior information and giving real-time alarms ensure the traceability and security controllability of the data transfer path; overall, the present invention provides an efficient and accurate abnormal access behavior recognition and response mechanism for the early tumor screening data sharing platform, effectively reducing the risks of data leakage and abuse.

[0132] Such as Figure 3 As shown, the second aspect of the present invention discloses a tumor early screening data sharing platform construction system 6 based on cloud computing. The tumor early screening data sharing platform construction system includes a memory 41 and a processor 52. A tumor early screening data sharing platform construction method program is stored in the memory 41. When the tumor early screening data sharing platform construction method program is executed by the processor 52, the steps of any one of the tumor early screening data sharing platform construction methods are implemented.

[0133] The third aspect of the present invention discloses a computer-readable storage medium. The computer-readable storage medium includes a tumor early screening data sharing platform construction method program. When the tumor early screening data sharing platform construction method program is executed by a processor, the steps of any one of the tumor early screening data sharing platform construction methods are implemented.

[0134] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0135] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0136] In addition, each functional unit in each embodiment of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0137] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks or optical discs and other various media that can store program codes.

[0138] Alternatively, if the above-mentioned integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of the various embodiments of the present invention. And the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks or optical discs and other various media that can store program codes.

[0139] The above are only the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. A method for constructing a tumor early screening data sharing platform based on cloud computing, characterized in that, Including: Obtain target multi-dimensional early screening data, use knowledge graph technology to conduct association modeling on the multi-dimensional early screening data, and generate a metadata index library; Encrypt sensitive fields in the metadata index library, and synchronously generate irreversible anonymous identifiers to replace the original identifiers; Arrange the metadata index library in the corresponding distributed storage system of the cloud platform according to tumor types, and use erasure code technology for cross-regional redundant backup; Build an access control smart contract based on blockchain technology, verify the qualifications of the requester through the digital identity authentication module, and after verifying the validity of the requester's digital certificate, dynamically decrypt and transmit data copies that meet the permission level; It also includes: Real-time collect the access traffic data of the tumor early screening data sharing platform; Extract multi-dimensional features from the collected access traffic data. If the deviation between the feature value and the data generated by the model exceeds the preset threshold, it is determined as a suspicious abnormal access behavior; Use the request source IP address of the suspicious abnormal access behavior and the target data node as nodes in the graph structure, and the access behavior as the edge to construct an abnormal behavior graph; Use graph neural network to perform multiple rounds of iterative learning on the abnormal behavior graph. In each round, update the node representation by aggregating the feature information of neighbor nodes, and calculate the association strength between nodes; Construct a dense subgraph according to the updated association strength. If the association strength between a certain IP address node and its neighbor nodes exceeds the preset threshold, mark it as the core node of the dense subgraph; Integrate the core node, its associated neighbor nodes, access behavior edges, and abnormal behavior characteristics to construct an abnormal access verification topology graph for the suspicious abnormal access behavior; Obtain the similarity between the abnormal access verification topology graph and the preset verification topology graph; if the similarity is greater than the preset similarity threshold, determine the suspicious abnormal access behavior as an abnormal access behavior; Encrypt and store the detailed information of the abnormal access behavior in the distributed audit log, and notify the administrator in real time through the intelligent alarm module.

2. A method for constructing a tumor early screening data sharing platform based on cloud computing according to claim 1, characterized in that, Use knowledge graph technology to conduct association modeling on multi-dimensional early screening data and generate a metadata index library. Specifically: Conduct feature analysis and processing on each multi-dimensional early screening data to obtain the key entities of each multi-dimensional early screening data; Through graph neural network iterative learning of the semantic association strength relationship between key entities, construct an initial knowledge graph according to the semantic association strength relationship; Based on the preset tumor early screening domain rule library, conduct logical verification on the management strength relationship between key entities in the initial knowledge graph, and identify and delete redundant relationships that do not conform to domain knowledge; Calculate the cosine similarity between each key entity in the initial knowledge graph; if the cosine similarity between two certain key entities is greater than the preset threshold, perform a pruning operation on any one of them; Use the rule inference engine to verify the missing relationships in the pruned knowledge graph, and supplement the missing key entities through entity attribute matching and semantic similarity calculation to generate an early screening data knowledge graph; Initialize several data indexes, generate a unique semantic identifier for each data index, and map each data index to the key entities in the early screening data knowledge graph to obtain a metadata index library.

3. A method for constructing a tumor early screening data sharing platform based on cloud computing according to claim 1, characterized in that, Encrypt the sensitive fields of the metadata index library, and synchronously generate irreversible anonymous identifiers to replace the original identifiers. Specifically: According to the preset sensitive information recognition rule library, match the fields in the metadata index library item by item. If the field name or content conforms to the sensitive information characteristics, mark it as a sensitive field; Perform format verification on the marked sensitive fields. If the field format conforms to the preset specification, further extract its key feature values; Use the hash algorithm to perform irreversible encryption processing on the extracted key feature values to generate a unique anonymous identifier; Map the generated anonymous identifier to the original identifier and store it in the encryption mapping table. At the same time, delete the original identifier field; If the field format does not conform to the preset specification, use the homomorphic encryption algorithm to encrypt it to ensure that calculations can still be performed on the encrypted data; Control the access rights to the encryption mapping table, only authorize specific modules to perform decryption operations when meeting the preset conditions, and record all access logs for auditing.

4. A method for constructing a tumor early screening data sharing platform based on cloud computing according to claim 3, characterized in that Use the hash algorithm to perform irreversible encryption processing on the extracted key feature values to generate a unique anonymous identifier. Specifically: Preprocess the extracted key feature values. If the feature value contains special characters or spaces, remove the invalid characters and uniformly convert it to the standardized string format; According to the preset hash algorithm configuration parameters, perform a hash operation on the standardized string to generate a hash value with a fixed length; Perform uniqueness verification on the generated hash value. If the hash value already exists in the anonymous identifier library, use the salt addition technique to add a random salt value to the original string and then perform the hash operation again until a unique hash value is generated; Map the generated unique hash value to the original feature value and store it in the encryption mapping table. At the same time, generate the corresponding anonymous identifier; If the original feature value is empty or invalid, skip the hash processing and mark it as an invalid record; Perform format standardization processing on the generated anonymous identifier to ensure that it conforms to the system identifier specification, and write it into the metadata index library to replace the original identifier field.

5. A method for constructing a tumor early screening data sharing platform based on cloud computing according to claim 1, characterized in that Arrange the metadata index library in the corresponding distributed storage system of the cloud platform according to the tumor type partition, and use the erasure code technology for cross-regional redundant backup. Specifically: Based on the preset tumor type classification rule library, perform type matching on the data entities in the metadata index library. If the data contains specific tumor markers or pathological characteristics, automatically assign it to the corresponding tumor type partition; Dynamically adjust the storage location of the data partition according to the real-time load status of the distributed storage nodes. If it is detected that the storage pressure of the target load node exceeds the threshold, migrate the data to the load node whose storage pressure does not exceed the threshold; Perform sharding processing on the data blocks of each tumor type partition. Use the erasure code algorithm to encode the data shards and parity blocks according to the preset ratio. If the sharded data volume does not meet the encoding requirements, supplement virtual padding data to complete the sharding; According to the cross-regional backup strategy, store the data shards and parity blocks in the storage nodes in different geographical regions respectively, ensuring that the data shards and parity blocks of the same partition are completely isolated geographically.

6. A method for constructing a tumor early screening data sharing platform based on cloud computing according to claim 1, characterized in that, Based on blockchain technology, an access control smart contract is built to verify the qualifications of the requester through a digital identity authentication module. After verifying the validity of the requester's digital certificate, a copy of the data that meets the permission level is dynamically decrypted and transmitted. Specifically: Based on the preset access rights policy template, create a smart contract instance in the blockchain network, define data access rights and decryption rules, and deploy the contract to distributed nodes; When the requester initiates a data access request, the smart contract automatically triggers the digital identity authentication module. If the digital certificate provided by the requester fails to pass the consensus verification of the blockchain node, the access process is immediately terminated and the abnormal log is recorded; For the verified requester, the permission label embedded in its certificate is matched with the corresponding permission level in the metadata index library. If the permission label does not match the minimum access level of the target data, a prompt indicating insufficient permissions is returned. For data requests that meet the permissions, the homomorphic encryption module is called to generate a dynamic decryption key. If the data needs to be shared across institutions, the target data shard is jointly decrypted through a multi-party secure computing protocol; The decrypted data copy is desensitized again according to the scope of authority and pushed to the requester through the encrypted transmission channel in the blockchain network. At the same time, the complete operation record is stored on the chain.

7. A method for constructing a tumor early screening data sharing platform based on cloud computing according to claim 1, characterized in that: The multidimensional early screening data includes medical images, pathological sections, gene sequencing data and electronic medical record texts.

8. A system for constructing a tumor early screening data sharing platform based on cloud computing, characterized in that, The system for building a tumor early screening data sharing platform includes a memory and a processor. The memory stores a tumor early screening data sharing platform building method program. When the tumor early screening data sharing platform building method program is executed by the processor, the steps of the tumor early screening data sharing platform building method as described in any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a method program for constructing a tumor early screening data sharing platform. When the method program for constructing a tumor early screening data sharing platform is executed by a processor, the steps of the method for constructing a tumor early screening data sharing platform as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Fine-grained security data sharing method for patient health record privacy protection

    CN116663047A

  • Application fusion system oriented to big data analysis

    CN117331995A

  • Medical data security sharing method and system based on block chain

    CN119357995A

  • Method and system for safely and rapidly storing image data of imaging department

    CN119517325A