Tumor early screening data sharing platform construction method and system based on cloud computing

By applying cloud computing, knowledge graph, encryption technology, distributed storage and blockchain technology on the tumor early screening data sharing platform, the problems of data silos, low sharing efficiency and insufficient privacy protection in the traditional data management model are solved, and efficient and secure multi-source heterogeneous data sharing and utilization are achieved.

CN119920488AActive Publication Date: 2025-05-02SHENZHEN RAPHA BIOTECHNOLOGY CO LTD

Patent Information

Application Number
CN202510403230.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-05-02
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The traditional tumor early screening data management model has problems such as data silos, low sharing efficiency and insufficient privacy protection, which is difficult to meet the needs of modern medical care for efficient utilization and secure sharing of large-scale, multi-source heterogeneous data.

Method used

The construction method of tumor early screening data sharing platform based on cloud computing is adopted. By obtaining multi-dimensional early screening data, using knowledge graph technology for correlation modeling, and generating metadata index database; encrypting sensitive fields to generate irreversible anonymous identifiers; data is partitioned by tumor type in a distributed storage system, and erasure coding technology is used for cross-region redundant backup; access control smart contracts are built based on blockchain technology, dynamically decrypting and transmitting data copies that meet the permission level.

Benefits of technology

The semantics and structure of data are realized, data privacy and security are ensured, data recovery capabilities and availability are improved, and through a secure data sharing mechanism, it supports the precise and intelligent development of early tumor screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920488A_ABST
    Figure CN119920488A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data, in particular to a cloud computing-based tumor early screening data sharing platform construction method and system. The sharing platform construction method comprises the following steps: carrying out association modeling on multi-dimensional early screening data by adopting a knowledge graph technology; encrypting the sensitive field of the metadata index database, and synchronously generating an irreversible anonymous identifier to replace the original identifier; partitioning and arranging the metadata index database in a distributed storage system corresponding to a cloud platform according to tumor types, and performing cross-region redundant backup by adopting an erasure code technology; and constructing an access control smart contract based on a block chain technology, verifying the qualification of a requester through a digital identity authentication module, and after verifying the validity of a digital certificate of the requester, dynamically decrypting and transmitting a data copy conforming to an authority level. Through organic fusion of the knowledge graph, the encryption technology, distributed storage and the block chain, a safe, efficient and reliable tumor early screening data sharing platform is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and in particular to a method and system for constructing a tumor early screening data sharing platform based on cloud computing. Background Art

[0002] In recent years, with the continuous increase in the incidence of cancer worldwide, the importance of early cancer screening in disease prevention and control has become increasingly prominent. The traditional early cancer screening data management model has problems such as data silos, low sharing efficiency, and insufficient privacy protection, which makes it difficult to meet the needs of modern medicine for efficient use and secure sharing of large-scale, multi-source heterogeneous data. In this context, cloud computing technology, with its advantages of elastic expansion, high availability, and on-demand services, provides a new technical architecture for early cancer screening data sharing. Therefore, building a cloud computing-based early cancer screening data sharing platform can not only break data silos and realize efficient integration and sharing of multi-source heterogeneous data, but also ensure the privacy security and compliance of data during the sharing process through advanced encryption and access control mechanisms, providing strong support for the precise and intelligent development of early cancer screening. Summary of the invention

[0003] The present invention overcomes the deficiencies of the prior art and provides a method and system for constructing a tumor early screening data sharing platform based on cloud computing.

[0004] To achieve the above-mentioned purpose, the technical solution adopted by the present invention is: The first aspect of the present invention discloses a method for constructing a tumor early screening data sharing platform based on cloud computing, comprising: Obtain target multi-dimensional early screening data, use knowledge graph technology to perform association modeling on the multi-dimensional early screening data, and generate a metadata index library; Encrypt sensitive fields in the metadata index library and simultaneously generate irreversible anonymous identifiers to replace the original identifiers; The metadata index library is partitioned by tumor type in the distributed storage system corresponding to the cloud platform, and erasure coding technology is used for cross-regional redundant backup; An access control smart contract is built based on blockchain technology. The qualifications of the requester are verified through a digital identity authentication module. After verifying the validity of the requester's digital certificate, a copy of the data that meets the permission level is dynamically decrypted and transmitted.

[0005] Preferably, knowledge graph technology is used to perform association modeling on multi-dimensional early screening data to generate a metadata index library, specifically: Perform feature analysis on each multi-dimensional early screening data to obtain key entities of each multi-dimensional early screening data; Iteratively learn the semantic association strength relationship between key entities through a graph neural network, and construct an initial knowledge graph based on the semantic association strength relationship; Based on the preset tumor early screening domain rule base, the management intensity relationship between key entities in the initial knowledge graph is logically verified to identify and delete redundant relationships that do not conform to domain knowledge; Calculate the cosine similarity between key entities in the initial knowledge graph; if the cosine similarity between two key entities is greater than the preset threshold, prune any of the key entities; Use the rule inference engine to verify missing relationships in the pruned knowledge graph, supplement missing key entities through entity attribute matching and semantic similarity calculation, and generate an early screening data knowledge graph; Initialize several data indexes, generate a unique semantic identifier for each data index, and map each data index with the key entities in the early screening data knowledge graph to obtain a metadata index library.

[0006] Preferably, the sensitive fields of the metadata index library are encrypted, and an irreversible anonymous identifier is simultaneously generated to replace the original identifier, specifically: According to the preset sensitive information identification rule library, the fields in the metadata index library are matched item by item. If the field name or content meets the sensitive information characteristics, it will be marked as a sensitive field; Perform format verification on the marked sensitive fields. If the field format meets the preset specifications, further extract its key feature values. The extracted key feature values ​​are irreversibly encrypted using a hash algorithm to generate a unique anonymous identifier; Map the generated anonymous identifier with the original identifier and store them in the encrypted mapping table, while deleting the original identifier field; If the field format does not meet the preset specifications, the homomorphic encryption algorithm is used to encrypt it to ensure that the data can still be calculated in an encrypted state; Access rights are controlled for the encrypted mapping table, and only specific modules are authorized to perform decryption operations when preset conditions are met. All access logs are recorded for auditing.

[0007] Preferably, a hash algorithm is used to perform irreversible encryption processing on the extracted key feature value to generate a unique anonymous identifier, specifically: Preprocess the extracted key feature values. If the feature values ​​contain special characters or spaces, remove the invalid characters and convert them into a standardized string format. According to the preset hash algorithm configuration parameters, the standardized string is hashed to generate a hash value of fixed length; Perform uniqueness check on the generated hash value. If the hash value already exists in the anonymous identifier library, use salting technology to add a random salt value to the original string and re-hash it until a unique hash value is generated. Map the generated unique hash value to the original feature value and store them in the encrypted mapping table, and generate the corresponding anonymous identifier at the same time; If the original feature value is empty or invalid, the hashing process is skipped and the record is marked as invalid; The format of the generated anonymous identifier is standardized to ensure that it complies with the system identifier specification, and it is written into the metadata index library to replace the original identification field.

[0008] Preferably, the metadata index library is partitioned by tumor type in a distributed storage system corresponding to the cloud platform, and erasure coding technology is used for cross-regional redundant backup, specifically: Based on the preset tumor type classification rule library, the data entities in the metadata index library are matched by type. If the data contains specific tumor markers or pathological features, it is automatically assigned to the corresponding tumor type partition; According to the real-time load status of the distributed storage nodes, the storage location of the data partition is dynamically adjusted. If it is detected that the storage pressure of the target load node exceeds the threshold, the data is migrated to the load node whose storage pressure does not exceed the threshold; The data blocks of each tumor type partition are fragmented, and the erasure coding algorithm is used to encode the data fragments and the check blocks according to the preset ratio. If the amount of fragmented data does not meet the coding requirements, virtual padding data is added to the complete fragment; According to the cross-region backup strategy, data shards and check blocks are stored in storage nodes in different geographical areas respectively, ensuring that data shards and check blocks of the same partition are completely isolated in terms of geographical location.

[0009] Preferably, an access control smart contract is constructed based on blockchain technology, and the qualification of the requester is verified through a digital identity authentication module. After verifying the validity of the digital certificate of the requester, a copy of the data that meets the permission level is dynamically decrypted and transmitted, specifically: Based on the preset access rights policy template, create a smart contract instance in the blockchain network, define data access rights and decryption rules, and deploy the contract to distributed nodes; When the requester initiates a data access request, the smart contract automatically triggers the digital identity authentication module. If the digital certificate provided by the requester fails to pass the consensus verification of the blockchain node, the access process is immediately terminated and the abnormal log is recorded; For the verified requester, the permission label embedded in its certificate is matched with the corresponding permission level in the metadata index library. If the permission label does not match the minimum access level of the target data, a prompt indicating insufficient permissions is returned. For data requests that meet the permissions, the homomorphic encryption module is called to generate a dynamic decryption key. If the data needs to be shared across institutions, the target data shard is jointly decrypted through a multi-party secure computing protocol; The decrypted data copy is desensitized again according to the scope of authority and pushed to the requester through the encrypted transmission channel in the blockchain network. At the same time, the complete operation record is stored on the chain.

[0010] Among them, the multidimensional early screening data includes medical images, pathological sections, gene sequencing data and electronic medical record texts.

[0011] The second aspect of the present invention discloses a cloud computing-based tumor early screening data sharing platform construction system, the tumor early screening data sharing platform construction system includes a memory and a processor, the memory stores a tumor early screening data sharing platform construction method program, when the tumor early screening data sharing platform construction method program is executed by the processor, any step of the tumor early screening data sharing platform construction method is implemented.

[0012] The third aspect of the present invention discloses a computer-readable storage medium, which includes a method program for constructing a tumor early screening data sharing platform. When the method program for constructing a tumor early screening data sharing platform is executed by a processor, the steps of any one of the methods for constructing a tumor early screening data sharing platform are implemented.

[0013] The present invention solves the technical defects existing in the background technology, and has the following beneficial effects: first, the target multi-dimensional early screening data is obtained, and the data is associated and modeled using knowledge graph technology to generate a metadata index library to realize the semanticization and structuring of the data; secondly, the sensitive fields in the metadata index library are encrypted, and an irreversible anonymous identifier is generated to replace the original identifier to ensure data privacy and security; then, the metadata index library is partitioned according to the tumor type in the distributed storage system of the cloud platform, and the erasure code technology is used for cross-regional redundant backup to improve the disaster tolerance and availability of the data; finally, an access control smart contract is constructed based on blockchain technology, and the qualification of the requester is verified through a digital identity authentication module. After the verification is passed, a copy of the data that meets the permission level is dynamically decrypted and transmitted to ensure the security and compliance of data sharing. The present invention constructs a safe, efficient and reliable tumor early screening data sharing platform through the organic integration of knowledge graph, encryption technology, distributed storage and blockchain, providing technical support for the sharing and application of tumor early screening data. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, drawings of other embodiments can be obtained based on these drawings without paying creative work.

[0015] Figure 1 A flowchart of the overall method for building a cloud computing-based tumor early screening data sharing platform; Figure 2 A partial flowchart of a method for building a cloud computing-based tumor early screening data sharing platform; Figure 3 A system block diagram for building a cloud computing-based cancer early screening data sharing platform. DETAILED DESCRIPTION

[0016] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0017] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0018] like Figure 1 As shown, the first aspect of the present invention discloses a method for constructing a tumor early screening data sharing platform based on cloud computing, comprising: S102, obtaining target multi-dimensional early screening data, using knowledge graph technology to perform association modeling on the multi-dimensional early screening data, and generating a metadata index library; S104, encrypting the sensitive fields of the metadata index library, and synchronously generating an irreversible anonymous identifier to replace the original identifier; S106, arranging the metadata index library in a distributed storage system corresponding to the cloud platform according to the tumor type partition, and using erasure coding technology for cross-regional redundant backup; S108. Build an access control smart contract based on blockchain technology, verify the qualifications of the requester through the digital identity authentication module, and after verifying the validity of the requester's digital certificate, dynamically decrypt and transmit a copy of the data that meets the permission level.

[0019] Among them, the multidimensional early screening data includes medical images, pathological sections, gene sequencing data and electronic medical record texts.

[0020] It should be noted that, first, the target multi-dimensional early screening data is obtained, and the knowledge graph technology is used to perform association modeling on the data, generate a metadata index library, and realize the semanticization and structuring of the data; secondly, the sensitive fields in the metadata index library are encrypted, and an irreversible anonymous identifier is generated to replace the original identifier to ensure data privacy and security; then, the metadata index library is partitioned according to the tumor type in the distributed storage system of the cloud platform, and the erasure code technology is used for cross-regional redundant backup to improve the disaster tolerance and availability of the data; finally, an access control smart contract is constructed based on blockchain technology, and the qualification of the requester is verified through the digital identity authentication module. After the verification is passed, the data copy that meets the permission level is dynamically decrypted and transmitted to ensure the security and compliance of data sharing. The present invention constructs a safe, efficient and reliable tumor early screening data sharing platform through the organic integration of knowledge graph, encryption technology, distributed storage and blockchain, which provides technical support for the sharing and application of tumor early screening data.

[0021] Preferably, knowledge graph technology is used to perform association modeling on multi-dimensional early screening data to generate a metadata index library, specifically: Perform feature analysis on each multi-dimensional early screening data to obtain key entities of each multi-dimensional early screening data; Among them, the key entities include but are not limited to patient name, key gene sequence features, key pathological text description features, and lesion area imaging features.

[0022] It should be noted that the format of the multidimensional early screening data is checked. If the data is unstructured text (such as a pathology report), natural language processing technology is used for entity recognition to extract key information such as patient name and pathological characteristics. If the data is structured data (such as a gene sequence), the feature fields are directly extracted. Secondly, for medical imaging data, lesion area detection is performed through deep learning models (such as convolutional neural networks) to extract imaging features.

[0023] Iteratively learn the semantic association strength relationship between key entities through a graph neural network, and construct an initial knowledge graph based on the semantic association strength relationship; It should be noted that the key entities and their initial relationships are represented as a graph structure, in which the entities are nodes, the initial relationships are edges, and an initial weight is assigned to each edge; secondly, a graph neural network model (such as GraphSAGE or GAT) is used to perform multiple rounds of iterative learning on the graph structure. In each round of iteration, the entity representation is updated by aggregating the feature information of neighboring nodes, and the semantic association strength between entities is calculated; then, the weight of the edge is adjusted according to the updated semantic association strength, and if the weight is lower than the preset threshold, the edge is removed; then, the graph neural network parameters are optimized through the back-propagation algorithm to ensure that the model can more accurately capture the semantic relationship between entities; finally, the optimized graph structure is output as an initial knowledge graph, in which the nodes represent the key entities and the edges represent the semantic association strength relationship, providing a basis for subsequent knowledge graph optimization.

[0024] Based on the preset tumor early screening domain rule base, the management intensity relationship between key entities in the initial knowledge graph is logically verified to identify and delete redundant relationships that do not conform to domain knowledge; Among them, the rule base in the field of early tumor screening is a set of rules built by relevant technical personnel specifically for the field of early tumor screening, which contains logical constraints on the relationship between entities, semantic association rules, and standardized definitions of domain knowledge. These rules are based on medical literature, clinical guidelines, and expert experience, and are used to describe the reasonable associations and constraints between key entities (such as gene mutations, pathological features, imaging features, etc.) in early tumor screening data. For example, the rule base contains logical rules such as "specific gene mutations are strongly associated with specific pathological features" or "benign tumors should not be directly associated with malignant gene mutations." Through the rule base, the entity relationships in the knowledge graph can be logically verified to ensure that they conform to the professional knowledge in the medical field, thereby improving the accuracy and reliability of the knowledge graph.

[0025] It should be noted that the rule base for the field of early cancer screening is loaded, which contains logical constraints on the relationship between entities (such as the association rules between "gene mutation type" and "pathological characteristics"); each edge in the initial knowledge graph is traversed, and if the semantic relationship of the edge conflicts with the logical constraints in the rule base (such as the unreasonable association between "benign tumors" and "malignant gene mutations"), it is marked as a relationship to be deleted; then, the marked redundant relationship is verified twice, and if it is confirmed that it does not conform to the domain knowledge, it is removed from the knowledge graph; then, the connectivity of the knowledge graph after removing the redundant relationship is checked, and if isolated nodes are found, the necessary relationship edges are supplemented according to the rule base; finally, the optimized knowledge graph is output to ensure that it conforms to the professional knowledge logic in the field of early cancer screening.

[0026] Calculate the cosine similarity between key entities in the initial knowledge graph; if the cosine similarity between two key entities is greater than the preset threshold, prune any of the key entities; It should be noted that the cosine similarity between key entities in the initial knowledge graph is calculated using the cosine similarity algorithm.

[0027] Use the rule inference engine to verify missing relationships in the pruned knowledge graph, supplement missing key entities through entity attribute matching and semantic similarity calculation, and generate an early screening data knowledge graph; It should be noted that based on the rule base in the field of early cancer screening, the inference rules for missing relationships between entities are defined (such as "specific gene mutations may lead to specific pathological characteristics"); the pruned knowledge graph is traversed, and if it is found that there is a potential relationship between the entities but they are not directly connected, the rule inference engine is triggered to predict the relationship; then, the attribute similarity between the entities is calculated through the entity attribute matching algorithm. If the similarity is higher than the preset threshold, the missing relationship edges are supplemented; then, the semantic similarity calculation model (such as BERT-based semantic embedding) is used to further verify the rationality of the supplemented relationship. If the semantic similarity conforms to the domain knowledge logic, the relationship is formally added to the knowledge graph; finally, the supplemented early screening data knowledge graph is output to ensure its completeness and accuracy.

[0028] Initialize several data indexes, generate a unique semantic identifier for each data index, and map each data index with the key entities in the early screening data knowledge graph to obtain a metadata index library.

[0029] It should be noted that based on the key entities in the early screening data knowledge graph, the core features of each data (such as patient ID, gene sequence, pathological characteristics, etc.) are extracted as the initialization content of the data index; the core features are irreversibly encrypted using a hash algorithm to generate a unique semantic identifier; the generated semantic identifier is bound to the corresponding data record and stored in the index mapping table; according to the semantic relationship of the key entities in the knowledge graph, the data index is mapped to the entity node, and a two-way association between the index and the entity is established; the mapped index information is integrated into a metadata index library to ensure that each piece of data can be quickly retrieved through the semantic identifier and accurately associated with the entity in the knowledge graph.

[0030] In summary, the present invention realizes the semantic association of multidimensional data through knowledge graph technology, solves the standardization problem of multi-source heterogeneous data; based on the collaborative optimization of graph neural network and rule reasoning, it ensures the accuracy and completeness of the knowledge graph, and provides reliable knowledge support for subsequent data mining and analysis; by generating unique semantic identifiers and metadata index libraries, it realizes efficient retrieval and cross-institutional sharing of data, and improves data utilization; at the same time, through logical verification and redundant relationship pruning, it reduces the impact of data noise on the analysis results, providing a high-quality data foundation for early cancer screening research.

[0031] Preferably, the sensitive fields of the metadata index library are encrypted, and an irreversible anonymous identifier is simultaneously generated to replace the original identifier, specifically: According to the preset sensitive information identification rule library, the fields in the metadata index library are matched item by item. If the field name or content meets the sensitive information characteristics (such as patient ID, gene locus, contact information, etc.), it will be marked as a sensitive field; Among them, the sensitive information identification rule base refers to a set of rules specifically used to identify and mark sensitive data, which contains the feature definition, format specification and matching rules of sensitive fields. It is used to accurately identify sensitive information (such as patient ID, gene loci, contact information, etc.) in the metadata index library. For example, the rule base defines rules such as "the ID number consists of 18 digits" or "the gene loci must comply with specific naming specifications." Through the rule base, the system can automatically detect and mark sensitive fields, provide a basis for subsequent encryption and anonymization processing, and ensure that data privacy protection meets compliance requirements.

[0032] Perform format verification on the marked sensitive fields. If the field format meets the preset specifications (such as ID card number, phone number, etc.), further extract its key feature values; The extracted key feature values ​​are irreversibly encrypted using a hash algorithm to generate a unique anonymous identifier; Map the generated anonymous identifier with the original identifier and store them in the encrypted mapping table, while deleting the original identifier field; If the field format does not meet the preset specifications, the homomorphic encryption algorithm is used to encrypt it to ensure that the data can still be calculated in an encrypted state; It should be noted that the fields that do not conform to the format specifications are cleaned, invalid characters are removed and the encoding format is unified to ensure that the data can be processed by the encryption algorithm; secondly, an applicable homomorphic encryption algorithm (such as Paillier or CKKS) is selected, and an encryption key pair is generated according to the preset key generation rules; then, the cleaned field data is processed in blocks. If the data length exceeds the maximum value supported by the algorithm, it is divided into multiple sub-blocks; then, each sub-block is encrypted using the public key to generate an encrypted data block, and the encrypted data blocks are sequentially spliced ​​into a complete encrypted field; finally, the encrypted field is stored in the metadata index library to ensure that it can still support calculation operations such as addition or multiplication in the encrypted state, and the decryption key is securely stored for use by the authorization module.

[0033] Access rights are controlled for the encrypted mapping table, and only specific modules are authorized to perform decryption operations when preset conditions are met. All access logs are recorded for auditing.

[0034] In summary, through sensitive information identification and encryption processing, we ensure that sensitive fields such as patient ID and gene forgery loci cannot be reversed during storage and transmission, reducing the risk of data leakage; irreversible anonymous identifiers are used to replace the original identifiers to protect privacy while maintaining data relevance; homomorphic encryption technology is used to support data calculation in an encrypted state, taking into account data security and availability; strict access control and audit log records further enhance the transparency and controllability of data access.

[0035] Preferably, a hash algorithm is used to perform irreversible encryption processing on the extracted key feature value to generate a unique anonymous identifier, specifically: Preprocess the extracted key feature values. If the feature values ​​contain special characters or spaces, remove the invalid characters and convert them into a standardized string format. According to the preset hash algorithm configuration parameters, the standardized string is hashed to generate a hash value of fixed length; Perform uniqueness check on the generated hash value. If the hash value already exists in the anonymous identifier library, use salting technology to add a random salt value to the original string and re-hash it until a unique hash value is generated. Among them, salting technology refers to adding a randomly generated string (called "salt value") to the original data (such as a string) during the hash operation to enhance the uniqueness and security of the hash result. Its core purpose is to prevent hash collisions (that is, different inputs generate the same hash value) and resist rainbow table attacks (reversely deriving the original data through pre-calculated hash values).

[0036] It should be noted that the generated hash value is compared with the existing hash value in the anonymous identifier library. If duplication is found, the salting technology processing is triggered; secondly, a random salt value is generated and appended to the end of the original string to ensure that the length and randomness of the salt value meet the preset security requirements; then, the salted string is re-hashed using the same hash algorithm to generate a new hash value; then, the new hash value is compared with the anonymous identifier library again. If duplication still exists, the above salting and hashing process is repeated until a unique hash value is generated; finally, the unique hash value is stored in the anonymous identifier library, and the corresponding mapping relationship between the original string and the salt value is recorded to ensure the accuracy of subsequent retrieval and verification.

[0037] Map the generated unique hash value to the original feature value and store them in the encrypted mapping table, and generate the corresponding anonymous identifier at the same time; If the original feature value is empty or invalid, the hashing process is skipped and the record is marked as invalid; The format of the generated anonymous identifier is standardized to ensure that it complies with the system identifier specification, and it is written into the metadata index library to replace the original identification field.

[0038] In summary, an irreversible anonymous identifier is generated through a hash algorithm to ensure that the original feature value cannot be reversed, thereby reducing the risk of data leakage; salting technology is used to solve the hash conflict problem and ensure the uniqueness of the anonymous identifier; the mapping relationship between the hash value and the original feature value is stored in an encrypted mapping table to support efficient retrieval and verification when necessary.

[0039] Preferably, the metadata index library is partitioned by tumor type in the distributed storage system corresponding to the cloud platform, and erasure coding technology is used for cross-regional redundant backup, such as Figure 2 As shown, specifically: S202, based on a preset tumor type classification rule library, type matching is performed on data entities in the metadata index library, and if the data contains specific tumor markers or pathological features, it is automatically assigned to the corresponding tumor type partition; Among them, the tumor type classification rule base refers to a set of rules specifically used to identify and classify tumor types, which contains matching rules and classification standards for key indicators such as tumor markers, pathological characteristics, and gene mutations. These rules are based on medical literature, clinical guidelines, and expert consensus, and are used to accurately identify tumor types (such as lung cancer, breast cancer, gastric cancer, etc.) in the metadata index library. For example, the rule base defines rules such as "EGFR gene mutation is associated with lung cancer" or "HER2 protein overexpression is associated with breast cancer."

[0040] S204, dynamically adjusting the storage location of the data partition according to the real-time load status of the distributed storage node, and if it is detected that the storage pressure of the target load node exceeds the threshold, migrating the data to the load node whose storage pressure does not exceed the threshold; S206, sharding the data blocks of each tumor type partition, encoding the data shards and the check blocks according to a preset ratio (such as a 4+2 mode) using an erasure coding algorithm, and if the amount of shard data does not meet the coding requirements, supplementing virtual padding data to complete the shard; It should be noted that the data blocks of each tumor type partition are sharded according to a preset size (such as 1MB). If the data volume of the last shard is insufficient, virtual padding data (such as all zero bytes) is added to the complete shard. According to the preset erasure code mode, the data shards and check blocks are encoded proportionally, where 4 data shards generate 2 check blocks to ensure that the original data can be recovered when any 2 shards are lost. The encoded data shards and check blocks are marked and stored in distributed storage nodes respectively, and the mapping relationship between the shards and the check blocks is recorded. Then the virtual padding data is marked to ensure that the padding part can be accurately identified and removed during data recovery to restore the original data content.

[0041] S208. According to the cross-region backup strategy, the data shards and the check blocks are stored in storage nodes in different geographical areas respectively, so as to ensure that the data shards and the check blocks in the same partition are completely isolated in terms of geographical location.

[0042] It should be noted that according to the preset cross-regional backup strategy, storage nodes in multiple geographical areas are selected to ensure that the geographical distance between the nodes meets the isolation requirements (such as at least 500 kilometers apart). The data shards and check blocks of each tumor type partition are sequentially allocated to storage nodes in different regions. The storage location of each shard and check block is recorded to ensure that it is not in the same area as other shards or check blocks of the same partition. Then, the availability and data integrity of the storage nodes are checked regularly. If a node in a certain area is found to be unavailable, the data recovery mechanism is triggered to reconstruct the lost data using shards and check blocks in other areas. Finally, the storage location record is updated to ensure that the new data shards and check blocks continue to meet the geographical isolation requirements.

[0043] In summary, by zoning and arranging tumor types, data classification management and efficient retrieval can be achieved, which is convenient for subsequent analysis and application; the dynamic adjustment mechanism based on real-time load status ensures the rational allocation and efficient utilization of storage resources and avoids node overload; the erasure code technology is used to shard and encode data, which not only improves storage efficiency but also enhances data fault tolerance; the cross-regional backup strategy ensures geographical isolation and redundancy of data, effectively coping with the risk of data loss caused by natural disasters or hardware failures.

[0044] Preferably, an access control smart contract is constructed based on blockchain technology, and the qualification of the requester is verified through a digital identity authentication module. After verifying the validity of the digital certificate of the requester, a copy of the data that meets the permission level is dynamically decrypted and transmitted, specifically: Based on the preset access rights policy template, create a smart contract instance in the blockchain network, define data access rights and decryption rules, and deploy the contract to distributed nodes; Among them, the access permission policy template is a set of predefined rules used to regulate the scope of data access permissions and operating conditions. Based on multi-dimensional factors such as roles, institutions, and data sensitivity, the template defines the access levels of different users or organizations to data (such as read-only, read-write, delete, etc.) and access restrictions (such as time range, data category, etc.). For example, the template may stipulate that "research institutions can only access de-identified data" or "medical institutions can access encrypted complete data but cannot download it." Through the access permission policy template, the system can automatically match user permissions with data access requests to ensure the compliance and security of data sharing. Distributed nodes refer to computers or servers that run independently in a distributed system. Each node has storage, computing and communication capabilities, and can work with other nodes to complete common tasks.

[0045] When the requester initiates a data access request, the smart contract automatically triggers the digital identity authentication module. If the digital certificate provided by the requester fails to pass the consensus verification of the blockchain node (such as the certificate expires or the signature is invalid), the access process is immediately terminated and the abnormal log is recorded; For the verified requester, the permission label embedded in its certificate is matched with the corresponding permission level in the metadata index library. If the permission label does not match the minimum access level of the target data, a prompt indicating insufficient permissions is returned. For data requests that meet the permissions, the homomorphic encryption module is called to generate a dynamic decryption key. If the data needs to be shared across institutions, the target data shard is jointly decrypted through a multi-party secure computing protocol; The decrypted data copy is desensitized again according to the scope of authority and pushed to the requester through the encrypted transmission channel in the blockchain network. At the same time, the complete operation record is stored on the chain.

[0046] It should be noted that, according to the permission level of the requester, the homomorphic encryption module is called to generate a dynamic decryption key to ensure that the key is only valid for the current request and is time-effective. If the data needs to be shared across institutions, the multi-party secure computing protocol is triggered to distribute the encrypted data shards to each participating institution. Each institution uses the local private key to partially decrypt the shards and generate an intermediate decryption result; then, the intermediate decryption result is aggregated to the coordination node through a secure communication channel, and the coordination node completes the final decryption; finally, the decrypted data shards are secondary desensitized according to the permission scope, and pushed to the requester through an encrypted transmission channel, while recording the decryption operation log for auditing.

[0047] In summary, the present invention automatically performs identity authentication, permission matching and data decryption through smart contracts, thereby achieving full-process security and controllability of data sharing.

[0048] In this embodiment, the method for constructing a tumor early screening data sharing platform may further include the following steps: Based on the fuzzy clustering algorithm, the tumor early screening data to be released in the tumor early screening data sharing platform are grouped and aggregated, and the statistical indicators of the grouped data (such as mean, proportion, frequency distribution, etc.) are calculated. If the statistical results contain sensitive information (such as extremely low frequency events), they are marked as data that needs to be protected; According to the preset differential privacy budget parameters, the utility function value of the data to be protected under the exponential mechanism is calculated, and the perturbation probability of each possible output result is calculated based on the utility function value to generate the perturbation probability distribution of the data to be protected; Based on the perturbation probability distribution, the data to be protected is randomly sampled to generate the perturbation result; Calculate the degree of overlap between the disturbed result and the original data. If the degree of overlap is greater than a preset overlap threshold, recalculate the disturbance probability distribution and perform sampling again until the degree of overlap is no greater than the preset overlap threshold. The statistical results after disturbance are published, and the noise addition parameters and operation logs are recorded to ensure subsequent traceability and auditing.

[0049] It should be noted that the fuzzy clustering algorithm allows data points to belong to multiple categories with a certain probability, which can better handle the uncertainty in early cancer screening data (such as the ambiguity of pathological characteristics), thereby generating more reasonable grouping results. This grouping aggregation not only reduces the data dimension, facilitating subsequent statistical calculations and privacy protection processing, but also reveals potential patterns and laws in the data (such as the association between specific gene mutations and tumor types), providing more accurate data support for early cancer screening research. At the same time, grouping aggregation helps to protect individual privacy when data is released and prevent extremely low-frequency events from being reversed.

[0050] It should be noted that the differential privacy budget parameter refers to the key value used to control the strength of privacy protection in differential privacy technology, and its size directly determines the degree of data disturbance and the level of privacy protection. The smaller the value, the greater the added noise, the higher the privacy protection strength, but the data availability will be reduced; the larger the value, the smaller the added noise, the higher the data availability, but the privacy protection strength is weakened. This parameter quantifies the risk of privacy leakage to ensure that the contribution of individual data cannot be reversed during the data release or sharing process, while trying to maintain the accuracy of the statistical results. The differential privacy budget parameter is a core tool for balancing privacy protection and data availability, and can be applied to data release, machine learning and other fields.

[0051] It should be noted that this method achieves a balance between privacy protection and usability of statistical results through the organic connection of data preprocessing, feature extraction, statistical calculation, noise addition and result verification.

[0052] In this embodiment, the method for constructing a tumor early screening data sharing platform may further include the following steps: Deploy a traffic collection module in the tumor early screening data sharing platform to capture access traffic data in real time; Perform multi-dimensional feature extraction on the collected traffic data (such as access frequency, data volume, time distribution, user behavior pattern, etc.). If the deviation between the feature value and the model-generated data exceeds the preset threshold, it is determined to be a suspicious abnormal access behavior. The request source IP address and target data node of suspicious abnormal access behavior are used as nodes in the graph structure, and the access behavior is used as the edge to construct an abnormal behavior graph; The abnormal behavior graph is iterated through multiple rounds of learning using a graph neural network. In each round, the node representation is updated by aggregating the feature information of neighboring nodes and calculating the correlation strength between nodes. A dense subgraph is constructed based on the updated association strength. If the association strength between an IP address node and its neighboring nodes exceeds a preset threshold, it is marked as a core node of the dense subgraph. By integrating the core node with its associated neighbor nodes, access behavior edges, and abnormal behavior features (such as access frequency and time distribution), an abnormal access verification topology diagram of suspicious abnormal access behaviors is constructed; Obtaining the similarity between the abnormal access verification topology map and the preset verification topology map; if the similarity is greater than a preset similarity threshold, determining the suspicious abnormal access behavior as an abnormal access behavior; The detailed information of the abnormal access behavior is encrypted and stored in the distributed audit log, and the administrator is notified in real time through the intelligent alarm module.

[0053] It should be noted that through real-time traffic collection and multi-dimensional feature extraction, early detection and accurate judgment of abnormal access behaviors are ensured, and the graph neural network is used to construct abnormal behavior graphs and identify dense subgraphs, which can efficiently capture abnormal patterns in complex access behaviors; through the similarity comparison of abnormal access verification topology graphs, the accuracy and reliability of abnormal behavior judgments are further improved, and the abnormal behavior information is encrypted and stored and alarms are issued in real time, ensuring the traceability and security controllability of the data flow path; on the whole, the present invention provides an efficient and accurate abnormal access behavior identification and response mechanism for the tumor early screening data sharing platform, effectively reducing the risk of data leakage and abuse.

[0054] like Figure 3 The second aspect of the present invention disclosed is a cloud computing-based cancer early screening data sharing platform construction system 6, the cancer early screening data sharing platform construction system includes a memory 41 and a processor 52, the memory 41 stores a cancer early screening data sharing platform construction method program, when the cancer early screening data sharing platform construction method program is executed by the processor 52, any step of the cancer early screening data sharing platform construction method is implemented.

[0055] The third aspect of the present invention discloses a computer-readable storage medium, which includes a method program for constructing a tumor early screening data sharing platform. When the method program for constructing a tumor early screening data sharing platform is executed by a processor, the steps of any one of the methods for constructing a tumor early screening data sharing platform are implemented.

[0056] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0057] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0058] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0059] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0060] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods of each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0061] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for constructing a tumor early screening data sharing platform based on cloud computing, characterized in that: include: Obtain target multi-dimensional early screening data, use knowledge graph technology to perform association modeling on the multi-dimensional early screening data, and generate a metadata index library; Encrypt sensitive fields in the metadata index library and simultaneously generate irreversible anonymous identifiers to replace the original identifiers; The metadata index library is partitioned by tumor type in the distributed storage system corresponding to the cloud platform, and erasure coding technology is used for cross-regional redundant backup; An access control smart contract is built based on blockchain technology. The qualifications of the requester are verified through a digital identity authentication module. After verifying the validity of the requester's digital certificate, a copy of the data that meets the permission level is dynamically decrypted and transmitted.

2. The method for constructing a cloud computing-based tumor early screening data sharing platform according to claim 1, characterized in that: The knowledge graph technology is used to perform association modeling on multi-dimensional early screening data and generate a metadata index library, specifically: Perform feature analysis on each multi-dimensional early screening data to obtain key entities of each multi-dimensional early screening data; Iteratively learn the semantic association strength relationship between key entities through a graph neural network, and construct an initial knowledge graph based on the semantic association strength relationship; Based on the preset tumor early screening domain rule base, the management intensity relationship between key entities in the initial knowledge graph is logically verified to identify and delete redundant relationships that do not conform to domain knowledge; Calculate the cosine similarity between key entities in the initial knowledge graph; if the cosine similarity between two key entities is greater than the preset threshold, prune any of the key entities; Use the rule inference engine to verify missing relationships in the pruned knowledge graph, supplement missing key entities through entity attribute matching and semantic similarity calculation, and generate an early screening data knowledge graph; Initialize several data indexes, generate a unique semantic identifier for each data index, and map each data index with the key entities in the early screening data knowledge graph to obtain a metadata index library.

3. The method for constructing a cloud computing-based tumor early screening data sharing platform according to claim 1, characterized in that: The sensitive fields of the metadata index library are encrypted, and irreversible anonymous identifiers are generated to replace the original identifiers. Specifically: According to the preset sensitive information identification rule library, the fields in the metadata index library are matched item by item. If the field name or content meets the sensitive information characteristics, it will be marked as a sensitive field; Perform format verification on the marked sensitive fields. If the field format meets the preset specifications, further extract its key feature values. The extracted key feature values ​​are irreversibly encrypted using a hash algorithm to generate a unique anonymous identifier; Map the generated anonymous identifier with the original identifier and store them in the encrypted mapping table, while deleting the original identifier field; If the field format does not meet the preset specifications, the homomorphic encryption algorithm is used to encrypt it to ensure that the data can still be calculated in an encrypted state; Access rights are controlled for the encrypted mapping table, and only specific modules are authorized to perform decryption operations when preset conditions are met. All access logs are recorded for auditing.

4. The method for constructing a cloud computing-based tumor early screening data sharing platform according to claim 3, characterized in that: The extracted key feature values ​​are irreversibly encrypted using a hash algorithm to generate a unique anonymous identifier, specifically: Preprocess the extracted key feature values. If the feature values ​​contain special characters or spaces, remove the invalid characters and convert them into a standardized string format. According to the preset hash algorithm configuration parameters, the standardized string is hashed to generate a hash value of fixed length; Perform uniqueness check on the generated hash value. If the hash value already exists in the anonymous identifier library, use salting technology to add a random salt value to the original string and re-hash it until a unique hash value is generated. Map the generated unique hash value to the original feature value and store them in the encrypted mapping table, and generate the corresponding anonymous identifier at the same time; If the original feature value is empty or invalid, the hashing process is skipped and the record is marked as invalid; The format of the generated anonymous identifier is standardized to ensure that it complies with the system identifier specification, and it is written into the metadata index library to replace the original identification field.

5. The method for constructing a cloud computing-based tumor early screening data sharing platform according to claim 1, characterized in that: The metadata index library is partitioned by tumor type in the distributed storage system corresponding to the cloud platform, and erasure coding technology is used for cross-regional redundant backup. Specifically: Based on the preset tumor type classification rule library, the data entities in the metadata index library are matched by type. If the data contains specific tumor markers or pathological features, it is automatically assigned to the corresponding tumor type partition; According to the real-time load status of the distributed storage nodes, the storage location of the data partition is dynamically adjusted. If it is detected that the storage pressure of the target load node exceeds the threshold, the data is migrated to the load node whose storage pressure does not exceed the threshold; The data blocks of each tumor type partition are fragmented, and the erasure coding algorithm is used to encode the data fragments and the check blocks according to the preset ratio. If the amount of fragmented data does not meet the coding requirements, virtual padding data is added to the complete fragment; According to the cross-region backup strategy, data shards and check blocks are stored in storage nodes in different geographical areas respectively, ensuring that data shards and check blocks of the same partition are completely isolated in terms of geographical location.

6. The method for constructing a cloud computing-based tumor early screening data sharing platform according to claim 1, characterized in that: Based on blockchain technology, an access control smart contract is built to verify the qualifications of the requester through a digital identity authentication module. After verifying the validity of the requester's digital certificate, a copy of the data that meets the permission level is dynamically decrypted and transmitted. Specifically: Based on the preset access rights policy template, create a smart contract instance in the blockchain network, define data access rights and decryption rules, and deploy the contract to distributed nodes; When the requester initiates a data access request, the smart contract automatically triggers the digital identity authentication module. If the digital certificate provided by the requester fails to pass the consensus verification of the blockchain node, the access process is immediately terminated and the abnormal log is recorded; For the verified requester, the permission label embedded in its certificate is matched with the corresponding permission level in the metadata index library. If the permission label does not match the minimum access level of the target data, a prompt indicating insufficient permissions is returned. For data requests that meet the permissions, the homomorphic encryption module is called to generate a dynamic decryption key. If the data needs to be shared across institutions, the target data shard is jointly decrypted through a multi-party secure computing protocol; The decrypted data copy is desensitized again according to the scope of authority and pushed to the requester through the encrypted transmission channel in the blockchain network. At the same time, the complete operation record is stored on the chain.

7. The method for constructing a cloud computing-based tumor early screening data sharing platform according to claim 1, characterized in that: The multidimensional early screening data includes medical images, pathological sections, gene sequencing data and electronic medical record texts.

8. A system for building a tumor early screening data sharing platform based on cloud computing, characterized in that: The system for building a tumor early screening data sharing platform includes a memory and a processor. The memory stores a tumor early screening data sharing platform building method program. When the tumor early screening data sharing platform building method program is executed by the processor, the steps of the tumor early screening data sharing platform building method as described in any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a method program for constructing a tumor early screening data sharing platform. When the method program for constructing a tumor early screening data sharing platform is executed by a processor, the steps of the method for constructing a tumor early screening data sharing platform as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Block chain transaction topological graph analysis method and device based on graph neural network

    CN113657896A

  • Fine-grained security data sharing method for patient health record privacy protection

    CN116663047A

  • Application fusion system oriented to big data analysis

    CN117331995A

  • Data security tracing method, system and device based on artificial intelligence

    CN118536093A

  • Medical data security sharing method and system based on block chain

    CN119357995A

Cited By

  • Data archiving processing method and system based on shutdown system

    CN120104569A

  • Internet of vehicles data sharing method and system based on dynamic behavior map and block chain

    CN120224180A

  • Digital pathological section data management system based on cloud computing

    CN120319382A

  • Information management platform and method supporting multi-center data fusion

    CN120511076A

  • Mobile phone file backup method and system based on intelligent hardware

    CN120540905A