Real world research database management method and system based on block chain
Through blockchain technology, hospital data is type-divided and hashed to store evidence, establish a quantum Merkle tree, add time stamp anchoring, smart contracts parse zk-SNARK proof, and establish cross-chain hash index, solving the problems of data silos and collaboration bottlenecks in real-world research, and realizing data transparency, credibility and shareability.
Patent Information
- Application Number
- CN202510740711.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The traditional artificial governance model is difficult to cope with the cleaning, standardization and cross-domain integration of massive unstructured data in real-world research, resulting in data silos and collaboration bottlenecks and research repeatability crisis.
Blockchain technology is used to collect hospital data, perform type division and hash evidence storage, establish a quantum Merkle tree, add time stamp anchoring, and smart contracts automatically parse zk-SNARK proofs in the block, and establish a cross-chain hash index to form a research database.
Improve data transparency, credibility, security and shareability, and promote the development of real-world research.
Smart Images

Figure CN120277160A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of blockchain forensics, and particularly to a method and system for managing a real-world research database based on blockchain. Background Art
[0002] With the rapid increase in data scale and complexity, the real-world research market scale is constantly expanding, and the data types have also expanded from structured electronic medical records to unstructured data such as Internet of Things devices and radiomics. This makes it difficult for traditional manual governance models to handle the cleaning, standardization, and cross-domain integration of massive data. Currently, the technology is facing challenges such as the rapid increase in data scale and complexity, data islands and collaboration bottlenecks, and the crisis of research reproducibility.
[0003] Therefore, the present invention provides a method and system for managing a real-world research database based on blockchain. Summary of the Invention
[0004] A method and system for managing a real-world research database based on blockchain provided by the present invention collect hospital-related data through blockchain technology, perform type classification and hash forensics on the data, upload the data digest to the blockchain and perform sharding processing, establish a quantum-resistant Merkle tree, add a timestamp anchor, automatically parse the zk-SNARK proof in the block by a smart contract, and establish a cross-chain hash index, and finally form a research database. This is beneficial to improving the transparency, credibility, security, and shareability of data, thereby promoting the development of real-world research.
[0005] The present invention provides a method for managing a real-world research database based on blockchain, including: Step 1: Collect hospital-related data, perform type classification on the related data, and generate corresponding hash forensics according to the type classification result; Step 2: Generate a digest based on the hash forensics and upload it to the blockchain, perform sharding processing on the digest, effectively verify the sharding result, and write the verified valid sharding information into a new block; Step 3: Establish a quantum-resistant Merkle tree based on the new block, add a timestamp anchor, automatically parse the zk-SNARK proof in the block by a smart contract, and establish a cross-chain hash index; Step 4: Set a hierarchical data release rule according to the cross-chain hash index to form a research database.
[0006] The present invention provides a method for managing a real-world research database based on blockchain, which collects hospital-related data, performs type classification on the related data, and generates corresponding hash forensics according to the type classification result, including: Dock to the interfaces of each business system of the hospital, and perform adaptation judgment on the interface status. Select the corresponding adaptation tool from the format-adaptation table according to the adaptation judgment result, and collect hospital-related data based on the adaptation tool; Perform type division on the relevant data, perform standard conversion based on the type division result, obtain the conversion result, and generate a hash deposit certificate according to the conversion result.
[0007] The present invention provides a method for managing a real-world research database based on blockchain. Performing type division on the relevant data, performing standard conversion based on the type division result, obtaining the conversion result, and generating a hash deposit certificate according to the conversion result, including: Input the relevant data into the feature analysis engine, parse the relevant data types and add corresponding type tags to obtain the type division result; Match the structured data in the type division result with the real-time terminology of the knowledge graph, and then perform unit intelligent conversion to obtain the first conversion result; Perform multi-modal AI annotation on the unstructured data in the type division result and perform cross-modal alignment to convert it into a standard result, and then obtain the second conversion result; Serialize the first conversion result and the second conversion result, calculate the hash value of the serialized result using the cryptographic hash algorithm, and then generate a hash deposit certificate.
[0008] The present invention provides a method for managing a real-world research database based on blockchain. Generate a digest and upload it to the blockchain according to the hash deposit certificate, perform sharding processing on the digest, perform effective verification on the sharding result, and write the verified effective sharding information into a new block, including: Extract the key fields of the hash deposit certificate and convert them into a compact JSON format to obtain a JSON string, calculate a secondary hash for the JSON string, and generate a chain identifier; Generate a digest according to the chain identifier, group the digest, obtain a time series group and a verification group, slice the time series group according to a fixed length, pad with zeros if insufficient, encode the slices, and calculate the metadata of all slices according to the encoding; Determine the verification group nodes according to the VRF random assignment. The time series group broadcasts the metadata to the verification group nodes, and each verification group node verifies a specific slice according to the VRF assignment to generate a signature proof; The time series group receives the signature proofs of all verification group nodes, and statistically calculates the effective slice ratio of the signature proofs. Based on the effective slice ratio, package the corresponding effective slice information to generate transaction data, submit the transaction to the blockchain, and perform malicious node processing on the blockchain to form network consensus, and write the effective slice information into a new block.
[0009] The present invention provides a method for managing a real-world research database based on blockchain. The verification group nodes are determined by random allocation based on VRF. The timing group broadcasts the metadata to the verification group nodes, and each verification group node verifies a specific shard according to the VRF allocation and generates a signature proof, including: Set a dynamic pledge mechanism according to the hospital level, and determine a variable candidate node list by combining the dynamic pledge mechanism, the medical dedicated blockchain network and hospital characteristics; Perform VRF calculation on the candidate nodes in the variable candidate node list based on hash evidence storage, and the nodes whose last two digits of the VRF output value are less than the preset value become the verification group, and the remaining nodes become the timing group; The timing group packs the metadata of the current time window and sends it to all verification group nodes through the medical private network. Each verification group node calculates the shard index according to its own VRF output value, and determines the specific shard verification right of each verification group node according to the shard index; The verification group node performs medical data verification according to the specific shard verification right and generates a signature proof after passing the verification.
[0010] The present invention provides a method for managing a real-world research database based on blockchain. An anti-quantum Merkle tree is established based on a new block, a timestamp anchor is added, the smart contract automatically parses the zk-SNARK proof in the block, and a cross-chain hash index is established, including: Take the verified specific shards in the new block as leaf nodes, use the XMSS algorithm to construct an anti-quantum Merkle tree, generate a root hash, and perform a timestamp anchor on the root hash; Use the smart contract to parse the pre-stored zk-SNARK proof in the new block, perform a one-way hash match on the specific shard, and obtain the first matching result; Verify the consistency between the root hash and the block header according to the Merkle tree to obtain the second verification result; If both the first matching result and the second verification result pass, a cross-chain hash index is generated.
[0011] The present invention provides a method for managing a real-world research database based on blockchain. Set a hierarchical data release rule according to the cross-chain hash index to form a research database, including: Construct an index key value according to the cross-chain hash index, write the index key value into the index mapping table of the target chain, and preset hierarchical conditions and corresponding hierarchical data release rules in the smart contract to form a research database.
[0012] The present invention provides a real-world research database management system based on blockchain, including: Hash evidence storage module: Collect hospital-related data, classify the related data, and generate corresponding hash evidence storage according to the classification result; Sharding module: Generate a digest for on-chain storage according to the hash-based evidence storage, perform sharding processing on the digest, conduct effective verification on the sharding results, and write the verified valid sharding information into a new block; Cross-chain module: Build a quantum-resistant Merkle tree based on the new block, add a timestamp anchor, have the smart contract automatically parse the zk-SNARK proof within the block, and establish a cross-chain hash index; Database module: Set hierarchical data release rules according to the cross-chain hash index to form a research database.
[0013] Compared with the prior art, the beneficial effects of the present application are as follows: Collect hospital-related data through blockchain technology, perform type classification and hash-based evidence storage, upload the data digest to the chain and perform sharding processing, build a quantum-resistant Merkle tree, add a timestamp anchor, have the smart contract automatically parse the zk-SNARK proof within the block, and establish a cross-chain hash index, and finally form a research database. This is beneficial to improving the transparency, credibility, security, and shareability of data, thereby promoting the development of real-world research.
[0014] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification and the drawings.
[0015] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings
[0016] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of a method for managing a real-world research database based on blockchain provided by an embodiment of the present invention; Figure 2 is a structural diagram of a system for managing a real-world research database based on blockchain provided by an embodiment of the present invention; Figure 3 is a flowchart of the data sharing process of a medical-specific blockchain network provided by an embodiment of the present invention. Detailed Embodiments
[0017] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to explain and illustrate the present invention and are not used to limit the present invention.
[0018] Example 1: The embodiment of the present invention provides a method for managing a real-world research database based on blockchain, as Figure 1 shown, including: Step 1: Collect hospital-related data, classify the related data, and generate corresponding hash deposits according to the classification results; Step 2: Generate a digest and upload it to the blockchain according to the hash deposit, perform sharding processing on the digest, effectively verify the sharding results, and write the verified valid sharding information into a new block; Step 3: Build a quantum-resistant Merkle tree based on the new block, add a timestamp anchor, the smart contract automatically parses the zk-SNARK proof in the block, and establish a cross-chain hash index; Step 4: Set a hierarchical data release rule according to the cross-chain hash index to form a research database.
[0019] In this embodiment, the hospital-related data refers to the data collected from each business system of the hospital, including patients' medical records, drug inventory information, laboratory test results, etc. For example, patients' diagnosis records, medication history, and laboratory test results.
[0020] In this embodiment, the classification refers to the process of classifying the collected data according to different types. For example, the data is divided into text data, numerical data, time series data, etc.
[0021] In this embodiment, the classification result refers to the result after data classification, which describes the type of each piece of data. For example, a patient information record is classified as text data.
[0022] In this embodiment, the hash deposit refers to using a hash function to process the data to generate a fixed-length string for verifying the integrity and authenticity of the data. For example, use the SHA-256 hash function to generate a hash value for the converted data and store this value on the blockchain.
[0023] In this embodiment, , where represents the hash value calculated according to the Pareto optimum; represents the intermediate hash value of structured data in the i-th iteration; represents the intermediate hash value of unstructured data in the i-th iteration; K represents an additional key; i represents the iteration number index; p represents a large prime number; represents the random coefficient of structured data; m represents the vector dimension of structured data; represents the non-linear activation function; represents the weight matrix of structured data; represents the vectorization of structured data; represents vector concatenation; Denotes the dynamic salt value for the i-th iteration; Denotes the vectorization of unstructured data; Denotes the weight matrix of unstructured data Denotes the random coefficient of unstructured data; Denotes the rounding function.
[0024] In this embodiment, generating a digest means performing a digest process on data to extract the key information of the data. This usually involves using a hash function or other digest algorithms. For example, using the SHA-256 hash function to generate a fixed-length digest for a large data file.
[0025] In this embodiment, the process of digest chaining is to extract key fields: extract key information from the hash deposit, such as the hash value and timestamp of the data; convert to compact JSON format: convert this key information to compact JSON format to reduce the data size; generate a JSON string: convert the JSON-formatted data to a string for hash calculation; calculate the secondary hash: apply a hash function to the JSON string to generate a hash value, which is the on-chain identifier; chain up: store the on-chain identifier on the blockchain as a permanent record of the data digest.
[0026] In this embodiment, the process of sharding is to generate a digest: generate a data digest according to the on-chain identifier; group: divide the digest into a time series group and a verification group. The time series group is responsible for broadcasting the data shards to the verification group, and the verification group is responsible for verifying the shards; sharding: shard the data in the time series group according to a fixed length, and fill with zeros if the data is insufficient; encoding: encode each shard, for example, using Base64 encoding, for easy transmission in the network; calculate metadata: calculate the metadata of all shards according to the encoded shards, such as the size and hash value of each shard.
[0027] In this embodiment, the process of effective verification is as follows: VRF allocation: Use a verifiable random function (VRF) to randomly allocate verification group nodes and determine the permission of each node to verify specific shards; Broadcast metadata: The timing group broadcasts the shards and metadata to all verification group nodes; Verify shards: Each verification group node verifies the specific shards according to the result of VRF allocation and generates a signature proof; Collect signature proofs: The timing group collects the signature proofs of all verification group nodes; Statistically analyze the proportion of valid shards: Statistically analyze the proportion of valid shards in the signature proofs; Package transaction data: Package the corresponding valid shard information into transaction data according to the proportion of valid shards; Submit transactions: Submit the transaction data to the blockchain; Handle malicious nodes: Handle malicious nodes on the blockchain to ensure network security; Reach consensus: Through the consensus mechanism of the blockchain, ensure that all nodes reach an agreement on the valid shard information; Write to a new block: Write the valid shard information to a new block to complete the permanent storage of the data.
[0028] In this embodiment, the valid shard information refers to the data and information contained in the valid shards during the verification process. For example, a shard containing valid transaction records.
[0029] In this embodiment, the quantum-resistant Merkle tree refers to a Merkle tree constructed using a quantum-resistant signature algorithm. It can resist attacks from quantum computers and ensure the security of data in the era of quantum computing. For example, a Merkle tree constructed using the XMSS algorithm.
[0030] In this embodiment, the new block refers to the block that is added to the blockchain after the verified specific shard information is packaged into transaction data, confirmed through the consensus mechanism of the blockchain network, and finally added to the blockchain. This new block contains the hash value and metadata of these shard information, as well as other possible relevant information, such as timestamps and signatures of block creators.
[0031] In this embodiment, timestamp anchoring refers to associating the hash value or digest of data with time information, usually using a timestamp service to ensure the existence and status of data at a specific time point. For example, associating the root hash of a Merkle tree with a timestamp service to prove that the data has not been tampered with at a specific time point.
[0032] In this embodiment, the zk-SNARK proof is a zero-knowledge proof that allows a prover to prove that a certain statement is true without revealing any specific information about the statement. For example, a medical data owner can prove their ownership of the data without revealing the actual content of the data.
[0033] In this embodiment, intelligent contract parsing refers to using intelligent contracts to automatically execute and verify the processing logic of data. For example, an intelligent contract is used to verify the consistency between the hash of medical data shards and the Merkle root hash.
[0034] In this embodiment, cross-chain hash indexing refers to establishing indexes between different blockchains to facilitate data sharing and interoperability. For example, a cross-chain hash index that associates medical data shards on different blockchains to facilitate data sharing and querying.
[0035] In this embodiment, the hierarchical data release rule refers to the rule defined in the intelligent contract to control the release and access rights of data according to the level and access conditions of the data. For example, a rule that allows researchers to access publicly level medical data, but additional authorization is required to access restricted level data.
[0036] In this embodiment, the research database refers to a well-organized data collection that contains data records and metadata for research purposes. For example, a database that contains medical records, patient information, and clinical research results to support medical research.
[0037] The working principle and beneficial effects of the above technical solution are as follows: Collect hospital-related data through blockchain technology, perform type classification and hash evidence storage, upload the data digest to the chain and perform sharding processing, establish an anti-quantum Merkle tree, add timestamp anchoring, the intelligent contract automatically parses the zk-SNARK proof in the block, and establish a cross-chain hash index, and finally form a research database. This helps to improve the transparency, credibility, security, and shareability of data, thus promoting the development of real-world research.
[0038] Embodiment 2: The embodiment of the present invention provides a method for managing a real-world research database based on blockchain, which collects hospital-related data, classifies the related data, and generates corresponding hash evidence storage according to the classification result, including: Dock with the interfaces of each business system of the hospital, and perform adaptation judgment on the interface status. Select the corresponding adaptation tool from the format-adaptation table according to the adaptation judgment result, and collect hospital-related data based on the adaptation tool; Classify the related data, perform standard conversion based on the classification result to obtain the conversion result, and generate hash evidence storage according to the conversion result.
[0039] In this embodiment, the interfaces of each business system of the hospital refer to the interfaces provided by different information systems within the hospital (such as the electronic medical record system, pharmacy management system, laboratory information system, etc.) for data exchange and communication. For example, the API interface of the electronic medical record system is used to query and update patient information.
[0040] In this embodiment, the interface status refers to indicators such as the health status, availability, performance, and security of the interface. For example, whether the interface is online, the response time, and the error rate.
[0041] In this embodiment, the adaptation judgment refers to the process of evaluating whether the interface can be compatible with the data acquisition system. For example, checking whether the protocol, data format, and data transmission method of the interface meet the requirements.
[0042] In this embodiment, the format - adaptation table is a mapping table that maps different interface formats to corresponding adaptation tools. This table is usually constructed based on the characteristics of the interface and the requirements of the data acquisition system. For example, a table that lists interface formats (such as HL7, FHIR, DICOM) and the corresponding adaptation tools (such as data converters, protocol converters).
[0043] In this embodiment, the adaptation tool refers to a tool used to convert data in different formats into a format that the data acquisition system can process. For example, a converter that converts data in HL7 format to FHIR format.
[0044] In this embodiment, the standard conversion refers to the process of converting different types of data into a unified format or standard. For example, converting dates in different formats into the ISO8601 standard format.
[0045] In this embodiment, the conversion result refers to the data after standard conversion. For example, a patient information record in a unified format after conversion.
[0046] The working principle and beneficial effects of the above - mentioned technical solution are as follows: By evaluating the compatibility of the hospital business system interface, selecting appropriate adaptation tools to collect data, classifying and standardizing the data conversion, and generating hash evidence to ensure data integrity and traceability, it improves data processing efficiency, promotes data sharing and collaboration, and enhances the repeatability and credibility of research.
[0047] Embodiment 3: The embodiment of the present invention provides a method for managing a real - world research database based on blockchain. It performs type division on the relevant data, conducts standard conversion based on the type - division result to obtain a conversion result, and generates hash evidence according to the conversion result, including: Input the relevant data into the feature analysis engine, parse the relevant data types and add corresponding type tags to obtain the type - division result; Match the structured data in the type - division result with the real - time terminology of the knowledge graph, and then perform unit intelligent conversion to obtain the first conversion result; Perform multi - modal AI annotation on the unstructured data in the type - division result and perform cross - modal alignment to convert it into a standard result, and then obtain the second conversion result; Serialize the first conversion result and the second conversion result, calculate the hash value of the serialized result using an encryption hashing algorithm, and then generate a hash deposit certificate.
[0048] In this embodiment, the feature analysis engine refers to a tool or software for analyzing and extracting data features, identifying the type, structure, and other attributes of data. For example, a natural language processing (NLP) engine is used to parse text data and extract keywords and entities.
[0049] In this embodiment, the type tag refers to a tag or classification used to identify the data type. These tags can help the system understand the meaning and use of the data. For example, for numerical data, the type tag can be numerical; for text data, the type tag can be text.
[0050] In this embodiment, structured data refers to data with a clear format and structure, such as data in a spreadsheet or records in a database. For example, a patient information table contains fields such as name, age, and gender.
[0051] In this embodiment, the real-time terminology matching of the knowledge graph refers to using the knowledge graph to understand and interpret terms and concepts in the data. This usually involves matching the terms in the data with the entities in the knowledge graph. For example, using a medical knowledge graph to parse the disease names and symptom descriptions in an electronic medical record.
[0052] In this embodiment, unit intelligent conversion refers to the process of converting the units in the data into standard units. This usually involves using rules or algorithms to identify and convert different measurement units. For example, converting body temperature from Celsius to Fahrenheit.
[0053] In this embodiment, the first conversion result refers to the structured data after unit intelligent conversion. For example, converting a patient's weight from kilograms to pounds.
[0054] In this embodiment, unstructured data refers to data without a clear format and structure, such as text, images, audio, etc. For example, doctors' diagnostic notes, X-ray images, and patient interview recordings.
[0055] In this embodiment, multi-modal AI annotation refers to the process of using artificial intelligence technology to annotate unstructured data. This usually involves using machine learning models to identify and mark different elements in the data. For example, using an image recognition model to annotate fractures in an X-ray image.
[0056] In this embodiment, cross-modal alignment refers to the process of aligning or synchronizing data in different modalities. This usually involves using algorithms to match the timestamps or events in data of different modalities. For example, aligning the timestamps of a patient interview recording with the event records in an electronic medical record.
[0057] In this embodiment, the second conversion result refers to the unstructured data after multi-modal AI annotation and cross-modal alignment. For example, keywords and entities are extracted from the patient interview recording and aligned with the text data in the electronic medical record.
[0058] In this embodiment, serialization refers to the process of converting data into a format that can be stored or transmitted. This usually involves converting a data structure into a byte stream or other format. For example, converting a patient information form into JSON format.
[0059] The working principle and beneficial effects of the above technical solution are as follows: The feature analysis engine is used to parse the data type and add tags, match the structured data with the knowledge graph for intelligent conversion, convert the unstructured data into standard results through multi-modal AI annotation and cross-modal alignment, and finally serialize the conversion results and calculate the hash value to generate a hash deposit certificate, improving the accuracy and reliability of the data.
[0060] Embodiment 4: The embodiment of the present invention provides a method for managing a real-world research database based on blockchain. Generate a digest based on the hash deposit certificate, perform sharding processing on the digest, effectively verify the sharding results, and write the verified valid sharding information into a new block, including: Extract the key fields of the hash deposit certificate and convert them into a compact JSON format to obtain a JSON string, calculate a secondary hash for the JSON string, and generate an on-chain identifier; Generate a digest based on the on-chain identifier, group the digest to obtain a time series group and a verification group, shard the time series group according to a fixed length, pad with zeros if insufficient, encode the shards, and calculate the metadata of all shards according to the encoding; Determine the verification group nodes according to VRF random allocation. The time series group broadcasts the metadata to the verification group nodes, and each verification group node verifies a specific shard according to the VRF allocation to generate a signature proof; The time series group receives the signature proofs of all verification group nodes, calculates the effective shard ratio of the signature proofs, packs the corresponding valid shard information based on the effective shard ratio to generate transaction data, submits the transaction to the blockchain, and performs malicious node processing on the blockchain to form a network consensus, and writes the valid shard information into a new block.
[0061] In this embodiment, the key fields refer to the most important information in the hash deposit certificate, which is used to identify and verify the integrity and authenticity of the data. For example, the hash value, timestamp, and data type of the data.
[0062] In this embodiment, the compact JSON format refers to representing data as a JSON-formatted string in a concise and efficient manner. This typically involves removing unnecessary whitespace and line breaks, as well as using short key names.
[0063] In this embodiment, a JSON string refers to a string that represents data in JSON format. JSON is a lightweight data interchange format that is easy for humans to read and machines to parse. For example, {key1:value1,key2:value2}.
[0064] In this embodiment, secondary hashing refers to applying a hash function to a JSON string again to generate a new hash value. This is typically used to generate a unique identifier for data. For example, using the SHA-256 hash function to generate a hash value for a JSON string.
[0065] In this embodiment, an on-chain identifier refers to a unique identifier used to identify data on a blockchain. This is typically a hash value or an address, such as an Ethereum address, used to identify a smart contract.
[0066] In this embodiment, a time series group refers to a grouping of data arranged in chronological order. This is typically used to record the order and timeliness of data. For example, a series of transaction records arranged in chronological order.
[0067] In this embodiment, a verification group refers to a set of nodes used to verify the integrity and authenticity of data. These nodes typically need to be authenticated and authorized. For example, a group of verification nodes participating in a blockchain network.
[0068] In this embodiment, fixed length means that the length of the data is fixed, and the insufficient part needs to be filled with zeros or other fillers. This is typically used to ensure the consistency and comparability of data. For example, a 128-bit hash value with zeros filled for the insufficient part.
[0069] In this embodiment, encoding refers to converting data into a specific format or encoding method. This typically involves using a specific encoding scheme or algorithm. For example, converting text data into UTF-8 encoding.
[0070] In this embodiment, metadata refers to data about data, such as the source, format, size, and other attributes of the data. For example, the metadata of a file includes the file name, size, creation date, etc.
[0071] In this embodiment, VRF random allocation refers to using a Verifiable Random Function (VRF) to randomly allocate tasks or resources. For example, using VRF to randomly allocate a verification task to a specific verification node.
[0072] In this embodiment, the verification group nodes refer to the nodes participating in the data verification process. These nodes usually need to be authenticated and authorized. For example, a medical node participating in a certain network.
[0073] In this embodiment, a specific shard refers to a part of the data for verification or storage. For example, a shard of a fixed size in a large data file.
[0074] In this embodiment, the signature proof refers to the proof using digital signature technology to prove the integrity and authenticity of the data. For example, a digital signature that signs the data digest using a private key.
[0075] In this embodiment, the valid shard ratio refers to the ratio of the valid shards to the total shards during the verification process. For example, among 100 shards, 95 shards are verified as valid, and the valid shard ratio is 95%.
[0076] In this embodiment, the transaction data refers to the data of transactions or operations conducted on the blockchain. For example, a certain network transaction, including the sender, the receiver, and the transaction amount.
[0077] In this embodiment, the network consensus refers to the process in which all nodes in the blockchain network reach an agreement on the data state. For example, the proof-of-work (PoW) consensus mechanism in a certain network.
[0078] In this embodiment, the malicious node handling refers to the process of identifying and handling malicious nodes in the blockchain network. For example, identifying and punishing a malicious node that attempts to double-spend in a certain network.
[0079] The working principle and beneficial effects of the above technical solution are as follows: Convert the key fields of the hash deposit into a compact JSON format, calculate the secondary hash to generate the on-chain identifier, generate the digest and group it into a time series group and a verification group, encode the shards of the time series group to calculate the metadata, randomly assign verification group nodes by VRF to verify the shards and generate the signature proof, the time series group counts the valid shard ratio and packages the transaction data to submit to the blockchain, conduct malicious node handling to form network consensus, and write to the new block to ensure the immutability and traceability of the data, and promote data sharing and collaboration.
[0080] Embodiment 5: The embodiment of the present invention provides a method for managing a real-world research database based on blockchain. According to the random assignment by VRF, the verification group nodes are determined, and the time series group broadcasts the metadata to the verification group nodes. Each verification group node verifies a specific shard according to the VRF assignment and generates a signature proof, including: Set a dynamic pledge mechanism according to the hospital level, and determine a variable candidate node list by combining the dynamic pledge mechanism, the medical dedicated blockchain network, and the hospital characteristics; Perform VRF calculations on candidate nodes in the variable candidate node list based on hash evidence storage. Nodes with the last two digits of the VRF output value less than the preset value form a verification group, and the remaining nodes form a timing group; The timing group packs the metadata of the current time window and sends it to all verification group nodes through the medical private network. Each verification group node calculates a shard index based on its own VRF output value and determines the specific shard verification right of each verification group node according to the shard index; The verification group nodes perform medical data verification according to the specific shard verification right and generate a signature proof after passing the verification.
[0081] In this embodiment, the hospital grade refers to different grades divided by hospitals according to factors such as their scale, service capabilities, and technical levels. For example, tertiary grade A hospitals, secondary grade B hospitals, etc. For instance, tertiary grade A hospitals, secondary grade B hospitals.
[0082] In this embodiment, the dynamic pledge mechanism refers to dynamically adjusting the number of tokens that nodes need to pledge according to certain characteristics of the nodes (such as reputation, performance, participation, etc.) to ensure the security and incentive compatibility of the network. For example, in a medical blockchain network, the pledge quantity is dynamically adjusted according to the hospital grade and participation.
[0083] In this embodiment, the medical dedicated blockchain network refers to a blockchain network designed specifically for medical data sharing and transactions, such as Figure 3 As shown, it provides features such as high security, privacy protection, and data immutability. For example, a medical data sharing platform based on Ethereum.
[0084] In this embodiment, the hospital characteristics refer to certain attributes or indicators of the hospital, such as the geographical location of the hospital, professional fields, number of patients, etc. For example, the location of the hospital, service scope, number of beds.
[0085] In this embodiment, the variable candidate node list refers to a list of candidate nodes that can change dynamically according to certain rules in a medical blockchain network. These nodes can be hospitals, research institutions, or other medical entities. For example, a candidate node list containing hospitals of different grades and medical research institutions.
[0086] In this embodiment, VRF calculation refers to using a verifiable random function (VRF) to calculate a certain input value (such as a global random seed) to generate a random output value and a proof, and other nodes can verify whether this output value is correct. For example, using VRF to calculate the global random seed to generate a random number and a proof.
[0087] In this embodiment, a node whose last two digits of the VRF output value are less than a preset value refers to a node in the VRF calculation where the last two digits of the output value are less than the preset threshold. These nodes will be selected as the verification group. For example, if the preset threshold is 50 and the VRF output value is 1234567890abcdef, the last two digits are ef, which is less than 50, so this node is selected as the verification group.
[0088] In this embodiment, the shard index refers to the index value used to identify the position of a data shard in the overall data. For example, the shard index of a file indicates the position of the shard in the entire file.
[0089] In this embodiment, the specific shard verification right refers to the right and ability of the verification group nodes to verify their specific shards. For example, the right of a verification node to verify the medical data within its shard.
[0090] In this embodiment, medical data verification refers to the inspection and confirmation of medical data to ensure the accuracy and integrity of the data. For example, verifying the accuracy of patient information, drug dosage, and other data in medical records.
[0091] In this embodiment, signature proof refers to the verification group nodes signing the data they have verified to prove the validity and authenticity of the data. For example, the verification node uses its private key to sign the medical record to generate a digital signature.
[0092] The working principle and beneficial effects of the above technical solution are as follows: A dynamic pledge mechanism is set according to the hospital level to generate a global random seed. Candidate nodes use their private keys for VRF calculation. Nodes whose last two digits of the output value are less than the preset value become the verification group, and the rest are the timing group. The timing group packs the metadata and sends it to the verification group. The verification group nodes calculate the shard index and verify the specific shard. After passing the verification, a signature proof is generated to ensure the integrity and immutability of the data, and to promote the sharing and collaboration of medical data.
[0093] Embodiment 6: The embodiment of the present invention provides a method for managing a real-world research database based on blockchain, which builds an anti-quantum Merkle tree based on a new block, adds a timestamp anchor, and an intelligent contract automatically parses the zk-SNARK proof in the block and establishes a cross-chain hash index, including: Taking the verified specific shards in the new block as leaf nodes, constructing an anti-quantum Merkle tree using the XMSS algorithm to generate a root hash, and performing a timestamp anchor on the root hash; Using an intelligent contract to parse the pre-stored zk-SNARK proof in the new block, performing a one-way hash match on the specific shards to obtain a first matching result; Verifying the consistency between the root hash and the block header according to the Merkle tree to obtain a second verification result; If both the first matching result and the second verification result are passed, a cross-chain hash index is generated.
[0094] In this embodiment, leaf nodes refer to the bottom-most nodes in the Merkle tree structure, which usually contain actual data shards or data hash values. For example, a hash value of a medical record shard.
[0095] In this embodiment, the XMSS algorithm is a digital signature algorithm resistant to quantum computing attacks, which can be used to construct a quantum-resistant Merkle tree. For example, the XMSS algorithm is used to generate signatures for leaf nodes.
[0096] In this embodiment, the root hash refers to the hash value at the top-most layer of the Merkle tree, which is a summary of the entire Merkle tree content and is used to verify the integrity and consistency of data. For example, the top hash value of the Merkle tree is used to verify the hash values of all leaf nodes.
[0097] In this embodiment, one-way hash matching refers to using a hash function to compare the hash values of two data shards to determine whether they are the same, without the need to compare the actual data content. For example, using the SHA-256 hash function to compare the hash values of two medical record shards.
[0098] In this embodiment, the first matching result refers to the result obtained through one-way hash matching, which indicates whether the hash values of two data shards are consistent. For example, if the hash values of two shards match, it indicates that the data has not been tampered with.
[0099] In this embodiment, the block header refers to the header information of each block in the blockchain, which contains metadata of the block, such as timestamp, difficulty target, hash value of the previous block, etc. For example, the header information of a certain network block contains information such as the creation time and difficulty target of the block.
[0100] In this embodiment, the second verification result refers to the result obtained by verifying the consistency between the root hash of the Merkle tree and the block header. For example, if the root hash of the Merkle tree is consistent with the hash value in the block header, it indicates that the block data has not been tampered with.
[0101] The working principle and beneficial effects of the above technical solution are as follows: taking the verified specific shards in the new block as leaf nodes, using the XMSS algorithm to construct a quantum-resistant Merkle tree, generating the root hash and timestamp anchoring; the smart contract parses the pre-stored zk-SNARK proof in the new block and performs one-way hash matching on the specific shards; verifying the consistency between the Merkle root hash and the block header; if both the matching and verification are passed, a cross-chain hash index is generated to promote data sharing between different blockchains.
[0102] Example 7: An embodiment of the present invention provides a method for managing a real-world research database based on blockchain. By setting hierarchical data release rules according to cross-chain hash indexes, a research database is formed, including: Construct index key values according to cross-chain hash indexes, write the index key values into the index mapping table of the target chain, and preset hierarchical conditions and corresponding hierarchical data release rules in the smart contract to form a research database.
[0103] In this embodiment, the index key value refers to a key-value pair used to uniquely identify a data record in a database or an index mapping table. In blockchain, this is usually a hash value or an address. For example, a key-value pair containing the hash value of a medical record shard and its location on the blockchain.
[0104] In this embodiment, the preset hierarchical condition refers to a condition set in the smart contract for dividing data into different levels or categories according to different data access requirements. For example, according to the sensitivity and privacy level of medical data, the data is divided into three levels: public, restricted, and confidential.
[0105] In this embodiment, the index mapping table refers to a data structure that maps index key values to the actual locations of data records, such as block locations in blockchain or row numbers in a database. For example, a table that maps the hash value of a medical record shard to its corresponding location on the blockchain.
[0106] The working principle and beneficial effects of the above technical solution are as follows: By establishing indexes between different blockchains through cross-chain hash indexes, it is convenient for data sharing and querying. By presetting hierarchical conditions and hierarchical data release rules in the smart contract, the security and privacy of data are protected. At the same time, a research database is formed, providing more convenient data access and analysis tools for researchers, and promoting data sharing and interoperability.
[0107] Example 8: An embodiment of the present invention provides a real-world research database management system based on blockchain, as Figure 2 shown, including: Hash evidence storage module: Collect hospital-related data, classify the related data, and generate corresponding hash evidence according to the classification result; Sharding module: Generate a digest and upload it to the chain according to the hash evidence, perform sharding processing on the digest, effectively verify the sharding result, and write the verified valid sharding information into a new block; Cross-chain module: Build a quantum-resistant Merkle tree based on the new block, add a timestamp anchor, the smart contract automatically parses the zk-SNARK proof in the block, and establishes a cross-chain hash index; Database module: Set hierarchical data release rules according to the cross-chain hash index to form a research database.
[0108] The working principle and beneficial effects of the above technical solution are as follows: Collect relevant hospital data through blockchain technology, classify and hash the data for evidence retention, upload the data digest to the blockchain and perform sharding processing, establish a quantum-resistant Merkle tree, add timestamp anchoring, have smart contracts automatically parse the zk-SNARK proofs within the blocks, and establish a cross-chain hash index, ultimately forming a research database. This is conducive to improving the transparency, credibility, security, and shareability of data, thereby promoting the development of real-world research.
[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for managing a real-world research database based on blockchain, characterized in that, Including: Step 1: Collect hospital-related data, classify the related data by type, and generate corresponding hash certificates based on the classification results; Step 2: Generate a digest and upload it to the chain according to the hash certificate, perform sharding on the digest, effectively verify the sharding results, and write the verified valid sharding information into a new block; Step 3: Build a quantum-resistant Merkle tree based on the new block, add a timestamp anchor, the smart contract automatically parses the zk-SNARK proof in the block, and builds a cross-chain hash index; Step 4: Set a hierarchical data release rule according to the cross-chain hash index to form a research database.
2. The method for managing a real-world research database based on blockchain according to claim 1, wherein Collect hospital-related data, classify the related data by type, and generate corresponding hash certificates based on the classification results, including: Connect to the interface of each business system in the hospital, judge the adaptation status of the interface, select the corresponding adaptation tool from the format-adaptation table according to the adaptation judgment result, and collect hospital-related data based on the adaptation tool; Classify the related data by type, perform standard conversion based on the classification result, obtain the conversion result, and generate a hash certificate according to the conversion result.
3. A method for managing a real-world research database based on blockchain according to claim 2, wherein Classify the related data by type, perform standard conversion based on the classification result, obtain the conversion result, and generate a hash certificate, including: Input the related data into the feature analysis engine, parse the related data type and add the corresponding type label to obtain the classification result; Match the structured data in the classification result with the real-time terminology of the knowledge graph, and then perform unit intelligent conversion to obtain the first conversion result; Perform multi-modal AI annotation on the unstructured data in the classification result and perform cross-modal alignment to convert it into a standard result, and then obtain the second conversion result; Serialize the first conversion result and the second conversion result, calculate the hash value of the serialized result using the cryptographic hash algorithm, and then generate a hash certificate.
4. A method for managing a real-world research database based on blockchain according to claim 1, characterized in that, Generate a digest and upload it to the chain according to the hash certificate, perform sharding on the digest, effectively verify the sharding results, and write the verified valid sharding information into a new block, including: Extract the key fields of the hash certificate and convert them into a compact JSON format to obtain a JSON string, calculate the secondary hash of the JSON string, and generate a chain identifier; Generate a digest according to the chain identifier, group the digest, obtain a time series group and a verification group, slice the time series group according to a fixed length, pad with zeros if it is insufficient, encode the slices, and calculate the metadata of all slices according to the encoding; Determine the verification group nodes according to the VRF random allocation. The time series group broadcasts the metadata to the verification group nodes. Each verification group node verifies a specific slice according to the VRF allocation and generates a signature proof; The time series group receives the signature proofs of all verification group nodes, calculates the effective slice ratio of the signature proofs, packs the corresponding valid slice information based on the effective slice ratio to generate transaction data, submits the transaction to the blockchain, and processes malicious nodes in the blockchain to form network consensus, and writes the valid slice information into a new block.
5. A method for managing a real-world research database based on blockchain according to claim 4, characterized in that, Determine the verification group nodes according to the VRF random allocation. The timing group broadcasts the metadata to the verification group nodes. Each verification group node verifies a specific shard according to the VRF allocation and generates a signature proof, including: Set a dynamic pledge mechanism according to the hospital level, and combine the dynamic pledge mechanism with the medical dedicated blockchain network and hospital characteristics to determine a variable candidate node list; Perform VRF calculation on the candidate nodes in the variable candidate node list based on hash evidence storage. The nodes whose last two digits of the VRF output value are less than the preset value become the verification group, and the remaining nodes become the timing group; The timing group packages the metadata of the current time window and sends it to all verification group nodes through the medical private network. Each verification group node calculates the shard index according to its own VRF output value and determines the specific shard verification right of each verification group node according to the shard index; The verification group node performs medical data verification according to the specific shard verification right and generates a signature proof after passing the verification.
6. A method for managing a real-world research database based on blockchain according to claim 5, characterized in that, Based on the new block, establish an anti-quantum Merkle tree, add a timestamp anchor. The smart contract automatically parses the zk-SNARK proof in the block and establishes a cross-chain hash index, including: Use the XMSS algorithm to construct an anti-quantum Merkle tree with the verified specific shards in the new block as leaf nodes, generate the root hash, and perform a timestamp anchor on the root hash; Use the smart contract to parse the pre-stored zk-SNARK proof in the new block, perform a one-way hash match on the specific shard, and obtain the first matching result; Verify the consistency between the root hash and the block header according to the Merkle tree to obtain the second verification result; If both the first matching result and the second verification result pass, generate a cross-chain hash index.
7. A method for managing a real-world research database based on blockchain according to claim 1, characterized in that, Set hierarchical data release rules according to the cross-chain hash index to form a research database, including: Construct index key values according to the cross-chain hash index, write the index key values into the index mapping table of the target chain, and preset hierarchical conditions and corresponding hierarchical data release rules in the smart contract to form a research database.
8. A real-world research database management system based on blockchain, characterized in that Including: Hash evidence storage module: Collect hospital-related data, classify the related data, and generate corresponding hash evidence storage according to the classification results; Sharding module: Generate a digest and upload it to the chain according to the hash evidence storage, perform sharding processing on the digest, effectively verify the sharding results, and write the verified valid shard information into the new block; Cross-chain module: Based on the new block, establish an anti-quantum Merkle tree, add a timestamp anchor. The smart contract automatically parses the zk-SNARK proof in the block and establishes a cross-chain hash index; Database module: Set hierarchical data release rules according to the cross-chain hash index to form a research database.
Citation Information
Patent Citations
Data evidence storage method and system based on block chain and interstellar file system
CN112084164A
Block chain evidence storage method and system based on isomorphic multi-chain architecture
CN113326317A
Cross-chain method between block chains and main block chain system
CN113704356A
Expandable block chain network model based on fragmentation
CN116668313A
Block chain data transaction method and device and storage medium
CN117437054A
Cited By
Block chain engineering contract evidence storage method based on BIM
CN120849439A
Log generation method and device based on block chain and storage medium
CN121118116A