Medical channel data checking method and system

By employing multimodal access, intelligent format adaptation, differential privacy desensitization, layered blockchain notarization, federated learning, and dynamic consensus mechanisms, this approach addresses issues such as heterogeneous data formats, privacy protection, cross-institutional collaboration, and anomaly detection in medical channel data verification, achieving efficient, secure, and accurate data verification throughout the entire process.

CN121786860APending Publication Date: 2026-04-03SHANGHAI BEITONG MEDICAL DEVICE MANAGEMENT CONSULTING CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing methods for verifying medical data suffer from problems such as heterogeneous data formats, difficulty in protecting privacy, insufficient cross-institutional collaboration capabilities, limited real-time response capabilities for anomaly detection, and low traceability and verification efficiency, making it difficult to achieve efficient, secure, and accurate data verification throughout the entire process.

Method used

Employing multimodal access and intelligent format adaptation, combined with differential privacy desensitization and hybrid encryption technologies, a layered blockchain for evidence storage is deployed. Based on a federated learning framework, cross-institutional model collaborative training is achieved, constructing a digital twin model for medical scenarios. Combining a real-time stream processing architecture and online learning model, a dynamic consensus mechanism is used for multi-institutional collaborative verification, generating tiered intelligent alerts, and tracing the source through blockchain evidence storage.

Benefits of technology

It achieves a unified standard mapping for data across all channels, ensuring data privacy and preventing leaks, improving data integration efficiency and verification accuracy, enhancing the generalization ability of AI quality assessment, enabling pre-verification of latent anomalies and rapid identification of explicit anomalies, and forming a closed-loop verification system with full-process traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786860A_ABST
    Figure CN121786860A_ABST
Patent Text Reader

Abstract

The invention discloses a medical channel data checking method and system, and relates to the technical field of medical data checking, omni-channel medical data is collected through multi-modal access, the omni-channel medical data is mapped to a unified standard through intelligent format adaptation, data privacy is guaranteed by adopting a differential privacy desensitization and hybrid encryption technology, and standardized encrypted data is output; grading according to data sensitivity and respectively writing the data into an alliance chain, a private chain or a distributed account book, generating a main hash chain type structure and a sub-hash chain type structure, and recording a full-life-cycle flow track of the data. Through multi-mode access and intelligent format adaptation, unified standard mapping of all-channel data is realized, and the data integration efficiency and the checking foundation quality are improved. Differential privacy desensitization and hybrid encryption technologies are combined with hierarchical block chain evidence storage, so that data privacy is not leaked, and compliance requirements are met. The federated learning framework realizes cross-mechanism model cooperative training, can improve the generalization ability of AI quality evaluation without collecting original data, and covers numerical anomaly, logic conflict and trend anomaly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data verification technology, and in particular to a method and system for verifying medical channel data. Background Technology

[0002] As the digital transformation of the healthcare industry progresses, the sources of medical data are becoming increasingly diverse, encompassing multiple channels such as medical institution records, insurance claims data, monitoring data from medical IoT devices, and drug distribution data. However, existing methods for verifying medical channel data have many shortcomings and are insufficient to meet practical application needs.

[0003] At the data level, the heterogeneous formats of multi-source data, with different institutions often using proprietary data standards, make data integration difficult, and the lack of format consistency directly affects the accuracy of verification. Regarding privacy protection, medical data contains a large amount of sensitive information. In traditional data sharing and verification processes, centralized data storage and transmission are prone to privacy leaks, and compliance is difficult to guarantee.

[0004] In terms of cross-institutional collaboration, data from various medical institutions are isolated, forming data silos. Traditional AI verification models rely on centralized data training, resulting in insufficient generalization ability and an inability to achieve cross-institutional collaborative quality assessment. Regarding anomaly detection, existing methods mostly focus on identifying explicit numerical anomalies, lacking the ability to predict conflicts in medical business logic and hidden risks in data flow. Furthermore, their real-time response capabilities are limited, making it difficult to quickly handle urgent anomalies.

[0005] Regarding traceability and verification efficiency, data modification records are easily lost or tampered with, making it difficult to trace abnormal data. The verification process is fixed and rigid, failing to dynamically adapt verification nodes and consensus mechanisms based on data importance and anomaly levels, resulting in low verification efficiency or insufficient rigor in verifying key data. Furthermore, existing methods lack a complete closed loop, with poor coordination between verification, correction, tracing, and result output, making it difficult to form standardized and complete verification results. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a method and system for verifying medical channel data. The technical solution adopted is as follows: The data verification method for medical channels includes the following steps: Step 1: Collect medical data from all channels through multimodal access, map it to a unified standard through intelligent format adaptation, and use differential privacy desensitization and hybrid encryption technology to protect data privacy, outputting standardized encrypted data; Step 2: Classify data according to sensitivity and write it into the consortium blockchain, private blockchain or distributed ledger respectively, generate a main hash and sub-hash chain structure to record the entire life cycle of data flow, and deploy layered exclusive smart contracts; Step 3: Implement cross-institutional model collaborative training based on the federated learning framework. Through anomaly detection, logic verification, trend analysis, and multi-model collaboration, output quality scores and grade labels. Step 4: Construct a digital twin model of the medical scenario to simulate data flow and pre-identify hidden anomalies. Combine a real-time stream processing architecture with an online learning model to generate hierarchical intelligent alerts.

[0007] Optionally, step 5 is also included: matching verification nodes through smart contracts, completing multi-institutional collaborative verification using a dynamic consensus mechanism, updating the hash chain and recording the modification trajectory after securely correcting the data, and outputting the corrected quality score and grading mark.

[0008] Optionally, step 6 is also included, which triggers the corresponding response level based on the risk level, and uses the data fingerprint chain stored on the blockchain to trace back the source of the risk and generate a traceability timeline.

[0009] Optionally, in step 1, multimodal access supports API, file upload, streaming interface, medical IoT devices, and the acquisition of DICOM format image data; intelligent format adaptation has a built-in adaptive parsing engine that automatically identifies the private data format of medical institutions and maps it to the FHIR or HL7 international medical data standard, supporting custom business field extensions; differential privacy desensitization uses an ε-differential privacy mechanism to inject noise into sensitive patient fields, and hybrid encryption uses a combination of the national cryptographic algorithm SM4 and AES-256 encryption scheme. The output standardized encrypted data must include data source identifier, acquisition timestamp, and format check code.

[0010] Optionally, in step 2, the data can be categorized into core, important, and ordinary levels based on sensitivity. The core level corresponds to patient identity information, core diagnostic conclusions, and surgical records; the important level corresponds to treatment details, medication records, and claim amounts; and the ordinary level corresponds to equipment operating parameters and non-sensitive statistical data. Core-level data is written to a cross-institutional consortium blockchain, important-level data is written to an institution's private blockchain, and ordinary-level data is stored using a distributed ledger and lightweight hash digest notarization. Both the main hash and sub-hashes are generated based on the SHA-256 algorithm. Sub-hashes are generated sequentially according to the flow nodes of the collection end, standardization end, verification end, and validation end, forming a chain-like traceability structure. Layered dedicated smart contracts include core-level contracts, important-level contracts, and ordinary-level contracts. Contracts automatically associate the hash storage and verification of data at the corresponding level.

[0011] Optionally, in step 3, federated learning uses the FedAvg framework to achieve cross-institutional model collaborative training. Each participating institution, as a federated node, only shares model parameters and uses a secure aggregation algorithm to encrypt and summarize feature vectors. Multi-model collaboration specifically includes: using isolated forest or LSTM autoencoder to identify numerical anomalies, which refer to test results that exceed clinically reasonable thresholds; verifying logical conflicts based on authoritative medical knowledge graphs; and using ARIMA or Prophet time series models to analyze trend anomalies. The quality score is divided into three levels from 0 to 100: credible, pending review, and abnormal. The credible level corresponds to 80-100 points, the pending review level corresponds to 60-79 points, and the abnormal level corresponds to less than 60 points. The hierarchical markers are associated one by one with the hash chain nodes in step 2, clarifying the initial flow of abnormal data.

[0012] Optionally, in step 4, the digital twin model of the medical scenario is constructed according to the core scenario. The scenario construction includes outpatient diagnosis and treatment, inpatient nursing, insurance claims and drug distribution, accurately mapping business processes, role permissions and diagnosis and treatment logic constraints. Pre-verification identifies hidden anomalies and generates early warning reports by simulating the entire data flow; real-time stream processing adopts the Apache Flink+Kafka architecture, and the online learning model uses the incremental SVM algorithm, which can dynamically adapt to new data distributions. The tiered intelligent alerts are divided into Level 1, Level 2, and Level 3 alerts based on urgency, financial risk relevance, statistical error relevance, and scope of impact. The alert content includes anomaly details, hash chain tracing nodes, recommended processing paths, and responsible parties.

[0013] Optionally, in step 5, the smart contract matches verification nodes based on a dual dimension of data type and anomaly level. Anomalies in emergency treatment are automatically assigned to the attending physician and department head, while anomalies in large insurance claims are assigned to hospital finance, insurance company underwriters, and regulatory agency specialists. The multimodal verification tool supports structured verification, unstructured verification, and AI-assisted verification. Structured verification involves checking to confirm the authenticity of the data, unstructured verification involves uploading supporting materials, including medical records and examination reports, and AI-assisted verification involves the system automatically matching similar cases to recommend verification conclusions. The dynamic consensus mechanism employs a single-node confirmation and multi-node supervision mode for ordinary anomalies, and a Byzantine fault-tolerant consensus mechanism for core-level data anomalies, requiring at least two-thirds of the participating nodes to complete digital signature confirmation. Data correction is only allowed to be performed by the original data provider. After correction, a new master hash and sub-hash are generated, and the old hash and modification record are associated and stored in the corresponding blockchain layer. The modification record includes the modifier, modification time, modification reason, and verification conclusion. Finally, the corrected quality score, grade mark, and modification trajectory traceability report are output.

[0014] Optionally, in step 6, the risk level and response level correspond one-to-one. High risk corresponds to a level 1 alarm, triggering an emergency response, immediately freezing the abnormal data flow, and pushing emergency consultation notifications to multiple parties. Risk tracing is based on the data fingerprint chain stored on the blockchain. By tracing the sub-hash in reverse, the operation records of the data at each flow node are traced. The operation records include the operator, operation time, and data status changes, generating a visual tracing timeline. Finally, a complete verification result is output, which includes the final data quality score, classification conclusion, anomaly handling details, risk tracing report, and blockchain storage compliance proof. The classification conclusion is either credible, pending review, or the anomaly has been corrected, realizing a closed-loop output of the entire process of collection, verification, correction, tracing, and conclusion.

[0015] The medical channel data verification system is used to implement medical channel data verification methods. The system includes modules for data collection and standardization, hierarchical blockchain, federated learning, digital twin, dynamic consensus and collaboration, risk response, and traceability and result output. The data acquisition and standardization module is used for multi-source data acquisition, format adaptation, de-identification and encryption, and output of standardized data; The layered blockchain module is used for data classification, hash-based notarization, and smart contract deployment. The federated learning module is used for collaborative model training and data quality assessment across institutions; The digital twin module is used for scenario simulation pre-verification and real-time anomaly detection and alarm; The dynamic consensus collaboration module is used to verify node matching, multi-agency collaborative verification, and data correction. The risk response module is used to trigger corresponding responses based on risk levels; The traceability and result output module is used for risk traceability and output of complete verification results; the system adopts a microservice distributed architecture to support cross-organizational collaboration and full-process security control.

[0016] In summary, the present invention has at least one of the following beneficial technical effects: This invention provides a method and system for verifying medical channel data. Through multimodal access and intelligent format adaptation, it achieves a unified standard mapping of data across all channels, improving data integration efficiency and the basic quality of verification. Differential privacy desensitization and hybrid encryption technology combined with layered blockchain notarization ensure data privacy while meeting compliance requirements. The federated learning framework enables cross-institutional model collaborative training, improving the generalization ability of AI quality assessment without the need to aggregate raw data, covering numerical anomalies, logical conflicts, and trend anomalies.

[0017] Digital twin models enable pre-verification of latent anomalies, while real-time stream processing architecture and online learning models ensure rapid identification of explicit anomalies and alert push notifications.

[0018] The dynamic consensus mechanism adapts the verification process according to the data type and anomaly level. Core data adopts a strict consensus mechanism, while ordinary data has a simplified verification process, balancing efficiency and rigor.

[0019] It achieves full-process traceability and closed-loop management. The blockchain hash chain records the entire lifecycle flow and modification trajectory of data. The risk tracing module accurately locates the source of anomalies and finally outputs a complete verification result including quality score, processing details, tracing report and compliance certificate, forming a closed-loop system of collection, verification, correction, tracing and conclusion. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating the medical channel data verification method of the present invention. Detailed Implementation

[0021] The present invention will be further described in detail below with reference to the accompanying drawings.

[0022] This invention discloses a method and system for verifying medical channel data.

[0023] Reference Figure 1 Example 1, Medical Channel Data Verification Method, includes the following steps: Step 1: Collect medical data from all channels through multimodal access, map it to a unified standard through intelligent format adaptation, and use differential privacy desensitization and hybrid encryption technology to protect data privacy, outputting standardized encrypted data; Step 2: Classify data according to sensitivity and write it into the consortium blockchain, private blockchain or distributed ledger respectively, generate a main hash and sub-hash chain structure to record the entire life cycle of data flow, and deploy layered exclusive smart contracts; Step 3: Implement cross-institutional model collaborative training based on the federated learning framework. Through anomaly detection, logic verification, trend analysis, and multi-model collaboration, output quality scores and grade labels. Step 4: Construct a digital twin model of the medical scenario to simulate data flow and pre-identify hidden anomalies. Combine a real-time stream processing architecture with an online learning model to generate hierarchical intelligent alerts.

[0024] Example 2 also includes step 5, which involves matching verification nodes through smart contracts, using a dynamic consensus mechanism to complete multi-agency collaborative verification, updating the hash chain and recording the modification trajectory after securely correcting the data, and outputting the corrected quality score and grading mark.

[0025] Example 3 also includes step 6, which involves using the data fingerprint chain stored on the blockchain to trace back the source of the risk and generate a traceability timeline based on the corresponding response level triggered by the risk level.

[0026] Example 4: In step 1, multimodal access supports API, file upload, streaming interface, medical IoT devices, and the acquisition of DICOM format image data; intelligent format adaptation has a built-in adaptive parsing engine that automatically identifies the private data format of medical institutions and maps it to the FHIR or HL7 international medical data standard, supporting custom business field extensions; differential privacy desensitization uses an ε-differential privacy mechanism to inject noise into sensitive patient fields, and hybrid encryption uses a combination of the national cryptographic algorithm SM4 and AES-256 encryption scheme. The output standardized encrypted data must include the data source identifier, acquisition timestamp, and format check code.

[0027] By adopting the above technical solutions, multimodal access achieves full-channel coverage of medical data by being compatible with various transmission methods such as APIs, files, streaming interfaces, medical IoT devices, and image data, overcoming the limitations of single access methods. Intelligent format adaptation relies on an adaptive parsing engine, automatically matching the private data structure of medical institutions through a pre-trained format recognition model, and then completing format conversion according to the field mapping rules of the FHIR or HL7 international standards. It also supports custom field extensions to adapt to personalized business needs. Differential privacy desensitization uses an ε-differential privacy mechanism to inject noise that conforms to the privacy budget into sensitive patient fields, achieving a balance between data availability and privacy. Hybrid encryption combines the national cryptographic algorithms SM4 and AES-256, leveraging the complementary strengths of the two algorithms to ensure data transmission and storage security. Finally, it outputs standardized data containing source identifiers, timestamps, and checksums, providing a unified and reliable input basis for subsequent verification.

[0028] The hierarchical blockchain evidence storage and traceability principle is based on the tiered classification of data sensitivity, taking into account the differences in privacy and importance of medical data. Different evidence storage strategies are adopted for core-level (highly sensitive), important-level (medium sensitive), and ordinary-level (low sensitive) data: core-level data is written to a cross-institutional consortium blockchain, ensuring immutability through multi-node consensus; important-level data is written to an institution's private blockchain, balancing security and management flexibility; ordinary-level data is stored using a distributed ledger + hash digest, reducing storage costs. The main hash and sub-hashes are generated based on the SHA-256 algorithm. Sub-hashes are generated sequentially and chained together according to the data flow nodes, forming a chain-like trajectory of "collection-standardization-verification-review," achieving full lifecycle traceability of data; the layered dedicated smart contracts automatically associate the hash storage, verification, and access control of the corresponding layer of data through preset logical rules, reducing manual intervention and improving evidence storage efficiency.

[0029] The principle of federated learning-based collaborative AI evaluation: Federated learning, based on the FedAvg framework, distributes training tasks to federated nodes across institutions. Each node trains its model using local data and only uploads updated parameter values. Parameters are encrypted and aggregated using a secure aggregation algorithm, avoiding the transmission of raw data across institutions and overcoming data silos. Multi-model collaboration achieves comprehensive quality assessment through functional complementarity: Isolation forests or LSTM autoencoders utilize unsupervised learning to identify numerical anomalies exceeding clinically reasonable thresholds; authoritative medical knowledge graphs integrate treatment guidelines and coding libraries, verifying logical conflicts such as diagnoses and medications, symptoms and conclusions through semantic matching; ARIMA or Prophet time-series models predict trends based on historical data, identifying abnormal trends such as sudden increases or decreases in data without medical basis. Quality scores are converted from the model's output anomaly probability into a 0-100 quantitative result, with graded labels associated with hash chain nodes, achieving initial localization of anomalous data.

[0030] The principles of digital twins and real-time anomaly detection in medical scenarios: Digital twin models map core business processes such as outpatient services, inpatient care, claims processing, and drug distribution to construct virtual scenarios including role permissions, treatment rules, and data flow logic. This simulates the entire data transmission process in real business operations, proactively identifying hidden anomalies such as "mismatch between symptoms and medication" and generating early warning reports. Real-time stream processing employs an Apache Flink + Kafka architecture. Kafka enables high-throughput data reception, while Flink, based on a streaming computing engine, achieves millisecond-level data processing. The online learning model uses an incremental SVM algorithm, dynamically updating model parameters based on new incoming data to adapt to changes in data distribution. Tiered intelligent alerts combine urgency and impact, classifying anomalies into different levels through preset rules, synchronously associating hash chain tracing nodes and handling suggestions to improve the targeting and efficiency of anomaly response.

[0031] The collaborative verification and correction principle of smart contracts constructs matching rules based on two dimensions: data type (diagnosis, claims, etc.) and anomaly level. It automatically assigns corresponding professional verification nodes to anomaly data in different scenarios, ensuring the professionalism of the verification entity. Multimodal verification tools meet the verification needs of different types of anomalies through structured selection, uploading of unstructured supporting materials, and AI-powered similar case recommendations. The dynamic consensus mechanism is designed differently based on data importance: ordinary anomalies use a "single-node confirmation + multi-node supervision" approach to balance efficiency and security; core-level data anomalies employ a Byzantine fault-tolerant consensus mechanism, requiring signature confirmation from at least two-thirds of the nodes to ensure the rigor of critical data verification. Data correction only grants operation permissions to the original data provider. After correction, a new hash is generated and associated with the old hash and modification records (person, time, reason, conclusion) for evidence storage, achieving transparency and traceability of the correction process.

[0032] The risk response and tracing principle links risk levels and response levels through preset mapping rules. High risk corresponds to a Level 1 alarm, triggering an emergency response (freezing data flow and multi-party consultation). Medium and low risks trigger regular and ordinary responses respectively, achieving a reasonable allocation of resources for anomaly handling. Risk tracing relies on a blockchain-based sub-hash chain structure. By traversing the sub-hashes of each flow node in reverse, it retrieves the operation records (personnel, time, data status) of the corresponding node and presents them in a visual timeline format, accurately locating the link and cause of the anomaly. The final complete verification result integrates the final quality assessment, grading conclusion, processing details, tracing report, and compliance proof, forming a closed-loop output of "collection-verification-correction-tracing-conclusion," meeting the requirements for the completeness and auditability of medical data verification.

[0033] The principles of multimodal and privacy enhancement are refined. Multimodal access expands the coverage of medical IoT devices (wearable devices, diagnostic devices) and DICOM image data acquisition, thus covering a wider range of medical data sources. The adaptive parsing engine improves the accuracy of recognizing data formats from different institutions by increasing the training of private format samples. The ε-differential privacy mechanism optimizes data availability under different privacy requirements by dynamically adjusting noise intensity. Hybrid encryption uses a combination of national cryptographic standards and international algorithms to balance domestic compliance requirements with international compatibility. The output data includes source identifiers, timestamps, and checksums, which are used for source tracing, time-series analysis, and data integrity verification, respectively, further enhancing the reliability of standardized data.

[0034] Example 5: In step 2, the data is divided into core level, important level, and ordinary level according to sensitivity; The core level corresponds to patient identity information, core diagnostic conclusions, and surgical records; the important level corresponds to treatment details, medication records, and claim amounts; and the ordinary level corresponds to equipment operating parameters and non-sensitive statistical data. Core-level data is written to a cross-institutional consortium blockchain, important-level data is written to an institution's private blockchain, and ordinary-level data is stored using a distributed ledger and lightweight hash digest notarization. Both the main hash and sub-hashes are generated based on the SHA-256 algorithm. Sub-hashes are generated sequentially according to the flow nodes of the collection end, standardization end, verification end, and validation end, forming a chain-like traceability structure. Layered dedicated smart contracts include core-level contracts, important-level contracts, and ordinary-level contracts. Contracts automatically associate the hash storage and verification of data at the corresponding level.

[0035] In Example 6, in step 3, federated learning uses the FedAvg framework to achieve cross-institutional model collaborative training. Each participating institution, as a federated node, only shares model parameters and uses a secure aggregation algorithm to encrypt and summarize feature vectors. Multi-model collaboration specifically includes: using isolated forest or LSTM autoencoder to identify numerical anomalies, which refer to test results that exceed clinically reasonable thresholds; verifying logical conflicts based on authoritative medical knowledge graphs; and using ARIMA or Prophet time series models to analyze trend anomalies. The quality score is divided into three levels from 0 to 100: credible, pending review, and abnormal. The credible level corresponds to 80-100 points, the pending review level corresponds to 60-79 points, and the abnormal level corresponds to less than 60 points. The hierarchical markers are associated one by one with the hash chain nodes in step 2, clarifying the initial flow of abnormal data.

[0036] Example 7: In step 4, the digital twin model of the medical scenario is constructed according to the core scenario. The scenario construction includes outpatient diagnosis and treatment, inpatient nursing, insurance claims and drug distribution, accurately mapping business processes, role permissions and diagnosis and treatment logic constraints. Pre-verification identifies hidden anomalies and generates early warning reports by simulating the entire data flow; real-time stream processing adopts the Apache Flink+Kafka architecture, and the online learning model uses the incremental SVM algorithm, which can dynamically adapt to new data distributions. The tiered intelligent alerts are divided into Level 1, Level 2, and Level 3 alerts based on urgency, financial risk relevance, statistical error relevance, and scope of impact. The alert content includes anomaly details, hash chain tracing nodes, recommended processing paths, and responsible parties.

[0037] By adopting the above technical solutions, the core of data sensitivity classification is to implement differentiated evidence storage strategies for different types of data based on the privacy protection needs and business importance of medical data. Core-level data includes the most sensitive information such as patient identity, core diagnoses, and surgical records. It is written to a cross-institutional consortium blockchain through a multi-institutional node consensus mechanism to ensure data immutability and trusted cross-institutional sharing. Important-level data covers sensitive information such as treatment details, medication records, and claim amounts. It is written to an institution's private blockchain to ensure data security and facilitate flexible internal management. Ordinary-level data includes low-sensitivity information such as equipment parameters and non-sensitive statistical data. It is stored using distributed ledgers and hash digests for lightweight evidence storage, which reduces storage and computing costs while meeting basic traceability needs.

[0038] Both the main hash and sub-hashes are generated based on the SHA-256 algorithm, which is irreversible and unique, ensuring the uniqueness of the data "fingerprint." Sub-hashes are generated and chained together sequentially from the collection end, standardization end, verification end, and validation end, forming a chain-like traceability structure. Each step leaves an immutable hash imprint, making the entire lifecycle of data traceable from generation to verification. Layered dedicated smart contracts pre-set differentiated logic according to the processing needs of different levels of data: core-level contracts include multi-party signature verification logic to ensure that critical data operations require confirmation from multiple institutions; important-level contracts have built-in modification permission control logic to limit the scope of data modification; and ordinary-level contracts record access logs to achieve basic operation auditing. Contracts automatically associate the hash storage and verification of data at the corresponding level, completing data storage and integrity verification without manual intervention, improving storage efficiency and reliability.

[0039] Federated learning utilizes the FedAvg framework to achieve cross-institutional model collaborative training. Its principle involves distributing model training tasks across federated nodes in participating institutions. Each node trains the model using only local medical data, and upon completion, uploads only updated model parameters, not the original data. A secure aggregation algorithm encrypts, summarizes, and merges the parameters uploaded by each node to generate a globally optimized model. This approach overcomes the limitations of data silos between institutions while avoiding the privacy risks associated with cross-institutional transmission of raw data.

[0040] Multi-model collaborative evaluation achieves comprehensive data quality verification through functional complementarity: Isolation Forest or LSTM autoencoders adopt an unsupervised learning mode, learning the normal data distribution characteristics based on historical data, and automatically identifying numerical anomalies when the test results exceed clinically reasonable thresholds; authoritative medical knowledge graphs integrate professional knowledge such as clinical treatment guidelines and ICD-10 diagnostic codes, and verify logical conflicts between diagnostic results and data such as medication type, symptoms and conclusions through semantic matching and logical reasoning; ARIMA or Prophet time series models predict data change trends based on the time series patterns of historical data, and identify abnormal trends such as a sudden increase in the usage of a certain treatment item in a short period of time without medical evidence.

[0041] The quality score is divided into three levels from 0 to 100. The anomaly probability and feature matching degree output by the model are converted into a quantitative score. A score of 80-100 indicates excellent data quality, a score of 60-79 indicates that further manual verification is needed, and a score below 60 indicates an anomaly, indicating a clear problem. The grade markers are associated one-to-one with the hash chain nodes in step 2. The markers can be used to directly locate the flow link where the abnormal data first appears, providing initial clues for subsequent tracing.

[0042] The digital twin model for healthcare scenarios is constructed based on four core scenarios: outpatient treatment, inpatient care, insurance claims, and drug distribution. Its principle is to accurately map the business processes (such as outpatient registration, consultation, examination, and prescription), role permissions (doctors, nurses, and claims adjusters), and treatment logic constraints (such as the correspondence between symptoms and medications, and the rules for calculating claims amounts) of each scenario using digital modeling technology, forming a virtual scenario consistent with real healthcare operations. During pre-verification, standardized data simulates the entire process flow in the twin model. The system automatically compares the data with the scenario logic, identifies hidden anomalies such as diabetic patients not being prescribed hypoglycemic drugs, and discrepancies between claim amounts and treatment fee standards, and generates early warning reports to proactively mitigate risks.

[0043] Real-time stream processing employs an Apache Flink and Kafka architecture. Kafka serves as a message queue to achieve high-throughput data reception and caching, while Apache Flink, based on a streaming computing engine, performs real-time data processing, ensuring that data processing latency is controlled within milliseconds. The online learning model uses an incremental SVM algorithm, which works by dynamically updating model parameters based on newly incoming data without retraining the entire model, quickly adapting to changes in data distribution, and ensuring the accuracy of real-time anomaly detection.

[0044] The tiered intelligent alert system sets levels based on urgency and scope of impact. Urgency is prioritized according to factors such as life safety, financial risk, and statistical error. Scope of impact is defined by whether it's across institutions, within a single institution, or within a single department. These two factors combine to form Level 1 (urgent), Level 2 (important), and Level 3 (general) alerts. Alert content includes anomaly details (such as abnormal data fields and values), hash chain tracing nodes (the initial stage of the anomaly's flow), recommended processing paths (such as assignment to the attending physician for review), and responsible parties (such as the data collection agency), ensuring relevant personnel can quickly grasp the anomaly information and take targeted measures.

[0045] In Example 8, in step 5, the smart contract matches verification nodes based on a dual dimension of data type and anomaly level. Emergency treatment anomalies are automatically assigned to attending physicians and department heads, while large insurance claim anomalies are assigned to hospital finance personnel, insurance company underwriters, and regulatory agency specialists. The multimodal verification tool supports structured verification, unstructured verification, and AI-assisted verification. Structured verification involves checking boxes to confirm data authenticity, unstructured verification involves uploading supporting materials, including medical records and examination reports, and AI-assisted verification involves the system automatically matching similar cases and recommending verification conclusions. The dynamic consensus mechanism employs a single-node confirmation and multi-node supervision mode for ordinary anomalies, and a Byzantine fault-tolerant consensus mechanism for core-level data anomalies, requiring at least two-thirds of the participating nodes to complete digital signature confirmation. Data correction is only allowed to be performed by the original data provider. After correction, a new master hash and sub-hash are generated, and the old hash and modification record are associated and stored in the corresponding blockchain layer. The modification record includes the modifier, modification time, modification reason, and verification conclusion. Finally, the corrected quality score, grade mark, and modification trajectory traceability report are output.

[0046] In Example 9, the risk level and response level in step 6 correspond one-to-one. High risk corresponds to a Level 1 alarm, triggering an emergency response, immediately freezing the abnormal data flow, and pushing emergency consultation notifications to multiple parties. Risk tracing is based on the data fingerprint chain stored on the blockchain. By traversing the sub-hash in reverse, the operation records of the data at each flow node are traced. The operation records include the operator, operation time, and data status changes, generating a visual tracing timeline. Finally, a complete verification result is output, which includes the final data quality score, classification conclusion, anomaly handling details, risk tracing report, and blockchain storage compliance proof. The classification conclusion is either credible, pending review, or the anomaly has been corrected, realizing a closed-loop output of the entire process of collection, verification, correction, tracing, and conclusion.

[0047] By adopting the above technical solution, in step 5, the smart contract matches and verifies nodes based on a dual dimension of data type and anomaly level. Emergency treatment anomalies are automatically assigned to attending physicians and department directors, while large insurance claim anomalies are assigned to hospital finance, insurance company underwriters, and regulatory agency specialists. The multimodal verification tool supports structured verification, unstructured verification, and AI-assisted verification. Structured verification involves checking to confirm the authenticity of the data, unstructured verification involves uploading supporting materials, including medical records and examination reports, and AI-assisted verification involves the system automatically matching similar cases to recommend verification conclusions. The dynamic consensus mechanism employs a single-node confirmation and multi-node supervision mode for ordinary anomalies, and a Byzantine fault-tolerant consensus mechanism for core-level data anomalies, requiring at least two-thirds of the participating nodes to complete digital signature confirmation. Data correction is only allowed to be performed by the original data provider. After correction, a new master hash and sub-hash are generated, and the old hash and modification record are associated and stored in the corresponding blockchain layer. The modification record includes the modifier, modification time, modification reason, and verification conclusion. Finally, the corrected quality score, grade mark, and modification trajectory traceability report are output.

[0048] In step 6, risk levels and response levels are correlated one-to-one. High risk corresponds to a Level 1 alarm, triggering an emergency response that immediately freezes the flow of abnormal data and pushes emergency consultation notifications to multiple parties. Risk tracing is based on the data fingerprint chain stored on the blockchain. By traversing the sub-hash in reverse, the operation records of the data at each flow node are traced. The operation records include the operator, operation time, and data status changes, generating a visual tracing timeline. Finally, a complete verification result is output, which includes the final data quality score, classification conclusion, anomaly handling details, risk tracing report, and blockchain storage compliance proof. The classification conclusion is either credible, pending review, or the anomaly has been corrected, achieving a closed-loop output of the entire process of collection, verification, correction, tracing, and conclusion.

[0049] Example 10, Medical Channel Data Verification System, used to implement a medical channel data verification method. The system includes a data collection and standardization module, a hierarchical blockchain module, a federated learning module, a digital twin module, a dynamic consensus collaboration module, a risk response module, and a traceability and result output module. The data acquisition and standardization module is used for multi-source data acquisition, format adaptation, de-identification and encryption, and output of standardized data; The layered blockchain module is used for data classification, hash-based notarization, and smart contract deployment. The federated learning module is used for collaborative model training and data quality assessment across institutions; The digital twin module is used for scenario simulation pre-verification and real-time anomaly detection and alarm; The dynamic consensus collaboration module is used to verify node matching, multi-agency collaborative verification, and data correction. The risk response module is used to trigger corresponding responses based on risk levels; The traceability and result output module is used for risk traceability and output of complete verification results; the system adopts a microservice distributed architecture to support cross-organizational collaboration and full-process security control.

[0050] The following specific embodiments illustrate the implementation principle of the present invention: Based on a regional medical collaboration platform, the participants include hospitals, medical insurance agencies, insurance companies, pharmaceutical distribution companies, and regulatory departments. The system needs to verify data from all channels, including outpatient treatment, inpatient care, medical insurance claims, pharmaceutical distribution, medical IoT monitoring, and image diagnosis. The system adopts a microservice distributed architecture and is deployed in a hybrid cloud environment. Specific implementation steps

[0051] Step 1: Standardization of multimodal data acquisition and privacy enhancement; The system collects data from hospital electronic health records, community clinics, drug delivery, wearable device monitoring, and DICOM format images through various methods including APIs, file uploads, and streaming interfaces. An intelligent format adaptation engine automatically identifies proprietary data formats and maps them according to FHIR or HL7 standards, supporting custom business field extensions. Sensitive patient fields are anonymized using an ε-differential privacy mechanism, and encrypted using a combination of the national cryptographic algorithms SM4 and AES-256, transmitted via the TLS 1.3 protocol, outputting standardized encrypted data containing source identifiers, timestamps, and checksums.

[0052] Step 2: Layered blockchain evidence storage and traceability; Data is categorized into core, important, and ordinary levels based on sensitivity. Core data is written to a cross-institutional consortium blockchain, important data is written to an institution's private blockchain, and ordinary data is stored using a distributed ledger and hash digest notarization. A master hash and sub-hashes are generated based on the SHA-256 algorithm, linked together by the flow nodes to form a chain structure. Layered, dedicated smart contracts containing different logics are deployed to automatically associate data hashes for storage and verification.

[0053] Step 3: Quality assessment of federated learning collaborative AI; Each participant acts as a federated node, sharing only model parameters and encrypting and aggregating feature vectors using a secure aggregation algorithm. Multi-model collaborative verification: Isolation forest identifies numerical anomalies exceeding clinically reasonable thresholds; medical knowledge graph verifies logical conflicts; and time-series model analyzes unfounded trend anomalies. A quality score is output from 0-100, categorized as trustworthy, pending review, and abnormal, with the grade markers associated with hash chain nodes.

[0054] Step 4: Digital Twin and Real-time Anomaly Detection; Digital twin models are constructed based on outpatient services, inpatient care, insurance claims, and drug distribution scenarios to simulate data flow, identify hidden anomalies, and generate early warning reports. The Apache Flink and Kafka architecture enables millisecond-level real-time stream processing, and the incremental SVM algorithm dynamically adapts to data distribution, generating level one to three intelligent alerts based on urgency and impact, including anomaly details, source nodes, and handling suggestions.

[0055] Smart contracts are matched with verification nodes based on data type and anomaly level. Emergency room anomalies are assigned to attending physicians and department heads, while large claim anomalies are assigned to multiple verification personnel. Multimodal verification is achieved through structured selection, uploading of unstructured supporting materials, and AI-powered similar case recommendations. Ordinary anomalies use single-node confirmation plus multi-node supervision, while core anomalies employ a Byzantine fault-tolerant consensus mechanism, requiring signature confirmation from at least two-thirds of the nodes. After the original data provider corrects the data, the system generates a new hash, associates the old hash with the modification record for evidence storage, and outputs the corrected quality score, grading label, and source tracing report.

[0056] High-risk scenarios trigger a Level 1 alert, necessitating an emergency response and freezing the flow of abnormal data. Medium- and low-risk scenarios trigger regular and normal responses respectively, with processing completed within specified timeframes. Based on the blockchain data fingerprint chain, the system reverse-traverses sub-hashes to trace the operation records of each flow node, generating a visual traceability timeline. The final output includes a complete verification result containing a final quality score, grading conclusion, anomaly handling details, a traceability report, and compliance proof, achieving a closed-loop process.

[0057] The refined technical features of Examples 4-9 have been incorporated into the above steps. Multimodal access covers all types of data, layered evidence storage and hash chain rules are implemented, federated learning and multi-model evaluation standards are executed, digital twin scenarios and real-time detection parameters are implemented, and verification node matching, consensus mechanisms and risk response strategies are put into practice.

[0058] The data collection and standardization module outputs standardized data, the layered blockchain module ensures data credibility and traceability, the federated learning module completes cross-institutional quality assessment, the digital twin module realizes anomaly pre-verification and real-time alerts, the dynamic consensus collaboration module promotes multi-institutional verification and correction, the risk response module triggers differentiated handling, and the traceability and result output module outputs complete verification results. All modules work together to ensure efficient and compliant verification.

[0059] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for verifying data from medical channels, characterized in that, Includes the following steps: Step 1: Collect medical data from all channels through multimodal access, map it to a unified standard through intelligent format adaptation, and use differential privacy desensitization and hybrid encryption technology to protect data privacy, outputting standardized encrypted data; Step 2: Classify data according to sensitivity and write it into the consortium blockchain, private blockchain or distributed ledger respectively, generate a main hash and sub-hash chain structure to record the entire life cycle of data flow, and deploy layered exclusive smart contracts; Step 3: Implement cross-institutional model collaborative training based on the federated learning framework. Through anomaly detection, logic verification, trend analysis, and multi-model collaboration, output quality scores and grade labels. Step 4: Construct a digital twin model of the medical scenario to simulate data flow and pre-identify hidden anomalies. Combine a real-time stream processing architecture with an online learning model to generate hierarchical intelligent alerts.

2. The medical channel data verification method according to claim 1, characterized in that, It also includes step 5, which involves matching verification nodes through smart contracts, using a dynamic consensus mechanism to complete multi-institutional collaborative verification, updating the hash chain and recording the modification trajectory after securely correcting the data, and outputting the corrected quality score and grade mark.

3. The medical channel data verification method according to claim 2, characterized in that, It also includes step 6, which involves triggering the corresponding response level based on the risk level, and using the data fingerprint chain stored on the blockchain to trace back the source of the risk and generate a traceability timeline.

4. The medical channel data verification method according to claim 3, characterized in that: Step 1's multimodal access supports API, file upload, streaming interface, medical IoT devices, and the acquisition of DICOM format image data; intelligent format adaptation has a built-in adaptive parsing engine that automatically identifies the private data formats of medical institutions and maps them to the FHIR or HL7 international medical data standards, supporting custom business field extensions; differential privacy desensitization uses an ε-differential privacy mechanism to inject noise into sensitive patient fields, and hybrid encryption uses a combination of the national cryptographic algorithm SM4 and AES-256 encryption scheme. The output standardized encrypted data must include data source identifier, acquisition timestamp, and format check code.

5. The medical channel data verification method according to claim 4, characterized in that: In step 2, the data is categorized into core, important, and ordinary levels based on sensitivity. The core level corresponds to patient identity information, core diagnostic conclusions, and surgical records; the important level corresponds to treatment details, medication records, and claim amounts. Standard-level corresponding equipment operating parameters and non-sensitive statistical data; Core-level data is written to a cross-institutional consortium blockchain, important-level data is written to an institution's private blockchain, and ordinary-level data is stored using a distributed ledger and lightweight hash digest notarization. Both the main hash and sub-hashes are generated based on the SHA-256 algorithm. Sub-hashes are generated sequentially according to the flow nodes of the collection end, standardization end, verification end, and validation end, forming a chain-like traceability structure. Layered dedicated smart contracts include core-level contracts, important-level contracts, and ordinary-level contracts. Contracts automatically associate the hash storage and verification of data at the corresponding level.

6. The medical channel data verification method according to claim 5, characterized in that: In step 3, federated learning uses the FedAvg framework to achieve cross-institutional model collaborative training. Each participating institution, as a federated node, only shares model parameters and uses a secure aggregation algorithm to encrypt and summarize feature vectors. Multi-model collaboration specifically includes: using isolated forest or LSTM autoencoder to identify numerical anomalies, which refer to test results that exceed clinically reasonable thresholds; verifying logical conflicts based on authoritative medical knowledge graphs; and using ARIMA or Prophet time series models to analyze trend anomalies. The quality score is divided into three levels from 0 to 100: credible, pending review, and abnormal. The credible level corresponds to 80-100 points, the pending review level corresponds to 60-79 points, and the abnormal level corresponds to less than 60 points. The hierarchical markers are associated one by one with the hash chain nodes in step 2, clarifying the initial flow of abnormal data.

7. The medical channel data verification method according to claim 6, characterized in that: In step 4, the digital twin model of the medical scenario is constructed according to the core scenario. The scenario construction includes outpatient treatment, inpatient nursing, insurance claims and drug distribution, accurately mapping business processes, role permissions and treatment logic constraints. Pre-verification identifies hidden anomalies and generates early warning reports by simulating the entire data flow; real-time stream processing adopts the Apache Flink+Kafka architecture, and the online learning model uses the incremental SVM algorithm, which can dynamically adapt to new data distributions. The tiered intelligent alerts are divided into Level 1, Level 2, and Level 3 alerts based on urgency, financial risk relevance, statistical error relevance, and scope of impact. The alert content includes anomaly details, hash chain tracing nodes, recommended processing paths, and responsible parties.

8. The medical channel data verification method according to claim 7, characterized in that: In step 5, the smart contract matches verification nodes based on a dual dimension of data type and anomaly level. Emergency treatment anomalies are automatically assigned to attending physicians and department heads, while large insurance claim anomalies are assigned to hospital finance, insurance company underwriters, and regulatory agency specialists. The multimodal verification tool supports structured verification, unstructured verification, and AI-assisted verification. Structured verification involves checking to confirm the authenticity of the data, unstructured verification involves uploading supporting materials, including medical records and examination reports, and AI-assisted verification involves the system automatically matching similar cases to recommend verification conclusions. The dynamic consensus mechanism employs a single-node confirmation and multi-node supervision mode for ordinary anomalies, and a Byzantine fault-tolerant consensus mechanism for core-level data anomalies, requiring at least two-thirds of the participating nodes to complete digital signature confirmation. Data correction is only allowed to be performed by the original data provider. After correction, a new master hash and sub-hash are generated, and the old hash and modification record are associated and stored in the corresponding blockchain layer. The modification record includes the modifier, modification time, modification reason, and verification conclusion. Finally, the corrected quality score, grade mark, and modification trajectory traceability report are output.

9. The medical channel data verification method according to claim 8, characterized in that: In step 6, risk levels and response levels are correlated one-to-one. High risk corresponds to a Level 1 alarm, triggering an emergency response that immediately freezes the flow of abnormal data and pushes emergency consultation notifications to multiple parties. Risk tracing is based on the data fingerprint chain stored on the blockchain. By traversing the sub-hash in reverse, the operation records of the data at each flow node are traced. The operation records include the operator, operation time, and data status changes, generating a visual tracing timeline. Finally, a complete verification result is output, which includes the final data quality score, classification conclusion, anomaly handling details, risk tracing report, and blockchain storage compliance proof. The classification conclusion is either credible, pending review, or the anomaly has been corrected, achieving a closed-loop output of the entire process of collection, verification, correction, tracing, and conclusion.

10. A medical channel data verification system, characterized in that: To implement the medical channel data verification method of claim 9, the system includes a data acquisition and standardization module, a hierarchical blockchain module, a federated learning module, a digital twin module, a dynamic consensus collaboration module, a risk response module, and a traceability and result output module. The data acquisition and standardization module is used for multi-source data acquisition, format adaptation, de-identification and encryption, and output of standardized data; The layered blockchain module is used for data classification, hash-based notarization, and smart contract deployment. The federated learning module is used for collaborative model training and data quality assessment across institutions; The digital twin module is used for scenario simulation pre-verification and real-time anomaly detection and alarm; The dynamic consensus collaboration module is used to verify node matching, multi-agency collaborative verification, and data correction. The risk response module is used to trigger corresponding responses based on risk levels; The traceability and result output module is used for risk traceability and output of complete verification results; the system adopts a microservice distributed architecture to support cross-organizational collaboration and full-process security control.

Citation Information

Cited By

  • A big data-based medical education platform intelligent monitoring method and system

    CN122222474A