Medical data source real-time processing and quality control method based on distributed computing

Optimizing multi-source heterogeneous medical data processing through distributed computing and federated learning algorithms, solving the problems of data processing delay, quality and privacy protection, and achieving efficient and secure medical data processing and sharing.

CN120448756APending Publication Date: 2025-08-08BEIJING GUANXIN MEDICAL SOFTWARE TECH CO LTD

Patent Information

Application Number
CN202510963883.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently process multi-source heterogeneous medical data, and there are problems such as delay in processing, incomplete data quality, insufficient privacy protection and data sharing difficulties, which affect medical decision-making and patient safety.

Method used

A multi-dimensional processing space is built using a distributed computing method, a federated learning algorithm is used to optimize data processing, and combined with edge computing and dynamic task allocation to achieve real-time, efficient and secure data processing.

Benefits of technology

It realizes efficient real-time processing of multi-source heterogeneous medical data, improves data accuracy and reliability, ensures privacy and security, and promotes the sharing and research of medical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448756A_ABST
    Figure CN120448756A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical data processing, and discloses a medical data source real-time processing and quality control method based on distributed computing. The method comprises the steps of obtaining and classifying multi-source heterogeneous medical data streams, constructing a distributed parallel processing framework, setting an initial constraint condition, executing distributed streaming computation to generate a primary processing strategy, and optimizing and generating a global consistency processing strategy by using a federated learning algorithm. The system comprises a plurality of modules such as a data access and classification module and a processing framework construction module. According to the method, multi-source heterogeneous medical data can be efficiently processed, real-time calculation is realized through a multi-dimensional processing space and dynamic task allocation, and the data quality and privacy are guaranteed by utilizing a federal learning optimization strategy. And equipment abnormity can be monitored, quality risks can be predicted, task scheduling can be managed, the accuracy, reliability and availability of medical data processing can be improved, and powerful support can be provided for medical decision making and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical data processing technology, and in particular to a real-time processing and quality control method for medical data sources based on distributed computing. Background Art

[0002] In today's era of rapid digital healthcare development, medical data is experiencing explosive growth, originating from diverse sources and possessing a complex nature. Multi-source, heterogeneous medical data encompasses a wide range of forms, from vital signs and medical imaging data collected by various medical devices to text data in electronic medical record systems. This data varies significantly in format and structure, as well as collection frequency and privacy levels. This poses numerous challenges to its effective processing and utilization.

[0003] Traditional medical data processing methods are mostly centralized, making them incapable of handling the real-time processing needs of large-scale, multi-source, heterogeneous data. In this centralized processing model, data is aggregated and processed at a single center, which can easily lead to processing bottlenecks and high processing latency. This makes it impossible to meet the requirements of medical scenarios with extremely high real-time requirements, such as intensive care and surgical monitoring. In these scenarios, real-time and accurate analysis of patient vital signs and other data is crucial; even the slightest delay can affect timely diagnosis and treatment of the patient's condition.

[0004] At the same time, the quality of medical data varies widely. Some data may be missing, erroneous, or inconsistent, potentially due to malfunctions in data collection equipment, interference during transmission, or manual data entry errors. Low-quality data can seriously impact the accuracy of medical decisions, leading to misdiagnoses and missed diagnoses, posing potential risks to patients' health and safety. For example, during disease diagnosis, erroneous test data can lead doctors to incorrect conclusions and, in turn, inappropriate treatment plans.

[0005] Furthermore, the privacy protection of medical data is becoming increasingly prominent. As the value of medical data continues to be explored, the risk of data leakage also increases. Patient medical data contains a wealth of sensitive information, such as personal identity, health status, and medical history. Once leaked, it would seriously infringe upon patients' privacy and trigger a crisis of trust. However, existing data processing technologies often have shortcomings in privacy protection, making it difficult to achieve efficient processing while ensuring data security.

[0006] Furthermore, data sharing and collaborative processing between different medical institutions face difficulties. Inconsistent data standards and diverse data storage formats across institutions lead to compatibility issues during data sharing, making effective integration and analysis difficult. This restricts the development of medical research and hinders the overall improvement of medical standards. For example, large-scale epidemiological studies require data collection from multiple medical institutions, but data disparities make data integration and analysis extremely complex, or even impossible. Summary of the Invention

[0007] The purpose of the present invention is to provide a real-time processing and quality control method for medical data sources based on distributed computing to solve the problems raised in the above background technology.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a method for real-time processing and quality control of medical data sources based on distributed computing, the method comprising: Acquire multi-source heterogeneous medical data streams and classify the data into structured data task sets based on preset dimensions, including data type, collection frequency, privacy level, and data source credibility; Construct a distributed parallel processing framework and define a multi-dimensional processing space based on the classification dimension of the data task set. The processing space includes a timing axis, a quality assessment axis, a node load axis, and a data correlation axis. Setting initial constraints based on data cleaning rules and real-time processing latency requirements, including data integrity thresholds, processing node resource usage limits, and cross-node communication redundancy; Perform distributed streaming computing on data task sets based on a multi-dimensional processing space, simulate data interaction and load balancing among multiple nodes through a dynamic task allocation engine, and generate preliminary processing strategies; The federated learning algorithm is used to dynamically optimize the preliminary processing strategy, adjust the data sharding rules and the coordination parameters between computing nodes, and generate a globally consistent processing strategy.

[0009] Preferably, the building of a distributed parallel processing framework includes: The time axis is divided into sliding windows synchronized with the data acquisition period, and each window is associated with an adaptive sampling rate adjustment factor; Define a data anomaly scoring matrix based on the quality assessment axis and integrate multimodal data consistency verification rules; A hardware resource dynamic monitoring module is embedded in the node load axis to correlate memory usage with the computing task migration decision tree in real time.

[0010] Preferably, the federated learning algorithm adopts a sharded model training architecture, including: The data privacy level and feature distribution difference are encoded into a gradient mask vector, and an objective function is defined to evaluate the convergence speed of cross-node model aggregation and the risk of information leakage; Local model update parameters are generated through a dynamic weight distribution strategy and differential privacy mechanism, and global model fusion is completed using a multi-party secure computing protocol.

[0011] Preferably, the distributed streaming computing includes: Establish an event-driven computing topology network and model the data preprocessing module, feature extraction engine, and quality detection unit as stateless service nodes; A lightweight consistency protocol is used to achieve fast synchronization of intermediate results between nodes, and a breakpoint resumption mechanism is designed to ensure the atomicity of task execution.

[0012] Preferably, the method further comprises: Deploy edge computing agent nodes to capture clock deviations and sampling rate fluctuations of data acquisition devices in real time; The time series pattern matching algorithm is used to identify abnormal timing segments in the data stream and trigger the data backtracking and repair process.

[0013] Preferably, the dynamic optimization includes: Build a deep reinforcement learning model with node computing load, data flow throughput and processing latency as the state space; The agent is trained through an asynchronous advantage action evaluation algorithm to generate task scheduling strategies to minimize the degree of global resource fragmentation.

[0014] Preferably, the definition of the quality assessment axis also includes: Integrate the medical ontology library into the data quality assessment model and use the graph attention network to mine the semantic relevance of cross-modal data; Dynamically adjust the sensitivity threshold of anomaly detection and the confidence interval of quality assessment based on the credibility of the data source.

[0015] Preferably, the method further comprises: Build a typical quality defect pattern library based on historical medical data characteristics, extract high-dimensional missing value combinations and cross-device data conflict templates; Use spatiotemporal convolutional networks to predict potential quality risk points in data streams and pre-generate data interpolation and correction solutions.

[0016] Preferably, the method further comprises: Use dynamic priority queues to manage real-time data streams and offline batch processing tasks, and build a two-tier scheduling strategy based on data timeliness requirements; Configure dedicated processing channels and redundant verification mechanisms for key vital signs data.

[0017] Preferably, the present invention further includes a real-time processing and quality control system for medical data sources based on distributed computing, the system comprising: Data access and classification module: used to obtain multi-source heterogeneous medical data streams and classify the data into structured data task sets based on preset dimensions such as data type, collection frequency, privacy level, and data source credibility; Processing framework construction module: Build a distributed parallel processing framework and define a multi-dimensional processing space based on the classification dimensions of the data task set, including the time series axis, quality assessment axis, node load axis, and data correlation axis; Constraint configuration module: sets initial constraints based on data cleansing rules and real-time processing latency requirements. These constraints include data integrity thresholds, processing node resource usage limits, and cross-node communication redundancy. Streaming computing module: performs distributed streaming computing on data task sets based on a multi-dimensional processing space, simulates data interaction and load balancing among multiple nodes through a dynamic task allocation engine, and generates preliminary processing strategies; Strategy optimization module: Uses federated learning algorithms to dynamically optimize the initial processing strategy, adjusts data sharding rules and coordination parameters between computing nodes, and generates a globally consistent processing strategy.

[0018] Compared with the prior art, the present invention has the following beneficial effects: In terms of data processing efficiency, by building a distributed parallel processing framework, defining a multi-dimensional processing space including a timing axis, a quality assessment axis, a node load axis, and a data association axis, and using a dynamic task allocation engine to simulate data interaction and load balancing between multiple nodes, efficient real-time processing of multi-source heterogeneous medical data streams is achieved. Compared with traditional centralized processing, processing delays are greatly reduced. For example, in intensive care scenarios, it can quickly process the vital signs data continuously generated by patients, allowing medical staff to promptly grasp changes in the patient's condition and buy precious time for emergency treatment. By dividing the timing axis into sliding windows synchronized with the data acquisition cycle and associating them with an adaptive sampling rate adjustment factor, the sampling strategy can be dynamically adjusted according to data changes, reducing unnecessary data processing while ensuring data accuracy, further improving processing efficiency.

[0019] For data quality control, integrating the medical ontology library into the data quality assessment model and using a graph attention network to mine the semantic relevance of cross-modal data can more accurately assess data quality. Dynamically adjusting the sensitivity threshold of anomaly detection and the confidence interval of quality assessment based on the credibility of the data source can effectively identify and process abnormal data. Building a library of typical quality defect patterns based on historical medical data and combining it with a spatiotemporal convolutional network to predict potential quality risk points and pre-generate correction plans can prevent data quality issues in advance and ensure the accuracy and reliability of medical data. This helps reduce misdiagnoses and missed diagnoses caused by data quality issues and improve the overall quality of medical services.

[0020] To protect privacy, the federated learning algorithm employs a sharded model training architecture, encoding data privacy levels and feature distribution differences as gradient mask vectors. Local model update parameters are generated through a dynamic weight allocation strategy and differential privacy mechanisms, and global model fusion is achieved using a multi-party secure computation protocol. This series of measures enables effective model training and optimization while protecting data privacy, reducing the risk of data leakage, allowing patients to confidently share medical data, and promoting the development of medical research and services.

[0021] When addressing issues with data collection equipment, edge computing agent nodes are deployed to capture clock deviations and sampling rate fluctuations in real time. Time series pattern matching algorithms are used to identify anomalies and trigger a retrospective repair process, ensuring the accuracy and continuity of data collection. Even in the event of a brief equipment failure or anomaly, the impact on data quality can be minimized.

[0022] Furthermore, a dynamic priority queue is used to manage real-time data streams and offline batch processing tasks, establishing a two-tier scheduling strategy. Dedicated processing channels and redundant verification mechanisms are configured for critical vital signs data. This ensures prioritized processing and accuracy of critical data, improving the reliability and stability of the entire system. In emergency situations, critical vital signs data can be quickly and accurately transmitted and processed, providing strong protection for patient safety. Furthermore, this flexible task scheduling and data management approach improves system resource utilization, enabling the system to better adapt to diverse medical data processing needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a working principle diagram of the real-time processing and quality control method of medical data sources based on distributed computing according to the present invention; Figure 2 Schematic diagram of the process of distributed streaming computing execution; Figure 3 This is a flow chart of edge computing agent node data monitoring and repair; Figure 4Flowchart of dynamic optimization task scheduling based on reinforcement learning. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0025] See also Figures 1-4 The present invention provides a real-time processing and quality control method for medical data sources based on distributed computing, and the specific implementation steps are as follows: Acquire multi-source, heterogeneous medical data streams, which may originate from medical devices across different hospital departments, electronic medical record systems, and other sources. Classify the data into structured data task sets based on pre-defined dimensions: data type (e.g., imaging data, laboratory data, vital sign data), collection frequency (real-time, scheduled), privacy level (high, medium, low), and data source credibility (based on data source reliability assessment criteria). For example, vital sign data collected in real time by an ECG monitor is categorized into a specific structured data task set due to its data type, high collection frequency, medium privacy level, and high data source credibility.

[0026] A distributed parallel processing framework is built, defining a multidimensional processing space based on the categorization of data task sets. This processing space includes a time series axis, a quality assessment axis, a node load axis, and a data correlation axis. This multidimensional processing space allows for the management and processing of data task sets from multiple dimensions, improving processing efficiency and quality.

[0027] Initial constraints are set based on data cleansing rules and real-time processing latency requirements. The data integrity threshold determines data integrity. For example, a certain type of test data integrity threshold is set at 95%. If the integrity threshold is below this threshold, the data is considered to be missing. The processing node resource usage limit specifies the maximum amount of resources each processing node can occupy when processing tasks, preventing excessive node resource consumption. Cross-node communication redundancy ensures the reliability of inter-node communication, allowing communication redundancy within a certain range.

[0028] Distributed streaming computing is performed on data task sets based on a multidimensional processing space. A dynamic task allocation engine simulates data interaction and load balancing across multiple nodes. Tasks are rationally assigned to each node for processing based on their load and data characteristics, generating a preliminary processing strategy. For example, when a node has a low load, more computing tasks are assigned to it to fully utilize resources.

[0029] A federated learning algorithm is used to dynamically optimize the initial processing strategy. By adjusting data sharding rules and coordination parameters between computing nodes, each node can work better together while protecting data privacy. Ultimately, a globally consistent processing strategy is generated, ensuring that the entire system achieves optimal processing results for medical data.

[0030] The present invention will be further described below in conjunction with Examples 1 to 6: Example

[0031] When building a distributed parallel processing framework, the following specific designs are performed on each axis: The time-series axis is divided into sliding windows synchronized with the data acquisition cycle. For example, if the data acquisition cycle for a medical device is 1 minute, the sliding window duration is also set to 1 minute. Each window is associated with an adaptive sampling rate adjustment factor, which dynamically adjusts the sampling rate based on data changes. For example, when data changes drastically, the sampling rate is appropriately increased to obtain more accurate data; when data is relatively stable, the sampling rate is reduced to reduce the amount of data processing. In actual applications, if ECG data fluctuates significantly over a period of time, the adaptive sampling rate adjustment factor will increase the sampling rate from the default 100 data points per second to 200 data points per second.

[0032] Based on the quality assessment axis, a data anomaly scoring matrix is defined. This matrix is used to quantify the degree of data anomaly, with elements in the matrix assigned values based on different data characteristics and rules. Furthermore, multimodal data consistency verification rules are integrated. For example, when simultaneously collecting patient imaging and test data, verification rules are used to check whether the organ status displayed in the image matches the relevant indicators in the test data. If the image shows liver lesions, but the liver function indicators in the test data are normal, an anomaly alert is triggered.

[0033] A dynamic hardware resource monitoring module is embedded in the node load axis. This module monitors the node's memory usage in real time and correlates this with a decision tree for computing task migration. When the memory usage reaches a certain threshold (e.g., 80%), the decision tree determines, based on pre-set rules, whether some computing tasks need to be migrated to other, less-loaded nodes to ensure stable system operation. For example, the decision tree can determine whether some image data preprocessing tasks on the current node, which require less real-time performance, should be migrated to a backup node to free up memory resources.

[0034] Example 2: In the sharded model training architecture used by the federated learning algorithm, the following specific operations are performed to effectively protect data privacy and optimize model training results: The data privacy level and feature distribution differences are encoded into a gradient mask vector. Data privacy levels are categorized based on the sensitivity of the medical data. For example, medical data involving patient identity information and core disease diagnoses has a higher privacy level, while routine vital sign data has a relatively lower privacy level. Feature distribution differences are measured by calculating the degree of dispersion of data features at different nodes. For example, the standard deviation of a test indicator value in each node's data is calculated. The larger the standard deviation, the greater the difference in feature distribution. Based on this information, a gradient mask vector is generated. Its purpose is to process the gradient during model training, protecting sensitive data during gradient propagation.

[0035] Define an objective function to evaluate the convergence speed and information leakage risk of cross-node model aggregation. The objective function formula is: .in, Represents the objective function value, which comprehensively reflects the effect of model aggregation; is the weight coefficient, which takes the value It determines the cross-node model aggregation convergence speed and information leakage risks The relative importance of the objective function. When more attention is paid to the convergence speed of the model, it can be appropriately increased. If you pay more attention to data privacy protection, that is, to reduce the risk of information leakage, then reduce value. It is measured by calculating the changes in parameter updates during multiple iterations of the model, such as calculating the Euclidean distance between the model parameters after each iteration and the parameters of the previous iteration. The smaller the distance change, the closer the model is to the convergence state. The evaluation is performed based on factors such as data privacy level and gradient mask vector. The higher the data privacy level, the higher the information leakage risk assessment value under the same gradient mask vector.

[0036] The local model update parameters are generated through the dynamic weight allocation strategy and differential privacy mechanism. The dynamic weight allocation strategy assigns different weights to nodes based on factors such as the amount of data and data quality of each node. Nodes with larger data volumes mean they contain more information and may contribute more to model training, so they are given higher weights. In terms of data quality, the weights of high-quality nodes are also increased accordingly, as measured by indicators such as data integrity and accuracy. The differential privacy mechanism adds noise that conforms to a specific distribution, such as Laplace noise, to the model update parameters. Assuming the original model update parameters are , the noise added is , then the parameters after differential privacy processing are , The size of is determined based on factors such as the privacy budget to protect data privacy. Finally, a multi-party secure computation protocol is used to complete global model fusion. During this process, each node processes local model parameters through encryption and obfuscation techniques, preventing the central node from accessing the original data of each node during model fusion, thus ensuring data security.

[0037] Example 3: In the distributed streaming computing process, the following specific implementation methods are adopted: Establish an event-driven computing topology network. Model the data preprocessing module, feature extraction engine, and quality inspection unit as stateless service nodes. When new medical data flows into the system, corresponding events are triggered. After receiving the data, the data preprocessing module performs different preprocessing operations for different types of data. For imaging data, image denoising is performed to remove noise generated during device acquisition or transmission using algorithms such as Gaussian filtering. For text-based medical data, data cleaning is performed to remove typos, special symbols, and other interfering information.

[0038] The feature extraction engine extracts specific features based on the data type. For ECG data, it extracts features such as heart rate, PR interval, and QT interval; for medical imaging data, it extracts image texture and shape features. The quality inspection unit tests the data according to pre-set quality standards. For example, for test data, it determines whether various indicators are within normal reference ranges.

[0039] A lightweight consensus protocol is used to quickly synchronize intermediate results between nodes. This protocol improves synchronization efficiency by reducing unnecessary network communication overhead. For example, when synchronizing intermediate results, only key feature data is transmitted, rather than the entire dataset. Furthermore, hash verification methods are used to ensure data accuracy. At the receiving end, the hash value of the received data is calculated and compared with the hash value of the sending end. If they match, the data transmission is considered correct.

[0040] A breakpoint-resume mechanism is designed to ensure the atomicity of task execution. When a node fails or is interrupted while processing a task, the system records the task's progress. This includes information such as the amount of data processed and the time point of processing. After the node returns to normal, the system resumes the task from the breakpoint based on the recorded information. For example, when processing a long period of ECG data, if a node fails midway through processing, subsequent processing will resume from the last data point processed before the failure, avoiding reprocessing of already processed data and improving system reliability and processing efficiency.

[0041] Example 4: The entire real-time processing and quality control system for medical data sources also includes the following important links: Deploy edge computing agent nodes to capture clock deviations and sampling rate fluctuations of data acquisition devices in real time. Edge computing agent nodes are closely connected to various data acquisition devices, such as bedside ECG monitors and blood pressure monitors, via wired or wireless connections. They continuously monitor device clock information and sampling rate data, comparing them with a standard clock source to obtain clock deviation information. For sampling rate fluctuations, if the preset sampling rate of an ECG monitor is 128 data points per second, if the actual sampling rate deviates from this value by more than the allowable range (e.g., plus or minus 5%) within a certain period of time, the edge computing agent node will promptly record and report this information.

[0042] A time series pattern matching algorithm is used to identify time series anomalies in data streams. The algorithm first builds a time series pattern library based on a large amount of historically normal medical data. For example, patient temperature data typically exhibits a certain pattern of fluctuation throughout the day. The algorithm learns and records this pattern as a normal pattern. When the real-time temperature data sequence does not match the normal pattern, such as a sudden, sharp rise or fall in temperature that does not conform to normal physiological changes, it is identified as a time series anomaly, triggering a data backtracking and repair process. During this process, the system retrieves data from the relevant time period from historical data storage based on the time of the abnormal data for comparison and repair. If data is missing or erroneous due to a sampling device failure, data from other related devices within the same time period may be referenced, or data interpolation algorithms may be used to supplement and correct the data.

[0043] Based on the characteristics of historical medical data, a library of typical quality defect patterns is constructed to extract high-dimensional missing value combinations and cross-device data conflict templates. Through in-depth analysis of large amounts of historical medical data, such as comprehensive research on years of test reports, imaging data, and medical records, it is discovered that certain specific test items are frequently missing simultaneously in the test data. These high-dimensional missing value combinations are recorded in the pattern library. For cross-device data conflicts, such as when the blood pressure values measured on the same patient by different blood pressure measuring devices vary significantly, exceeding the reasonable error range, this situation is saved as a cross-device data conflict template. A spatiotemporal convolutional network is used to predict potential quality risk points in data streams. The spatiotemporal convolutional network combines information from both temporal and spatial dimensions to effectively analyze the time series characteristics of medical data and the spatial correlations between data from different devices. For example, by combining vital sign data collected by different devices over a period of time, the network can predict the likelihood of data quality issues and identify potential risks in advance. Once a potential quality risk point is predicted, the system will pre-generate a data interpolation and correction plan, such as using mean interpolation, linear regression interpolation and other methods to interpolate possible missing data. For possible erroneous data, corrections will be made based on historical data patterns and relevant medical knowledge.

[0044] Example 5: The following specific approach is used to dynamically optimize the initial processing strategy and refine the definition of the quality assessment axes: A deep reinforcement learning model is constructed, using node computational load, data flow throughput, and processing latency as the state space. Node computational load is measured by monitoring hardware resource metrics such as the node's CPU utilization and memory usage. For example, if CPU utilization exceeds 80% for a prolonged period, it indicates a high computational load on the node. Data flow throughput refers to the amount of data processed by the system per unit time and is calculated by counting the size or amount of medical data processed within a certain time interval. Processing latency is the time interval from data entry to completion of processing. The processing latency is calculated by accurately recording the timestamp of data entry and the timestamp of processing completion. These metrics are used as input to the deep reinforcement learning model, which continuously learns from the state of the environment to make optimal decisions.

[0045] The agent is trained using an asynchronous dominant action evaluation algorithm to generate task scheduling strategies. This algorithm allows the agent to learn and update at different time points, accelerating model training. During training, the agent attempts different task scheduling actions, such as assigning a data processing task to different compute nodes. It evaluates the effectiveness of each action by observing changes in state after each action, such as whether the node's computing load decreases, data flow throughput increases, and processing latency decreases. After repeated attempts and learning, the agent gradually generates a task scheduling strategy that minimizes global resource fragmentation. For example, when a system has multiple compute nodes and some are overloaded while others are idle, the agent learns to assign appropriate tasks to idle nodes, balancing the load across nodes and improving overall system performance.

[0046] Regarding the definition of quality assessment axes, a medical ontology library is integrated into the data quality assessment model. Medical ontology libraries contain a wealth of medical knowledge, such as disease diagnostic criteria and medical term definitions. Integrating this into the assessment model enables the model to leverage this knowledge to make more accurate judgments when assessing data quality. A graph attention network is used to mine semantic correlations in cross-modal data. The graph attention network can automatically learn important relationships between data modalities. For example, when processing patient imaging data and textual diagnosis data, the network can discover associations between lesion features in the images and disease descriptions in the textual diagnoses, enabling a more comprehensive assessment of data quality. The sensitivity threshold for anomaly detection and the confidence interval for quality assessment are dynamically adjusted based on the credibility of the data source. For data sources with higher credibility, the sensitivity threshold for anomaly detection is appropriately lowered to avoid misclassification of normal data; for data sources with lower credibility, the sensitivity threshold is increased to enhance data detection. Furthermore, the confidence interval for quality assessment is adjusted based on the credibility of the data source. For data sources with higher credibility, the confidence interval is narrower, indicating more reliable assessment results; for data sources with lower credibility, the confidence interval is wider, reflecting the uncertainty of the assessment results.

[0047] Example 6: In the process of data management and task scheduling, the following methods are used to ensure the efficiency and accuracy of medical data processing: Dynamic priority queues are used to manage real-time data streams and offline batch processing tasks, and a two-tier scheduling strategy is constructed based on data timeliness requirements. For real-time data streams, priority is assigned based on the data's urgency and importance. For example, vital sign data from patients undergoing surgery is directly related to the safe conduct of the surgery and therefore has the highest priority. Meanwhile, routine physical examination data collection has less timeliness requirements and therefore has a lower priority. In the first tier of this two-tier scheduling strategy, tasks are assigned to different queues based on data priority, with high-priority data entering the high-priority queue and low-priority data entering the low-priority queue. The second tier selects tasks from the priority queues based on the real-time availability of system resources. When system resources are sufficient, tasks in the high-priority queue are prioritized. When system resources are limited, low-priority tasks are appropriately scheduled while ensuring the processing of high-priority tasks.

[0048] Dedicated processing channels and redundant verification mechanisms are configured for critical vital sign data. Critical vital sign data, including heart rate, blood pressure, and blood oxygen saturation, are crucial to patient safety. Dedicated processing channels utilize independent hardware and network links to ensure fast and stable transmission and processing of this data. For example, in a hospital's intensive care unit, dedicated high-speed network lines are deployed for critical vital sign data to prevent conflicts and interference with other data. The redundant verification mechanism ensures data accuracy by performing multiple checks on critical vital sign data. Various verification methods, such as parity and CRC, are employed. If errors are detected during verification, the system promptly corrects or re-collects data. For example, when continuously monitoring a patient's heart rate, if a CRC check reveals an error in the heart rate data at a specific moment, the system automatically triggers a re-collection and compares and analyzes the data with previous heart rate data to ensure accurate data, providing a reliable basis for diagnosis and treatment.

[0049] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0050] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for real-time processing and quality control of medical data sources based on distributed computing, characterized in that: include: Acquire multi-source heterogeneous medical data streams and classify the data into structured data task sets based on preset dimensions, including data type, collection frequency, privacy level, and data source credibility; Construct a distributed parallel processing framework and define a multi-dimensional processing space based on the classification dimension of the data task set. The processing space includes a timing axis, a quality assessment axis, a node load axis, and a data correlation axis. Setting initial constraints based on data cleaning rules and real-time processing latency requirements, including data integrity thresholds, processing node resource usage limits, and cross-node communication redundancy; Perform distributed streaming computing on data task sets based on a multi-dimensional processing space, simulate data interaction and load balancing among multiple nodes through a dynamic task allocation engine, and generate preliminary processing strategies; The federated learning algorithm is used to dynamically optimize the preliminary processing strategy, adjust the data sharding rules and the coordination parameters between computing nodes, and generate a globally consistent processing strategy.

2. The method for real-time processing and quality control of medical data sources according to claim 1, characterized in that: The construction of the distributed parallel processing framework includes: The time axis is divided into sliding windows synchronized with the data acquisition period, and each window is associated with an adaptive sampling rate adjustment factor; Define a data anomaly scoring matrix based on the quality assessment axis and integrate multimodal data consistency verification rules; A hardware resource dynamic monitoring module is embedded in the node load axis to correlate memory usage with the computing task migration decision tree in real time.

3. The method for real-time processing and quality control of medical data sources according to claim 1, characterized in that: The federated learning algorithm adopts a sharded model training architecture, including: The data privacy level and feature distribution difference are encoded into a gradient mask vector, and an objective function is defined to evaluate the convergence speed of cross-node model aggregation and the risk of information leakage; Local model update parameters are generated through a dynamic weight distribution strategy and differential privacy mechanism, and global model fusion is completed using a multi-party secure computing protocol.

4. The method for real-time processing and quality control of medical data sources according to claim 1, characterized in that: The distributed streaming computing includes: Establish an event-driven computing topology network and model the data preprocessing module, feature extraction engine, and quality detection unit as stateless service nodes; A lightweight consistency protocol is used to achieve fast synchronization of intermediate results between nodes, and a breakpoint resumption mechanism is designed to ensure the atomicity of task execution.

5. The real-time processing and quality control method for medical data sources according to claim 1, characterized in that: The method further comprises: Deploy edge computing agent nodes to capture clock deviations and sampling rate fluctuations of data acquisition devices in real time; The time series pattern matching algorithm is used to identify abnormal timing segments in the data stream and trigger the data backtracking and repair process.

6. The method for real-time processing and quality control of medical data sources according to claim 1, characterized in that: The dynamic optimization includes: Build a deep reinforcement learning model with node computing load, data flow throughput and processing latency as the state space; The agent is trained through an asynchronous advantage action evaluation algorithm to generate task scheduling strategies to minimize the degree of global resource fragmentation.

7. The method for real-time processing and quality control of medical data sources according to claim 2, characterized in that: The definition of the quality assessment axis also includes: Integrate the medical ontology library into the data quality assessment model and use the graph attention network to mine the semantic relevance of cross-modal data; Dynamically adjust the sensitivity threshold of anomaly detection and the confidence interval of quality assessment based on the credibility of the data source.

8. The method for real-time processing and quality control of medical data sources according to claim 1, characterized in that: The method further comprises: Build a typical quality defect pattern library based on historical medical data characteristics, extract high-dimensional missing value combinations and cross-device data conflict templates; Use spatiotemporal convolutional networks to predict potential quality risk points in data streams and pre-generate data interpolation and correction solutions.

9. The method for real-time processing and quality control of medical data sources according to claim 1, characterized in that: The method further comprises: Use dynamic priority queues to manage real-time data streams and offline batch processing tasks, and build a two-tier scheduling strategy based on data timeliness requirements; Configure dedicated processing channels and redundant verification mechanisms for key vital signs data.

10. A real-time processing and quality control system for medical data sources based on distributed computing, characterized in that: include: Data access and classification module: used to obtain multi-source heterogeneous medical data streams and classify the data into structured data task sets based on preset dimensions such as data type, collection frequency, privacy level, and data source credibility; Processing framework construction module: Build a distributed parallel processing framework and define a multi-dimensional processing space based on the classification dimensions of the data task set, including the time series axis, quality assessment axis, node load axis, and data correlation axis; Constraint configuration module: sets initial constraints based on data cleansing rules and real-time processing latency requirements. These constraints include data integrity thresholds, processing node resource usage limits, and cross-node communication redundancy. Streaming computing module: performs distributed streaming computing on data task sets based on a multi-dimensional processing space, simulates data interaction and load balancing among multiple nodes through a dynamic task allocation engine, and generates preliminary processing strategies; Strategy optimization module: Uses federated learning algorithms to dynamically optimize the initial processing strategy, adjusts data sharding rules and coordination parameters between computing nodes, and generates a globally consistent processing strategy.

Citation Information

Patent Citations

  • Task processing optimization system based on cloud computing and medical big data

    CN111694651A

  • Federal learning method based on EMD distance fusion multi-source heterogeneous data

    CN113139603A

  • Operation business risk analysis method and device, equipment and storage medium

    CN119204677A

  • Multi-source heterogeneous data processing method, device and system

    CN119311375A

  • Power grid real-time data reasoning analysis system and method based on Rete algorithm optimization

    CN119917253A

Cited By

  • Medical insurance data management method and system based on cloud platform

    CN121544404A

  • Inflammation signal expression profile fused nerve injury prognosis prediction system and method

    CN121565473A

  • A system and method for predicting the prognosis of neurological injury by fusing inflammatory signal expression profiles.

    CN121565473B

  • Multi-source heterogeneous data integration and intelligent quality control system and method for community medical treatment

    CN122245818A