Data processing method and system for artificial intelligence driven medical diagnosis and treatment

The AI-driven medical diagnosis and treatment system solves the problem of integrating and analyzing multi-source heterogeneous medical data, enabling personalized health risk prediction and treatment recommendations, and improving the accuracy and efficiency of diagnosis and treatment.

CN120998475BActive Publication Date: 2026-08-04FUJIAN PROVINCIAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FUJIAN PROVINCIAL HOSPITAL
Filing Date
2025-10-24
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

The heterogeneity and complexity of medical data make it difficult for doctors to make comprehensive and accurate diagnoses and develop personalized treatment plans. Existing technologies are unable to effectively integrate and analyze heterogeneous medical data from multiple sources, leading to misdiagnosis, missed diagnosis, and poor treatment outcomes.

Method used

The AI-driven medical diagnosis and treatment system includes modules for multimodal medical data acquisition, data cleaning and standardization, deep extraction of medical features, construction of a multidimensional health status space, real-time medical data fusion, and personalized health risk prediction. It generates individualized treatment recommendations through technologies such as deep convolutional neural networks and temporal attention mechanisms.

Benefits of technology

It enables seamless integration and efficient utilization of multi-source medical data, improving diagnostic accuracy and personalized treatment effectiveness, and reducing the suffering and waste of resources caused by misdiagnosis and unsuitable treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998475B_ABST
    Figure CN120998475B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of medical data processing system, and particularly relates to a data processing method and system for medical diagnosis and treatment driven by artificial intelligence. The present application comprises a multi-modal medical data acquisition module for acquiring and processing multi-source heterogeneous medical data of a patient to generate a standardized data set; a medical feature deep extraction module for performing multi-dimensional feature extraction on the data set to construct a dynamic evolution feature matrix; a multi-dimensional health state space construction module for constructing a multi-dimensional health state space of the patient according to the matrix to determine a key medical early warning index set; a real-time medical data fusion module for mapping real-time data to the space to generate real-time health risk factors; and a medical risk prediction and decision module for establishing a personalized model according to the real-time health risk factors to output a disease occurrence probability and generate an individualized treatment suggestion. The system solves the problems of difficult medical data processing, incomplete diagnosis analysis, and lack of individualization in treatment schemes, and improves the accuracy of diagnosis and the pertinence of treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data processing system technology, and in particular to artificial intelligence-driven data processing methods and systems for medical diagnosis and treatment. Background Technology

[0002] In today's medical field, medical data is experiencing explosive growth. With the widespread use of various medical devices and the deep application of information technology in the healthcare industry, a massive amount of medical data resources have been formed, ranging from patient medical records and laboratory reports to physiological data collected in real time by wearable devices. However, this data is characterized by its multi-source and heterogeneous nature; different medical institutions and different devices generate data with different formats and standards, which poses significant challenges to the effective utilization of medical data.

[0003] Traditional medical diagnosis and treatment rely primarily on the individual experience and expertise of doctors. When faced with complex conditions, doctors need to comprehensively analyze various aspects of the patient's information, including symptom descriptions, medical history, and test results. However, due to the complexity and dispersion of data, doctors struggle to fully and accurately grasp the patient's health status, leading to misdiagnosis or missed diagnosis. For example, diagnosing rare or complex diseases may require consulting a large amount of medical literature and past cases, but doctors often lack the time to conduct a comprehensive search and analysis, resulting in low diagnostic efficiency.

[0004] In the past, treatment plans were often formulated using relatively universal strategies, failing to fully consider the individual differences of each patient. Different patients, due to variations in genes, lifestyle habits, and physical conditions, may respond vastly to the same treatment method. For example, certain drugs may be highly effective for some patients, but may cause serious adverse reactions in others. Furthermore, with the continuous advancement of medical technology, new treatment methods and drugs are constantly emerging, making it difficult for doctors to quickly and accurately select the most suitable treatment plan for each patient.

[0005] In the process of processing medical data, data quality is a critical issue that cannot be ignored. Raw medical data often contains noise, missing values, and mislabeled information, which seriously affects its analytical and utilization value. Moreover, due to the lack of effective data cleaning and standardization methods, data from different sources is difficult to integrate and compare, limiting the application of medical data in clinical decision-making, medical research, and other fields. Against this backdrop, there is an urgent practical need to develop a data processing system capable of efficiently processing multi-source heterogeneous medical data to achieve accurate diagnosis and personalized treatment. Summary of the Invention

[0006] The main objective of this invention is to provide a data processing method and system for medical diagnosis and treatment driven by artificial intelligence, aiming to solve the technical problems in the prior art.

[0007] This invention proposes an artificial intelligence-driven data processing system for medical diagnosis and treatment, comprising:

[0008] The multimodal medical data acquisition module is used to acquire multi-source heterogeneous medical data of patients within a preset time period, and to perform data cleaning and standardization processing on the multi-source heterogeneous medical data to generate a standardized multimodal medical dataset.

[0009] The medical feature deep extraction module is used to extract multi-dimensional features from the standardized multimodal medical dataset to obtain an initial medical feature set, and to construct a dynamic evolution feature matrix based on the time-varying patterns of each feature in the initial medical feature set.

[0010] The multidimensional health status space construction module is used to construct a multidimensional spatial representation of the patient's health status based on the dynamic evolution feature matrix, calculate the distribution density of historical medical events in the multidimensional space, and determine a set of key medical early warning indicators based on the distribution density.

[0011] The real-time medical data fusion module is used to receive real-time collected patient medical data, extract real-time medical features and form real-time feature vectors, map the real-time feature vectors to the multi-dimensional health state space, and generate real-time health risk factors.

[0012] The medical risk prediction and decision-making module is used to establish a personalized health risk prediction model based on the real-time health risk factors, the initial medical feature set, and the dynamic evolution feature matrix, output the probability of disease occurrence, and generate individualized treatment suggestions.

[0013] Preferably, the multimodal medical data acquisition module is specifically used for:

[0014] Collect multi-source heterogeneous medical data, including imaging data, time-series physiological parameter data, laboratory test data, and clinical text data;

[0015] The multi-source heterogeneous medical data is subjected to outlier removal, noise filtering, and missing value imputation to enhance data quality.

[0016] The enhanced multi-source heterogeneous medical data is aligned and normalized according to data modality to generate a standardized multimodal medical dataset.

[0017] Preferably, the medical feature depth extraction module is specifically used for:

[0018] A deep convolutional neural network is used to extract multi-scale medical features from the standardized multimodal medical dataset;

[0019] Medical features at different scales are fused and spliced ​​together to form an initial medical feature set;

[0020] Calculate the rate of change and trend of change of each feature in the initial medical feature set within a specified time window;

[0021] The rate of change and trend of change of each feature are organized according to time series to construct a dynamic evolution feature matrix.

[0022] Preferably, the multidimensional health state space construction module is specifically used for:

[0023] The number of dimensions of the multidimensional health state space is determined based on the number of feature dimensions of the dynamic evolution feature matrix.

[0024] Calculate the distribution density of the corresponding features at the time of historical diagnostic events in the multidimensional health state space;

[0025] A density threshold is set based on the distribution density, and features whose distribution density exceeds the density threshold are selected as the key medical early warning indicator set.

[0026] Preferably, the real-time medical data fusion module is specifically used for:

[0027] Feature extraction is performed on real-time collected patient medical data to generate real-time medical features;

[0028] The real-time medical features are constructed into real-time feature vectors according to the organization format of the dynamic evolution feature matrix;

[0029] The real-time feature vector is projected onto the multidimensional health state space, and its spatial distance with the set of key medical early warning indicators is calculated.

[0030] Real-time health risk factors are generated based on the spatial distance.

[0031] Preferably, the medical risk prediction and decision-making module establishes a personalized health risk prediction model as follows:

[0032] The real-time health risk factors are fused with the initial medical feature set;

[0033] The dynamic evolution feature matrix is ​​reconstructed using a temporal attention mechanism with weights.

[0034] A risk prediction network is trained based on the fused features and the reconstructed feature matrix.

[0035] Output the probability value of the patient's disease occurrence and generate a personalized treatment suggestion plan.

[0036] Preferably, the system further includes:

[0037] The medical knowledge graph construction module is used to integrate medical knowledge bases and clinical guidelines to build a medical knowledge graph that includes the relationships between diseases, symptoms, drugs, and treatment plans.

[0038] The medical risk prediction and decision-making module is also used to match and verify the individualized treatment recommendations with the medical knowledge graph.

[0039] Preferably, the system further includes:

[0040] The multi-center medical data collaboration module is used to aggregate standardized multimodal medical datasets from multiple medical institutions, perform federated learning training, and update the personalized health risk prediction model.

[0041] The medical risk prediction and decision-making module is also used to make predictions using an updated personalized health risk prediction model.

[0042] Preferably, the system further includes:

[0043] The diagnosis and treatment process traceability module is used to record the implementation process and effect feedback of the individualized treatment recommendations, forming closed-loop medical data;

[0044] The medical risk prediction and decision-making module is also used to continuously optimize the personalized health risk prediction model based on the closed-loop medical data.

[0045] Preferably, the present invention also includes an artificial intelligence-driven data processing method for medical diagnosis and treatment, the method comprising all modules and method flows of the artificial intelligence-driven data processing system for medical diagnosis and treatment described above.

[0046] The beneficial effects of this invention are as follows:

[0047] The system, through its multimodal medical data acquisition module, can acquire multi-source, heterogeneous medical data from patients within a preset time period, and perform data cleaning and standardization to generate a standardized multimodal medical dataset. This solves the problem of traditional medical data being difficult to integrate and utilize due to inconsistent formats and standards. For example, previously, medical records from different hospitals had different formats, and the units and reference ranges for laboratory indicators were also inconsistent, making it extremely difficult for doctors to analyze patient data from different hospitals. This system can unify and standardize this complex and diverse data, providing a reliable foundation for subsequent analysis and enabling truly seamless integration and efficient utilization of medical data.

[0048] The medical feature deep extraction module extracts multi-dimensional features from a standardized multimodal medical dataset to obtain an initial medical feature set, and constructs a dynamically evolving feature matrix based on the changes in these features over time. Compared to traditional single-dimensional or static feature analysis, this provides a more comprehensive and dynamic reflection of a patient's health status. For example, in monitoring the development of a patient's chronic disease, traditional methods may only focus on fixed indicators at a few time points, while this system, through the dynamic evolving feature matrix, can clearly present the continuous changing trends of various indicators over time. Doctors can more accurately determine the stage and pattern of disease progression and detect potential health risks in advance.

[0049] The multidimensional health status space construction module constructs a multidimensional spatial representation of the patient's health status based on a dynamic evolution feature matrix, calculates the distribution density of historical medical events in the multidimensional space, and determines a set of key medical early warning indicators. This innovative multidimensional space construction method breaks through the limitations of traditional two-dimensional or simple dimensional analysis. Taking the diagnosis of cardiovascular diseases as an example, traditional methods may only rely on a few indicators such as blood pressure and heart rate, while this system comprehensively considers numerous factors in the multidimensional space, including the patient's age, family medical history, lifestyle habits, multiple blood indicators, and the dynamic changes of these indicators. This greatly improves the accuracy and comprehensiveness of diagnosis, enabling more precise identification of key medical early warning indicators and providing strong support for early disease warning.

[0050] The real-time medical data fusion module receives real-time patient medical data, extracts real-time medical features and forms real-time feature vectors, which are then mapped to a multi-dimensional health status space to generate real-time health risk factors. This enables doctors to promptly grasp the patient's current health risk status. For example, during a patient's hospitalization, various monitoring devices collect data in real time, and the system can quickly transform this data into real-time health risk factors. Doctors can then adjust treatment plans promptly based on these factors to address sudden changes in health status and prevent the condition from worsening.

[0051] The medical risk prediction and decision-making module establishes a personalized health risk prediction model based on real-time health risk factors, an initial medical feature set, and a dynamic evolutionary feature matrix. This model outputs the probability of disease occurrence and generates individualized treatment recommendations. Compared to traditional general treatment plans, this system fully considers the unique circumstances of each patient. Taking cancer treatment as an example, different patients have different genetic characteristics and physical tolerance levels. Traditional treatment plans are often "one-size-fits-all," resulting in poor efficacy. However, this system, through its personalized health risk prediction model, can accurately calculate the probability of disease occurrence for each cancer patient and formulate the most suitable treatment plan for them, greatly improving treatment effectiveness, reducing unnecessary waste of medical resources, and minimizing the suffering and side effects caused by unsuitable treatment plans. Attached Figure Description

[0052] Figure 1 This is a timing diagram of the AI-driven medical diagnosis and treatment data processing system described in this invention.

[0053] Figure 2 This is a flowchart for feature extraction and evolution matrix construction.

[0054] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0055] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0056] like Figure 1 As shown, this application provides an AI-driven data processing system for medical diagnosis and treatment, comprising: acquiring multi-source heterogeneous medical data of patients within a preset time period through a multimodal medical data acquisition module, which performs data cleaning and standardization to generate a standardized multimodal medical dataset; a medical feature deep extraction module subsequently extracting multi-dimensional features from the dataset to generate an initial medical feature set, and constructing a dynamically evolving feature matrix based on the changing patterns of each feature over time; a multi-dimensional health state space construction module using this matrix to establish a multi-dimensional spatial representation of the patient's health state, and determining a set of key medical early warning indicators by calculating the distribution density of historical medical events in the space; a real-time medical data fusion module receiving real-time acquired patient medical data, extracting real-time medical features to form a real-time feature vector, and mapping this vector to the multi-dimensional health state space to generate real-time health risk factors; and a medical risk prediction and decision-making module combining real-time health risk factors, the initial medical feature set, and the dynamically evolving feature matrix to establish a personalized health risk prediction model, outputting the probability of disease occurrence and generating individualized treatment recommendations. The entire system achieves efficient data processing and intelligent decision support through modular design.

[0057] Example 1: See Figure 2The multimodal medical data acquisition module executes a comprehensive process for acquiring and preprocessing patient data. This module connects to CT / MRI equipment, ECG monitors, laboratory information systems, and electronic medical record systems via digital interfaces within medical institutions, acquiring multi-source heterogeneous medical data covering imaging data, time-series physiological parameter data, laboratory test data, and clinical text data. Imaging data includes DICOM format computed tomography images and 3D reconstruction data; time-series physiological parameter data involves continuously monitored vital signs such as heart rate, blood pressure, and blood oxygen saturation; laboratory test data covers blood biochemical indicators and microbial culture results; and clinical text data includes physician-written medical records and nursing documents. The data acquisition time period is set to the patient's most recent 12 months of historical records to ensure clinical relevance across the time span. Immediately after acquisition, a data cleaning process is initiated: outlier removal uses a quantile-based outlier detection algorithm to automatically mark and remove data points exceeding medically reasonable ranges; noise filtering applies a Butterworth low-pass filter to physiological signals to eliminate power frequency interference and motion artifacts; and missing value imputation uses the K-nearest neighbor algorithm to intelligently fill in similar patterns in similar patient data. After data quality enhancement, the standardization phase begins: time alignment is performed using the R-wave peak of the ECG signal as the reference point, synchronizing the timestamps of other modal data; normalization processing adjusts the window width and window level for imaging data, and uses log transformation to eliminate dimensional differences for laboratory data, ultimately generating a standardized multimodal medical dataset with unified dimensions and storing it in the medical database.

[0058] The medical feature deep extraction module performs deep analysis on the standardized dataset. This module is configured with a multi-stream convolutional neural network architecture. The image data processing branch uses the 3DResNet-50 network to extract ground-glass opacities and lesion volume changes in lung CT images. The physiological signal branch uses a 1DCNN combined with LSTM layers to capture abnormal patterns of the RR interval in heart rate variability. The test data branch identifies key indicators such as the albumin / globulin ratio through a fully connected network. The text branch uses a BERT model to extract entity features such as "dyspnea" and "fever" from clinical records. The feature extraction process generates feature maps at four scales: 128×128 pixels for macroscopic tissue features, 64×64 pixels for vascular distribution features, 32×32 pixels for cellular features, and 16×16 pixels for microstructural features. Multi-scale feature fusion adopts a channel attention mechanism. After calculating the weights of each channel through the SE module, the features are weighted and concatenated to form an initial medical feature set containing image texture features, physiological fluctuation features, test numerical features, and textual semantic features. This feature set is further subjected to dynamic evolution analysis: for time-series features such as blood glucose levels, the rate of change (first derivative) and trend of change (second derivative) within a 24-hour sliding window are calculated; for tumor volume features, a 72-hour window is used to analyze the growth rate. The dynamic attributes of all features are organized into a dynamic evolution feature matrix with minute-level timestamps. The row vectors of the matrix represent hourly feature snapshots, and the column vectors store triples of feature values, rates of change, and trend directions. Finally, a 1200×150-dimensional spatiotemporal feature matrix is ​​constructed to comprehensively characterize the disease evolution trajectory.

[0059] The construction of the feature matrix involves rigorous medical logic verification. Dynamic analysis of imaging features is cross-checked with lesion progression reports annotated by radiologists, while the trend calculation of physiological features references the fluctuation threshold standards published by the American Heart Association. Laboratory feature change rate calibration adopts the allowable total error range of the Clinical Laboratory Standards Institute (CLSI) to ensure that the feature extraction results meet evidence-based medicine requirements. The dynamically evolving feature matrix is ​​transmitted to the hospital's private cloud platform via a medical-specific encryption protocol, enabling real-time updates and rapid retrieval within an in-memory computing framework, providing high-precision spatiotemporal feature support for subsequent health status modeling. Strict adherence to HIPAA medical privacy protection guidelines ensures that patient identity information is anonymized before feature extraction. The data processing server is deployed in the hospital's local computer room, physically isolated from external networks, and all data transmission uses AES-256 encryption to guarantee information security. The training data for the feature extraction model comes from anonymized medical datasets approved by the ethics committee, and the model weights are iteratively updated quarterly based on the latest clinical research findings to ensure the scientific rigor and timeliness of the medical features.

[0060] Example 2: The multidimensional health state space construction module processes the dynamic evolution feature matrix. This module analyzes the dimensional structure of the matrix to determine the mathematical representation of the health state space. The dynamic evolution feature matrix contains 150 feature dimensions corresponding to the definition of spatial coordinate axes. The space construction adopts a Euclidean geometric model, with each coordinate axis corresponding to the evolution trajectory of a specific medical feature. For example, the X-axis maps the trend of blood glucose change rate, the Y-axis represents the tumor volume growth rate, and the Z-axis is associated with the fluctuation amplitude of inflammatory factors. The mapping process of historical medical events extracts timestamps from the diagnostic record database of the hospital information system, including the occurrence time of critical events such as myocardial infarction and diabetic ketoacidosis. The feature vectors of the corresponding time points are located in the dynamic evolution feature matrix as spatial coordinate points. The distribution density calculation deploys a Gaussian kernel density estimation algorithm to generate a probability density distribution field centered on each historical event coordinate point. The bandwidth parameter is adaptively adjusted according to the standard deviation of the feature values. The density calculation covers the entire 150-dimensional space to form a probability cloud map.

[0061] Spatial density analysis employs a hierarchical scanning strategy, dividing the 150-dimensional space into 15 clinical subsystems, including cardiovascular, metabolic, and oncology subsystems, based on medical significance. The density distribution of each subsystem is calculated independently. Density thresholds are set using ternary statistical analysis, taking the 75th percentile of each subsystem's density value as the critical line. An early warning flag is triggered when the density of a specific region exceeds the critical value for that subsystem. The selection of key medical early warning indicators is combined with a clinical expert rule base. For example, in the cardiovascular subsystem, the feature combination of "ST segment elevation density > threshold" and "troponin rise slope density > threshold" is included in the early warning set. The early warning indicator set is stored in a tree structure, with the root node representing the disease classification (e.g., respiratory diseases) and leaf nodes storing specific feature threshold conditions (e.g., blood oxygen saturation change rate > 5% / hour). A real-time verification mechanism is introduced during spatial construction. When new historical events continuously occur, the system automatically triggers spatial reconstruction: the coordinates of new events are added to the point set, and the density distribution is recalculated. If this causes a change in the density field morphology of a subsystem exceeding a preset tolerance (e.g., density peak offset > 10%), the subsystem boundary is redefined, and the threshold is updated. The dynamic maintenance of the early warning indicator set is achieved through version control. A new version of the indicator set is generated after each spatial reconstruction, while the old version is retained in the medical records for retrospective analysis of treatment efficacy. All spatial coordinate data is stored using double-precision floating-point numbers, and the density field is archived in a compressed voxel grid format, supporting spatiotemporal range retrieval using a medical-specific query language.

[0062] The multidimensional health status space is clinically validated and integrated with the hospital's clinical decision support system. When the spatial model identifies high-risk areas, it automatically triggers the medical order review process. For example, if the oncology subsystem detects that a patient's feature vector is located in a high-density area of ​​bone metastasis, the system will push a CT bone window review recommendation to the attending physician's workstation. The spatial model is evaluated quarterly by the clinical audit committee, which adjusts the density algorithm parameters by comparing the consistency between the early warning indicator set and actual disease records to ensure that the model's predictions meet the requirements of evidence-based medicine. The spatial data visualization component generates a 3D heat map for multidisciplinary consultations, allowing physicians to observe the relative positions of patient feature points in the historical event density field through virtual reality devices. The module's computing resource allocation adopts a medical priority scheduling strategy, automatically preempting GPU cluster resources for real-time spatial mapping when emergency patient data is input. Historical data processing utilizes idle computing power at night for batch spatial reconstruction, and all computation processes are recorded on medical blockchain nodes for traceability. Spatial model update packages are distributed to various clinical terminals through a secure channel within the hospital's intranet, maintaining the continuity of early warning services during updates and ensuring that real-time monitoring in the intensive care unit is not affected by system upgrades.

[0063] A tertiary-level Class A hospital's cardiology department admitted a 58-year-old male patient with a 10-year history of hypertension and a 5-year history of diabetes. He was admitted due to persistent chest pain for 2 hours. The medical feature deep extraction module generated a dynamically evolving feature matrix from the patient's multimodal data within 72 hours of admission. This matrix contained a 125-dimensional feature time series, with data collected every 30 minutes, forming a feature matrix of 150 time points. The matrix included key indicators such as ST-segment elevation on electrocardiogram (feature dimensions 1-5), troponin I rise slope (dimensions 6-10), heart rate variability SDNN value (dimensions 11-15), and blood pressure variability coefficient (dimensions 16-20).

[0064] The multidimensional health status space construction module first analyzes the dimensional structure of the feature matrix to determine the construction of a 125-dimensional health status space. Each dimension corresponds to the temporal evolution trajectory of a medical feature; for example, dimension 1 specifically maps the ST-segment elevation amplitude in the anterior leads, and dimension 6 records the hourly relative change rate of troponin I. The system extracts data on patients diagnosed with acute myocardial infarction in the past three years from the hospital's cardiology specialty database, mapping the feature values ​​corresponding to the occurrence of 217 confirmed events to this 125-dimensional space: for each event, based on its occurrence time, 125 feature values ​​from the feature matrix at the corresponding time are taken as spatial coordinate points. Spatial density calculation uses adaptive bandwidth Gaussian kernel density estimation, and density distribution is calculated separately for the anterior wall infarction subgroup (89 cases) in the myocardial infarction feature set. The system detects a significant high-density region in the feature subspace composed of dimension 1 (ST-segment elevation amplitude) and dimension 6 (troponin I rise slope): the case distribution density reaches its peak when ST-segment elevation is ≥2mm and troponin I rises ≥0.2ng / mL per hour. The density threshold is set using the percentile method, taking the 80th percentile of the density value of the subspace as the critical threshold. Feature combinations exceeding this threshold are included in the key medical early warning indicator set.

[0065] A clinical expert verification mechanism was incorporated into the generation of the early warning indicator set: the chief physician of cardiology reviewed the feature combinations screened by the system, confirming that the combination of three indicators—ST segment elevation amplitude, troponin I rise rate, and frequency of ventricular arrhythmias—had the highest clinical relevance. The myocardial infarction early warning rule ultimately generated by the system includes multi-layered judgment logic: the primary condition is ST segment elevation ≥2mm combined with troponin I rise rate ≥0.2ng / mL / h; the secondary conditions include a sudden drop in heart rate variability of more than 30% or the occurrence of frequent premature ventricular contractions. After the early warning system was put into clinical use, when real-time data of newly admitted patients was input, the system mapped their feature vectors to a 125-dimensional health status space. In one monitoring session, a 63-year-old female patient's feature coordinates were found to be in a high-density area: ST segment elevation 2.3mm (dimensional 1 value), troponin I rise rate 0.25ng / mL / h (dimensional 6 value), and a 35% decrease in heart rate variability from baseline (dimensional 11 value). The system immediately triggered a red alert and sent an emergency consultation request to the cardiovascular intensive care unit. After admission, the physician confirmed the patient had an acute anterior wall myocardial infarction and immediately activated the catheterization lab to prepare for emergency interventional treatment. The spatial model is continuously and dynamically updated: when the hospital admits new variant cases of myocardial infarction (such as inferior wall infarction combined with right ventricular infarction), the system automatically adds the new case's characteristic coordinates to the spatial point set. Spatial reconstruction is performed monthly, recalculating the density distribution and adjusting the warning thresholds. After one reconstruction, a unique distribution pattern was observed in the diabetic myocardial infarction patient population: ST-segment elevation was often not significant, but troponin levels rose at a faster rate. Based on this, the system generated a specific warning threshold for diabetic patients, improving diagnostic sensitivity for this population. All spatial computation data is stored in a medical-grade encrypted format, with access restricted to the cardiovascular specialist team. The spatial visualization interface supports 3D projection viewing, allowing physicians to select any three feature dimensions to observe the case distribution pattern. Warning history records are automatically compared with the final diagnostic results, generating a warning accuracy report for analysis by the Medical Quality Improvement Committee.

[0066] Example 3: The real-time medical data fusion module continuously receives streaming data from bedside monitoring devices. This module is configured with a high-throughput data pipeline to process hundreds of physiological parameter readings per second generated by devices such as ECG monitors, ventilators, and intravenous infusion pumps. A real-time feature extraction engine is deployed on a medical edge computing node, employing a lightweight convolutional neural network to analyze ST segment shift features in ECG waveforms, calculating the instantaneous coefficient of variation of respiratory rate through a recurrent neural network, and extracting the fluctuation entropy value of the blood pressure sequence using a sliding window statistical method. The generation of real-time medical features follows millisecond-level temporal constraints, with each feature vector accompanied by a timestamp accurate to milliseconds and a device source identifier. The feature vector construction process strictly aligns with the dimensional specifications of the dynamically evolving feature matrix. The 150-dimensional feature slots are filled according to a preset mapping table: dimensions 1-30 are allocated to ECG features (such as QRS complex area), dimensions 31-60 store respiratory parameters (such as inspiratory time ratio), dimensions 61-90 record hemodynamic parameters (such as arterial compliance), and the remaining dimensions are reserved for urgent laboratory test results (such as blood gas analysis values).

[0067] The projection of real-time feature vectors onto the multidimensional health state space employs an orthogonal decomposition algorithm, projecting the 150-dimensional vectors onto a feature subspace composed of principal components of historical data. Spatial distance calculation utilizes an improved Mahalanobis distance metric to avoid computational biases caused by differences in the dimensions of different features.

[0068] ;

[0069] in: This represents a real-time risk distance scalar. This represents the real-time feature vector at time t. It is the mean vector of the key medical early warning indicator set in the corresponding dimension. It is the inverse matrix of the covariance matrix of the early warning indicator set. This distance value is converted into a health risk factor in the range of 0-1 using the sigmoid function. The risk factor is updated every second and written to the real-time medical database. The clinical interpretation level of the risk factor is set as follows: the range of 0-0.3 is marked as green low risk, the range of 0.3-0.7 triggers a yellow moderate warning, and the range of 0.7 and above is a red high risk, which immediately initiates the alarm process.

[0070] The medical risk prediction and decision-making module employs a multimodal fusion architecture. The fusion of real-time health risk factors and the initial medical feature set is achieved through a gating attention mechanism, with risk factors serving as gating signals to weightedly filter relevant patterns from historical features. A temporal attention mechanism reconstructs the dynamically evolving feature matrix, calculating the cosine similarity between the feature vector at each historical time point and the current real-time state as attention weights, and then weighted summing to generate a context-aware feature representation. The risk prediction network uses a deep residual structure. The input layer receives the fused feature tensor, the hidden layer contains three 512-unit LSTM layers to capture temporal dependencies, and the output layer generates a disease probability distribution using a softmax function. The generation of personalized treatment recommendations is integrated with a clinical guideline rule engine: for coronary heart disease risk prediction, the ACC / AHA drug treatment rule tree is automatically invoked when the probability value exceeds 0.75; for diabetes risk prediction, insulin dosage adjustment recommendations are generated based on ADA guidelines. All treatment recommendations are pushed to the physician workstation via the HL7 protocol for confirmation and execution. The model training process employs a time-retention cross-validation strategy, training the network with the first 80% of the patient's time data and validating the prediction accuracy with the last 20%. Network parameter optimization utilizes an adaptive moment estimation algorithm, with the learning rate dynamically adjusted based on the validation set loss. Model deployment employs containerization technology to ensure consistency of the computing environment, and the prediction service response time is controlled within 200 milliseconds to meet clinical real-time requirements. The treatment suggestion generation module incorporates a safety verifier, filtering suggestions with potential medication risks against a drug contraindication database to ensure that the output plan conforms to the individual patient's contraindications.

[0071] A 67-year-old male patient with sepsis was admitted to the intensive care unit of a general hospital. He presented with fever, tachycardia, and hypotension after surgery for an abdominal infection. The real-time medical data fusion module continuously received vital sign data from the bedside monitor: heart rate, blood pressure, and blood oxygen saturation parameters were collected every 2 seconds; mean arterial pressure readings were obtained from the ductus arteriosus every 5 minutes; and body temperature trends were recorded hourly. The feature extraction engine simultaneously processed multi-source data streams: extracting the low-frequency to high-frequency power ratio of heart rate variability from the electrocardiogram signal, calculating the vascular tension index from the blood pressure waveform, and deriving the fever rate index from the body temperature curve.

[0072] The real-time feature vector construction strictly follows the dimensional specifications of historical data training. The LF / HF ratio (heart rate variability) is filled into the 23rd dimension, the vascular tension index into the 47th dimension, the fever rate into the 89th dimension, and the latest procalcitonin level into the 112th dimension. Vector projection maps the current 125-dimensional features to a sepsis warning space constructed from data of sepsis patients admitted over the past three years. Spatial distance calculation identifies the distance between the current vector and the center point of the sepsis exacerbation cluster as only 0.31 Mahalanobis distance units, significantly lower than the warning threshold of 1.5. The system immediately generates a health risk factor of 0.83, triggering a red high-risk alarm. The medical risk prediction and decision-making module initiates multi-dimensional analysis: fusing the real-time risk factor with the patient's initial feature set upon admission. The initial features include basic immune function indicators, chronic disease history coding, and surgical trauma severity scores. The temporal attention mechanism heavily weighted the patient's characteristic changes over the past 6 hours, paying particular attention to the decreasing trend of mean arterial pressure from 72 mmHg to 65 mmHg and the increasing process of lactate level from 2.1 mmol / L to 3.8 mmol / L. The risk prediction network output a septic shock probability of 0.79, exceeding the intervention threshold of 0.7. The individualized treatment recommendation module referenced the Sepsis Rescue Movement guidelines: initially recommending a 30 ml / kg crystalloid resuscitation regimen, while simultaneously suggesting intravenous norepinephrine to maintain a mean arterial pressure ≥65 mmHg, and recommending a broad-spectrum antibiotic regimen of meropenem combined with vancomycin. The system automatically checked the drug contraindication database to confirm that the patient had no history of vancomycin allergy and that renal function indicators were within acceptable ranges. The treatment plan was implemented immediately after confirmation by the attending physician, and the nurse precisely controlled the infusion rate of vasoactive drugs using an intelligent infusion pump. The real-time monitoring system continuously tracked the treatment response: central venous pressure was measured every 15 minutes after fluid resuscitation, and capillary refill time was continuously monitored after the administration of vasoactive drugs. When the system detects a decrease in the patient's lactate level from 3.8 mmol / L to 2.4 mmol / L within 4 hours, and a decrease in heart rate from 125 beats / min to 98 beats / min, the risk factor is automatically adjusted to 0.45. The warning level is downgraded to yellow (moderate risk), but the recommendation to continue antibiotic treatment remains in effect. All treatment process data is recorded in real time and transmitted to the hospital's data center, forming a complete digital trajectory of sepsis treatment.

[0073] Example 4: The medical knowledge graph construction module integrates multi-source medical knowledge bases. This module extracts disease diagnosis and treatment guidelines from the UpToDate clinical decision-making system, obtains drug interaction information from the FDA drug instruction database, and constructs a disease symptom association network from the NCBI Medical Subject Headings (MeSH). Knowledge extraction uses a BERT-based entity recognition model to parse medical literature, identifying clinical entities such as "acute myocardial infarction," "aspirin," and "percutaneous coronary intervention." Relation extraction uses graph neural networks to mine semantic connections between entities, establishing association rules such as "disease-treatment plan," "drug-contraindications," and "symptom-examination method." The graph is stored using the Neo4j graph database, and node attributes include medical metadata such as disease ICD-10 codes, drug ATC classification codes, and treatment plan evidence levels.

[0074] The matching and verification process between individualized treatment recommendations and the knowledge graph employs a two-way verification process: forward verification checks whether the drug treatment in the recommendations matches disease guidelines. When the system generates a recommendation for "using ticagrelor in patients with acute ST-segment elevation myocardial infarction," it automatically queries the knowledge graph for the therapeutic association between this disease and P2Y12 inhibitors; reverse verification checks for contraindication conflicts. If the patient has a history of active bleeding, the system traverses the association path between "ticagrelor" and "bleeding risk" in the knowledge graph and triggers an alert. The matching and verification results generate a structured report, marking compliant recommendations that have passed verification and risky recommendations that conflict. The multi-center medical data collaboration module connects the data centers of three hospitals, aggregating standardized multimodal medical datasets. Data aggregation follows the FHIR standard conversion format, and patient identity information is transmitted anonymously through hash encryption. The federated learning architecture adopts a horizontal partitioning model, where each hospital locally trains the neural network weight parameters of the risk prediction model and exchanges gradient updates through a secure multi-party computation protocol. The model aggregation server executes a weighted average algorithm: weighting coefficients are assigned based on the data volume of each center (center A: 45%, center B: 30%, center C: 25%), and global model parameters are updated monthly. The updated personalized health risk prediction model incorporates noise protection using differential privacy technology and is deployed to the inference servers of each hospital after review by the medical ethics committee. Key indicators for monitoring during the multi-center data collaboration process are shown in Table 1.

[0075] Table 1: Monitoring Table of Multi-center Federated Learning Operational Indicators

[0076] When the medical risk prediction and decision-making module loads the updated model, it performs a version compatibility check to ensure seamless integration of input / output interfaces with the hospital's existing systems. The prediction service employs an A / B testing strategy to gradually switch to the new model, initially comparing prediction results with 10% of patient traffic before full implementation after clinicians confirm accuracy. A complete audit log is recorded for the model inference process, including key steps such as input feature values, prediction probability calculation paths, and treatment plan generation logic, for regular review by the medical quality management department. A continuous update mechanism for the knowledge graph monitors the latest research findings in medical journals. When *The New England Journal of Medicine* publishes clinical trials of novel anticoagulant drugs, the system automatically extracts key conclusions and updates the treatment rules in the knowledge graph. A multi-center data collaboration mechanism establishes an anomaly detection mechanism. When the data distribution uploaded by a hospital deviates from the overall benchmark by more than 3 standard deviations, the node's participation in federated learning is automatically suspended, and a data quality verification process is triggered. All medical data transmission and storage comply with HIPAA security standards, and encryption keys are centrally managed by the hospital's information security department, ensuring the secure and reliable operation of the medical intelligent system.

[0077] Example 5: The treatment process traceability module automatically collects treatment implementation data through medical IoT devices. This module connects to a smart infusion pump to obtain drug infusion rate and duration parameters, extracts physician-prescribed medical order execution records from the electronic medical record system, and collects physiological parameter responses during treatment through bedside monitoring. Taking anticoagulation therapy after percutaneous coronary intervention (PCI) for coronary artery disease patients as an example: the system records the dosing time and dosage of aspirin 100mg once daily and ticagrelor 90mg twice daily, while simultaneously monitoring changes in laboratory indicators such as platelet aggregation rate and bleeding time. The effect feedback data includes the daily ST segment regression amplitude on ECG within 30 days post-procedure, dynamic changes in myocardial enzyme spectrum, and the degree of chest pain relief reported by the patient. All data is stored in a medical data warehouse in time series to form a structured treatment trajectory.

[0078] The integration of closed-loop medical data adopts a time-series database architecture. Each record contains four-dimensional attributes: a timestamp accurate to milliseconds, treatment item coding using the LOINC standard, physiological parameter values ​​with measurement device identifiers, and efficacy evaluation indicating the clinician's verification status. A data association engine automatically matches treatment measures with corresponding effect feedback; for example, it associates the ticagrelor administration time point with subsequent platelet inhibition rate test results to establish a dosing-pharmacokinetics response curve. Data quality verification rules reject records with logical conflicts. For example, a data review process is triggered if bleeding time is significantly prolonged without recording anticoagulant adjustments. The model optimization of the medical risk prediction and decision-making module uses an incremental learning framework. Closed-loop medical data is preprocessed and transformed into model training samples. The implementation effect of treatment recommendations is quantified as a reward signal: a positive sample with complete adherence and significant efficacy receives a reward of +1, while a negative sample with partial adherence and adverse reactions receives a reward of -1. Model parameter updates use a stochastic gradient descent algorithm, focusing on adjusting weight parameters that significantly affect treatment effect prediction. The optimization process preserves the model's original knowledge structure and uses regularization terms to prevent overfitting to a small number of recent samples.

[0079] Taking insulin dosage adjustment for diabetic patients as an example: the system initially recommends 12 units of aspart insulin subcutaneously injected before each of the three meals. The traceability module records the blood glucose monitoring values ​​for the subsequent 72 hours to form an effect feedback dataset. When a post-lunch blood glucose level is detected to be consistently higher than 13.9 mmol / L and a hypoglycemic event occurs before breakfast, the incremental learning algorithm adjusts the dosage prediction model: reducing the pre-breakfast dose weight coefficient and increasing the post-lunch dose sensitivity parameter. The optimized model generates a new regimen of 10 units before breakfast, 14 units before lunch, and 11 units before dinner. This regimen is then approved by an endocrinologist and put into clinical use. The treatment traceability data supports multi-dimensional efficacy analysis, and clinicians can query the response rate statistics of specific treatment regimens in different populations. The system provides a treatment pathway backtracking function, allowing patients to trace back to all treatment operations within the last 30 days when adverse reactions occur, assisting the medical team in analyzing causal relationships. All data operation records comply with FDA 21 CFR Part 11 electronic record specifications, and the audit trail log records in detail the data modification time, operator identity, and reason for modification.

[0080] The model optimization process establishes clinical safety boundaries, ensuring that no parameter adjustments exceed the extreme ranges specified in the drug's instructions. When optimization recommendations differ significantly from clinical guidelines, the system automatically initiates a multidisciplinary consultation request, where a team of medical experts assesses the rationality of the algorithm's recommendations. The optimized model version runs in shadow mode, comparing its predictions with actual clinical decisions in parallel, continuously collecting efficacy data until statistical significance is achieved before official implementation. This entire traceability and optimization cycle is integrated into the hospital's quality management system, generating a monthly performance report for the medical AI system, which is submitted to the Medical Technology Committee for clinical efficacy evaluation.

[0081] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. An artificial intelligence driven data processing system for medical diagnosis and treatment, characterized in that, include: The multimodal medical data acquisition module is used to acquire multi-source heterogeneous medical data of patients within a preset time period, and to perform data cleaning and standardization processing on the multi-source heterogeneous medical data to generate a standardized multimodal medical dataset. The medical feature deep extraction module is used to extract multi-dimensional features from the standardized multimodal medical dataset to obtain an initial medical feature set, and to construct a dynamic evolution feature matrix based on the time-varying patterns of each feature in the initial medical feature set. The multidimensional health status space construction module is used to construct a multidimensional spatial representation of the patient's health status based on the dynamic evolution feature matrix, calculate the distribution density of historical medical events in the multidimensional space, and determine a set of key medical early warning indicators based on the distribution density. The real-time medical data fusion module is used to receive real-time collected patient medical data, extract real-time medical features and form real-time feature vectors, map the real-time feature vectors to the multi-dimensional health state space, and generate real-time health risk factors. The medical risk prediction and decision-making module is used to establish a personalized health risk prediction model based on the real-time health risk factors, the initial medical feature set, and the dynamic evolution feature matrix, output the probability of disease occurrence, and generate individualized treatment suggestions. The real-time medical data fusion module is specifically used for: Feature extraction is performed on real-time collected patient medical data to generate real-time medical features; The real-time medical features are constructed into real-time feature vectors according to the organization format of the dynamic evolution feature matrix; The real-time feature vector is projected onto the multidimensional health state space, and its spatial distance with the set of key medical early warning indicators is calculated. Real-time health risk factors are generated based on the spatial distance; The medical risk prediction and decision-making module establishes a personalized health risk prediction model as follows: The real-time health risk factors are fused with the initial medical feature set; The dynamic evolution feature matrix is ​​reconstructed using a temporal attention mechanism with weights. A risk prediction network is trained based on the fused features and the reconstructed feature matrix. Output the probability value of the patient's disease occurrence and generate a personalized treatment suggestion plan. 2.The data processing system of artificial intelligence driven medical diagnosis and treatment of claim 1, wherein, The multimodal medical data acquisition module is specifically used for: Collect multi-source heterogeneous medical data, including imaging data, time-series physiological parameter data, laboratory test data, and clinical text data; The multi-source heterogeneous medical data is subjected to outlier removal, noise filtering, and missing value imputation to enhance data quality. The enhanced multi-source heterogeneous medical data is aligned and normalized according to data modality to generate a standardized multimodal medical dataset. 3.The data processing system of artificial intelligence driven medical diagnosis and treatment according to claim 2, wherein, The medical feature deep extraction module is specifically used for: A deep convolutional neural network is used to extract multi-scale medical features from the standardized multimodal medical dataset; Medical features at different scales are fused and spliced ​​together to form an initial medical feature set; Calculate the rate of change and trend of change of each feature in the initial medical feature set within a specified time window; The rate of change and trend of change of each feature are organized according to time series to construct a dynamic evolution feature matrix. 4.The data processing system of artificial intelligence driven medical diagnosis and treatment according to claim 3, wherein, The multidimensional health state space construction module is specifically used for: The number of dimensions of the multidimensional health state space is determined based on the number of feature dimensions of the dynamic evolution feature matrix. Calculate the distribution density of the corresponding features at the time of historical diagnostic events in the multidimensional health state space; A density threshold is set based on the distribution density, and features whose distribution density exceeds the density threshold are selected as the key medical early warning indicator set. 5.The data processing system of artificial intelligence driven medical diagnosis and treatment according to claim 4, wherein, Also includes: The medical knowledge graph construction module is used to integrate medical knowledge bases and clinical guidelines to build a medical knowledge graph that includes the relationships between diseases, symptoms, drugs, and treatment plans. The medical risk prediction and decision-making module is also used to match and verify the individualized treatment recommendations with the medical knowledge graph. 6.The data processing system of artificial intelligence driven medical diagnosis and treatment according to claim 5, wherein, Also includes: The multi-center medical data collaboration module is used to aggregate standardized multimodal medical datasets from multiple medical institutions, perform federated learning training, and update the personalized health risk prediction model. The medical risk prediction and decision-making module is also used to make predictions using an updated personalized health risk prediction model. 7.The data processing system of artificial intelligence driven medical diagnosis and treatment according to claim 6, wherein, Also includes: The diagnosis and treatment process traceability module is used to record the implementation process and effect feedback of the individualized treatment recommendations, forming closed-loop medical data; The medical risk prediction and decision-making module is also used to continuously optimize the personalized health risk prediction model based on the closed-loop medical data.

8. A data processing method for medical diagnosis and treatment driven by artificial intelligence, characterized in that, It includes all modules and method flows of the AI-driven medical diagnosis and treatment data processing system as described in any one of claims 1 to 7.