Multi-mode intelligent medical teaching training system and data processing method

Through the multimodal intelligent medical teaching and training system, the CNN, RNN and GNN branches are used to extract and fuse multi-source heterogeneous medical data features, and the existing system's problems in synchronization accuracy, model adaptation and cross-professional collaboration are solved, and the teaching effect and analysis capabilities are improved.

CN120373905AInactive Publication Date: 2025-07-25NANJING COLLEGE OF INFORMATION TECH

Patent Information

Application Number
CN202510493411.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When processing multi-source heterogeneous data, the medical e-teaching training system has problems such as insufficient synchronization accuracy, difficulty in adapting to the model architecture to the correlation characteristics of timing physiological signals and spatial, data augmentation methods cannot generate pathological samples that conform to medical laws, and lack of cross-professional collaboration support, resulting in poor teaching results.

Method used

A multimodal intelligent medical teaching training system is adopted, including medical data acquisition, preprocessing, quality assessment, multi-branch neural network feature extraction and feature fusion modules. The time frequency, timing and spatial features are extracted respectively by CNN, RNN and GNN branches, and a unified representation is formed through feature fusion. Combined with teaching effect evaluation and early warning analysis of critical patients, the data processing process is optimized.

Benefits of technology

The synchronization accuracy and teaching effect of multi-source heterogeneous data are improved, the ability to analyze complex physiological states is improved, cross-modal information complementarity and cross-professional collaboration are achieved, and students' comprehensive innovation ability is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373905A_ABST
    Figure CN120373905A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal intelligent medical teaching practical training system and a data processing method. The system comprises a medical data acquisition module used for acquiring multi-source heterogeneous medical data; the preprocessing module is used for preprocessing the collected medical data, including data synchronization; the medical data quality evaluation module is used for evaluating the medical data quality by adopting a quality scoring function to obtain a data quality evaluation result, and performing feature extraction on the corresponding medical data if the evaluation result is higher than a set value; the feature extraction module is used for performing medical feature extraction on the medical data after quality evaluation by using a multi-branch neural network, the multi-branch neural network comprises a CNN branch, an RNN branch and a GNN branch, and time-frequency features, time-sequence features and spatial features in the medical features are extracted respectively; and the feature fusion module is used for performing feature fusion on the medical features of each branch obtained by the feature extraction module by adopting a fusion function, and combining the multi-modal features extracted by different branches to form uniform feature representation. According to the system and the method, the multi-source heterogeneous data synchronization precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-modal intelligent medical teaching and training method and system, belonging to the technical field of comprehensive experimental teaching in higher vocational colleges, undergraduate and postgraduate education. Background Art

[0002] With the rapid development of intelligent medicine, the cross-integration of medical electronics and artificial intelligence has become increasingly close. In teaching practice, there are many deficiencies in traditional medical electronics training systems. In terms of data acquisition, existing systems mostly use single sensors or homogeneous sensor combinations, and cannot effectively process multi-modal data with a sampling rate difference exceeding 1000:1, resulting in insufficient synchronization accuracy of high-frequency ECG signals and low-frequency body temperature data; in terms of model architecture, most use a single network structure to process medical images, and it is difficult to adapt to temporal physiological signals and spatial correlation features, restricting the system's analysis ability for complex physiological states; in terms of data augmentation, traditional methods such as oversampling or SMOTE technology cannot generate pathological samples that conform to medical laws, limiting the model's recognition ability for rare cases; in terms of teaching applications, experimental projects are often limited to a single professional perspective, lacking cross-professional collaboration support, and it is difficult to cultivate students' comprehensive innovation ability.

[0003] Currently, the field of medical artificial intelligence is developing rapidly. For example, in multi-modal fusion, the cross-modal alignment technology of medical images and text reports based on CLIP has significantly improved the feature expression ability; in terms of model lightweighting, knowledge distillation technology is applied in the deployment of medical devices, and the accuracy loss of 8-bit fixed-point inference is controlled within 1.2%; in terms of teaching digitization, the application of digital twin technology in clinical training has realized the linkage between virtual operations and physical devices. Patent CN113724853A discloses a smart medical system based on deep learning, and patent CN115470856A discloses a multi-modal data fusion method and application based on semantic information content. However, in the sampling and subsequent data processing of multi-source heterogeneous data, the design of medical AI models faces problems such as multi-modal feature extraction, temporal alignment, and model interpretability. Special network structures are required for feature extraction of different modal medical data, and the temporal alignment and feature fusion of multi-source medical data are complex. For example, the phase difference between electrocardiogram and blood pressure waveforms is about 0.1 - 0.3s. The above problems make it necessary to further improve the training effect of medical electronics teaching and training systems. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: in the process of medical electronics teaching and training, how to process the collected multi-source heterogeneous data, improve the synchronization accuracy of multi-source heterogeneous data, and further enhance the effect of medical teaching and training.

[0005] To solve the above technical problems, the present invention provides a multi-modal intelligent medical teaching and training system, including the following modules:

[0006] Medical data acquisition module: used to collect multi-source heterogeneous medical data;

[0007] Preprocessing module: preprocess the collected medical data, including data synchronization;

[0008] Medical data quality assessment module: use a quality scoring function to evaluate the quality of medical data, obtain the data quality assessment result, and if the assessment result is higher than the set value, perform feature extraction on the corresponding medical data;

[0009] Feature extraction module: use a multi-branch neural network to extract medical features from the medical data after quality assessment. The multi-branch neural network includes a CNN branch, an RNN branch, and a GNN branch, which respectively extract time-frequency features, time-series features, and spatial features in the medical features;

[0010] Feature fusion module: use a fusion function to perform feature fusion on the medical features of each branch obtained by the feature extraction module, combine the multi-modal features extracted by different branches, and form a unified feature representation.

[0011] The foregoing multi-modal intelligent medical teaching and training system further includes:

[0012] Teaching effect assessment module: evaluate the teaching effect of the virtual-real combined medical scenario during the medical teaching process according to the medical feature fusion data.

[0013] Severe patient warning analysis module: based on the multi-modal medical data after data fusion, use a scoring function to perform real-time monitoring of the multi-organ functions of severe patients and perform warning analysis.

[0014] Data processing performance assessment module: use an assessment model to evaluate the data processing performance of medical training data. The data processing performance of medical training data includes data processing efficiency and effect indicators. The data processing includes data synchronization and data monitoring, providing a basis for further optimization of medical training processing.

[0015] Optimization processing module: further optimize the synchronized multi-modal data stream and the multi-branch neural network to reduce latency and resource consumption.

[0016] In the medical data acquisition module of the foregoing multi-modal intelligent medical teaching and training system, the multi-source heterogeneous medical data includes: electrocardiogram, electroencephalogram, electromyogram, blood oxygen, blood pressure, body temperature, respiration, motion state, and environmental parameters.

[0017] In the medical data quality assessment module, the quality scoring function is , where is the weight coefficient of the -th type of medical data, is the quality score of the -th type of medical data, is the diagnostic time window weighting function, is the clinical integrity indicating function, is the total number of medical data categories.

[0018] For the aforementioned multi-modal intelligent medical teaching and training system, in the feature extraction module, in the CNN branch, a multi-scale convolution structure is adopted to extract time-frequency features of high-frequency signals. The feature extraction function of the CNN branch is , where represents the output features extracted by the CNN branch, represents using the -th convolution kernel with a convolution kernel length to convolve the input signal to obtain the features, is the weight coefficient of the -th convolution kernel, is the bias term.

[0019] For the aforementioned multi-modal intelligent medical teaching and training system, in the feature extraction module, the RNN branch adopts a bidirectional GRU structure to process the temporal features of physiological parameter sequences and models the temporal dependence relationship of vital signs through a gating mechanism; the calculation process of the RNN branch includes: , , , ; where, is the update gate, is the reset gate, is the candidate hidden state, is the current hidden state; is the Sigmoid activation function, , , are the weight matrices of the update gate, reset gate, and candidate hidden state respectively, and the symbol represents element-wise multiplication of vectors.

[0020] For the aforementioned multi-modal intelligent medical teaching and training system, in the feature extraction module, the GNN branch constructs a graph attention network structure based on the anatomical spatial relationship of electrocardiogram / electroencephalogram multi-leads, calculates the association weights between different lead nodes through an attention mechanism, and updates the node spatial features, expressed as:

[0021]

[0022] ;

[0023] Among them, is the ReLU activation function with leakage, is the learnable attention weight vector, is the linear transformation matrix of node features, represents node 's input feature vector, represents node 's updated feature vector. The symbol means concatenating and vectorially, is the set of neighbor nodes of node .

[0024] For the aforementioned multimodal intelligent medical teaching and training system, in the feature fusion module, the fusion function is , = 3, where the weight , and corresponding weight parameters are configured for different disease states. The weight configuration of is adopted for arrhythmia detection, and the weight configuration of is adopted for heart failure early warning, and the weight configuration of is adopted for stroke;

[0025] MLP refers to a multi-layer fully connected neural network, which is used to process the concatenated feature vectors and output the weights of each modality feature;

[0026] represents the learnable parameters in the multi-layer fully connected neural network MLP, including weights and biases.

[0027] For the aforementioned multimodal intelligent medical teaching and training system, in the teaching effect evaluation module, a comprehensive scoring calculation function is used to fuse multiple evaluation indicators. The comprehensive scoring calculation function is , where , , are the weight coefficient one, weight coefficient two, and weight coefficient three respectively, , , are the calculation results of the virtual medical scenario training, physical operation training, and clinical case analysis scoring functions respectively;

[0028] The virtual medical scenario training scoring function is , where is the completion score of the th virtual training project, is the project weight coefficient, is the total number of virtual training projects;

[0029] The scoring function for hands-on training is , where is the completion quality score of the th hands-on operation project, is the time coefficient, is the operation weight coefficient, is the total number of hands-on operation projects;

[0030] The scoring function for clinical case analysis and evaluation is , where is the analysis score of the th clinical case, is the application ability coefficient, is the case weight coefficient, is the total number of clinical cases.

[0031] For the aforementioned multi-modal intelligent medical teaching and training system, in the data preprocessing module, the multi-modal medical data is synchronously processed for alignment of different modal medical data on the time axis, including:

[0032] Inject medical signal synchronization marks during the acquisition of medical data for calibration and alignment of multi-source medical data, expressed as: , where is the synchronization reference time, is the marking time interval, is the introduced random offset, represents the synchronization mark time point inserted for the th time;

[0033] Perform medical parameter cascaded adaptive filtering, and the formula is , where is the filtering output of the th type of signal at discrete time , is the input value of the th type of signal at time , is the th coefficient of the adaptive filter for the th type of signal, is the signal correction factor based on medical knowledge, is the coefficient length of the filter;

[0034] Perform physiological signal phase locking to further align the phases of different modal signals at the cycle level. The calculation formula is , where and are the signals respectively Sum signal The occurrence time point of the characteristic event within a certain synchronization period and are respectively the cycle lengths of signal and signal ; Indicates the phase difference of signal relative to signal .

[0035] For the aforementioned multi-modal intelligent medical teaching and training system, in the critical patient warning analysis module, the scoring function is , and the formula is as follows:

[0036]

[0037] Among them, is the function score of the th organ, is the weight coefficient of the th organ function, is the time sensitivity coefficient, is the inflammation factor correction index; is the function trend score of the th organ, and the formula is , and the vertical bar means taking the absolute value;

[0038] is the organ-specific time weight, is the th change amount of the organ function score between adjacent time points; the early warning reserved time is calculated by the formula , where is the critical intervention threshold of the multi-organ function score, and the corresponding critical time is , is the safety factor.

[0039] For the aforementioned multi-modal intelligent medical teaching and training system, in the data processing performance evaluation module, the evaluation model includes:

[0040] The balance degree of medical multi-modal data processing, and the formula is , where is the number of medical data types, is the processing efficiency of the th type of medical data, is the importance weight of the th type of medical data, is a medical relevance index, used to characterize the relevance degree of different data types in medical decision-making;

[0041] Clinical model scheduling efficiency, the formula is , where is the number of clinical models participating in the scheduling, is the clinical importance index of the th clinical model, is the th inference time of the clinical model, is the maximum inference time among all clinical models, is the th GPU / CPU utilization rate of the clinical model;

[0042] Medical data processing throughput, the formula is , where is the computing power of the th processing unit, is the difficulty factor of the th type of disease, is the number of processing scenarios considered;

[0043] System response efficiency, the formula is , where is the theoretically shortest response time, is the actual response time;

[0044] Diagnostic accuracy rate, the formula is , where is the number of samples with correct diagnosis, is the total number of samples in the diagnostic test.

[0045] The aforementioned multi-modal intelligent medical teaching and training system, in the optimization processing module, includes:

[0046] Clinical priority dynamic scheduling, the formula is , where is the weight of the th type of medical modality, is the generation frequency of medical data in this modality, is the th clinical importance index corresponding to the th type of data, is the processing priority value of the

[0047] Medical AI model lightweighting, the formula is , where is the basic distillation loss weight, is the clinical knowledge distillation loss weight, For the output of the teacher model and the output of the student model the Kullback–Leibler divergence between them, where and represent the output probability distributions of the teacher model and the student model respectively, is the task loss function of the model, is the loss term to ensure the ability of the student model to identify key clinical features, is the total loss function of knowledge distillation.

[0048] A data processing method for a multi-modal intelligent medical teaching and training system, comprising the following steps:

[0049] Step 1: For collecting multi-source heterogeneous medical data;

[0050] Step 2: Preprocessing the collected medical data, including data synchronization;

[0051] Step 3: Using a quality scoring function to evaluate the quality of medical data, obtaining a data quality evaluation result. If the evaluation result is higher than the set value, feature extraction is performed on the corresponding medical data;

[0052] Step 4: Using a multi-branch neural network to extract medical features from the quality-evaluated medical data. The multi-branch neural network includes a CNN branch, an RNN branch, and a GNN branch, which respectively extract time-frequency features, temporal features, and spatial features in the medical features;

[0053] Step 5: Using a fusion function to perform feature fusion on the medical features of each branch obtained by the feature extraction module, combining the multi-modal features extracted by different branches to form a unified feature representation.

[0054] The foregoing data processing method for a multi-modal intelligent medical teaching and training system further includes:

[0055] Step 6: According to the medical feature fusion data, evaluating the teaching effect of the virtual-real combined medical scenario during the medical teaching process.

[0056] Step 7: Based on the multi-modal medical data after data fusion, using a scoring function to perform real-time monitoring of the multi-organ functions of critically ill patients and perform early warning analysis.

[0057] Step 8: Using an evaluation model to evaluate the processing performance of medical training data. The processing performance of medical training data includes data processing efficiency and effect indicators. The data processing includes data synchronization and data monitoring, providing a basis for further optimization of medical training processing.

[0058] Step Nine: Further optimize the synchronized multi-modal data stream and the multi-branch neural network to reduce latency and resource consumption.

[0059] Beneficial effects achieved by the present invention: The multi-modal intelligent medical teaching and training system of the present invention uniformly collects various physiological signal types, performs adaptive sampling synchronization and data quality assessment on the collected medical data, and uses a CNN / RNN / GNN multi-branch neural network for heterogeneous data feature extraction and multi-level feature fusion, improving the effect of medical teaching and training.

[0060] Meanwhile, the present invention further optimizes the synchronized multi-modal data stream and the multi-branch neural network to further reduce the latency and resource consumption of the medical teaching and training system, providing a new solution for improving the effect of medical teaching and training. Description of the Drawings

[0061] Figure 1 It is a schematic structural diagram of a multi-modal intelligent medical teaching and training system in Embodiment 1 of the present invention;

[0062] Figure 2 It is a flowchart of a data processing method of a multi-modal intelligent medical teaching and training system in Embodiment 4 of the present invention;

[0063] Figure 3 It is a schematic diagram of the data flow direction in the process of feature extraction and data fusion using a multi-branch deep learning network in Embodiment 4 of the present invention. Detailed Embodiments

[0064] The following further elaborates on the present invention in conjunction with the drawings.

[0065] Embodiment 1

[0066] As Figure 1 shown, this embodiment provides a multi-modal intelligent medical teaching and training system, including the following modules:

[0067] Medical data acquisition module: used to collect multi-source heterogeneous medical data, including electrocardiogram, electroencephalogram, electromyogram, blood oxygen, blood pressure, body temperature, respiration, motion state, and environmental parameters;

[0068] In medical electronic teaching and scientific research scenarios, medical data has heterogeneous characteristics, including continuous waveform data such as electrocardiogram (ECG), electromyography (EMG), and electroencephalogram (EEG), interval sampling data such as blood pressure, blood oxygen, and body temperature, motion state data such as acceleration, angular velocity, and posture, and environmental parameter data such as temperature, humidity, light, and air pressure. Different types of data have significant differences in sampling characteristics. The sampling rate of bioelectric signals is greater than 500Hz, and high-precision sampling must be guaranteed; the sampling rate of physiological signals is within the range of 10-500Hz, and sampling stability must be guaranteed; the sampling rate of physiological parameters is less than 10Hz, and the accuracy of timing sampling must be guaranteed; at the same time, trigger collection of specific events is also required, such as human movement, changes in body position, etc.

[0069] There are significant differences in the feature expression of multi-source medical data. In terms of time domain features, ECG signals include P waves (0.08-0.12s), QRS complex waves (0.06-0.10s), T waves (0.16-0.20s) and other features, and EEG signals include α waves (8-13Hz), β waves (13-30Hz) and other features; in terms of frequency domain features, the main frequency band of ECG is 0.5-40Hz, the main frequency band of EEG is 0.5-30Hz, and the main frequency band of EMG is 20-500Hz; in terms of spatial features, it includes 12-lead ECG spatial distribution, multi-lead EEG spatial distribution, multi-point EMG spatial arrangement, etc. The above feature differences put forward higher requirements for the data acquisition module of the medical training system.

[0070] In this embodiment, modular sensors are used for data collection, including bioelectric signal sensors, physiological parameter sensors, motion state sensors and environmental parameter sensors, wherein the sampling rate of the bioelectric signal sensor is greater than 500Hz, and the accuracy is not less than 16 bits; the sampling rate of the physiological parameter sensor is in the range of 10-500Hz, and the accuracy is not less than 12 bits; the sampling rate of the motion state sensor is greater than 100Hz, and the accuracy is not less than 14 bits; the sampling rate of the environmental parameter sensor is greater than 1Hz, and the accuracy is not less than 10 bits.

[0071] Preprocessing module: preprocesses the collected medical data, including data synchronization, signal filtering, noise removal, baseline drift correction and outlier removal in sequence;

[0072] The raw data is filtered, denoised, baseline drift corrected, outliers removed, etc. to improve signal quality and achieve high-precision synchronization technology. At the same time, the triple mechanism of hardware clock synchronization, software timestamp synchronization and adaptive delay compensation is used to ensure the accurate synchronization of multi-source medical data, with a synchronization accuracy better than 0.1ms.

[0073] Medical data quality assessment module: The quality of medical data is evaluated using a quality scoring function to obtain the data quality assessment result. The collected data is judged whether it meets the requirements of subsequent data processing based on the data quality assessment result. If the assessment result is higher than the set value, feature extraction is performed on the corresponding medical data; if the assessment result is lower than the set value, it indicates that the data quality is insufficient and collection improvement or data compensation is required to ensure the reliability of subsequent feature extraction and analysis processing.

[0074] The quality scoring function is , where is the weight coefficient of the th type of medical data, is the quality score of the th type of medical data, is the diagnostic time window weighting function, is the clinical integrity indication function, is the total number of medical data categories.

[0075] The diagnostic time window weighting function takes the value of 1.2 in the emergency care monitoring scenario, 1.0 in the routine care monitoring scenario, 0.9 in the long-term monitoring scenario, and 0.8 in the health assessment scenario.

[0076] Clinical integrity indication function takes the value of 1.0 when the data in the critical interval is complete, 0.6 when there is a missing value in the non-critical interval, 0.3 when there is a missing value in the critical interval but interpolation is possible, and 0.0 when there is a serious missing value in the critical interval.

[0077] Feature extraction module: A multi-branch neural network is used to extract medical features from the medical data after quality assessment. The multi-branch neural network includes a CNN branch, an RNN branch, and a GNN branch, which extract time-frequency features, temporal features, and spatial features in the medical features respectively;

[0078] In the CNN branch, a multi-scale convolution structure is used to extract time-frequency features from high-frequency signals. The high-frequency signals include ECG / EMG waveforms with a sampling rate of 500 - 1000Hz, and the time-frequency features include waveform morphology features.

[0079] The feature extraction function of the CNN branch is , where represents the output features extracted by the CNN branch, represents using a convolution kernel with a length of for the th convolution kernel to convolve the input signal to obtain the features, is the The weight coefficients of a convolution kernel is the bias term; the length of the convolution kernel is designed according to the electrocardiogram waveform characteristics: for QRS complex detection, for P wave detection, for T wave detection, for ST segment change detection.

[0080] The RNN branch adopts a bidirectional GRU structure to process the temporal characteristics of physiological parameter sequences (0.1–100 Hz), and models the temporal dependence of vital signs through a gating mechanism; the calculation process of the RNN branch includes: , , , ; where is the update gate, is the reset gate, is the candidate hidden state, is the current hidden state; is the Sigmoid activation function, , , are the weight matrices of the update gate, reset gate, and candidate hidden state respectively, and the symbol represents element-wise multiplication of vectors (Hadamard product). For different types of physiological signals, different numbers of GRU hidden units and sequence lengths can be configured. For example, the heart rate sequence uses 64 hidden units and the sequence length is 300, the blood pressure sequence uses 32 hidden units and the length is 180, the blood oxygen saturation sequence uses 16 hidden units and the length is 120, and the body temperature sequence uses 8 hidden units and the length is 60.

[0081] The GNN branch constructs a Graph Attention Network (GAT) structure based on the anatomical spatial relationship of electrocardiogram / electroencephalogram multi-leads, calculates the association weights between different lead nodes through an attention mechanism, and updates the node spatial features, expressed as:

[0082]

[0083] ;

[0084] where is the leaky ReLU activation function, is the learnable attention weight vector, is the linear transformation matrix of the node features, Represents the input feature vector of node , represents the updated feature vector of node . The symbol represents concatenating and into a vector. is the set of neighbor nodes of node . In the GNN branch, nodes represent individual measurement channels or sensor acquisition points in multi-lead (such as electrocardiogram or electroencephalogram) data. Each node reflects the original signal features of a single lead or sensor, and the edge weights calculated through the graph attention mechanism reflect the anatomical or functional relationships between nodes. The edge weights represent the weight coefficients of the connection strength between different nodes.

[0085] Feature fusion module: A fusion function is used to perform feature fusion on the medical features of each branch obtained by the feature extraction module, combining the multi-modal features extracted from different branches to form a unified feature representation for subsequent early warning analysis and teaching effect evaluation, thereby improving the accuracy of analysis and achieving cross-modal information complementarity.

[0086] Adopt a dynamic weighted feature fusion mechanism considering disease diagnosis associations. The fusion function is , = 3, where the weight . Corresponding weight parameters are configured for different disease states. For arrhythmia detection, the weight configuration of is adopted. For heart failure early warning, the weight configuration of is adopted. For stroke, the weight configuration of is adopted.

[0087] MLP refers to a multi-layer fully connected neural network, which is used to process the concatenated feature vector and output the weights of each modal feature;

[0088] represents the learnable parameters in the multi-layer fully connected neural network MLP, including weights and biases, which are used to adjust the mapping relationship of the network.

[0089] Figure 3 is a schematic diagram of the structure and data flow of the multi-branch deep learning model fusion system, including three parallel neural network branches and a feature fusion mechanism. Figure 3 In

[0090] For different application scenarios such as arrhythmia detection, brain function assessment, and hemodynamic analysis, using the above fusion method, the fusion accuracies reach 98.5%, 97.8%, and 96.9% respectively.

[0091] For different medical application scenarios, corresponding multimodal fusion algorithms are adopted to make full use of information from different modalities. In the arrhythmia detection scenario, a multimodal fusion method based on electrocardiogram (ECG), blood oxygen, and blood pressure signals is used, and the fusion function is defined as , where , , are the weight coefficients of the ECG signal, blood oxygen signal, and blood pressure signal respectively, , , are the feature vectors extracted from the ECG, blood oxygen, and blood pressure signals respectively. By dynamically adjusting the weights of each modal feature according to the pathological characteristics of arrhythmia, accurate identification of arrhythmia is achieved, and the detection accuracy reaches 98.5%.

[0092] In the brain dysfunction assessment scenario, three modal features of electroencephalogram (EEG), eye movement, and facial expression are fused for comprehensive analysis. The fusion function is defined as , where , , are the weight coefficients of the EEG feature, eye movement feature, and facial expression feature respectively; represents the EEG feature extracted from the EEG power spectrum, etc., which is used to evaluate the state of brain nerve activity, represents the feature extracted from eye movement tracking, which is used to evaluate visual attention and cognitive function, represents the feature extracted from facial expression (such as facial symmetry), which is used to evaluate the state of facial nerve function. Through the complementary fusion of multimodal features, a comprehensive assessment of brain dysfunction is achieved, and the assessment accuracy reaches 96.8%.

[0093] In the hemodynamic analysis scenario, three physiological signals of ECG, blood pressure, and blood oxygen are fused to achieve comprehensive analysis of the cardiovascular circulation state. The fusion function can adopt a form similar to that in arrhythmia detection, for example . The meanings of the above symbols are the same as those defined in the arrhythmia detection scenario. By dynamically adjusting the weight coefficients of each modal signal, the hemodynamic state can be accurately characterized, and the assessment of the patient's circulatory function can be achieved, and the analysis accuracy reaches 95.2%. It can be seen that for different disease characteristics, different multimodal fusion strategies can be selected, but the overall fusion framework remains the same.

[0094] Teaching effect evaluation module: According to the medical feature fusion data, evaluate the teaching effect of the virtual-real combined medical scenario during the medical teaching process.

[0095] In this embodiment, a comprehensive score calculation function is used to fuse multiple evaluation indicators to quantify the comprehensive performance of learners. The comprehensive score calculation function is , where , , are the first weight coefficient, the second weight coefficient, and the third weight coefficient respectively, , , are the calculation results of the virtual medical scenario training, physical operation training, and clinical case analysis scoring functions respectively. Through the comprehensive score function, the evaluation results of the three links of virtual simulation training, physical operation training, and clinical case analysis are weighted and fused to obtain a comprehensive evaluation of the teaching and training performance of learners.

[0096] The evaluation results can be used to reverse-verify the effectiveness of the multi-modal training system, forming a complete teaching closed-loop.

[0097] The virtual medical scenario training scoring function is , where is the completion score of the th virtual training project, is the project weight coefficient, is the total number of virtual training projects, which can objectively reflect the operation performance of learners in the virtual environment.

[0098] The physical operation training scoring function is , where is the completion quality score of the th physical operation project, is the time coefficient, is the operation weight coefficient, is the total number of physical operation projects, taking into account both operation quality and time efficiency.

[0099] The clinical case analysis evaluation scoring function is , where is the analysis score of the th clinical case, is the application ability coefficient, is the case weight coefficient, is the total number of clinical cases, which can comprehensively evaluate the clinical application ability of learners.

[0100] In terms of teaching evaluation, the evaluation results have a significant correlation with clinical practice performance, and the correlation coefficient > 0.85, which can accurately reflect the theoretical mastery degree and practical operation ability of learners, providing a quantitative basis for teaching improvement.

[0101] The data preprocessing module also includes: synchronizing multimodal medical data for further alignment of different-modal medical data on the time axis, ensuring the accuracy of multi-source medical data fusion, including:

[0102] First, inject medical signal synchronization markers during the acquisition of medical data for calibrating and aligning multi-source medical data, expressed as: , where is the synchronization reference time, is the marker time interval, is the introduced random offset, represents the th inserted synchronization marker time point.

[0103] By introducing synchronization markers, it is ensured that physiological signals with different sampling rates can be accurately aligned, and the synchronization error is controlled within 1 ms.

[0104] Second, perform medical parameter cascaded adaptive filtering, and the formula is , where is the filtering output of the th type of signal at discrete time , is the input value of the th type of signal at time , is the th coefficient of the adaptive filter of the th type of signal, is the signal correction factor based on medical knowledge, is the coefficient length of the filter.

[0105] The adaptive filtering process can adjust parameters according to the spectral differences of different physiological signals, effectively improving the signal quality and increasing the signal-to-noise ratio by no less than 20 dB.

[0106] Third, perform physiological signal phase locking to further align the phases of different-modal signals at the cycle level. The calculation formula is , where and are the characteristic event occurrence time points of signal and signal within a certain synchronization cycle. For example, when signal is an electrocardiogram signal, can be set to represent the time when the electrocardiogram R wave appears, and are the characteristic event occurrence time points of signal and signal The cycle length, such as the time interval between adjacent characteristic events; Indicates a signal Relative to the signal The phase difference, expressed in degrees.

[0107] Through the above synchronization process, precise phase alignment of multi-source physiological signals is achieved, and the relative phase error of cross-modal signals is controlled within ±5°. High-precision synchronization of multi-modal medical data is realized, providing an accurate and reliable time reference for feature fusion and teaching evaluation.

[0108] Severe patient warning analysis module: Based on the multi-modal medical data after data fusion, a scoring function is used to monitor the real-time multi-organ functions of severe patients and conduct warning analysis, so as to realize the real-time evaluation and early warning of the functional states of multiple important organs of severe patients.

[0109] The scoring function is , and the formula is as follows:

[0110]

[0111] Where is the function score of the th organ, is the weight coefficient of the th organ function, is the time sensitivity coefficient, is the inflammation factor correction index; is the function trend score of the th organ, and the formula is . In this formula, the vertical bar means to take of the absolute value. Therefore, the fraction

[0112]

[0113] is actually of the sign function, and its result is:

[0114] When ,

[0115] When , ;

[0116] Where is the organ-specific time weight, is the change amount of the th organ function score between adjacent time points, that is, the current value minus the previous moment value; the early warning reserve time The calculation formula is , where is the critical intervention threshold for the multi-organ function score. Reaching the threshold indicates that the patient's condition has become extremely critical and immediate intervention is required. The corresponding critical time is , is the safety factor.

[0117] Data processing performance evaluation module: An evaluation model is used to evaluate the data processing performance of medical training. The data processing performance of medical training includes data processing efficiency and effect indicators. The data processing includes data synchronization and data monitoring, providing a basis for further optimization of medical training. The evaluation model includes:

[0118] Balance degree of medical multi-modal data processing, and the formula is , where is the number of medical data types, is the th class of medical data processing efficiency, is the th class of medical data importance weight, is the medical relevance index, which is used to characterize the degree of relevance of different data types in medical decision-making. This indicator is used to measure the balance degree of different types of medical data processing. The closer the value is to 1, the more balanced the processing of various data is.

[0119] Clinical model scheduling efficiency, and the formula is , where is the number of clinical models participating in the scheduling, is the th clinical model's clinical importance index, is the th clinical model's inference time, is the maximum inference time among all clinical models, is the th clinical model's GPU / CPU utilization rate. By maximizing this indicator, the overall system performance can be improved.

[0120] Medical data processing throughput, and the formula is , where is the th processing unit (referring to the computing resource unit for data processing, including CPU, GPU or dedicated hardware accelerator)'s computing power, such as the amount of data that can be processed per second, is the th class of disease type's difficulty factor, is the number of processing scenarios under consideration. This metric represents the maximum rate at which the system can process medical data, which is jointly determined by the computing power of each processing module and the complexity of the diseases being processed. The minimum value of the above product is taken as the evaluation metric.

[0121] System response efficiency, with the formula , where is the theoretically shortest response time, is the actual response time. Response efficiency reflects the ability to respond to inputs in real time, and the larger the value, the faster the response.

[0122] Diagnostic accuracy rate, with the formula , where is the number of samples correctly diagnosed, is the total number of samples in the diagnostic test. Diagnostic accuracy rate is used to measure the accuracy of diagnostic results, and the value closer to 1 indicates a more accurate diagnosis.

[0123] Optimization of processing modules: Further optimize the synchronized multi-modal data stream and multi-branch neural network to further reduce latency and resource consumption, improve overall performance, and at the same time ensure that the diagnostic accuracy rate does not decrease, including:

[0124] Dynamic scheduling of clinical priorities, with the formula , where is the weight of the th type of medical modality, is the generation frequency of medical data under this modality, is the th clinical importance index corresponding to the type of data, is the th processing priority value of the type of data. Through the dynamic scheduling of clinical priorities, the processing order of different data types can be dynamically adjusted according to clinical needs to ensure that key data is processed first.

[0125] Lightweighting of the medical AI model, with the formula , where is the basic distillation loss weight, is the clinical knowledge distillation loss weight, is the output of the teacher model and the output of the student model The Kullback–Leibler divergence between them, where and represent the output probability distributions of the teacher model and the student model respectively, is the task loss function of the model, is the loss term to ensure the ability of the student model to identify key clinical features, It is the total loss function for knowledge distillation. Through this model compression method, while maintaining the diagnostic performance of the model as much as possible, the computational resource requirements of the model are significantly reduced.

[0126] Example 2

[0127] In the feature extraction module, the CNN branch can adopt a variant structure of ResNet for time-frequency feature extraction. The variant structure of ResNet includes 5 convolutional layers and 2 fully connected layers, with an output feature dimension of 256, and at the same time extracts the waveform morphology features in medical data.

[0128] Example 3

[0129] In the feature extraction module, the RNN branch uses a bidirectional LSTM (BiLSTM) structure to implement temporal feature extraction, including 3 stacked bidirectional LSTM units with a hidden layer dimension of 128, which is used to extract the temporal features of medical data.

[0130] Example 4

[0131] As Figure 2 shown, this embodiment provides a data processing method for a multi-modal intelligent medical teaching and training system, including the following steps: Step 1, collect multi-source heterogeneous medical data, including electrocardiogram, electroencephalogram, electromyogram, blood oxygen, blood pressure, body temperature, respiration, motion state, and environmental parameters;

[0132] In the scenarios of medical electronics teaching and scientific research, medical data has heterogeneous characteristics, including continuous waveform data such as electrocardiogram (ECG), electromyogram (EMG), electroencephalogram (EEG), etc., interval sampling data such as blood pressure, blood oxygen, body temperature, etc., motion state data such as acceleration, angular velocity, posture, etc., and environmental parameter data such as temperature, humidity, light, air pressure, etc. There are significant differences in the sampling characteristics of different types of data. Among them, the sampling rate of bioelectric signals is greater than 500Hz, and high-precision sampling needs to be ensured; the sampling rate of physiological signals is in the range of 10 - 500Hz, and sampling stability needs to be ensured; the sampling rate of physiological parameters is less than 10Hz, and timing sampling accuracy needs to be ensured; at the same time, trigger-based acquisition of specific events such as human movement and body position change is also required.

[0133] There are significant differences in the feature expressions of multi-source medical data. In terms of time-domain features, electrocardiogram (ECG) signals contain features such as P waves (0.08 - 0.12 s), QRS complexes (0.06 - 0.10 s), and T waves (0.16 - 0.20 s), while electroencephalogram (EEG) signals contain features such as alpha waves (8 - 13 Hz) and beta waves (13 - 30 Hz); in terms of frequency-domain features, the main frequency band of ECG is 0.5 - 40 Hz, the main frequency band of EEG is 0.5 - 30 Hz, and the main frequency band of electromyogram (EMG) is 20 - 500 Hz; in terms of spatial features, it includes the spatial distribution of 12-lead ECG, the spatial distribution of multi-channel EEG, and the spatial arrangement of multi-point EMG. The above-mentioned feature differences pose higher requirements for the data acquisition module of the medical training system.

[0134] In this embodiment, the sensors adopt a modular design, including bioelectric signal sensors, physiological parameter sensors, motion state sensors, and environmental parameter sensors. Among them, the sampling rate of bioelectric signal sensors is greater than 500 Hz, and the accuracy is not less than 16 bits; the sampling rate of physiological parameter sensors is in the range of 10 - 500 Hz, and the accuracy is not less than 12 bits; the sampling rate of motion state sensors is greater than 100 Hz, and the accuracy is not less than 14 bits; the sampling rate of environmental parameter sensors is greater than 1 Hz, and the accuracy is not less than 10 bits.

[0135] Step two, preprocess the collected medical data, including data synchronization, which successively includes signal filtering, noise removal, baseline drift correction, and outlier removal;

[0136] Process the original data such as filtering, denoising, baseline drift correction, and outlier removal to improve the signal quality and achieve high-precision synchronization technology. At the same time, through the triple mechanisms of hardware clock synchronization, software timestamp synchronization, and adaptive delay compensation, ensure the precise synchronization of multi-source medical data, and the synchronization accuracy is better than 0.1 ms.

[0137] Step three, use a quality scoring function to evaluate the quality of medical data to obtain the data quality evaluation result. Use the data quality evaluation result to judge whether the collected data meets the requirements of subsequent data processing. If the evaluation result is higher than the set value, perform feature extraction on the corresponding medical data; if the evaluation result is lower than the set value, it indicates that the data quality is insufficient and it is necessary to improve the acquisition or perform data compensation to ensure the reliability of feature extraction and analysis processing in subsequent steps.

[0138] The quality scoring function is , where is the weight coefficient of the th type of medical data, is the quality score of the th type of medical data, is the diagnostic time window weighting function, is the clinical integrity indication function, is the total number of medical data categories.

[0139] The diagnostic time window weighting function takes a value of 1.2 in the emergency care monitoring scenario, 1.0 in the routine care monitoring scenario, 0.9 in the long-term monitoring scenario, and 0.8 in the health assessment scenario.

[0140] Clinical integrity indication function takes a value of 1.0 when the data in the critical interval is complete, 0.6 when there is a missing value in the non-critical interval, 0.3 when there is a missing value in the critical interval but interpolation is possible, and 0.0 when there is a serious missing value in the critical interval.

[0141] Step 4: Use a multi-branch neural network to extract medical features from the quality-assessed medical data. The multi-branch neural network includes a CNN branch, an RNN branch, and a GNN branch, which extract time-frequency features, temporal features, and spatial features from the medical features respectively;

[0142] In the CNN branch, a multi-scale convolutional structure is used to extract time-frequency features from high-frequency signals. The high-frequency signals include ECG / EMG waveforms with a sampling rate of 500 - 1000 Hz, and the time-frequency features include waveform morphology features.

[0143] The feature extraction function of the CNN branch is , where represents the output features extracted by the CNN branch, represents the feature obtained by convolving the input signal using the th convolutional kernel with a convolutional kernel length of , is the th weight coefficient of the convolutional kernel, is the bias term; the convolutional kernel length is designed according to the electrocardiogram waveform features: is used for QRS complex detection, is used for P wave detection, is used for T wave detection, is used for ST segment change detection.

[0144] The RNN branch adopts a bidirectional GRU structure to process the temporal features of physiological parameter sequences (0.1 - 100 Hz), and models the temporal dependence of vital signs through a gating mechanism; the calculation process of the RNN branch includes: , , , ; where, For the update gate, For the reset gate, For the candidate hidden state, For the current hidden state; For the Sigmoid activation function, 、 、 Are the weight matrices of the update gate, reset gate, and candidate hidden state respectively. The symbol Indicates element-wise multiplication of vectors (Hadamard product). For different types of physiological signals, different numbers of GRU hidden units and sequence lengths can be configured. For example, the heart rate sequence uses 64 hidden units and the sequence length is 300, the blood pressure sequence uses 32 hidden units and the length is 180, the blood oxygen saturation sequence uses 16 hidden units and the length is 120, and the body temperature sequence uses 8 hidden units and the length is 60.

[0145] The GNN branch constructs a Graph Attention Network (GAT) structure based on the anatomical spatial relationship of electrocardiogram / electroencephalogram multi-leads, calculates the association weights between nodes of different leads through the attention mechanism, and updates the node spatial features, expressed as:

[0146]

[0147] ;

[0148] Among them, Is the ReLU activation function with leakage, Is the learnable attention weight vector, Is the linear transformation matrix of the node features, Represents the node 's input feature vector, Represents the node 's updated feature vector. The symbol Indicates concatenating and into a vector, Is the set of neighbor nodes of the node . In the GNN branch, nodes represent individual measurement channels or sensor acquisition points in multi-lead (such as electrocardiogram or electroencephalogram) data. Each node reflects the original signal characteristics of a single lead or sensor, and the edge weights calculated by the graph attention mechanism reflect the anatomical or functional relationships between nodes.

[0149] Step 5: Use a fusion function to perform feature fusion on the branch medical features obtained in Step 4, combine the multi-modal features extracted from different branches, form a unified feature representation for subsequent early warning analysis and teaching effect evaluation, thereby improving the accuracy of analysis and achieving cross-modal information complementarity.

[0150] Adopt a dynamic weighted feature fusion mechanism considering disease diagnosis associations, and the fusion function is , = 3, where the weight , and corresponding weight parameters are configured for different disease states. For arrhythmia detection, weight configuration is adopted, and for heart failure early warning, weight configuration is adopted. For stroke, weight configuration is adopted.

[0151] MLP refers to a multi-layer fully connected neural network, which is used to process the concatenated feature vectors and output the weights of each modal feature;

[0152] represents the learnable parameters in the multi-layer fully connected neural network MLP, including weights and biases, which are used to adjust the mapping relationship of the network.

[0153] Figure 2 is a schematic diagram of the structure and data flow of a multi-branch deep learning model fusion system, including three parallel neural network branches and a feature fusion mechanism. Figure 2 Different colors are used in

[0154] to distinguish the various components of the system: yellow represents the input data, red represents the CNN branch, green represents the RNN branch, blue represents the GNN branch, purple represents the feature fusion layer, orange represents the disease-specific fusion algorithm, and gray represents the interpretability design and system optimization module. The thick arrow represents the main data flow, indicating the process of the multi-branch deep learning model working together to achieve intelligent analysis of medical data.

[0155] For different application scenarios such as arrhythmia detection, brain function evaluation, and hemodynamic analysis, the above fusion method is adopted, and the fusion accuracies reach 98.5%, 97.8%, and 96.9% respectively. , where , , are the weight coefficients of the electrocardiogram signal, blood oxygen signal, and blood pressure signal respectively, , , They respectively represent the feature vectors extracted from electrocardiogram, blood oxygen and blood pressure signals. By dynamically adjusting the weights of the features of each modality according to the pathological features of arrhythmia, the accurate identification of arrhythmia is realized, and the detection accuracy rate reaches 98.5%.

[0156] In the scenario of brain dysfunction assessment, the features of three modalities, electroencephalogram, eye movement and facial expression, are fused for comprehensive analysis. The fusion function is defined as , where 、 、 are the weight coefficients of electroencephalogram features, eye movement features and facial expression features respectively; represents the electroencephalogram features extracted from the electroencephalogram power spectrum, etc., which is used to evaluate the state of brain nerve activity, represents the features extracted by eye movement tracking, which is used to evaluate visual attention and cognitive function, represents the features extracted from facial expressions (such as facial symmetry), which is used to evaluate the state of facial nerve function. Through the complementary fusion of multi-modal features, the comprehensive assessment of brain dysfunction is realized, and the assessment accuracy rate reaches 96.8%.

[0157] In the scenario of hemodynamic analysis, the electrocardiogram, blood pressure and blood oxygen physiological signals are fused to realize the comprehensive analysis of the cardiovascular circulation state. The fusion function can adopt a form similar to that in arrhythmia detection, for example . The meanings of the above symbols are the same as those defined in the arrhythmia detection scenario. By dynamically adjusting the weight coefficients of each modality signal, the hemodynamic state can be accurately characterized, and the assessment of the patient's circulatory function can be realized, and the analysis accuracy rate reaches 95.2%. It can be seen that for different disease characteristics, different multi-modal fusion strategies can be selected, but the overall fusion framework remains the same.

[0158] Step 6: According to the medical feature fusion data, evaluate the teaching effect of the virtual-real combined medical scenario during the medical teaching process.

[0159] In this embodiment, a comprehensive score calculation function is used to fuse multiple evaluation indicators to quantify the comprehensive performance of learners. The comprehensive score calculation function is , where 、 、 are the first weight coefficient, the second weight coefficient and the third weight coefficient respectively, 、 、 are the calculation results of the virtual medical scenario training, physical operation training and clinical case analysis scoring functions respectively. Through the comprehensive score function, the evaluation results of the three links of virtual simulation training, physical operation training and clinical case analysis are weighted and fused to obtain a comprehensive evaluation of the teaching and training performance of learners.

[0160] The evaluation results of this step can be used to reverse-verify the effectiveness of the multi-modal training system, forming a complete teaching closed-loop.

[0161] The scoring function for virtual medical scenario training is , where is the completion score of the th virtual training item, is the item weight coefficient, is the total number of virtual training items, which can objectively reflect the operation performance of learners in the virtual environment.

[0162] The scoring function for physical operation training is , where is the completion quality score of the th physical operation item, is the time coefficient, is the operation weight coefficient, is the total number of physical operation items, which comprehensively considers the operation quality and time efficiency.

[0163] The scoring function for clinical case analysis evaluation is , where is the analysis score of the th clinical case, is the application ability coefficient, is the case weight coefficient, is the total number of clinical cases, which can comprehensively evaluate the clinical application ability of learners.

[0164] In terms of teaching evaluation, the evaluation results have a significant correlation with clinical practice performance, with a correlation coefficient > 0.85, which can accurately reflect the theoretical mastery level and practical operation ability of learners, providing a quantitative basis for teaching improvement.

[0165] In the second step, it also includes: synchronously processing multi-modal medical data for further alignment of different-modal medical data on the time axis, ensuring the accuracy of multi-source medical data fusion, including:

[0166] First, inject medical signal synchronization marks during the collection of medical data for calibrating and aligning multi-source medical data, expressed as: , where is the synchronization reference time, is the marking time interval, is the introduced random offset, represents the th inserted synchronization mark time point.

[0167] By introducing synchronization markers, it is ensured that physiological signals with different sampling rates can be accurately aligned, and the synchronization error is controlled within 1 ms.

[0168] Secondly, perform medical parameter cascaded adaptive filtering. The formula is , where is the filtered output of the th type of signal at discrete time , is the input value of the th type of signal at time , is the th coefficient of the adaptive filter for the th type of signal, is the signal correction factor based on medical knowledge, is the coefficient length of the filter.

[0169] The adaptive filtering process can adjust parameters according to the spectral differences of different physiological signals, effectively improve the signal quality, and increase the signal-to-noise ratio by no less than 20 dB.

[0170] Thirdly, perform physiological signal phase locking to further align the phases of different modality signals at the cycle level. The calculation formula is , where and are the occurrence time points of characteristic events of signals and signal within a certain synchronization cycle. For example, when signal is an electrocardiogram signal, can be set to represent the time when the electrocardiogram R wave appears, and are the cycle lengths of signals and signal , such as the time interval between adjacent characteristic events; represents the phase difference of signal relative to signal , expressed in degrees.

[0171] Through the above synchronization processing, precise phase alignment of multi-source physiological signals is achieved, and the relative phase error of cross-modal signals is controlled within ±5°. High-precision synchronization of multi-modal medical data is realized, providing an accurate and reliable time reference for feature fusion and teaching evaluation.

[0172] Step 7: Based on the multi-modal medical data after data fusion, use a scoring function to perform real-time monitoring of the multiple organ functions of critically ill patients and conduct early warning analysis, so as to achieve real-time assessment and early warning of the functional states of multiple important organs of critically ill patients.

[0173] ​The scoring function is , and the formula is as follows:

[0174]

[0175] Where is the function score of the th organ, is the weight coefficient of the th organ function, is the time sensitivity coefficient, is the inflammatory factor correction index; is the function trend score of the th organ, and the formula is . In this formula, the vertical bar means to take the absolute value. Therefore, the fraction

[0176]

[0177] is actually the sign function, and its result is:

[0178] When ,

[0179] When , ;

[0180] Where is the organ-specific time weight, is the change in the th organ function score between adjacent time points, that is, the current value minus the previous value; the early warning reserve time is calculated by the formula , where is the critical intervention threshold of the multi-organ function score. Reaching the threshold indicates that the patient's condition is extremely critical and immediate intervention is required. The corresponding critical time is , is the safety factor.

[0181] Step eight, use the evaluation model to evaluate the medical training data processing performance. The medical training data processing performance includes data processing efficiency and effect indicators. The data processing includes data synchronization and data monitoring, providing a basis for further optimization of medical training processing. The evaluation model includes:

[0182] The balance degree of medical multi-modal data processing, and the formula is , where is the number of medical data types, is the Processing efficiency of medical data For the importance weight of medical data is the medical relevance index, which is used to characterize the degree of relevance of different data types in medical decision-making. This indicator is used to measure the balance degree of different types of medical data processing, and the closer the value is to 1, the more balanced the processing of various data is.

[0183] Clinical model scheduling efficiency, the formula is where is the number of clinical models participating in the scheduling, is the clinical importance index of the th clinical model, is the th clinical model's inference time, is the maximum inference time among all clinical models, is the th clinical model's GPU / CPU utilization rate. By maximizing this indicator, the overall performance of the system can be improved.

[0184] Medical data processing throughput, the formula is where is the computing power of the th processing unit (referring to the computing resource unit for data processing, including CPU, GPU or dedicated hardware accelerator), such as the amount of data that can be processed per second, is the th disease type's difficulty factor, is the number of processing scenarios considered. This indicator represents the maximum rate at which the system can process medical data, which is jointly determined by the computing power of each processing module and the complexity of the diseases being processed, and the minimum value of the above product is taken as the evaluation indicator.

[0185] System response efficiency, the formula is where is the theoretically shortest response time, is the actual response time. Response efficiency reflects the ability to respond to inputs in real time, and the larger the value, the faster the response.

[0186] Diagnostic accuracy rate, the formula is where is the number of samples correctly diagnosed, is the total number of samples in the diagnostic test. Diagnostic accuracy rate is used to measure the accuracy of diagnostic results, and the closer the value is to 1, the more accurate the diagnosis is.

[0187] Step 9: Further optimize the synchronized multimodal data stream and the multi-branch neural network to further reduce latency and resource consumption, improve the overall performance, and ensure that the diagnostic accuracy does not decrease, including:

[0188] Clinical priority dynamic scheduling, the formula is , where is the weight of the th type of medical modality, is the generation frequency of medical data under this modality, is the clinical importance index corresponding to the th type of data, is the processing priority value of the th type of data. Through the dynamic scheduling of clinical priorities, the processing order of different data types can be dynamically adjusted according to clinical needs to ensure that key data is processed first.

[0189] Lightweight medical AI model, the formula is , where is the basic distillation loss weight, is the clinical knowledge distillation loss weight, is the output of the teacher model and the output of the student model The Kullback–Leibler divergence between them, where and represent the output probability distributions of the teacher model and the student model respectively, is the task loss function of the model, is the loss term to ensure the ability of the student model to identify key clinical features, is the total loss function of knowledge distillation. Through this model compression method, while maintaining the diagnostic performance of the model as much as possible, the computational resource requirements of the model are significantly reduced.

[0190] Through data processing optimization, while maintaining the diagnostic accuracy, the response time is reduced by 40%, the resource consumption is reduced by 50%, and the response time of medical staff intervention is shortened by 43%, significantly improving the patient's medical experience. Through the teaching evaluation and optimization method combining virtual and real, the improvement of medical teaching effect and the improvement of medical service quality are realized, providing strong support for medical talent cultivation and clinical practice.

[0191] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0192] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0194] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principles of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A multi-modal intelligent medical teaching and training system, characterized in that, It includes the following modules: Medical data acquisition module: used to acquire multi-source heterogeneous medical data; Preprocessing module: preprocess the acquired medical data, including data synchronization; Medical data quality assessment module: use a quality scoring function to evaluate the quality of medical data, obtain the data quality assessment result, and if the assessment result is higher than the set value, perform feature extraction on the corresponding medical data; Feature extraction module: use a multi-branch neural network to extract medical features from the medical data after quality assessment. The multi-branch neural network includes a CNN branch, an RNN branch, and a GNN branch, which respectively extract time-frequency features, temporal features, and spatial features in the medical features; Feature fusion module: use a fusion function to perform feature fusion on the medical features of each branch obtained by the feature extraction module, and combine the multi-modal features extracted by different branches to form a unified feature representation.

2. The multimodal intelligent medical teaching and training system according to claim 1, characterized in that, It also includes: Teaching effect evaluation module: evaluate the teaching effect of the virtual-real combined medical scenario during the medical teaching process based on the medical feature fusion data; Critical patient warning analysis module: based on the multi-modal medical data after data fusion, use a scoring function to perform real-time monitoring of the multi-organ functions of critical patients and conduct warning analysis; Data processing performance evaluation module: use an evaluation model to evaluate the data processing performance of medical training data. The data processing performance of medical training data includes data processing efficiency and effect indicators. The data processing includes data synchronization and data monitoring, providing a basis for further optimization of medical training processing; Optimization processing module: further optimize the synchronized multi-modal data stream and the multi-branch neural network to reduce latency and resource consumption.

3. The multimodal intelligent medical teaching and training system according to claim 1, wherein In the medical data acquisition module, the multi-source heterogeneous medical data includes: electrocardiogram, electroencephalogram, electromyogram, blood oxygen, blood pressure, body temperature, respiration, motion state, and environmental parameters; In the medical data quality assessment module, the quality scoring function is , where is the weight coefficient of the th type of medical data, is the quality score of the th type of medical data, is the diagnostic time window weighting function, is the clinical integrity indication function, is the total number of medical data categories.

4. A multimodal intelligent medical teaching and training system according to claim 1, characterized in that, In the feature extraction module, in the CNN branch, a multi-scale convolutional structure is adopted to extract time-frequency features of high-frequency signals. The feature extraction function of the CNN branch is , where represents the output features extracted by the CNN branch, represents the convolution of the input signal using the -th convolutional kernel with a convolutional kernel length to obtain the features, is the weight coefficient of the-th convolutional kernel, and is the bias term.​ 5. A multimodal intelligent medical teaching and training system according to claim 1, characterized in that, In the feature extraction module, the RNN branch adopts a bidirectional GRU structure to process the temporal features of the physiological parameter sequence and model the temporal dependence relationship of vital signs through a gating mechanism; The calculation process of the RNN branch includes: ; ; ; ; Among them, is the update gate, is the reset gate, is the candidate hidden state, is the current hidden state; is the Sigmoid activation function, 、 、 are the weight matrices of the update gate, the reset gate, and the candidate hidden state respectively. The symbol denotes element-wise multiplication of vectors.

6. The multimodal intelligent medical teaching and training system according to claim 1, wherein In the feature extraction module, the GNN branch constructs a graph attention network structure based on the anatomical spatial relationship of multiple leads of electrocardiogram / electroencephalogram, calculates the association weights between different lead nodes through an attention mechanism, and updates the node spatial features, expressed as: ; ; Among them, is the ReLU activation function with leakage, is the learnable attention weight vector, is the linear transformation matrix of node features, represents node 's input feature vector, represents node 's updated feature vector. The symbol means to and are concatenated as vectors, is the set of neighbor nodes of node .

7. A multimodal intelligent medical teaching and training system according to claim 1, characterized in that, In the feature fusion module, the fusion function is , = 3; Among them, the weights , corresponding weight parameters are configured for different disease states. For arrhythmia detection, the weight configuration of is adopted. For heart failure early warning, the weight configuration of is adopted. For stroke, the weight configuration of is adopted; MLP refers to a multi-layer fully connected neural network, which is used to process the concatenated feature vectors and output the weights of each modal feature; Represent learnable parameters in a multi-layer fully connected neural network MLP, including weights and biases.

8. The multimodal intelligent medical teaching and training system according to claim 2, wherein In the teaching effect evaluation module, a comprehensive scoring calculation function is used to integrate multiple evaluation indicators. The comprehensive scoring calculation function is , where , , are the first weight coefficient, the second weight coefficient, and the third weight coefficient respectively, , , are the calculation results of the virtual medical scenario training, physical operation training, and clinical case analysis scoring functions respectively; The virtual medical scenario training scoring function is , where is the completion score of the th virtual training item, is the item weight coefficient, is the total number of virtual training items; The scoring function for hands-on training is , where is the completion quality score of the th hands-on operation item, is the time coefficient, is the operation weight coefficient, is the total number of hands-on operation items; The clinical case analysis evaluation scoring function is , where is the analysis score of the th clinical case, is the application ability coefficient, is the case weight coefficient, is the total number of clinical cases.

9. A multi-modal intelligent medical teaching and training system according to claim 1, characterized in that, In the data preprocessing module, the multi-modal medical data is synchronized for alignment on the time axis of different modal medical data, including: Inject a medical signal synchronization marker when collecting medical data for calibrating and aligning multi-source medical data, expressed as: , where is the synchronization reference time, is the marker time interval, is the introduced random offset, represents the th inserted synchronization marker time point; Perform medical parameter cascaded adaptive filtering, and the formula is , where is the filtering output of the th type of signal at discrete time , is the input value of the th type of signal at time , is the th coefficient of the adaptive filter of the th type of signal, is the signal correction factor based on medical knowledge, is the coefficient length of the filter; Perform physiological signal phase locking to further align the phases of different modality signals at the cycle level. The calculation formula is , where and are the occurrence time points of characteristic events of signal and signal within a certain synchronization cycle, and are the cycle lengths of signal and signal respectively; represents the phase difference of signal relative to signal .

10. A multimodal intelligent medical teaching and training system according to claim 2, characterized in that, In the critical patient warning analysis module, the scoring function is , and the formula is as follows: ; Among them, is the functional score of the th organ, is the weight coefficient of the th organ function, is the time sensitivity coefficient, is the inflammatory factor correction index; is the functional trend score of the th organ, and the formula is , and the vertical bar means taking the absolute value; is the organ-specific time weight, is the change in the th organ function score between adjacent time points; the early warning reserved time is calculated as , where is the critical intervention threshold of the multiple organ function score, and the corresponding critical time is , is the safety factor.

11. A multimodal intelligent medical teaching and training system according to claim 2, characterized in that, In the data processing performance evaluation module, the evaluation model includes: The balance degree of medical multi-modal data processing, the formula is , where is the number of medical data types, is the processing efficiency of the th type of medical data, is the importance weight of the th type of medical data, is the medical relevance index, which is used to characterize the degree of relevance of different data types in medical decision-making; Clinical model scheduling efficiency, with the formula being , where is the number of clinical models participating in scheduling, is the th clinical importance index of the clinical model, is the th inference time of the clinical model, is the maximum inference time among all clinical models, is the th GPU / CPU utilization rate of the clinical model; Medical data processing throughput, with the formula , where is the computing power of the th processing unit, is the difficulty factor for the th type of disease, is the number of processing scenarios considered; System response efficiency, the formula is , where is the theoretically shortest response time, is the actual response time; Diagnostic accuracy rate, with the formula being , where is the number of samples with correct diagnosis, is the total number of samples in the diagnostic test.

12. A multimodal intelligent medical teaching and training system according to claim 2, characterized in that, In the optimization processing module, it includes: Clinical priority dynamic scheduling, the formula is , where is the weight of the th type of medical modality, is the generation frequency of medical data under this modality, is the clinical importance index corresponding to the th type of data, is the processing priority value of the th type of data; The lightweighting of the medical AI model has the formula , where is the weight of the basic distillation loss, is the weight of the clinical knowledge distillation loss, is the output of the teacher model and is the Kullback–Leibler divergence between the output of the student model, where and represent the output probability distributions of the teacher model and the student model respectively, is the task loss function of the model, is the loss term to ensure the ability of the student model to identify key clinical features, is the total loss function of knowledge distillation.

13. A data processing method for a multi-modal intelligent medical teaching and training system, characterized in that It includes the following steps: Step 1: used to acquire multi-source heterogeneous medical data; Step 2: preprocess the acquired medical data, including data synchronization; Step 3: use a quality scoring function to evaluate the quality of medical data, obtain the data quality assessment result, and if the assessment result is higher than the set value, perform feature extraction on the corresponding medical data; Step 4: Use a multi-branch neural network to extract medical features from the quality-assessed medical data. The multi-branch neural network includes a CNN branch, an RNN branch, and a GNN branch, which respectively extract time-frequency features, temporal features, and spatial features from the medical features; Step 5: Adopt a fusion function to perform feature fusion on the medical features of each branch obtained by the feature extraction module, combine the multi-modal features extracted by different branches, and form a unified feature representation.

14. The data processing method of a multi-modal intelligent medical teaching and training system according to claim 13, characterized in that, It also includes: Step 6: According to the medical feature fusion data, evaluate the teaching effect of the virtual-real combined medical scenario during the medical teaching process; Step 7: Based on the multi-modal medical data after data fusion, use a scoring function to perform real-time monitoring of the multi-organ functions of critically ill patients and conduct early warning analysis; Step 8: Use an evaluation model to evaluate the processing performance of medical training data. The processing performance of medical training data includes data processing efficiency and effect indicators. The data processing includes data synchronization and data monitoring, providing a basis for further optimization of medical training processing; Step 9: Further optimize the synchronized multi-modal data stream and the multi-branch neural network to reduce latency and resource consumption.

Citation Information

Patent Citations

  • Intelligent medical system based on deep learning

    CN113724853A

  • Multi-modal data fusion method based on semantic information amount and application

    CN115470856A

Cited By

  • Network security state characterization method, network security intelligent decision-making method and network security intelligent decision-making device

    CN120750632A