Chest pain classification method and system based on multi-modal data fusion and deep learning model

Through multimodal data fusion and deep learning models, combined with chest X-ray and electrocardiogram preprocessing and enhancement, and dynamic adjustment of expert weights, the limitations and interpretability problems of single-modality data processing in chest pain classification are solved, and efficient and accurate chest pain diagnosis support is achieved.

CN120670936AActive Publication Date: 2025-09-19AFFILIATED HOSPITAL OF NANTONG UNIV

Patent Information

Application Number
CN202510701353.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-19
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing deep learning models have limitations in single-modality data processing, data scarcity and imbalance, and lack of interpretability in chest pain classification, resulting in inaccurate classification results and insufficient robustness.

Method used

By adopting multimodal data fusion and deep learning models, preprocessing and enhancing chest X-ray images and electrocardiograms, a multi-scale visual model and a multi-scale time series network are constructed. Combined with a dynamic gated hybrid expert system, expert weights are dynamically adjusted and clinical knowledge is introduced to generate an explainable diagnostic report.

Benefits of technology

It achieves efficient fusion of multimodal data, improves the accuracy and robustness of chest pain classification, and provides explainable diagnostic support to meet the timeliness requirements of emergency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670936A_ABST
    Figure CN120670936A_ABST
Patent Text Reader

Abstract

The invention relates to a chest pain classification method and system based on multi-modal data fusion and a deep learning model, and belongs to the technical field of intelligent medical auxiliary diagnosis. The method comprises the following steps: firstly, carrying out heart region segmentation and pathological feature enhancement preprocessing on a chest radiograph image, extracting anatomical features by adopting a multi-scale visual model which is optimized and improved aiming at the chest radiograph, meanwhile, carrying out multi-lead time sequence calibration and ST-segment waveform marking on an electrocardiogram signal, and capturing electrophysiological time sequence features through a multi-scale one-dimensional convolutional network; then, a dynamic gating hybrid expert system is constructed, an anatomy expert, an electrophysiology expert, a multi-modal association expert and a critical value expert process specific modal features respectively, and a gating network dynamically calculates expert weights based on real-time vital signs and pathological features; according to the method, by integrating multi-modal data and combining a deep learning technology, rapid and accurate chest pain classification support and report generation can be provided for clinicians, and the method has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a chest pain classification method and system based on multimodal data fusion and deep learning models, belonging to the technical field of intelligent medical auxiliary diagnosis. Background Art

[0002] Chest pain is a common and important symptom in thoracic surgery. It often involves multiple potential causes, including but not limited to coronary heart disease, myocardial infarction, lung disease, and other thoracic diseases. Accurate diagnosis of chest pain is crucial for timely management and treatment, especially in the case of acute chest pain, where rapid identification of the cause directly impacts patient safety and treatment outcomes. Currently, the classification and diagnosis of chest pain typically relies on physician experience and a range of traditional examination methods, such as electrocardiograms (ECGs), chest X-rays, computed tomography (CT), and laboratory tests. While these methods provide valuable clinical information, they often have limitations. First, the clinical manifestations of chest pain can be highly similar, and symptoms of chest pain due to different etiologies may not be clearly distinguishable. Therefore, it is difficult to accurately determine the cause in a short period of time based solely on physician experience or a single examination result. Second, traditional imaging and testing methods rely on manual analysis, and the results are limited by the physician's expertise and workload, which may lead to missed or misdiagnoses. Furthermore, certain acute heart conditions, such as myocardial infarction, require early diagnosis and urgent treatment. Therefore, rapid diagnosis of the cause of chest pain is crucial to improving patient survival.

[0003] With the rapid development of artificial intelligence (AI) technology, particularly the widespread application of deep learning in medical imaging and signal processing, research on chest pain classification is gradually moving towards automation. Deep learning can extract complex feature information from medical images such as chest X-rays and electrocardiograms, analyze and classify them, overcoming the limitations of traditional manual diagnosis. In recent years, an increasing number of studies have attempted to use deep learning techniques, particularly models such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to automatically analyze chest image and electrocardiogram data and identify the cause and type of chest pain. Automated diagnostic systems based on deep learning can significantly improve the efficiency and accuracy of chest pain diagnosis, providing real-time decision support for clinicians.

[0004] However, despite the significant progress made in medical imaging and signal processing by deep learning, existing methods still face several challenges. First, most existing methods focus on single-modality data processing and cannot effectively fuse information from multiple sources (such as chest X-ray and electrocardiogram data), which limits the accuracy and robustness of the classification results. Second, the training of existing deep learning models mostly relies on a large amount of labeled data, and the chest pain classification task often faces the problem of data scarcity or data imbalance, resulting in poor performance of the model when dealing with some more special causes of chest pain. More importantly, although deep learning technology has outstanding performance in classification accuracy, it lacks sufficient interpretability, which is particularly critical for medical applications. Clinicians need to be able to understand the decisions made by the model in order to trust and adopt the system's recommendations in complex medical scenarios. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides a chest pain classification method and system based on multimodal data fusion and deep learning models. First, the chest X-ray image is preprocessed with cardiac region segmentation and pathological feature enhancement. An optimized and improved multi-scale visual model (Vit-Med) for chest X-ray is used to extract anatomical features. At the same time, multi-lead timing calibration and ST segment waveform labeling are performed on the electrocardiogram signal, and electrophysiological timing features are captured through a multi-scale one-dimensional convolutional network. Subsequently, a dynamic gated hybrid expert system (DG-MOE) is constructed, in which anatomical experts, electrophysiological experts, multimodal association experts, and critical value experts respectively process specific modal features, and the gating network dynamically calculates expert weights based on real-time vital signs (such as heart rate and blood pressure) and pathological features (such as ST segment offset).

[0006] The technical solutions of the present invention are as follows:

[0007] A chest pain classification method based on multimodal data fusion and deep learning model, the steps are as follows:

[0008] (1) Preprocessing and enhancing chest X-ray images and electrocardiogram multimodal data;

[0009] (2) anatomical-temporal dual-stream feature extraction;

[0010] (3) Clinical knowledge guides the dynamic integration of MOE;

[0011] (4) Diagnostic reasoning driven by large models.

[0012] According to the preferred embodiment of the present invention, in step (1), the specific steps are:

[0013] To address the heterogeneity and noise issues of medical multimodal data, this step constructs a preprocessing pipeline guided by clinical prior knowledge. This preprocesses chest radiographs and uses a pretrained U-Net network to segment key anatomical regions, including the heart and lung fields, generating binary masks to guide ROI standardization. Adaptive window width and position adjustment (window position 40HU, window width 400HU) is used to enhance the contrast of mediastinal structures, and pathology-preserving data enhancement (±15° rotation and elastic deformation of lung field texture) is combined to improve model robustness.

[0014] Electrocardiogram (ECG) image preprocessing involves aligning 12-lead signals based on dynamic time warping (DTW), and eliminating baseline drift using wavelet threshold denoising (db6 wavelet, 5-layer decomposition). Key ST segment / T wave waveforms are detected using an LSTM-CRF model, generating a time series annotation matrix (e.g., marking the start / end of the ST segment), and applying elastic time warping (ETW) enhancement technology to simulate heart rate variability. After preprocessing, the data is stored in tensor form: 512×512×1 for chest X-rays and 12 leads × 2560 sampling points for ECGs.

[0015] According to the preferred embodiment of the present invention, in step (2), the specific process is:

[0016] (21) A multi-scale anatomical perception pyramid network is constructed based on the anatomical features of chest X-ray images. Global and local features are extracted from chest X-ray images, and cross-block attention fusion is performed. The attention of the pathological area of ​​the chest X-ray image is then enhanced, and the medical semantic position encoding of the chest X-ray image is performed.

[0017] (22) A time series feature extraction method is constructed for electrocardiogram, and a multi-scale time series convolutional pyramid network is constructed. Feature extraction is performed from three aspects, namely short-time feature extraction, medium-time feature extraction, and long-time feature modeling. Feature fusion is then performed, and a total loss function is constructed.

[0018] According to the preferred embodiment of the present invention, in step (21), the specific steps are:

[0019] A multi-scale anatomical-aware pyramid network was constructed based on the anatomical features of chest X-ray images. Global features were extracted from chest X-ray images, which were segmented into 16×16 pixel blocks (each 256×256 area) and fed into a standard Vision Transformer (ViT) module. A multi-head self-attention mechanism was used to capture the overall structure of the chest cavity (such as the cardiac outline and lung field distribution). Local features were then extracted from the chest X-rays using a dilated convolution (dilation = 2) on 8×8 pixel blocks (each 128×128 area) to enhance the capture of detailed features such as vascular texture and calcifications.

[0020] Cross-Patch Attention Fusion: Global features (16×16 blocks) are concatenated with local features (8×8 blocks), and features of different scales are fused through the Cross-Patch Attention (CPA) module. The formula is:

[0021]

[0022] in, is the query matrix of global features, is the key and value matrix of local features, M ROI ∈{0, 1} N×M is the anatomical mask matrix, which only allows feature interactions in the heart / lung field region, d k = d / h is the dimension of each attention head. Through multi-scale fusion, it can simultaneously capture the global anatomical structure and local pathological details, thereby improving the sensitivity to small lesions (such as 2mm lung nodules);

[0023] The attention of the pathological area of ​​the chest X-ray image is enhanced, the anatomical mask is generated, and the pre-trained U-Net segmentation network outputs the binary mask M of key areas such as the heart and the hilum patho ∈{0, 1} H×W , attention bias injection, in the self-attention calculation, the anatomical mask is used as a bias term to strengthen the model's attention to the pathological area:

[0024]

[0025] in, is a block-level mask matrix, which is set to 1 if the block contains a pathological area, and 0 otherwise; λ is a learnable parameter with an initial value of 0.5. The attention intensity of the pathological area is adjusted through backpropagation. By enhancing the attention to the pathological area, the model's attention to areas such as lung field texture is improved;

[0026] Perform medical semantic position coding on chest X-ray images. Predefine six anatomical landmarks (such as the cardiac apex c1 = (0.42, 0.67) and the aortic arch apex c2 = (0.38, 0.45)) as anatomical key points. Perform weighted coding on the inverse distance and calculate the distance between the center coordinates p = (x, y) of each image block and the anatomical key points to generate the semantic position coding:

[0027]

[0028] Among them, c k ∈[0, 1} 2 is the normalized coordinate of the kth key point, is a learnable weight matrix that maps distance to feature space, ||pc k||2 is the Euclidean distance, which measures the proximity of the image patch to the anatomical landmark.

[0029] According to the preferred embodiment of the present invention, in step (22), the specific steps are:

[0030] A time series feature extraction (MS-TCN module) is constructed for the electrocardiogram, and a multi-scale time series convolutional pyramid network is built. Feature extraction is performed from three aspects: short-time feature extraction, medium-time feature extraction, and long-time feature modeling. Among them, short-time feature extraction uses a 5ms convolution kernel (kernel size = 128 sampling points) to capture QRS wave details (such as R wave peak and Q wave depth); medium-time feature extraction uses a 15ms convolution kernel (kernel size = 384 sampling points) to analyze ST segment morphological change trends (such as elevation / depression slope); long-time feature modeling uses a bidirectional LSTM network (hidden layer 128 dimensions) to capture rhythm features such as heart rate variability (HRV). On this basis, feature fusion is performed, and the multi-scale features are spliced ​​and then reduced in dimension through 1×1 convolution. The formula is:

[0031] F fused =Conv1D 1×1 (Concat(F short , F mid , F long )) (4)

[0032] in, is a feature tensor of different time scales, and then a clinical waveform constraint loss function is constructed to align the ST segment features and constrain the ST segment features f predicted by the model. ST The actual value marked by the doctor The error is:

[0033]

[0034] in, is the ST segment feature of the i-th sample output by the model, In order to extract the true ST segment features (labeling interval ±20ms) through the waveform annotation matrix, the features are constrained by waveform classification consistency, and the KL divergence is used to ensure that the waveform probability distribution p of the model output is wave Distribution with annotations Consistent:

[0035]

[0036] in, is the QRS / ST / T wave probability predicted by the model, is the real waveform distribution of the annotation;

[0037] Construct a total loss function to jointly optimize feature alignment and distribution consistency:

[0038]

[0039] Among them, α=0.1 is the weight to balance the two losses. is the loss value of the ST segment, is the loss between the image and the annotation.

[0040] According to the preferred embodiment of the present invention, in step (3), the specific process is:

[0041] (31) Construct four specialized expert models to handle specific modalities or tasks respectively;

[0042] (32) Construction of dynamic gating network;

[0043] (33) Perform multimodal feature fusion and splice features of multiple expert models.

[0044] According to the preferred embodiment of the present invention, step (31) is specifically as follows: constructing an anatomical expert model, based on the 3D-ResNet50 network structure, inputting multi-scale features of the chest X-ray (768 dimensions), and outputting anatomical abnormality scores (such as cardiac hypertrophy index and lung field translucency), for example, when the chest X-ray detects mediastinal widening (>3cm) or signs of pulmonary edema, it is activated; constructing an electrophysiological expert model, based on Transformer-ECG, inputting ECG timing features (256 dimensions), and outputting electrophysiological features such as ST segment offset and QRS complex width. Parameters, forced activation when the ST segment deviation of any ECG lead is ≥1mm; construct a multimodal association expert model, based on the graph attention network (GAT), model the cross-modal association between the chest X-ray area (node) and the ECG waveform (edge), and activate it when the confidence difference between the chest X-ray and ECG features is >20%; construct a critical value expert model based on a lightweight CNN (4-layer convolution), input vital signs (heart rate, blood pressure) and pathological markers (such as troponin), and force activation when the heart rate is >120 beats / min or the systolic blood pressure is <90mmHg.

[0045] According to the preferred embodiment of the present invention, step (32) is specifically as follows:

[0046] The gating network dynamically adjusts the expert weights according to real-time clinical data. It is divided into two parts: data-driven and rule-driven. The first is data-driven weight calculation, which inputs the multimodal feature h in and vital signs (heart rate, blood pressure, blood oxygen), build a gating function:

[0047]

[0048] Among them, W g To dynamically calculate the weights, σ ​​is a variable parameter, and then the weights are modified by rules, imposing hard constraints when specific clinical events are detected:

[0049]

[0050] Among them, the gating function is different under different constraint conditions. When ST elevation is greater than or equal to 2 mm and k = 2, the gating constraint is 0.7; when heart rate is greater than 120 and k = 4, the gating constraint is 0.5; in other cases, it is Then the two weights are synthesized by the synthesis method:

[0051]

[0052] Among them, σ is the Sigmoid function, which limits the weight to the interval [0,1], and α is a learnable parameter that balances the contribution of data and rules (initial value 0.6). is the gating network weight matrix.

[0053] According to the preferred embodiment of the present invention, step (33) is specifically as follows:

[0054] Perform multimodal feature fusion, splice features of multiple expert models, and perform splicing based on residual connections:

[0055]

[0056] h final =h in +W r ·h expert (12)

[0057] Among them, E k (·) is the feature extraction function of expert k, is the residual projection matrix, is the final fused feature vector with a dimension of 1024.

[0058] According to the preferred embodiment of the present invention, in step (4), the specific process is:

[0059] Construct an input feature concatenation method and convert the 1024-dimensional h output from step (3) into final Feature vector, 768-dimensional vector t encoded with the patient's electronic medical record text (chief complaint, medical history, etc.) text To splice:

[0060] h joint =W c Concat(h final , t text ) (13)

[0061] in, is a dimensionality reduction matrix (1792=1024+768), Concat is a concatenation method, and the output is

[0062] Diagnostic decision generation uses causal attention decoding to generate diagnostic content. The LLaMA-Med Transformer decoder is used to generate the diagnostic logic chain. Attention calculation introduces medical knowledge constraints:

[0063]

[0064] Among them, M rule ∈{0, 1} L×L is the clinical rule mask matrix, L is the sequence length, d k =2048 / 32=64 is the dimension of each attention head;

[0065] Structured output generation, the model output consists of three parallel headers: disease probability header, treatment recommendation header, and differential diagnosis header. The disease probability header is:

[0066] p disease =Softmax(W p ·h in ) (15)

[0067] in, is the four-category weight (pulmonary embolism, coronary heart disease, normal, aortic stratum);

[0068] The disposal suggestion header is:

[0069] s action =Sigmoid(W a ·h in ) (16)

[0070] in, There are 12 standard actions (e.g., “nitroglycerin sublingual administration”, “catheterization laboratory activation”, etc.);

[0071] The differential diagnosis heads are:

[0072]

[0073] Among them, ε DDx It is an embedding vector library containing 30 chest pain related diseases.

[0074] Generate a natural language decision chain based on the above steps:

[0075] Generate report templates that are in line with clinical thinking.

[0076] A chest pain classification system based on multimodal data fusion and deep learning model, including:

[0077] A preprocessing module is used to preprocess and enhance chest X-ray images and electrocardiogram multimodal data;

[0078] Extraction module, used for anatomical-temporal dual-stream feature extraction;

[0079] Fusion module, used to guide the dynamic fusion of MOE based on clinical knowledge;

[0080] Reasoning module, used for large model-driven diagnostic reasoning.

[0081] The beneficial effects of the present invention are:

[0082] The present invention combines multimodal data such as chest X-rays and electrocardiograms, analyzes the anatomical features of chest X-rays and the timing features of electrocardiograms in real time through a dynamic gating network, adaptively adjusts expert weights based on vital signs (e.g., the weight of electrophysiology experts is increased to 70% when the ST segment is elevated), uses multimodal comparative learning to achieve cross-modal semantic alignment, and embeds clinical rule constraints (e.g., a heart rate >120 beats / min triggers a critical value expert) to ensure diagnostic logic compliance. At the same time, it integrates a large medical model to generate an interpretable report containing a bimodal heat map and a natural language reasoning chain. By integrating multimodal data and combining deep learning technology, the method of the present invention can provide clinicians with fast and accurate chest pain classification support and report generation, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0084] The present invention will be further described below with reference to embodiments and accompanying drawings, but is not limited thereto.

[0085] Example 1:

[0086] This implementation provides a chest pain classification method based on multimodal data fusion and deep learning model. The steps are as follows:

[0087] (1) Preprocess and enhance the chest X-ray image and electrocardiogram multimodal data. The specific steps are as follows:

[0088] To address the heterogeneity and noise issues of medical multimodal data, this step constructs a preprocessing pipeline guided by clinical prior knowledge. This preprocesses chest radiographs and uses a pretrained U-Net network to segment key anatomical regions, including the heart and lung fields, generating binary masks to guide ROI standardization. Adaptive window width and position adjustment (window position 40HU, window width 400HU) is used to enhance the contrast of mediastinal structures, and pathology-preserving data enhancement (±15° rotation and elastic deformation of lung field texture) is combined to improve model robustness.

[0089] Electrocardiogram (ECG) image preprocessing involves aligning 12-lead signals based on dynamic time warping (DTW), and eliminating baseline drift using wavelet threshold denoising (db6 wavelet, 5-layer decomposition). Key ST segment / T wave waveforms are detected using an LSTM-CRF model, generating a time series annotation matrix (e.g., marking the start / end of the ST segment), and applying elastic time warping (ETW) enhancement technology to simulate heart rate variability. After preprocessing, the data is stored in tensor form: 512×512×1 for chest X-rays and 12 leads × 2560 sampling points for ECGs.

[0090] (2) Anatomy-temporal dual-stream feature extraction, the specific process is as follows:

[0091] (21) A multi-scale anatomical perception pyramid network is constructed based on the anatomical features of chest X-ray images. Global features of chest X-ray images are extracted and divided into 16×16 pixel blocks (each block has a 256×256 area). The images are then fed into a standard Vision Transformer (ViT) module. The overall structure of the chest cavity (such as the heart contour and lung field distribution) is captured through a multi-head self-attention mechanism. Local features of the chest X-ray are then extracted. Dilated convolution (dilation = 2) is used to process 8×8 pixel blocks (each block has a 128×128 area) to enhance the ability to capture detailed features such as vascular texture and calcification points.

[0092] Cross-Patch Attention Fusion: Global features (16×16 blocks) are concatenated with local features (8×8 blocks), and features of different scales are fused through the Cross-Patch Attention (CPA) module. The formula is:

[0093]

[0094] in, is the query matrix of global features, is the key and value matrix of local features, M ROI ∈{0, 1} N×M is the anatomical mask matrix, which only allows feature interactions in the heart / lung field region, d k = d / h is the dimension of each attention head. Through multi-scale fusion, it can simultaneously capture the global anatomical structure and local pathological details, thereby improving the sensitivity to small lesions (such as 2mm lung nodules);

[0095] The attention of the pathological area of ​​the chest X-ray image is enhanced, the anatomical mask is generated, and the pre-trained U-Net segmentation network outputs the binary mask M of key areas such as the heart and the hilum patho ∈{0, 1} H×W , attention bias injection, in the self-attention calculation, the anatomical mask is used as a bias term to strengthen the model's attention to the pathological area:

[0096]

[0097] in, is a block-level mask matrix, which is set to 1 if the block contains a pathological area, and 0 otherwise; λ is a learnable parameter with an initial value of 0.5. The attention intensity of the pathological area is adjusted through backpropagation. By enhancing the attention to the pathological area, the model's attention to areas such as lung field texture is improved;

[0098] Perform medical semantic position coding on chest X-ray images. Predefine six anatomical landmarks (such as the cardiac apex c1 = (0.42, 0.67) and the aortic arch apex c2 = (0.38, 0.45)) as anatomical key points. Perform weighted coding on the inverse distance and calculate the distance between the center coordinates p = (x, y) of each image block and the anatomical key points to generate the semantic position coding:

[0099]

[0100] Among them, c k ∈[0, 1] 2 is the normalized coordinate of the kth key point, is a learnable weight matrix that maps distance to feature space, ||pc k ||2 is the Euclidean distance, which measures the proximity of the image patch to the anatomical landmark;

[0101] (22) A time series feature extraction (MS-TCN module) is constructed for the electrocardiogram, and a multi-scale time series convolutional pyramid network is constructed. Feature extraction is performed from three aspects, namely short-time feature extraction, medium-time feature extraction, and long-time feature modeling. Among them, short-time feature extraction is a 5ms convolution kernel (kernel size = 128 sampling points) to capture QRS wave details (such as R wave peak and Q wave depth); medium-time feature extraction is a 15ms convolution kernel (kernel size = 384 sampling points) to analyze the trend of ST segment morphology changes (such as elevation / depression slope); long-time feature modeling is a bidirectional LSTM network (hidden layer 128 dimensions) to capture rhythm features such as heart rate variability (HRV). On this basis, feature fusion is performed, and the multi-scale features are spliced ​​and reduced in dimension through 1×1 convolution. The formula is:

[0102] F fused =Conv1D 1×1 (Concat(F short , F mid , F long )) (4)

[0103] in, is a feature tensor of different time scales, and then a clinical waveform constraint loss function is constructed to align the ST segment features and constrain the ST segment features f predicted by the model. STThe actual value marked by the doctor The error is:

[0104]

[0105] in, is the ST segment feature of the i-th sample output by the model, In order to extract the true ST segment features (labeling interval ±20ms) through the waveform annotation matrix, the features are constrained by waveform classification consistency, and the KL divergence is used to ensure that the waveform probability distribution p of the model output is wave Distribution with annotations Consistent:

[0106]

[0107] in, is the QRS / ST / T wave probability predicted by the model, is the real waveform distribution of the annotation;

[0108] Construct a total loss function to jointly optimize feature alignment and distribution consistency:

[0109]

[0110] Among them, α=0.1 is the weight to balance the two losses. is the loss value of the ST segment, is the loss between the image and the annotation.

[0111] (3) Clinical knowledge guides the dynamic integration of MOE. The specific process is as follows:

[0112] (31) Construct four specialized expert models to handle specific modalities or tasks respectively;

[0113] An anatomical expert model was constructed based on the 3D-ResNet50 network structure. The model took in multi-scale features of the chest X-ray (768 dimensions) and output anatomical abnormality scores (such as cardiac hypertrophy index and lung field translucency). For example, the model was activated when the chest X-ray detected mediastinal widening (>3cm) or signs of pulmonary edema. An electrophysiological expert model was constructed based on the Transformer-ECG model. The model took in ECG time series features (256 dimensions) and output electrophysiological parameters such as ST segment deviation and QRS complex width. When any ECG lead Forced activation when the ST segment deviation is ≥1mm; construct a multimodal association expert model based on the graph attention network (GAT) to model the cross-modal association between the chest X-ray area (node) and the ECG waveform (edge), and activate it when the confidence difference between the chest X-ray and ECG features is >20%; construct a critical value expert model based on a lightweight CNN (4-layer convolution), input vital signs (heart rate, blood pressure) and pathological markers (such as troponin), and force activation when the heart rate is >120 beats / min or the systolic blood pressure is <90mmHg.

[0114] (32) Construction of dynamic gating network;

[0115] The gating network dynamically adjusts the expert weights according to real-time clinical data. It is divided into two parts: data-driven and rule-driven. The first is data-driven weight calculation, which inputs the multimodal feature h in and vital signs (heart rate, blood pressure, blood oxygen), build a gating function:

[0116]

[0117] Among them, W g To dynamically calculate the weights, σ ​​is a variable parameter, and then the weights are modified by rules, imposing hard constraints when specific clinical events are detected:

[0118]

[0119] Among them, the gating function is different under different constraint conditions. When ST elevation is greater than or equal to 2 mm and k = 2, the gating constraint is 0.7; when heart rate is greater than 120 and k = 4, the gating constraint is 0.5; in other cases, it is Then the two weights are synthesized by the synthesis method:

[0120]

[0121] Among them, σ is the Sigmoid function, which limits the weight to the interval [0,1], and α is a learnable parameter that balances the contribution of data and rules (initial value 0.6). is the gating network weight matrix.

[0122] (33) Perform multimodal feature fusion, splice features of multiple expert models, and perform splicing based on residual connections:

[0123]

[0124] h final =h in +W r ·h expert (12)

[0125] Among them, E k (·) is the feature extraction function of expert k, is the residual projection matrix, is the final fused feature vector with a dimension of 1024.

[0126] (4) Diagnostic reasoning driven by large models.

[0127] Construct an input feature concatenation method and convert the 1024-dimensional h output from step (3) into final Feature vector, 768-dimensional vector t encoded with the patient's electronic medical record text (chief complaint, medical history, etc.) text To splice:

[0128] h joint =W c Concat(h final , t text ) (13)

[0129] in, is a dimensionality reduction matrix (1792=1024+768), Concat is a concatenation method, and the output is

[0130] Diagnostic decision generation uses causal attention decoding to generate diagnostic content. The LLaMA-Med Transformer decoder is used to generate the diagnostic logic chain. Attention calculation introduces medical knowledge constraints:

[0131]

[0132] Among them, M rule ∈{0, 1} L×L is the clinical rule mask matrix, L is the sequence length, d k =2048 / 32=64 is the dimension of each attention head;

[0133] Structured output generation, the model output consists of three parallel headers: disease probability header, treatment recommendation header, and differential diagnosis header. The disease probability header is:

[0134] p disease=Softmax(W p ·h in ) (15)

[0135] in, is the four-category weight (pulmonary embolism, coronary heart disease, normal, aortic stratum);

[0136] The disposal suggestion header is:

[0137] s action =Sigmoid(W a ·h in ) (16)

[0138] in, There are 12 standard actions (e.g., “nitroglycerin sublingual administration”, “catheterization laboratory activation”, etc.);

[0139] The differential diagnosis heads are:

[0140]

[0141] Among them, ε DDx It is an embedding vector library containing 30 chest pain related diseases.

[0142] Generate a natural language decision chain based on the above steps:

[0143] Generate report templates that are in line with clinical thinking.

[0144] [Key Findings]

[0145] -ECG: ST segment elevation 2.3mm in leads V1-V4 (weighted by electrophysiologist 0.82)

[0146] - Chest X-ray: Positive hilar butterfly sign (anatomy expert weight 0.71)

[0147] [Diagnosis conclusion]

[0148] - Acute anterior myocardial infarction (94.7%), TIMI risk score 5

[0149] [Disposal suggestions]

[0150] 1. Immediately activate the catheterization lab (98% confidence)

[0151] 2. Chewed aspirin 300mg (95% confidence level)

[0152] [Differential diagnosis]

[0153] - Pulmonary embolism (3.2%): negative D-dimer can rule it out

[0154] Ultimately, through the deep integration of the semantic understanding capabilities of the large model and the medical knowledge base, intelligent mapping from multimodal data to clinical decisions was achieved, ensuring high accuracy while meeting the timeliness requirements of emergency scenarios.

[0155] Example 2:

[0156] A chest pain classification system based on multimodal data fusion and deep learning model, including:

[0157] A preprocessing module is used to preprocess and enhance chest X-ray images and electrocardiogram multimodal data;

[0158] Extraction module, used for anatomical-temporal dual-stream feature extraction;

[0159] Fusion module, used to guide the dynamic fusion of MOE based on clinical knowledge;

[0160] Reasoning module, used for large model-driven diagnostic reasoning.

Claims

1. A chest pain classification method based on multimodal data fusion and deep learning model, characterized in that: Here are the steps: (1) Preprocessing and enhancing chest X-ray images and electrocardiogram multimodal data; (2) anatomical-temporal dual-stream feature extraction; (3) Clinical knowledge guides the dynamic integration of MOE; (4) Diagnostic reasoning driven by large models.

2. The chest pain classification method based on multimodal data fusion and deep learning model according to claim 1, characterized in that: In step (1), the specific steps are: Chest X-ray images are preprocessed, and key anatomical regions, including the heart and lung fields, are segmented using a pretrained U-Net network to generate binary masks. Adaptive window width and window position adjustment is used to enhance the contrast of mediastinal structures, and pathology-preserving data enhancement is combined to improve model robustness. The electrocardiogram (ECG) was preprocessed by aligning the 12-lead signals based on dynamic time warping and using wavelet threshold denoising to eliminate baseline drift. The ST segment / T wave key waveforms were detected using the LSTM-CRF model to generate a time series annotation matrix, and elastic time warping enhancement technology was applied to simulate heart rate variability. After preprocessing, the data was stored in tensor form.

3. The chest pain classification method based on multimodal data fusion and deep learning model according to claim 2, characterized in that: In step (2), the specific process is: (21) A multi-scale anatomical perception pyramid network is constructed based on the anatomical features of chest X-ray images. Global and local features are extracted from chest X-ray images, and cross-block attention fusion is performed. The attention of the pathological area of ​​the chest X-ray image is then enhanced, and the medical semantic position encoding of the chest X-ray image is performed. (22) A time series feature extraction method is constructed for electrocardiogram, and a multi-scale time series convolutional pyramid network is constructed. Feature extraction is performed from three aspects, namely short-time feature extraction, medium-time feature extraction, and long-time feature modeling. Feature fusion is then performed, and a total loss function is constructed.

4. The chest pain classification method based on multimodal data fusion and deep learning model according to claim 3, characterized in that: In step (21), the specific steps are: A multi-scale anatomical-aware pyramid network was constructed based on the anatomical features of chest X-ray images. Global features were extracted from chest X-ray images, which were segmented into 16×16 pixel blocks and fed into a standard Vision Transformer (ViT) module. A multi-head self-attention mechanism was used to capture the overall structure of the chest cavity. Local features were then extracted from the chest X-rays using dilated convolutions to process 8×8 pixel blocks. Cross-block attention fusion: Global features are concatenated with local features, and features of different scales are fused through the cross-block attention module. The formula is: in, is the query matrix of global features, is the key and value matrix of local features, M ROI ∈{0, 1} N×M is the anatomical mask matrix, which only allows feature interactions in the heart / lung field region, d k = d / h is the dimension of each attention head, which captures both global anatomical structure and local pathological details through multi-scale fusion; The attention of the pathological area of ​​the chest X-ray image is enhanced, the anatomical mask is generated, and the pre-trained U-Net segmentation network outputs the binary mask M of the heart and hilum area. patho ∈{0, 1} H×W , attention bias injection, in the self-attention calculation, the anatomical mask is used as a bias term to strengthen the model's attention to the pathological area: in, is a block-level mask matrix, which is set to 1 if the block contains a pathological area, otherwise it is set to 0; λ is a learnable parameter; Perform medical semantic position coding on chest X-ray images, predefine 6 anatomical landmarks as anatomical key points, perform weighted coding on the inverse distance, calculate the distance between the center coordinates p = (x, y) of each image block and the anatomical key point, and generate semantic position coding: Among them, c k ∈[0, 1] 2 is the normalized coordinate of the kth key point, is a learnable weight matrix that maps distance to feature space, ||pc k ||2 is the Euclidean distance, which measures the proximity of the image patch to the anatomical landmark.

5. The chest pain classification method based on multimodal data fusion and deep learning model according to claim 4, characterized in that: In step (22), the specific steps are: For electrocardiogram (ECG), we construct a time series feature extraction and a multi-scale time series convolutional pyramid network. We extract features from three aspects: short-time feature extraction, medium-time feature extraction, and long-time feature modeling. Short-time feature extraction uses a 5ms convolution kernel to capture QRS wave details; medium-time feature extraction uses a 15ms convolution kernel to analyze ST segment morphological changes; and long-time feature modeling uses a bidirectional LSTM network to capture heart rate variability features. Based on this, we perform feature fusion, concatenating multi-scale features and performing dimensionality reduction through 1×1 convolution. The formula is: F fused =Conv1D 1×1 (Concat(F short ,F mid ,F long )) (4) in, is a feature tensor of different time scales, and then a clinical waveform constraint loss function is constructed to align the ST segment features and constrain the ST segment features f predicted by the model. ST The actual value marked by the doctor The error is: in, is the ST segment feature of the i-th sample output by the model, In order to extract the real ST segment features through the waveform annotation matrix, the waveform classification consistency constraints are imposed on the features, and the KL divergence is used to ensure that the waveform probability distribution p of the model output is wave Distribution with annotations Consistent: in, is the QRS / ST / T wave probability predicted by the model, is the marked real waveform distribution; Construct a total loss function to jointly optimize feature alignment and distribution consistency: Among them, α=0.1 is the weight to balance the two losses. is the loss value of the ST segment, is the loss between the image and the annotation.

6. The chest pain classification method based on multimodal data fusion and deep learning model according to claim 5, characterized in that: In step (3), the specific process is: (31) Construct four specialized expert models to handle specific modalities or tasks respectively; (32) Construction of dynamic gating network; (33) Perform multimodal feature fusion and splice features of multiple expert models.

7. The chest pain classification method based on multimodal data fusion and deep learning model according to claim 6, characterized in that: Step (31) is specifically as follows: construct an anatomical expert model based on the 3D-ResNet50 network structure, input multi-scale features of the chest X-ray, and output an anatomical abnormality score; construct an electrophysiological expert model based on Transformer-ECG, input ECG timing features, output electrophysiological parameters such as ST segment offset and QRS complex width, and force activation when the ST segment offset of any ECG lead is ≥1mm; construct a multimodal association expert model based on the graph attention network, model the cross-modal association between the chest X-ray area and the ECG waveform, and activate when the confidence difference between the chest X-ray and ECG features is >20%; construct a critical value expert model based on a lightweight CNN, input vital signs and pathological markers, and force activation when the heart rate is >120 beats / min or the systolic blood pressure is <90mmHg.

8. The chest pain classification method based on multimodal data fusion and deep learning model according to claim 7, characterized in that: Step (32) is specifically: The gating network dynamically adjusts the expert weights according to real-time clinical data. It is divided into two parts: data-driven and rule-driven. The first is data-driven weight calculation, which inputs the multimodal feature h in and vital signs Build the gating function: Among them, W g To dynamically calculate the weights, σ ​​is a variable parameter, and then the weights are modified by rules, imposing hard constraints when specific clinical events are detected: Among them, the gating function is different under different constraint conditions. When ST elevation is greater than or equal to 2 mm and k = 2, the gating constraint is 0.7; when heart rate is greater than 120 and k = 4, the gating constraint is 0.5; in other cases, it is Then the two weights are synthesized by the synthesis method: Among them, σ is the Sigmoid function, which limits the weight to the interval [0,1], and α is a learnable parameter that balances the contribution of data and rules. is the gating network weight matrix; Preferably, step (33) is specifically: Perform multimodal feature fusion, splice features of multiple expert models, and perform splicing based on residual connections: h final =h in +W r ·h expert (12) Among them, E k (·) is the feature extraction function of expert k, is the residual projection matrix, is the final fused feature vector with a dimension of 1024.

9. The chest pain classification method based on multimodal data fusion and deep learning model according to claim 8, characterized in that: In step (4), the specific process is: Construct an input feature concatenation method and convert the 1024-dimensional h output from step (3) into final Feature vector, 768-dimensional vector t encoding the patient's electronic medical record text text To splice: h joint =W c ·Concat(h final ,t text ) (13) in, is the dimensionality reduction matrix, Concat is the concatenation method, and the output is Diagnostic decision generation uses causal attention decoding to generate diagnostic content. The LLaMA-Med Transformer decoder is used to generate the diagnostic logic chain. Attention calculation introduces medical knowledge constraints: Among them, M rule ∈{0, 1} L×L is the clinical rule mask matrix, L is the sequence length, d k =2048 / 32=64 is the dimension of each attention head; Structured output generation, the model output consists of three parallel headers: disease probability header, treatment recommendation header, and differential diagnosis header. The disease probability header is: p disease =Softmax(W p ·h in ) (15) in, is the weight of the four categories; The disposal suggestion header is: s action =Sigmoid(W a ·h in ) (16) in, There are 12 standard disposals; The differential diagnosis heads are: Among them, ε DDx An embedding vector library containing 30 chest pain-related diseases; Generate a natural language decision chain based on the above steps: Generate report templates that are in line with clinical thinking.

10. A chest pain classification system based on multimodal data fusion and deep learning model, characterized in that: include: A preprocessing module is used to preprocess and enhance chest X-ray images and electrocardiogram multimodal data; Extraction module, used for anatomical-temporal dual-stream feature extraction; Fusion module, used to guide the dynamic fusion of MOE based on clinical knowledge; Reasoning module, used for large model-driven diagnostic reasoning.

Citation Information

Patent Citations

  • Heart hypertrophy multi-label detection system based on multi-modal deep learning

    CN115281688A

  • Heart failure diagnosis auxiliary method based on multi-modal data fusion

    CN116451068A

  • Medical decision-oriented multi-modal data dynamic fusion and labeling method and system

    CN119377894A

  • Device for machine learning-supported medical segmentation and analysis

    DE202024106059U1

  • Arrhythmia classification using correlation image

    US20200375490A1

Cited By

  • Automatic lung lesion detection and diagnosis method based on CT image

    CN120953287A

  • Lung ventilation-perfusion development area image segmentation and quantitative analysis system

    CN120953301A

  • Intelligent assessment method and system for coronary artery stenosis based on chest radiography

    CN121306446A