AI (aortic dissection) auxiliary diagnosis method based on multi-modal information fusion

By using an AI diagnostic method that integrates multimodal information fusion and combines CTA images and clinical indicator data, a dual-branch feature extraction network is established. This solves the problems of low information integration efficiency and poor diagnostic consistency in the diagnosis of aortic dissection, and enables rapid and accurate disease type identification.

CN121724932APending Publication Date: 2026-03-24FIRST AFFILIATED HOSPITAL OF DALIAN MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies for diagnosing aortic dissection suffer from problems such as low information integration efficiency, poor diagnostic consistency, and high rate of missed diagnoses. In particular, they are insufficient in multi-source data integration and scenario adaptability, making it difficult to meet the needs of rapid decision-making in emergency situations.

Method used

An AI-assisted diagnostic method based on multimodal information fusion is adopted. By combining CTA images and clinical indicator data through a dual-branch feature extraction network, an image feature extraction branch, a clinical indicator feature extraction branch, and a cross-modal fusion module are established to obtain multi-scale fusion features and global clinical feature sequences. The network is trained using a loss function to assist clinicians in judging the disease type.

Benefits of technology

It significantly improves the predictive accuracy and speed of aortic dissection diagnosis, enabling rapid and accurate disease type probabilities, enhancing diagnostic consistency and accuracy, and meeting the rapid decision-making needs of emergency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121724932A_ABST
    Figure CN121724932A_ABST
Patent Text Reader

Abstract

The invention discloses an aortic dissection AI auxiliary diagnosis method based on multi-modal information fusion, and the method comprises the steps: building a dual-branch feature extraction network, obtaining a multi-scale fusion feature finally outputted by an image branch according to an enhanced image and an image feature extraction branch, and carrying out the multi-scale fusion feature extraction. According to the clinical index set aligned to the image acquisition time and the clinical index feature extraction branches, obtaining a fused global clinical feature sequence; obtaining a final multi-modal feature after fusion according to a multi-scale fusion feature finally output by the image branch, the fused global clinical feature sequence and a cross-modal fusion module so as to obtain a probability of a predicted disease type based on clinical data of a user; after the loss function is adopted to train the double-branch feature extraction network, the probability of the predicted disease type can be obtained based on the clinical data of the to-be-diagnosed user, and a clinician is assisted to judge the disease type of the to-be-diagnosed user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart medical technology, and in particular to an AI-assisted diagnostic method for aortic dissection based on multimodal information fusion. Background Technology

[0002] Aortic dissection (AD) is a rapidly progressing and fatal vascular disease. Early diagnosis and accurate classification directly determine treatment options and patient prognosis. Currently, the mainstream clinical diagnostic process still relies on physicians manually integrating multi-source medical information: physicians must review independent reports such as CT angiography (CTA), echocardiography, laboratory tests (e.g., D-dimer), and electrocardiograms (ECG), combining these with the Stanford classification criteria (Type A involves the ascending aorta, Type B is limited to the descending aorta) for subjective judgment. To standardize this process, the International Aortic Dissection Registry (IRAD) has proposed a clinical pathway that includes history taking, image review, and laboratory verification. The 2022 American Heart Association / American College of Cardiology (AHA / ACC) guidelines for chest pain also emphasize the necessity of cross-validation of multi-dimensional data. However, this traditional model relies heavily on physician experience and has gradually revealed significant limitations in real clinical scenarios: First, information integration efficiency is low. Manual comparison and logical correlation of multi-source data require reviewing and repeatedly verifying each document, often taking more than 25 minutes per patient, which is insufficient to meet the need for rapid decision-making during the "golden period" of emergency care. Second, diagnostic consistency is poor. Different physicians have low accuracy in judging the Stanford classification due to differences in experience, especially in complex cases (such as those with organ ischemia or anatomical variations). Third, the rate of missed diagnoses is high. Taking abdominal pain type AD as an example, its symptoms lack specificity and are easily confused with gastrointestinal diseases. During manual interpretation, the risk of misdiagnosis increases significantly due to scattered focus or omission of key features.

[0003] To overcome the efficiency bottleneck of manual diagnosis, single-modal medical imaging AI systems have been gradually applied to the field of AD-assisted diagnosis in recent years. This type of technology is based on a single image data source (such as CTA) and uses deep learning models to train and achieve automatic identification and segmentation of aortic structures (such as intimal flaps, true and false lumens). A typical example is the CTA aortic dissection AI analysis system developed by United Imaging Intelligence. Compared to traditional models, it has improved in speed and accuracy of image structure recognition, but it still has fundamental defects due to its "single-modal" design logic: First, the problem of modal isolation is prominent. It only focuses on image data and ignores the dynamic changes of D-dimer in the blood. D-dimer is a highly sensitive biomarker for acute AD (sensitivity > 95%) and has key value in excluding low-risk patients, but information on its abnormally elevated or normal levels is not included in the model analysis, resulting in incomplete diagnostic criteria. Second, it has weak scenario adaptability. Existing models are mostly trained based on standardized CTA data from specific centers. The sensitivity drops sharply for atypical patients (such as those with obesity, renal insufficiency leading to poor image quality, or anatomical variations combined with congenital heart disease), making it difficult to meet the needs of multi-center and heterogeneous clinical scenarios.

[0004] In summary, existing technologies have two major contradictions: traditional manual diagnostic methods are inefficient in integrating multi-source information and are highly subjective, making it difficult to meet the speed and accuracy requirements of AD emergency scenarios; while single-modal AI systems improve image processing efficiency, they cannot provide full-dimensional diagnostic support covering "morphology-function-biochemistry-electrophysiology" due to modal fragmentation, lack of dynamic functions and insufficient scenario adaptability. Summary of the Invention

[0005] This invention discloses an AI-assisted diagnostic method for aortic dissection based on multimodal information fusion, in order to overcome the above-mentioned technical problems.

[0006] To achieve the above objectives, the technical solution of the present invention is as follows: A multimodal information fusion-based AI-assisted diagnostic method for aortic dissection includes the following steps: S1: Acquire the user's clinical data, including the patient's CTA 2D images and clinical indicator data; S2: Based on the CTA two-dimensional image, obtain a segmented image of the region of interest containing the aortic arch and related vascular structures to obtain an enhanced image; S3: Preprocess the clinical indicator data to obtain a set of clinical indicators aligned to the image acquisition time; wherein, the clinical indicator data includes continuous indicators, categorical indicators, and time-series signal processing indicators; S4: Establish a dual-branch feature extraction network, including an image feature extraction branch, a clinical indicator feature extraction branch, and a cross-modal fusion module; based on the enhanced image and the image feature extraction branch, obtain the multi-scale fusion features finally output by the image branch; based on the clinical indicator set aligned to the image acquisition time and the clinical indicator feature extraction branch, obtain the fused global clinical feature sequence; and based on the multi-scale fusion features finally output by the image branch, the fused global clinical feature sequence, and the cross-modal fusion module, obtain the fused final multi-modal features to obtain the probability of predicted disease type based on the user's clinical data; S5: A loss function is used to train the dual-branch feature extraction network. Based on the clinical data of the user to be diagnosed, the trained dual-branch feature extraction network can obtain the probability of the predicted disease type based on the clinical data of the user to be diagnosed, so as to assist clinicians in judging the disease type of the user to be diagnosed.

[0007] Furthermore, the method for obtaining the enhanced image is as follows: S21: Based on the CTA two-dimensional image, an automatic segmentation algorithm is used to obtain a segmented image of the region of interest containing the aortic arch and related vascular structures; S22: Based on the segmented image of the region of interest containing the aortic arch and related vascular structures, a size-normalized image is obtained using the following formula:

[0008] In the formula: This represents the image after size normalization; Indicates will Images were unified to a resolution of 512. The normalization function for 512; S23: Based on the image after size normalization, an image enhancement method is used to obtain an enhanced image.

[0009] Furthermore, S3 includes: S31: Preprocess the clinical indicator data, including: For continuous metrics, standardization is performed using the Z-score:

[0010] In the formula: This represents the standardized index value, where the standardized data follows a distribution with a mean of 0 and a standard deviation of 1. This indicates the value of the original continuous index; Indicators The mean of the entire sample or training set; Indicators Standard deviation; For categorical metrics, a learnable embedding is used to map them into a vector representation: e c =Embed( x c ) In the formula: e c This represents the vector representation of a categorical indicator; Embed( x c ) represents a learnable embedding function; x c This represents the original categorical indicator; For time-series signal processing metrics, obtain fixed-length sliding window segments:

[0011] In the formula: Indicates from time t The beginning of the continuous A time sequence segment composed of several time points; Indicates time Observed values; Indicates the length of the sliding window; Indicates time; S32: Based on the preprocessed clinical indicator data, a time alignment function is used to obtain a set of clinical indicators aligned to the image acquisition time.

[0012] Furthermore, the method for obtaining the multi-scale fusion features of the final output of the image branch is as follows: S401: Based on the enhanced image, obtain an image block of a first size and an image block of a second size, wherein the second size < the first size; S402: A linear embedding layer is used to obtain the feature vectors of the first-size image patch and the second-size image patch, respectively, based on the first-size image patch and the second-size image patch. The formula used is as follows:

[0013]

[0014] In the formula: Indicates the first Feature vectors of image blocks of the second size; Indicates the first Feature vectors of image blocks of the second size; , All of these represent trainable weights; Indicates the first A second-sized image block; Indicates the first A second-sized image block; , Both represent bias values ​​in the linear embedding layer; This represents the feature vector transformation function used to straighten spatial dimensions into a single line. S403: Based on the feature vectors of the first-size image patch and the second-size image patch, and using a cross-scale attention pyramid mechanism, attention-based features of the first-size image patch and attention-based features of the second-size image patch are obtained to acquire the multi-scale fusion features of the final output of the image branch. The formula used is as follows:

[0015] In the formula: This indicates the concatenation of feature dimensions; and All are learnable parameters; d is the output feature dimension. This represents the multi-scale fusion features of the final output of the image branch; The sequence length of the output features associated with the number of patches; These represent the first-size image patch features and the second-size image patch features based on attention, respectively.

[0016] Furthermore, the formula used to obtain the fused global clinical feature sequence is as follows:

[0017] In the formula: This represents the fused global clinical feature sequence; Indicates the multi-head self-attention layer; This represents a vector encoded as a continuous index aligned to the image acquisition time using MLP. A vector representation of the categorical index aligned to the image acquisition time; This represents a segment of time-series signal processing metrics output aligned to the image acquisition time of the time-series pyramid. The length of the fused global clinical feature sequence; Indicates dimension.

[0018] Furthermore, the method for obtaining the final multimodal features after fusion is as follows: S421: Retrieve the key features, value features, query features of CTA two-dimensional image plots and the key features, value features, and query features of clinical indicator data:

[0019]

[0020] In the formula: For CTA 2D image query features; For clinical indicator data query features; , , , , , All are linear transformation matrices; Key features for CTA 2D image maps; Key features for clinical indicator data; Value features for CTA 2D image maps; Value features for clinical indicator data; S422: Obtain guided attention results from image to clinical indicator direction and guided attention results from clinical indicator to image direction; obtain bidirectional interaction features, using the following formula:

[0021]

[0022]

[0023] In the formula: This serves as a guide for attention towards clinical indicators from imaging findings. The attention calculation function; This serves as a guide for attention from clinical indicators to imaging findings. It features two-way interaction; S423: Obtain the fused multimodal features using the following formula:

[0024] In the formula: The final multimodal features after fusion are used for classification; For learnable gating coefficients; It is a single-modal splicing feature; in,

[0025]

[0026] In the formula: For the Sigmoid function; For gating layer weights; For bias; S424: Based on the fused multimodal features, the predicted disease type probability is obtained using the following formula:

[0027]

[0028] In the formula: For classifying logits vectors; The final multimodal features after fusion; The normalized computation function for the layer; The bias of the normalized computation function for the layer; For the predicted first The probability of such diseases; This is the normalization function.

[0029] Furthermore, the loss function is expressed as follows:

[0030] In the formula: This is the balance coefficient; Weighted label smoothing cross-entropy loss; Total loss; The loss is the feature center loss; in,

[0031]

[0032] In the formula: k is the index of the feature center, which is also the index of the disease type; The weights for disease types, where, , The number of samples for the k-th disease category; Label smoothing is used to smooth the label. To predict the probability that the outcome is the kth type of disease, This is the label smoothing coefficient; This represents the function that determines whether the predicted label index is consistent with the feature center.

[0033] In the formula: The total number of samples in a single patch; This refers to the sample index within a single patch. For the first The fusion feature vector of each sample; For the first k Characteristic centers for each disease type.

[0034] Beneficial Effects: This invention provides an AI-assisted diagnostic method for aortic dissection based on multimodal information fusion. It establishes a dual-branch feature extraction network. Based on the enhanced image and imaging feature extraction branches, it obtains the multi-scale fusion features output by the imaging branch. Based on the clinical indicator set aligned to the image acquisition time and the clinical indicator feature extraction branch, it obtains the fused global clinical feature sequence. Furthermore, based on the multi-scale fusion features output by the imaging branch, the fused global clinical feature sequence, and the cross-modal fusion module, it obtains the final fused multimodal features to obtain the probability of predicted disease type based on the user's clinical data. After training the dual-branch feature extraction network using a loss function, it can obtain the probability of predicted disease type based on the clinical data of the user to be diagnosed, assisting clinicians in judging the disease type of the user. This invention significantly improves prediction accuracy by using the user's CTA two-dimensional image and clinical indicator data as input data and jointly analyzing the CTA two-dimensional image and clinical indicators. The dual-branch feature extraction network can quickly obtain prediction results, improving the accuracy of aortic dissection classification. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a flowchart of the AI-assisted diagnosis method for aortic dissection based on multimodal information fusion according to the present invention; Figure 2 This is a schematic diagram of the image feature extraction branch structure of the present invention; Figure 3 This is a schematic diagram of the clinical indicator feature extraction branch structure in an embodiment of the present invention; Figure 4 This is a schematic diagram of the cross-modal fusion module in an embodiment of the present invention. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] This embodiment introduces an AI-assisted diagnostic method for aortic dissection based on multimodal information fusion, including the following steps: Figure 1 As shown: S1: Acquire the patient's clinical data, including the user's CTA 2D images and clinical indicator data; Specifically, based on historical clinical diagnostic records, clinical data from multiple patients are acquired and processed as sample data. In this embodiment, the CTA two-dimensional image is as follows: ,in, Represents a two-dimensional CTA image; Represent real numbers; Indicates the height of the CTA 2D image; This represents the width of the CTA 2D image; the task of this embodiment is to learn a mapping: f θ :( I,X → y ,in, y Aortic dissection classification label; f θ :( I,X This indicates that the type of disease is determined by combining the user's CTA two-dimensional image and the user's clinical indicator data set. S2: Based on the CTA two-dimensional image, obtain a segmented image of the region of interest containing the aortic arch and related vascular structures to obtain an enhanced image; Preferably, the method for obtaining the enhanced image is as follows: S21: Based on the CTA two-dimensional image, an automatic segmentation algorithm is used to obtain a segmented image of the region of interest containing the aortic arch and related vascular structures; Specifically, from the original CTA 2D image, an automatic segmentation algorithm is used to locate and extract the Region of Interest (ROI) containing the aortic arch and related vascular structures to reduce redundant information and focus attention on the lesion site. The formula used is as follows: (1) In the formula: A segmented image representing a region of interest containing the aortic arch and related vascular structures; Represents the existing segmentation model; Represents a two-dimensional CTA image; S22: Based on the segmented image of the region of interest containing the aortic arch and related vascular structures, obtain a size-normalized image; Specifically, the ROI is uniformly scaled to a fixed resolution (512). (512 pixels), and the contrast of vascular tissue can be enhanced by adjusting the window width / window level.

[0039] (2) In the formula: This represents the image after size normalization; Indicates will Images were unified to a resolution of 512. The normalization function for 512; S23: Based on the image after size normalization, use an image enhancement method to obtain an enhanced image; Specifically, during the training phase, rotations (±10) are applied randomly. Operations such as translation (±5 pixels), affine transformation, brightness and contrast adjustment are performed to improve the robustness and generalization ability of the model.

[0040] (3) In the formula: This represents the enhanced image; Indicates image enhancement methods; S3: Preprocess the clinical indicator data to obtain a set of clinical indicators aligned to the image acquisition time; wherein, the clinical indicator data includes continuous indicators, categorical indicators, and time-series signal processing indicators; In this embodiment, the clinical indicator data includes, but is not limited to, blood biochemical indicators (D-dimer, C-reactive protein, etc.), electrocardiogram characteristics, blood pressure pulse pressure, and heart rate variability. Among them, blood biochemical indicators are continuous indicators, while ECG waveforms, dynamic blood pressure change curves, electrocardiogram characteristics, blood pressure pulse pressure, and heart rate variability are time-series signal processing indicators.

[0041] Preferably, S3 includes: S31: Preprocess the clinical indicator data, including: For continuous metrics, standardization is performed using the Z-score: (4) In the formula: This represents the standardized index value, where the standardized data follows a distribution with a mean of 0 and a standard deviation of 1, so that characteristics of different dimensions can be compared; This indicates the value of a raw, continuous indicator, such as the "C-reactive protein concentration" or "blood pressure value" of a particular measurement. Indicators The mean of the entire sample or training set is used to measure the central tendency of the indicator. Indicators The standard deviation is used to reflect the fluctuation range of the indicator; For categorical metrics, they are mapped to vector representations using learnable embeddings: e c =Embed( x c (5) In the formula: e c This represents the vector representation of the categorical indicator, used for subsequent fusion with continuous indicators or image features; Embed( x c ) represents a learnable embedding function used to calculate discrete categorical indicators, such as "gender", "blood type", "patent history type", etc. x c This represents the original categorical indicator; For time-series signal processing metrics: For time-series metrics (such as ECG waveforms, ambulatory blood pressure change curves), a fixed-length sliding window is used to segment the data, and missing values ​​are interpolated and filled in. (6) In the formula: Represents the continuous sequence starting from time t. A time series segment consisting of 10 time points is used as model input to capture time dependencies; Represents the observed value at time t (such as blood pressure, heart rate, or ECG sampling point at a certain time); This indicates the sliding window length, which is the number of time steps contained in each time segment; Indicates time; S32: Based on the preprocessed clinical indicator data, obtain a set of clinical indicators aligned to the image acquisition time; taking time-series signal processing indicators as an example, the formula used is as follows: (7) In the formula: This represents a set of clinical indicators aligned to the image acquisition time, ensuring consistency with image features in the time dimension, thereby enabling multimodal synchronous analysis. This represents a time alignment function used to align the timeline of clinical data with the time points of images. Common strategies include nearest time matching, interpolation, or window averaging. Indicates the acquisition time point of the corresponding image (CT, MRI, X-ray, etc.); Specifically, by aligning all clinical indicators (including continuous, categorical, and time-series indicators) to the image acquisition time point on the time axis, the temporal consistency of multimodal data is ensured.

[0042] S4: Establish a dual-branch feature extraction network, including an image feature extraction branch, a clinical indicator feature extraction branch, and a cross-modal fusion module; based on the enhanced image and the image feature extraction branch, input the enhanced image into the image feature extraction branch to obtain the multi-scale fusion features finally output by the image branch; based on the clinical indicator set aligned to the image acquisition time and the clinical indicator feature extraction branch, input the clinical indicator set aligned to the image acquisition time into the clinical indicator feature extraction branch to obtain the fused global clinical feature sequence; and based on the multi-scale fusion features finally output by the image branch, the fused global clinical feature sequence, and the cross-modal fusion module, input the multi-scale fusion features finally output by the image branch and the fused global clinical feature sequence into the cross-modal fusion module to obtain the final fused multi-modal features; to obtain the predicted disease type probability; Preferably, such as Figure 2 As shown, the method for obtaining the multi-scale fusion features of the final output of the image branch is as follows: S401: Based on the enhanced image, obtain an image block of a first size and an image block of a second size. Wherein, the second dimension is less than the first dimension; Specifically, such as Figure 1 As shown, the ROI region in the enhanced CTA 2D image is simultaneously divided into a large patch (image block of the first size) and a small patch (image block of the second size). The large patch is used to capture the overall vascular morphology, while the small patch is used to capture local lesion details. In this embodiment, the sizes of the large patch and the small patch are 32 × 32 and 8 × 8, respectively. The sets of large patches and small patches are represented as follows: , (8) In the formula: This represents the set of image patches of the first size, i.e., the large-scale patch set; This indicates the first large-size patch (e.g., 32×32), which mainly captures macroscopic features such as overall vascular structure, morphology, and branching patterns. This represents the NL-th large-scale patch; This represents the total number of large-scale patches; Represents a small-scale patch set; This indicates the first small-scale patch; This represents the Nth large-scale patch; This represents the total number of small-scale patches; S402: A linear embedding layer is used to obtain the feature vectors of the first-size image block and the second-size image block, respectively, based on the first-size image block and the second-size image block. Specifically, each scale patch is mapped into a sequence of feature vectors through an independent linear embedding layer. The formulas used to obtain the feature vectors of the first-size image patch and the second-size image patch are expressed as follows: (9) (10) In the formula: Indicates the first Feature vectors of image blocks of the second size; Indicates the first Feature vectors of image blocks of the second size; , All of these represent trainable weights; Indicates the first A second-sized image block; Indicates the first A second-sized image block; , Both represent bias values ​​in the linear embedding layer; This represents the feature vector transformation function used to straighten spatial dimensions into a single line. S403: Based on the feature vectors of the first-size image patch and the second-size image patch, and using the cross-scale attention pyramid mechanism, obtain the attention-based features of the first-size image patch and the attention-based features of the second-size image patch to obtain the multi-scale fusion features of the final output of the image branch; Specifically, for , Dynamic positional encoding is used to determine the position of feature vectors. Each scale's patch sequence is input into an independent ViT encoder (containing multi-head self-attention + feedforward network) to extract global features within the scale. A cross-scale attention pyramid (CSAP) mechanism is introduced in each Transformer layer, enabling large-scale features to guide the attention distribution of small-scale features, and vice versa.

[0043] (11) (12) In the formula: This represents attention from a large-scale query to a small-scale key / value, used to allow global structural information to guide the aggregation of local features; This represents attention from small-scale queries to large-scale key / value pairs, enabling reverse feedback from details to macro-level features; Represents a large-scale query matrix; Represents a large-scale key matrix; Represents a large-scale Value matrix; Represents a small-scale query matrix; Represents a small-scale key matrix; Represents the small-scale Value matrix; Indicates transpose; The feature dimension 'd' is used for scaling to prevent excessive attention weights. After cross-scale attention interaction, large-scale features will be... (Attention-based first-size image patch features) and small-scale features (Attention-based second-size image patch features) are concatenated along the feature dimension and mapped to a unified dimension via linear projection to obtain the final output features of the image branch: (13) In the formula: This indicates the concatenation of feature dimensions; and All are learnable parameters; d is the output feature dimension. This represents the multi-scale fusion features output by the imaging branch, which are used for subsequent modality fusion with clinical indicator features. The sequence length of the output features associated with the number of patches; These represent the first-size image patch features and the second-size image patch features based on attention, respectively. Preferred, such as Figure 3 As shown, the method used to obtain the fused global clinical feature sequence is as follows:

[0044] In the formula: X clin This represents a set of clinical indicator data for patients after pretreatment. Specifically, dynamic temporal signals are used to extract multi-scale temporal features through the existing multi-layer 1D-CNN+Dilated Convolution Temporal Pyramid Network (TPN).

[0045] (14) In the formula: Indicates the first l Temporal characteristics representation of layers; Indicates the first lDilated convolution achieves a greater sense of depth through dilated convolution. Indicates the index of the dilated convolutional layer; Indicates the initial time segment; Specifically, static and categorical features are encoded using MLP / Embedding, concatenated with the temporal pyramid output, and then the dependencies between different clinical feature types are captured through a multi-head self-attention layer (MHSA) to generate a global clinical feature vector. F clin .

[0046] (15) In the formula: This represents the fused global clinical feature sequence; This represents a multi-head self-attention layer, used to model dependencies between different clinical feature types; This represents a vector encoded as a continuous index aligned to the image acquisition time using MLP. A vector representation of the categorical index aligned to the image acquisition time; This represents a segment of time-series signal processing metrics output aligned to the image acquisition time of the time-series pyramid. The length of the fused global clinical feature sequence; Indicates dimension; Preferably, the method for obtaining the final multimodal features after fusion is as follows: S421: Obtain the key features, value features, query features of CTA two-dimensional image maps and the key features, value features, and query features of clinical indicator data; like Figure 4 As shown, let the image branch output feature be... Clinical branch output characteristics are , The length of the multi-scale fused features in the final output of the image branch is used to obtain their respective Query, Key, and Value through linear mapping:

[0047]

[0048] In the formula: For CTA 2D image query features; For clinical indicator data query features; , , , , , All are linear transformation matrices; Key features for CTA 2D image maps; Key features for clinical indicator data; Value features for CTA 2D image maps; Value features for clinical indicator data; S422: Obtain guided attention results from image to clinical indicator direction and guided attention results from clinical indicator to image direction; obtain bidirectional interaction features, using the following formula: Specifically, using image features as the query and clinical features as the key / value pair, we calculate the image → clinical guided attention features:

[0049] In the formula: This serves as a guide for attention towards clinical indicators; The attention calculation function; Calculate clinical-to-image guided attention features using clinical features as the query and imaging features as the key / value pair:

[0050] Concatenate the bidirectional guiding features:

[0051] In the formula: S423: Obtain the fused multimodal features using the following formula: (This refers to bidirectional interaction features; S423 is the formula for obtaining multimodal features after fusion.)

[0052] In the formula: The final multimodal features after fusion are used for classification; The learnable gating coefficients are constrained to [0, 1] by the Sigmoid function; It is a single-modal splicing feature; in,

[0053]

[0054] In the formula: For the Sigmoid function; For gating layer weights; For bias; This embodiment will These will serve as input features for subsequent classification heads and auxiliary tasks.

[0055] S424: Obtain the predicted disease type probability based on the fused multimodal features; Specifically, in this embodiment, the classification is a four-class classification task, and the formula used to obtain the predicted disease type probability is as follows:

[0056]

[0057] In the formula: For classifying logits vectors; The final multimodal features after fusion; The normalized computation function for the layer; The bias of the normalized computation function for the layer; This is the predicted probability of the k-th type of disease; This is the normalization function; Specifically, the probability of the kth type of disease correspond{ A,B, Atypical , The predicted probability of the comparison.

[0058] S5: A loss function is used to train the dual-branch feature extraction network. Based on the clinical data of the person to be diagnosed, the trained dual-branch feature extraction network can obtain the probability distribution of the predicted disease type of the person to be diagnosed. This helps clinicians to judge the disease type of the person to be diagnosed by combining other actual conditions of the patient.

[0059] Preferably, the loss function is expressed as follows:

[0060] In the formula: This is the balance coefficient; Weighted label smoothing cross-entropy loss; Total loss; The loss is the feature center loss; in,

[0061]

[0062] In the formula: k is the index of the feature center, which is also the index of the disease type, and its number is consistent with the number of disease types; The weights for disease types, where, , The number of samples for the k-th disease category; Labels smoothed using labelsmoothing; To predict the probability of the k-th type of disease, This is the label smoothing coefficient; This represents the function that determines whether the predicted label index is consistent with the feature center; it is 1 when they are consistent.

[0063]

[0064]

[0065] In the formula: This represents the total number of samples in the Patch. This refers to the sample index within a single patch. For the first The fusion feature vector of each sample; For the first k Characteristic centers for each disease type; In this embodiment, the disease type k ∈{1 , 2 , 3 , 4}, corresponding to {A, B, atypical, control} respectively. Specifically, in aortic dissection, the terms "Type A", "Type B", "atypical", and "control" represent different disease classifications and groupings in clinical studies: Type A aortic dissection: the most urgent and dangerous type, involving the ascending aorta. Regardless of the specific location of the tear (intima tear), as long as the dissection affects the ascending aorta, it belongs to Type A; Type B aortic dissection: relatively less urgent and dangerous than Type A, but still a critical disease. The dissection only involves the descending aorta (i.e., starting distal to the opening of the left subclavian artery) and does not involve the ascending aorta. This includes uncomplicated Type B and complicated Type B. Uncomplicated Type B: the first choice is drug treatment (strict control of blood pressure and heart rate) and close observation; Complicated Type B: when complications such as rupture, poor perfusion, persistent pain, or rapid enlargement of the aneurysm occur, emergency TEVAR surgery or open surgery is required. Atypical aortic dissection: This mainly refers to intramural hematoma and penetrating aortic ulcer. When the intramural hematoma or ulcer involves the ascending aorta, it is treated as type A; when it is limited to the descending aorta, it is treated as type B. "Control" is a term used in medical research. In clinical studies or statistical analyses of aortic dissection, a healthy control group is a group of people established for comparison with a patient group. It usually consists of healthy individuals or patients with other non-aortic diseases. The purpose is to more clearly identify the unique characteristics of aortic dissection patients (such as specific biomarkers, gene expression, risk factors, imaging differences, etc.) through comparison.

[0066] A specific application scenario of this embodiment is as follows: Case Name: Rapid AI-Assisted Diagnosis of Suspected Aortic Dissection Patients in the Emergency Department User information: 58-year-old male, sudden onset of tearing chest pain accompanied by radiating back pain, blood pressure fluctuating (170 / 95 mmHg → 150 / 80 mmHg). Input data: Aortic CTA: Dynamic sequences of the chest and abdomen (arterial / venous phases) Echocardiography: Dilated ascending aorta (42 mm in diameter) D-dimer: 5500 ug / L Clinical text: Chief complaint: "Sudden onset of chest pain for 2 hours," with a 12-year history of hypertension. Solution implementation process: Multimodal data fusion analysis: CTA: Segmenting the three-dimensional structure of the aorta, detecting the intima-lamina (accuracy 99.2%), and the location of the false lumen (indicated by the arrow above). Ultrasound: Quantification of blood flow velocity at the aortic root (peak velocity 4.2 m / s) D-dimer: 5500 ug / L. D-dimer has high sensitivity in diagnosing acute aortic dissection and is of great significance in ruling out other diagnoses.

[0067] Text semantic analysis: The NLP module identifies keywords such as "tearing pain" and "blood pressure fluctuations" and generates high-risk feature vectors. Decision fusion and diagnostic output: Fusion network output probability: 92.7% positive probability of aortic dissection (Stanford type A) Risk level: Level III (High risk, surgery required within 1 hour) Visualization report: Mark the intimal tear (22mm from the opening of the left subclavian artery) and the thrombotic area of ​​the false lumen.

[0068] Comparative experimental design: Control group: Traditional single-modal diagnostics: AI systems based solely on CTA Initial clinical consultation: Three emergency cardiac physicians independently analyzed the same data. Evaluation indicators: The average value of the single-modal AI doctor group in this embodiment is [indicator]. Sensitivity (%) 98.5 82.3 89.1 Specificity (%) 96.2 91.7 93.4 False negative rate (%) 1.5 17.7 10.9 Diagnosis time (minutes): 4.26.825+ Stanford classification accuracy: 95.876.483.2 Test set: 217 suspected Alzheimer's disease (AD) patients admitted to the cardiac emergency department and cardiovascular surgery department of a tertiary hospital in 2024 (148 confirmed AD cases and 69 non-AD cases). Typical comparison results Example 1: Users with atypical symptoms (avoiding false negatives) Patient data: Female, 36 years old, presenting only with abdominal pain and syncope (no typical chest pain). Single-modal AI (CTA only): No endothelial flap detected → Negative diagnosis (actually DeBakey type III) This embodiment's solution: Ultrasound detected abnormal abdominal aortic flow velocity (3.8 m / s). D-dimer: 7400ug / L Text analysis showed that the keyword "sudden fainting" triggered a high-risk warning. → Correct diagnosis of positive (probability 87.3%) Example 2: Calcification artifact interference (false positive avoidance) Patient data: Male, 65 years old, CTA showed extensive calcification of the aortic wall. Single-modal AI: Misclassifying calcifications as endometrial flaps → False positive diagnosis This embodiment's solution: Ultrasound confirmed that the vessel wall was continuous and without separation. D-dimer: 413ug / L → Corrected to negative (false positive probability reduced to 2.1%) This embodiment presents an AI-assisted diagnostic method for aortic dissection based on multimodal information fusion. A dual-branch feature extraction network is established. Based on the enhanced image and imaging feature extraction branches, the multi-scale fusion features output by the imaging branch are obtained. Based on the clinical indicator set aligned to the image acquisition time and the clinical indicator feature extraction branch, the fused global clinical feature sequence is obtained. Furthermore, based on the multi-scale fusion features output by the imaging branch, the fused global clinical feature sequence, and the cross-modal fusion module, the final fused multimodal features are obtained to acquire the probability of a predicted disease type based on the user's clinical data. After training the dual-branch feature extraction network using a loss function, the probability of a predicted disease type can be obtained based on the clinical data of the user to be diagnosed, assisting clinicians in determining the disease type of the user.

[0069] This embodiment uses user-generated two-dimensional CTA images and clinical indicator data as input data. It employs a pathophysiological-driven cross-modal interaction mechanism, jointly analyzing CTA images and clinical indicators, to map the aortic dissection development mechanism (intima tear → false lumen expansion → abnormal organ perfusion) into a multimodal analysis logic chain. A cross-modal causal inference engine is established: CTA structural segmentation + hematological D-dimer changes + ultrasound blood flow impact vector analysis, significantly improving prediction accuracy. Simultaneously, an asynchronous parallel processing pipeline is used, breaking the traditional serial processing mode. Modality-related parallel computing channels are designed: aortic CTA + hematological D-dimer + cardiac ultrasound + clinical text, achieving dynamic decision fusion. This enables extremely rapid prediction results, providing clinicians with highly reliable evidence for disease diagnosis and maximizing the golden window for emergency treatment. A dual-branch feature extraction network predicts the probability of disease type, and an adaptive patch resolution Vision Transformer (ViT) is introduced in the imaging branch. The encoder introduces a multi-layer temporal pyramid network into the clinical branch and achieves deep interaction of cross-modal features through a modality-guided residual gating fusion module, thereby improving the accuracy of aortic dissection classification.

[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for AI-assisted diagnosis of aortic dissection based on multimodal information fusion, characterized in that, Includes the following steps: S1: Acquire the user's clinical data, including the patient's CTA 2D images and clinical indicator data; S2: Based on the CTA two-dimensional image, obtain a segmented image of the region of interest containing the aortic arch and related vascular structures to obtain an enhanced image; S3: Preprocess the clinical indicator data to obtain a set of clinical indicators aligned to the image acquisition time; wherein, the clinical indicator data includes continuous indicators, categorical indicators, and time-series signal processing indicators; S4: Establish a dual-branch feature extraction network, including an image feature extraction branch, a clinical indicator feature extraction branch, and a cross-modal fusion module; based on the enhanced image and the image feature extraction branch, obtain the multi-scale fusion features finally output by the image branch; based on the clinical indicator set aligned to the image acquisition time and the clinical indicator feature extraction branch, obtain the fused global clinical feature sequence; and based on the multi-scale fusion features finally output by the image branch, the fused global clinical feature sequence, and the cross-modal fusion module, obtain the fused final multi-modal features to obtain the probability of predicted disease type based on the user's clinical data; S5: A loss function is used to train the dual-branch feature extraction network. Based on the clinical data of the user to be diagnosed, the trained dual-branch feature extraction network can obtain the probability of the predicted disease type based on the clinical data of the user to be diagnosed, so as to assist clinicians in judging the disease type of the user to be diagnosed.

2. The AI-assisted diagnostic method for aortic dissection based on multimodal information fusion according to claim 1, characterized in that, The method for obtaining the enhanced image is as follows: S21: Based on the CTA two-dimensional image, an automatic segmentation algorithm is used to obtain a segmented image of the region of interest containing the aortic arch and related vascular structures; S22: Based on the segmented image of the region of interest containing the aortic arch and related vascular structures, a size-normalized image is obtained using the following formula: In the formula: This represents the image after size normalization; Indicates will Images were unified to a resolution of 512. The normalization function for 512; S23: Based on the image after size normalization, an image enhancement method is used to obtain an enhanced image.

3. The AI-assisted diagnostic method for aortic dissection based on multimodal information fusion according to claim 1, characterized in that, S3 includes: S31: Preprocess the clinical indicator data, including: For continuous metrics, standardization is performed using the Z-score: In the formula: This represents the standardized index value, where the standardized data follows a distribution with a mean of 0 and a standard deviation of 1. This indicates the value of the original continuous index; Indicators The mean of the entire sample or training set; Indicators Standard deviation; For categorical metrics, a learnable embedding is used to map them into a vector representation: e c =Embed( x c ) In the formula: e c This represents the vector representation of a categorical indicator; Embed( x c ) represents a learnable embedding function; x c Represents the original categorical index; For time-series signal processing metrics, obtain fixed-length sliding window segments: In the formula: Indicates from time t The beginning of the continuous A time sequence segment composed of several time points; Indicates time Observed values; Indicates the length of the sliding window; Indicates time; S32: Based on the preprocessed clinical indicator data, a time alignment function is used to obtain a set of clinical indicators aligned to the image acquisition time.

4. The AI-assisted diagnostic method for aortic dissection based on multimodal information fusion according to claim 1, characterized in that, The method for obtaining the multi-scale fusion features of the final output of the image branch is as follows: S401: Based on the enhanced image, obtain an image block of a first size and an image block of a second size, wherein the second size < the first size; S402: A linear embedding layer is used to obtain the feature vectors of the first-size image patch and the second-size image patch, respectively, based on the first-size image patch and the second-size image patch. The formula used is as follows: In the formula: Indicates the first Feature vectors of image blocks of the second size; Indicates the first Feature vectors of image blocks of the second size; , All of these represent trainable weights; Indicates the first A second-sized image block; Indicates the first A second-sized image block; , Both represent bias values ​​in the linear embedding layer; This represents the feature vector transformation function used to straighten spatial dimensions into a single line. S403: Based on the feature vectors of the first-size image patch and the second-size image patch, and using a cross-scale attention pyramid mechanism, attention-based features of the first-size image patch and attention-based features of the second-size image patch are obtained to acquire the multi-scale fusion features of the final output of the image branch. The formula used is as follows: In the formula: This indicates the concatenation of feature dimensions; and All are learnable parameters; d is the output feature dimension. This represents the multi-scale fusion features of the final output of the image branch; The sequence length of the output features associated with the number of patches; These represent the first-size image patch features and the second-size image patch features based on attention, respectively. express 3D space.

5. The AI-assisted diagnostic method for aortic dissection based on multimodal information fusion according to claim 4, characterized in that, The formula used to obtain the fused global clinical feature sequence is as follows: In the formula: This represents the fused global clinical feature sequence; This indicates a multi-head self-attention layer; This represents a vector encoded as a continuous index aligned to the image acquisition time using MLP. A vector representation of the categorical index aligned to the image acquisition time; This represents a segment of time-series signal processing metrics output aligned to the image acquisition time of the time-series pyramid. The length of the fused global clinical feature sequence; Indicates dimension; for 3D space.

6. The AI-assisted diagnostic method for aortic dissection based on multimodal information fusion according to claim 5, characterized in that, The method for obtaining the final multimodal features after fusion is as follows: S421: Retrieve the key features, value features, query features of CTA two-dimensional image plots and the key features, value features, and query features of clinical indicator data: In the formula: For CTA 2D image query features; For clinical indicator data query features; , , , , , All are linear transformation matrices; Key features for CTA 2D image maps; Key features for clinical indicator data; Value features for CTA 2D image maps; Value features for clinical indicator data; S422: Obtain guided attention results from image to clinical indicator direction and guided attention results from clinical indicator to image direction; obtain bidirectional interaction features, using the following formula: In the formula: This serves as a guide for attention towards clinical indicators; The attention calculation function; This serves as a guide for attention from clinical indicators to imaging findings. It features two-way interaction; S423: Obtain the fused multimodal features using the following formula: In the formula: The final multimodal features after fusion are used for classification; For learnable gating coefficients; It is a single-modal splicing feature; in, In the formula: For the Sigmoid function; For gating layer weights; For bias; S424: Based on the fused multimodal features, the predicted disease type probability is obtained using the following formula: In the formula: For classifying logits vectors; The final multimodal features after fusion; The normalized computation function for the layer; The bias of the normalized computation function for the layer; For the predicted first The probability of such diseases; This is the normalization function; It is a 4-dimensional space; express The weight.

7. The AI-assisted diagnostic method for aortic dissection based on multimodal information fusion according to claim 1, characterized in that, The loss function is expressed as follows: In the formula: This is the balance coefficient; Weighted label smoothing cross-entropy loss; Total loss; The loss is the feature center loss; in, In the formula: k is the index of the feature center, which is also the index of the disease type; The weights for disease types, where, , The number of samples for the k-th disease category; Label smoothing is used to smooth the label. To predict the probability that the outcome is the kth type of disease, This is the label smoothing coefficient; This represents the function that determines whether the predicted label index is consistent with the feature center. In the formula: The total number of samples in a single patch; This refers to the sample index within a single patch. For the first The fusion feature vector of each sample; For the first k Characteristic centers for each disease type.