Cross-modal distillation coupled coronary plaque classification method
Through the cross-modal distillation coupling method, the high-resolution characteristics of IVUS are used to enhance the discriminant ability of CTA images, solving the problems of insufficient spatial resolution and artifact interference of CTA images, achieving high-accuracy plaque classification, and maintaining the non-invasive clinical advantages.
Patent Information
- Application Number
- CN202510228809.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-28
AI Technical Summary
In the prior art, CTA images have low spatial resolution and are susceptible to artifact interference, while high-precision IVUS requires invasive operation, which may bring discomfort and potential complication risks to patients.
The cross-modal distillation coupling method is adopted to separate the anatomical structural features and artifact noise of CTA images through the adaptive feature decoupling module, design a hierarchical knowledge distillation mechanism, use IVUS high-resolution characteristics to establish a multi-scale feature mapping relationship, and realize modal invariance feature extraction through dynamic contrast learning strategies.
It greatly improves the accuracy of plaque classification of CTA images, overcomes the problems of insufficient spatial resolution and artifact interference of CTA images, and avoids the risks brought about by invasive operations of IVUS, and maintains the non-invasive clinical advantages.
Smart Images

Figure CN120164020A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence, deep learning, and computer vision content, and in particular to a coronary plaque classification method coupled with cross-modal distillation. Background Art
[0002] Accurate classification of coronary artery plaques is a key link in the diagnosis and treatment of cardiovascular diseases. Currently, clinical diagnosis mainly relies on two imaging techniques of CT angiography (CTA) and intravascular ultrasound (IVUS). Traditional methods are mostly based on single-modal feature analysis. Among them, CTA has the advantage of non-invasiveness, but has inherent defects such as insufficient spatial resolution and susceptibility to calcification artifacts, resulting in an identification accuracy of only 68-75% for key features such as fibrous plaques and lipid cores. Although IVUS can provide high-resolution cross-sectional images at the 0.1 mm level, it requires invasive catheter operations, with a 3-5% risk of vascular injury and patient tolerance problems.
[0003] Currently, Zreik et al. proposed a deep learning-based coronary artery plaque classification method, which uses CTA images for automatic extraction and classification of plaque features. This method extracts features from CTA images through a convolutional neural network (CNN) and combines clinical data for plaque classification, achieving good classification results. However, it does not fully utilize the high-resolution characteristics of IVUS. Due to problems such as insufficient spatial resolution and interference from calcification artifacts in CTA images, the identification accuracy of key features such as fibrous plaques and lipid cores is limited. Especially in complex lesion areas, traditional methods are difficult to effectively separate anatomical structure features from artifact noise.
[0004] The present invention is a coronary plaque classification method coupled with cross-modal distillation, aiming to solve the problems that the spatial resolution of CTA is relatively low and it is easily interfered by artifacts, while high-precision IVUS requires introducing a probe into the blood vessel, which may cause discomfort and potential complication risks to patients. In order to integrate the advantages of these two technologies of CTA and IVUS, an artificial intelligence technology is used to model the mapping relationship between CTA and IVUS. Based on the automatic delineation of coronary blood vessels and plaques, the high-resolution characteristics of IVUS are used to overcome the limitations of CTA and endow CTA images with the discriminant ability of IVUS.
[0005] The present invention innovatively constructs a cross-modal distillation coupling framework: First, an adaptive feature decoupling module is used to separate the anatomical structure features and artifact noises of CTA images; then, a hierarchical knowledge distillation mechanism is designed to establish a multi-scale feature mapping relationship by utilizing the high-resolution characteristics of IVUS; finally, a dynamic contrast learning strategy is introduced to construct cross-modal positive and negative sample pairs in the latent space, and modality-invariant feature extraction is achieved through a learnable similarity metric function. This method breaks through the limitations of traditional single-modal analysis, greatly improves the plaque classification accuracy of CTA images, and at the same time maintains the non-invasive clinical advantages. Summary of the Invention
[0006] The object of the present invention: The spatial resolution of coronary CTA is low and it is easily interfered by artifacts. While high-precision IVUS requires introducing a probe into the blood vessel, which may cause discomfort and potential complication risks to patients. The present invention adopts a cross-modal distillation coupling method to directly learn the discrimination ability of IVUS from CTA images, and combines the clues of CTA itself to comprehensively evaluate and judge coronary plaques, improving the accuracy of medical diagnosis.
[0007] The technical solution of the present invention is as follows:
[0008] A cross-modal distillation coupling method for coronary plaque classification, the steps are as follows:
[0009] Step 1: Use a backbone network to extract the image features of CT angiography (CTA) and intravascular ultrasound (IVUS);
[0010] Given the CTA modality input and the IVUS modality input The backbone network extracts the feature maps of I CTA and I IVUS where N represents the number of CTA or IVUS pictures of the input sample, H and W respectively represent the height and width of the image, and C represents the channel dimension of the image. and
[0011] Step 2: Use a feature enhancement module to obtain a feature-enhanced representation vector, and use a cross-modal distillation module to achieve the transfer of IVUS modality knowledge to the CTA modality through feature alignment;
[0012] V CTA and V IVUS respectively obtain the feature-enhanced representation vectors V' CTA and V' IVUS after passing through the feature enhancement module Transformer1; V CTA passes through Transformer2, and learns the discriminative clues in V' IVUS through the cross-modal distillation module to obtain V'' CTA .
[0013] During the processing of Transformer2, a feature-based cross-modal loss function is used to achieve knowledge distillation from the IVUS modality to the CTA modality, minimizing the Euclidean distance between the enhanced feature vectors of these two modalities. The Euclidean distance expression is as follows:
[0014]
[0015] Step 3: Use a cross-modal coupling module to fuse multi-modal features;
[0016] The cross-modal coupling module includes a cross-modal self-attention module and a cross-modal cross-attention module. The cross-modal self-attention module includes a normalization layer Norm, a multi-head self-attention MHSA, a convolutional feed-forward network Conv-FFN, and a connector. The cross-modal cross-attention module includes a normalization layer Norm, a multi-head cross-attention MHCA, a coupling connection unit, a convolutional feed-forward network Conv-FFN, and a connector.
[0017] The cross-modal coupling module first performs reconstruction and self-reinforcement through two cross-modal self-attention modules, then enters the cross-modal cross-attention module for feature fusion, and finally outputs the features V cls and V' cls .
[0018] Furthermore, the specific process of the cross-modal self-attention module is expressed by the following formula: For the input X, the output Y can be obtained through the following formula:
[0019] X' = MHSA(Norm(X)) + X
[0020] Y = Conv-FFN(Norm(X')) + X'
[0021] where X is the input V' CTA 、V' IVUS and V'' CTA to the cross-modal self-attention module. Y is the output of the cross-modal self-attention module, which acts on the cross-modal cross-attention module. Let the outputs of the cross-modal self-attention module be respectively. The output of the cross-modal self-attention module is set to
[0022] Furthermore, the specific process of the cross-modal cross-attention module is as follows: The cross-modal cross-attention module includes two branches, and each branch has its own data input. Among them, is the data input of one branch, or Data input for another branch. After passing through the normalization layer, the results are respectively fed into the multi-head cross-attention layer of its own branch and the multi-head cross-attention layer of another branch. The results after passing through the multi-head cross-attention layer are concatenated with the input of this branch. The concatenated results of the two branches are input into the coupling connection unit for coupling connection. The results of the coupling connection are then successively passed through the normalization layer and the convolutional feed-forward network and concatenated with the results of the coupling connection; when the data input of another branch is the output of the cross-modal cross-attention module is V′ cls ; when the data input of another branch is the output of the cross-modal cross-attention module is V cls .
[0023] Furthermore, the specific process of the coupling connection unit is as follows: First, concatenate V′ CTA and V″ CTA along the channel, and then use the channel attention mechanism to fuse the two branches to obtain the fused feature vector V cls . The channel attention mechanism is implemented by a multi-layer perceptron containing two layers of 1×1 convolutions, and its calculation process is:
[0024] V cls = MLP(Cat(V′ CTA , V″ CTA ))·Cat(V′ CTA , V″ CTA )
[0025] where Cat(·) represents concatenation along the channel, and MLP(·) represents the multi-layer perceptron.
[0026] Similarly, V′ CTA and V′ IVUS are fused in the same way to obtain the feature vector V′ cls .
[0027] Step four: Establish spatial consistency constraints to model the spatial context relationship between slices
[0028] First, sample positive and negative sample pairs, and then use the contrast loss function to constrain the spatial semantic continuity of V cls and V′ cls so that the model prediction results satisfy the local continuity assumption. The expression of this loss function is as follows,
[0029]
[0030] where is the feature vector at the current position, is the positive sample with the same category as x randomly sampled from the x spatial neighborhood, Negative samples that are randomly sampled from the x - space neighborhood and are different from the x category.
[0031] After semantic continuity calculation, the L2 loss function is combined to assist in regularizing the feature vectors. The auxiliary loss is as follows.
[0032]
[0033] When y is a positive sample, b is 1, and when y is a negative sample, b is 0. The auxiliary loss considers additional spatial information to avoid overfitting of the model. The total spatial consistency loss is:
[0034]
[0035] Step Five: Use a constrained classification network for classification;
[0036] The classification features V cls and V′ cls processed in Step Four are passed through a fully - connected layer, and the cosine similarity is used to constrain the consistency of V cls and V′ cls respectively, and the cross - entropy losses and are used to regularize V cls and V′ cls to make the predicted category approach the true value category. The cosine similarity formula is:
[0037]
[0038] The beneficial effects of the present invention: Through the cross - modal distillation coupling framework, the present invention utilizes the high - resolution characteristics of IVUS to enhance the discriminative ability of CTA images, overcomes the inherent defects of insufficient spatial resolution and artifact interference of CTA images, and at the same time avoids the risk of vascular injury and patient discomfort caused by the invasive operation of IVUS. In addition, the present invention introduces a dynamic contrast learning strategy, and realizes the extraction of modality - invariant features by constructing cross - modal coupled sample pairs. The introduction of spatial semantic continuity constraints and contrast learning loss functions further enhances the model's ability to model local continuity and spatial context relationships, and improves the robustness of the classification results. The present invention not only greatly improves the plaque classification accuracy of CTA images, but also maintains the non - invasive clinical advantage, providing strong technical support for the precise diagnosis and treatment of cardiovascular diseases. Brief Description of the Drawings
[0039] Figure 1 is the overall flow chart.
[0040] Figure 2 is the schematic diagram of the cross - modal coupling unit.
[0041] Figure 3 Schematic diagram of the coupling connection unit Specific implementation manner
[0042] A coronary plaque classification method based on cross-modal distillation coupling, the steps are as follows:
[0043] Construct a network framework: The network framework of the present invention consists of a dual-branch structure, which processes CTA and IVUS modal inputs respectively. The backbone network uses a convolutional neural network (CNN) to extract initial features, followed by a feature enhancement module (Transformer1), a cross-modal distillation module, a cross-modal coupling module, and a constrained classification network. Among them, the cross-modal distillation module realizes the transfer of IVUS modal knowledge to the CTA modal through feature alignment, and the coupling module fuses multi-modal features through an attention mechanism. Finally, plaque classification is completed by combining spatial semantic continuity constraints and a classifier.
[0044] Step 1: Use the backbone network to extract the image features of CT angiography (CTA) and intravascular ultrasound (IVUS);
[0045] Given the CTA modal input and the IVUS modal input The backbone network extracts I CTA and I IVUS feature maps and where N represents the number of CTA or IVUS images of the input sample, H and W represent the height and width of the image respectively, and C represents the channel dimension of the image.
[0046] Step 2: Use the feature enhancement module to obtain feature-enhanced representation vectors, and use the cross-modal distillation module to realize the transfer of IVUS modal knowledge to the CTA modal through feature alignment
[0047] V CTA 、V IVUS After passing through the feature enhancement module Transformer1, the feature-enhanced representation vectors V' CTA and V' IVUS (The Transformer1 module is used to generate V' CTA , retaining the original information contained in the CTA modal); V CTA Passes through Transformer2 and learns the discriminative clues in V' IVUS to obtain V'' CTA .
[0048] During the processing of Transformer2, a feature-based cross-modal loss function is used to achieve knowledge distillation from the IVUS modality to the CTA modality, minimizing the Euclidean distance between the enhanced feature vectors of these two modalities. The expression of the Euclidean distance is as follows:
[0049]
[0050] Step 3: Adopt a cross-modal coupling module to fuse multi-modal features
[0051] Since the CTA and IVUS modalities are heterogeneous, in order to make the feature information between V′ CTA and V″ CTA compatible, the present invention designs a cross-modal coupling mechanism to achieve the alignment and fusion of the information of the two, and obtain classification features Through cross-modal coupling, the CTA modality information and the information distilled from the IVUS modality can be organically combined, avoiding modal information conflicts and contaminations caused by simple operations such as splicing.
[0052] The cross-modal coupling module includes a cross-modal self-attention module and a cross-modal cross-attention module. Both the cross-modal self-attention module and the cross-modal cross-attention module adopt a multi-head design, and use Conv-FFN to replace the fully connected layer of the attention module. At the same time, a coupling connection module is used to replace the concatenation operation in the cross-attention module to better couple the features. The cross-modal self-attention module includes a normalization layer Norm, a multi-head self-attention MHSA, a convolutional feed-forward network Conv-FFN, and a connector. The cross-modal cross-attention module includes a normalization layer Norm, a multi-head cross-attention MHCA, a coupling connection unit, a convolutional feed-forward network Conv-FFN, and a connector.
[0053] The cross-modal coupling module first performs reconstruction and self-reinforcement through two cross-modal self-attention modules, then enters the cross-modal cross-attention module for feature fusion, and finally outputs the features V cls and V′ cls .
[0054] Further, the specific process of the cross-modal self-attention module is expressed by the following formula: For the input X, the output Y can be obtained through the following formula:
[0055] X′ = MHSA(Norm(X)) + X
[0056] Y = Conv-FFN(Norm(X′)) + X′
[0057] where X is the input V′ CTA 、V′ IVUS and V″CTA , where Y is the output of the cross-modal self-attention module and acts on the cross-modal cross-attention module. Let the outputs of the cross-modal self-attention module be respectively. The output of the cross-modal self-attention module is set to
[0058] Although Transformer can model long-range correlations well, in this task, local correlations will reveal more valuable information. The fully connected part of Transformer only performs separate channel transformations on each element and lacks attention to the local environment. Therefore, in the present invention, the FFN is adjusted by convolution, and a 3×3 convolutional layer is used to replace the original two fully connected layers, and a dimension same as the original FFN is generated in the way of "3×3 convolutional layer - batch normalization - activation".
[0059] Furthermore, the specific process of the cross-modal cross-attention module is as follows: The cross-modal cross-attention module includes two branches, and each branch has its own data input. Among them, is the data input of one branch, or is the data input of the other branch. After passing through the normalization layer, the results are respectively sent to the multi-head cross-attention layer of its own branch and the multi-head cross-attention layer of the other branch. The results after passing through the multi-head cross-attention layer are connected with the input of this branch. The results after connecting the two branches are input into the coupling connection unit for coupling connection. The results of the coupling connection are then successively passed through the normalization layer and the convolutional feed-forward network and then connected with the results of the coupling connection; when the data input of the other branch is , the output of the cross-modal cross-attention module is V′ cls ; when the data input of the other branch is , the output of the cross-modal cross-attention module is V cls .
[0060] Furthermore, the specific process of the coupling connection unit is as follows: First, V′ CTA and V″ CTA are concatenated along the channel, and then the channel attention mechanism is used to fuse the two branches to obtain the fused feature vector V cls . The channel attention mechanism is implemented by a multi-layer perceptron containing two layers of 1×1 convolution, and its calculation process is:
[0061] V cls = MLP(Cat(V′ CTA , V″ CTA ))·Cat(V′ CTA , V″ CTA )
[0062] Where Cat(·) represents concatenation along the channel, and MLP(·) represents multi-layer perceptron.
[0063] Similarly, V′ CTA and V′ IVUS The feature vector V′ is obtained by fusion in the same way cls .
[0064] Step 4: Establish spatial consistency constraints to model the spatial contextual relationship between slices
[0065] Considering that plaque lesions occupy a certain space along the direction of blood vessels, the prediction results of adjacent slices are continuous. Using this prior knowledge helps to improve the robustness of classification. The present invention establishes spatial consistency constraints to model the spatial contextual relationship between slices. First, the semantic continuity of the generated feature vector is calculated. Specifically, positive and negative sample pairs are sampled first, and then V is constrained by the contrast loss function. cls and V′ cls The spatial semantic continuity of makes the model prediction results meet the local continuity assumption. The loss function expression is as follows:
[0066]
[0067] in is the feature vector of the current position, is a positive sample of the same category as x randomly sampled from the spatial neighborhood of x, is a negative sample of different category from x randomly sampled in the neighborhood of x space. This loss function can bring the positive sample closer in the feature space and push the negative sample away. After calculating the semantic continuity, the L2 loss function is combined to assist in regularizing the feature vector. The auxiliary loss is as follows:
[0068]
[0069] When y is a positive sample, b is 1, and when y is a negative sample, b is 0. The auxiliary loss takes into account additional spatial information to avoid overfitting of the model. The total spatial consistency loss for:
[0070]
[0071] Step 5: Use constrained classification network for classification
[0072] The classification feature V processed in step 4 cls , V′ cls After the fully connected layer, cosine similarity is used to constrain V cls and V′ cls The consistency of , and the cross entropy loss is used respectively and To regularize V cls and V' cls , so that the predicted category approaches the true category. The cosine similarity formula is as follows:
[0073]
[0074] Experimental content:
[0075] The present invention uses the currently publicly available CTA and IVUS data sets and the data provided by the cooperating hospitals. The CTA data set includes 200 patients, with 25 CTA slices for each patient, totaling approximately 5,000 images; the IVUS data set includes 150 patients, with 20 cross-sectional slice data for each patient, totaling 3,000 images. All of the above have been labeled by professional doctors.
[0076]
[0077] During the experiment, the initial learning rate was set to 1e-4, the batch size was 16, the Adam optimizer was used, and the weight decay was 1e-5. The evaluation metrics included classification accuracy (Accuracy), precision (Precision), recall (Recall), F1 score, and AUC (area under the curve). The experimental results showed that the overall accuracy of the model reached 89.5%, significantly higher than the 68%-75% of the traditional single-modal method; the AUC value was improved to 0.93-0.94, better than the traditional 0.8-0.9. At the same time, the model maintained the advantage of non-invasive diagnosis.
Claims
1. A coronary plaque classification method based on cross-modal distillation coupling, characterized in that: Here are the steps: Step 1: Use the backbone network to extract the image features of CT angiography CTA and intravascular ultrasound IVUS; Given a CTA modal input and IVUS modality input Backbone Network Extraction I CTA and I IVUS Feature map and Wherein, N represents the number of CTA or IVUS images of the input sample, H and W represent the height and width of the image respectively, and C represents the channel dimension of the image; Step 2: Use the feature enhancement module to obtain the feature enhanced representation vector, and use the cross-modal distillation module to achieve the migration of IVUS modality knowledge to CTA modality through feature alignment; V CTA 、V IVUS After the feature enhancement module Transformer1, the feature enhanced representation vector V′ is obtained CTA and V′ IVUS ; V CTA After Transformer2, V′ is learned through the cross-modal distillation module IVUS The discriminant clues in the CTA ; In the process of Transformer2 processing, a feature-based cross-modal loss function is used to realize the knowledge distillation from IVUS modality to CTA modality, minimizing the Euclidean distance between the enhanced feature vectors of these two modalities. The Euclidean distance expression is as follows: Step 3: Use the cross-modal coupling module to fuse multi-modal features; The cross-modal coupling module includes a cross-modal self-attention module and a cross-modal cross-attention module, the cross-modal self-attention module includes a normalization layer Norm, a multi-head self-attention MHSA, a convolutional feedforward network Conv-FFN and a connector, and the cross-modal cross-attention module includes a normalization layer Norm, a multi-head cross-attention MHCA, a coupling connection unit, a convolutional feedforward network Conv-FFN and a connector; The cross-modal coupling module is first reconstructed and self-reinforced through two cross-modal self-attention modules, then enters the cross-modal cross-attention module for feature fusion, and finally outputs the feature V for classification. cls and V′ cls ; Step 4: Establish spatial consistency constraints to model the spatial contextual relationship between slices; First sample positive and negative sample pairs, and then constrain V through the contrast loss function cls and V′ cls The spatial semantic continuity of the model makes the prediction results satisfy the local continuity assumption; the loss function expression is as follows: in is the feature vector of the current position, is a positive sample of the same category as x randomly sampled from the spatial neighborhood of x, Negative samples of different categories from x are randomly sampled from the spatial neighborhood of x; After the semantic continuity calculation, the L2 loss function is used to assist in regularizing the feature vector. The auxiliary loss is as follows: When y is a positive sample, b is 1, and when y is a negative sample, b is 0; the auxiliary loss takes into account additional spatial information to avoid overfitting of the model; the total spatial consistency loss for: Step 5: Use constrained classification network for classification; The classification feature V processed in step 4 cls , V′ cls After the fully connected layer, cosine similarity is used to constrain V cls and V′ cls The consistency of , and the cross entropy loss is used respectively and To regularize V cls and V′ cls , so that the predicted category is close to the true value category; the cosine similarity formula is:
2. The method for coronary plaque classification based on cross-modal distillation coupling according to claim 1, characterized in that: The specific process of the cross-modal self-attention module is expressed as follows: For input X, the output Y can be obtained through the following formula: X′=MHSA(Norm(X))+X Y = Conv-FFN(Norm(X′))+X′ Where X is the input V′ of the cross-modal self-attention module CTA , V′ IVUS and V″ CTA , Y is the output of the cross-modal self-attention module, acting on the cross-modal cross-attention module; the output of the cross-modal self-attention module is set to be 3. The coronary plaque classification method based on cross-modal distillation coupling according to claim 1, characterized in that: The specific process of the cross-modal attention module is as follows: the cross-modal attention module includes two branches, each of which has its own data input, wherein: Data input for a branch, or is the data input of another branch; after passing through the normalization layer, the results are respectively transmitted to the multi-head cross attention layer of its own branch and the multi-head cross attention layer of another branch, and the results after the multi-head cross attention layer are connected with the input of the branch. The results after the connection of the two branches are input to the coupling connection unit for coupling connection, and the results of the coupling connection are then connected with the results of the coupling connection after passing through the normalization layer and the convolutional feedforward network in turn; when the data input of the other branch is When , the output of the cross-modal attention module is V′ cls ; When the data input of another branch is When , the output of the cross-modal attention module is V cls .
4. The method for coronary plaque classification based on cross-modal distillation coupling according to claim 3, characterized in that: The specific process of the coupling connection unit is as follows: first, V' is spliced along the channel. CTA and V″ CTA , and then use the channel attention mechanism to fuse the two branches to obtain the fused feature vector V cls ; The channel attention mechanism is implemented by a multi-layer perceptron containing two layers of 1×1 convolution, and its calculation process is: In cls =MLP(Cat(V′ CTA "In" CTA ))·Cat(V′ CTA "In" CTA ) Where Cat(·) represents concatenation along the channel, and MLP(·) represents multi-layer perceptron; Similarly, V′ CTA and V′ IVUS The feature vector V′ is obtained by fusion in the same way cls .
Citation Information
Patent Citations
A method for calculating coronary blood flow reserve fraction based on porous medium theory
CN109106348A
System and method for image registration in medical imaging system
US20170301080A1
IVUS enabled contrast agent imaging with dual energy x-ray
WO2024068502A1
Pedestrian attribute cross-modal alignment method based on complete attribute identification enhancement
WO2024114185A1