Coronary artery disease early risk prediction and typing method based on deep learning
By combining the MedSAM encoder, multi-scale state-space feature encoder and hierarchical feature reconstruction module, and using the multimodal feature alignment optimization loss module, the problem of global features and local details in coronary CTA image segmentation is solved, and efficient and accurate early risk prediction and classification of coronary artery disease is achieved.
Patent Information
- Application Number
- CN202510814627.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies have difficulty in simultaneously capturing global features and local details in coronary artery CTA image segmentation, and have difficulty in effectively utilizing multi-scale information, resulting in insufficient segmentation accuracy and robustness.
The MedSAM encoder is combined with a multi-scale state-space feature encoder and a hierarchical feature reconstruction module, and the multimodal feature alignment is used to optimize the loss module to weaken the distribution difference between image features and text features. The multi-scale convolutional attention mechanism and the fusion of medical image and text features are used to improve the segmentation accuracy and robustness.
It achieves efficient and accurate segmentation and risk assessment of coronary artery CTA images, significantly improves segmentation accuracy and model robustness, reduces computational costs, and improves the accuracy of early risk prediction and classification of coronary artery disease.
Smart Images

Figure CN120766947A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a deep learning-based early risk prediction and typing method for coronary artery disease, belonging to the technical field of medical image processing. BACKGROUND
[0002] Coronary artery disease (CAD) is one of the main causes of cardiovascular disease deaths worldwide, and its early diagnosis and risk assessment are of great significance for improving patient outcomes. In recent years, with the rapid development of medical imaging technology, coronary CT angiography (CTA) has become an important non-invasive diagnostic tool, which can provide high-resolution cardiac vascular images and provide the possibility for early detection, risk assessment and typing of coronary artery disease.
[0003] However, although CTA technology has significant advantages in the diagnosis of coronary artery disease, how to accurately extract valuable information from a large amount of CTA images is still a problem to be solved. Traditional image analysis methods rely on manual annotation and feature extraction, which not only consumes time and effort, but also is easily affected by human factors, making it difficult to meet the clinical demand for rapid and accurate diagnosis. In recent years, deep learning technology has made breakthroughs in the field of medical image analysis, providing a new approach to automatic diagnosis and risk prediction of coronary artery disease.
[0004] In the field of deep learning, image segmentation is one of the key technologies for medical image analysis. Although traditional fully supervised pixel-level segmentation methods can provide high-precision segmentation results, they require a large amount of pixel-level annotation data, which is often difficult to obtain in the medical image field. In order to improve the segmentation efficiency and reduce the dependence on annotation data, researchers have begun to explore new methods and technologies. Among them, segmentation networks based on encoder-decoder architecture (such as U-Net and its variants) have made significant achievements in medical image segmentation, but these methods still have certain limitations in dealing with complex coronary CTA images, such as difficulty in capturing global features and local details simultaneously, and difficulty in effectively utilizing multi-scale information.
[0005] To overcome these limitations, researchers have begun to explore new technical means. As an emerging deep learning architecture, state space model has strong global feature capturing ability and long-range dependency modeling ability, which can effectively alleviate the computational bottleneck caused by quadratic complexity of traditional network models (such as Transformer). In addition, with the development of multi-modal learning, the technology of image-text fusion that combines image information with text information has gradually attracted attention. For example, the CLIP model as a basic large model has been widely used in the medical field for the fusion and pairing of clinical text and image information, showing good adaptability and flexibility. Through multi-modal feature alignment optimization loss, the difference between image feature distribution and text feature distribution can be further reduced, thereby providing more discriminative identification scores for early risk prediction and typing of coronary artery disease.
[0006] In the field of medical image segmentation, MedSAM encoder as a widely used benchmark model has been proven to perform well in various medical image segmentation tasks due to its strong feature extraction capability. However, existing methods often fail to fully utilize multi-scale features and global context information when processing coronary CTA images, resulting in insufficient segmentation accuracy and robustness. SUMMARY
[0007] To improve image segmentation accuracy and thus improve the accuracy of early risk prediction and typing of coronary artery disease, the present invention provides a deep learning-based method for early risk prediction and typing of coronary artery disease, which is as follows:
[0008] Step 1: Obtain a cardiac angiography CTA image;
[0009] Step 2: Process the cardiac angiography CTA image using a cardiac angiography CTA image segmentation model to complete pixel-level segmentation of the cardiac angiography CTA image;
[0010] Step 3: Use a medical image-text feature fusion module and multi-modal feature alignment optimization loss to weaken the difference between image feature distribution and text feature distribution, and complete early risk prediction and typing of coronary artery disease;
[0011] The cardiac angiography CTA image segmentation model includes a MedSAM encoder, a multi-scale state space feature encoder, and a hierarchical feature reconstruction module. The multi-scale state space feature encoder includes four stages of QSSBlock feature encoders, and the hierarchical feature reconstruction module includes a multi-scale feature fusion module and a multi-scale convolution attention module. The hierarchical feature reconstruction module reconstructs the image features extracted by the MedSAM encoder and the multi-scale state space feature encoder.
[0012] Optionally, the QSSBlock feature encoder comprises a Mamba residual block and a spatial gating block.
[0013] The calculation formula of the Mamba residual block is represented as:
[0014] Y1=BN(Conv3(Conv1(Y′)))
[0015] Y2=Conv3(BN(DW(Y1)))+Y1
[0016] Y=SS2D(LN(Y2))
[0017] Wherein, Y' represents the input feature of the Mamba residual block, BN represents the batch normalization operation, Conv3 represents the convolution layer with the convolution kernel size of 3x3, Conv1 represents the convolution layer with the convolution kernel size of 1x1, DW represents the depth convolution, LN represents the layer normalization operation, and SS2D represents the SS2D block.
[0018] The calculation formula of the spatial gating block is represented as:
[0019] Z2=σ(PW(DW(LN(Z1))*DW(LN(Z1))))
[0020] Z=Z1+Z2
[0021] Wherein, Z1 represents the input feature of the spatial gating block, PW represents the point-wise convolution layer, DW represents the depth separable convolution, LN represents the layer normalization operation, and σ(·) represents the Sigmoid function.
[0022] Optionally, the multi-scale feature fusion module comprises an up-convolution module and a group attention module, the up-convolution module is used for up-sampling the feature map of the current stage to match the size and resolution of the feature map of the next jump connection, and the group attention module uses the high-level semantic feature obtained by the up-convolution module to enhance the low-level semantic feature transmitted through the jump connection, while suppressing irrelevant features.
[0023] Optionally, the formula of the multi-modal feature alignment optimization loss is:
[0024]
[0025] Wherein, U(V, T) represents all possible transmission matrices, n and m represent the class and region vector subscripts respectively, P ij represents the feature distribution, M ij quantifies the difference between the i-th image and the j-th text, and λ is a penalty term related to the distribution P.
[0026] Optionally, the up-convolution module comprises, connected in sequence: a depth separable convolution layer, a batch normalization layer, a ReLU activation function and a convolution layer.
[0027] Optionally, the calculation formula of the grouping attention module is represented as:
[0028] Q(g,x)=R(BN(GC g (g))+R(BN(GC x (x))
[0029] F(g,x)=σ(Conv1(R(Q(g,x))))*x+x
[0030] Wherein, sigma (·) represents a Sigmoid function, BN represents a batch normalization layer, R represents a residual connection, GC g (g) and GC x (x) represent grouping convolution operations on input g and x, Q(g,x) and F(g,x) represent operation results.
[0031] Optionally, the multi-scale convolution attention module firstly aggregates local information through a deep convolution, then captures multi-scale context information by using a multi-branch deep convolution, simulates the relationship between different channels by using a 1x1 convolution, and finally the output of the 1x1 convolution is used as an attention weight to directly perform matrix multiplication on the input of the multi-scale convolution attention module for weighting.
[0032] A second object of the application is to provide a deep learning-based coronary artery disease early risk prediction and typing system, which comprises:
[0033] An image acquisition module configured to acquire a cardiac angiography CTA image;
[0034] An image segmentation module configured to process the cardiac angiography CTA image by using the cardiac angiography CTA image segmentation model, and complete pixel-level segmentation of the cardiac angiography CTA image;
[0035] A prediction and typing module configured to use the medical image text feature fusion module and the multi-modal feature alignment optimization loss to weaken the difference between image feature distribution and text feature distribution, and complete coronary artery disease early risk prediction and typing.
[0036] The cardiac angiography CTA image segmentation model comprises a MedSAM encoder, a multi-scale state space feature encoder and a hierarchical feature reconstruction module, the multi-scale state space feature encoder comprises four-stage QSSBlock feature encoders, and the hierarchical feature reconstruction module comprises a multi-scale feature fusion module and a multi-scale convolution attention module; the hierarchical feature reconstruction module reconstructs image features extracted by the MedSAM encoder and the multi-scale state space feature encoder.
[0037] A third object of the present application is to provide a computer device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, characterized in that the processor implements the deep learning-based coronary artery disease early risk prediction and typing method according to any one of the preceding items when executing the computer program.
[0038] A fourth object of the present application is to provide a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program implements the deep learning-based coronary artery disease early risk prediction and typing method according to any one of the preceding items when executed by a processor.
[0039] The present application has the following advantages:
[0040] The deep learning-based coronary artery disease early risk prediction and typing system provided by the present application realizes efficient and accurate segmentation and risk assessment of coronary CTA images by innovatively combining a multi-scale state space feature encoder, a hierarchical feature reconstruction module and a multi-modal feature alignment optimization loss module.
[0041] The system utilizes the powerful feature extraction capability of the MedSAM encoder and combines the multi-scale state space feature encoder to effectively capture global features and long-range correlations in the cardiac angiography CTA image, and at the same time, through the feature fusion and convolution attention mechanism in the hierarchical feature reconstruction module, the multi-scale information is fully utilized, which significantly improves the segmentation accuracy and the robustness and generalization ability of the model. In addition, the multi-modal feature alignment optimization loss module in the system effectively reduces the distribution difference between image features and text features by fusing the medical image text feature fusion module, realizes the deep fusion of image and clinical text information, provides more discriminative recognition scores for the early risk prediction and typing of coronary artery disease, reduces the computational cost, and improves the clinical application value of the model.
[0042] The application not only provides a new technical means for coronary CTA image segmentation and risk assessment, but also promotes the development of multi-modal fusion technology in the field of medical image processing, provides strong support for early diagnosis, risk stratification and personalized treatment of coronary artery disease, and has wide application prospect and important practical value. BRIEF DESCRIPTION OF DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0044] Figure 1 A flow chart of a coronary artery disease early risk prediction and typing method based on deep learning provided by the present application.
[0045] Figure 2 A whole model architecture diagram of a coronary artery disease early risk prediction and typing method based on deep learning provided by the present application.
[0046] Figure 3 A structural schematic diagram of a multi-scale state space feature encoder provided by the present application.
[0047] Figure 4 A structural schematic diagram of a hierarchical feature reconstruction module provided by the present application. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0049] Embodiment one:
[0050] The embodiment provides a coronary artery disease early risk prediction and typing method based on deep learning, which comprises the following steps:
[0051] Step 1: acquiring a cardiac angiography CTA image;
[0052] Step 2: processing the cardiac angiography CTA image by using a cardiac angiography CTA image segmentation model to complete pixel-level segmentation of the cardiac angiography CTA image;
[0053] Step 3: using a medical image text feature fusion module and a multi-modal feature alignment optimization loss to weaken the difference between image feature distribution and text feature distribution, and completing coronary artery disease early risk prediction and typing.
[0054] The structure of the cardiac angiography CTA image segmentation model is shown in the following formula (I) and includes a MedSAM encoder, a multi-scale state space feature encoder, and a hierarchical feature reconstruction module. Figure 2 The structure of the MedSAM encoder is shown in the following formula (II) and includes a convolutional layer, a batch normalization layer, a depthwise separable convolutional layer, a batch normalization layer, a convolutional layer, a residual connection, a layer normalization layer, and an SS2D block.
[0055] The MedSAM encoder is a widely used benchmark model in the field of medical image segmentation and has strong feature extraction capability.
[0056] The structure of the multi-scale state space feature encoder is shown in the following formula (III) and includes a QSSBlock feature encoder in four stages. Figure 3 The QSSBlock feature encoder is used to capture global features in the cardiac angiography CTA image and establish long-range correlations between instances to address the shortcoming of mutual independence between instances.
[0057] As shown in the following formula (IV), the QSSBlock feature encoder includes a Mamba residual block and a spatial gating block. Figure 3 The Mamba residual block includes, in sequence, two convolutional layers, a batch normalization layer, a depthwise separable convolutional layer, a batch normalization layer, a convolutional layer, a residual connection, a layer normalization layer, and an SS2D block.
[0058] Y1=BN(Conv3(Conv1(Y′)))
[0059] Y2=Conv3(BN(DW(Y1)))+Y1
[0060] Y=SS2D(LN(Y2))
[0061] Wherein Y' represents the input feature of the Mamba residual block, BN represents the batch normalization operation, Conv3 represents the convolutional layer with a convolution kernel size of 3x3, Conv1 represents the convolutional layer with a convolution kernel size of 1x1, DW represents the depthwise convolution, LN represents the layer normalization operation, and SS2D represents the SS2D block.
[0062] The spatial gating block includes, in sequence, a layer normalization layer, two depthwise separable convolutional layers, a sigmoid activation function, a matrix multiplication unit, a pointwise convolutional layer, and a residual connection.
[0063] Z2=σ(PW(DW(LN(Z1))*DW(LN(Z1))))
[0064] Z=Z1+Z2
[0065] Wherein Z1 represents the input feature of the spatial gating block, PW represents the pointwise convolutional layer, DW represents the depthwise separable convolution, LN represents the layer normalization operation, and σ(·) represents the Sigmoid function.
[0066] This spatial gating block can capture more global features while only slightly increasing the computational cost. It uses residual splicing to more efficiently reflow gradients, and by retaining and utilizing the spatial structural information of the image, it reduces the computational cost and significantly improves performance.
[0067] The structure of the hierarchical feature reconstruction module is as follows Figure 4 As shown, the system consists of a sequentially connected multi-scale convolutional attention module and three multi-scale feature fusion modules. The multi-scale convolutional attention module acquires MedSAM encoded features and QSSBlock features from the fourth stage to aggregate local information and capture multi-scale contextual information. The three multi-scale feature fusion modules acquire QSSBlock features from the first, second, and third stages, respectively. The hierarchical feature reconstruction module optimizes performance and computational efficiency. Leveraging the multi-scale feature fusion module and the multi-scale convolutional attention module, the system significantly enhances feature maps through multi-scale convolution and grouped convolutional attention. This module is very effective in capturing complex spatial relationships while focusing on salient areas.
[0068] like Figure 4 As shown in the figure, the multi-scale feature fusion module includes an up-convolution module and a grouped attention module. The up-convolution module is used to upsample the feature map of the current stage to match the size and resolution of the feature map of the next jump connection; the grouped attention module is used to use the high-level semantic features obtained by the up-convolution module to enhance the low-level semantic features transmitted through the jump connection and suppress irrelevant features.
[0069] The upconvolution module consists of a depthwise separable convolutional layer, a batch normalization layer, a ReLU activation function, and a convolutional layer, connected in sequence. This upconvolution module progressively upsamples the feature maps of the current stage to match the size and resolution of the feature maps of the next skip connection. The upconvolution module first upsamples the feature maps by a factor of 2. The upscaled feature maps are then enhanced by applying a 3×3 depthwise convolution, a batch normalization layer, and a ReLU activation function. Finally, a 1×1 convolution is used to reduce the number of channels to match the next stage.
[0070] The group attention module includes two group convolution layers, one batch normalization layer, three ReLU activation functions, two residual connections, one convolution layer, one sigmoid activation function, and one matrix multiplication. The group attention module gradually combines the feature maps with the attention coefficients learned by the network, thereby improving the activation of relevant features and the suppression of irrelevant features. The input variables g and y are processed by applying separate 1x1 group convolutions. The convolution features are then normalized using batch normalization and merged through element addition. The generated feature maps are activated through ReLU, after which a 1x1 convolution and a batch normalization layer are applied to obtain a single-channel feature map. The resulting single-channel feature map is then passed through a Sigmoid activation function to produce attention coefficients. The output of this transformation scales the input features x through element multiplication, and finally, x is connected in residual connection with this transformation, which can reduce the impact of high-level semantic features on low-level semantic features, thereby avoiding performance degradation of the overall model. The calculation formula is as follows:
[0071] Q(g, x) = R(BN(GC g (g)) + R(BN(GC x (x))
[0072] F(g, x) = σ(Conv1(R(Q(g, x)))) * x + x
[0073] where σ(·) represents the Sigmoid function, BN represents the batch normalization layer, R represents the residual connection, GC g (g) and GC x (x) represent group convolution operations on inputs g and x, Q(g, x) and F(g, x) represent the operation results.
[0074] The multi-scale convolution attention module includes four depthwise separable convolution layers, one residual connection, one matrix multiplication unit, and one convolution layer. The multi-scale convolution attention module first uses a 5x5 depthwise convolution to aggregate local information, then enters a multi-branch depthwise convolution to capture multi-scale context information, and then enters a 1x1 convolution to simulate the relationship between different channels. The output of the 1x1 convolution is used as the attention weight, and the input of the multi-scale convolution attention module is directly multiplied by the matrix to be weighted.
[0075] To utilize the CLIP model's capabilities in image-text alignment, the CLIP model is introduced in the medical image-text feature fusion module, and relevant clinical text information is combined with the segmentation image to improve the accuracy of early risk prediction and classification of coronary artery disease. Specifically, the CLIP is used to calculate the segmentation image refinement region features generated by the segmentation model and the text encoder of the list of coronary artery disease categories. The calculation formula is as follows:
[0076]
[0077] where S denotes the alignment matrix, V denotes the region feature, T denotes the text feature, v, t denote different learnable feature projections on image and text features, and τ is an adjustable parameter. On this basis, the different image-text pair representations are calculated by cross-entropy loss, and the image-text pairs with the same semantics are fused, and the formula is as follows:
[0078]
[0079] where Y ij denotes the coronary artery disease category corresponding to the image, S ij denotes the text encoder calculation result, and n and m denote the category and region vector indexes respectively. This method improves the coronary artery disease classification by using the alignment capability of the medical image-text feature fusion module at the image region level.
[0080] The multi-modal feature alignment optimization loss reduces the computational cost by using the idea of contrast learning, weakens the difference between image feature distribution and text feature distribution, and improves the accuracy of coronary artery disease classification, and the formula is as follows:
[0081]
[0082] where U(V, T) denotes all possible transmission matrices, n and m denote the category and region vector indexes respectively, P ij denotes the feature distribution, and M ij quantifies the difference between the i-th image and the j-th text, and λ is a penalty term related to the distribution P.
[0083] Embodiment Two:
[0084] The embodiment provides a training method of a coronary artery disease early risk prediction and typing model, including the following steps:
[0085] Step 1: Obtain a coronary artery disease CTA image and clinical information dataset; 70% of the dataset is used for training; 10% of the dataset is used for testing; 20% of the dataset is used for verification; all images are adjusted to 512x512 pixels; use opencv to read the original image and convert it to a pixel matrix.
[0086] Step 2: Train a deep learning-based coronary artery disease early risk prediction and typing model, and the structure of the model is the structure recorded in Embodiment One.
[0087] The training loss function comprises a stage loss, a fusion loss and a final target loss, wherein the stage loss is used to fully utilize the pixel-level label to enhance the supervision ability, the fusion loss is used to optimize the high-level semantic information and edge detail information of the extracted features, and the final target loss is used to guide the network segmentation learning.
[0088] The calculation formula of the stage loss is as follows:
[0089]
[0090] Y t = I (t) * Y t-1 + (1-I (t) ) * Y t n denotes label information, denotes a calculation result, t denotes a stage number, and I is an indication function.
[0091] The calculation formula of the fusion loss is as follows:
[0092]
[0093] Y t = I (t) * Y t-1 + (1-I (t) ) * Y t n denotes label information, denotes a calculation result after fusion, and I is an indication function.
[0094] The calculation formula of the final target loss is as follows:
[0095]
[0096] Y t = I (t) * Y t-1 + (1-I (t) ) * Y t denotes a stage loss, and L fuse. denotes a fusion loss.
[0097] Further, referring to Figure 2 , in order to utilize the function of CLIP in image-text alignment, a CLIP model is introduced in the medical image-text feature fusion module, and relevant clinical text information is combined with the segmentation image, so as to improve the early risk prediction and classification accuracy of coronary artery disease. Specifically, a text encoder of a list of coronary artery disease categories is used to calculate the segmentation image refinement region features generated by the segmentation model. The calculation formula is as follows:
[0098]
[0099] wherein S denotes an alignment matrix, V denotes a region feature, T denotes a text feature, v and t denote different learnable feature projections on the image and text features, and τ is an adjustable parameter. On this basis, cross-entropy loss is used to calculate different image-text pair representations, and image-text pairs with the same semantics are fused, and the formula is as follows:
[0100]
[0101] where Y ij represents the coronary artery disease category corresponding to the image, S ij represents the text encoder calculation result, n and m represent the category and region vector indexes respectively. This method improves the coronary artery disease classification by utilizing the alignment capability of the medical image text feature fusion module at the image region level.
[0102] The multi-modal feature alignment optimization loss reduces the operation cost by using the contrast learning idea, weakens the difference between the image feature distribution and the text feature distribution, and improves the coronary artery disease classification accuracy, and the formula is as follows:
[0103]
[0104] where U(V, T) represents all possible transmission matrices, n and m represent the category and region vector indexes respectively, P ij represents the feature distribution, M ij quantifies the difference between the ith image and the jth text, and λ is a penalty term related to the distribution P.
[0105] Embodiment three:
[0106] The embodiment provides a coronary artery disease early risk prediction and classification system based on deep learning, which comprises:
[0107] An image acquisition module configured to acquire a cardiac angiography CTA image;
[0108] An image segmentation module configured to process the cardiac angiography CTA image by using a cardiac angiography CTA image segmentation model to complete pixel-level segmentation of the cardiac angiography CTA image;
[0109] A prediction and classification module configured to weaken the difference between the image feature distribution and the text feature distribution by using a medical image text feature fusion module and a multi-modal feature alignment optimization loss, and complete coronary artery disease early risk prediction and classification.
[0110] The cardiac angiography CTA image segmentation model comprises a MedSAM encoder, a multi-scale state space feature encoder and a hierarchical feature reconstruction module, the multi-scale state space feature encoder comprises four-stage QSSBlock feature encoders, and the hierarchical feature reconstruction module comprises a multi-scale feature fusion module and a multi-scale convolution attention module; the hierarchical feature reconstruction module reconstructs the image features extracted by the MedSAM encoder and the multi-scale state space feature encoder.
[0111] Some steps in the embodiment of the application can be realized by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk.
[0112] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for early risk prediction and classification of coronary artery disease based on deep learning, characterized in that: The method comprises: Step 1: Acquire cardiac angiography CTA images; Step 2: Processing the cardiac angiography CTA image using a cardiac angiography CTA image segmentation model to complete pixel-level segmentation of the cardiac angiography CTA image; Step 3: Utilize the medical image-text feature fusion module and multimodal feature alignment to optimize the loss, reduce the difference between image feature distribution and text feature distribution, and complete the early risk prediction and classification of coronary artery disease; The cardiac angiography (CTA) image segmentation model includes: a MedSAM encoder, a multi-scale state-space feature encoder, and a hierarchical feature reconstruction module. The multi-scale state-space feature encoder includes a four-stage QSSBlock feature encoder. The hierarchical feature reconstruction module includes a multi-scale feature fusion module and a multi-scale convolutional attention module. The hierarchical feature reconstruction module reconstructs the image features extracted by the MedSAM encoder and the multi-scale state-space feature encoder.
2. The method for early risk prediction and classification of coronary artery disease based on deep learning according to claim 1, characterized in that: The QSSBlock feature encoder includes: a Mamba residual block and a spatial gating block; The calculation formula of the Mamba residual block is expressed as: Y1=BN(Conv3(Conv1(Y′))) Y2=Conv3(BN(DW(Y1)))+Y1 Y=SS2D(LN(Y2)) Wherein, Y′ represents the input features of the Mamba residual block, BN represents the batch normalization operation, Conv3 represents the convolution layer with a convolution kernel size of 3x3, Conv1 represents the convolution layer with a convolution kernel size of 1x1, DW represents depthwise convolution, LN represents the layer normalization operation, and SS2D represents the SS2D block; The calculation formula of the spatial gating block is expressed as: Z2=σ(PW(DW(LN(Z1))*DW(LN(Z1)))) Z=Z1+Z2 Wherein, Z1 represents the input feature of the spatial gating block, PW represents the point-by-point convolution layer, DW represents the depth-wise separable convolution, LN represents the layer normalization operation, and σ(·) represents the Sigmoid function.
3. The method for early risk prediction and classification of coronary artery disease based on deep learning according to claim 1, characterized in that: The multi-scale feature fusion module includes an upconvolution module and a group attention module. The upconvolution module is used to upsample the feature map of the current stage to match the size and resolution of the feature map of the next jump connection; the group attention module uses the high-level semantic features obtained by the upconvolution module to enhance the low-level semantic features transmitted through the jump connection, while suppressing irrelevant features.
4. The method for early risk prediction and classification of coronary artery disease based on deep learning according to claim 1, characterized in that: The formula for the multimodal feature alignment optimization loss is: Where U(V, T) represents all possible transmission matrices, n and m represent the category and region vector subscripts respectively, P ij represents the feature distribution, M ij quantifies the difference between the i-th image and the j-th text, and λ is a penalty term related to the distribution P.
5. The method for early risk prediction and classification of coronary artery disease based on deep learning according to claim 3, characterized in that: The upconvolution module includes: a depthwise separable convolution layer, a batch normalization layer, a ReLU activation function and a convolution layer connected in sequence.
6. The method for early risk prediction and classification of coronary artery disease based on deep learning according to claim 3, characterized in that: The calculation formula of the group attention module is expressed as: Q(g,x)=R(BN(GC g (g))+R(BN(GC x (x)) F(g,x)=σ(Conv1(R(Q(g,x))))*x+x Where σ(·) represents the Sigmoid function, BN represents the batch normalization layer, R represents the residual connection, and GC g (g) and GC x (x) represents the grouped convolution operation on the inputs g and x, and Q(g,x) and F(g,x) represent the operation results.
7. The method for early risk prediction and classification of coronary artery disease based on deep learning according to claim 3, characterized in that: The multi-scale convolutional attention module first aggregates local information through deep convolution, then uses multi-branch deep convolution to capture multi-scale contextual information, and uses 1×1 convolution to simulate the relationship between different channels. Finally, the output of the 1×1 convolution is used as the attention weight, and the input of the multi-scale convolutional attention module is directly matrix multiplied for weighting.
8. A deep learning-based early risk prediction and classification system for coronary artery disease, characterized by: The system comprises: An image acquisition module is configured to acquire a cardiac angiography (CTA) image; an image segmentation module configured to process the CTA image using the CTA image segmentation model to perform pixel-level segmentation of the CTA image; The prediction and classification module is configured to use the medical image-text feature fusion module and multimodal feature alignment optimization loss to reduce the difference between image feature distribution and text feature distribution to complete early risk prediction and classification of coronary artery disease; The cardiac angiography (CTA) image segmentation model includes: a MedSAM encoder, a multi-scale state-space feature encoder, and a hierarchical feature reconstruction module. The multi-scale state-space feature encoder includes a four-stage QSSBlock feature encoder. The hierarchical feature reconstruction module includes a multi-scale feature fusion module and a multi-scale convolutional attention module. The hierarchical feature reconstruction module reconstructs the image features extracted by the MedSAM encoder and the multi-scale state-space feature encoder.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the method for early risk prediction and classification of coronary artery disease based on deep learning as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements a method for early risk prediction and classification of coronary artery disease based on deep learning as described in any one of claims 1 to 7.