A method and system for predicting the benefit of interventional therapy for patients with coronary occlusive lesions based on multi-modal images
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-11
AI Technical Summary
[0002]目前,冠状动脉慢性完全闭塞病变(CTO)介入治疗(PCI)风险高、难度大、并发症多,即便手术成功,仍有部分患者无法改善生活质量与预后;术前缺乏准确、高效的患者获益预测手段,难以精准筛选适宜治疗人群
Smart Images

Figure CN122552048A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interventional therapy benefit prediction technology, and more specifically to a method and system for predicting interventional therapy benefits in patients with coronary artery occlusion based on multimodal imaging. Background Technology
[0002] Currently, percutaneous coronary intervention (PCI) for chronic total occlusion (CTO) is high-risk, difficult, and has many complications. Even if the surgery is successful, some patients still cannot improve their quality of life and prognosis. There is a lack of accurate and efficient means of predicting patient benefits before the operation, making it difficult to accurately select suitable patients for treatment.
[0003] Coronary CTO lesions require assessment of internal occlusion features, myocardial function, tissue characteristics, and collateral circulation. Multimodal imaging (coronary angiography, CTA, cardiac MRI) features are complex and difficult to extract, easily generating useless or indistinguishable features. Existing multimodal image feature fusion techniques are immature and cannot efficiently utilize complementary information from different modalities. Traditional covariance estimation suffers from instability, singularity, and confounding effects, and is difficult to integrate with end-to-end neural network training, thus failing to support accurate prediction.
[0004] Therefore, how to propose a method and system for predicting the benefits of interventional treatment for patients with coronary artery occlusion based on multimodal imaging, and overcome the shortcomings of existing technologies, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for predicting the benefit of interventional treatment in patients with coronary artery occlusion based on multimodal imaging, and an intelligent prediction scheme for interventional treatment benefit that efficiently processes multimodal cardiovascular images, solving the challenges of clinical screening and evaluation. To achieve the above objectives, the present invention adopts the following technical solution: A method for predicting the benefit of interventional treatment in patients with coronary artery occlusion based on multimodal imaging includes: Multimodal images are acquired, and multimodal features are extracted from the multimodal images using a multimodal image feature extraction model; A multimodal image feature fusion model is constructed to fuse the multimodal features; A model for predicting the benefit of interventional treatment for patients with coronary artery occlusion is constructed based on a visual representation model of the covariance matrix. The multimodal imaging data to be tested is acquired and input into the interventional treatment benefit prediction model for patients with coronary artery occlusion to obtain the benefit prediction results.
[0006] Optionally, the multimodal images include: coronary angiography, coronary CTA, and cardiac MRI images.
[0007] Optionally, the multimodal image feature extraction model includes: Acquire the input image and extract an initial pixel-level representation from the input image; The initial pixel-level representation is grouped according to semantics to obtain a semantic-level representation; Semantic-level representations are processed using a feature enhancement module and a feature interaction module guided by a relation matrix. The initial pixel-level representation is enhanced by the processed semantic-level representation, resulting in a pixel-level representation with semantic relationships.
[0008] Optionally, the multimodal image feature fusion model includes: a spatial decoupling feature fusion module, a spatial feature fusion module, a spatial cross-attention feature fusion module, and a feature convergence layer; the three modules respectively perform multimodal local fusion within the same spatial step, global fusion of all spatial steps and all modalities, and cross-modal complementary fusion with the dominant modality as the query; the feature convergence layer performs weighted convergence of the outputs of the three modules and outputs the multimodal fused features. Let the outputs of the three parallel modules be respectively... , , Then the feature aggregation layer Calculate the weights of the three branches and obtain the final multimodal fusion feature map. .
[0009] Optionally, the visual representation model based on the covariance matrix includes: a sparse inverse covariance estimation module based on a neural network and an iterative sparse inverse covariance estimation module. The sparse inverse covariance estimation module based on the neural network is designed as a structured layer of the neural network. The input is a set of local descriptors of the neural network extracted from the convolutional feature map, and the output is a sparse inverse covariance matrix. The iterative sparse inverse covariance estimation module is used for end-to-end training of the sparse inverse covariance estimation module, and solves the optimization problem of the sparse inverse covariance matrix in the forward and backward propagation steps.
[0010] Optionally, the sparse inverse covariance estimation module based on neural networks includes: Extract a set of local descriptors from the convolutional feature maps and calculate the sample-based covariance matrix. ,make This represents the corresponding sparse inverse covariance matrix; The off-diagonal terms are used to capture the direct correlation between different descriptor components, and they are zero if the two components are independent after eliminating the effects of confounding variables. Through the By applying SPD constraints and sparse priors, the problem is solved by maximizing the log-likelihood of the data with a penalty. The estimation problem guides the assessment of the connectivity of sparse graphs, and the optimal solution is: ; in, This represents the corresponding sparse inverse covariance matrix. It is based on the sample covariance matrix. , and Let the determinant, trace, and sum of the vectorized matrix be represented respectively. Norm; Item pair Applying structural sparsity, Used to control the trade-off between sparsity and log-likelihood estimation.
[0011] Optionally, the iterative sparse inverse covariance estimation module includes: The iterative sparse inverse covariance estimation module performs end-to-end training on the sparse inverse covariance estimation module. set up It is the objective function of the optimal solution. Through the Optimize by taking the derivative: ; in, and Each contains The positive and negative parts are optimized using projective gradient descent.
[0012] Optionally, the optimization via projected gradient descent includes: Multimodal feature map fusion Remodeled into a local descriptor data matrix Where N is the number of spatial locations or local descriptors in the fused feature map. For the fusion feature descriptor at the nth spatial location; Centralized processing is performed to obtain ,in , Represent a vector of length N consisting entirely of 1s; calculate the sample covariance matrix based on the centered local descriptor data matrix. and will Input the sparse inverse covariance estimation module for subsequent iterative optimization.
[0013] Optionally, the iterative sparse inverse covariance estimation module includes: In estimation At this point, the Newton-Schulz iterative method is used to approximate the inverse matrix, with a convergence condition applied. Through the The trace is normalized, that is inverse square root matrix The square of the trace is normalized to invert it, that is... The accuracy matrix is obtained. ; make The iSICE iteration begins by projecting gradient descent onto the gradient of SICE, dividing the sparse inverse covariance matrix S into its positive and negative parts: Then apply the PGD steps respectively; The optimal solution is rewritten in two parts: ; Then take a PGD step to update. and : ; in, A function used to reproject the gradient of PGD onto the feasible region of each boundary constraint, a constant. For learning rate, , This is used to control the decay of the learning rate; use and Composition of the current Estimated value: ; in, Used to ensure matrix It is symmetrical. It serves as input for the visual representation model subsequently used for profit prediction.
[0014] Optionally, a system for predicting the benefit of interventional treatment for patients with coronary artery occlusion based on multimodal images includes: acquiring multimodal images and extracting multimodal features from the multimodal images using a multimodal image feature extraction model; A multimodal image feature fusion model is constructed to fuse the multimodal features; A model for predicting the benefit of interventional treatment for patients with coronary artery occlusion is constructed based on a visual representation model of the covariance matrix. The multimodal imaging data to be tested is acquired and input into the interventional treatment benefit prediction model for patients with coronary artery occlusion to obtain the benefit prediction results.
[0015] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a method and system for predicting the benefit of interventional treatment for patients with coronary artery occlusion based on multimodal imaging, which has the following beneficial effects: This invention proposes a method for predicting the benefit of interventional treatment for patients with coronary artery occlusion based on multimodal imaging, comprising: acquiring multimodal images; extracting multimodal features from the multimodal images using a multimodal image feature extraction model; constructing a multimodal image feature fusion model to fuse the multimodal features; constructing a benefit prediction model for interventional treatment of patients with coronary artery occlusion based on a visual representation model of the covariance matrix; acquiring the multimodal image data to be tested, inputting it into the benefit prediction model for interventional treatment of patients with coronary artery occlusion, and obtaining the benefit prediction result.
[0016] This invention constructs a multimodal image feature extraction model based on an intervention-driven relationship network. It automatically filters useless / indistinguishable features using a diagnostic deletion paradigm, accurately identifying semantically relevant and highly discriminative pixel-level features. This model can be seamlessly integrated with existing frameworks, improving feature representation quality. Furthermore, it constructs a novel cross-modal fusion framework based on a spatially decoupled feature algorithm. Utilizing three core modules—a spatially decoupled feature fusion module, a spatial feature fusion module, and a spatial cross-attention feature fusion module—it decouples and integrates features at different time steps, efficiently utilizing modal complementary information and significantly improving the accuracy and efficiency of multimodal data processing. Finally, it constructs a visual representation model based on the covariance matrix for predicting the benefit of interventional treatment in patients with coronary artery occlusion. This model utilizes a neural network-based sparse inverse covariance estimation (SICE) module to eliminate confounding effects, accurately capture direct correlations between features, and improve the interpretability and stability of feature representation. The iterative sparse inverse covariance estimation (iSICE) module enables end-to-end trainability and supports GPU parallel computing, solving the problems of slow speed and inability to backpropagate in traditional methods. The intelligent prediction model can accurately assess the benefits of interventional treatment for CTO patients before surgery, assist in the clinical screening of the optimal treatment population, reduce surgical risks, and improve patient prognosis and the efficiency of medical resource allocation. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1 This invention provides a flowchart illustrating a method for predicting the benefits of interventional treatment for patients with coronary artery occlusion based on multimodal imaging.
[0019] Figure 2 The diagram shows the network structure framework of the intervention-driven relationship provided by this invention.
[0020] Figure 3 The structural framework diagram of the multimodal image feature fusion model provided by the present invention is shown.
[0021] Figure 4 This is an example diagram illustrating partial correlation for understanding the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] This invention discloses a method for predicting the benefit of interventional treatment in patients with coronary artery occlusion based on multimodal imaging, such as... Figure 1 As shown, it includes: Multimodal images are acquired, and multimodal features are extracted from the multimodal images using a multimodal image feature extraction model; A multimodal image feature fusion model is constructed to fuse the multimodal features; A model for predicting the benefit of interventional treatment for patients with coronary artery occlusion is constructed based on a visual representation model of the covariance matrix. The multimodal imaging data to be tested is acquired and input into the interventional treatment benefit prediction model for patients with coronary artery occlusion to obtain the benefit prediction results.
[0024] In a specific implementation, a method for predicting the benefit of interventional treatment in patients with coronary artery occlusion based on multimodal imaging is proposed, which includes the following: Step 1: Develop a model for feature extraction and fusion of different modalities of cardiovascular imaging for coronary artery occlusion (CTO) lesions. The model is developed by applying intervention-driven relationship algorithms and spatial decoupling feature algorithms to complete the feature extraction and fusion tasks for different modalities of cardiovascular imaging for CTO lesions.
[0025] (1) Construct a multimodal image feature extraction model for coronary artery occlusion lesions
[0026] Clinical assessment of coronary artery CTO lesions requires high precision, necessitating accurate evaluation of the internal characteristics of the occluded lesion, myocardial motion function, myocardial tissue features, and collateral circulation using imaging techniques before treatment. Therefore, the imaging features of coronary artery occlusion lesions are complex and difficult to extract. To address this, this invention proposes an intervention-driven relationship network to solve the problems of generating useless and indistinguishable features during feature extraction, thereby achieving efficient extraction of cardiovascular imaging features from different modalities such as coronary angiography, coronary CTA, and cardiac MRI.
[0027] The intervention-driven relationship network proposed in this invention utilizes a diagnostic deletion paradigm to guide the construction of relationships between pixel semantic categories and pixels, thereby enhancing features. Through the interaction between semantic categories and pixels, the network accurately identifies semantically relevant and discriminative features, thus avoiding the extraction of useless and indistinguishable features. The intervention-driven relationship network is as follows: Figure 2 As shown, given the input image This ultimately results in an enhanced pixel-level representation. .
[0028] This network can be seamlessly integrated with existing feature extraction frameworks, enhancing the extracted pixel-level feature representations. The initial pixel-level representations are divided into semantic-level representations to facilitate pixel relationship modeling. This is achieved first using a feature enhancement module and a relation matrix... The guided feature interaction module processes semantic-level representations. Among these, the relation matrix... Learned through the proposed diagnostic deletion paradigm. Relationship matrix. The semantic-level representation carries the relationships between semantic categories, thus the resulting semantic-level representation is a semantically interactive representation. Subsequently, the processed semantic-level representation is used to enhance the initial pixel-level representation, thereby obtaining a pixel-level representation with semantic relationships.
[0029] 1) First use the backbone network From an input image Extracting high-level pixel representation ,in, and Indicates the height and width of the image. This refers to the number of feature channels. To facilitate seamless integration of this paradigm with existing feature extraction frameworks, this invention proposes the following pseudo-label generation process: ; (1); in, This represents a recognition function used to activate the pixel representation, while It is a convolutional layer that projects pixel representations onto a probability space. for The predicted probabilities of each category, This is a pseudo-tag.
[0030] 2) Since directly modeling pixel-level relationships using the diagnostic deletion paradigm is computationally infeasible, this invention simplifies pixel-level relationship modeling to semantic-level relationship modeling. That is, by aggregating corresponding input pixel representations using the same pseudo-labels, the model is... Grouped into semantic-level representations , (2); in, , It is the input image The number of categories that exist in it. It is a weighted summation operation that sums the pixel representations with... The corresponding probabilities are combined as weights. For spatial location ( , The pixel-level feature vector at position () is given by the given information. This indicates that the value at that position is taken across all channels. For position ( , The pseudo-label at the location.
[0031] 3) A feature enhancement module was further introduced to improve... The ability to recognize pixels helps enhance pixel representation. Specifically, formula (2) is modified as follows: (3); in, This represents a feature concatenation operation. This represents the discriminant vector for each class in the training dataset. Constructing by learning the representation of the dataset for each category This invention will class k. The update strategy is simplified to: ; in, To update momentum, the default value is set to 0.1. Before training, this invention simply... Initialize as a matrix of all zeros. ( The number of feature channels is calculated as follows: ; in, Indicates the first The mean of the feature vectors obtained by the class under the constraints of the real mask, where GT represents the real segmentation mask. This indicates that the extracted pixel representations with the same real label are averaged.
[0032] 4) Following the feature enhancement module, this invention proposes using a semantic-level relation matrix. make The matrix is updated through mutual interaction to enrich itself, and is updated using the proposed diagnostic deletion paradigm. The process of the feature interaction module can be represented as follows: ; Among them, the present invention adopts To represent matrix multiplication, It is an enhanced semantic-level representation. Used for conversion To adapt The shape, and is described as: ; Among them, the introduction To control two in There is no interaction between semantically weaker representations. It is a threshold used to identify these weak relationships. This indicates that the elements actually appearing in the current image and being aggregated are... The original number of the semantic category. In actual calculation, if... Less than t, the present invention will Let it be negative infinity. Then, Rearranged to the original pixel representation shape to enhance... , ; in, yes Medium category The corresponding matrix index. It is a feature representation initialized as a zero matrix, based on the pseudo-label at each pixel position using... The text indicates that it is filled. For all dimensions of the feature in this row, that is, the first... The complete feature vector corresponding to the row.
[0033] 5) Use To enhance the final prediction , ; ; in, This invention employs Self-Attention, a self-attention mechanism, to further process enhanced pixel representations in order to balance the diversity and discriminative power of similar pixels. Final pixel prediction The result is: ; in, For an 8x upsampling operation, Implemented by convolutional layers, used to generate pixel class probabilities. This is a bilinear interpolation operation. During training, the overall objective function of the backpropagation algorithm is defined as: ; in, and They are respectively and GT, Cross-entropy loss between and GT It is balance and Hyperparameters.
[0034] 6) In order to calculate A diagnostic deletion paradigm was proposed. First, It is initialized as an identity matrix. This invention... Randomly delete a semantic-level representation from the middle to obtain The distribution of the deleted representations is as follows . No. The deletion probability is represented by each semantic category. This is related to the number of times the category has been deleted, and can be expressed as: ; in, For category Inverse frequency weight, It is a category The number of times it has been deleted. As input, pixel predictions are calculated by reusing formulas (6)-(10). And further calculations were performed respectively. and Pixel-level cross-entropy loss and Then, from Extracting semantic categories The loss values are as follows: ; right Perform the same operation to obtain .
[0035] 7) Based on the two sets of loss values extracted and , can be The deleted number Semantic categories and reserved categories The present invention models the relationships between them. Specifically, the present invention calculates... and Mean and variance between: ; ; in, and They represent and The average value. Equation (14) shows that considering the mean and variance of the loss variation is helpful for pixel recognition. To utilize this model relationship, this invention maintains two relationship matrices during training. and And update it to: ; ; in, and These are the momentum of two matrices, both empirically set to 0.1. and For the relation matrix, by calculating formulas (6)-(8), we can obtain the following results respectively. and Finally, combining them, we obtain the formula used in equation (9). , .
[0036] Step 2: Construct a multimodal imaging feature fusion model for coronary artery occlusion lesions
[0037] Treatment plans for coronary artery occlusion (CTO) primarily rely on the internal features of the occluded lesion and collateral circulation information provided by coronary angiography and coronary CTA. Furthermore, information such as the functional and histological characteristics of the myocardium supplied by the affected vessels, provided by cardiac MRI, is also invaluable in guiding clinical treatment of CTO lesions. Multimodal imaging can provide more information in the preoperative assessment of coronary CTO lesions and is very helpful in predicting the benefit of interventional treatment for patients with coronary CTO. However, the fusion of multimodal image features for coronary occlusion lesions still presents technical challenges. Therefore, this invention proposes a cross-modal fusion framework for improving the processing performance of multimodal image data from coronary angiography, coronary CTA, and cardiac MRI.
[0038] Specifically, this framework is a novel cross-modal fusion framework, comprising a spatially decoupled feature fusion module, a spatial feature fusion module, a spatial cross-attention feature fusion module, and a feature convergence layer. The logical relationship between these three modules is as follows: the spatially decoupled feature fusion module acts as a local branch, modeling the correspondence between different modalities within the same spatial unit; the spatial feature fusion module acts as a global branch, modeling the overall dependency relationship between all spatial units and all modalities; and the spatial cross-attention feature fusion module acts as a complementary branch, using a preset primary modality as the query and introducing complementary information provided by other modalities. The three branches share the features output by the aforementioned multimodal image feature extraction model, undergo adaptive weighted integration in the feature convergence layer, and finally output multimodal fusion features for the covariance matrix visual representation model. The design of this framework aims to significantly improve the accuracy and efficiency of data processing by decoupling and reintegrating feature information at different time steps and effectively utilizing complementary information between multimodalities through a cross-attention mechanism.
[0039] Specifically, it includes: a spatial decoupling feature fusion module, a spatial feature fusion module, and a spatial cross-attention feature fusion module, as well as a feature aggregation layer connected to the outputs of the three modules. The spatial decoupling feature fusion module performs local self-attention fusion of features from different modalities within the same spatial step size. The spatial feature fusion module performs masked global self-attention fusion across all spatial steps and all modalities. The spatial cross-attention feature fusion module performs cross-modal complementary information fusion using the main modal feature as the query and the other modal features as keys and values. The feature aggregation layer weightedly aggregates the fused features output from the three modules into a multimodal fusion feature.
[0040] Specifically, let the m-th modality image of the i-th patient be... The corresponding modal feature maps are obtained through the multimodal image feature extraction model. Where m=1,...,M. The feature sources of the three types of fusion modules are all... However, the input organization methods before entering each module are different: the spatial decoupling feature fusion module uses modal features within the same spatial unit and learnable fusion labels to form a local input sequence, and the output... The spatial feature fusion module uses all spatial units, all modal features, and the learnable fusion labels corresponding to the spatial units to form a global input sequence, and outputs... The spatial cross-attention feature fusion module uses preset main modal features as queries and other modal features as keys and values to perform cross-attention fusion, outputting... .
[0041] Specifically, will Expanded according to spatial location or spatial unit Where T represents the number of spatial units, This represents the modal feature at the t-th spatial unit. Describes the linear mapping function for the m-th mode. Indicates spatial location embedding, This represents modality embedding. Based on the same set of feature sources mentioned above, Figure 3 The spatial decoupling feature fusion module of A in the middle uses the multimodal features and fusion labels within each spatial unit t to form a local sequence, and the output is denoted as , Figure 3 The spatial feature fusion module of B consists of a global sequence composed of all T spatial units, all M modal features, and T fusion labels, and the output is denoted as , Figure 3 The spatial cross-attention feature fusion module in C constructs queries using preset primary modal features and constructs keys and values using other modal features. The output is denoted as... .
[0042] The feature aggregation layer is based on Calculate the weights of the three branches and output the final multimodal fusion feature map. Where GAP(·) represents global average pooling and LN(·) represents layer normalization.
[0043] (1) Spatial decoupling feature fusion module
[0044] like Figure 3 As shown in Figure A, the spatial decoupling feature fusion module is used to complete multimodal local fusion within the same spatial unit. For the t-th spatial unit of the i-th patient, the modal features are combined with learnable fusion labels. Concatenate into the input sequence .
[0045] Then Input by A self-attention fusion processor composed of Transformer encoder blocks, the computation of the l-th layer is as follows: ; ; Where MHSA represents multi-head self-attention mechanism, LN represents layer normalization, and MLP represents multilayer perceptron. Finally, the th... The representations corresponding to the learnable fusion labels in the layer output sequence serve as the spatial decoupling fusion features of that spatial unit. .
[0046] Rearrange the outputs of all spatial units in spatial order to obtain a spatially decoupled and fused feature map. This branch emphasizes the local correspondence between modalities such as coronary angiography, coronary CTA, and cardiac MRI within the same spatial unit.
[0047] (2) Spatial feature fusion module
[0048] like Figure 3 As shown in Figure B, the spatial feature fusion module is used to complete the global fusion of all spatial units and all modalities. This module no longer performs fusion only within a single spatial unit, but instead combines the learnable fusion labels corresponding to T spatial units and the T×M modal features to form a global input sequence. To constrain the output of the t-th spatial unit to focus on its associated spatial units and multimodal features, this module sets an attention mask in the self-attention process. and through Update the masked Transformer encoder block. ; ; Here, Masked-MHSA represents a multi-head self-attention mechanism with an attention mask. The global fusion output corresponding to the t-th spatial unit is... The outputs of all spatial units are rearranged to obtain a spatial feature fusion feature map. This branch is used to supplement the spatial decoupling feature fusion module, which only focuses on local correspondences, and to capture dependency information across spatial units and modes as a whole.
[0049] (3) Spatial Cross-Attention Feature Fusion Module
[0050] like Figure 3 As shown in Figure C, the spatial cross-attention feature fusion module is used to complete the complementary fusion of the primary mode and the auxiliary mode. Let... The primary modality is preset and can be set to coronary CTA, coronary angiography, or cardiac MRI according to task requirements. The primary modality features provide queries, while other modal features provide keys and values. For the t-th spatial unit, the following is defined: ; ; , ; in, , and These represent linear mapping functions for the query, key, and value, respectively. The l-th layer cross-attention fusion includes main modality self-attention update and cross-modality attention update. ; ; ; MHCA stands for Multi-head Cross-Attention Mechanism. After layer cross-attention, the spatial cross-attention fusion feature of the t-th spatial unit is obtained. The outputs of all spatial units are rearranged to obtain a spatial cross-attention fusion feature map. This branch explicitly defines the direction of information flow as an auxiliary mode to supplement the main mode, thereby enhancing information in the main mode that is related to occlusive lesions, collateral circulation, myocardial function, or tissue characteristics but is underexpressed.
[0051] (4) Convergence of three-branch features and coupling with the covariance visual representation model
[0052] based on Figure 3 China A Figure 3 China B and Figure 3 The three outputs obtained from C , and This invention adaptively calculates the weights of the three branches through a feature convergence layer. It outputs the final multimodal fusion feature map. Where GAP(·) represents global average pooling, , and These represent the weights of the three fusion branches, and The final multimodal fusion feature map is as follows: .
[0053] Then Remodeled into a local descriptor data matrix Where N is the number of spatial locations or local descriptors in the fused feature map. For the fusion feature descriptor at the nth spatial location; Centralized processing is performed to obtain ,in And calculate the sample covariance matrix. .Should As input to the visual representation model based on the covariance matrix, the sparse inverse covariance matrix is estimated by SICE / iSICE. Therefore, the three-branch fusion model is responsible for obtaining multimodal complementary features, and the SICE / iSICE module further estimates the direct correlation structure after deconfounding. Together, they constitute a predictive model for the benefit of interventional treatment in patients with coronary artery occlusion.
[0054] 3.2.2 Develop an intelligent prediction model for the benefit of interventional treatment in patients with coronary artery occlusion based on multimodal cardiovascular imaging.
[0055] This section mainly focuses on the multimodal image feature extraction and feature fusion model in Scheme 1, applying the sparse covariance estimation algorithm to solve the multimodal image feature prediction task, and establishing an intelligent prediction model for the benefit of interventional treatment for patients with coronary artery CTO lesions.
[0056] Step 3: Construct an intelligent prediction model for the benefit of interventional treatment in patients with coronary artery occlusion based on multimodal cardiovascular imaging.
[0057] Interventional treatment of coronary artery CTO lesions is high-risk, difficult, and prone to complications. Even with successful recanalization, some CTO patients still experience no improvement in quality of life or prognosis. Accurate pre-operative prediction of whether patients will benefit from percutaneous coronary intervention (PCI) is of significant clinical importance. However, there is currently no effective solution for predicting the benefits of multimodal image fusion features for coronary CTO lesions. Therefore, this invention proposes a visual representation model based on the covariance matrix for intelligent prediction of the benefits of interventional treatment for patients with coronary CTO lesions.
[0058] Specifically, the visual representation model comprises two parts: a neural network-based Sparse Inverse Covariance Estimation (SICE) module and an Iterative Sparse Inverse Covariance Estimation (iSICE) module. The neural network-based SICE module defines SICE as a novel structured layer within the neural network. Its input is a set of local descriptors extracted from the convolutional feature map, and its output is a sparse inverse covariance matrix. The Iterative Sparse Inverse Covariance Estimation module, designed to ensure the end-to-end trainability of SICE, addresses the aforementioned matrix optimization problem in the forward and backward propagation steps. This module includes two key steps: estimation of the accuracy matrix and estimation of the sparse inverse covariance matrix.
[0059] (1) Sparse Inverse Covariance Estimation Module (SICE) based on Neural Networks.
[0060] Visual representations based on covariance matrices have demonstrated their effectiveness in image classification through pairwise correlations of different channels in convolutional feature maps. However, pairwise correlations become misleading once another channel is associated with the two channels of interest, introducing confounding effects. For this situation, partial correlation is the correct measure. It regresses the effects of other variables from two variables and then calculates the correlation of their residuals. Partial correlation can be conveniently obtained by calculating the inverse covariance matrix, also known as the precision matrix in the statistical community.
[0061] like Figure 4 As shown, taking a 3D case as an example, the green box indicates that for d>3, calculating the partial correlation requires inverting the covariance matrix. Assuming the sample size... Number of channels In the 3D case, x and y are projected onto a plane perpendicular to z. Then, ( and (Similar calculations can be made). The "residual" of the projection. and It can be calculated as shown in the figure, where, , ( (This can be calculated in a similar way). Through this operation, unlike ordinary covariance (e.g., the pairwise correlation between x and y corresponding to channels), the partial correlation between variables x and y removes the influence of the confounding variable z. Based on this, a visual representation for image classification based on the inverse covariance matrix is proposed. This invention studies the inverse covariance matrix from the perspective of image classification tasks and proposes the SICE module.
[0062] SICE aims to improve covariance estimation by leveraging prior knowledge. The benefits of this approach are twofold: first, it mitigates the instability or singularity of the covariance matrix from a small number of high-dimensional feature vector samples; second, it improves covariance estimation from a limited sample size. To incorporate prior knowledge, SICE switches from the covariance matrix to its inverse. In principle, the covariance matrix captures the obvious pairwise correlations, i.e., indirect correlations, between feature components. In contrast, the inverse of the covariance matrix can characterize the direct correlations (i.e., partial correlations) between two feature components by regressing the remaining features. Using the inverse covariance matrix not only helps interpret the fundamental relationship between two features but also allows for the convenient incorporation of sparse priors. Estimation of the inverse covariance matrix effectively eliminates the misleading effects of confounding factors, thus benefiting image classification tasks.
[0063] To better utilize SICE in image classification tasks, this invention incorporates it into neural networks. Specifically, it assumes a set of local descriptors of the neural network extracted from convolutional feature maps, and calculates the sample-based covariance matrix. .let This represents the corresponding sparse inverse covariance matrix. The off-diagonal terms capture the direct correlation between different descriptor components. They are zero if two components are independent after eliminating the effects of confounding variables. In the literature, by analyzing... By implementing SPD constraints and imposing sparse priors, the problem can be effectively solved by maximizing the penalized log-likelihood of the data. The problem of estimating the connectivity of sparse graphs is called the SICE problem. The optimal solution to this problem is called SICE. Therefore, mathematically, SICE can be defined as follows: ; in, It is based on the sample covariance matrix. , and Let the determinant, trace, and sum of the vectorized matrix be represented respectively. Norm.
[0064] In order to obtain reliable and accurate SICE, Item pair Structural sparsity was applied. The trade-off between sparsity and log-likelihood estimation is controlled. Formula The problem in this case is convex and can be solved using readily available software packages. However, due to... The penalty is that the target is non-smooth. The current software package cannot be used with neural network module layers for backpropagation training. However, it still suffers from slow speed; therefore, the iSICE method below is designed to improve trainability.
[0065] (2) Iterative sparse inverse covariance estimation module (iSICE) is used for end-to-end training of SICE.
[0066] set up It is a formula The objective function. By means of We optimize by taking the derivative, as follows: (twenty three); in, and Each contains The positive and negative parts. Equation (23) can be optimized by projected gradient descent, which supports backpropagation on GPUs and can leverage GPU parallel computing to improve speed.
[0067] In the optimization of formula (23), the present invention uses the above-mentioned multimodal fusion feature map. As input to the visual representation model. Remodeled into a local descriptor data matrix Where N is the number of spatial locations or local descriptors. This is the fused feature descriptor at the nth spatial location. Centralized processing is performed to obtain ,in, Further calculate the sample covariance matrix. and will As input to the SICE / iSICE iterative estimation.
[0068] In a specific embodiment, the key steps of the sparse inverse covariance estimation module based on neural networks are as follows: The first key step is the precision matrix. The estimate.
[0069] The Newton-Schulz iterative method is popular because it can quickly approximate the square root of a matrix on a GPU. In contrast, in estimation... At this time, the Newton-Schulz iterative method is used to quickly approximate the inverse matrix. This process requires the application of a convergence condition. Therefore, the present invention, through the... The trace is normalized, that is Then, the inverse square root matrix. The square of the trace is normalized to invert it, that is... .
[0070] The second key step is the sparse inverse covariance matrix. The estimate.
[0071] Given the result obtained in the previous step (To standardize the notation, let) This invention begins the iteration of iSICE by projecting gradient descent onto the gradient of SICE (i.e., formula (22)), as shown in formula (23), dividing S into its positive and negative parts: ; Then apply the PGD steps to them respectively.
[0072] In this case, by The simplification of the sparsity constraint imposed by the norm, i.e. The gradient can be assumed to be , The gradient can also be assumed to be Therefore, this invention first rewrites formula (23) into two parts: ; Then, this invention uses a PGD step to update. and : ; in, It is a function that reprojects the gradient of PGD onto the feasible region of each boundary constraint (one for non-negative values). , one for non-positive ).constant It is the learning rate, and Control the decay of the learning rate.
[0073] Finally, using and Composition of the current Estimated value: (27); in, Ensure matrix It is symmetrical (and is also an intermediate estimate of SICE). As input to the visual representation model for subsequent benefit prediction, Used to represent the direct correlation structure between multimodal fusion features after removing confounding effects.
[0074] Furthermore, the multimodal image feature fusion and the visual representation model based on the covariance matrix are constructed in a logical cascade sequence of "multimodal feature extraction—three-branch spatial fusion—sparse inverse covariance representation—benefit classification". Let the multimodal image of the i-th patient be... The features of each modality are obtained by the multimodal image feature extraction model. Where m=1,...,M. The three branches are... As common feature sources, different input organization forms are constructed respectively: the spatial decoupling feature fusion module constructs local input sequences within the same spatial unit and obtains... The spatial feature fusion module constructs a global input sequence covering all spatial units and all modalities and obtains... The spatial cross-attention feature fusion module constructs a cross-attention input consisting of a main modality query and auxiliary modality keys, and obtains... ; Calculated through feature aggregation layer Calculate the weights of the three branches and output the final multimodal fusion feature map. .Will Remodeled into a local descriptor data matrix Where N is the number of spatial locations or local descriptors in the fused feature map. For the fusion feature descriptor at the nth spatial location; Centralized processing is performed to obtain ,in And calculate the sample covariance matrix. Based on SICE / iSICE Sparse inverse covariance estimation yields ,in, Characterize the direct correlation between features of the fused image after removing confounding effects. Finally, take... ,pass The probability of benefit from interventional treatment is obtained based on a threshold. Output the profit prediction results.
[0075] Furthermore, the iSICE algorithm proposed in this invention starts with a dense precision matrix and applies the above steps iteratively N times. To facilitate adjustment of the learning rate, the algorithm first performs trace normalization on the input matrix, and then reverses the trace normalization after completion. Because As a symmetric matrix, this invention takes Vectorize the upper triangular part and the diagonal part to obtain and will Input a fully connected prediction layer.
[0076] Based on the estimated value The process of determining the prediction results is as follows: ,in, This represents the probability of benefit for the i-th patient after undergoing interventional treatment for coronary artery occlusion. and For prediction layer parameters, This is the Sigmoid function. If... ≥ Then output the prediction result. =1 indicates that the patient is expected to benefit from the interventional treatment; if < Then output the prediction result. =0 indicates that the patient's expected benefit is insufficient or non-existent. The preset classification threshold can be set to 0.5 by default, or it can be determined based on the sensitivity and specificity of the validation set.
[0077] When multiple levels of benefit are required, the aforementioned Sigmoid prediction layer is replaced with a Softmax prediction layer, i.e. and with Determine whether the i-th patient belongs to the high-benefit, moderate-benefit, or low-benefit category. Therefore, the estimated value... It serves not only as an intermediate result of the sparse inverse covariance matrix, but also as a visual representation input for the benefit prediction model, used to obtain explicit prediction results of interventional treatment benefits.
[0078] The iSICE algorithm proposed in this invention is derived from a dense precision matrix. Start by applying the above steps N times. To facilitate adjusting the learning rate... The algorithm first performs... Perform trace normalization, and then reverse the trace normalization after completion. Otherwise, Must be based on Scaling the value of the largest eigenvalue is somewhat impractical when running neural network modules end-to-end across multiple mini-batch runs. Because... It is a symmetric matrix. In this invention, only the upper triangular part (plus the diagonal part) is taken and classified through a fully connected layer.
[0079] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0080] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for predicting the benefit of intervention treatment for a patient with coronary occlusion lesions based on multi-modal images, characterized in that, include: Multimodal images are acquired, and multimodal features are extracted from the multimodal images using a multimodal image feature extraction model; A multimodal image feature fusion model is constructed to fuse the multimodal features; A model for predicting the benefit of interventional treatment for patients with coronary artery occlusion is constructed based on a visual representation model of the covariance matrix. The multimodal imaging data to be tested is acquired and input into the interventional treatment benefit prediction model for patients with coronary artery occlusion to obtain the benefit prediction results.
2. The method of claim 1, wherein the multi-modal images comprise: Coronary angiography, coronary CTA, and cardiac MRI images.
3. The method for predicting the benefit of interventional treatment for patients with coronary artery occlusion based on multimodal imaging according to claim 1, characterized in that, The multimodal image feature extraction model includes: Acquire the input image and extract an initial pixel-level representation from the input image; The initial pixel-level representation is grouped according to semantics to obtain a semantic-level representation; Semantic-level representations are processed using a feature enhancement module and a feature interaction module guided by a relation matrix. The initial pixel-level representation is enhanced by the processed semantic-level representation, resulting in a pixel-level representation with semantic relationships.
4. The method of claim 1, wherein the method further comprises: determining a benefit of the intervention therapy for the patient based on the predicted benefit of the intervention therapy for the patient. The multimodal image feature fusion model includes: a spatial decoupling feature fusion module, a spatial feature fusion module, and a spatial cross-attention feature fusion module.
5. The method of claim 1, wherein the method further comprises: determining a benefit of the intervention therapy for the patient based on the predicted benefit of the intervention therapy for the patient. The visual representation model based on the covariance matrix includes: a sparse inverse covariance estimation module based on a neural network and an iterative sparse inverse covariance estimation module. The sparse inverse covariance estimation module based on the neural network is designed as a structured layer of the neural network. The input is a set of local descriptors of the neural network extracted from the convolutional feature map, and the output is a sparse inverse covariance matrix. The iterative sparse inverse covariance estimation module is used for end-to-end training of the sparse inverse covariance estimation module, and solves the optimization problem of the sparse inverse covariance matrix in the forward and backward propagation steps.
6. The method of claim 5, wherein the method further comprises: The neural network-based sparse inverse covariance estimation module includes: Extract a set of local descriptors from the convolutional feature maps and calculate the sample-based covariance matrix. ,make This represents the corresponding sparse inverse covariance matrix; The off-diagonal terms are used to capture the direct correlation between different descriptor components, and they are zero if the two components are independent after eliminating the effects of confounding variables. Through the By applying SPD constraints and sparse priors, the problem is solved by maximizing the log-likelihood of the data with a penalty. The estimation problem guides the assessment of the connectivity of sparse graphs, and the optimal solution is: ; in, This represents the corresponding sparse inverse covariance matrix. It is based on the sample covariance matrix. , and Let the determinant, trace, and sum of the vectorized matrix be represented respectively. Norm; Item pair Applying structural sparsity, Used to control the trade-off between sparsity and log-likelihood estimation.
7. The method for predicting the benefit of interventional treatment in patients with coronary artery occlusion based on multimodal imaging according to claim 6, characterized in that, The iterative sparse inverse covariance estimation module includes: The iterative sparse inverse covariance estimation module performs end-to-end training on the sparse inverse covariance estimation module. Let be the objective function of the optimal solution, Optimization is performed by taking the derivative of the optimization: ; wherein, and respectively comprise the positive and negative parts are optimized by projected gradient descent.
8. The method for predicting the benefit of interventional treatment in patients with coronary artery occlusion based on multimodal imaging according to claim 7, characterized in that, The optimization via projected gradient descent includes: Reshape the multimodal fusion feature map into a local descriptor data matrix; The local descriptor data matrix is centered. The sample covariance matrix is calculated based on the centered local descriptor data matrix, and then input into the sparse inverse covariance estimation module for subsequent iterative optimization.
9. The method for predicting the benefit of interventional treatment in patients with coronary artery occlusion based on multimodal imaging according to claim 8, characterized in that, The iterative sparse inverse covariance estimation module includes: In estimation At this point, the Newton-Schulz iterative method is used to approximate the inverse matrix, with a convergence condition applied. Through the The trace is normalized, that is inverse square root matrix The square of the trace is normalized to invert it, that is... The accuracy matrix is obtained. ; Let The iteration of iSICE is started by projecting gradient descent on the gradient of SICE, splitting the sparse inverse covariance matrix S into its positive and negative parts: Then apply the PGD steps respectively; The optimal solution is rewritten in two parts: ; Then take a step PGD to update and : ; wherein, a function for re-projecting the gradient of the PGD to the feasible region of each boundary constraint, constant is a learning rate, , for controlling the decay of the learning rate; By means of and comprising the current estimated value: ; wherein for ensuring that the matrix is symmetric.
10. A system for predicting the benefit of interventional treatment for patients with coronary artery occlusion based on multimodal imaging, characterized in that, include: Multimodal images are acquired, and multimodal features are extracted from the multimodal images using a multimodal image feature extraction model; A multimodal image feature fusion model is constructed to fuse the multimodal features; A model for predicting the benefit of interventional treatment for patients with coronary artery occlusion is constructed based on a visual representation model of the covariance matrix. The multimodal imaging data to be tested is acquired and input into the interventional treatment benefit prediction model for patients with coronary artery occlusion to obtain the benefit prediction results.