Regression positioning network prediction method for stent model in coronary intervention surgery
Through positioning regression deep learning model architecture, the swin transformer network is used to improve the global context modeling ability, solve the objectivity problem of interventional stent model selection, realize the accurate prediction of interventional instruments and the recommendation of stent models, and improve the effectiveness of the surgery.
Patent Information
- Application Number
- CN202211006628.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-08-22
AI Technical Summary
The existing interventional stent model selection lacks objectivity and rigor, and it is difficult to provide fine and reliable quantitative results relying on the visual evaluation of surgeons. The stenosis detector only provides information on the location of stenosis lesions, which lacks practical application value.
The deep learning model architecture of positioning regression is adopted, including positioning network, feature extraction network and regression network, and the swin transformer network is used to improve the global context modeling ability, extract narrow area features through the positioning network, and learn the mapping relationship between image features and predicted values of interventional instruments through the regression network, and trained in combination with SIOU and weighted mean square error loss function.
It realizes accurate prediction of interventional devices, helps doctors quickly locate the location of stenosis lesions and recommends appropriate stent models, reduces the probability of postoperative cardiovascular adverse events and improves the effectiveness of the surgery.
Smart Images

Figure CN115457108B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical imaging and computer technology, and particularly relates to a regression positioning network prediction method for stent models in coronary heart disease intervention surgery. Background Art
[0002] Coronary artery intervention is an important treatment for coronary heart disease. During percutaneous coronary intervention (PCI), quantitative prediction of interventional stents under the guidance of X-ray coronary angiography is extremely important. Currently, the selection of interventional stent models relies entirely on the surgeon's visual assessment, which makes the selection of stent models lacking objectivity and rigor. The prediction of interventional stent models is based on further clinical research on quantitative morphological indicators of coronary artery stenosis. Despite long-term research on the precise quantification of coronary arteries, current model solutions often struggle to provide precise and reliable quantitative results. This is because relying solely on the model's self-attention mechanism makes it difficult to increase the model's attention to the small area target of stenosis lesions.
[0003] With the development of deep learning, lesion detection has been applied to medical image processing. Methods based on deep convolutional neural networks (CNNs) rely on data-driven features and have achieved excellent performance in detecting biomedical targets. Recently, Kun Pang et al. proposed a stenosis detection network that maximizes the use of temporal information by designing a sequence feature fusion module and a sequence consistency alignment module. The sequence feature fusion module fuses all candidate box features and enhances these features using temporal information. The sequence consistency alignment module optimizes the initial results using coronary artery displacement information and image features of adjacent images, thereby achieving end-to-end detection of coronary artery stenosis. The development of coronary artery stenosis detection research shows that simplifying the detection process and improving detection accuracy are the key directions of development. However, current stenosis detectors can only provide physicians with information on the location of stenotic lesions, providing limited guidance in clinical practice and lacking practical application value.
[0004] Many deep learning tasks have been developed for coronary X-ray angiography images, including vessel segmentation, 3D reconstruction, and lesion classification. Only a few works have focused on quantifying morphological stenosis metrics. Among these, only a few have focused on quantifying coronary artery stenosis. D. Zhang et al. proposed a multi-view hierarchical attention learning model (HEAL) that integrates spatiotemporal attention on coronary artery video data for contrastive learning to achieve direct quantification of coronary artery stenosis without the need for intermediate segmentation or reconstruction. W. Xue et al. proposed a semi-automatic method for multi-type cardiac index estimation. This network comprises a deep convolutional autoencoder and a multi-output convolutional neural network. The joint learning of these two networks effectively enhances the expressive power of image representations relative to cardiac index, resulting in accurate and reliable cardiac index estimates. Furthermore, other stenosis quantification methods have used 3D CT angiography and IVUS angiography images, achieving stenosis quantification through methods such as convolutional networks or XGBoost. However, the models used in these works are insufficient for global feature modeling. In this regard, our work on tree-structured stenosis feature extraction is related to the Transformer model. Transformer models are currently very popular in computer vision tasks. Their excellent ability to model long-range dependencies offsets the inherent limitations of convolutional neural networks and demonstrates excellent performance in many tasks. For example, Chen et al. combined Transformers with U-Net, significantly improving experimental accuracy in medical image segmentation. Transformer models have been successful in many medical applications. However, Transformer models have not yet been applied to the quantification of narrow morphological metrics.
[0005] In order to achieve accurate prediction of interventional devices, it is necessary to study a regression positioning network prediction method for stent models in coronary intervention surgery. Summary of the Invention
[0006] The purpose of the present invention is to provide a positioning regression deep learning model architecture, which consists of a positioning network, a feature extraction network and a regression network to achieve end-to-end quantitative prediction of interventional devices. The positioning network is used to extract the stenosis area in advance to avoid the influence of a large amount of redundant background on the quantitative prediction of interventional devices, which solves the problem of detecting stenotic blood vessels in small areas to a certain extent. At the same time, the swin transformer network is used as the backbone to enhance the global context modeling capability of the model, so as to better extract the features of the stenotic blood vessels after positioning. In order to obtain quantitative prediction indicators of interventional devices, we use a fully connected layer as a regression network to learn the mapping relationship between image features and interventional device prediction values.
[0007] To achieve the above object, the technical solution of the present invention is as follows:
[0008] A regression positioning network prediction method for stent model in coronary intervention surgery includes the following steps:
[0009] Step 1: Obtain the data required for the experiment; doctors mark the location and size of the stenotic lesion area in the image, and randomly divide the marked data into a training set, a validation set, and a test set;
[0010] Step 2: Establish a positioning regression model. The positioning regression model consists of three parts: positioning network, feature extraction network, and regression network. The positioning network in the model is trained using the training set and validation set using the deep learning method. The SIOU loss function used is as follows:
[0011]
[0012] Where Δ is the distance loss Ω is the shape loss, and IOU represents the intersection-over-union ratio of the target box in the label to the target box area predicted by the model;
[0013] Step 3: When the performance of the model training set and the validation set is close and meets the requirements, the positioning network training is completed. After the training, the performance of the positioning network is verified on the test set. If it meets the requirements, the next step of training can be carried out. If not, the network hyperparameters are adjusted and the positioning network is retrained.
[0014] Step 4: Freeze the trained positioning network parameters, build a feature extraction network and a regression network, and combine them with the positioning network to jointly train the overall model. The trained model can directly predict the stenosis diameter, normal diameter, and stenosis length of the coronary angiography image to recommend the appropriate interventional stent size. The regression network uses a weighted mean square error loss L mse Perform supervised training, where y i ∈{y1,y2...y M} represents the tag value, represents the predicted value of the model;
[0015]
[0016] Step 5: Use samples from the test set to test the model, compare the model prediction results with the labels made by the doctors, verify the effectiveness of the overall positioning regression model, determine the final model and use it for clinical testing to help doctors select interventional surgery stent models.
[0017] A further improvement of the technical solution of the present invention is that: complete DICOM data is collected from the hospital, the average age of the subjects is: 47-71 years old, the image data size is 512×512, the pixel pitch is 0.258mm / pixel, and the views of the collected data include the foot view CAU, the cranial view CRA, the left anterior oblique view LAO, the left anterior oblique foot view LAO CAU, the left anterior oblique cranial view LAOCRA, the right anterior oblique view RAO, the right anterior oblique foot view RAO CAU and the right anterior oblique cranial view RAO CRA; among them, LAO CRA and RAOCRA mainly observe the middle and distal segments of the left anterior descending artery; LAO CAU and RAO CAU mainly observe the proximal segments of the left anterior descending artery and the left circumflex artery; LAO mainly observes the proximal segment of the right coronary artery, and LAO CRA mainly observes the distal segment of the right coronary artery.
[0018] A further improvement of the technical solution of the present invention is that the data sets are grouped according to the needs of different body position data in clinical application scenarios, and the doctor marks the position and size of the stenotic lesion area in the image. The size information includes the blood vessel diameter at the stenosis, the diameter of the normal blood vessel at the stenosis, and the length of the stenosis.
[0019] A further improvement of the technical solution of the present invention is that the specific structure of the positioning network in step 2 is divided into three parts: convlution backbone, region proposal and ROI pooling, which are respectively responsible for extracting image features, generating proposed target boxes and adjusting target boxes. The training set and validation set are used for training on the constructed positioning network, where the training set is used to adjust the network parameters, and the validation set is used for performance testing after each network parameter adjustment.
[0020] A further improvement of the technical solution of the present invention is that the feature extraction network in step 4 adopts the swintransformer module, and local windows are added to the sub-attention mechanism inside the module to obtain deep image prediction capabilities and high-resolution image prediction capabilities.
[0021] A further improvement of the technical solution of the present invention is that the regression network in step 4 consists of a global average pooling layer and three fully connected layers, each fully connected layer includes a BN layer and a relu nonlinear activation function. Since the image features are processed by the leading network, the task complexity of the regression network is reduced. The mean square error is used as the loss function of the network in the regression network, and the weight term λi∈{λ1,λ2,λ3} is added to the loss function. The loss function formula is as follows:
[0022]
[0023] Where N represents the number of prediction tasks, f(x i) represents the predicted value of the model's i-th task, Y i Represents the label value of the i-th task.
[0024] A further improvement of the technical solution of the present invention is that: during the training phase, the positioning network parameters trained in step 3 are frozen, the overall model is jointly trained, and training strategies such as random cropping, rotation, and mosaic enhancement are used to expand the data set to improve the accuracy and robustness of the model.
[0025] Due to the adoption of the above technical solution, the technical effects achieved by the present invention are as follows:
[0026] This invention utilizes a multi-task deep learning approach to achieve end-to-end prediction of interventional stent size on coronary angiography images. This allows physicians to quickly locate coronary artery stenosis and predict the size of the stent required for interventional surgery. This helps physicians select the most appropriate stent size, reducing the risk of postoperative adverse cardiovascular events and improving surgical effectiveness. Compared to directly predicting interventional stent size, this model significantly improves prediction accuracy and robustness, thus possessing strong clinical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is the overall framework diagram of the present invention;
[0028] Figure 2 It is the swin module in the regression network;
[0029] Figure 3 is a flow chart of the present invention;
[0030] Figure 4 To locate the Convlution Backbone structure of the network;
[0031] Figure 5 To locate the network region proposal structure;
[0032] Figure 6 It is the ROI pooling structure of the positioning network. DETAILED DESCRIPTION
[0033] The present invention will be further described below with reference to specific embodiments and accompanying drawings.
[0034] The specific regression positioning network prediction method for stent model in coronary intervention surgery includes the following steps:
[0035] Step 1: Obtain the data required for the experiment;
[0036] Complete DICOM data was collected from a hospital. The average age of the subjects ranged from 47 to 71 years. The image size was 512×512, with a pixel pitch of 0.258 mm / pixel. The acquired data included the foot view (CAU), cranial view (CRA), left anterior oblique view (LAO), left anterior oblique foot view (LAO CAU), left anterior oblique cranial view (LAO CRA), right anterior oblique view (RAO), right anterior oblique foot view (RAO CAU), and right anterior oblique cranial view (RAO CRA). The LAO CRA and RAO CRA primarily observed the middle and distal segments of the left anterior descending artery. The LAO CAU and RAO CAU primarily observed the proximal left anterior descending artery and left circumflex artery, while the LAO primarily observed the proximal right coronary artery. The LAO CRA primarily observed the distal segment of the right coronary artery. The dataset was grouped based on the requirements for data from different body positions in clinical application scenarios. Physicians annotated the location and dimensions of the stenotic lesions in the images. Dimensional information included the vessel diameter at the stenosis site, the diameter of the normal vessel at the stenosis site, and the length of the stenosis site. The labeled data is divided into training set, validation set and test set using random sampling method to facilitate the performance verification of subsequent models.
[0037] Step 2: Build and train the positioning network
[0038] In order to meet the needs of stent model prediction in PCI, a positioning regression model was designed, and its framework is as follows: Figure 1 As shown, it consists of three parts: positioning network, feature extraction network, and regression network. The prediction problem is simply defined as formula (1):
[0039]
[0040]
[0041] Where X and Y are the training data set and the true label respectively. The model goal is to learn the mapping relationship from the feature representation r(X) of the lesion to the size Y of the interventional stent. However, since the coronary angiography image contains a large amount of organ artifact contour information, it is not conducive to the model to directly locate and learn the stenotic lesion area. Therefore, it is necessary to strike a balance between the ability to extract global features and the ability to focus on local features in model design. Therefore, the Faster RCNN detection network is introduced to convert the mapping process of formula (1) into formula (2), where det is the detection network that converts the 512×512 image to 224×224, and J is the objective function. Since stenotic lesions usually only occupy tens to hundreds of pixels, it is extremely difficult to directly achieve accurate detection of small targets for 512×512 images. Therefore, the FPN structure is introduced in the convlution backbone part of the positioning network, such as Figure 4As shown in the figure, the fusion of high-resolution shallow layers and semantically rich deep layers greatly improves the detection model's ability to detect small targets.
[0042] The overall model adopts the form of step-by-step training, firstly building and training the positioning network. Figure 1 As shown in the original text detection network section of the figure, the specific structure is divided into three parts: the convolution backbone, region proposal generation, and region of interest pooling (ROI pooling). These are responsible for extracting image features, generating target bounding boxes, and adjusting target bounding boxes, respectively. The constructed localization network is trained using a training set and a validation set. Only the training set is used to adjust network parameters, while the validation set is used to test performance after each network parameter adjustment. The SIOU loss function used in training is as follows:
[0043]
[0044] Where Δ is the distance loss, Ω is the shape loss, and IOU represents the intersection-over-union ratio of the target box in the label to the target box predicted by the model. Compared to other loss functions, SIOU has a stronger ability to locate the target and can accelerate model convergence.
[0045] Step 3: Locate network performance test
[0046] When the model's performance on the training set and validation set is close and meets the requirements, the positioning network training is considered complete. After training, the positioning network's performance is tested on the test set. If it meets the requirements, the next step of training can be carried out. If not, the network hyperparameters should be adjusted and the positioning network should be retrained.
[0047] Step 4: Overall model construction and training
[0048] This step requires building a feature extraction network and a regression network and splicing them with the positioning network for joint training. It is worth noting that the present invention embeds a random cropping module ( Figure 1 In ImageAugmentation, the image is randomly cropped based on the bounding boxes [x, y, h, w] of the target bounding box obtained by the localization network. This random cropping of the original image centered around the lesion bounding box obtained by the localization network reduces the model's reliance on data volume and improves its translation invariance. Furthermore, it is crucial to design the entire architecture to eliminate the periodic accumulation of errors caused by the biases of the detection network.
[0049] In order to obtain accurate prediction results, the Swin transformer network is embedded into the entire model framework as a feature extraction network. The Swin transformer module is the key architecture of the Swin transformer network. Its architecture is as follows Figure 2 As shown in the figure, see Z.Liu et al., "Swin Transformer: Hierarchical Vision Transformerusing Shifted Windows," 2021IEEE / CVF International Conference on Computer Vision (ICCV), 2021, pp.9992-10002, doi:10.1109 / ICCV48922.2021.00986. This structure obtains powerful image neighborhood connection and local spatial feature extraction capabilities by retaining the prior knowledge of the convolutional neural network. Unlike the traditional transformer block, the swin transformer module reduces the computational complexity of image vision tasks while retaining strong global context modeling capabilities by adding local windows (local windows) to the self-attention calculation window, thereby obtaining deep image prediction capabilities and high-resolution image prediction capabilities.
[0050] The regression network consists of a global average pooling layer and three fully connected layers. Each fully connected layer includes a batch normalization layer and a relu nonlinear activation function. Because processing image features through the leading network reduces the task complexity of the regression network, mean squared error is used as the loss function for the network in the regression network. Experiments have found that the magnitudes of three important indicators in the model prediction results vary. Therefore, a weight term λi∈{λ1,λ2,λ3} is added to the loss function to balance the loss function preferences for each prediction target. The loss function formula is as follows:
[0051]
[0052] Where N represents the number of prediction tasks, f(x i ) represents the predicted value of the model's i-th task, Y i Represents the label value of the i-th task.
[0053] During the training phase, the localization network parameters trained in step 3 are frozen, and the entire model is jointly trained. Training strategies such as random cropping, rotation, and mosaic enhancement are used to expand the dataset and improve the model's accuracy and robustness. The trained model can directly predict the stenosis diameter, normal diameter, and stenosis length from coronary angiography images to recommend appropriate interventional stent sizes.
[0054] Step 5: Test the performance of the overall model
[0055] Similar to step 3, the effectiveness of the overall positioning regression model is verified, and the final model is determined and used for clinical testing to help doctors select interventional stent models.
[0056] Example:
[0057] Using the above method, 500 coronary angiography images were used as a training set, 50 as a validation set, and 40 as a test set to train and test the positioning regression model.
[0058] In the experiment, the applicant compared four commonly used neural network regression models, namely models A, B, C, and D. Model A uses resnet50 as the regression model backbone, for details, see K.He, X.Zhang, S.Ren and J.Sun, "Deep Residual Learning for Image Recognition," 2016IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp.770-778, doi:10.1109 / CVPR.2016.90. Model B uses densenet121 as the regression model backbone, for details, see G.Huang, Z.Liu, L.VanDerMaaten and KQWeinberger, "Densely Connected Convolutional Networks," 2017IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp.2261-2269, doi:10.1109 / CVPR.2017.243. It is worth noting that in order to improve the performance of Model A and Model B, this application adds SE blocks to ResNet50 and DenseNet121. Model C uses the visual attention network as the backbone of the feature extraction network, see Guo, MH, Lu, CZ, Liu, ZN, Cheng, MM, Hu, SM: Visual Attention Network.arXiv e-prints (2022) arXiv:2202.09741 for details, and embeds it into the position regression network architecture. In addition, to verify the effectiveness of our proposed method, we also compare it with the method of D. Zhang et al. (Method D), see Zhang, D., Yang, G., Zhao, S., Zhang, Y., Ghista, D., Zhang, H., Li, S.: Direct quantification of coronary artery stenosis through hierarchical attentive multi-view learning. IEEE Transactions on Medical Imaging 39 (2020) 4322–4334.
[0059] After testing, the positioning regression network architecture proposed in this application achieved accurate interventional device prediction results, where the mean and variance of the stenosis diameter, normal blood vessel diameter at the stenosis, and stenosis length were 0.1425±0.1389, 0.2357±0.1117, and 3.1586±1.4067, respectively, and the average Mae value was 1.1789±0.5524mm. Among them, the average MAE of the proposed coarse positioning regression network structure was the smallest, and its mean and variance were 5 times and 2 times the image pixel spacing (0.258mm / pixel), respectively. At the same time, the method of the present invention also achieved the best PCC index. The effectiveness of the positioning regression network is shown in Table 1:
[0060] Table 1 Comparison of results of various methods
[0061]
[0062]
[0063] Table 1 shows that the model method of this application aims to solve clinical problems, solves the problem of finely predicting the size of interventional devices, and can provide clinicians with a reference for the model of interventional surgical stents.
[0064] In summary, this invention utilizes a multi-task deep learning approach to achieve end-to-end prediction of interventional stent models using coronary angiography images. This allows physicians to quickly locate coronary artery stenosis and predict the appropriate stent model for interventional surgery. This helps physicians select the most appropriate stent model, reducing the risk of postoperative adverse cardiovascular events and improving surgical effectiveness.
Claims
1. A regression positioning network prediction method for stent model in coronary intervention surgery, characterized by The steps include: Step 1: Obtain the data required for the experiment; doctors mark the location and size of the stenotic lesion area in the image, and randomly divide the marked data into a training set, a validation set, and a test set; Step 2: Establish a positioning regression model. The positioning regression model consists of three parts: positioning network, feature extraction network, and regression network. The positioning network in the model is trained using the training set and validation set using the deep learning method. The SIOU loss function used is as follows: Where Δ is the distance loss, Ω is the shape loss, and IOU represents the intersection-over-union ratio of the target box in the label to the target box predicted by the model. The specific structure of the localization network in step 2 is divided into three parts: convolution backbone, region proposal, and ROIpooling. These are responsible for extracting image features, generating proposed target boxes, and adjusting target boxes, respectively. The constructed localization network is trained using a training set and a validation set. The training set is used to adjust the network parameters, and the validation set is used to test the performance after each network parameter adjustment. Step 3: When the performance of the model training set and the validation set meets the requirements, the positioning network training is completed. After the training, the performance of the positioning network is verified on the test set. If it meets the requirements, the next step of training can be carried out. If not, the network hyperparameters are adjusted and the positioning network is retrained. Step 4: Freeze the trained positioning network parameters, build a feature extraction network and a regression network, and combine them with the positioning network to jointly train the overall model. The trained model is used to predict the three indicators of stenosis diameter, normal diameter, and stenosis length of coronary angiography images to recommend the appropriate interventional stent size. The regression network uses a weighted mean square error loss L mse Perform supervised training, where λ i ∈{λ1,λ2...λ M } represents the weighted value, y i ∈{y1,y2...y M } represents the tag value, represents the predicted value of the model; In step 4, the feature extraction network uses the Swin transformer module. The module adds local windows to the sub-attention mechanism to obtain deep image prediction capabilities and high-resolution image prediction capabilities. In step 4, the regression network consists of a global average pooling layer and three fully connected layers. Each fully connected layer includes a BN layer and a relu nonlinear activation function. The mean square error is used as the loss function of the network in the regression network. Step 5: Use samples from the test set to test the model, compare the model prediction results with the labels made by the doctors, verify the effectiveness of the overall positioning regression model, determine the final model and use it for clinical testing to help doctors select interventional surgery stent models.
2. The method for predicting stent model using a regression positioning network in coronary intervention according to claim 1, characterized in that: Complete DICOM data were collected from the hospital. The average age of the subjects was 47-71 years old. The image data size was 512×512, with a pixel pitch of 0.258 mm / pixel. The views of the collected data included the foot view (CAU), cranial view (CRA), left anterior oblique view (LAO), left anterior oblique foot view (LAO CAU), left anterior oblique cranial view (LAO CRA), right anterior oblique view (RAO), right anterior oblique foot view (RAO CAU), and right anterior oblique cranial view (RAO CRA). Among them, LAO CRA and RAO CRA mainly observed the middle and distal segments of the left anterior descending artery; LAOCAU and RAO CAU mainly observed the proximal segments of the left anterior descending artery and left circumflex artery; LAO mainly observed the proximal segment of the right coronary artery; and LAO CRA mainly observed the distal segment of the right coronary artery.
3. The method for predicting stent model using a regression positioning network in coronary intervention according to claim 1, characterized in that: The data sets are grouped based on the demand for different body position data in clinical application scenarios, and doctors mark the location and size of the stenotic lesion area in the image. The size information includes the diameter of the blood vessel at the stenosis, the diameter of the normal blood vessel at the stenosis, and the length of the stenosis.
4. The method for predicting stent model using a regression positioning network in coronary intervention according to claim 1, characterized in that: Freeze the localization network parameters trained in step 3, jointly train the entire model, and use random cropping, rotation, and mosaic enhancement training strategies to expand the dataset and improve the accuracy and robustness of the model.
Citation Information
Patent Citations
Device, system, and medium for guiding implantation of stent into blood vessel
CN110664524A
Method for predicting narrow blood vessel size and instrument size based on Swin-T
CN114052762A