Attention mechanism under feature fusion of multi-period brain cta collateral circulation scoring method
By employing a feature fusion method based on an attention mechanism, collateral circulation scoring is performed using multi-period CTA data. This addresses the issues of insufficient scoring accuracy and low efficiency of manual feature extraction in traditional methods, achieving more efficient collateral circulation scoring and prognostic assessment, and assisting in clinical diagnosis.
Patent Information
- Application Number
- CN202211334054.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-10-28
AI Technical Summary
Existing technologies struggle to effectively utilize multi-period CTA data for collateral circulation scoring, resulting in insufficient scoring accuracy. Furthermore, traditional methods rely on manual feature extraction, which is inefficient and cannot adapt to changes in collateral circulation across different periods.
A multi-stage brain CTA collateral circulation scoring method based on attention mechanism feature fusion is adopted. By designing four independent single-branch networks and a global attention mechanism, sparse features are compressed, efficient and abundant vascular features are extracted, and multi-stage feature fusion is performed. Finally, the feature is input into a linear classifier for classification.
It improves the accuracy and prognostic efficiency of multi-stage CTA collateral circulation scoring, reduces the limitations of manual feature extraction, and provides richer semantic and feature information to assist clinicians in diagnosis and treatment decisions.
Smart Images

Figure CN115601343B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer-aided medical treatment, in particular to a multi-period brain CTA collateral circulation scoring method based on feature fusion under attention mechanism. BACKGROUND
[0002] The establishment degree of cerebral collateral circulation is closely related to the prognosis of patients with acute ischemic stroke. However, due to the complex structure of cerebral vessels, the diversity of scoring standards, and the poor consistency of doctors in judging the results, the evaluation of collateral circulation requires high professional experience of physicians. Therefore, it is of great significance to use computer-aided diagnosis technology to evaluate the establishment of collateral circulation in patients with ischemic stroke in clinical practice.
[0003] At present, the methods for computer-aided scoring of cerebral collateral circulation mainly include the following three types.
[0004] The first type is the traditional method based on matrix statistical analysis. This type of method models the blood vessels and simulates the evaluation method on the existing clinical protocol. Mumu proposed a matrix completion algorithm based on Frank-Wolfe optimization, which calculates the ratio of unfilled blood vessels to estimated complete blood vessels to obtain an automatic collateral circulation evaluation method in ischemic stroke. In 2017, Su et al. proposed a graph-based registration algorithm, which measures the length of the extracted blood vessel structure after blood vessel segmentation, and also measures the volume ratio of the occluded side and the opposite side of the blood vessel. Finally, a collateral classification model is established. However, this method requires manual registration of the brain map, which is relatively complex and has low accuracy.
[0005] The second type is based on traditional machine learning methods. In the process of computer-aided medical image diagnosis, radiomics extracts a large number of quantitative image features from massive medical image data with the help of computers, and then uses machine learning or statistical learning methods to quantitatively analyze a large number of radiomics features, and finally uses them for disease classification and prognosis prediction. Radiomics provides a quantitative diagnostic model for clinical practice, improving the accuracy of clinical diagnosis. In the past, radiologists could only make semi-quantitative descriptions of lesions, such as density uniformity and shape regularity. These pathological descriptions mainly rely on the experience of clinicians and are not reliable. Therefore, it is often necessary to repeat the expert review to determine the qualitative description of the lesion. Radiomics relies on computer-aided methods to quantitatively analyze lesion features and use a large number of quantitative values to describe lesion information. Radiomics can obtain feature information that traditional clinical analysis and traditional imaging analysis cannot obtain. In 2017, Xiao et al. first used 4D CTA to score collateral circulation by using low-rank decomposition, principal component analysis, and support vector machine methods. Due to the rapid development of radiomics, medical images contain many high-dimensional features that cannot be observed by the naked eye. In 2020, Aktar used non-contrast computed tomography (NCCT) radiomics features and machine learning to automatically assess collateral circulation. In the method combining radiomics and machine learning, statistical feature information focuses on manually extracting features such as blood vessel texture and blood perfusion. However, manual feature extraction has limitations. Moreover, the location of ischemia in each patient is not fixed because of the uniqueness of the vascular structure system. There is no obvious boundary line for collateral circulation in CTA data, and it is difficult to draw an ROI region on the CTA image. Therefore, radiomics is not suitable for CTA data.
[0006] The third type is a method based on a deep learning framework. The high-dimensional feature representation capability obtained based on a deep convolutional neural network is stronger than the low-dimensional feature obtained based on traditional manual delineation. The application of deep learning technology in medical images is in a development boom, and the convolutional neural network in particular has made a great contribution to medical image processing. In order to obtain more deep data in medical pictures and improve the accuracy of judgment, researchers often choose deep learning technology for research. The network structure of deep learning and medical image processing are systematically arranged, and the application in medical diagnosis is quite widespread. The algorithm of artificial intelligence has produced excellent results in medical image tasks. Ryan first proposed an artificial intelligence algorithm for multi-view CTA scoring in December 2021, which takes the CTA of the transverse plane, the sagittal plane and the coronal plane together as input, and identifies the collateral filling condition of the patient from the perspective of two classification and three classification, which achieves good accuracy. However, this study can only solve the single-period CTA collateral circulation scoring. Menon showed in the study that the use of multi-period CTA to predict the prognosis effect is better than the use of single-period CTA data and perfusion CT data. Multi-period CTA data contain more rich collateral vessel information, but are limited by the existing network framework and cannot be adapted to multi-period data. There is no progress in the research on multi-period CTA data collateral circulation rating based on deep learning. SUMMARY
[0007] The purpose of the present application is to solve the problems existing in the prior art, and to provide a multi-period brain CTA collateral circulation scoring method based on feature fusion under an attention mechanism, which can better adapt to the range of collateral circulation at different periods and simulate the process of continuously observing the CTA of each period of the patient by the physician in the judgment process, and finally realize the classification of multi-period brain CTA collateral circulation, effectively improve the prognosis judgment efficiency, avoid the limitations of manual extraction of image features in traditional methods, and play an auxiliary role in the diagnosis of doctors in clinical practice, and have a guiding role in the clinical treatment decision of collateral circulation stroke.
[0008] In order to achieve the above-mentioned purpose, the following technical solutions are implemented:
[0009] The multi-period brain CTA collateral circulation scoring method based on feature fusion under the attention mechanism comprises the following steps:
[0010] S1, convert the original CTA-dicom sequence into a pseudo-RGB image, perform data preprocessing by maximum density projection, and use translation, rotation and displacement methods to expand the original data set for offline data enhancement;
[0011] S2, four independent single-branch networks are designed for data features of multiple time periods, and a feature densification module based on attention mechanism is used to compress redundant information in sparse features and extract efficient rich vascular features;
[0012] S3, the designed single-branch multi-feature fusion and multi-period multi-feature fusion module is used to fuse multi-period CTA images and enrich data features;
[0013] S4, a global attention mechanism module is added to the fusion network to extract global features;
[0014] S5, the filtered fusion features are input into a linear classifier to classify the patient's collateral circulation.
[0015] Preferably, the step S1 comprises:
[0016] Step S11, set the HU threshold, when processing the head CTA image, adjust the window width of CTA to 350 and the window level to 70, and the gray value HU value of the contrast agent vessel is 500;
[0017] Step S12, the original data is normalized to [0, 255], and the original data is converted from DICOM to PNG image;
[0018] Step S13, project the CTA sequence images of the same period along the axial direction, and convert the CTA sequence images into CTA-MIP images;
[0019] Step S14, the CTA-MIP images pass through the feature reuse module, the original two-dimensional pixel matrix is broadcasted along the new space axis, and the data samples are enhanced, and the specific steps are batch processing of the image matrix, including rotation, flipping and displacement.
[0020] Preferably, the four independent single-branch network structures in the step S2 are the same, but do not share parameters, and each of the four independent single-branch networks is composed of three 3×3 convolution, four basic feature extraction blocks, four local attention mechanism feature extraction modules and a single-period feature up-sampling module.
[0021] Preferably, the local attention residual module is arranged in each of the four independent single-branch networks, and the local attention residual module is composed of a residual module, a grouped convolution and a layer normalization.
[0022] Preferably, in the step S3,
[0023] Feature map F mn First, the 3×3 transpose convolution layer is up-sampled, and then expanded to a feature map F′ with a length and width of 128 by a bilinear interpolation algorithm mn Feature map F mnThe feature map extracted by the nth local attention mechanism feature extraction module of the mth independent single-branch network;
[0024] All independent feature maps F' in the four independent single-branch networks mn are fused, which is referred to as single-period feature fusion, including the steps of:
[0025]
[0026] V m m1, m2 m3 m4 )
[0027] Wherein, F' is a real number mn , c×h×w , represents a transposed convolution operation, the convolution kernel size is 3, the step is 2, Concate represents a concatenation operation, V m is a feature value set after concatenation on a single branch;
[0028] The image features of the four periods are finally fused by pixel-level addition, and the formula is as follows:
[0029] V all = Add (V1, V2, V3, V4)
[0030] Wherein, Add represents an element addition operation, V all is a feature value set after concatenation of the features of the four independent single-branch networks, V1, V2, V3, and V4 represent the feature sets on the 1st-4th branches respectively;
[0031] The network after the four-period fusion module is composed of two 3x3 convolution, global attention mechanism feature extraction blocks, and finally outputs the prediction result under the multi-period deep feature fusion.
[0032] Preferably, in step S4, the fused feature map M k is a real number c×h×w respectively through a global pooling layer and an average pooling layer, the pooling layer is compressed along the channel axis, and then the channel dimension is spliced to obtain a feature map;
[0033] The above process is defined as:
[0034]
[0035] G max = MAX d (M k )
[0036] wherein G avg represents a two-dimensional matrix after the average pooling layer, i represents a spatial index, all index positions in the space s are selected in turn, G max represents a two-dimensional matrix after the maximum spatial pooling layer, MAX d represents a spatial maximum pooling operation;
[0037] G avg and the matrix G max The feature matrix set with doubled channel number is obtained through the splicing operation, then a 3*3 standard convolution layer is connected, the features are aggregated through the convolution layer, and finally the weight matrix Z of each channel is obtained according to the importance degree of each channel Si g moid activation function, and the process is defined as:
[0038] Z = sigma (f 3×3 [G avg ; G max ])
[0039] wherein Z e R 1×h×w , f 3×3 represents a convolution operation with a convolution kernel size of 3 and a step of 1, and the activation function sigmoid is denoted by a symbol sigma.
[0040] Preferably, in the step S5, the weight matrix Z of the channel is weighted with each channel of the feature map M k to obtain a global attention feature map G(fe k ), and the process is defined as follows:
[0041] G(fe k ) = M k *Z
[0042] After the global feature extraction module, the feature map will pass through the average pooling layer and the fully connected layer, and finally the class with the maximum probability value output by SoftMax is taken as the final scoring result.
[0043] Compared with the prior art, the application has the beneficial effects that:
[0044] The application proposes a classification algorithm of a multi-branch feature fusion multiple attention mechanism network for automatically scoring collateral circulation, local and global attention mechanisms are added to residual blocks in deep and shallow layers of the network, and the network will adaptively learn collateral circulation features according to different degree collateral circulation mapping and attention mapping weights; secondly, the network fuses attention extraction features at different periods, so that the differences of collateral features at different periods are fused and paid attention to, compared with single-period collateral features, the fused features are more rich; in summary, scoring on multi-period CTA images is more comprehensive than single-period CTA images, has higher accuracy, and has more rich semantic information and feature information; it has great significance for computer-aided cerebral collateral circulation diagnosis technology research. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 is an algorithm framework of the application;
[0046] Figure 2 is a CTA data preprocessing flowchart;
[0047] Figure 3 is a local attention residual module detail;
[0048] Figure 4 is a feature fusion module, a) local feature fusion module detail b) global feature fusion module detail;
[0049] Figure 5 is a global feature attention module;
[0050] Figure 6 is a four-period CTA-MIP data set profile;
[0051] Figure 7 is a ROC curve of non-good collateral score. DETAILED DESCRIPTION
[0052] The application will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the application and not to limit the scope of the application. In addition, it should be understood that those skilled in the art can make various modifications or modifications to the application after reading the content taught by the application, and these equivalent forms also fall within the scope defined by the application.
[0053] Example 1: As shown in the attached Figure 1 , the application described is a multi-period brain CTA collateral circulation scoring method under the attention mechanism feature fusion, including the following steps:
[0054] S1, convert the original CTA-dicom sequence into a pseudo-RGB image, perform data preprocessing by maximum density projection, and use translation, rotation and displacement methods to expand the original data set, perform offline data enhancement, the purpose is to enrich the sample size.
[0055] S2, four independent single-branch networks are designed for multi-period data features, and a feature dense module based on attention mechanism is used to compress redundant information in sparse features and extract efficient rich vascular features.
[0056] S3, by designing a single-branch multi-feature fusion and multi-period multi-feature fusion module, multi-period CTA images are fused to enrich data features.
[0057] S4, a global attention mechanism module is added, which is similar to adding a local feature extraction attention mechanism in the pre-fusion network and adding a global feature extraction attention mechanism in the post-fusion network. The mixed attention mechanism module helps to strengthen global feature extraction.
[0058] S5, the filtered fusion features are input into a linear classifier to classify the patient's collateral circulation.
[0059] Preferably, in step 1, the storage format of the original CTA medical image is DICOM (digital imaging and communications in medicine). Since the data format is inconsistent with the input format required by the deep learning network, data preprocessing is first needed. The preprocessing steps of the original data CTA image are as follows:
[0060] Step S11, set the HU (Hounsfield unit) threshold, when processing the head CTA image, adjust the CTA window width to 350 and the window level to 70, and the gray value HU value of the contrast agent vessel is close to 500; Step S12, the original data is normalized to [0, 255], and the original data is converted from DICOM to PNG image; Step S13, the CTA sequence images of the same period are projected along the axial direction, and the CTA sequence image is converted into CTA-MIP image; Step S14, the CTA-MIP image passes through the feature reuse module, the original two-dimensional pixel matrix is broadcast along the new space axis, and the data samples are enhanced. The specific steps are batch processing of the image matrix, including rotation, flipping and displacement, etc.
[0061] Through the above data preprocessing, the main advantages are as follows: (1) using the maximum density projection algorithm, converting the sequence image into a single sample, enhancing the collateral vessel features, and more intuitively observing the vascular features, while avoiding missing small blood vessels, and the required storage space is smaller than the original sequence. (2) The enhancement method simulates the sample of the real scene, expands the data sample, and avoids network overfitting without spending time on manual labeling.
[0062] The above preprocessing procedure for the CTA shadow data packet is as follows: Figure 2As shown, first, the experimental data DICOM format files of the four period head CTA are converted into pixel matrix.
[0063] The CT value of the CTA image data is contained in the DICOM file, the CT value is a kind of measuring unit for determining the density of human tissue or organ, also called Hounsfield unit (HU), and the conversion relationship between pixel value and CT value is represented by formula (1):
[0064] HU=PV×RS+RI (1)
[0065] Wherein PV represents pixel value (Pixel Value), RS represents rescale slope value (Rescale Slope), and RI represents rescale intercept value (Rescale Intercept), and RS and RI information values are included in the DICOM file.In this experiment, in order to observe the blood vessel information, RS=1 and RI=-1024 are taken.
[0066] In medical images, different ranges of CT values can reflect different tissues or organs, and the present application limits HU to [10, 500], i.e.the lower limit of CT value is CT min =10, and the upper limit of CT value is CT max =500.Because the range of pixel value in the image is [0, 255], we need to normalize the CT value, and map the CT value in the visible range by formula (2).
[0067]
[0068] Wherein x is the original CT value, and x' is the pixel value after normalization processing.The DICOM format is finally converted into PNG format image, and a CTA cutout slice is converted into a 2D matrix through this step.In order to highlight the blood vessel features in the original image, facilitate input and realize feature reuse, the single-channel 2D matrix is copied to the new channel until the three-channel image is converted.
[0069] Then, the cutout data is subjected to maximum density projection, the maximum pixel value of the CTA sequence image of a single period is retained along the axial direction, the CTA sequence is converted into CTA-MIP (Maximum Intensity Projection) image, and finally the CTA sequence of four periods is converted into four MIP image data, and a patient only saves four CTA-MIP images of four periods.
[0070] Preferably, since the CTA-MIP image side branch area may exist in large vessels and the morphology is complex, the computer-aided scoring method only uses single-period images, does not combine the characteristics of richer multi-period blood vessel information, and the side branch circulation contains small blood vessels, which may cause false detection during manual detection. The application adopts two kinds of mixed attention mechanisms to enhance single-branch side branch circulation image features and fusion features, respectively. The module can also fully excavate image semantic information and effectively utilize multi-period large and small vessel information. As shown in the accompanying Figure 1 The four independent single-branch networks before fusion are composed of four structure-same but parameter-unshared independent single-branch networks, and the independent single-branch network (CCA-NET) is composed of three 3×3 convolutions, four basic feature extraction blocks, four local attention mechanism feature extraction modules and single-period feature up-sampling modules.
[0071] Further, in the CTA-based side branch circulation three-classification task, some redundant information is contained in the original image, which will interfere with the decision. After adding the attention mechanism, the network dynamically learns the region to focus on in the feature extraction stage. In the training process of the network, the model will learn different weight information to focus on more critical decision-making features, which will play a positive role in clinical diagnosis. The attention mechanism in the network is divided into two types, local attention mechanism and global attention mechanism. The local attention mechanism is set in the network before the four-period fusion, that is, the four independent feature extraction networks, and the local attention residual module is arranged in the four independent single-branch networks, as shown in the accompanying Figure 3 The local attention residual module is composed of a residual module, a group convolution and a layer normalization.
[0072] The output feature map of each BaseBlock basic residual block is defined as fe k ∈R c×h×w As input, 2 convolution kernels 3×3 are convolved to fully excavate the deep semantic information of the image, and then group normalization is used to speed up the convergence of the neural network, fe k After encoding, the deep feature map D(fe k ) with strong semantic information is obtained. k According to the design of the residual module, the new feature FE k is aggregated.
[0073] W(fe k )=σ(D(fe k )+fe k ) (3)
[0074] In formula (3), the activation function sigmoid is represented by the symbol σ, and the aggregated new feature FE k ∈R c×h×w c is the number of channels, h is the image height, w is the image width, k∈{1,2,...,c} is the index based on the channel direction, and the weighting factor is related to F. i Each pixel is multiplied and then weighted to obtain a local attention feature map L(fe). k The input features are ultimately transformed into nonlinear features, a process defined as follows:
[0075] L(fe k ) = W(fe k )*fe k (4)
[0076] In formula (4), * indicates that the corresponding pixels at the same position of the two matrices are multiplied.
[0077] Preferred options are listed below. Figure 4 As shown, F 11 ,F 12 ,F 13 ,F 14 The feature maps extracted by the four local attention modules in the first independent single-branch network are further labeled as F. mn Feature map F mn First, the data is upsampled using a 3×3 transposed convolutional layer, and then augmented to a feature map F′ with dimensions of 128 using bilinear interpolation. mn This invention designs a concatenation operation to combine all independent feature maps F′ from four single-branch networks. mn The fusion process, known as single-period feature fusion, involves the following steps:
[0078]
[0079] V m =Concate(F′ m1 F′ m2 F′ m3 F′ m4 (6)
[0080] In formula (5), F′ mn ∈R c×h×w , This indicates a transpose convolution operation with a kernel size of 3 and a stride of 2. In formula (6), Concate represents a concatenation operation, and V m It is the set of feature values after splicing on a single branch. The splicing operation merges the feature values, increasing the number of channels.
[0081] The image features of the four periods are finally fused by pixel-level addition, and the formula is as follows:
[0082] V all =Add(V1,V2,V3,V4) (7)
[0083] In formula (7), Add represents an element addition operation, V all is a feature value set after feature splicing of the four-branch network, V1, V2, V3, and V4 represent feature sets on the 1st-4th branches, respectively.
[0084] The network after the four-period fusion module is composed of two 3x3 convolution and global attention mechanism feature extraction blocks, and finally outputs the prediction result under the multi-period deep feature fusion.
[0085] Preferably, as shown in the accompanying drawings Figure 5 The global attention mechanism is arranged in the network after the four-period fusion, and the fused feature map M k ∈R c×h×w is subjected to global pooling and average pooling layers, respectively, the pooling layers are compressed along the channel axis, and then the channel dimension is spliced to obtain a feature map.
[0086] The above process is defined as:
[0087]
[0088] G max =MAX d (M k ) (9)
[0089] In formulas (8) and (9), G avg represents a two-dimensional matrix after the average pooling layer, i represents a spatial index, all index positions in the space s are selected in turn, G max represents a two-dimensional matrix after the maximum spatial pooling layer, MAX d represents a spatial maximum pooling operation.
[0090] G avg and the matrix G max are obtained by concatenation operation to obtain a feature matrix set with doubled channel number, then a 3x3 standard convolution layer is connected, the features are aggregated through the convolution layer, and finally the importance of each channel is learned to obtain a weight matrix Z of each channel by a sigmoid activation function, and the process is defined as:
[0091] Z=σ(f 3×3 [G avg ;G max ]) (10)
[0092] In formula (10), Z e R 1×h×w , f 3×3 represents the convolution operation, the convolution kernel size is 3, the step is 1, and the activation function sigmoid is represented by a symbol σ. Finally, the channel weight matrix Z is weighted with each channel of the feature map M k , to obtain the global attention feature map G(fe k ), which is defined as follows:
[0093] G(fe k ) = M k *Z (II)
[0094] After passing through the global feature extraction module, the feature map will pass through the average pooling layer and the fully connected layer, and finally the class with the maximum probability value output by SoftMax is taken as the final scoring result.
[0095] Example 2: Experimental Part
[0096] 1. Dataset, training settings and experimental equipment
[0097] The experimental data is from the Affiliated Hospital of Chongqing Medical University, the images of emergency one-stop CTP / CTA examination of all patients are imported into a workstation (Vitrea, fX, 1.0, Canon Medical Systems Corporation, Japan), a “deconvolution method based on SVD+ algorithm brain perfusion” protocol is adopted, and the software system automatically labels the inflow artery and outflow vein for post-processing. Doctors screen out four volume data packages of arterial phase, arteriovenous phase, venous phase and late venous phase in the CTA sequence images of all patients according to the perfusion time-density curve, and the volume data package of each phase image is subtracted from the corresponding head CT volume data package, so as to obtain the head CTA silhouette image data package without bone in each phase. The images are scored by two doctors with five years of experience in the imaging department, the doctors in the application adopt the improved ASITN / SIR dynamic multi-phase CTA grading scale to evaluate, the grade of collateral score is divided into good collateral circulation (label 2), there is no obvious change in the blood flow filling degree of the blood supply area on both sides, general collateral circulation (label 1), the blood flow filling degree of the ischemic side is 70% or more of the normal side, and poor collateral circulation (label 0), the blood flow filling degree of the ischemic side is less than 70% of the normal side. Doctors use Tan collateral circulation scoring system to evaluate, the grade of collateral score is divided into good collateral circulation, general collateral circulation and poor collateral circulation. Good collateral circulation means that 100% of the bypass supply fills the occluded MCA area. Poor collateral circulation means that 0% to 50% of the collateral supply fills the occluded MCA area. Moderate collateral circulation means that 50% to 100% of the collateral supply fills the occluded MCA area. If the scoring results are inconsistent, consistent scoring results are obtained through discussion. The experimental data includes 173 patients, of which 41 patients have poor collateral circulation, 50 patients have general collateral circulation, and 82 patients have good collateral circulation. After data preprocessing, the data set is expanded to 11 times of the original data set, and the data set is randomly divided into a training set and a test set, wherein the training set accounts for 70% of the total number of samples, the test set accounts for 30% of the total number of samples, the sample resolution is 256x256, and the data distribution is shown in Table 1. In the experiment, the parameter batch_size is set to 4, the Adam optimization algorithm is adopted, the initial learning rate in the first 100 epochs is set to 0.002, the exponential decay rate is set to 0.98, and the drop period is set to 1. The experiment is performed on a workstation equipped with a Windows system, an intelcorei7-7700 CPU and an NVIDIA GeForce RTX 2070 SUPER GPU. The deep learning open source framework pytorch1.7.0 is used, and the experiment is performed in the environment of python3.8.5, so as to verify the feasibility and effectiveness of the method.
[0098] Table 1: Side branch dataset distribution division / example
[0099]
[0100] 2. Evaluation index
[0101] To fully prove the effectiveness of the present application, a comprehensive experiment was designed and the accuracy, precision, recall (sensitivity), F1 score and ROC curve were used as indicators for performance evaluation. The specific mathematical definitions of each indicator are as follows:
[0102]
[0103]
[0104]
[0105]
[0106] In formulas (12-15), Accuracy, Precision, Recall and F1-score represent accuracy, precision, recall and F1-score, respectively. TP is the number of correctly classified positive examples, TN is the number of correctly classified negative examples, FP is the number of incorrectly classified positive examples, and PN is the number of incorrectly classified negative examples. The accuracy is the proportion of all correctly predicted classification samples in the total samples. The precision indicates the proportion of correctly predicted samples in the results predicted by the model. The recall indicates the proportion of correctly predicted samples in the results with the true label. According to the precision and recall, the F1-score can be calculated. The larger the value of the above four evaluation indicators, the better the performance of the classification model. AUC is the area under the receiver operating characteristic (ROC) curve, and the closer the value of AUC to 1, the higher the authenticity, and the higher the authenticity of the classification algorithm.
[0107] 3. Experimental results and analysis
[0108] 3.1 Classification performance and ablation experiment
[0109] The experiment used the Monte Carlo 5-fold cross-validation method to test the present application, and Table 1 shows the performance indicators of each fold.
[0110] Table 1: Indicators of each fold under 5-fold cross-validation / %
[0111]
[0112] To verify the effectiveness of the local feature fusion module, the local feature attention module and the global feature attention module in the method of the present application, the local feature fusion module, the local feature attention module and the global feature attention module are removed respectively, and the main network is still retained as the deep and shallow feature extraction network and the final prediction is performed. The experimental results are shown in Table 2. After removing only the local feature fusion module, the performance indicators all decrease, and the accuracy decreases by 1.75%, indicating that the fusion of the feature maps of the deep and shallow layers of the network helps to utilize more rich context information to improve the accuracy of network classification. After removing the local feature attention module and the global feature attention module respectively, the overall performance of the network decreases significantly in each indicator. After removing only the local feature attention module, retaining the main network and the global attention module as the feature extraction network and performing the final prediction, the accuracy decreases by 5.34%, indicating that the use of the local feature attention mechanism can enhance the effective use of the features of the collateral circulation region. In order to prove the role of the global feature attention module proposed in the present application in the CCA4CTA model, the performance of the model is also verified when the global feature attention module is removed. The experimental results are shown in Table 3. After removing the global feature attention module, the accuracy of the model decreases by 2.94, proving that this kind of attention method is beneficial to mining global feature information, reducing redundant information and helping the model to learn effectively.
[0113] Table 2: Ablation experiment under different modules / %
[0114]
[0115] 3.2 Classification performance comparison experiment
[0116] Since there is no similar method to the present application in the existing research on collateral circulation scoring of multi-period CTA-MIP images, in order to verify the effectiveness of the method of the present application, the current several mainstream classification networks are applied to the task of the present application, and the network CCA-Net before fusion is replaced by the current mainstream classification network. Under the condition that other conditions remain unchanged, the performance indicators of the research method are trained and sorted. Table 3 shows the performance of the present application and other several classification networks in each performance indicator. The accuracy, sensitivity, specificity and F1 score of CCA-NET are 90.42%, 78.43%, 84.19% and 81.19% respectively. As can be seen from Table 2, the accuracy, recall rate and F1 score are higher than those of other classification networks, indicating that the model of the present application has good performance in detecting positive samples and the model is robust.
[0117] Table 3: Comparison of related indicators of different methods / %
[0118]
[0119] 3.3 Single time period and multi time period method experiment comparison
[0120] Since the current clinical collateral circulation score is mainly divided into single time period input and multi time period input two cases, at present, there is no related research using multi time period CTA-MIP as input by using deep learning algorithm, the network structure in the research
[17] is taken as the network structure of single time period image, and compared with the input of multi time period image, and the results are shown in table 4. It can be seen that the network structure similar to the residual structure with attention mechanism is more beneficial to extract the context feature information of collateral circulation.
[0121] Table 4: Experimental results under different input strategies / %
[0122]
[0123] Figure 7 It is the ROC curve for the non-good case of collateral score, in the figure (a) takes multi time period as input, (b) takes artery phase as single time period input, it can be seen that the ROC curve obtained by the method of the application taking multi time period as input is more stable and smooth, and the AUC value is closer to 1, indicating that the classification effect of the method of the application is more significant, further proving the superiority of the algorithm of the application.
[0124] In view of the current lack of effective method for computer-aided multi time period brain CTA collateral circulation automatic scoring, the application designs a reasonable pretreatment scheme for CTA image data; according to the characteristics that multi time period CTA contains more rich spatial domain information, a multi time period fusion algorithm is proposed to fuse the deep and shallow information of four time periods; at the same time, according to the characteristics that collateral circulation contains many small blood vessels, an adaptive attention mechanism is designed to extract features of the image, so that the network can better distinguish the difference between different evaluation categories. Through the algorithm of the application for scoring multi time period CTA collateral circulation, the classification effect of the method of the application is better than that of the classical network, and the best performance is obtained under a plurality of indexes, the method can be used as a method for early detection, and assists doctors in diagnosis and treatment in clinic, at the same time, the method also provides a novel idea for classification of multi time period medical image data, and has far-reaching significance.
Claims
1. A method for multi-period brain CTA collateral circulation scoring under attention mechanism, characterized in that, The method comprises the steps of: S1, converting the original CTA-dicom sequence into a pseudo-RGB image, pre-processing the data by maximum density projection, and expanding the original data set by translation, rotation and displacement method for offline data enhancement; S2, four independent single-branch networks are designed for the data characteristics of multiple periods, and a feature dense module based on attention mechanism is used to compress redundant information in sparse features and extract efficient rich vessel features; Among them, the four independent single-branch network structures are the same, but do not share parameters, and the four independent single-branch networks are composed of three 3*3 convolutions, four basic feature extraction blocks, four local attention mechanism feature extraction modules and single-period feature up-sampling modules; Local attention residual modules are arranged in the four independent single-branch networks, and the local attention residual modules are composed of residual modules, grouped convolutions and layer normalization; S3, through the designed single-branch multi-feature fusion and multi-period multi-feature fusion module, the multi-period CTA images are fused to enrich the data features; Feature map F mn First, up-sampling is performed through a 3x3 transpose convolutional layer, and then expanded to a feature map F' with a length and width of 128 by a bilinear interpolation algorithm mn Feature map F mn Feature map extracted by the nth local attention mechanism feature extraction module of the mth independent single-branch network Fusion of all independent feature maps F' in the four independent single-branch networks mn is performed, referred to as single-stage feature fusion, comprising the steps of: V m =Concate(F′ m1 ,F′ m2 ,F′ m3 ,F′ m4 ) wherein F′ mn ∈ R c×h×w , denotes a transpose convolution operation with a kernel size of 3 and a stride of 2, Concate denotes a concatenation operation, V m is a set of feature values after concatenation on a single branch; The image features of the four periods are finally fused by pixel-level addition, and the formula is as follows: V all = Add(V1, V2, V3, V4) wherein, Add denotes an element addition operation, V all is a feature value set after feature splicing of four independent single-branch networks, V1, V2, V3, and V4 respectively denote feature sets on the 1st-4th branches; The network after the four-period fusion module is composed of two 3*3 convolutions, global attention mechanism feature extraction blocks, and finally outputs the prediction results under the fusion of multi-period deep features; S4, a global attention mechanism module is added to the fused network; S5, the filtered fusion features are input into a linear classifier to classify the patient's collateral circulation.
2. The method of claim 1, wherein the method is characterized by, The step S1 comprises: Step S11, set the HU threshold, when processing the head CTA image, adjust the CTA window width to 350 and the window level to 70, and the gray value HU value of the contrast agent vessel is 500; Step S12, the original data is normalized to [0, 255], and the original data is converted from DICOM to PNG image; Step S13, project the CTA sequence images of the same period along the axial direction, and convert the CTA sequence into CTA-MIP image; Step S14, the CTA-MIP image passes through the feature reuse module, the original two-dimensional pixel matrix is broadcasted along the new space axis, and the data samples are enhanced, and the specific steps are batch processing of the image matrix, including rotation, flipping and displacement.
3. The method of claim 1, wherein the method is characterized by, In step S4, the fused feature map M k ∈R c×h×w After the global pooling and average pooling layers respectively, the pooling layer compresses the matrix by performing a pooling operation along the channel axis, and then the channel dimension is spliced to obtain a feature map. The above process is defined as: G max = MAX d (M k ) wherein G avg represents a two-dimensional matrix after the average pooling layer, i represents a spatial index, all index positions in the space s are selected in turn, G max represents a two-dimensional matrix after the maximum spatial pooling layer, MAX d represents a spatial maximum pooling operation; G avg and matrix G max The channel number doubled feature matrix set is obtained by splicing operation, and then a 3 × 3 standard convolution layer is connected. The features are aggregated through the convolution layer, and finally the importance of each channel is learned. The weight matrix Z of each channel is obtained by Sigmoid activation function. The process is defined as: Z = σ(f 3×3 [ G avg ; G max ]) where Z ∈ R 1×h×w , f 3×3 denotes the convolution operation with a kernel size of 3 and a stride of 1, and the activation function sigmoid is denoted by the symbol σ.
4. The method of claim 3, wherein the method is characterized by, In step S5, the weight matrix Z of the channel is weighted with the feature map M k of each channel to obtain the global attention feature map G(fe k ), and the process is defined as follows: G(fe k ) = M k *Z After the global feature extraction module, the feature map will pass through the average pooling layer and the full connection layer, and finally the class with the maximum probability value is output by SoftMax as the final scoring result.
Citation Information
Patent Citations
Multi-branch feature fusion remote sensing scene image classification method based on attention mechanism
CN112861978A
Food material image classification model establishment method based on attention and depth feature fusion
CN114898360A