Large vascular occlusion detection method and device based on end-to-end deep learning model

By using an improved Mask R-CNN model and a three-branch attention mechanism, combined with lightweight convolutional blocks and a pre-trained ResNet50 backbone network, the problem of fast and accurate detection of vascular occlusion in NCCT images was solved, achieving end-to-end automated detection and improving the diagnostic efficiency of acute ischemic stroke.

CN121707922APending Publication Date: 2026-03-20YANGZHOU FIRST PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511672679.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing medical image processing methods struggle to quickly and accurately identify and locate vascular occlusion areas in non-contrast CT images, especially in patients with acute ischemic stroke, leading to delayed treatment and low diagnostic efficiency.

Method used

An improved Mask R-CNN model is used, combining a feature pyramid network and a small target attention module, along with lightweight convolutional blocks and a pre-trained ResNet50 backbone network. A three-branch attention mechanism is used to detect vascular occlusion in NCCT images, achieving end-to-end automated detection.

Benefits of technology

It significantly shortens the detection time for vascular occlusion, improves detection accuracy and the automation level of the model, enabling rapid and accurate identification and localization of vascular occlusion areas in primary healthcare institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121707922A_ABST
    Figure CN121707922A_ABST
Patent Text Reader

Abstract

The invention discloses a large vessel occlusion detection method and device based on an end-to-end deep learning model, and the method comprises the steps: obtaining an NCCT image and an angiography of a patient with acute ischemic stroke, carrying out the screening and marking of the NCCT image, and generating an NCCT data set with a blood vessel segmentation label; training an improved Mask R-CNN model based on the NCCT data set, wherein the improved Mask R-CNN model comprises a feature pyramid network and a small target attention module; detecting and segmenting blood vessel segments based on the trained Mask R-CNN model, and outputting an NCCT slice after multi-channel fusion of each blood vessel segment; the NCCT slice is input into a classification model based on CNN-TriFuse ResNet, the result of whether the blood vessel in the NCCT slice is occluded or not is output, and the classification model based on CNN-TriFuse ResNet comprises a lightweight convolution block, a pre-trained ResNet50 backbone network and a three-branch attention mechanism module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image-assisted diagnostic technology, specifically to a method and device for detecting large vessel occlusion based on an end-to-end deep learning model. Background Technology

[0002] Acute ischemic stroke (AIS) is a common cerebrovascular disease in clinical practice, mainly caused by sudden occlusion of cerebral blood vessels leading to insufficient blood supply to the brain tissue and subsequent neurological deficits. For AIS patients, timely and accurate diagnosis and intervention can significantly improve survival rates and quality of life. Non-contrast CT (NCCT), as a rapid, non-invasive, and widely used imaging technique, is the preferred diagnostic tool for acute ischemic stroke patients. However, the diagnosis of vascular occlusion in NCCT images faces several challenges, mainly in the following aspects: NCCT images typically show density changes in ischemic lesions but lack clear vascular contours, especially the markers of vascular occlusion, which are often blurred and difficult to identify directly from CT images. Furthermore, occluded areas often appear as low-contrast "small targets" with small areas and significant background noise, resulting in low accuracy for traditional image processing and analysis methods in detecting these small targets. Currently, less than 5% of patients with large vessel occlusion worldwide receive intravenous thrombolysis (IVT) and endovascular thrombectomy (EVT) within the eligible treatment window. Ultimately, this is mainly due to inconsistencies in local expertise, delays in vascular examination time, and differences between institutions in the interpretation of vascular imaging. In addition, initiating inter-hospital communication for the triage and transfer of patients with large vessel occlusion to EVT centers is also extremely challenging. Even in experienced systems, an average delay of approximately 142 minutes is frequently observed. Although existing medical image processing methods are gradually incorporating deep learning technology, most methods mainly focus on processing enhanced CTA imaging detection, which is highly dependent on the medical conditions of local hospitals and requires at least 40 minutes of preparation time. For acute ischemic stroke, a specific field with extremely high requirements for time and accuracy, there is still a lack of an end-to-end, automated, and fast deep learning model to directly identify and locate vascular occlusion areas on NCCT. Summary of the Invention

[0003] The embodiments described herein provide a method, apparatus, and computer-readable storage medium storing a computer program for detecting large vessel occlusion based on an end-to-end deep learning model. By introducing an improved Mask R-CNN model, the detection capability for small targets (such as vessel segments or small occlusion regions) is enhanced. Furthermore, by introducing lightweight convolutional blocks, a pre-trained ResNet50 backbone network, and a three-branch attention mechanism into the vessel occlusion detection model, rapid and accurate determination of the presence of vessel occlusion in NCCT images is achieved, greatly shortening the examination time and making it easy to promote to more grassroots medical institutions.

[0004] According to a first aspect of the present invention, a method for detecting large vessel occlusion based on an end-to-end deep learning model is provided, comprising: acquiring NCCT images and vascular imaging of patients with acute ischemic stroke, and filtering and labeling the NCCT images to generate an NCCT dataset with vascular segment labels; training an improved Mask R-CNN model based on the NCCT dataset, the improved Mask R-CNN model including a feature pyramid network and a small target attention module; detecting and segmenting vascular segments based on the trained Mask R-CNN model, and outputting NCCT slices after multi-channel fusion for each vascular segment; inputting the NCCT slices into a classification model based on CNN-TriFuseResNet, and outputting the result of whether the vascular segment is occluded, the classification model based on CNN-TriFuseResNet including lightweight convolutional blocks, a pre-trained ResNet50 backbone network and a three-branch attention mechanism module.

[0005] In some embodiments of this disclosure, acquiring NCCT images and vascular imaging of patients with acute ischemic stroke, and filtering and labeling the NCCT images to generate an NCCT dataset with vascular segmentation labels includes: acquiring NCCT images and vascular imaging of patients with acute ischemic stroke, wherein the vascular imaging is any one or more of computed tomography angiography, magnetic resonance angiography, and digital subtraction angiography; filtering the acquired NCCT images and vascular imaging, wherein the filtered image data satisfies that the time interval between NCCT images and vascular imaging does not exceed 72 hours. Angiography confirmed unilateral occlusion of the anterior circulation artery in the intracranial cavity, and the responsible vessel for this symptom was the occluded vessel confirmed by imaging. No intravenous thrombolysis or intravascular recanalization was performed between the NCCT images and the angiography examination. The vascular segments were labeled in the screened NCCT images and corrected based on the angiography. The labeled vascular segments included the C6 and C7 segments of the internal carotid artery, the M1 and M2 segments of the middle cerebral artery, and the A1 and A2 segments of the anterior cerebral artery. The labeled NCCT images were standardized to generate an NCCT dataset with vascular segment labels.

[0006] In some embodiments of this disclosure, training an improved Mask R-CNN model based on the NCCT dataset includes: using cropmaxROI to crop different blood vessel segments in the NCCT dataset into multiple patches, with each patch cropped based on the maximum diameter at both ends of the blood vessel segment; inputting the multiple patches into the improved Mask R-CNN model for training to obtain the trained improved Mask R-CNN model.

[0007] In some embodiments of this disclosure, the improved Mask R-CNN model adds a small target attention module after the output of each layer of the feature pyramid network. The small target attention module includes a channel attention unit, a spatial attention unit, a high-frequency enhancement unit, and a feature fusion unit. For high-resolution feature maps, the weight coefficient of the high-frequency enhancement unit is increased, and for low-resolution feature maps, the weight coefficient of the channel attention unit is increased. The channel attention unit calculates the statistical information of each channel on the input feature map using a global average pooling layer, and generates channel-dimensional attention weights through a multilayer perceptron and activation function. The spatial attention unit calculates channel-dimensional average pooling and max pooling on the input feature map, concatenates the two feature maps containing spatial information, and generates spatial-dimensional attention weights through convolution and activation function. The high-frequency enhancement unit extracts edge and detail features from the input feature map through multi-scale convolution, and adds the weighted edge and detail features to the original feature map through residual connections to generate an enhanced feature map. The feature fusion unit adaptively fuses the output features of the channel attention unit, spatial attention unit, and high-frequency enhancement unit.

[0008] In some embodiments of this disclosure, the detection and segmentation of blood vessel segments based on the trained Mask R-CNN model, and the output of multi-channel fused NCCT slices for each blood vessel segment, includes: using the trained improved Mask R-CNN model to detect the boundaries of different blood vessel segments and generate segmentation masks; merging the segmentation masks belonging to the same blood vessel segment to generate multi-channel fused NCCT slices, wherein the NCCT slices contain information of multiple blood vessel segments.

[0009] In some embodiments of this disclosure, a lightweight convolutional block is used to extract low-order feature maps of NCCT slices through a three-layer convolutional neural network. The first convolutional neural network has a kernel size of 3×3 and 64 channels, the second convolutional neural network has a kernel size of 3×3 and 128 channels, and the third convolutional neural network has a kernel size of 3×3 and 3 channels. Each convolutional neural network is followed by a batch normalization layer and a ReLU activation function. A pre-trained ResNet50 backbone network is used to extract high-order semantic features from the low-order feature maps extracted by the lightweight convolutional block through a residual structure. A three-branch attention mechanism module is used to weight the high-order semantic features output by the ResNet50 backbone network across width, height, and depth to output an attention-enhanced feature map. A global average pooling and classification layer is used to compress the attention-enhanced feature map into a one-dimensional feature vector, and a Softmax classifier is used to output the predicted probability of occlusion.

[0010] In some embodiments of this disclosure, the classification model based on CNN-TriFuseResNet is trained using a transfer learning strategy, which includes: freezing the weights of the first 20 layers of the ResNet50 backbone network pre-trained on the ImageNet dataset, and fine-tuning the parameters of subsequent layers of the model based on the cross-entropy loss function and the Adam optimizer.

[0011] In some embodiments of this disclosure, the method further includes: extracting activation values ​​of all feature maps in the last convolutional layer of the trained CNN-TriFuseResNet-based classification model, weighting and summing each feature map using the weights of the fully connected layer to generate a heatmap; and marking the predicted occlusion regions on the heatmap by visualizing the convolutional neural network to generate an occlusion detection and localization report.

[0012] According to a second aspect of the present invention, a large vessel occlusion detection device based on an end-to-end deep learning model is provided. The device includes at least one processor and at least one memory storing a computer program. When the computer program is executed by the at least one processor, the device causes the following actions: acquiring NCCT images and vascular imaging of patients with acute ischemic stroke, filtering and labeling the NCCT images to generate an NCCT dataset with vascular segment labels; training an improved Mask R-CNN model based on the NCCT dataset, the improved Mask R-CNN model including a feature pyramid network and a small target attention module; detecting and segmenting vascular segments based on the trained Mask R-CNN model, outputting NCCT slices after multi-channel fusion for each vascular segment; inputting the NCCT slices into a CNN-TriFuseResNet-based classification model, outputting the result of whether the vascular segment in the NCCT slice is occluded, the CNN-TriFuseResNet-based classification model including lightweight convolutional blocks, a pre-trained ResNet50 backbone network, and a three-branch attention mechanism module.

[0013] According to a third aspect of this disclosure, a computer-readable storage medium storing a computer program is provided, wherein the computer program, when executed by a processor, implements the steps of the large vessel occlusion detection method based on an end-to-end deep learning model according to the first aspect of this disclosure.

[0014] The large vessel occlusion detection method and apparatus based on an end-to-end deep learning model according to embodiments of this disclosure employs an improved Mask R-CNN model with a small target attention module added after the feature pyramid network to perform vessel segmentation detection on NCCT slices. This improves sensitivity to detailed features, accurately locates vessel segments, and enhances the accuracy of vessel occlusion detection by utilizing vessel segmentation information. This scheme uses an end-to-end deep learning framework, automating the entire process from data input to occlusion detection result output. During training, the model can self-learn and optimize feature extraction and classification steps, further improving the model's accuracy and efficiency.

[0015] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0016] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application. In the drawings: Figure 1 A flowchart illustrating a method for detecting large vessel occlusion based on an end-to-end deep learning model according to an embodiment of the present invention is shown. Figure 2 This is a schematic diagram of the structure of an improved Mask R-CNN model according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram illustrating the calculation process of the small target attention module according to an embodiment of the present disclosure; Figure 4 This is a schematic diagram of the structure of a classification model based on CNN-TriFuseResNet according to an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of a three-branch attention mechanism module according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram of the vascular occlusion detection results according to an embodiment of the present disclosure; Figure 7 This is a schematic block diagram of a large blood vessel occlusion detection device 700 based on an end-to-end deep learning model according to an embodiment of the present disclosure. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are also within the scope of protection of this disclosure.

[0018] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the specification and in the relevant art, and shall not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein.

[0019] This disclosure provides a method for detecting large vessel occlusion based on an end-to-end deep learning model. By using an improved MaskR-CNN model and a classification model based on CNN-TriFuseResNet, it can automatically identify and locate the occlusion area of ​​the vessel from NCCT images, providing doctors with fast and accurate auxiliary diagnosis.

[0020] Figure 1 This is a flowchart illustrating a large blood vessel occlusion detection method based on an end-to-end deep learning model according to an embodiment of this disclosure. (Refer to...) Figure 1As shown, in step S102, NCCT images and vascular imaging of patients with acute ischemic stroke are first acquired, and the NCCT images are filtered and labeled to generate an NCCT dataset with vascular segment labels.

[0021] According to some embodiments of this disclosure, by analyzing consecutive cases of acute unilateral anterior circulation LVO (large vessel occlusion), patients meeting specific criteria can be selected for study, and NCCT images and vascular imaging of the occluded patients can be obtained. Inclusion criteria include: age ≥18 years, confirmed acute ischemic stroke upon re-examination (i.e., short symptom onset time, and cerebral ischemia caused by cerebral vascular problems). Vascular imaging (any one or more of MRA / CTA / DSA, MRA: magnetic resonance angiography, CTA: computed tomography angiography, DSA: digital subtraction angiography) confirms unilateral intracranial anterior circulation large artery occlusion, characterized by interrupted arterial visualization and non-visualization of distal vessels. This indicates that vascular imaging confirms unilateral anterior cerebral circulation (internal carotid artery, anterior cerebral artery, etc.) vascular occlusion, and blood flow distal to the occluded segment is not visualized. The responsible vessel for this symptom is the occluded vessel confirmed by imaging, meaning the occluded vessel on imaging is closely related to the occurrence of symptoms. Non-contrast CT scans (NCCT) performed within 24 hours of symptom onset, with image quality sufficient for subsequent evaluation, are crucial for the diagnosis of acute stroke. The time interval between NCCT images and angiography should not exceed 72 hours to ensure timely diagnosis and avoid affecting interpretation due to excessive time span. No intravenous thrombolysis or endovascular recanalization was performed between NCCT and angiography (MRA / CTA / DSA) examinations to ensure that interventions for restoring blood flow do not affect the interpretation of the study results.

[0022] Patients will be excluded if the NCCT image slice thickness is ≥5mm or the image quality is poor, making it impossible to clearly display the occluded vessel or assess its imaging characteristics. For vessels distal to segments M2 and A2 (branch vessels) with small diameters, thrombectomy stents (such as Solitaire and Trevo) are difficult to anchor precisely, easily leading to vessel damage; interventional recanalization is rarely recommended. Clinically, the requirement for accurate localization of occluded vessels is not high. Therefore, in this embodiment, the vessel segmentation includes segments of the internal carotid artery C6 and C7, the middle cerebral artery M1 and M2, and the anterior cerebral artery A1 and A2. Patients will be excluded if the occlusion occurs outside the middle cerebral artery M2 segment or the distal anterior cerebral artery A2.

[0023] These inclusion and exclusion criteria are designed to screen patients who meet specific clinical conditions, ensuring that study subjects have consistent characteristics, while excluding cases with other complications or those difficult to assess on imaging. Patient NCCT scans can be performed using the following CT equipment: Siemens Somatom Definition AS+ 128-slice spiral CT, Siemens third-generation dual-source Force CT, and GE's 128-slice spiral CT. These CT devices provide high-quality images suitable for the assessment of acute stroke or vascular occlusion. Raw scan parameters include: Slice thickness: 2.0–3.0 mm, suitable for detailed brain scans while balancing scan time and image quality; Number of slices: 36–56, indicating the range of areas covered and the detail of the image; more slices result in higher resolution images; Tube voltage: 120 kV, a commonly used CT scan voltage to ensure clear image quality, especially in visualizing brain structures; Tube current: 270–450 mA, the current level determines the radiation dose and image quality during the scan. Higher current can improve image quality, but it will increase the radiation dose. Matrix: 512×512, this matrix size represents the image resolution, providing clearer details, suitable for subsequent image analysis.

[0024] To facilitate further analysis and model training, vascular segments can be labeled in the selected NCCT images and corrected based on vascular imaging. The labeled vascular segments include the C6 and C7 segments of the internal carotid artery, the M1 and M2 segments of the middle cerebral artery, and the A1 and A2 segments of the anterior cerebral artery. For example, during the labeling process, two neuroradiologists with 5 and 10 years of clinical experience independently completed the labeling and reached a consensus through discussion. ROI (Region of Interest, i.e., the relevant site of vascular occlusion) delineation can be completed using ITK-SNAP software and corrected based on comparative analysis of NCCT images and vascular imaging to ensure accurate correspondence between the occlusion area and the vessel course. Finally, the labeled NCCT images are standardized to ensure consistency in image resolution, contrast, and other factors, generating an NCCT dataset with vascular segment labels.

[0025] Subsequently, in step S104, an improved Mask R-CNN model is trained based on the NCCT dataset. The improved Mask R-CNN model includes a feature pyramid network and a small object attention module.

[0026] According to one embodiment of this disclosure, cropmaxROI can be used to crop different vascular segments in an NCCT dataset into multiple patches (rectangular slices), with each patch cropped based on the maximum diameter at both ends of the vascular segment. Specifically, the input to cropmaxROI mainly includes labeled vascular segment ROIs (regions of interest), which can be binary masks or polygonal annotations representing vascular occlusion or vascular orientation regions. For each vascular segment region, principal component analysis or least squares fitting linear methods are used to calculate the principal axis of the vessel. The principal axis roughly represents the vascular orientation, i.e., the direction of the vessel from its starting point to its ending point. The vascular ROI is scanned along the principal axis, and the width of each section is recorded. The maximum diameter at the starting and ending points is found, ensuring that the patch covers the widest area of ​​the vessel and preventing the vessel edges from being cropped. A rectangular region is generated with the vascular principal axis as the center line. The long side of the rectangular region is along the vascular principal axis, and the short side takes the maximum diameter value at both ends of the vessel, covering the entire vascular segment while ensuring that the width at both ends is not reduced. Finally, multiple rectangular slices are cropped from the original image of the NCCT dataset using the calculated rectangular region. Multiple patches are input into the improved Mask R-CNN model for training, resulting in the trained improved Mask R-CNN model.

[0027] In the traditional Mask R-CNN architecture, the Feature Pyramid Network (FPN) is used to extract feature maps at different scales to better identify multi-scale objects in images. However, the standard output of the FPN suffers from accuracy issues when processing small objects, especially when identifying fine details. This is because low resolution in high-level feature maps leads to a loss of detail, resulting in poor performance for small objects.

[0028] Figure 2 This is a schematic diagram of the structure of an improved Mask R-CNN model according to embodiments of the present disclosure. (Refer to...) Figure 2As shown, FPN enhances the model's adaptability to targets of different sizes through multi-scale feature fusion, improving the ability to detect small targets. ROI Align reduces quantization errors and improves localization accuracy through finer coordinate adjustment. The Mask branch directly generates pixel-level segmentation masks for each instance, rather than simply selecting the target region. The improved Mask R-CNN model adds small target attention modules (SAM1, SAM2, SAM3, SAM4) after each layer's output of the Feature Pyramid Network (FPN). Through a collaborative enhancement mechanism of three sub-modules (channel attention unit, spatial attention unit, and high-frequency enhancement unit), different aspects of the feature map are strengthened, enabling it to more effectively capture the details of small targets. This module uses differentiated parameter configurations for feature maps of different resolutions: for high-resolution feature maps, the weight coefficients of the high-frequency enhancement unit are increased; for low-resolution feature maps, the weight coefficients of the channel attention unit are increased.

[0029] In step S106, the trained Mask R-CNN model is used to detect and segment blood vessel segments, and outputs NCCT slices of each blood vessel segment after multi-channel fusion.

[0030] The improved Mask R-CNN model, after training, takes a cropped NCCT image as input. The model outputs the boundaries and segmentation masks of the vessel segments. Segmentation masks belonging to the same vessel segment are then merged to generate multi-channel fused NCCT slices containing information from multiple vessel segments.

[0031] Figure 3 This is a schematic diagram illustrating the calculation process of the small target attention module according to an embodiment of this disclosure. (Refer to...) Figure 3 As shown, the small target attention module includes a channel attention unit, a spatial attention unit, a high-frequency enhancement unit, and a feature fusion unit.

[0032] The channel attention unit is used to calculate the statistical information of each channel on the input feature map through a global average pooling layer, and generates the channel-dimensional attention weights through a multilayer perceptron and activation function: X_c = σ(MLP(Pool(F)))⊙ X, where σ represents the Sigmoid function, MLP represents the multilayer perceptron, Pool represents the pooling operation, X represents the input feature map, and ⊙ represents element-wise multiplication.

[0033] The spatial attention unit is used to calculate the channel-dimensional average pooling and max pooling of the input feature maps. It concatenates two feature maps containing spatial information and then generates spatial attention weights through convolution and activation functions. For example, concatenating two average-pooled feature maps Avg(X) and max-pooled feature maps Max(X) and then passing them through a 7x7 convolutional layer and a sigmoid function generates spatial attention weights: W_s = σ(Conv([Avg(X); Max(X)]))⊙X_c, where W_s is the spatial attention weight, Conv is the convolution operation, Avg is the average value calculated for each channel, and Max is the maximum value calculated for each channel.

[0034] The high-frequency enhancement unit is used to extract edge and detail features from the input feature map through multi-scale convolution, and then adds the weighted edge and detail features to the original feature map to generate an enhanced feature map. For example, high-frequency feature extraction is performed on the input feature map X through 3x3 convolutional layers: Edge(X) and Detail(X), which extract the edge and detail features in the image, respectively. The enhanced features are then added to the original features through weighted and residual connections to obtain the enhanced feature map X_h.

[0035] To effectively fuse the features generated by the three sub-modules, the small target attention module employs an adaptive feature fusion mechanism. The fusion process, based on the weights of each module, combines features from channel attention, spatial attention, and high-frequency enhancement using learnable parameters for weighted fusion. The feature fusion unit adaptively fuses the output features of the channel attention unit, spatial attention unit, and high-frequency enhancement unit. The fused feature is: Y = X + Calib(αX_c + βX_s + γX_h), where α, β, and γ are adaptive fusion weights normalized by Softmax, and Calib is a feature calibration function used to further adjust the fused features to ensure their effectiveness. In practical applications, feature maps at different resolutions have different characteristics. The small target attention module dynamically adjusts its parameter configuration according to the feature map resolution: for high-resolution feature maps (i.e., low-level features), the weight coefficient of the high-frequency enhancement unit is increased because high-resolution feature maps contain more detailed information, and edge and texture information is crucial for small targets. For low-resolution feature maps (i.e., high-level features), the weight coefficient of the channel attention unit is increased because, at low resolutions, channel-level feature representation is particularly important for small target detection.

[0036] This adaptive weight configuration can be adjusted according to the actual needs of the training data to optimize the fusion effect of each module. Therefore, this approach can significantly improve the performance of Mask R-CNN when processing small objects. The small object attention module enhances small objects in multiple dimensions, enabling the network to better capture and express the detailed features of small objects, thereby improving the overall detection accuracy, especially when dealing with small objects with high noise or low contrast.

[0037] Finally, in step S108, the NCCT slices are input into the CNN-TriFuseResNet-based classification model, which outputs the result of whether the blood vessels in the NCCT slices are occluded. The CNN-TriFuseResNet-based classification model includes lightweight convolutional blocks, a pre-trained ResNet50 backbone network, and a three-branch attention mechanism module.

[0038] Figure 4 This is a schematic diagram of the structure of a classification model based on CNN-TriFuseResNet according to an embodiment of this disclosure. (Refer to...) Figure 4 As shown, the first part of the CNN-TriFuseResNet-based classification model is a custom lightweight convolutional block that extracts low-order feature maps from NCCT slices through a three-layer convolutional neural network. Each layer has a 3×3 kernel size and 64, 128, and 3 channels, respectively. After the convolution operation, each layer is followed by batch normalization and ReLU activation, which helps stabilize the training process and accelerate model convergence. The output of feature extraction is the convolutional feature map, which can be represented as: y=f(W x+b) Where W and b are the convolution kernel weights and biases, respectively, and x is the input NCCT slice. It is a convolution operation, f( ) is the ReLU activation function.

[0039] After initial low-level feature extraction, the pre-trained ResNet50 backbone network is used to perform hierarchical extraction of high-order semantic features from the feature maps through residual structures. ResNet50 utilizes residual structures to address the vanishing gradient problem common in deep networks, enhancing the model's learning ability. The residual block format is as follows: H(x) = F(x, W) + x Here, F(x,W) represents the nonlinear transformation of the input features, x is the input NCCT slice, and H(x) is the output of the residual block. This structure enables the network to effectively capture more complex semantic features.

[0040] To further improve the model's sensitivity to subtle features in low-contrast regions, especially occluded regions, a three-branch attention mechanism module was introduced. Figure 5 This is a schematic diagram of the structure of a three-branch attention mechanism module according to an embodiment of this disclosure. (Refer to...) Figure 5 As shown, this module weights the feature map from three dimensions: height, width, and channels, and outputs an attention-enhanced feature map. The calculation process of the attention mechanism can be represented as follows: F′=F Aw Ah Ac Where F is the nonlinear transformation of the input NCCT slice, and Aw, Ah, and Ac represent the attention weights for width, height, and channel dimensions, respectively. This is an element-wise multiplication operation. With attention enhancement, the model can focus more on details in areas related to vascular occlusion, thereby improving the prediction accuracy for low-contrast regions.

[0041] The feature maps enhanced by the three-branch attention mechanism are input into a global average pooling (GAP) layer. The GAP layer compresses the features into a single feature vector by averaging them across the spatial dimensions. After GAP, Dropout (with a scale of 0.1) is added to prevent overfitting and improve the model's generalization ability.

[0042] Finally, the predicted occlusion probability is output through a fully connected (FC) layer. The Softmax function is used to output the probability distribution of the classes. P(y|x)=Softmax(Wz+b) Where z is the pooled feature vector, and W and b are the weights and biases of the fully connected layer. The goal of this classifier is to determine whether there is vascular occlusion in the input slice.

[0043] According to one embodiment of this disclosure, to accelerate convergence and reduce training time, a transfer learning strategy is employed to train a classification model based on CNN-TriFuseResNet. Specifically, the weights of the first 20 layers of the ResNet50 backbone network pre-trained on the ImageNet dataset are frozen, and the parameters of subsequent layers are fine-tuned based on the cross-entropy loss function and the Adam optimizer. The cross-entropy loss function is:

[0044] Where yi represents the true label and yi′ represents the model's predicted probability. The Adam optimizer is used with a small learning rate, a batch size of 32, and 30 training iterations. The training objective is to balance computational efficiency and classification accuracy, and to maximize the ability to extract subtle image features from NCCT slices.

[0045] After training, the model can automatically determine whether an input NCCT slice is occluded. Specific applications include: extracting activation values ​​from the last convolutional layer of a CNN-TriFuseResNet-based classification model after training; weighting and summing each feature map using the weights of the fully connected layers to generate a heatmap; and visualizing the convolutional neural network to annotate the predicted occlusion regions on the heatmap, generating an occlusion detection and localization report.

[0046] Figure 6 This is a schematic diagram of the vascular occlusion detection results according to an embodiment of this disclosure. (Refer to...) Figure 6 As shown, through a visual convolutional neural network (CAM), the model can generate a visual heatmap that marks the predicted occlusion area. This spatial localization capability allows the model not only to determine whether an occlusion exists, but also to pinpoint specific vessel segments. This visualization provides doctors with a more intuitive image reference. Based on the heatmap generated by CAM, the model can generate a detailed diagnostic report, indicating the specific location of the occlusion and providing classification information on whether an occlusion has occurred. This not only improves diagnostic efficiency but also provides doctors with precise localization, guiding subsequent treatment.

[0047] Figure 7 This is a schematic block diagram of a large blood vessel occlusion detection device 700 based on an end-to-end deep learning model according to an embodiment of the present disclosure. Figure 7 As shown, the device 700 may include a processor 710 and a memory 720 storing a computer program. When the computer program is executed by the processor 710, the device 700 is made capable of performing actions such as... Figure 1 The steps of the large vessel occlusion detection method 100 based on an end-to-end deep learning model are shown below. In one example, device 700 can be a computer device or a cloud computing node. Device 700 can: acquire NCCT images and vascular imaging of patients with acute ischemic stroke, and filter and label the NCCT images to generate an NCCT dataset with vessel segment labels; train an improved Mask R-CNN model based on the NCCT dataset, the improved Mask R-CNN model including a feature pyramid network and a small target attention module; detect and segment vessel segments based on the trained Mask R-CNN model, and output NCCT slices after multi-channel fusion for each vessel segment; input the NCCT slices into a CNN-TriFuseResNet-based classification model, and output the result of whether the vessel in the NCCT slice is occluded, the CNN-TriFuseResNet-based classification model including lightweight convolutional blocks, a pre-trained ResNet50 backbone network, and a three-branch attention mechanism module.

[0048] In some embodiments of this disclosure, the device 700 can acquire NCCT images and vascular imaging of patients with acute ischemic stroke. The vascular imaging can be any one or more of computed tomography angiography, magnetic resonance angiography, and digital subtraction angiography. The acquired NCCT images and vascular imaging are screened, and the screened image data meets the following requirements: the time interval between the NCCT images and vascular imaging does not exceed 72 hours; the vascular imaging confirms unilateral occlusion of the anterior circulation artery in the intracranial cavity; the responsible vessel for the current symptoms is the occluded vessel confirmed by imaging; and no intravenous thrombolysis or intravascular recanalization treatment was performed between the NCCT images and the vascular imaging examination. The vascular segments are labeled in the screened NCCT images and corrected based on the vascular imaging. The labeled vascular segments include the vascular segment labels of the internal carotid artery C6 and C7 segments, the middle cerebral artery M1 and M2 segments, and the anterior cerebral artery A1 and A2 segments. The labeled NCCT images are standardized to generate an NCCT dataset with vascular segment labels.

[0049] In some embodiments of this disclosure, the device 700 can use cropmaxROI to crop different blood vessel segments in the NCCT dataset into multiple patches, with each patch cropped based on the maximum diameter at both ends of the blood vessel segment; the multiple patches are then input into an improved Mask R-CNN model for training to obtain the trained improved Mask R-CNN model.

[0050] In some embodiments of this disclosure, the device 700 can use a trained improved Mask R-CNN model to detect the boundaries of different vascular segments and generate segmentation masks; the segmentation masks belonging to the same vascular segment are merged to generate multi-channel fused NCCT slices, which contain information of multiple vascular segments.

[0051] In some embodiments of this disclosure, the device 700 can extract the activation values ​​of all feature maps in the last convolutional layer of a CNN-TriFuseResNet-based classification model after training, use the weights of the fully connected layer to perform a weighted summation on each feature map to generate a heatmap, and use the visualization convolutional neural network to mark the predicted occlusion regions on the heatmap to generate an occlusion detection and localization report.

[0052] In embodiments of this disclosure, processor 710 may be, for example, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a processor based on a multi-core processor architecture, etc. Memory 720 may be any type of memory implemented using data storage technologies, including but not limited to random access memory, read-only memory, semiconductor-based memory, flash memory, disk storage, etc.

[0053] Furthermore, in embodiments of this disclosure, the device 700 may also include an input device 730, such as a keyboard or mouse, for inputting acquired NCCT images and vascular imaging, etc. Additionally, the device 700 may also include an output device 740, such as a display, for outputting the probability of NCCT slice occlusion, the specific location of the occlusion area, thermal images, and other determination results.

[0054] In other embodiments of this disclosure, a computer-readable storage medium storing a computer program is also provided, wherein the computer program, when executed by a processor, is capable of performing the following functions: Figure 1 The steps of the large vessel occlusion detection method 100 based on an end-to-end deep learning model are shown.

[0055] In summary, the large vessel occlusion detection method and apparatus based on an end-to-end deep learning model according to embodiments of this disclosure employs an improved Mask R-CNN model with a small target attention module added after the feature pyramid network to perform vessel segmentation detection on NCCT slices. This improves sensitivity to detailed features, accurately locates vessel segments, and enhances the accuracy of vessel occlusion detection by utilizing vessel segmentation information. This scheme uses an end-to-end deep learning framework, automating the entire process from data input to occlusion detection result output. During training, the model can self-learn and optimize feature extraction and classification steps, further improving the model's accuracy and efficiency.

[0056] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatuses and methods according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0057] Unless otherwise expressly indicated by the context, the singular form of words used herein and in the appended claims includes the plural form, and vice versa. Thus, when referring to the singular, the plural form of the corresponding term is generally included. Similarly, the terms “comprising” and “including” shall be interpreted as including rather than exclusively. Likewise, the terms “including” and “or” shall be interpreted as including unless such interpretation is expressly prohibited herein. Where the term “example” is used herein, particularly when it follows a set of terms, the “example” is merely exemplary and illustrative and should not be considered exclusive or extensive.

[0058] Further aspects and scope of adaptation become apparent from the description provided herein. It should be understood that various aspects of this application may be implemented individually or in combination with one or more other aspects. It should also be understood that the descriptions and specific embodiments herein are for illustrative purposes only and are not intended to limit the scope of this application.

[0059] Several embodiments of this disclosure have been described in detail above. However, it is obvious that those skilled in the art can make various modifications and variations to the embodiments of this disclosure without departing from the spirit and scope of this disclosure. The scope of protection of this disclosure is defined by the appended claims.

Claims

1. A method for detecting large vessel occlusion based on an end-to-end deep learning model, characterized in that, include: Acquire NCCT images and vascular imaging of patients with acute ischemic stroke, and filter and annotate the NCCT images to generate an NCCT dataset with vascular segment labels. An improved Mask R-CNN model was trained based on the NCCT dataset. The improved Mask R-CNN model includes a feature pyramid network and a small target attention module. Based on the trained Mask R-CNN model, the system detects and segments blood vessels, and outputs NCCT slices of each blood vessel segment after multi-channel fusion. The NCCT slices are input into a classification model based on CNN-TriFuseResNet, which outputs the result of whether the blood vessels in the NCCT slices are occluded. The classification model based on CNN-TriFuseResNet includes lightweight convolutional blocks, a pre-trained ResNet50 backbone network, and a three-branch attention mechanism module.

2. The method for detecting large vessel occlusion based on an end-to-end deep learning model according to claim 1, characterized in that, The process of acquiring NCCT images and vascular imaging of patients with acute ischemic stroke, and filtering and labeling the NCCT images to generate an NCCT dataset with vascular segment labels includes: Acquire NCCT images and vascular imaging of patients with acute ischemic stroke, wherein the vascular imaging is any one or more of computed tomography angiography, magnetic resonance angiography and digital subtraction angiography. The acquired NCCT images and vascular imaging were screened. The screened image data met the following conditions: the time interval between the NCCT images and vascular imaging did not exceed 72 hours; the vascular imaging confirmed unilateral intracranial anterior circulation aortic occlusion; the responsible vessel for this symptom was the occluded vessel confirmed by imaging; and no intravenous thrombolysis or intravascular recanalization treatment was performed between the NCCT images and the vascular imaging examination. In the selected NCCT images, vascular segments were marked and corrected based on vascular imaging. The marked vascular segments included the internal carotid artery C6 and C7 segments, the middle cerebral artery M1 and M2 segments, and the anterior cerebral artery A1 and A2 segments. The labeled NCCT images are standardized to generate an NCCT dataset with blood vessel segment labels.

3. The method for detecting large vessel occlusion based on an end-to-end deep learning model according to claim 1, characterized in that, The improved Mask R-CNN model trained based on the NCCT dataset includes: The cropmaxROI method was used to crop different blood vessel segments in the NCCT dataset into multiple patches, with each patch cropped based on the maximum diameter at both ends of the blood vessel segment. The multiple patches are input into the improved Mask R-CNN model for training, resulting in the trained improved Mask R-CNN model.

4. The method for detecting large vessel occlusion based on an end-to-end deep learning model according to claim 1, characterized in that, The improved Mask R-CNN model adds a small target attention module after the output of each layer of the feature pyramid network. The small target attention module includes a channel attention unit, a spatial attention unit, a high-frequency enhancement unit, and a feature fusion unit. For high-resolution feature maps, the weight coefficient of the high-frequency enhancement unit is increased, and for low-resolution feature maps, the weight coefficient of the channel attention unit is increased. The channel attention unit is used to calculate the statistical information of each channel of the input feature map through a global average pooling layer, and to generate channel-dimensional attention weights through a multilayer perceptron and activation function. The spatial attention unit is used to calculate the average pooling and max pooling of the channel dimension for the input feature map respectively, and to generate spatial dimension attention weights by concatenating two feature maps containing spatial information and then performing convolution and activation functions. The high-frequency enhancement unit is used to extract edge features and detail features from the input feature map through multi-scale convolution, and to add the weighted edge features and detail features to the original feature map through residual connection to generate an enhanced feature map; The feature fusion unit is used to adaptively fuse the output features of the channel attention unit, spatial attention unit, and high-frequency enhancement unit.

5. The method for detecting large vessel occlusion based on an end-to-end deep learning model according to claim 1, characterized in that, The NCCT slices of each blood vessel segment, obtained by detecting and segmenting blood vessels using the trained Mask R-CNN model, and outputting multi-channel fused slices of each blood vessel segment, include: The improved Mask R-CNN model was trained to detect the boundaries of different blood vessel segments and generate segmentation masks. The segmentation masks belonging to the same vascular segment are merged to generate a multi-channel fused NCCT slice, which contains information about multiple vascular segments.

6. The method for detecting large vessel occlusion based on an end-to-end deep learning model according to claim 1, characterized in that, The lightweight convolutional block is used to extract low-order feature maps from the NCCT slices through a three-layer convolutional neural network. The first layer of the convolutional neural network has a kernel size of 3×3 and 64 channels, the second layer has a kernel size of 3×3 and 128 channels, and the third layer has a kernel size of 3×3 and 3 channels. Each layer of the convolutional neural network is followed by a batch normalization layer and a ReLU activation function. The pre-trained ResNet50 backbone network is used to extract high-order semantic features from the low-order feature maps extracted by the lightweight convolutional block through a residual structure. The three-branch attention mechanism module is used to weight the high-order semantic features output by the ResNet50 backbone network across width, height, and depth to output an attention-enhanced feature map. The global average pooling and classification layer is used to compress the attention-enhanced feature map into a one-dimensional feature vector, and output the predicted probability of occlusion through a Softmax classifier.

7. The method for detecting large vessel occlusion based on an end-to-end deep learning model according to claim 1, characterized in that, The classification model based on CNN-TriFuseResNet is trained using a transfer learning strategy, which includes: The weights of the first 20 layers of a ResNet50 backbone network pre-trained on the ImageNet dataset are frozen, and the parameters of subsequent layers of the model are fine-tuned based on the cross-entropy loss function and the Adam optimizer.

8. The method for detecting large vessel occlusion based on an end-to-end deep learning model according to claim 1, characterized in that, The method further includes: The activation values ​​of all feature maps are extracted in the last convolutional layer of the trained CNN-TriFuseResNet-based classification model. The weights of the fully connected layers are used to sum the values ​​of each feature map to generate a heatmap. The predicted occlusion areas are marked on the heat map by visualizing the convolutional neural network, and an occlusion detection and localization report is generated.

9. A large blood vessel occlusion detection device based on an end-to-end deep learning model, characterized in that, The device includes: At least one processor; and At least one memory storing a computer program; When the computer program is executed by the at least one processor, the device performs the steps of the large vessel occlusion detection method based on an end-to-end deep learning model according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the large vessel occlusion detection method based on an end-to-end deep learning model according to any one of claims 1 to 8.