An intracranial aneurysm image detection method, system, device and storage medium

By adopting the multi-scale axial residual projection network (MARP-Net) architecture, the problems of low detection accuracy, long training cycle and low detection efficiency in the prior art are solved, and more efficient and accurate detection and segmentation of intracranial aneurysm are achieved.

CN119991681BActive Publication Date: 2025-06-20SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510479431.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-06-20
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing intracranial aneurysm detection methods based on 3D TOF-MRA technology have problems with low detection accuracy, long training cycle and low detection efficiency, especially when dealing with complex backgrounds and small targets.

Method used

A new method for detecting intracranial aneurysm images is proposed, using a multi-scale axial residual projection network (MARP-Net) architecture, and through maximum density projection, multiple filtering processing and feature extraction modules, the detection accuracy and efficiency are improved.

Benefits of technology

It significantly improves the accuracy of intracranial aneurysm detection and segmentation, and significantly improves detection efficiency, allowing more efficient processing of complex backgrounds and small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991681B_ABST
    Figure CN119991681B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of medical image segmentation, and relates to a method, system, device and storage medium for detecting intracranial aneurysm images. The present invention designs a new network architecture, namely a multi-scale axial residual projection network, to improve the accuracy of intracranial aneurysm detection and segmentation. In terms of model input, the present invention adopts preprocessing steps including matched filtering, Gaussian filtering and Laplacian filtering to enhance vascular structures and edge information and reduce the demand for computing resources. In terms of model structure, the present invention uses the DWR module to achieve multi-scale feature extraction and enhance the model's segmentation ability for small targets; in addition, the ACRE module is used to enhance the foreground and suppress the background through the axial attention module and the context relationship encoder module, improving the network's processing ability for the edge information of the foreground. The present invention can significantly improve the accuracy of intracranial aneurysm detection and segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image segmentation, and particularly relates to a method, a system, a device and a storage medium for intracranial aneurysm image detection. Background Art

[0002] The three-dimensional time-of-flight magnetic resonance angiography (3D TOF-MRA) technology has the advantages of safety and non-invasiveness because it does not use contrast agents, and has great application potential in the examination and diagnosis of intracranial aneurysms, and can timely detect unruptured intracranial aneurysms in early diagnosis. However, when detecting aneurysms using 3D TOF-MRA images, there are problems such as insufficient display of small aneurysms, long examination time, long training time, and high memory occupancy. Moreover, with the progress of imaging technology, the number of layers of three-dimensional medical images has been increasing, and the workload of doctors' manual film reading has increased significantly. The high-intensity film reading work will reduce the sensitivity of radiologists' diagnosis.

[0003] The maximum intensity projection (MIP) image generated by rotating and projecting 3D TOF-MRA images retains the density information of the original image to a great extent, and has fewer layers than the original 3D image, reducing the requirements for computing resources. However, the location of intracranial aneurysms is random, and there are problems such as blood vessel occlusion. Intracranial aneurysms are only easy to observe at certain projection angles, and there are differences in the cerebrovascular structures of different people, making the optimal projection angle random.

[0004] In addition, the current intracranial aneurysm detection algorithms can be divided into two categories: traditional algorithms and deep learning algorithms. Traditional algorithms usually detect based on one or more features of intracranial aneurysms. Deep learning algorithms can use depth information to improve the detection performance of intracranial aneurysms. By training network models with a large amount of data, the detection effect of intracranial aneurysms is more robust than that of traditional algorithms. For the research on deep learning algorithms for intracranial aneurysms, the dimension of the input image is the primary factor considered in intracranial aneurysm detection. Because 3D images contain more useful information, but require more memory and training time compared to 2D images.

[0005] In the research of intracranial aneurysms based on 3D images, most algorithm implementations adopt the basic structure of an encoder and a decoder. Improvements usually focus on adjusting the network structure, modifying the size of convolutional kernels, extracting and fusing multi-scale features, adding attention mechanisms, and improving loss functions, etc., to enhance the detection performance of the algorithm. In addition, some researchers also combine traditional algorithms with deep learning algorithms. Based on the original deep learning model, they embed the pre-processing or post-processing methods of traditional algorithms into the overall detection process to improve the detection accuracy of the algorithm. There are also some researchers who combine two or more lightweight deep learning models and propose a multi-stage learning strategy to achieve fine segmentation and detection of intracranial aneurysms.

[0006] To sum up, traditional algorithms have fewer restrictions on the number of training samples, and their detection efficiency for intracranial aneurysms is better than that of deep learning methods. However, deep learning detection algorithms have stronger feature extraction capabilities and better detection effects for data with sufficient sample quantities and complex intracranial aneurysm features. Moreover, with the improvement of computer computing power, CAD-assisted diagnosis applications based on deep learning are becoming more and more widespread. Nevertheless, there are still certain drawbacks in the current deep learning-based intracranial aneurysm algorithms, namely, low detection accuracy of intracranial aneurysms, long training cycles of intracranial aneurysm detection algorithms, and low detection efficiency, etc. Summary of the Invention

[0007] The purpose of the present invention is to propose a method for detecting intracranial aneurysm images, which can improve the accuracy of intracranial aneurysm detection and segmentation by proposing a new network structure, and significantly improve the detection efficiency.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] A method for detecting intracranial aneurysm images includes the following steps:

[0010] Step 1. Perform a maximum intensity projection operation on 3D TOF-MRA images to obtain MIP images;

[0011] Step 2. Perform various different filtering processes on the generated MIP images, adjust the sizes of the filtered images to be unified, and then splice the unified-sized images to obtain pre-processed images;

[0012] Step 3. Input the pre-processed images into a pre-built intracranial aneurysm image detection model based on a multi-scale axial residual projection network for intracranial aneurysm detection to obtain intracranial aneurysm detection results;

[0013] The detection model implements multi-scale feature extraction based on the DWR module, and uses the correlation-based axial context relationship encoder ACRE as the axial attention mechanism and residual connection to strengthen the feature map, improve the segmentation ability for small targets and complex backgrounds, achieve deeper feature fusion and information extraction, and enable the detection model to focus on the segmentation of small targets.

[0014] In addition, based on the above intracranial aneurysm image detection method, the present invention also proposes an intracranial aneurysm image detection system corresponding to the intracranial aneurysm image detection method, which adopts the following technical solutions:

[0015] An intracranial aneurysm image detection system includes the following modules:

[0016] The MIP image acquisition module is used to perform a maximum intensity projection operation on the 3D TOF-MRA image to obtain the MIP image;

[0017] The image preprocessing module is used to perform various different filtering processes on the generated MIP image, adjust the sizes of the filtered images to be unified, and then splice the images with unified sizes to obtain a preprocessed image;

[0018] And the prediction module is used to input the preprocessed image into the intracranial aneurysm image detection model based on the multi-scale axial residual projection network built in advance for intracranial aneurysm detection to obtain the intracranial aneurysm detection result;

[0019] The detection model implements multi-scale feature extraction based on the DWR module, and uses the correlation-based axial context relationship encoder ACRE as the axial attention mechanism and residual connection to strengthen the feature map, improve the segmentation ability for small targets and complex backgrounds, achieve deeper feature fusion and information extraction, and enable the detection model to focus on the segmentation of small targets.

[0020] In addition, based on the above intracranial aneurysm image detection method, the present invention also proposes a computer device, which includes a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it is used to implement the above intracranial aneurysm image detection method.

[0021] In addition, based on the above intracranial aneurysm image detection method, the present invention also proposes a computer-readable storage medium, on which a program is stored. When the program is executed by the processor, it is used to implement the above intracranial aneurysm image detection method.

[0022] The present invention has the following advantages:

[0023] As described above, the present invention relates to a method for detecting intracranial aneurysm images, which designs a new network architecture, namely the multi-scale axial residual projection network MARP-Net, to improve the accuracy of intracranial aneurysm detection and segmentation. In terms of model input, by performing maximum intensity projection on the original 3D TOF-MRA to obtain the MIP image, the density information of the original image can be retained, the number of image layers can be reduced, the requirement for computing resources can be lowered, and the detection efficiency can be improved. In addition, the present invention adopts preprocessing steps including matched filtering, Gaussian filtering, and Laplace filtering to enhance the vascular structure and edge information. By optimizing these preprocessing steps, noise and redundant information can be reduced, the requirement for computing resources can be decreased, the training efficiency of the model can be improved, and the training cycle can be shortened. In terms of the model structure, the present invention uses the DWR module to achieve multi-scale feature extraction and enhance the model's segmentation ability for small targets. In addition, the present invention uses the ACRE module to enhance the foreground and suppress the background through the axial attention module and the context relationship encoder module, improve the network's processing ability for the edge information of the foreground, and proposes to adopt a more computationally efficient proxy attention mechanism in the ACRE module. Through the design of the network architecture, the processing ability for the preprocessed image is further optimized, and the requirement for computing resources is reduced. Through the above improvements in model input and model network structure, the present invention can significantly improve the accuracy of intracranial aneurysm detection and segmentation, and the detection efficiency is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is the overall flowchart of the method for detecting intracranial aneurysm images in the embodiment of the present invention;

[0025] Figure 2 It is a schematic diagram of performing maximum intensity projection on the 3D TOF-MRA image;

[0026] Figure 3 It is the schematic diagram of the bilinear interpolation method in the embodiment of the present invention;

[0027] Figure 4 It is the network architecture diagram of the multi-scale axial residual projection network MARP_Net built in the embodiment of the present invention;

[0028] Figure 5 It is the network structure diagram of the improved DWR module in the embodiment of the present invention;

[0029] Figure 6 It is the network structure diagram of the improved ACRE attention module in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:

[0031] Embodiment 1

[0032] As Figure 1 shown, this embodiment describes an intracranial aneurysm image detection method. First, a maximum intensity projection operation is performed on the original 3D TOF-MRA image. After performing various different filtering processes on the generated MIP image, the image quality is optimized to improve the detection accuracy, suppress noise interference, and solve the inherent defects of the MIP technology. Then, the preprocessed image formed after resizing and stitching is input into a pre-built intracranial aneurysm segmentation network with multi-scale features and axial residual projection enhancement (Multi-scale Axial Residual Projection Network, MARP_Net), simply referred to as the multi-scale axial residual projection network, which is the intracranial aneurysm image detection model. Inside the model, the DWR module is used to extract multi-scale features from the preprocessed image, effectively enhancing the feature representation ability. Among them, in view of the problem that the existing DWR module adopts a multi-branch structure with a fixed dilation rate, which cannot adapt to the dynamic requirements of different target scales, the parallel calculation of multiple branches increases the number of parameters and the computational cost, and the direct splicing of the outputs of different dilation rate branches may lead to insufficient multi-scale feature fusion, this invention proposes an improved DWR module, which adopts a frequency-adaptive dynamic dilation rate AdaDR, changing the fixed dilation rate to a dynamically adjusted dilation rate to balance the effective bandwidth and the receptive field. Subsequently, the ACRE module is used to dynamically adjust the weight distribution of different positions in the feature map by using the axial attention mechanism. However, the traditional attention mechanism adopted in the ACRE module usually needs to calculate the interactions between all features, which will lead to a very large amount of calculation in the high-dimensional feature space and lack effective regularization means, easily resulting in overfitting of the model and poor generalization ability. In view of this problem, this invention improves the ACRE module and adopts a proxy attention mechanism. By introducing proxy features, the features are first dimension-reduced or simplified, and then the attention weights are calculated, significantly reducing the amount of calculation in the high-dimensional feature space, improving the computational efficiency of the model, and the calculation of the attention weights through proxy features can play a certain regularization effect, reducing the risk of overfitting of the model, thereby highlighting the foreground target and suppressing background noise interference, while strengthening the edge information of the features. Finally, the features from different channels are deeply fused (that is, the result after axial attention processing is fused with the foreground and background features to supplement the foreground and edge information of the features, obtain the processed features, and get an accurate final segmentation result), generating an accurate final segmentation result.

[0033] As Figures 1 to 6 shown, the intracranial aneurysm image detection method in this embodiment includes the following steps:

[0034] Step 1. Perform a maximum intensity projection operation on the 3D TOF-MRA image to obtain a MIP image.

[0035] As Figure 2 shows a schematic diagram of the maximum intensity projection process. Maximum intensity projection is generated by calculating the maximum density pixels encountered along each ray of the scanned object. When the light passes through the volume data, the pixels with the maximum density are saved and projected onto a two-dimensional plane, thus forming a MIP image, which can well display changes such as stenosis and dilation of blood vessels.

[0036] Perform a maximum intensity projection on the 3D TOF-MRA image and detect intracranial aneurysms on the obtained MIP image.

[0037] Perform a maximum intensity projection on the original 3D TOF-MRA image to generate a MIP image. This step can preserve the density information of the original image, reduce the number of image layers, and lower the requirements for computing resources.

[0038] There are quite a lot of black background areas in the original MIP image, which will occupy a large amount of computing resources, and are of no help to the feature learning of intracranial aneurysms. Moreover, there is interference from high-frequency noise, affecting the image clarity.

[0039] Therefore, preprocessing is performed on the MIP image before feeding it into the MARP_Net model (as shown in Step 2 below). Utilizing the anatomical prior knowledge that intracranial aneurysms attach to blood vessels, the possible regions where intracranial aneurysms may appear are located through blood vessels, which can reduce the computing cost while reducing the interference of targets such as the skull, thereby effectively improving the detection accuracy.

[0040] Step 2. Perform various different filtering processes on the generated MIP image, adjust the size of each filtered image to be unified, and then splice the images with unified size to obtain a preprocessed image.

[0041] In this embodiment, the filtering methods, for example, include three types: matched filtering, Gaussian filtering, and Laplacian filtering. The bilinear interpolation method is used to adjust the size of each image after each filtering process to be unified.

[0042] A matched filter is a filter used to detect specific signals or patterns. It detects whether there is a part in the signal that matches the filter template by performing a correlation operation with the input signal. In image processing, a matched filter can be used to detect specific image features or patterns. Use a matched filter to enhance the contrast of blood vessels and suppress background noise. Highlight the blood vessel structure through a matched filter to provide a clear blood vessel image for subsequent feature extraction and image segmentation.

[0043] The kernel of the matched filter It is expressed as follows:

[0044] ; where represents the position coordinates of pixels in the original image, represents the range of the filter cross-section intensity, and L represents the length of the blood vessel.

[0045] For the input original image , the image after matched filtering processing has the following formula:

[0046] .

[0047] Where and represent the relative position offset of the filter template in the image, that is, the moving position of the filter template in the image. For example, when and , the filter kernel is completely aligned with a certain pixel point in the image; when and , the filter kernel moves one pixel position horizontally relative to a certain pixel in the image.

[0048] By calculating the convolution sum at each pixel position in the image, the image after matched filtering is obtained.

[0049] The Gaussian filter is a linear smoothing filter used to remove noise in the image. It achieves the smoothing effect by convolving the image with the Gaussian function. The Gaussian function has locality in space and can effectively remove high-frequency noise in the image. Applying the Gaussian filter to remove high-frequency noise in the image and smooth the image provides a clearer image basis for feature extraction.

[0050] Assume the template of the Gaussian filter is , then for the input image its image after Gaussian filtering processing has the following formula:

[0051] .

[0052] Where the Gaussian function has the formula .

[0053] is the standard deviation of the Gaussian function, which determines the smoothing degree of the filter, and its value range is between 0.5 and 2.0.

[0054] When When it is small, the Gaussian filter has a low smoothing degree, mainly removing high-frequency noise in the image while retaining more image details; when is large, the Gaussian filter has a high smoothing degree, which can remove more noise in the image, but at the same time, it will also remove more image details. Since this model mainly targets small objects, is taken as 1 to retain more image details while removing high-frequency noise in the image.

[0055] The Laplace filter is an edge detection filter used to detect edges in an image. It achieves edge detection by calculating the second derivative of the image. The Laplace filter can detect the abrupt parts in the image, thereby extracting the edge features of the image. Using the Laplace filter to enhance the edge information in the image helps with subsequent image segmentation and feature extraction. For the input original image , the image after being processed by the Laplace filter is:

[0056] .

[0057] Matched filtering can detect specific features or patterns in an image, thereby extracting important information in the image; Gaussian filtering can remove high-frequency noise in the image, thereby improving the quality and clarity of the image; Laplace filtering can detect the edge features in the image, thereby extracting important information in the image. By combining these three types of filtering for filtering processing, the present invention can comprehensively extract various features and information in the image, thereby improving the effect and accuracy of image processing.

[0058] After the original image is processed by these three types of filtering, the obtained feature maps each have different characteristics and advantages, and can provide rich features and information for subsequent image processing and analysis.

[0059] Define the images after being processed by matched filtering, Gaussian filtering, and Laplace filtering as respectively. Use the interpolation method to adjust the sizes of the images after the three types of filtering so that they have the same size.

[0060] The target image is adjusted to the same size (W, H) using the bilinear interpolation method to obtain better image quality. The bilinear interpolation method can effectively reduce the aliasing and mosaic phenomena at the image edges through two linear interpolations (in the x direction and y direction), generating a smoother result than the nearest neighbor interpolation. And the bilinear interpolation decomposes the two-dimensional interpolation into two one-dimensional linear interpolations, and the computational complexity is , which is significantly lower than that of bicubic interpolation and is more computationally efficient.

[0061] For each pixel in the target image, find the four nearest pixels in the source image, i.e., and estimate the value of the target pixel by the weighted average of these four pixels, as Figure 3 shown.

[0062] Assume that the coordinates of the target pixel in the source image are (x, y), and x and y are not necessarily integers. The four nearest pixels are found by the following four formulas , , , :

[0063] ;

[0064] ;

[0065] ;

[0066] .

[0067] where represents the source image, and the value of the target pixel can be calculated by the following formula:

[0068] .

[0069] where , , denotes rounding down, denotes rounding up.

[0070] Stitch the three filtered images of the same size to obtain a preprocessed image.

[0071] The image after filtering is a single-channel grayscale image. Create an empty three-channel image, copy the single-channel data to each channel to obtain a multi-channel image, and stitch three multi-channel images of the same scale to obtain a preprocessed image.

[0072] Annotate the preprocessed image, and use the annotated image as the training dataset for training the detection model MARP_Net in step 3 below. Among them, in the aneurysm detection task, the manually segmented label values are 0 for the normal blood vessel area, 1 for the aneurysm area, and 2 for the blood vessel boundary area.

[0073] Step 3. Input the preprocessed image into the pre-built multi-scale axial residual projection network, i.e., the detection model MARP_Net, for intracranial aneurysm detection to obtain the intracranial aneurysm detection result.

[0074] AsFigure 4 As shown, the detection model MARP_Net adopts a modular architecture design and consists of four core components: an encoder, a parallel decoder (Partial Decoder), a DWR module, and an ACRE module.

[0075] The encoder is based on an improved Res-UNet framework, using the ResNet-101 deep residual network as the backbone network for multi-level semantic feature extraction. Through the design of the deep residual structure, the network can effectively capture rich image semantic information, especially enhancing the feature representation ability of tiny anatomical structures, laying a foundation for subsequent accurate segmentation. The decoder part adopts a lightweight design concept and selectively fuses dense features at all levels through a partial decoding mechanism. This strategy significantly reduces the computational complexity while maintaining the aneurysm segmentation accuracy, enhancing the practicality and deployment efficiency of the model.

[0076] To further improve the feature expression ability, the network introduces an expandable residual module (Dilated Weighted Residual, DWR). This module realizes the fusion of multi-scale context information through adjustable dilated residual units, effectively solving the segmentation problem caused by the tiny target size and blurred boundaries in intracranial aneurysm medical images. By capturing context information at different scales, the DWR module can more accurately identify and segment the aneurysm area, improving the segmentation accuracy and robustness.

[0077] The Axial Context Relation Encoder (ACRE) module, based on correlation, realizes the adaptive optimization of the feature map by dynamically modeling spatial dependency relationships.

[0078] In the intracranial aneurysm segmentation task, the ACRE module can capture the spatial associations of aneurysms in different axial directions, further enhancing the richness and accuracy of feature representation, thus improving the overall segmentation effect.

[0079] The preprocessed image is fed into the MARP_Net network. The MARP_Net network captures context information at different scales with the help of the DWR module and uses the ACRE module as the axial attention mechanism and residual connection. The DWR module captures context information at different scales through a multi-scale feature pyramid, providing rich feature representations for subsequent segmentation tasks. The ACRE module further enhances the correlation and consistency of features through the modeling of axial spatial dependencies. The combination of the two can capture and express the features of aneurysms more comprehensively at different scales and different axes. The model of the present invention can achieve the collaborative optimization of multi-scale features and axial spatial dependencies through the combined use of the DWR module and the ACRE module, thereby further improving the segmentation effect. This model constructs a segmentation framework suitable for small targets and complex scenes through the multi-scale feature decomposition of the DWR module and the axial relationship modeling of the ACRE encoder. The technical core lies in the collaborative design of multi-scale dilated convolution and axial residual attention, which not only strengthens the perception of local details but also ensures the stability of the deep network through residual connections.

[0080] The following will further elaborate on the network structure and signal processing flow of the MARP_Net constructed in the present invention in conjunction with the attached Figure 4 First, preprocess the original image, extract features from the image processed by the matched filter (MF), Gaussian filter (GF), and Laplace filter (LF), and splice them into a fixed-length feature representation form.

[0081] The network architecture includes an encoder, a parallel decoder PD, a DWR module, and an ACRE module. Among them, there are three DWR modules and three ACRE modules. When the preprocessed image is fed into the detection model MARP_Net, its processing flow is as follows:

[0082] The preprocessed image passes through the encoder, and ordinary convolution is used to extract features. The image after two convolution processes is used as the first-level input layer and input into the first-level DWR module to extract and fuse multi-scale features to obtain the first-level transmission layer.

[0083] This level's transmission layer serves as the next-level input layer and passes through the next-level DWR module again.

[0084] Let the image output after passing through the first-level DWR module be , and use the first-level transmission layer as the second-level input layer. Then the image serves as the second-level input image , and is input into the second-level DWR module to obtain the image .

[0085] The second-level transmission layer is used as the third-level input layer, that is, the image output by the second-level DWR module As the input image of the third pole ,image Input to the third-level DWR module for processing to obtain the image .

[0086] A parallel decoder is used to aggregate the first, second, and third input layers, that is, the image Aggregation is performed, and the calculation formula is: , thus obtaining a global feature map containing rich information .

[0087] The ACRE module consists of an axial attention module and a contextual relation encoder module.

[0088] The axial attention module uses the attention mechanism to set different weights for the foreground and background of the feature map according to the difference in the importance of features in the image, so as to strengthen the foreground and suppress the background.

[0089] The context encoder module mainly targets the result feature map of the network at this level Expand the process and calculate the foreground of the image ,background Feature map after axial attention processing The contextual relationship between them.

[0090] Given that the pixel values ​​of intracranial aneurysms with protruding intracranial blood vessels are significantly different from those of the surrounding background, and the global feature map Only the approximate area of ​​the intracranial aneurysm can be captured. The features output by the DWR module , are processed by the ACRE module at the same level to obtain the corresponding feature maps, which are defined as feature maps .

[0091] The ACRE module can solve the problem that traditional multi-scale medical image segmentation algorithms have poor processing capabilities for lesion edges and details, which can also help the network segment small-area targets.

[0092] The global feature map The output characteristics of the third level Add together to get the third local feature map , the third local feature map The output characteristics of the second-level ACRE module Add together to get the second local feature map .

[0093] The second local feature map Add the output features of the ACRE module at the first level to obtain the first local feature map , and finally process it using the Sigmoid activation function to obtain the intracranial aneurysm image segmentation result.

[0094] The DWR module is a multi-scale feature enhancement module designed for medical image segmentation tasks. Its core idea is to achieve differential capture and adaptive optimization of context information through multi-branch dilated convolution and dynamic weighted residual fusion.

[0095] The existing DWR module adopts a multi-branch structure with a fixed dilation rate. However, the fixed dilation rate cannot adapt to the dynamic requirements of different target scales. The multi-branch parallel calculation increases the number of parameters and computational cost, and the direct splicing of the outputs of different dilation rate branches may lead to insufficient multi-scale feature fusion. Therefore, the present invention improves the traditional DWR module and adopts a dynamic dilation rate AdaDR based on frequency adaptation to change the fixed dilation rate to a dynamically adjusted dilation rate.

[0096] The structure of the improved DWR module is as shown in Figure 5 and its processing flow is as follows:

[0097] First, centered at the position , extract the local window of the input image (for example ), where represents the local window size and represents the number of channels. Perform discrete Fourier transform on it to obtain the frequency domain representation , which is expressed as:

[0098] .

[0099] Among them, represents the complex output array of the DFT, and represent its height and width, represents the coordinates of the feature map , and the normalized frequencies in the height and width dimensions are given by and .

[0100] According to the Nyquist frequency , define the high-frequency set that cannot be captured by the current dilation rate :

[0101] .

[0102] Sum the frequency components in to obtain the high-frequency power:

[0103] 。

[0104] For the input feature map , spatial context features are extracted through a lightweight convolutional layer (1×1 convolution) to generate an intermediate feature map. A convolutional layer for predicting the dilation rate (with parameters , usually a 3×3 convolution) is used to predict the intermediate feature map, and a dilation rate map with the same size as the input feature map is output , and the ReLU activation function is used to ensure that the dilation rate is non - negative:

[0105] 。

[0106] During the training process, the parameters are optimized through the following objective :

[0107] 。

[0108] Among them, are the top 25% pixels with the highest high - frequency power (such as object edges), are the bottom 25% pixels with the lowest high - frequency power (such as background or object centers). For the input feature map , in the AdaDR part, the dilation rate is dynamically adjusted according to the local frequency (frequency distribution within the local window) to balance the receptive field and the effective bandwidth, and the dilated convolution formula is obtained:

[0109] 。

[0110] Among them, is the value at position in the output feature map, is the size of the convolution kernel, is the predefined sampling offset, is the dynamic dilation rate at each position .

[0111] Traditional fixed designs are difficult to adapt to the differences between high - frequency and low - frequency components of the input feature map. The dynamic adjustment of the dilation rate adopted in the present invention uses a smaller dilation rate in the high - frequency regions (such as edges and textures) to retain details, and a larger dilation rate in the low - frequency regions (such as backgrounds) to expand the receptive field, thereby balancing the effective bandwidth and the receptive field.

[0112] Adding to the input image gives the final output result 。

[0113] Such as Figure 6As shown in Figure 1, the ACRE module consists of an axial attention module and a contextual relationship encoder module. The axial attention module adopts a proxy attention mechanism, which efficiently integrates key features and contextual information in the image by introducing a proxy vector, thereby enhancing the model's ability to capture image details, while reducing computational costs and improving the model's flexibility and interpretability.

[0114] Context encoder module, for this level of network The result feature is expanded to calculate the foreground of the image ,background Feature map after axial attention processing By supplementing and enhancing the details of the edge of the feature map, the network's ability to process the edge information of the foreground is improved, thereby improving the network's ability to segment the edge of the image.

[0115] Traditional attention mechanisms usually need to calculate the interactions between all features, which will result in a very large amount of calculation in high-dimensional feature space, and lack effective regularization methods, which can easily lead to model overfitting and poor generalization ability.

[0116] Therefore, the present invention adopts a proxy attention mechanism, by introducing proxy features, first reducing or simplifying the features, and then calculating the attention weights. This can significantly reduce the amount of calculation in high-dimensional feature space, improve the computational efficiency of the model, and calculate the attention weights by proxy features, which can have a certain regularization effect and reduce the risk of overfitting of the model.

[0117] Proxy Attention is a method that optimizes the efficiency of traditional attention calculations by introducing learnable proxy parameters. The core idea is to replace the high-dimensional interactions of original features with a small number of proxy vectors, thereby reducing the computational complexity while maintaining the ability to focus on key features.

[0118] Specifically, the processing flow of the axial attention module is as follows:

[0119] For input features , with the proxy vector , perform similarity calculation:

[0120] .

[0121] in, is the similarity matrix, N is the number of spatial locations, d is the number of channels, k is the number of agents, and d is the feature dimension. In the image segmentation task, k is set to be related to the number of target categories, and the temperature coefficient Used to stabilize gradients;

[0122] The proxy vectors are weighted and aggregated using the similarity matrix S to generate proxy context features , and the formula is as follows:

[0123]

[0124] The original features and the proxy context features are fused to obtain the fused features , in order to retain details and enhance semantic associations:

[0125] .

[0126] Among them, represents layer normalization processing, is the original feature, i.e., the input feature.

[0127] Assume the input image , whose shape is . The input feature is decomposed into two sub-features, which are decomposed along the height axis and the width axis respectively. Then, the feature is obtained through proxy attention, and the formula is as follows:

[0128] .

[0129] Among them, represents the batch size, represents the number of channels, represents the height, represents the width, represents decomposition along the height axis, represents decomposition along the width axis.

[0130] The intermediate output result of the deeper layer in the MARP_Net network , through processing, the foreground and the background can be obtained:

[0131] ;

[0132] .

[0133] Then, the feature is fused with the foreground and the background to supplement the foreground and edge information of the feature and obtain the processed feature , and the formula is expressed as follows:

[0134] ;

[0135] ;

[0136] 。

[0137] By means of the improved ACRE module, the problem that traditional multi-scale medical image segmentation algorithms have weak processing capabilities for lesion edges and details can be solved, so as to help the network segment small-area targets.

[0138] Intracranial aneurysms usually account for a small part of the images containing them, and the incorrect use of the loss function will lead to the problem of class imbalance. The cross-entropy loss function is used to test the similarity between the prediction results and the results obtained using the manually segmented masks. It is easy to make the loss reach a local minimum, which causes the model to focus on the background area during training and makes it difficult to accurately predict the lesion area. The Dice loss function is suitable for solving the problem of class imbalance.

[0139] In the designed MARP-Net model, the calculation of the loss function can adopt a combined loss function, combining the Dice loss and the binary cross-entropy loss, so as to optimize the segmentation accuracy. The Dice loss function is defined as follows:

[0140] 。

[0141] where n is the number of label classes, is the model prediction value, i.e., the aneurysm detection result, is the manually segmented label value, is a small smoothing constant to prevent the denominator from becoming zero and the gradient from vanishing. In the aneurysm detection task, the manually segmented label value is that the normal blood vessel area is marked as 0, the aneurysm area is marked as 1, and the blood vessel boundary area is marked as 2.

[0142] The convergence rate of the Dice loss function becomes lower in the later stage of training. During the learning process, due to the large data variance, instability is likely to occur, making it difficult to improve the segmentation accuracy of this method.

[0143] Therefore, in this embodiment, a weighted combination of the Dice and entropy loss functions is used, and the corresponding expression is:

[0144] 。

[0145] where α is the fractional weight used to balance the contributions of the Dice and cross-entropy loss functions, on a scale of 0 - 1.

[0146] During training, a previously prepared training dataset (i.e., multiple 3D TOF-MRA images are obtained and preprocessed according to steps 1 and 2 to obtain preprocessed images) is used, and the Dice loss and binary cross entropy loss are combined to optimize the model to optimize the accuracy of segmentation.

[0147] After the model is trained, the trained model is actually deployed.

[0148] After acquiring the 3D TOF-MRA image in real time, first use step 1 to obtain the MIP image, then use step 2 to perform multiple filtering processes on the MIP image, adjust the size to a uniform size, and splice to form a pre-processed image. The pre-processed image is input into the MARP-Net model to perform intracranial aneurysm detection and obtain the intracranial aneurysm detection result.

[0149] The present invention performs maximum density projection operation on the original 3D TOF-MRA image to generate a MIP image, and uses a joint processing process such as matched filtering, Gaussian filtering and Laplace filtering to enhance the vascular structure and edge information, providing a clearer input image for the network architecture, so that the DWR module and ACRE module can extract features and perform segmentation more effectively. Next, the DWR module enhances the segmentation ability of small targets through multi-scale feature extraction, which provides the ACRE module with richer feature information, enabling it to more accurately enhance the foreground and suppress the background through the axial attention and contextual relationship encoder modules. The ACRE module further enhances the model's segmentation ability for small targets by enhancing the edge information of the foreground. This synergy enables the model to more accurately identify boundaries when detecting and segmenting intracranial aneurysms, improving overall performance.

[0150] Example 2

[0151] This embodiment 2 describes an intracranial aneurysm image detection system, which is based on the same inventive concept as the intracranial aneurysm image detection method described in the above embodiment 1.

[0152] An intracranial aneurysm image detection system includes the following modules:

[0153] An MIP image acquisition module is used to perform a maximum density projection operation on the 3D TOF-MRA image to obtain an MIP image;

[0154] An image preprocessing module is used to perform various filtering processes on the generated MIP image, and adjust the size of each image after the filtering process to a uniform size, and then splice the images after the uniform size to obtain a preprocessed image;

[0155] And a prediction module, which is used to input the preprocessed image into a pre-established intracranial aneurysm image detection model based on a multi-scale axial residual projection network for intracranial aneurysm detection to obtain an intracranial aneurysm detection result;

[0156] The detection model realizes multi-scale feature extraction based on the DWR module, and uses the axial context relationship encoder ACRE based on correlation as the axial attention mechanism and residual connection to strengthen the feature map, improve the segmentation ability for small targets and complex backgrounds, achieve deeper feature fusion and information extraction, and enable the detection model to focus on the segmentation of small targets.

[0157] It should be noted that in the intracranial aneurysm image detection system of this Embodiment 2, the implementation processes of the functions and roles of each functional module can be found in detail in the implementation processes of the corresponding steps of the method in the above Embodiment 1, and will not be elaborated here.

[0158] Embodiment 3

[0159] This Embodiment 3 describes a computer device.

[0160] The computer device includes a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it is used to implement the steps of the intracranial aneurysm image detection method in the above Embodiment 1.

[0161] In this embodiment, the computer device is any device or apparatus with data processing capabilities, which will not be elaborated here.

[0162] Embodiment 4

[0163] This Embodiment 4 describes a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it is used to implement the steps of the intracranial aneurysm image detection method in the above Embodiment 1.

[0164] The computer-readable storage medium can be an internal storage unit of any device or apparatus with data processing capabilities, such as a hard disk or memory, or an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device.

[0165] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the guidance of this specification fall within the substantial scope of this specification and should be protected by the present invention.

Claims

1. A method for detecting intracranial aneurysm images, characterized in that: The steps include: Step 1. Perform maximum intensity projection on the 3D TOF-MRA image to obtain a MIP image; Step 2. Apply various filtering processes to the generated MIP image, resize the filtered images to a uniform size, and then splice the images with uniform size to obtain the preprocessed image; Step 3. Input the preprocessed image into a pre-built intracranial aneurysm image detection model based on a multi-scale axial residual projection network to perform intracranial aneurysm detection and obtain intracranial aneurysm detection results; The detection model implements multi-scale feature extraction based on the DWR module, and uses the correlation-based axial contextual encoder, namely the ACRE module, as the axial attention mechanism and residual connection to strengthen the feature map. The DWR module is obtained by improving the traditional fixed expansion rate DWR module, which is changed to adopt a dynamic expansion rate based on frequency adaptation, and the fixed expansion rate is changed to a dynamically adjusted expansion rate to balance the effective bandwidth and the receptive field; The ACRE module consists of an axial attention module and a contextual relationship encoder module; the axial attention module adopts a proxy attention mechanism, which introduces a proxy vector to efficiently integrate key features and contextual information in the image, thereby enhancing the model's ability to capture image details; The context encoder module processes the result features of the network at this level and calculates the contextual relationship between the foreground and background of the image and the feature map after axial attention processing.

2. The intracranial aneurysm image detection method according to claim 1, characterized in that: In step 2, the filtering processing methods include at least matched filtering, Gaussian filtering and Laplace filtering, and the images after each filtering processing are adjusted to a uniform size using bilinear interpolation; The filtered image is a single-channel grayscale image. Create a three-channel empty image, copy the single-channel data to each channel to get a multi-channel image, and splice the three multi-channel images of the same scale to get the preprocessed image.

3. The intracranial aneurysm image detection method according to claim 1, characterized in that: In step 3, the detection model includes an encoder, a parallel decoder, a DWR module and an ACRE module, wherein there are three DWR modules and three ACRE modules, and the processing flow of the detection model is as follows: The preprocessed image passes through the encoder, uses two convolutions to extract features, and inputs the convolution-processed image into the first-level DWR module as the first-level input layer to extract and fuse multi-scale features to obtain the first-level transmission layer; The first-level transmission layer is used as the second-level input layer and input into the second-level DWR module to obtain the second-level transmission layer; the second-level transmission layer is used as the third-level input layer and input into the third-level DWR module for processing; A parallel decoder is used to aggregate the first, second and third input layers to obtain a global feature map; the features output by each level of DWR module are processed by the ACRE module at the corresponding level to obtain the corresponding feature map; The global feature map is added to the output feature of the third-level DWR module to obtain a third local feature map, and the third local feature map is added to the output feature of the second-level ACRE module to obtain a second local feature map; The second local feature map is added to the output feature of the first-level ACRE module to obtain the first local feature map. Finally, the first local feature map is processed using the Sigmoid activation function to obtain the intracranial aneurysm image segmentation result.

4. The intracranial aneurysm image detection method according to claim 1, characterized in that: The processing flow of the improved DWR module is as follows: First, the location Centered on the input image, extract the local window , perform discrete Fourier transform on it and get the frequency domain representation , where s represents the local window size and C represents the number of channels; According to the Nyquist frequency , the definition cannot be affected by the current expansion rate Captured high frequency collection for: ; in, denote the normalized frequencies in height and width dimensions respectively; right The frequency components within are summed to obtain the high frequency power : ; For the input feature map , extract spatial context features through a 1×1 convolution, generate an intermediate feature map, and use a parameter of , a 3×3 convolutional layer with predicted dilation rate, predicts the intermediate feature map, and outputs a dilation rate map of the same size as the input feature map , and the ReLU activation function is used to ensure that the expansion rate is non-negative: ; During training, the parameters are optimized by the following objectives : ; in, The first 25% of pixels with the highest high-frequency power are the edge positions of objects. The last 25% of pixels with the lowest high-frequency power are the background or the center of the object; for the input feature map In the AdaDR part, the dilation rate is dynamically adjusted according to the frequency distribution in the local window to balance the receptive field and the effective bandwidth, and the dilated convolution formula is obtained: ; in, is the position in the output feature map The value of is the size of the convolution kernel, is a predefined sampling offset, Each location Dynamic expansion rate; Dynamically adjust the dilation rate. Use a smaller dilation rate in high-frequency areas where edges and textures are located to preserve details, and a larger dilation rate in low-frequency areas where the background is located to expand the receptive field. With the input image Add to get the final output result , the formula is: .

5. The intracranial aneurysm image detection method according to claim 1, characterized in that: In step 3, after the proxy attention mechanism is introduced, the processing flow of the axial attention module is as follows: For input features , and the proxy vector , perform similarity calculation: ; in, is the similarity matrix, N is the number of spatial locations, d is the number of channels, k is the number of agents, and d is the feature dimension. In the image segmentation task, k is set to be related to the number of target categories, and the temperature coefficient Used to stabilize gradients; Use the similarity matrix S to perform weighted aggregation on the agent vectors to generate agent context features , the formula is as follows: ; Combine the original features with the proxy context features Fusion, get fusion features , to preserve details and enhance semantic associations: ; in, Representation layer normalization processing, is the original feature, i.e., the input feature; Assume that the input image , whose shape is , the input feature Decomposed into two sub-features, respectively along the height Axis and Width The axis is decomposed, and the features are obtained through proxy attention , the formula is as follows: ; in, represents the batch size, Represents the number of channels, Represents height, Represents the width, Indicates along the height The axis is decomposed. Indicates along the width Axis decomposition.

6. An intracranial aneurysm image detection system for implementing the intracranial aneurysm image detection method according to claim 1, characterized in that: The intracranial aneurysm image detection system includes the following modules: An MIP image acquisition module is used to perform a maximum density projection operation on the 3D TOF-MRA image to obtain an MIP image; An image preprocessing module is used to perform various filtering processes on the generated MIP image, and adjust the size of each image after the filtering process to a uniform size, and then splice the images after the uniform size to obtain a preprocessed image; and a prediction module, which is used to input the preprocessed image into a pre-built intracranial aneurysm image detection model based on a multi-scale axial residual projection network to perform intracranial aneurysm detection and obtain intracranial aneurysm detection results; The detection model implements multi-scale feature extraction based on the DWR module, and uses the correlation-based axial contextual relationship encoder ACRE as the axial attention mechanism and residual connection to strengthen the feature map, improve the segmentation ability of small targets and complex backgrounds, achieve deeper feature fusion and information extraction, and enable the detection model to focus on the segmentation of small targets.

7. A computer device comprising a memory and one or more processors; an executable code is stored in the memory; characterized in that: When the processor executes the executable code, it is used to implement the steps of the intracranial aneurysm image detection method described in any one of claims 1 to 5.

8. A computer-readable storage medium having a program stored thereon; characterized in that: When the program is executed by a processor, it is used to implement the steps of the intracranial aneurysm image detection method described in any one of claims 1 to 5 above.

Citation Information

Patent Citations

  • Thyroid nodule segmentation method fusing global reasoning and MLP architecture

    CN115018780A

  • Intracranial aneurysm image detection method, system, equipment and medium

    CN117392137A