Deep residual network-based pulmonary embolism image segmentation method and device
Through the multi-layer feature extraction, fusion and attention-guiding mechanism of the deep residual network, the problem of insufficient feature extraction in pulmonary embolism image segmentation is solved, and the recognition of pulmonary embolism region with higher accuracy and robustness is achieved, and the accuracy of clinical diagnosis is improved.
Patent Information
- Application Number
- CN202510534774.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
When the existing pulmonary embolism image segmentation method deals with complex features such as small targets, blurred edges and diverse morphology, there are problems of insufficient feature extraction and information loss, resulting in inaccurate segmentation results and difficult to meet clinical diagnosis needs.
The pulmonary embolism image segmentation method based on deep residual network is adopted, and the multi-layer feature extraction, feature fusion and attention guidance mechanism is combined with the comprehensive loss function to optimize the model parameters to improve the feature expression ability and small-object recognition ability.
It improves the accuracy and robustness of pulmonary embolism image segmentation, can more accurately identify pulmonary embolism areas in complex morphology, reduce misdiagnosis and misdiagnosis, and provide reliable auxiliary diagnostic tools.
Smart Images

Figure CN120451542A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning technology, and in particular to a method and device for pulmonary embolism image segmentation based on a deep residual network. Background Art
[0002] Pulmonary embolism (PE) is a common and potentially fatal lung disease, primarily caused by a blood clot blocking the pulmonary artery or its branches, often caused by a deep vein thrombosis that breaks away and enters the pulmonary circulation. Clinically, early diagnosis and timely treatment of PE are crucial for improving patient survival. However, because PE symptoms vary and can be easily confused with other conditions, accurate diagnosis based on clinical manifestations alone is difficult, and imaging is often required for auxiliary diagnosis.
[0003] Computed tomography angiography (CTPA) of the pulmonary arteries is currently the most commonly used and reliable imaging technique for detecting pulmonary embolism. It can visually demonstrate the patency of the pulmonary artery and its branches in three-dimensional images. When interpreting CTPA images, physicians typically need to review numerous images frame by frame to locate possible embolic areas and determine their morphology, size, and distribution. This process is not only time-consuming and laborious, but also susceptible to the physician's subjective experience, leading to the risk of misdiagnosis and missed diagnosis.
[0004] In recent years, artificial intelligence technologies, especially deep learning-based image segmentation algorithms, have been widely studied and applied in the field of medical image processing. Automatic segmentation of pulmonary embolism regions by training convolutional neural networks (CNNs) can improve diagnostic efficiency while maintaining accuracy. However, existing image segmentation methods still face many challenges when processing pulmonary embolism images. On the one hand, pulmonary embolism regions often exhibit complex features such as small targets, blurred edges, and diverse morphologies, resulting in traditional convolutional networks being unable to adequately learn small samples and detailed information. On the other hand, although deep network models have strong feature extraction capabilities, they are prone to information loss and gradient dissipation, which affect the accuracy of segmentation results.
[0005] Therefore, there is an urgent need for a self-learning network model that integrates feature aggregation mechanism and attention mechanism to enhance the feature expression ability and small target recognition ability of pulmonary embolism images, thereby improving the accuracy and robustness of automatic segmentation and providing a more reliable auxiliary diagnostic tool for clinicians. Summary of the Invention
[0006] Based on this, it is necessary to provide a pulmonary embolism image segmentation method and device based on deep residual network to address the above technical problems.
[0007] In a first aspect, the present application provides a pulmonary embolism image segmentation method based on a deep residual network, comprising:
[0008] Inputting the training set images with the pulmonary embolism marked areas into a preset deep residual network to extract multi-layer image features of the training set images;
[0009] performing feature fusion on the multi-layer image features to generate a segmented image with a pulmonary embolism prediction area;
[0010] Decoding and reconstructing the multi-layer image features based on an attention guidance mechanism to generate a multi-layer attention map with a pulmonary embolism prediction area;
[0011] Constructing a comprehensive loss function according to the segmented image, the multi-layer attention map, and the training set image, and optimizing the model parameters of the deep residual network based on the comprehensive loss function;
[0012] The optimized deep residual network is used to segment and annotate the pulmonary embolism area in the segmented image.
[0013] In one embodiment, extracting multi-layer image features of the training set images includes:
[0014] The low-level detail features and high-level semantic features of the training set images are extracted through the deep convolution kernel of the deep residual network.
[0015] In one embodiment, performing feature fusion on the multi-layer image features to generate a segmented image with a predicted pulmonary embolism area includes:
[0016] Performing feature fusion on the multi-layer image features to generate a fused unified feature map;
[0017] Extracting row / column vectors from the unified feature map, and classifying the row / column vectors;
[0018] Based on the classification results of the row / column vectors, a segmented image with a predicted pulmonary embolism area is generated.
[0019] In one embodiment, the decoding and reconstruction of the multi-layer image features based on the attention guidance mechanism to generate a multi-layer attention map with a pulmonary embolism prediction area includes:
[0020] Upsampling the multi-layer image features and splicing them with the image features of the corresponding layers to obtain multi-layer spliced image features;
[0021] Based on the attention guidance mechanism, the attention map of the spliced image features of each layer is extracted, and the attention map of each layer is upsampled to the image size of the next decoding layer.
[0022] In one embodiment, after uniformly upsampling the attention maps of each layer to the image size of the next decoding layer and the size of the training set image, the method further includes:
[0023] The uniformly upsampled attention map is combined with the next decoding layer image to construct the loss function.
[0024] In one embodiment, a comprehensive loss function is constructed based on the segmented image, the multi-layer attention map, and the training set image, including:
[0025] Constructing a first loss function according to the segmented image and the training set image;
[0026] Constructing a second loss function based on the multi-layer attention map and the training set images;
[0027] A comprehensive loss function is constructed by the first loss function and the second loss function.
[0028] In one embodiment, after optimizing the model parameters of the deep residual network based on the comprehensive loss function, the method further includes:
[0029] The validation set images with pulmonary embolism annotated areas are input into the optimized deep residual network to verify the degree of optimization of the model parameters of the deep residual network.
[0030] In a second aspect, the present application also provides a pulmonary embolism image segmentation device based on a deep residual network, comprising:
[0031] A feature extraction module is used to input the training set images with pulmonary embolism annotated areas into a preset deep residual network to extract multi-layer image features of the training set images;
[0032] a feature fusion module, configured to fuse the multi-layer image features to generate a segmented image with a pulmonary embolism prediction area;
[0033] an attention self-learning module, configured to decode and reconstruct the multi-layer image features based on an attention guidance mechanism to generate a multi-layer attention map with a pulmonary embolism prediction area;
[0034] A parameter optimization module, configured to construct a comprehensive loss function based on the segmented image, the multi-layer attention map, and the training set image, and optimize the model parameters of the deep residual network based on the comprehensive loss function;
[0035] The segmentation and annotation module is used to use the optimized deep residual network to segment and annotate the pulmonary embolism area in the segmented image.
[0036] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method steps described in the first aspect when executing the computer program.
[0037] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method steps described in the first aspect.
[0038] In a fifth aspect, the present application further provides a computer program product, which includes a computer program that, when executed by a processor, implements the method steps described in the first aspect.
[0039] The pulmonary embolism image segmentation method based on a deep residual network (DRN) disclosed in this application extracts multi-layer image features by feeding training set images with annotated pulmonary embolism regions into a pre-set DRN. This method fully preserves the information expression capability of different semantic levels, facilitating accurate identification of the complex morphology of pulmonary embolism regions. Furthermore, by fusing the multi-layer image features, the method not only effectively supplements the semantic information in high-level features with the edge details in low-level features, but also improves the spatial localization accuracy and structural integrity of the segmented image. Furthermore, by incorporating an attention-guided mechanism to decode and reconstruct the multi-layer image features, the method significantly strengthens the network's response to key regions, highlighting the significance of pulmonary embolism regions, and thus improving the accuracy and reliability of the segmentation results. Furthermore, a comprehensive loss function is constructed based on the segmented images, the multi-layer attention map, and the training set images. During the joint optimization process, the model focuses on both semantic segmentation performance and attention learning effects, improving network training efficiency and convergence. Finally, by optimizing the model parameters of the DRN, the optimized network is capable of effectively identifying and accurately annotating pulmonary embolism regions in the segmented images, demonstrating strong generalization and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 1 is a flow chart of a pulmonary embolism image segmentation method based on a deep residual network in one embodiment;
[0041] Figure 2 1 is a schematic diagram of the encoding and decoding process of pulmonary embolism images based on a deep residual network in one embodiment;
[0042] Figure 3 Schematic diagram of the module structure of a pulmonary embolism image segmentation device in one embodiment;
[0043] Figure 4 Schematic diagram of the module structure of a pulmonary embolism image segmentation device in one embodiment;
[0044] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0046] The present invention provides a method for pulmonary embolism image segmentation based on feature aggregation attention self-learning. This method can be widely used in medical image-assisted diagnosis systems, and is particularly suitable for computer-assisted detection and segmentation of lung diseases such as pulmonary embolism (PE). In modern medical diagnosis and treatment, pulmonary embolism, as an acute pulmonary vascular disease with rapid onset, rapid progression, and high mortality rate, urgently requires high-precision and high-efficiency image processing technology to assist doctors in early identification and treatment intervention.
[0047] In real-world applications, medical imaging platforms typically receive large volumes of chest CT images uploaded by hospital radiology departments. The platform's servers must use automated algorithms to segment and analyze the lung structures within these images, identifying potential lesion areas and generating diagnostic assistance results. These images are high-resolution, feature complex distributions, and contain diverse morphologies of embolic lesions. Traditional segmentation algorithms have significant limitations in boundary recognition, target focus, and generalization capabilities, hindering their clinical effectiveness.
[0048] The image segmentation method provided in this application is based on feature aggregation and attention self-learning mechanism. By constructing a multi-scale decoding network structure and introducing a spatial attention map to guide the optimization of key areas, it can effectively improve the recognition accuracy and segmentation accuracy of pulmonary embolism lesion areas. This method is not only suitable for pulmonary embolism auxiliary detection systems, but can also be extended to other scenarios that require high-precision lesion positioning, such as lung nodule screening, pneumonia area identification, and tumor image analysis. As the core computing node, the platform server can deploy this segmentation method to realize automatic image processing, lesion report generation, and multi-source image result fusion, thereby improving the intelligence level of the medical imaging platform, reducing the burden of repetitive work on doctors, and helping to achieve more efficient and accurate auxiliary diagnosis and treatment services.
[0049] The pulmonary embolism image segmentation method based on deep residual network disclosed in this application can be as follows Figure 1 As shown, the following steps are included:
[0050] Step 101: Input the training set images with the pulmonary embolism marked areas into a preset deep residual network to extract multi-layer image features of the training set images.
[0051] During implementation, the system can first select medical images containing annotated areas of pulmonary embolism to construct a training set, which can provide clear boundary information of the lesion area and provide reliable supervision signals for subsequent model training. The training set images can be used to train the deep learning model so that the deep learning model can learn to extract the potential features of the pulmonary embolism area in the training set images, so as to facilitate the use of the deep learning model to identify the pulmonary embolism area in the medical image. Specifically, the deep learning model can adopt a deep residual network model. The deep residual network is a deep neural network structure widely used in computer vision tasks, which can effectively solve the common gradient disappearance or degradation problems in traditional deep convolutional neural networks. By introducing residual connections, the network structure can improve the trainability and feature expression ability of the model while maintaining the continuity of information transmission. Specifically, the residual structure allows the input information of a certain layer to be directly transmitted to the subsequent layer through a short-circuit path, which helps the deep network converge faster and slows down the problem of gradual attenuation of information in multi-layer networks.
[0052] By feeding the training set images into a preset deep residual network, the network can perform multi-level feature encoding on the training set images to obtain multi-layer image features with rich representation capabilities, thus providing important basic support for subsequent pulmonary embolism region segmentation and boundary extraction. In this way, by introducing a deep residual network to perform multi-layer feature extraction on the training images, not only can the integrity and hierarchy of feature expression be improved, but also the model's ability to adapt to morphological changes in the lesion area can be enhanced, thereby improving the overall performance and generalization ability of the pulmonary embolism image segmentation model.
[0053] It is worth mentioning that in one specific embodiment, in order to improve the accuracy and robustness of the image segmentation model in identifying pulmonary embolism areas, before training the deep residual network, the medical image data in the public dataset can first be uniformly preprocessed, region extracted, and data segmented. The specific steps are as follows:
[0054] Step 1: The collected pulmonary embolism medical imaging data (such as CT pulmonary angiography images) undergoes unified data preprocessing to ensure the stability of subsequent model training and the consistency of image features. Specifically, the original image can first be adjusted in terms of window width and window position to enhance the contrast of the pulmonary artery region, making the embolic region easier to identify; the image is then cropped to remove irrelevant edge regions and retain key anatomical structures relevant to diagnosis; the image is then grayscale normalized to unify the image grayscale value range to [0,1], reducing the impact of different equipment or shooting parameters on image quality; the image is also resized and resampled by interpolation to ensure consistent resolution for all images; and denoising can also be applied to the image, such as using Gaussian filtering and non-local mean filtering, to suppress noise interference and retain key texture details. Through the above processing steps, the data input to the model can be ensured to have high clarity and consistency, which is conducive to improving the accuracy of subsequent feature extraction and segmentation.
[0055] Step 2: After completing the data preprocessing for the image, the discriminative information area in the image can be further extracted. Specifically, the anatomical area where the pulmonary artery is located can be extracted from the entire CT image as the region of interest (ROI) for subsequent learning and positioning of the pulmonary embolism area. Among them, the method of extracting ROI may include but is not limited to anatomical template-based matching, density-based region growing, or preliminary screening based on traditional image segmentation methods (such as Otsu, threshold segmentation, etc.), or the use of a pre-trained model for pulmonary artery area detection. Through this step, the interference of redundant background information on the segmentation model can be effectively reduced, and the training efficiency can be improved.
[0056] In one embodiment, the multi-layer image feature extraction process may be specifically as follows: low-level detail features and high-level semantic features of the training set images are extracted through deep convolution kernels of a deep residual network.
[0057] In practice, in order to effectively utilize key information at different scales in the training set images, the deep residual network can be designed to contain multiple convolutional modules, covering convolution kernels with different receptive fields, to achieve hierarchical extraction of image features. Multi-layer image features can include at least low-level detail features and high-level semantic features. Among them, low-level detail features are mainly extracted by shallow convolutional layers, focusing on subtle but important information in the image, such as blood vessel edges, lung texture changes, local contrast, etc. These features mainly retain spatial structure and detailed morphology, and play a key role in identifying the boundaries of pulmonary embolism; high-level semantic features are extracted by deeper convolutional layers. These layers have larger receptive fields and can capture image context and abstract semantic information in a larger range, such as the shape pattern and relative position of the lesion area, which helps the network accurately judge potential embolic areas.
[0058] Furthermore, in order to improve the fusion effect of the above-mentioned different-level features, the residual network structure also introduces jump connections, allowing shallow and deep feature information to cross-circulate in the network, thereby enhancing the overall feature expression integrity.
[0059] It is understandable that in image segmentation tasks, low-level detail features help locate edges and texture structures, while high-level semantic features help discern target categories and global context. The hierarchical structure and jump connection mechanism of the deep residual network are more suitable for the "blurred boundaries + complex morphology" characteristics of medical images, and can effectively solve the problem of shallow features being too local and deep features lacking detail. In actual operation, typical residual network structures such as ResNet, ResNeXt, and HRNet can be selected, and parameters and depth can be adjusted according to the characteristic distribution of pulmonary embolism images, so that the network has a stronger adaptability to interference factors such as irregular morphology and low contrast in lung images.
[0060] Furthermore, to effectively train and evaluate the model's performance, the image data extracted from the ROI can be partitioned to produce the required training and validation sets. The training set can be used for iterative learning of model parameters, while the validation set can be used for accuracy monitoring and early stopping during model training. The specific partitioning ratio can be set based on the size of the dataset, such as an 8:2 or 7:3 split.
[0061] Step 102 : performing feature fusion on the multi-layer image features to generate a segmented image with a predicted pulmonary embolism area.
[0062] In this embodiment, the system can perform feature fusion operations based on multi-layer image features extracted through a deep residual network. It is not difficult to understand that multi-layer image features cover different semantic levels and spatial details. If a single-layer feature is directly used to identify pulmonary embolism areas, the accuracy of model recognition may be reduced due to incomplete information or expression bias. To this end, a feature fusion strategy is used to integrate features from different levels, which helps to fully utilize the information contained in various features and improve the robustness and accuracy of image segmentation. Specifically, feature fusion is a method of merging features from multiple sources or multiple levels in spatial or channel dimensions. In this embodiment, the fusion method can include but is not limited to upsampling, feature splicing, attention weighting, pixel-by-pixel weighting and other implementation paths, which can be flexibly configured based on the actual training results. Through this feature fusion operation, a segmented image can be finally generated, in which each pixel or pixel block can carry a predicted label for pulmonary embolism, indicating whether the area is likely to be pulmonary embolism, thereby intuitively showing the lesion area identified by the model.
[0063] In this way, the deep residual network model can not only use detailed features to locate the edge of the lesion, but also use semantic features to determine the target category, thereby improving the accuracy of the pulmonary embolism prediction area. The segmented image generated by feature fusion has medical auxiliary diagnosis value and provides a visual basis for doctors' judgment.
[0064] In one embodiment, the above-mentioned feature fusion process for generating a segmented image can be specifically as follows: multi-layer image features are fused to generate a unified fused feature map; row / column vectors in the unified feature map are extracted and the row / column vectors are classified; and based on the classification results of the row / column vectors, a segmented image with a predicted pulmonary embolism area is generated.
[0065] In implementation, in order to further improve the recognition accuracy of pulmonary embolism areas, the feature fusion process not only stops at the coarse fusion of image features, but also further generates a unified fusion feature map and introduces a vector classification mechanism to enhance the model's ability to understand the distribution direction and location pattern of embolism.
[0066] Specifically, multiple layers of feature maps can be first aligned in spatial dimensions (e.g., using upsampling or interpolation methods) and then fused into a feature map of uniform scale and dimension through feature stacking or channel weighting. This fused feature map integrates feature expressions from shallow and deep layers and serves as a comprehensive basis for segmentation and discrimination. To more finely analyze the spatial distribution characteristics of pulmonary embolism regions in images, the system can further extract row / column vectors from the fused unified feature map. Row / column vectors refer to the process of dividing the feature map horizontally or vertically into several row vectors or column vectors, each of which represents the feature combination information within the row / column region. Subsequently, a classification network can be introduced to discriminate these row / column vectors one by one to determine whether a pulmonary embolism region exists in that row / column. Classification methods can use fully connected layers, convolutional classifiers, attention modules, and other methods. The specific implementation can be flexibly configured based on computing resources and model complexity. Ultimately, the row / column vector classification results can be reconstructed into a two-dimensional heat map or binary segmentation map, which is then used as the output of the pulmonary embolism prediction region. In this way, by converting the fused feature map into row / column vectors and performing classification processing, on the one hand, the spatial pattern can be converted into sequence features for discrimination, which is conducive to improving the model's capture accuracy of pulmonary embolism areas; on the other hand, it is also convenient for the model to identify "slender" and "along the blood vessels" pulmonary embolism structures, significantly optimizing the recognition ability of areas with blurred edges and complex morphology.
[0067] It is worth mentioning that the above method converts the original image segmentation problem (H×W) in two-dimensional space into a one-dimensional vector classification problem to simplify the learning task and improve positioning accuracy. The fused unified feature map is transformed in dimension to extract the feature vectors on each row or column (i.e., row vectors / column vectors), and these row / column vectors are independently classified through a classification submodule (such as a fully connected layer or a convolution classifier). The classification result of each row / column vector can be regarded as a predicted label at the corresponding spatial position. This method essentially converts the image segmentation task into multiple parallel row / column classification tasks, thereby significantly improving the ability to recognize fine-grained boundaries during the training phase.
[0068] Here, feature aggregation is only used during the training phase. By deconstructing the complex image segmentation learning process into multiple low-dimensional classification problems, it improves the network's efficiency in learning spatial context. During the testing phase, this module is not explicitly used, but its learning results are reflected in the parameter adjustments of the backbone network, indirectly improving the accuracy of image segmentation.
[0069] Step 103: decode and reconstruct the multi-layer image features based on the attention guidance mechanism to generate a multi-layer attention map with the pulmonary embolism prediction area.
[0070] In implementation, in order to restore the fused high-dimensional deep image features to a spatial image representation consistent with the resolution of the original image, the system can use a decoder structure to perform layer-by-layer decoding and reconstruction operations on the fused feature map, and introduce an attention guidance mechanism to enhance the network's ability to focus on key areas (i.e., pulmonary embolism areas). Specifically, the decoder structure can be composed of an upsampling operation and a transposed convolution module. The upsampling operation can be implemented through bilinear interpolation, that is, while maintaining the continuity of the image feature structure, the spatial dimension of the feature map is expanded, such as upsampling the 64×64 feature map to 128×128 to gradually restore the original image scale; the transposed convolution can be used to structurally enlarge and reconstruct the image through learnable parameters. It is generally composed of "1×1 convolution + 3×3 transposed convolution + 1×1 convolution". Among them, the first 1×1 convolution is mainly used for channel number adjustment, the 3×3 transposed convolution is used to perform feature map scale enlargement, and the second 1×1 convolution is used for feature fusion and redundancy compression. In addition, in each layer of decoding, the decoder can concatenate the upsampled feature map of the current layer with the feature map of the corresponding layer in the encoder to compensate for the loss of high-frequency information (such as edge details) during the encoding process.
[0071] Furthermore, after the transposed convolution operation of each layer of the above decoder module is completed, the attention mechanism can be introduced to perform spatial attention modeling. The main process can compress the three-dimensional feature map (length × width × channel) output by the convolution into a two-dimensional length × width attention map. In this process, compression can be performed along the channel dimension through methods such as average pooling and maximum pooling to extract the significant response of each channel in the spatial dimension. The above two-dimensional attention map can reflect the importance of each spatial position. Furthermore, in order to facilitate subsequent fusion and judgment, the attention map can be input into the Sigmoid activation function for normalization processing, so that the attention value is limited to between 0 and 1, thereby highlighting the high response area (possible area of pulmonary embolism) and suppressing irrelevant background areas. Among them, the Sigmod activation function can be x is the feature map.
[0072] Based on the above content, the generation process of multi-layer attention maps can include: upsampling multi-layer image features, splicing them with the image features of the corresponding layers to obtain multi-layer spliced image features; extracting the attention map of each layer of spliced image features based on the attention guidance mechanism, and upsampling each layer of attention map to the image size of the next decoding layer.
[0073] It can be understood that when decoding each layer, the upsampled decoded feature map will be spliced with the output of the encoder at the same layer, which not only improves the consistency of semantic and detail features, but also strengthens the complementarity of contextual information. The spliced image features of each layer are sent to the attention guidance module for processing separately, and the attention map of the layer is extracted. The attention map can respectively characterize the pulmonary embolism attention area at different scales and provide support for multi-scale perception. Taking into account the inconsistency of the spatial size of each layer's attention map, in order to ensure the accuracy of subsequent cross-layer supervision or alignment processing, the attention map of each layer is upsampled so that its size is consistent with the spatial size of the image features of the decoder of the next layer. Bilinear interpolation can be used here to expand the size to ensure the comparability and fusion ability between multi-scale attention maps.
[0074] Furthermore, based on the spatial alignment between the generated attention maps at each layer and the image features of the decoder layer at the next layer, a hierarchical supervision loss function can be constructed to improve the effectiveness of the attention mechanism during training. Specifically, the uniformly upsampled attention maps are compared pixel by pixel with the corresponding decoder image features at the next layer. By calculating the difference between the two, a loss function can be constructed to measure the spatial guidance ability of the attention map.
[0075] Step 104: construct a comprehensive loss function based on the segmented image, the multi-layer attention map, and the training set image, and optimize the model parameters of the deep residual network based on the comprehensive loss function.
[0076] In implementation, in order to improve the accuracy and robustness of pulmonary embolism image segmentation, a multi-source supervision mechanism can be introduced into the training of the deep residual network, and a comprehensive loss function can be constructed to comprehensively constrain the model performance and improve the prediction accuracy and convergence stability.
[0077] The comprehensive loss function may include a first loss function and a second loss function, the first loss function may be constructed based on the segmented image and the training set image, and the second loss function may be constructed based on the multi-layer attention map and the training set image.
[0078] Specifically, the first loss function can be constructed by converting the image segmentation task into a row-by-row classification task to obtain the segmented image, where each row of the image is treated as an independent classification result. It is then compared with the annotation results of the corresponding row in the training set, and the cross-entropy loss or Dice loss is calculated as the first loss function, which can be denoted as Lossa. This utilization of the row characteristics of the image structure enhances the segmentation model's understanding of linear spatial structure, and is particularly suitable for medical image features in which pulmonary embolism areas present a mixed distribution of linear and blocky features.
[0079] For the second loss function, in order to ensure that the area extracted by the attention guidance mechanism is indeed focused on the key lesion area, supervision constraints can be constructed for the multi-layer attention map output by the network. Since the attention map is a multi-layer output (such as 3 layers), the attention map generated by each layer is different from the labeled image in the training set in terms of spatial size. Bilinear upsampling can be used to uniformly adjust the attention maps of each layer to the same size as the original Figure 1 The upsampled attention map is compared with the training set image element by element, and the mean square error loss function is used to measure the difference between the two. The losses corresponding to the attention maps of each layer are then accumulated in sequence to form the overall attention loss term, that is, the second loss function Lossb = attLoss1 + attLoss2 + attLoss3.
[0080] Based on the above processing, the model can not only be optimized at the final segmentation result level, but also perceive the target area in advance during the attention generation stage of the intermediate layer, thereby significantly improving the model's sensitivity in identifying small targets (such as pulmonary embolism).
[0081] Finally, the above two loss functions constitute the dual guidance signals for model training. The first loss function Lossa and the second loss function Lossb can be weightedly fused or simply added to construct the final comprehensive loss function Loss = αLossa + βLossb.
[0082] During the training process, with the comprehensive loss function as the target, the back-propagation algorithm is used to continuously optimize the model parameters at all levels of the deep residual network until the comprehensive loss function no longer converges, thereby achieving end-to-end learning of the model.
[0083] Please refer to Figure 2 As shown in Figure 1, the input layer is a 512×512 training set image. After the window width and window position are adjusted, the deep residual network is input. In the encoding stage, the input image is downsampled 5 times (maximum pooling or strided convolution), 256×256×4→128×128×64→64×64×128→32×32×256→16×16×512. The size of the feature map is halved step by step, and the number of channels increases from 4 to 512. The spatial resolution is compressed step by step, and the number of channels is expanded to extract multi-scale features. At the same time, each After upsampling to 512×512, features are fused. In the decoding stage, transposed convolutions are used for progressive upsampling: 16×16×512 → 32×32×256 → 64×64×128 → 128×128×64 → 256×256×4 → 512×512×1. Each decoder layer receives the output feature map of the previous layer (e.g., 16×16×512), fuses the encoder features of the same scale (e.g., 16×16×512 skip connections), and adjusts the number of channels through 1×1 convolutions. Under the attention mechanism, attention losses Lossb of adjacent layers are calculated through bilinear interpolation upsampling, including attLoss1 (16×16 layer), attLoss2 (32×32 layer), and attLoss3 (64×64 layer).
[0084] Step 105: Use the optimized deep residual network to segment and annotate the pulmonary embolism area in the image to be segmented.
[0085] In another embodiment, after optimizing the model parameters of the deep residual network based on the comprehensive loss function, validation set images with annotated pulmonary embolism areas can be input into the optimized deep residual network to verify the degree of optimization of the model parameters of the deep residual network.
[0086] In implementation, after the model parameters are optimized based on the comprehensive loss function, in order to ensure that the training results have good generalization ability and clinical adaptability, a validation set image evaluation mechanism can be further introduced to verify the performance of the deep residual network.
[0087] Specifically, based on the aforementioned validation set, the validation set image can be input into the optimized deep residual network so that the deep residual network outputs the corresponding pulmonary embolism prediction area image (i.e., the segmentation result image). Afterwards, the segmentation result image can be compared with the annotation content in the validation set image to calculate performance evaluation indicators such as accuracy, recall rate, F1 score, IoU (Intersection over Union), so as to judge the performance of the current training model on non-training samples, that is, to verify whether the optimized model parameters have good generalizability and practical applicability. In this way, by introducing the validation set, not only the integrity of the model development process can be improved, but also it is ensured that the final model has the ability to identify pulmonary embolism areas in real application scenarios, providing reliable support for clinical medicine auxiliary diagnosis.
[0088] The pulmonary embolism image segmentation method based on a deep residual network (DRN) disclosed in this application extracts multi-layer image features by feeding training set images with annotated pulmonary embolism regions into a pre-set DRN. This method fully preserves the information expression capability of different semantic levels, facilitating accurate identification of the complex morphology of pulmonary embolism regions. Furthermore, by fusing the multi-layer image features, the method not only effectively supplements the semantic information in high-level features with the edge details in low-level features, but also improves the spatial localization accuracy and structural integrity of the segmented image. Furthermore, by incorporating an attention-guided mechanism to decode and reconstruct the multi-layer image features, the method significantly strengthens the network's response to key regions, highlighting the significance of pulmonary embolism regions, and thus improving the accuracy and reliability of the segmentation results. Furthermore, a comprehensive loss function is constructed based on the segmented images, the multi-layer attention map, and the training set images. During the joint optimization process, the model focuses on both semantic segmentation performance and attention learning effects, improving network training efficiency and convergence. Finally, by optimizing the model parameters of the DRN, the optimized network is capable of effectively identifying and accurately annotating pulmonary embolism regions in the segmented images, demonstrating strong generalization and practical application value.
[0089] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0090] Based on the same inventive concept, Figure 3 The embodiment of the present application further provides a pulmonary embolism image segmentation device 300 based on a deep residual network, comprising:
[0091] A feature extraction module 301 is configured to input a training set image with a pulmonary embolism annotated area into a preset deep residual network to extract multi-layer image features of the training set image;
[0092] A feature fusion module 302 is configured to fuse the multi-layer image features to generate a segmented image with a predicted pulmonary embolism area;
[0093] an attention self-learning module 303 for decoding and reconstructing the multi-layer image features based on an attention guidance mechanism to generate a multi-layer attention map with a pulmonary embolism prediction area;
[0094] A parameter optimization module 304 is configured to construct a comprehensive loss function based on the segmented image, the multi-layer attention map, and the training set image, and optimize the model parameters of the deep residual network based on the comprehensive loss function;
[0095] The segmentation and annotation module 305 is used to segment and annotate the pulmonary embolism area of the image to be segmented using the optimized deep residual network.
[0096] In one embodiment, the feature extraction module 301 is specifically configured to:
[0097] The low-level detail features and high-level semantic features of the training set images are extracted through the deep convolution kernel of the deep residual network.
[0098] In one embodiment, the feature fusion module 302 is specifically configured to:
[0099] Performing feature fusion on the multi-layer image features to generate a fused unified feature map;
[0100] Extracting row / column vectors from the unified feature map, and classifying the row / column vectors;
[0101] Based on the classification results of the row / column vectors, a segmented image with a predicted pulmonary embolism area is generated.
[0102] In one embodiment, the attention self-learning module 303 is specifically configured to:
[0103] Upsampling the multi-layer image features and splicing them with the image features of the corresponding layers to obtain multi-layer spliced image features;
[0104] Based on the attention guidance mechanism, the attention map of the spliced image features of each layer is extracted, and the attention map of each layer is upsampled to the image size of the next decoding layer.
[0105] In one embodiment, the attention self-learning module 303 is further configured to:
[0106] The uniformly upsampled attention map is combined with the next decoding layer image to construct the loss function.
[0107] In one embodiment, the parameter optimization module 304 is specifically configured to:
[0108] Constructing a first loss function according to the segmented image and the training set image;
[0109] Constructing a second loss function based on the multi-layer attention map and the training set images;
[0110] A comprehensive loss function is constructed by the first loss function and the second loss function.
[0111] In one embodiment, Figure 4 As shown, the pulmonary embolism image segmentation device 300 further includes:
[0112] The model verification module 306 is used to input the verification set images with the pulmonary embolism annotated areas into the optimized deep residual network to verify the degree of optimization of the model parameters of the deep residual network.
[0113] In one embodiment, a computer device is provided, whose internal structure diagram can be as follows: Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange data between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0114] Those skilled in the art will understand that Figure 5The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0115] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0116] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0117] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0118] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. For purposes of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The processors involved in the various embodiments provided herein may be general-purpose processors, central processing units (CPUs), graphics processors (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like, without limitation thereto.
[0119] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0120] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A pulmonary embolism image segmentation method based on deep residual network, characterized in that: include: Inputting the training set images with the pulmonary embolism marked areas into a preset deep residual network to extract multi-layer image features of the training set images; performing feature fusion on the multi-layer image features to generate a segmented image with a pulmonary embolism prediction area; Decoding and reconstructing the multi-layer image features based on an attention guidance mechanism to generate a multi-layer attention map with a pulmonary embolism prediction area; Constructing a comprehensive loss function according to the segmented image, the multi-layer attention map, and the training set image, and optimizing the model parameters of the deep residual network based on the comprehensive loss function; The optimized deep residual network is used to segment and annotate the pulmonary embolism area in the segmented image.
2. The method according to claim 1, characterized in that The extracting multi-layer image features of the training set image includes: The low-level detail features and high-level semantic features of the training set images are extracted through the deep convolution kernel of the deep residual network.
3. The method according to claim 1, characterized in that Performing feature fusion on the multi-layer image features to generate a segmented image with a pulmonary embolism prediction area, including: Performing feature fusion on the multi-layer image features to generate a fused unified feature map; Extracting row / column vectors from the unified feature map, and classifying the row / column vectors; Based on the classification results of the row / column vectors, a segmented image with a predicted pulmonary embolism area is generated.
4. The method according to claim 1, wherein The decoding and reconstruction of the multi-layer image features based on the attention guidance mechanism to generate a multi-layer attention map with a pulmonary embolism prediction area includes: Upsampling the multi-layer image features and splicing them with the image features of the corresponding layers to obtain multi-layer spliced image features; Based on the attention guidance mechanism, the attention map of the spliced image features of each layer is extracted, and the attention map of each layer is upsampled to the image size of the next decoding layer.
5. The method according to claim 4, characterized in that After uniformly upsampling each layer of attention map to the image size of the next decoding layer and the size of the training set image, it also includes: The uniformly upsampled attention map is combined with the next decoding layer image to construct the loss function.
6. The method according to claim 1, characterized in that Based on the segmented image, the multi-layer attention map and the training set image, a comprehensive loss function is constructed, including: Constructing a first loss function according to the segmented image and the training set image; Constructing a second loss function based on the multi-layer attention map and the training set images; A comprehensive loss function is constructed by the first loss function and the second loss function.
7. The method according to claim 1, characterized in that After optimizing the model parameters of the deep residual network based on the comprehensive loss function, the method further includes: The validation set images with pulmonary embolism annotated areas are input into the optimized deep residual network to verify the degree of optimization of the model parameters of the deep residual network.
8. A pulmonary embolism image segmentation device based on deep residual network, characterized in that: include: A feature extraction module is used to input the training set images with pulmonary embolism annotated areas into a preset deep residual network to extract multi-layer image features of the training set images; a feature fusion module, configured to fuse the multi-layer image features to generate a segmented image with a pulmonary embolism prediction area; an attention self-learning module, configured to decode and reconstruct the multi-layer image features based on an attention guidance mechanism to generate a multi-layer attention map with a pulmonary embolism prediction area; A parameter optimization module, configured to construct a comprehensive loss function based on the segmented image, the multi-layer attention map, and the training set image, and optimize the model parameters of the deep residual network based on the comprehensive loss function; The segmentation and annotation module is used to use the optimized deep residual network to segment and annotate the pulmonary embolism area in the segmented image.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Medical image segmentation method fusing multi-scale residual attention
CN116563204A
Medical image segmentation method based on multi-scale cross-layer attention fusion network
CN117152433A
Deep learning-based pulmonary embolism segmentation method
CN118212411A
Pancreatic tumor image segmentation method and system based on reinforcement learning and attention
WO2023221954A1