Pneumonia focus segmentation method and system, electronic equipment and storage medium
By combining the dual-encoder U-shaped network of dual-layer routing attention Vision Transformer and residual convolutional neural network, bidirectional feature fusion and wavelet transform downsampling, the problem of insufficient global and local information fusion in pneumonia lesion segmentation is solved, and a more efficient and accurate lesion segmentation effect is achieved.
Patent Information
- Application Number
- CN202510035235.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-09
AI Technical Summary
The prior art is difficult to effectively integrate global and local information in pneumonia lesions segmentation, resulting in a lack of consistency in segmentation results and artifacts may occur. The training effect of Transformer on small data sets is not ideal, especially in detail capture.
A dual encoder U-shaped network that combines a two-layer routing attention Vision Transformer encoder and a residual convolutional neural network encoder is used to fuse global and local features through a bidirectional feature fusion module, combining wavelet transform downsampling and a jump connection module that shares attention weights to improve segmentation accuracy and robustness.
It effectively improves the accuracy and robustness of pneumonia lesions, can better adapt to the characteristics of different lesions, reduce artifacts and blur boundary problems, and improves the detailed performance of segmentation results.
Smart Images

Figure CN119941755A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image segmentation, and in particular to a pneumonia lesion segmentation method, system, electronic device and storage medium. Background Art
[0002] Pneumonia is a common respiratory disease with high morbidity and mortality. Pneumonia can be roughly divided into bacterial pneumonia, fungal pneumonia, mycoplasma pneumonia and viral pneumonia. Among them, viral pneumonia is a common type of pneumonia with a high transmission rate in my country. It often occurs in winter and summer. It is mainly transmitted through droplets and direct contact, invading the airways and alveoli to multiply, causing cell death and immune system overreaction, thereby aggravating lung damage. From December 2019 to November 2024, viral pneumonia caused more than seven million deaths worldwide. Typical symptoms of viral pneumonia include fever, cough, myalgia and dyspnea. Severe patients may develop complications such as acute respiratory distress syndrome. Although reverse transcription polymerase chain reaction is considered the gold standard for confirming virality in the medical field, its detection sensitivity is limited, resulting in false negative results, which is not conducive to the isolation and treatment of patients. In contrast, medical imaging, especially chest CT scans, are considered to have an irreplaceable role in the diagnosis and assessment of pneumonia. CT images can not only show subtle lesions in patients with early infection, but also intuitively present typical imaging features such as ground-glass opacity, consolidation, and septal thickening. Segmenting the infected area can not only assist doctors in conducting rapid and accurate quantitative analysis, but also provide data support for tracking disease progression and evaluating efficacy.
[0003] Medical image segmentation technology has experienced a gradual development from traditional methods to deep learning models. Early segmentation methods mainly relied on techniques such as thresholding, region growing, and edge detection. These methods are effective in specific cases, but are easily affected by noise and image complexity. After 2000, graph theory, clustering, and classification methods were gradually introduced, and segmentation performance was improved, but it was still difficult to capture global context information. In 2015, Long et al. proposed the fully convolutional network (FCN), laying the foundation for the application of deep learning in semantic segmentation. FCN uses convolutional layers instead of traditional fully connected layers, which can perform end-to-end pixel-level predictions and greatly improve segmentation accuracy. Subsequently, the U-Net structure proposed by Ronneberger et al. further optimized the segmentation performance. U-Net uses a symmetrical encoder-decoder architecture to capture global context and local details at the same time, and uses jump connections to achieve the fusion of low-level features and high-level features, greatly improving segmentation accuracy. Multiple extensions and variants based on U-Net have also emerged. For example, the U-Net++ proposed by Zhou et al. enhances the feature fusion capability, the attention mechanism introduced by Oktay et al. improves the model's attention to the region of interest, and the 3D U-Net proposed by Cicek et al. provides a solution for 3D medical image segmentation. These innovations continue to promote the advancement of medical image segmentation technology, improving segmentation accuracy and application scope. At the same time, Transformer has shown great potential in capturing complex texture information with its powerful global feature modeling capabilities. Vaswani et al. first proposed the Transformer model, which marked an important breakthrough in the field of deep learning. Chen et al. proposed TransUNet, which embeds Transformer into the U-Net framework to capture global contextual information, thereby achieving efficient segmentation in complex scenes. Cao et al. developed Swin-Unet, which is based on the Transformer with a sliding window mechanism and can significantly reduce the computational complexity while ensuring global feature extraction. In order to cope with the limitations of small data sets in medical images, Valanarasu et al. proposed Medical Transformer (MedT), which achieves efficient feature capture through the axial attention mechanism.
[0004] Many scholars have conducted in-depth research on pneumonia lesion segmentation and proposed many practical methods. Among them, the best one was proposed by Fan et al. in 2020. The specific method is as follows:
[0005] Step 1: Obtain pneumonia lesion images in 2D format: Since medical CT is mostly in 3D format, when performing lesion segmentation, the 3D image is first preprocessed into a 2D image.
[0006] Step 2: Infnet model training: The main task of pneumonia segmentation is to train a deep network model to identify and segment the pneumonia area, which can be divided into the following three steps:
[0007] Feature encoding: During the downsampling process, CNN feature encoding is used to fuse each stage to complete feature extraction.
[0008] Feature decoding: The decoder uses upsampling to expand the image dimension step by step for the features extracted by the encoder. The features extracted by the encoder at each stage are fused through the parallel partial decoder and input into the decoder. At the same time, the jump connection of the reverse attention mechanism is used to fuse with the features of the encoder at each stage to segment the lesion area.
[0009] Loss function Backpropagation training model: The loss is calculated using the cross entropy loss function.
[0010] In the pneumonia lesion segmentation method, a single method is difficult to cope with the diverse features in pneumonia lesion segmentation. CNN often lacks the ability to model global information, and the training effect of Transformer on small data sets is not ideal, especially in terms of capturing details. At present, the pneumonia segmentation task still faces the following problems: (1) The lesion area of the lung lesion image has the characteristics of blurred edges and high variability. The early lesion area is small, and the lesion area will continue to expand in the later development. For large-area pneumonia lesions, high-level features may pay more attention to the global pattern and ignore the differences in details inside the lesion, resulting in a lack of consistency in the segmentation results within the lesion area and possible artifacts. (2) In the lung lesion segmentation task, the combination of global information and local information is crucial. The effective fusion of these two types of information can help the model extract richer, more diverse and key features. However, the fusion method of simply splicing or adding global and local features may not fully utilize the advantages of the two features. Summary of the invention
[0011] In view of the shortcomings existing in the above problems, the present invention provides a pneumonia lesion segmentation method, system, electronic device and storage medium.
[0012] To achieve the above object, the present invention provides a pneumonia lesion segmentation method, comprising:
[0013] Obtain 2D images of pneumonia CT;
[0014] Training a lesion segmentation model based on the two-dimensional image;
[0015] Using the trained lesion segmentation model to perform lesion segmentation on the input new CT image;
[0016] Among them, the lesion segmentation model includes two encoders, a jump connection module and a decoder, one of the encoders is a two-layer routing attention Vision Transformer encoder, and the other encoder is a residual convolutional neural network encoder. The two-layer routing attention Vision Transformer encoder obtains the global features of the two-dimensional image, and the residual convolutional neural network encoder obtains the local features of the two-dimensional image. The global features and the local features are fused to obtain bidirectional features. The feature fusion formula is:
[0017] X fi =CF i (B i ||C i )+B i ||C i , 0≤i≤n
[0018] CF i (X) = FB3 (FB2 (FB1 (CA (X))))
[0019] FB i (X) = Con 1×1 (RELU(BN(Con 1×1 (X))))+X, 0≤i≤n
[0020] Where: X fi is the fusion feature of each stage, B i and C i They are the features from the two-layer routing attention VisionTransformer encoder and the residual convolutional neural network encoder; CF i is the channel fusion module, CA is the SE attention mechanism, and i takes the value of 2.
[0021] Preferably, a residual wavelet downsampling module is introduced into the dual-layer routing attention Vision Transformer encoder, and the sampling formula of the residual wavelet downsampling module is:
[0022] Haar(X) = {y L ,y H},y H ={y HL ,y LH ,y HH};
[0023]
[0024] Y=Conv 1×1 (ReLU(BN(Conv 1×1 (X haar ))))+Xhaar ;
[0025] Where: Haar(X) represents the division of the feature map into components in different directions; L is the low frequency component, y HL is the high-frequency component in the horizontal direction, y LH is the high frequency component in the vertical direction, y HH is the high-frequency component in the diagonal direction; X haar It is the feature map after concatenating the four directional components; Y represents the feature representation learning of the obtained feature map and outputs the required dimension.
[0026] Preferably, the formula of the skip connection module is:
[0027] W channel =σ(Conv(F pool_channel ));
[0028] W spatial =σ(Conv(F pool_spatial) );
[0029] F enc_att =(W channel +W spatial )·F enc
[0030] F dec_att =(W channel +W spatial )·F dec ;
[0031] F SASC =F enc_att +F dec_att +F enc +F dec ;
[0032] Where: W channel and W spatial are the channel and attention weights, F enc_att and F dec_att are applied to the encoder and decoder features respectively, and the features obtained are enhanced by attention; in order to ensure the integrity of the original feature information, F SASC It is the output feature after the fusion of original features and enhanced features.
[0033] Preferably, a COVID-19-CT-Seg dataset training set is constructed based on the two-dimensional image.
[0034] The present application also provides a pneumonia lesion segmentation system, comprising:
[0035] An acquisition module, used for acquiring a two-dimensional image of pneumonia CT;
[0036] A training module, used for training a lesion segmentation model based on the two-dimensional image;
[0037] An input module, used to perform lesion segmentation on an input new CT image using the trained lesion segmentation model;
[0038] Among them, the lesion segmentation model includes two encoders, a jump connection module and a decoder, one of the encoders is a two-layer routing attention Vision Transformer encoder, and the other encoder is a residual convolutional neural network encoder. The two-layer routing attention Vision Transformer encoder obtains the global features of the two-dimensional image, and the residual convolutional neural network encoder obtains the local features of the two-dimensional image. The global features and the local features are fused to obtain bidirectional features. The feature fusion formula is:
[0039] X fi =CF i (B i ||C i )+B i ||C i , 0≤i≤n
[0040] CF i (X) = FB3 (FB2 (FB1 (CA (X))))
[0041] FB i (X) = Con 1×1 (RELU(BN(Con 1×1 (X))))+X, 0≤i≤n
[0042] Where: X fi is the fusion feature of each stage, B i and C i They are the features from the two-layer routing attention VisionTransformer encoder and the residual convolutional neural network encoder; CF i is the channel fusion module, CA is the SE attention mechanism, and i takes the value of 2.
[0043] The present invention also provides an electronic device, comprising at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the above method.
[0044] The present invention also provides a storage medium storing a computer program executable by an electronic device, and when the program runs on the electronic device, the electronic device executes the above method.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] The present invention effectively integrates residual convolutional neural network, VisionTransformer based on double-layer routing mechanism attention, wavelet transform and other technologies to further perform the task of lesion segmentation in pneumonia CT images. Compared with other models, this model effectively improves the accuracy of pneumonia lesion segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flow chart of the pneumonia lesion segmentation method of the present invention;
[0048] Figure 2 It is a schematic diagram of the overall architecture of the lesion segmentation model in the present invention;
[0049] Figure 3 It is a schematic diagram of the dual-layer routing attention Vision Transformer architecture in the present invention;
[0050] Figure 4 It is a schematic diagram of the bidirectional feature fusion module architecture in the present invention;
[0051] Figure 5 It is a schematic diagram of the architecture of the wavelet transform downsampling module in the present invention;
[0052] Figure 6 It is a schematic diagram of the jump connection module architecture in the present invention. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0054] Reference Figure 1 The present invention provides a pneumonia lesion segmentation method, comprising:
[0055] Obtain 2D images of pneumonia CT;
[0056] Training the lesion segmentation model based on 2D images;
[0057] Use the trained lesion segmentation model to perform lesion segmentation on the input new CT image;
[0058] Among them, the lesion segmentation model includes two encoders, a skip connection module and a decoder. One encoder is a two-layer routing attention Vision Transformer encoder, and the other encoder is a residual convolutional neural network encoder. The two-layer routing attention Vision Transformer encoder obtains the global features of the two-dimensional image, and the residual convolutional neural network encoder obtains the local features of the two-dimensional image. The global features and local features are fused to obtain bidirectional features.
[0059] Reference Figure 2 , this application proposes a dual encoder U-type network based on residual convolution and double-layer routing attention Vision Transformer, a lesion segmentation model for pneumonia CT images. It consists of three parts, including encoder, jump connection, and decoder. Unlike Infnet, it abandons the use of a simple residual convolutional neural network, but uses a double-layer routing attention Vision Transformer and a residual convolutional neural network as encoders, and uses a two-way feature fusion method to obtain global context representation and rich local information. At the same time, the jump connection part uses shared attention weights to better adjust the semantic information of high and low dimensions. Not only that, we abandon the simple downsampling method and use downsampling based on wavelet transform to better retain the detail information of the features, so that the model can achieve better segmentation effects.
[0060] In this embodiment, since the size of the lesions at different stages in the pneumonia lesion segmentation example is highly variable, the simple use of convolutional neural networks and Vision Transformer will lead to the loss of feature information. Because the encoder uses residual convolutional neural networks and Vision Transformer based on double-layer routing attention, it can not only extract local and global information, but also reduce the complexity of calculation. Figure 3 shown.
[0061] Reference Figure 4, since the residual convolutional neural network focuses more on extracting local information, the Vision Transformer based on the two-layer routing attention focuses more on extracting global information. In order to better fuse the information of the two, unlike the previous unilateral supplement of the information extracted by the Vision Transformer to the convolutional neural network encoder part, a bidirectional feature fusion module is used to fuse the information of the two and then transmit it back to the two encoders respectively. ViT captures semantic information in a global scope and has better noise resistance to local interference such as artifacts. CNN can avoid ViT's neglect of local features, identify lesions in detail areas, and reduce the possibility of false detection of artifacts. This complementary feature enables the model to stably segment the lesion area in the presence of noise or artifacts. At the same time, for larger and obvious lesions, the advantages of ViT's global features are more obvious. For small lesions or lesions with irregular shapes, CNN's ability to extract local details is particularly important. Bidirectional fusion allows the model to dynamically adjust the weights of local and global features to better adapt to the characteristics of different lesions.
[0062] Specifically, the feature fusion formula is:
[0063] X fi =CF i (B i ||C i )+B i ||C i , 0≤i≤n
[0064] CF i (X) = FB3 (FB2 (FB1 (CA (X))))
[0065] FB i (X) = Con 1×1 (RELU(BN(Con 1×1 (X))))+X, 0≤i≤n
[0066] Where: X fi is the fusion feature of each stage, B i and C i They are the features from the two-layer routing attention VisionTransformer encoder and the residual convolutional neural network encoder; CF i is the channel fusion module, CA is the SE attention mechanism, and i takes the value of 2.
[0067] Reference Figure 5In the pneumonia lesion segmentation task, a residual wavelet downsampling module was introduced into the encoder part of the Vision Transformer based on dual-layer routing attention to improve the feature extraction ability of the model. Residual wavelet downsampling divides the input features into low-frequency and high-frequency components through frequency domain decomposition, which not only retains the global structural information of the lesion, but also captures the boundary details and texture features, thereby enhancing the model's sensitivity to small lesions and complex boundaries. In addition, compared with traditional downsampling methods, residual wavelet downsampling can reduce the feature size while retaining key information, thereby significantly reducing computational complexity and improving model efficiency. In this way, the model can dynamically fuse multi-scale information during segmentation, alleviate the problem of fuzzy boundaries, and improve segmentation accuracy and robustness while suppressing noise.
[0068] Specifically, a residual wavelet downsampling module is introduced in the two-layer routing attention Vision Transformer encoder. The sampling formula of the residual wavelet downsampling module is:
[0069] Haar(X) = {y L ,y H},y H ={y HL ,y LH , yHH};
[0070]
[0071] Y=Conv 1×1 (ReLU(BN(Conv 1×1 (X haar ))))+X haar ;
[0072] Where: Haar(X) represents the division of the feature map into components in different directions; V L is the low frequency component, y HL is the high-frequency component in the horizontal direction, y LH is the high frequency component in the vertical direction, y HH is the high-frequency component in the diagonal direction; X haar It is the feature map after concatenating the four directional components; Y represents the feature representation learning of the obtained feature map and outputs the required dimension.
[0073] Reference Figure 6In the pneumonia lesion segmentation task, the feature hierarchical mismatch between the high-level semantic features of the encoder and the low-level spatial details of the decoder limits the efficiency of feature fusion, especially in preserving boundaries and enhancing details. To solve this problem, a skip connection module with shared attention weights is proposed. By introducing an interactive shared attention mechanism, the shared attention weights of the channel and spatial dimensions are dynamically generated to improve the alignment of high-level semantics and low-level detail features. The skip connection module with shared attention weights achieves efficient collaborative feature fusion, which not only improves the propagation of global semantic information, but also focuses on enhancing boundary and detail representation. This design significantly improves the accuracy of pneumonia lesion segmentation, especially in the segmentation of fuzzy boundaries and small lesion areas.
[0074] Specifically, the formula of the skip connection module is:
[0075] W channel =σ(Conv(F pool_chanel ));
[0076] W spatial =σ(Conv(F pool_spatial ));
[0077] F enc_att =(W channel +W spatial )·F enc
[0078] F dec_att =(W channel +W spatial )·F dec ;
[0079] F SASC =F enc_att +F dec_att +F enc +F dec ;
[0080] Where: W channel and W spatial are the channel and attention weights, F enc_att and F dec_att are applied to the encoder and decoder features respectively, and the features obtained are enhanced by attention; in order to ensure the integrity of the original feature information, F SASC It is the output feature after the fusion of original features and enhanced features.
[0081] In this embodiment, the present application establishes a suitable model. In the encoder part, the advantages of convolutional neural networks in local feature extraction are combined with the ability of Vision Transformer based on a two-layer routing attention mechanism in global dependency modeling, which can accurately capture the boundary of lesions and complex texture information. At the same time, the bidirectional feature fusion module of the channel dimension is used to make full use of global and local features, thereby significantly improving the detail performance of the segmentation results. A downsampling module based on wavelet transform is used to prevent the loss of early small lesion area features. In the jump connection part, by dynamically adjusting the fusion weights of low-dimensional encoding features and high-dimensional decoding features, the feature transfer process is optimized, the semantic gap problem between the encoder and the decoder is effectively alleviated, and the segmentation accuracy of the lesion area is improved. Finally, the final model is obtained by training on the public dataset COVID-19-CT-Seg dataset and inputting a two-dimensional image of pneumonia CT with a size feature of 256×256. The trained model is used to segment the lesions of the input new CT image.
[0082] The present application also provides a pneumonia lesion segmentation system, comprising:
[0083] An acquisition module, used for acquiring a two-dimensional image of pneumonia CT;
[0084] A training module, used for training a lesion segmentation model based on two-dimensional images;
[0085] An input module is used to perform lesion segmentation on an input new CT image using a trained lesion segmentation model;
[0086] Among them, the lesion segmentation model includes two encoders, a skip connection module and a decoder. One encoder is a two-layer routing attention Vision Transformer encoder, and the other encoder is a residual convolutional neural network encoder. The two-layer routing attention Vision Transformer encoder obtains the global features of the two-dimensional image, and the residual convolutional neural network encoder obtains the local features of the two-dimensional image. The global features and local features are fused to obtain bidirectional features. The feature fusion formula is:
[0087] X fi =CF i (B i ||C i )+B i ||C i , 0≤i≤n
[0088] CF i (X) = FB3 (FB2 (FB1 (CA (X))))
[0089] FB i (X) = Con 1×1(RELU(BN(Con 1×1 (X))))+X, 0≤i≤n
[0090] Where: X fi is the fusion feature of each stage, B i and C i They are the features from the two-layer routing attention VisionTransformer encoder and the residual convolutional neural network encoder; CF i is the channel fusion module, CA is the SE attention mechanism, and i takes the value of 2.
[0091] The present invention also provides an electronic device, comprising at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the above method.
[0092] The present invention also provides a storage medium storing a computer program executable by an electronic device. When the program runs on the electronic device, the electronic device executes the above method.
[0093] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A pneumonia lesion segmentation method, characterized in that: include: Obtain 2D images of pneumonia CT; Training a lesion segmentation model based on the two-dimensional image; Using the trained lesion segmentation model to perform lesion segmentation on the input new CT image; Among them, the lesion segmentation model includes two encoders, a jump connection module and a decoder, one of the encoders is a two-layer routing attention Vision Transformer encoder, and the other encoder is a residual convolutional neural network encoder. The two-layer routing attention Vision Transformer encoder obtains the global features of the two-dimensional image, and the residual convolutional neural network encoder obtains the local features of the two-dimensional image. The global features and the local features are fused to obtain bidirectional features. The feature fusion formula is: X fi =CF i (B i ||C i )+B i |C i ,0≤i≤n CF i (X)=FB3(FB2(FB1(CA(X)))) FB i (X)=With 1×1 (RELU(BN(With 1×1 (X))))+X,0≤i≤n Where: X fi is the fusion feature of each stage, B i and C i They are the features from the two-layer routed attention Vision Transformer encoder and the residual convolutional neural network encoder; CF i is the channel fusion module, CA is the SE attention mechanism, and i takes the value of 2.
2. The pneumonia lesion segmentation method according to claim 1, characterized in that: The residual wavelet downsampling module is introduced into the dual-layer routing attention VisionTransformer encoder, and the sampling formula of the residual wavelet downsampling module is: Haar(X)={y L ,and H },and H ={and HL ,and LH ,and HH }; Y=Conv 1×1 (ReLU(BN(Conv 1×1 (X haar ))))+X haar Where: Haar(X) represents the division of the feature map into components in different directions; L is the low frequency component, y HL is the high-frequency component in the horizontal direction, y LH is the high frequency component in the vertical direction, y HH is the high-frequency component in the diagonal direction; X haar It is the feature map after concatenating the four directional components; Y represents the feature representation learning of the obtained feature map and outputs the required dimension.
3. The pneumonia lesion segmentation method according to claim 2, characterized in that: The formula of the skip connection module is: W channel =σ(Conv(F pool_channel )); W spatial =σ(Conv(F pool_spatial )); F enc_att =(W channel +W spatial )·F enc F dec_att =(W channel +W spatial )·F dec ; F SASC =F enc_att +F dec_att +F enc +F dec Where: W channel and W spatial are the channel and attention weights, F enc_att and F dec_att are applied to the encoder and decoder features respectively, and the features obtained are enhanced by attention; in order to ensure the integrity of the original feature information, F SASC It is the output feature after the fusion of original features and enhanced features.
4. The pneumonia lesion segmentation method according to claim 1, characterized in that: The COVID-19-CT-Seg dataset training set is constructed based on the two-dimensional images.
5. The method for segmenting pneumonia lesions according to claim 4, characterized in that The image size feature in the training set is 256×256.
6. A pneumonia lesion segmentation system, characterized in that: include: An acquisition module, used for acquiring a two-dimensional image of pneumonia CT; A training module, used for training a lesion segmentation model based on the two-dimensional image; An input module, used to perform lesion segmentation on an input new CT image using the trained lesion segmentation model; Among them, the lesion segmentation model includes two encoders, a jump connection module and a decoder, one of the encoders is a two-layer routing attention Vision Transformer encoder, and the other encoder is a residual convolutional neural network encoder. The two-layer routing attention Vision Transformer encoder obtains the global features of the two-dimensional image, and the residual convolutional neural network encoder obtains the local features of the two-dimensional image. The global features and the local features are fused to obtain bidirectional features. The feature fusion formula is: X fi =CF i (B i ||C i )+B i |C i ,0≤i≤n CF i (X)=FB3(FB2(FB1(CA(X)))) FB i (X)=With 1×1 (RELU(BN(With 1×1 (X))))+X,0≤i≤n Where: X fi is the fusion feature of each stage, B i and C i They are the features from the two-layer routed attention Vision Transformer encoder and the residual convolutional neural network encoder; CF i is the channel fusion module, CA is the SE attention mechanism, and i takes the value of 2.
7. An electronic device, characterized in that: The method comprises at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the method according to any one of claims 1 to 5.
8. A storage medium, characterized in that: It stores a computer program executable by an electronic device. When the program runs on the electronic device, the electronic device executes the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Distraction-based Swin Unet variant medical image segmentation method
CN116524187A
Pneumonia CT image segmentation method, device and equipment
CN116579982A
Lung CT image segmentation method based on Transform and convolutional neural network
CN116739985A
Hyperspectral image classification method based on mixed covariance attention and cross-layer fusion Transform
CN117576473A
Viral pneumonia focus segmentation method based on domain self-adaption and multi-scale feature fusion
CN117830327A
Cited By
Medical image focus identification method and system based on double-path attention mechanism
CN120525860A