Medical image segmentation method and system based on CNN and ViT hybrid architecture

By combining the advantages of CNN and ViT, a medical image segmentation method with a hybrid architecture is proposed, which solves the difficulties in the existing technology in complex shape and boundary recognition, and achieves a balance between high precision and low computing costs, which is suitable for a variety of medical image segmentation tasks.

CN120107607AActive Publication Date: 2025-06-06NANCHANG HANGKONG UNIVERSITY

Patent Information

Application Number
CN202510593299.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-06
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

Existing medical image segmentation technology has the problem of precise identification and segmentation when dealing with complex shapes and boundaries, especially in blurred edges and low contrast areas, and the calculation cost is high, making it difficult to achieve a balance between high precision and low complexity.

Method used

A medical image segmentation method based on the hybrid architecture of CNN and ViT is proposed, combining the local feature extraction capabilities of CNN and the global context modeling capabilities of ViT, and enhancing feature extraction and semantic correlation through the multi-scale feature extraction module and semantic enhancement module, and refine it through the edge uncertainty guidance module to improve segmentation accuracy and calculation efficiency.

Benefits of technology

It significantly improves the accuracy of medical image segmentation, especially the segmentation effect of tiny lesions, low-contrast areas and boundary blur areas, reduces calculation costs, and improves the applicability and real-timeness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107607A_ABST
    Figure CN120107607A_ABST
Patent Text Reader

Abstract

The invention specifically discloses a medical image segmentation method and system based on a CNN and ViT hybrid architecture, and relates to the technical field of medical image processing. The method comprises the following steps: S1, acquiring a plurality of types of medical image data sets, preprocessing the medical image data sets, and dividing the medical image data sets into a training set, a verification set and a test set in proportion; s2, a segmentation model based on a CNN and ViT hybrid architecture is constructed, and the segmentation model takes a U-Net network as a basic framework and comprises an encoder, a decoder and an edge optimization module; s3, performing segmentation model training based on the divided training data set; s4, calculating a loss function of the medical image segmentation model, updating parameters by using an Adam optimizer, and storing an optimal model weight in a training process; and S5, testing the data of the test set by using the trained segmentation model. According to the method, the problems that the CNN receptive field is limited and the ViT calculation cost is high are effectively relieved, and the segmentation precision of the focus area is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a medical image segmentation method and system based on a CNN and ViT hybrid architecture. Background Art

[0002] As a key technology in the field of computer vision, medical image segmentation aims to accurately divide lesions or regions of interest from medical images to provide support for quantitative and qualitative analysis. This technology plays an extremely important role in medical auxiliary diagnosis, treatment plan planning, and prognosis judgment. Although deep learning-based methods have achieved remarkable results, the characteristics of medical images themselves, such as morphological complexity (such as the irregular fractal structure of invasive tumors), low signal-to-noise ratio interference (for example, the grayscale of lesions and background in ultrasound images is highly overlapped), and the multi-scale feature mismatch of tiny lesions are still the main obstacles to the current development of technology.

[0003] With the continuous innovation of deep learning technology, convolutional neural network (CNN) has become the mainstream architecture in the field of medical image segmentation with its excellent local feature extraction ability. It can automatically mine deep features from images without manual feature design or additional intervention. However, CNN also has obvious limitations. Its fixed convolution kernel and limited receptive field result in poor global context modeling capabilities, and small target lesions are very likely to be missed during the segmentation process. Specifically, it is mainly reflected in two aspects: first, the convolution kernel parameters are determined in the training stage, and lack the ability to adaptively adjust when facing different input images; second, the convolution kernel can only model the pixel relationship in the local area and cannot perceive the global information, which seriously restricts the accurate recognition and segmentation of lesion shape and boundary pixels (such as skin lesions and liver tumor boundaries). Therefore, in the task of medical image segmentation, mining the long-range dependency between image pixels is crucial to improving segmentation accuracy.

[0004] In order to overcome these shortcomings of CNN, many medical image segmentation methods based on visual transformer (ViT) have emerged. ViT can obtain a large receptive field and effectively capture the long-range dependencies between different image patches. However, the self-attention mechanism in ViT has significant drawbacks. It usually requires a large amount of data for training to achieve ideal performance, and its computational complexity increases quadratically with the length of the input sequence, which is contrary to the real-time requirements of clinical applications. In addition, compared with CNN, ViT is less sensitive to local positions, and the segmentation quality will be seriously affected when processing small target areas and medical images with complex and changeable morphology.

[0005] Although the recently proposed method based on the hybrid architecture of CNN and ViT has achieved leading results in benchmark tests, it is still difficult to deal with blurred edges and low-contrast areas, and is often accompanied by high computational costs, making it difficult to find an ideal balance between high-precision segmentation and model complexity. Therefore, an innovative network architecture and implementation method are urgently needed to overcome the above difficulties and promote further improvement of medical image segmentation performance. Summary of the invention

[0006] The purpose of the present invention is to propose a medical image segmentation method and system based on a hybrid architecture of CNN and ViT, which combines the advantages of CNN and ViT, effectively alleviates the problems of limited receptive field of CNN and expensive computational cost of ViT, and captures local spatial information and global context information while closely combining channel information modeling capabilities, more efficiently extracts medical image features, enhances the correlation between different semantic levels, improves the segmentation accuracy of lesion areas, and provides a more accurate and reliable solution in the field of medical image segmentation.

[0007] To achieve the above objectives, the present invention proposes a medical image segmentation method based on a hybrid architecture of CNN and ViT, and the specific steps are as follows: Step S1, obtaining multiple types of medical image data sets, preprocessing the original data, and generating preprocessed images; dividing the processed data into a training set, a validation set, and a test set in proportion; Step S2, constructing a segmentation model based on a hybrid architecture of CNN and ViT, wherein the segmentation model is based on a U-Net network as a basic framework, including an encoder, a decoder, and an edge optimization module; Step S3, performing segmentation model training based on the divided training data set; Step S4, calculating the loss function of the medical image segmentation model, using the Adam optimizer to update the parameters, and saving the optimal model weights during the training process; Step S5: Use the trained segmentation model to test the test set data.

[0008] Preferably, in step S2, the encoder includes four encoding layers, each layer includes a multi-scale feature extraction module and a semantic enhancement module; the decoder includes four decoding layers, each layer also includes a multi-scale feature extraction module and a semantic enhancement module; a jump connection is used between the encoding layer and the decoding layer in the same layer, and frequency domain information is introduced at the bottleneck layer; the multi-scale feature extraction module extracts multi-scale feature semantic information from the input image, and the semantic enhancement module enhances the correlation between different semantic information; the output of the last decoding layer is upsampled to obtain a coarse segmentation result, which is then refined by the edge uncertainty guidance module to obtain a final output prediction result; In the encoder stage, the input feature map After downsampling through the multi-scale feature extraction module and the semantic enhancement module, it is input to the next encoding layer. After downsampling of the feature map layer by layer, the spatial resolution decreases layer by layer, and the number of channels increases layer by layer. The downsampling formula at the encoder stage is as follows: ; ; ; ; in, is a 3×3 convolution operation, is the semantic enhancement module, is a multi-scale feature extraction module. is the output of the first encoding layer, is the output of the second encoding layer, is the output of the third encoding layer, is the output of the fourth encoding layer.

[0009] Preferably, in step S2, additional frequency domain features are introduced at the bottleneck layer, and the input is decomposed into low-frequency and high-frequency components by wavelet transform WT, and reconstructed by inverse wavelet transform IWT after deep convolution extraction of features, and the formula is as follows: ; in, is the output result after inverse wavelet transform, is the weight vector after the convolution kernel of size k×k.

[0010] Preferably, in step S2, in the decoder stage, the input of each decoding layer is the output feature map of the previous decoding layer, and the upsampling result is combined with the output feature map of the same encoding layer by element point summation for feature fusion. After the feature map is upsampled layer by layer, the spatial resolution is restored layer by layer, the number of channels is reduced layer by layer, and finally the segmentation model generates a coarse segmentation result.

[0011] Preferably, in step S2, the edge uncertainty guidance module optimizes and reshapes the uncertainty pixels in the boundary fuzzy area in the coarse segmentation result, specifically by generating a refined mask by combining shallow features, uncertainty mapping and local feature matching, optimizing the coarse segmentation result, and obtaining the final predicted segmentation result.

[0012] Preferably, the shallow features are outputs of the first encoding layer; The uncertainty mapping is calculated based on the probability map of the rough segmentation result, and the formula is as follows: ; in, For uncertain mapping, the value range is ; , Represent the maximum and minimum probability values ​​of the pixels respectively. Calculate for exponential function; The local feature is generated by weighted local averaging of the uncertain mapping, and the formula is as follows: ; ; in, is a local feature, represents the normalization factor, is a small value that prevents the denominator from reaching zero, is a pixel, is the pixel after refinement, The pixel belongs to the target category The probability of It is a shallow feature. Pixel The neighborhood area; Refined masks are generated by combining shallow features, uncertainty mapping, and local feature matching , the calculation formula is as follows: ; ; in, For pixels The confidence score at is the confidence module; The rough segmentation result is optimized to obtain the final predicted segmentation result. The formula is as follows: ; in, To predict the segmentation result.

[0013] Preferably, the confidence module includes 3×3 convolution, Relu activation function and residual connection; wherein, 3×3 convolution is applied on the main branch for feature extraction, and 3×3 convolution is also applied on the residual branch for feature extraction, and finally the main branch output and the residual branch output are added element by element; the Relu activation function is applied before the convolution operation.

[0014] Preferably, in step S4, the loss function of the segmentation model is calculated, and the formula is as follows: ; ; ; in, is the sum of all pixels in the input image, , denote the predicted value and the target value respectively. is the total loss, is the binary cross entropy loss, is the dice loss, , is the weight assigned to each loss.

[0015] The present invention also provides a medical image segmentation system based on a CNN and ViT hybrid architecture, comprising: The data acquisition and preprocessing module is used to acquire multiple types of medical image data sets, preprocess the original data, generate preprocessed images, and divide the processed data into training sets, validation sets, and test sets in proportion; A model building module, used to build a segmentation model based on a hybrid architecture of CNN and ViT, wherein the model is based on a U-Net network and includes an encoder, a decoder, and an edge optimization module; The model training module is used to train the model based on the divided training data set, calculate the loss function of the medical image segmentation model, update the parameters using the Adam optimizer, and save the optimal model weights during the training process; The model testing and evaluation module is used to test the test set data using the trained segmentation model and evaluate the model performance; The storage module is used to store medical image data, model parameters and intermediate calculation results.

[0016] Preferably, the encoder constructed by the model building module includes four encoding layers, each layer includes a multi-scale feature extraction module and a semantic enhancement module; the decoder includes four decoding layers, each layer also includes a multi-scale feature extraction module and a semantic enhancement module; jump connections are used between the encoding layer and the decoding layer at the same layer, and frequency domain information is introduced at the bottleneck layer.

[0017] Therefore, the present invention proposes a medical image segmentation method and system based on a hybrid architecture of CNN and ViT, which has the following beneficial effects: (1) Significantly improved segmentation accuracy: The hybrid architecture combining CNN and ViT can simultaneously capture local spatial information and global contextual information, and significantly improves the segmentation accuracy of tiny lesions, low-contrast areas, and areas with blurred boundaries.

[0018] (2) Enhanced long-range dependency modeling capabilities: The introduction of ViT’s self-attention mechanism effectively compensates for the shortcomings of traditional CNN in global information modeling and improves the ability to recognize complex shapes and boundaries.

[0019] (3) Good edge refinement effect: The designed edge enhancement perception module can optimize the fuzzy boundary areas in the segmentation results, and significantly improve the segmentation quality when dealing with unclear boundaries or small targets.

[0020] (4) Computational efficiency optimization: Combining the local feature extraction of CNN and the global modeling capability of ViT, compared with the pure ViT method, it has higher computational efficiency, avoids high computational costs, and is more suitable for scenarios with high clinical real-time requirements.

[0021] (5) Wide applicability: The hybrid architecture is highly versatile and can be applied to a variety of medical image segmentation tasks, such as liver tumor segmentation, cell segmentation, and skin lesion segmentation. It can achieve high-precision segmentation in different types of medical images.

[0022] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 A schematic diagram of a medical image segmentation method based on a CNN and ViT hybrid architecture of the present invention; Figure 2 A flow chart of feature extraction by the multi-scale feature extraction module in the present invention; Figure 3 The figure is a flow chart of the refinement operation of the edge uncertainty guidance module in the present invention. DETAILED DESCRIPTION

[0024] In order to make the technical solutions, advantages and purposes of the present invention clearer, the technical solutions of the embodiments of the present invention are clearly and completely described below. The described embodiments are part of the embodiments of the present invention, not all of them. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.

[0025] Unless otherwise defined, technical or scientific terms used in the present invention shall have the common meanings understood by one having ordinary skills in the field to which the present invention belongs.

[0026] Embodiment 1 like Figure 1 FIG. 1 is a flow chart of a medical image segmentation method based on a hybrid architecture of CNN and ViT according to the present invention, and the specific steps are as follows: S1. Obtain multiple types of medical image data sets, preprocess the original data, and generate preprocessed images; divide the processed data into a training set, a validation set, and a test set in proportion; S2. Build a segmentation model based on the hybrid architecture of CNN and ViT. The segmentation model uses the U-Net network as the basic framework, including encoder, decoder and edge optimization module; like Figure 2 As shown, the encoder includes four encoding layers, each of which includes a multi-scale feature extraction module and a semantic enhancement module; the decoder includes four decoding layers, each of which also includes a multi-scale feature extraction module and a semantic enhancement module; a jump connection is used between the encoding layer and the decoding layer in the same layer, and frequency domain information is introduced at the bottleneck layer; the multi-scale feature extraction module extracts multi-scale feature semantic information from the input image, and the semantic enhancement module enhances the correlation between different semantic information; the output of the last decoding layer is upsampled to obtain a coarse segmentation result, which is then refined through the edge uncertainty guidance module to obtain the final output prediction result; In the encoder stage, the input feature map After downsampling through the multi-scale feature extraction module and the semantic enhancement module, it is input to the next encoding layer. After downsampling of the feature map layer by layer, the spatial resolution decreases layer by layer, and the number of channels increases layer by layer. The downsampling formula at the encoder stage is as follows: ; ; ; ; in, is a 3×3 convolution operation, is the semantic enhancement module, is a multi-scale feature extraction module. is the output of the first encoding layer, is the output of the second encoding layer, is the output of the third encoding layer, is the output of the fourth encoding layer.

[0027] Additional frequency domain features are introduced at the bottleneck layer, and the input is decomposed into low-frequency and high-frequency components using wavelet transform WT. After deep convolution extracts features, it is reconstructed using inverse wavelet transform IWT. The formula is as follows: ; in, is the output result after inverse wavelet transform, is the weight vector after the convolution kernel of size k×k.

[0028] , , , , Through , , and DWConv is used to obtain multi-scale local information. After convolution, each subset , i =1,2,3,4,5, the output is accumulated in the subsequent stage, which realizes the advance of fine-grained multi-scale features, expands the receptive field of the model, and reduces the loss of feature information flow. The specific formula is as follows: ; Finally, The multi-scale features obtained by splicing them according to the channel dimension express: ; in, , Represent the final output and the output of each feature subset respectively, represents the squeeze excitation module, Indicates the concatenation operation according to the channel dimension. is the convolution process, For the The output feature map of each subset after the corresponding convolution processing.

[0029] In the decoder stage, the input of each decoding layer is the upsampling result of the output feature map of the previous decoding layer and the output feature map of the same encoding layer. The feature fusion is performed by element-wise summation. After the feature map is upsampled layer by layer, the spatial resolution is restored layer by layer, and the number of channels is reduced layer by layer. Finally, the segmentation model generates a coarse segmentation result.

[0030] like Figure 3 As shown in the figure, the edge uncertainty guidance module optimizes and reshapes the uncertainty pixels in the fuzzy boundary area of ​​the coarse segmentation result. Specifically, it generates a refined mask by combining shallow features, uncertainty mapping and local feature matching, optimizes the coarse segmentation result, and obtains the final predicted segmentation result.

[0031] The shallow features are the output of the first encoding layer; The uncertainty mapping is calculated based on the probability map of the coarse segmentation result, and the formula is as follows: ; in, For uncertain mapping, the value range is ; , Represent the maximum and minimum probability values ​​of the pixels respectively. Calculate for exponential function; The local feature is generated by weighted local averaging of uncertain mapping, and the formula is as follows: ; ; in, is a local feature, represents the normalization factor, is a small value that prevents the denominator from reaching zero, is a pixel, is the pixel after refinement, The pixel belongs to the target category The probability of It is a shallow feature. Pixel The neighborhood area; Refined masks are generated by combining shallow features, uncertainty mapping, and local feature matching , the calculation formula is as follows: ; ; in, For pixels The confidence score at is the confidence module; The rough segmentation result is optimized to obtain the final predicted segmentation result. The formula is as follows: ; in, To predict the segmentation result.

[0032] The confidence module includes 3×3 convolution, Relu activation function and residual connection; among them, 3×3 convolution is applied on the main branch for feature extraction, and 3×3 convolution is also applied on the residual branch for feature extraction. Finally, the main branch output and the residual branch output are added element by element; the Relu activation function is applied before the convolution operation.

[0033] S3, performing segmentation model training based on the divided training data set; S4. Calculate the loss function of the medical image segmentation model, use the Adam optimizer to update the parameters, and save the optimal model weights during the training process; Calculate the loss function of the segmentation model, the formula is as follows: ; ; ; in, is the sum of all pixels in the input image, , denote the predicted value and the target value respectively. is the total loss, is the binary cross entropy loss, is the dice loss, , is the weight assigned to each loss.

[0034] S5. Use the trained segmentation model to test the test set data.

[0035] Embodiment 2 The present invention also provides a medical image segmentation system based on a CNN and ViT hybrid architecture, comprising: The data acquisition and preprocessing module is used to acquire multiple types of medical image data sets, preprocess the original data, generate preprocessed images, and divide the processed data into training sets, validation sets, and test sets in proportion; The model building module is used to build a segmentation model based on the hybrid architecture of CNN and ViT. The segmentation model is based on the U-Net network framework, including encoder, decoder and edge optimization module; The encoder constructed by the model building module includes four encoding layers, each of which includes a multi-scale feature extraction module and a semantic enhancement module; the decoder includes four decoding layers, each of which also includes a multi-scale feature extraction module and a semantic enhancement module; jump connections are used between the encoding layer and the decoding layer on the same layer, and frequency domain information is introduced at the bottleneck layer.

[0036] The model training module is used to train the model based on the divided training data set, calculate the loss function of the medical image segmentation model, update the parameters using the Adam optimizer, and save the optimal model weights during the training process; The model testing and evaluation module is used to test the test set data using the trained segmentation model and evaluate the model performance; The storage module is used to store medical image data, model parameters and intermediate calculation results.

[0037] It is worth noting that the contents not elaborated in detail in the present invention are all prior art and are well known to those skilled in the art.

[0038] Therefore, the present invention provides a medical image segmentation method and system based on a CNN and ViT hybrid architecture, which effectively alleviates the problems of limited CNN receptive field and expensive ViT computational cost by combining the advantages of CNN and ViT, improves the segmentation accuracy of the lesion area, and provides a more accurate and reliable solution for the field of medical image segmentation.

[0039] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A medical image segmentation method based on a hybrid architecture of CNN and ViT, characterized in that: Here are the steps: Step S1, obtaining multiple types of medical image data sets, preprocessing the original data, and generating preprocessed images; dividing the processed data into a training set, a validation set, and a test set in proportion; Step S2, constructing a segmentation model based on a hybrid architecture of CNN and ViT, wherein the segmentation model is based on a U-Net network as a basic framework, including an encoder, a decoder, and an edge optimization module; Step S3, performing segmentation model training based on the divided training data set; Step S4, calculating the loss function of the medical image segmentation model, using the Adam optimizer to update the parameters, and saving the optimal model weights during the training process; Step S5: Use the trained segmentation model to test the test set data.

2. The medical image segmentation method based on a hybrid architecture of CNN and ViT according to claim 1, characterized in that: In step S2, the encoder includes four encoding layers, each of which includes a multi-scale feature extraction module and a semantic enhancement module; the decoder includes four decoding layers, each of which also includes a multi-scale feature extraction module and a semantic enhancement module; a skip connection is used between the encoding layer and the decoding layer in the same layer, and frequency domain information is introduced at the bottleneck layer; The multi-scale feature extraction module extracts multi-scale feature semantic information from the input image, and the semantic enhancement module enhances the correlation between different semantic information; the output of the last decoding layer is upsampled to obtain a coarse segmentation result, which is then refined by the edge uncertainty guidance module to obtain a final output prediction result; In the encoder stage, the input feature map After downsampling through the multi-scale feature extraction module and the semantic enhancement module, it is input to the next encoding layer. After downsampling of the feature map layer by layer, the spatial resolution decreases layer by layer, and the number of channels increases layer by layer. The downsampling formula at the encoder stage is as follows: ; ; ; ; in, is a 3×3 convolution operation, is the semantic enhancement module, is a multi-scale feature extraction module. is the output of the first encoding layer, is the output of the second encoding layer, is the output of the third encoding layer, is the output of the fourth encoding layer.

3. The medical image segmentation method based on the hybrid architecture of CNN and ViT according to claim 1, characterized in that: In step S2, additional frequency domain features are introduced at the bottleneck layer, and the input is decomposed into low-frequency and high-frequency components using wavelet transform WT. After deep convolution extracts features, it is reconstructed by inverse wavelet transform IWT. The formula is as follows: ; in, is the output result after inverse wavelet transform, is the weight vector after the convolution kernel of size k×k.

4. The medical image segmentation method based on a hybrid architecture of CNN and ViT according to claim 1, characterized in that: In step S2, in the decoder stage, the input of each decoding layer is the output feature map of the previous decoding layer, which is upsampled and fused with the output feature map of the same encoding layer by element point summation. After the feature map is upsampled layer by layer, the spatial resolution is restored layer by layer, and the number of channels is reduced layer by layer. Finally, the segmentation model generates a coarse segmentation result.

5. The medical image segmentation method based on a hybrid architecture of CNN and ViT according to claim 1, characterized in that: In step S2, the edge uncertainty guidance module optimizes and reshapes the uncertainty pixels in the boundary fuzzy area in the coarse segmentation result, specifically by combining shallow features, uncertainty mapping and local feature matching to generate a refined mask, optimize the coarse segmentation result, and obtain the final predicted segmentation result.

6. The medical image segmentation method based on the hybrid architecture of CNN and ViT according to claim 5, characterized in that: The shallow features are the output of the first encoding layer; The uncertainty mapping is calculated based on the probability map of the rough segmentation result, and the formula is as follows: ; in, For uncertain mapping, the value range is ; , Represent the maximum and minimum probability values ​​of the pixels respectively. Calculate for exponential function; The local feature is generated by weighted local averaging of the uncertain mapping, and the formula is as follows: ; ; in, is a local feature, represents the normalization factor, is a small value that prevents the denominator from reaching zero, is a pixel, is the pixel after refinement, The pixel belongs to the target category The probability of It is a shallow feature. Pixel The neighborhood area; Refined masks are generated by combining shallow features, uncertainty mapping, and local feature matching , the calculation formula is as follows: ; ; in, For pixels The confidence score at is the confidence module; The rough segmentation result is optimized to obtain the final predicted segmentation result. The formula is as follows: ; in, To predict the segmentation result.

7. The medical image segmentation method based on the hybrid architecture of CNN and ViT according to claim 6, characterized in that: The confidence module includes 3×3 convolution, Relu activation function and residual connection; wherein, 3×3 convolution is applied on the main branch for feature extraction, and 3×3 convolution is also applied on the residual branch for feature extraction, and finally the main branch output and the residual branch output are added element by element; the Relu activation function is applied before the convolution operation.

8. The medical image segmentation method based on a hybrid architecture of CNN and ViT according to claim 1, characterized in that: In step S4, the loss function of the segmentation model is calculated, and the formula is as follows: ; ; ; in, is the sum of all pixels in the input image, , denote the predicted value and the target value respectively. is the total loss, is the binary cross entropy loss, is the dice loss, , is the weight assigned to each loss.

9. A medical image segmentation system based on a hybrid architecture of CNN and ViT, characterized in that: include: The data acquisition and preprocessing module is used to acquire multiple types of medical image data sets, preprocess the original data, generate preprocessed images, and divide the processed data into training sets, validation sets, and test sets in proportion; A model building module, used to build a segmentation model based on a CNN and ViT hybrid architecture, wherein the segmentation model is based on a U-Net network and includes an encoder, a decoder, and an edge optimization module; The model training module is used to train the model based on the divided training data set, calculate the loss function of the medical image segmentation model, update the parameters using the Adam optimizer, and save the optimal model weights during the training process; The model testing and evaluation module is used to test the test set data using the trained segmentation model and evaluate the model performance; The storage module is used to store medical image data, model parameters and intermediate calculation results.

10. The medical image segmentation system based on the CNN and ViT hybrid architecture according to claim 9, characterized in that: The encoder constructed by the model building module includes four encoding layers, each of which includes a multi-scale feature extraction module and a semantic enhancement module; the decoder includes four decoding layers, each of which also includes a multi-scale feature extraction module and a semantic enhancement module; jump connections are used between the encoding layer and the decoding layer on the same layer, and frequency domain information is introduced at the bottleneck layer.

Citation Information

Patent Citations

  • CNN and Transform fusion-based colonoscope polyp image segmentation method

    CN115018824A

  • Transform and U-Net combined medical image liver segmentation method and system

    CN115965633A

  • Food image segmentation method based on discrete wavelet attention network

    CN116630964A

  • Medical image segmentation model construction method based on CNN and SwinTransform hybrid coding

    CN118521784A

  • Multi-modal image segmentation method based on multi-scale feature extraction and lossless information conversion

    CN118570466A

Cited By

  • Medical image segmentation method based on multi-scanning visual state space

    CN120953621A