Medical Image Segmentation Device Based on Global Information Perception

Through the global information-aware medical image segmentation device, the CAB and AAFM modules provide a semantic bridge when fusion of codec features is used, which solves the problem of lack of global interaction relationship modeling in the prior art, and realizes more precise medical image segmentation, especially the clear segmentation of tissues and organs in CT images.

CN116258933BActive Publication Date: 2025-07-29XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310238744.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-13
Publication Date
2025-07-29
Estimated Expiration
2043-03-13

AI Technical Summary

Technical Problem

Existing medical image segmentation methods are difficult to effectively capture long-distance dependence information, and ignore the global interaction between feature maps at different semantic levels, resulting in inaccurate segmentation results, especially in medical images with complex boundary textures, especially in CT images, which are difficult to segment tissues and organs.

Method used

Using a medical image segmentation device based on global information perception, the correlation between high and low dimensional feature maps is modeled through CAB with low computational complexity, combined with CAB and AAFM modules to provide a semantic bridge when encoding and decoding feature fusion, alleviate the semantic gap problem, and aligning the feature receptive fields at each level through hollow convolution to achieve adaptive fusion of multi-level features.

Benefits of technology

It improves the accuracy of medical image segmentation, especially in CT images with complex boundary textures, which can better segment the boundaries of tissues and organs and provide clearer segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116258933B_ABST
    Figure CN116258933B_ABST
Patent Text Reader

Abstract

The present invention provides a medical image segmentation device based on global information perception. A scanning head scans medical images; a memory stores medical images, a medical image sample set, and a medical image segmentation network model; a processing card uses the trained medical image segmentation network model to segment the medical images, obtaining a segmentation result map with clear boundaries. The medical image segmentation network model of the present invention models the per-pixel correlation between high- and low-dimensional feature maps through a CAB with low computational complexity, realizing seamless fusion of low-dimensional detailed information and high-dimensional semantic information during the feature encoding process; the CAB provides a semantic bridge during the encoding and decoding feature fusion to alleviate the semantic gap problem. In addition, the feature fusion module AAFM aligns the receptive fields of features at all levels through dilated convolution, and realizes calibration of significant regions of features at all levels in the spatial dimension through a feature fusion activation method. Therefore, the present invention can provide a more accurate segmentation result for medical images with complex boundary textures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical devices, and particularly relates to a medical image segmentation device based on global information perception. Background Art

[0002] With the development of technology, medical devices are widely used. Due to the large population base in China and the relatively high medical treatment pressure, good medical devices can reduce the reception pressure of hospitals and improve the work efficiency of doctors at the same time. Some medical devices provide more intuitive reference images for patients and doctors through medical images. And medical imaging is also an important discipline in the medical field.

[0003] In the field of medical imaging, there are many imaging devices, such as B-ultrasound devices, CT devices, and X-ray scanning devices. These devices obtain medical images through scanning, providing intuitive references for doctors. In CT images, different tissues and organs are presented with different CT values for doctors' reference. However, CT images are single-channel grayscale images, and the range of CT values is much larger than the human visual perception range, resulting in difficulties for doctors in differentiating adjacent tissues with blurred boundaries and similar visual features during the interpretation process. Therefore, image preprocessing is required to increase the inter-class difference between tissues and organs, providing richer information for subsequent processing. Currently, most deep learning-based medical image automatic segmentation methods mainly use U-Net as the basic network framework and introduce modules such as dense connections and attention mechanisms for improvement in the network. However, due to the local characteristics of convolutional calculations, the above methods are unable to capture long-range dependency information. Different organs exhibit variability in shape and size due to individual differences, have complex internal textures, and have blurred inter-class boundaries with surrounding tissues, requiring comprehensive consideration of global context information and local detail features to obtain accurate segmentation results. In recent years, some studies have introduced Transformer into the medical image segmentation task, using the multi-head self-attention mechanism to model feature context information. Cao et al. designed Swin-UNet, which uses Transformer to replace the convolutional module in U-Net for feature extraction, achieving accurate segmentation of abdominal CT images and cardiac MRI images. The UNTER proposed by Ali et al. samples 3D medical images into token sequences, uses Transformer to replace the encoder to enhance the network's context information modeling ability, and simultaneously uses skip connections to fuse multi-scale features for segmentation result prediction. Experiments show that UNTER has achieved excellent performance in brain tumor and spleen segmentation tasks. Some teams have attempted to combine the respective advantages of CNN and Transformer to improve the segmentation performance of the network model. Chen et al. combined Transformer and CNN, and the deep embedding of the encoder with the Transformer structure constitutes TransUNet, and its effectiveness was verified on abdominal CT datasets and cardiac MRI datasets. The MBT-Net proposed by Zhang et al. applies a hybrid residual transformer feature extraction module, giving full play to the advantages of convolutional calculations and transformers in local details and global semantics, and achieving accurate segmentation of corneal endothelial cells.

[0004] Both Swin-Unet designed by Cao et al. and UNTER proposed by Ali et al. extract image features based on a pure Transformer structure. However, the Transformer lacks the ability to model local detailed information and lacks the characteristics of translational invariance and inductive bias, resulting in relatively rough segmentation results for the edge details of the target area in the segmentation method based on a pure Transformer. In addition, the Transformer has high requirements for the storage space of the computer during the calculation process. Although the segmentation model combining CNN and Transformer reduces the computational burden by embedding the self-attention mechanism in the deep layer of the CNN, using the self-attention mechanism only in the deep layer of the CNN cannot model the context information of features such as shape and texture in the shallow fine-grained information. In addition, most current methods only focus on the global context relationship of the feature map itself and ignore the global interaction relationship between feature maps at different semantic levels. Modeling these global interaction relationships plays an important role in bridging the semantic gap between features of different semantic dimensions. Therefore, how to promote the seamless fusion of multi-level features by utilizing cross-scale dependence relationships, so as to better incorporate global and local information to enhance the representation ability of the medical image segmentation network, and design an image segmentation device or an image region detection device is a technical problem to be solved urgently. Summary of the Invention

[0005] To solve the above problems existing in the prior art, the present invention provides a medical image segmentation device based on global information perception. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0006] The present invention provides a medical image segmentation device based on global information perception, including:

[0007] A scanning head for collecting medical images of a user's predetermined part through scanning;

[0008] A memory connected to the scanning head through a communication line for storing the medical images, a medical image sample set of a predetermined part collected in advance, and a medical image segmentation network model based on global information perception constructed in advance;

[0009] A processing card disposed on a computing device for performing the following processes:

[0010] Iteratively training the medical image segmentation network model using the medical image sample set, and optimizing the weight parameters of the medical image segmentation network model using a defined deep supervision loss function and an optimizer during the training process to obtain an optimal weight parameter segmentation network model; taking the medical image as an image to be segmented and inputting it into the optimal weight parameter segmentation network model to obtain a segmentation result map with clear distinguishing boundaries;

[0011] A display, which communicates remotely with the computing device wirelessly or by wire, is configured to display a segmentation result map with clear boundaries.

[0012] The present invention provides a computing device for implementing the specific processes executed by a processing card.

[0013] The present invention provides a medical image segmentation device based on global information perception. It includes a scanning head for scanning and acquiring medical images of a user's predetermined part; a memory for storing medical images, a pre-acquired medical image sample set of the predetermined part, and a pre-constructed medical image segmentation network model based on global information perception; a processing card for segmenting the image to be segmented using the trained medical image segmentation network model to obtain a segmentation result map with clear boundaries, and displaying it through a display. The medical image segmentation network model of the present invention models the pixel-by-pixel correlation between high- and low-dimensional feature maps through a CAB with low computational complexity, achieving seamless fusion of low-dimensional detailed information and high-dimensional semantic information during the feature encoding process; on the other hand, it uses the CAB to provide a semantic bridge during the encoder-decoder feature fusion to alleviate the semantic gap problem. In addition, an independent feature fusion module AAFM is set outside the decoder to achieve adaptive fusion of multi-level features in the decoder, thereby providing comprehensive and rich basis for the prediction task. The AAFM aligns the receptive fields of each level of features through dilated convolution and calibrates the significant regions of each level of features in the spatial dimension through the way of feature fusion-activation. Therefore, the present invention can provide more accurate segmentation results for medical images with complex boundary textures.

[0014] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. Description of the Drawings

[0015] Figure 1 It is a schematic structural diagram of a medical image segmentation device based on global information perception of the present invention;

[0016] Figure 2 It is a schematic diagram of a medical image segmentation network model based on global information perception of the present invention;

[0017] Figure 3 It is a schematic diagram of a global enhancement convolution module of the present invention;

[0018] Figure 4 It is a schematic diagram of a global spatial attention module of the present invention;

[0019] Figure 5 It is a schematic diagram of a cross-attention mechanism module of the present invention;

[0020] Figure 6 It is a schematic diagram of a fully adaptive feature fusion module of the present invention. Detailed Description of the Embodiment

[0021] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0022] As Figure 1 shown, the present invention provides a medical image segmentation device based on global information perception, including:

[0023] A scanning head, configured to collect medical images of a user's predetermined part through scanning;

[0024] A memory, connected to the scanning head through a communication line, configured to store the medical images, a medical image sample set of a predetermined part collected in advance, and a medical image segmentation network model based on global information perception pre-constructed;

[0025] A processing card, disposed on a computing device, configured to perform the following processes:

[0026] Iteratively train the medical image segmentation network model using the medical image sample set, and optimize the weight parameters of the medical image segmentation network model using a defined deep supervision loss function and an optimizer during the training process to obtain an optimal weight parameter segmentation network model; use the medical images as images to be segmented, and input them into the optimal weight parameter segmentation network model to obtain a segmentation result map with clear separation boundaries;

[0027] Taking chest CT as an example, the present invention can collect chest CT images as the original data set, and delineate the thymic epithelial tumor area in the data set; use a three-channel pseudo-color image preprocessing method to map the original data set into a three-channel pseudo-color image data set. Divide the preprocessed data set into a training set and a test set according to a ratio of 4:1; set the initial learning rate of the network, the learning rate decay method, the number of network iterations, the optimization method and the optimizer; use the medical image sample set to train the network model, and use the test set images to evaluate the segmentation effect of the model after training.

[0028] A display, remotely communicating with the computing device wirelessly or by wire, configured to display the segmentation result map.

[0029] Embodiment 2

[0030] As an optional embodiment of the present invention, the processing card is further configured to:

[0031] Perform three-channel pseudo-color map preprocessing on each medical image sample in the medical image sample set to obtain a three-channel pseudo-color image corresponding to each medical image sample;

[0032] The three-channel pseudo-color map preprocessing process is as follows:

[0033] (1) Read each medical image sample in the original DICOM format from the memory, and map its pixel values to CT values in Hounsfield units;

[0034] (2) Based on the CT window technology, overlay the mediastinal window and the window corresponding to the predetermined part on each medical image sample respectively to obtain the mediastinal window image and the predetermined part window image;

[0035] It should be noted that: if the predetermined part is bone, the window corresponding to the predetermined part is the bone window; if it is the abdomen, it is the abdominal window; if it is the lungs, it is the lung window.

[0036] (3) Add the mediastinal window and the predetermined part window images obtained by overlaying pixel by pixel and take the average to obtain the average window image;

[0037] (4) Map the CT values of the mediastinal window image, the average window image and the predetermined part window image to the range of 0-255, and stack them in the channel dimension in sequence to obtain a three-channel pseudo-color image.

[0038] The present invention can integrate the imaging manifestations of thymic epithelial tumors and surrounding structures under the CT window of CT images by using the three-channel pseudo-color image preprocessing method, so as to highlight the intra-class features of thymic epithelial tumors and the inter-class differences from surrounding tissues, and provide reliable and rich information for the subsequent segmentation network model.

[0039] Reference Figure 2 As shown, the medical image segmentation network model stored in the memory is constructed based on an encoder-decoder structure, which includes an initial module containing a residual structure, an encoder, a decoder, a global attention module, an adaptive feature fusion module, and a segmentation result output layer;

[0040] Among them, the encoder in the encoder-decoder structure includes 4 global enhanced convolutional modules, the decoder includes 4 convolutional modules corresponding one by one to the four global enhanced convolutional modules, and there is a global spatial attention module between the global enhanced convolutional module in the encoder and the corresponding convolutional module in the decoder. The input end of the encoder is connected to the initial module, and the input end of the initial module inputs the image to be segmented; there are neural layers for max-pooling operation and bicubic interpolation between the initial module, the first global enhanced convolutional module to the fourth global enhanced convolutional module. There is a deconvolution layer between the fourth global enhanced convolutional module and the fourth global spatial attention module; there is a deconvolution layer between the i-th convolutional module in the decoder and the i-1-th global spatial attention module; the output ends of the four convolutional modules of the decoder are all connected to the adaptive feature fusion module, and the output of the adaptive feature fusion module is connected to the input of the segmentation result output layer.

[0041] The initial module is used to map the image to be segmented from the image space to the feature space;

[0042] The neural layer for max pooling operation is used to perform max pooling operation on the image features of the previous layer and send them to the next global enhancement convolutional module;

[0043] The neural layer for bicubic linear interpolation is used to perform bicubic linear interpolation on the image features of the previous layer and send them to the next global enhancement convolutional module;

[0044] Each global enhancement convolutional module is used to perform global information modeling based on the image features obtained by max pooling operation and the image features obtained by bicubic linear interpolation, and output to the global spatial attention module;

[0045] Each global spatial attention module is used to supplement the low-dimensional detailed features in the global information to the high-dimensional semantic features in the decoder in a semantically consistent manner;

[0046] Each convolutional module in the decoder is used to perform convolution on the image features output by the global spatial attention module and send them to the adaptive feature fusion module;

[0047] The transposed convolution layer is used to perform transposed convolution on the input image features and send them to the connected global spatial attention module;

[0048] The adaptive feature fusion module is used to fuse the image features output by all convolutional modules and input them to the output layer;

[0049] The segmentation result output layer is used to output the channel fusion image features from multiple feature channels, and each channel corresponds to a segmentation target category.

[0050] The number of feature channels of the segmentation result output layer of the present invention is the number of segmentation target categories + 1, and the convolutional layer for segmenting the target / background can be used as the final segmentation result output layer.

[0051] Embodiment III

[0052] As an optional embodiment of the present invention, referring to Figure 3 and Figure 5 as shown, the global enhancement convolutional module includes a neural layer for bicubic linear interpolation, a neural layer for max pooling operation, a residual module, a first cross-attention module, and a self-attention module;

[0053] The neural layer for bicubic linear interpolation downsamples the feature matrix F of the previous level by a factor of two using bicubic linear interpolation to obtain the feature map X h ;

[0054] The neural layer of the max pooling operation performs a max pooling operation on the feature matrix F of the previous level to obtain the feature map X r ’;

[0055] The residual module models the significant information in the feature map X r ’ to obtain the semantic feature X r ;

[0056] The first cross-attention module uses the cross-attention mechanism to calculate the global dependency between X r and X h ; and concatenates with X r in the channel dimension; and uses a 1×1 convolution operation for feature fusion and channel dimension reduction to obtain the feature map X;

[0057] The self-attention module uses the self-attention mechanism to model the global information of the feature map X:

[0058]

[0059] Among them, the feature map is grouped and calculated according to the channel dimension. The size of a feature map is W×H×C, where C is the number of channels. C is divided into 4 groups for separate calculation, d = C / 4; if divided into 1 group, d = C; the feature value at position i in the feature map X is denoted as x i , and the feature value at position j is denoted as x j ;

[0060] And output the global information to the connected global spatial attention module.

[0061] Example 4

[0062] As an optional embodiment of the present invention, referring to Figure 4 and Figure 5 shown, the global spatial attention module includes a second cross-attention module and a spatial attention module;

[0063] The second cross-attention module is used to calculate the global dependency between and using the cross-attention mechanism to obtain CA(D,L);

[0064] Among them, D is a low-semantic dimension feature map containing fine-grained detail information from the encoder, and L is a high-semantic dimension feature map containing coarse-grained semantic information from the decoder;

[0065] The spatial attention module uses the spatial attention mechanism to highlight the regions related to the segmentation target in CA(D,L) in the spatial dimension:

[0066]

[0067]

[0068] Among them, ω ψ , ω x and ω g are three linear transformations, b g and b ψ are the corresponding bias values, σ1 and σ2 are the ReLu and Sigmoid activation functions respectively, represents dot product, represents dot addition;

[0069] Superimpose AT(D,L) and L in the channel dimension to obtain the output result, and send the output result to the connected convolution module;

[0070] HA(D,L) = AT(AT(D,L),)

[0071] Among them, CAT(…) represents concatenating the feature matrices in the channel dimension.

[0072] The calculation formula of the cross-attention mechanism is:

[0073]

[0074] Among them, X h and X r are feature matrices at different semantic levels in the segmentation network respectively; Q(·), K(·), V(·) are three 1×1 convolution operations, used to represent the information at each coordinate point in the feature matrix; and are coordinate matrices, used to supplement the coordinate point position information in the cross-attention calculation process; represents downsampling the feature matrix in the spatial dimension; d is the depth of the channel dimension of the feature matrix.

[0075] The global enhanced convolution module overcomes the limitation of the lack of global context model ability in the convolution operation and realizes the effective and tight fusion of high-dimensional and low-dimensional features. The global spatial attention perception module establishes a semantic bridge by aligning the information in the encoder-decoder features through CAB, thus realizing the effective fusion of the encoder-decoder features. In addition, the spatial attention mechanism module assigns higher weights to the task-related regions, so that the feature map can provide information for the target task more purposefully.

[0076] The cross-attention mechanism proposed by the present invention can generalize the explicit modeling method used in calculating long-distance dependencies to the representation of global correlation relationships between different-dimensional features. The establishment of a semantic bridge can be achieved by modeling the per-pixel mutual relationship between different feature maps. At the same time, the receptive fields can be aligned and the significant regions can be corrected during multi-level feature fusion, which will provide reliable information for the accurate prediction of the semantic category at each pixel point.

[0077] Example Five

[0078] Reference Figure 6 As shown, the adaptive feature fusion module includes an interpolation sampling module, four channel attention modules, four dilated attention modules, an activation gate module, and a concatenation module;

[0079] The interpolation sampling module performs three interpolation upsamplings on four feature maps DF = {df1, df2, …, df n}, (n = 4) to unify the sizes of the feature maps, and outputs the four unified-size feature maps DF U = {df1 U , df2 U , …, df n U}, (n = 4) to the channel attention module;

[0080] The channel attention module corrects the channels of the four feature maps to obtain the feature map DF U—SE ;

[0081] The dilated attention module corrects the feature map DF U—SE in the spatial dimension through dilated convolutions with different dilation rates to obtain the feature map M;

[0082] The activation gate module performs element-wise addition on M = {m1, m2, …, m n} in the channel dimension, and uses the ReLU and Sigmoid functions to obtain a significant region attention map, and multiplies it element-wise with DF U—SE to obtain A;

[0083]

[0084]

[0085] where A = {a1, a2, …, a n};

[0086] The concatenation module stacks the feature matrices output by each activation gate module in the channel dimension to obtain AC, and outputs AC to the segmentation result output layer.

[0087] The adaptive feature fusion module of the present invention adaptively fuses multi-level features in an interactive manner, making full use of the complementary information of features in different dimensions.

[0088] Example Six

[0089] As an optional embodiment of the present invention, the process of defining the deep supervision loss function is as follows:

[0090] (1) Use cross entropy loss and dice loss to construct a preliminary loss function:

[0091] L=L dice +L ce

[0092]

[0093]

[0094] Among them gt i With p i The gold standard and segmentation network prediction results are outlined, L ce With L dice They are cross entropy loss function and dice loss function respectively;

[0095] (2) Constructing a deep supervision loss function:

[0096]

[0097] Among them, L i is DF={df1,df2,…,df n},(n=4) The loss value of the segmentation result, L A is the loss value of the segmentation result obtained by AC, α i and β are weight coefficients.

[0098] The present invention provides a computing device for implementing a specific process executed by a processing card.

[0099] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0100] Although the present application is described herein with reference to various embodiments, those skilled in the art will be able to understand and implement other variations of the disclosed embodiments in practicing the claimed application by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality.

[0101] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A medical image segmentation device based on global information perception, characterized in that, Comprising: A scanning head for acquiring medical images of a predetermined part of a user by scanning; A memory connected to the scanning head through a communication line for storing the medical images, a set of medical image samples of a predetermined part acquired in advance, and a pre-constructed medical image segmentation network model based on global information perception; the medical image segmentation network model based on global information perception stored in the memory is constructed based on an encoder-decoder structure, which includes an initial module containing a residual structure, an encoder, a decoder, a global attention module, an adaptive feature fusion module, and a segmentation result output layer; Among them, the encoder in the encoder-decoder structure includes 4 global enhancement convolution modules, the decoder includes 4 convolution modules corresponding one by one to the four global enhancement convolution modules, and there is a global spatial attention module between the global enhancement convolution module in the encoder and the corresponding convolution module in the decoder. The input end of the encoder is connected to the initial module, and the input end of the initial module inputs the image to be segmented; there are neural layers for performing max pooling operations and bicubic bilinear interpolation between the initial module, the first to the fourth global enhancement convolution modules. There is a deconvolution layer between the fourth global enhancement convolution module and the fourth global spatial attention module; there is a deconvolution layer between the i-th convolution module in the decoder and the (i - 1)-th global spatial attention module; the output ends of the four convolution modules of the decoder are all connected to the input of the adaptive feature fusion module, and the output of the adaptive feature fusion module is connected to the input of the segmentation result output layer; The global enhancement convolution module includes a neural layer for bicubic bilinear interpolation, a neural layer for max pooling operation, a residual module, a first cross-attention module, and a self-attention module; The neural layer for bicubic linear interpolation takes the feature matrix from the previous level F and performs bicubic linear interpolation for downsampling by a factor of two to obtain the feature map ; The neural layer of the max pooling operation processes the feature matrix from the previous level F to perform the max pooling operation and obtain the feature map ; Residual module, which models the significant information in the feature map to obtain semantic features ; Cross-attention module, which calculates the global dependency relationship between and using the cross-attention mechanism; concatenates with along the channel dimension; and performs feature fusion and channel dimension reduction using a 1×1 convolution operation to obtain the feature map ; Self-attention module, using the self-attention mechanism to model the global information of the feature map : ; Among them, the feature maps are grouped and calculated according to the channel dimension. The size of a feature map is W×H×C, where C is the number of channels. C is divided into 4 groups for separate calculations. If d = C / 4, and if it is divided into 1 group, d = C; the feature value at position i in the feature map X is denoted as , and the feature value at position j is denoted as ; and are the coordinate matrices; represents downsampling the feature matrix in the spatial dimension; And output the global information to the connected global spatial attention module; A processing card, arranged on a computing device, for performing the following processes: Iteratively train the medical image segmentation network model using the medical image sample set, and optimize the weight parameters of the medical image segmentation network model using the defined deep supervision loss function and optimizer during the training process to obtain an optimal weight parameter segmentation network model; use the medical image as the image to be segmented and input it into Most the optimal weight parameter segmentation network model to obtain a segmentation result map with clear distinguishing boundaries; A display, remotely communicating with the computing device wirelessly or wiredly, for displaying a segmentation result map with clear boundaries.

2. The medical image segmentation device based on global information perception according to claim 1, wherein The processing card is further used for: Performing three-channel pseudo-color map preprocessing on each medical image sample in the medical image sample set to obtain a three-channel pseudo-color image corresponding to each medical image sample; The process of three-channel pseudo-color map preprocessing is as follows: (1) Read each medical image sample in the original DICOM format from the memory and map its pixel values to CT values in the Hounsfield unit; (2) Based on the CT window technology, superimpose the mediastinal window and the predetermined part window on each medical image sample respectively to obtain a mediastinal window image and a predetermined part window image; (3) Add the mediastinal window and the predetermined part window images obtained by superimposition pixel by pixel and take the average to obtain an average window image; (4) Map the CT values of the mediastinal window image, the average window image, and the predetermined part window image to the range of 0 - 255, and stack them in the channel dimension in sequence to obtain a three-channel pseudo-color image.

3. The medical image segmentation device based on global information perception according to claim 1, characterized in that, The initial module is used to map the image to be segmented from the image space to the feature space; A neural layer for max pooling operation, which is used to perform max pooling operation on the image features of the previous layer and send them to the next global enhancement convolutional module; A neural layer for bicubic bilinear interpolation, which is used to perform bicubic bilinear interpolation on the image features of the previous layer and send them to the next global enhancement convolutional module; Each global enhancement convolutional module is used to perform global information modeling based on the image features obtained by max pooling operation and the image features obtained by bicubic bilinear interpolation, and output to the global spatial attention module; Each global spatial attention module is used to supplement the low-dimensional detailed features in the global information to the high-dimensional semantic features in the decoder in a semantically consistent manner; Each convolutional module in the decoder is used to perform convolution on the image features output by the global spatial attention module and send them to the adaptive feature fusion module; A transposed convolutional layer is used to perform transposed convolution on the input image features and send them to the connected global spatial attention module; The adaptive feature fusion module is used to fuse the image features output by all convolutional modules and input them to the output layer; The segmentation result output layer is used to output the channel fusion image features from multiple feature channels, and each channel corresponds to a segmentation target category.

4. The medical image segmentation device based on global information perception according to claim 3, characterized in that, The global spatial attention module includes a second cross-attention module and a spatial attention module; A second cross-attention module for calculating the global dependencies between ; Among them, is a low-semantic-dimension feature map containing fine-grained detail information from the encoder, is a high-semantic-dimension feature map containing coarse-grained semantic information from the decoder; Spatial attention module, using spatial attention mechanism to highlight in the spatial dimension the regions related to the segmentation target in: Among them, , and are three linear transformations, and are the corresponding bias values, and are the ReLu and Sigmoid activation functions respectively, represents dot product, represents dot addition; Combine with in the channel dimension to obtain an output result, and send the output result to the connected convolutional module; Among them, represents concatenating the feature matrices in the channel dimension.

5. A medical image segmentation device based on global information perception according to claim 1 or 4, characterized in that, The calculation formula of the cross-attention mechanism is: Among them, , , are three 1×1 convolution operations, used to represent the information at each coordinate point in the feature matrix; and are coordinate matrices, used to supplement the coordinate point position information in the cross-attention calculation process; indicates downsampling the feature matrix in the spatial dimension; is the depth of the channel dimension of the feature matrix.

6. The medical image segmentation device based on global information perception according to claim 4, wherein The adaptive feature fusion module includes an interpolation sampling module, four channel attention modules, four dilated attention modules, an activation gate module, and a concatenation module; Interpolation sampling module, for four feature maps with different sizes Perform three times of interpolation upsampling to unify the size of the feature maps, and output the four feature maps with unified size To the channel attention module; Channel attention module, which corrects the channels of four feature maps to obtain a feature map ; Hole attention module, for the feature map Perform correction in the spatial dimension through dilated convolutions with different dilation rates to obtain the feature map ; Activation gate module, which performs point addition in the channel dimension, and uses ReLU and Sigmoid functions to obtain a significant region attention map, and multiplies it with to obtain ; ; Among them, ; The concatenation module obtains, by stacking in the channel dimension, the feature matrices output by each activation gate module , and outputs the AC to the segmentation result output layer.

7. The medical image segmentation device based on global information perception according to claim 6, characterized in that The process of defining the deep supervision loss function is: (1) Use cross-entropy loss and dice loss to construct a preliminary loss function: in and The gold standard and segmentation network prediction results are outlined respectively. and They are cross entropy loss function and dice loss function respectively; (2) Construct the deep supervision loss function: Among them, is the segmentation result obtained by and the loss value calculated by the preliminary loss function, is the loss value of the segmentation result obtained by and and are weight coefficients.

8. A computing device, characterized in that, The specific process executed by the medical image segmentation device based on global information perception for implementing any one of claims 1 to 7.

Citation Information

Patent Citations

  • Medical image segmentation method based on deep learning

    CN111145170A

  • Polyp segmentation method combining attention U-shaped network and multi-scale feature fusion

    CN114820635A