A lightweight segmentation method and system for medical images
Through downsampling of multimodal medical images, multi-scale feature extraction and fusion, the problem of manual diagnosis error is solved, automated lesion or organ segmentation is realized, and diagnostic accuracy is improved.
Patent Information
- Application Number
- CN202510833286.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
In the existing medical imaging diagnosis, the reliance on manual diagnosis leads to errors, affecting the diagnostic accuracy.
By obtaining multimodal medical images, downsampling is performed to obtain shallow and deep feature maps, multi-scale feature extraction and fusion are performed, combined with upsampling and feature fusion, segmentation results are generated to automatically identify the size, shape and location of the target lesions or organs.
It realizes automatic segmentation of medical images, reduces manual errors, and improves diagnostic accuracy.
Smart Images

Figure CN120339306B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer model technology, and in particular to a lightweight segmentation method and system for medical images. Background Art
[0002] With the rapid development of medical devices, doctors can now diagnose and treat diseases through medical imaging. For example, in the diagnosis and treatment of tumors, doctors can use CT (Computed Tomography), MRI (Nuclear Magnetic Resonance Imaging), and PET (Positron Emission Tomography). However, while a variety of medical imaging techniques can be used for disease diagnosis, the diagnosis still relies on the doctor's experience to identify information such as the tumor's location, which can affect diagnostic accuracy due to human error. Summary of the Invention
[0003] The purpose of the embodiments of the present application is to provide a lightweight segmentation method and system for medical images to solve the problem of errors in manual diagnosis during medical image diagnosis. The specific technical solution is as follows:
[0004] In a first aspect of an embodiment of the present application, a lightweight segmentation method for medical images is provided, the method comprising:
[0005] Acquire multimodal medical images to be identified;
[0006] Downsampling the multimodal medical image to be identified to obtain a shallow feature map and a deep feature map;
[0007] Extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features;
[0008] Upsampling the multi-scale features and fusing them with the shallow feature map to obtain features to be segmented;
[0009] Segmentation is performed according to the segmentation features to obtain a segmentation result, wherein the segmentation result includes size, shape and position information of the target lesion or organ.
[0010] In a possible implementation, obtaining a multimodal medical image to be identified includes:
[0011] Acquiring an original multimodal image, wherein the original multimodal image includes at least one of a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, and a positron emission tomography (PET) image;
[0012] spatially registering the original multimodal image with a preset template to obtain a registered multimodal image;
[0013] The registered multimodal image is normalized and resized to obtain the multimodal medical image to be identified.
[0014] In a possible implementation, downsampling the multimodal medical image to be identified to obtain a shallow feature map and a deep feature map includes:
[0015] Extracting shallow features of the multimodal medical image to be identified through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the first sampling layer to obtain the shallow feature map;
[0016] Through the dynamic kernel mixer and dual-domain collaborative attention mechanism module in the second sampling layer, deep features of the multimodal medical image to be identified are extracted to obtain the deep feature map.
[0017] In a possible implementation, extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features includes:
[0018] Extracting multi-scale context features from the deep feature map by extending the channel compression mixer;
[0019] The extracted context features are spliced to obtain the multi-scale features.
[0020] In a possible implementation, performing segmentation according to the segmentation feature to obtain a segmentation result includes:
[0021] generating a probability map according to the segmentation features through an activation function;
[0022] Generate a binary mask based on the probability map by presetting a threshold corresponding to the lesion category;
[0023] According to the binary mask, the size and / or position information of the target lesion is obtained through dilation and / or erosion.
[0024] A second aspect of the embodiments of the present application provides a lightweight segmentation system for medical images, the system comprising:
[0025] An image acquisition module, used to acquire multimodal medical images to be identified;
[0026] A downsampling module, configured to downsample the multimodal medical image to be identified to obtain a shallow feature map and a deep feature map;
[0027] A feature fusion module is used to extract and fuse multi-scale features from the deep feature map to obtain multi-scale features;
[0028] An upsampling module is used to upsample the multi-scale features and fuse them with the shallow feature map to obtain the features to be segmented;
[0029] The feature segmentation module is used to perform segmentation according to the segmentation features to obtain a segmentation result, wherein the segmentation result includes the size, shape and position information of the target lesion or organ.
[0030] In one possible embodiment, the image acquisition module is specifically configured to acquire an original multimodal image, wherein the original multimodal image includes at least one of a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, and a positron emission tomography (PET) image; spatially register the original multimodal image with a preset template to obtain a registered multimodal image; and normalize and resize the registered multimodal image to obtain the multimodal medical image to be identified.
[0031] In one possible embodiment, the downsampling module is specifically used to extract shallow features of the multimodal medical image to be identified through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the first sampling layer to obtain the shallow feature map; and to extract deep features of the multimodal medical image to be identified through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the second sampling layer to obtain the deep feature map.
[0032] In a possible implementation, the feature fusion module is specifically configured to extract multi-scale context features from the deep feature map by extending a channel compression mixer; and to splice the extracted context features to obtain the multi-scale features.
[0033] In one possible embodiment, the feature segmentation module is specifically used to generate a probability map based on the segmentation features through an activation function; generate a binary mask based on the probability map through a preset threshold corresponding to the lesion category; and obtain the size and / or location information of the target lesion through expansion and / or corrosion based on the binary mask.
[0034] Another aspect of the present application provides an electronic device, including:
[0035] Memory for storing computer programs;
[0036] The processor is configured to implement any of the above-mentioned lightweight segmentation methods for medical images when executing the program stored in the memory.
[0037] In another aspect of an embodiment of the present application, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, any of the above-mentioned lightweight segmentation methods for medical images is implemented.
[0038] In another aspect of the embodiments of the present application, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute any of the above-mentioned lightweight segmentation methods for medical images.
[0039] Beneficial effects of the embodiments of the present application:
[0040] The embodiment of the present application provides a lightweight segmentation method and system for medical images, the method comprising: obtaining a multimodal medical image to be identified; downsampling the multimodal medical image to be identified to obtain a shallow feature map and a deep feature map; extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features; upsampling the multi-scale features and fusing them with the shallow feature map to obtain features to be segmented; segmenting according to the segmentation features to obtain a segmentation result, wherein the segmentation result includes the size, shape and position information of the target lesion or organ. Through the scheme of the embodiment of the present application, after obtaining the multimodal medical image to be identified, the deep feature map and the shallow feature map can be obtained by downsampling, and then the multi-scale features are extracted and fused to obtain multi-scale features, and then the features to be segmented are obtained by upsampling and feature fusion, so as to obtain the size and / or position information of the target lesion through feature segmentation. Therefore, not only can automated segmentation be achieved through the model, but also the size and / or position information of the target lesion can be obtained.
[0041] Of course, it is not necessary to achieve all the advantages described above at the same time when implementing any product or method of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.
[0043] Figure 1 A schematic diagram of a flow chart of a lightweight segmentation method for medical images provided in an embodiment of the present application;
[0044] Figure 2 A schematic diagram of a process for obtaining multimodal images provided in an embodiment of the present application;
[0045] Figure 3An architecture diagram of LM-UNet provided in an embodiment of the present application;
[0046] Figure 4 A generation flow chart of a dynamic core mixer provided in an embodiment of the present application;
[0047] Figure 5 A module architecture diagram of the ECCM provided in an embodiment of the present application;
[0048] Figure 6 An architectural diagram of a DDC provided in an embodiment of the present application;
[0049] Figure 7 An architectural diagram of the ACI-L provided in an embodiment of the present application;
[0050] Figure 8 A schematic diagram of a feature segmentation process provided in an embodiment of the present application;
[0051] Figure 9 A schematic structural diagram of a lightweight medical image segmentation system provided in an embodiment of the present application;
[0052] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field based on this application are within the scope of protection of this application.
[0054] In a first aspect of the embodiments of the present application, a lightweight segmentation method for medical images is first provided. Figure 1 , Figure 1 A schematic flow chart of a lightweight segmentation method for medical images provided in an embodiment of the present application, the method comprising:
[0055] Step S11, obtaining a multimodal medical image to be identified;
[0056] Step S12, downsampling the multimodal medical image to be identified to obtain a shallow feature map and a deep feature map;
[0057] Step S13, extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features;
[0058] Step S14, upsampling the multi-scale features and fusing them with the shallow feature map to obtain features to be segmented;
[0059] Step S15 , performing segmentation according to the segmentation features to obtain a segmentation result, wherein the segmentation result includes size, shape and position information of the target lesion or organ.
[0060] Corresponding to step S11 above, the multimodal medical images in the embodiments of the present application may include multiple types of medical images. Specifically, they may include CT (Computed Tomography), MRI (Nuclear Magnetic Resonance Imaging), and PET (Positron Emission Tomography). It should be noted that the multimodal medical images in the embodiments of the present application are medical images targeting the same lesion location. In one example, when the solution of the embodiments of the present application is applied to tumor diagnosis, the multimodal medical images are multimodal medical images targeting the location of the tumor.
[0061] Corresponding to the above-mentioned step S12, in an embodiment of the present application, the multimodal medical image to be identified is downsampled, and downsampling can be performed through multiple downsampling layers to obtain shallow feature maps and deep feature maps. Specifically, high-resolution image features can be collected through one sampling layer, and then low-resolution image features can be collected through another sampling layer. In one case, the two sampling layers in the embodiment of the present application can include DKM (dynamic kernel mixer) and DDC (dual-domain collaborative attention mechanism), and downsampling is performed through the dynamic kernel mixer and the dual-domain collaborative attention mechanism.
[0062] Corresponding to step S13 above, when extracting and fusing multi-scale features from the deep feature map, multi-scale features can be first extracted from the deep feature map. Specifically, the extraction and fusion can be performed using ECCM (Extended Channel Compression Mixer). In one embodiment, when extracting and fusing multi-scale features, the deep feature map can first be convolved with empty segments, and then the convolution results can be fused to obtain scale features. Feature fusion can be performed using various methods, such as feature concatenation and summation.
[0063] Corresponding to step S14 above, when upsampling the multi-scale features and fusing them with the shallow feature map, the multi-scale features can be upsampled first, and then the upsampling results are fused with the shallow feature map. During upsampling, the multi-scale features can be restored to a feature map of a preset size, and then spliced with the shallow feature map to obtain the features to be segmented. In actual use, the features to be segmented can refer to logits, which are the input of the segmentation layer and the output of the previous layer of the segmentation layer. The segmentation layer can perform segmentation based on the logits to obtain the segmentation result.
[0064] Corresponding to step S15 above, the segmentation result includes size, shape, and location information of the target lesion or organ. For example, the segmentation result may be the size, shape, and location information of an organ such as the liver in an image, or the size, shape, and location information of a lesion such as a tumor. When performing segmentation based on the segmentation features, feature segmentation can be achieved by generating a probability map and performing binary masking using a threshold, thereby obtaining the size and / or location information of the target lesion.
[0065] It can be seen that through the solution of the embodiment of the present application, after obtaining the multimodal medical image to be identified, deep feature maps and shallow feature maps can be obtained by downsampling, and then multi-scale features can be extracted and fused to obtain multi-scale features. Then, features to be segmented can be obtained through upsampling and feature fusion, and the size and / or location information of the target lesion can be obtained through feature segmentation. In this way, not only can automated segmentation be achieved through the model, but also the size and / or location information of the target lesion can be obtained.
[0066] In a possible implementation, the multimodal medical image to be identified is obtained, see Figure 2 , Figure 2 A schematic diagram of a process for obtaining multimodal images provided in an embodiment of the present application includes:
[0067] Step S21, acquiring an original multimodal image, wherein the original multimodal image includes at least one of a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, and a positron emission tomography (PET) image;
[0068] Step S22, spatially registering the original multimodal image with the preset template to obtain a registered multimodal image;
[0069] Step S23 , normalizing and resizing the registered multimodal image to obtain the multimodal medical image to be identified.
[0070] The original multimodal images in the embodiments of the present application include at least one of: computed tomography (CT) images, magnetic resonance imaging (MRI) images, and positron emission tomography (PET) images. Specifically, the original multimodal images, such as HU value (Houinsky unit) sequences of CT and multi-contrast sequences of MRI, can be stored in DICOM (Digital Imaging and Communications in Medicine) format. The corresponding data structure can be a three-dimensional array I∈R H×W×Cm , where H / W is the spatial size; C m is the modality number, such as C for MRI with four modalitiesm =4; R is a real number.
[0071] In one case, the original multimodal image and the preset template are spatially registered to obtain the registered multimodal image. The SyN (a typical denial of service attack) algorithm of the ANTs (a database) library can be used to align the multimodal image to a standard template, such as the MNI152 (a map) brain map, and output the registered image I reg ∈R H×W×Cm , ensuring spatial consistency of anatomical structures.
[0072] In one case, when normalizing and resizing the registered multimodal image, the normalization can be performed using a preset specification. For example, for a CT image, the HU value can be truncated to [-200, 300] and normalized to [0, 1], and the formula is: norm =clip(I reg ,-200,300)-(-200) / 500, where clip is the loss function; in another example, for MRI data, Z-Score normalization can be used to eliminate field strength inhomogeneity. In the embodiment of the present application, when resizing, the image can be resized to a fixed size H′×W′×D′, such as 256×256×64, 3D (three-dimensional) image, thereby outputting a four-dimensional tensor X∈R B×Cm×H′×W′×D′ , where H′, W′, and D′ represent height, width, and depth, respectively; B is the batch size. In actual use, the CPU (Central Processing Unit) can execute the registration logic, while the GPU (Graphics Processing Unit) can accelerate interpolation and normalization calculations, such as tensor operations in PyTorch CUDA (a parallel computing platform and programming model for PyTorch). In this application, modality and dynamic image processing are decoupled from spatiotemporal features (optical flow motion compensation) through a modality-specific kernel generation mechanism (e.g., a PET metabolic zone weight of 0.72). This achieves a 3.1% cross-modality improvement and a 48.7% reduction in dynamic sequence error (e.g., cardiac MRI myocardial wall segmentation error reduced from 2.3 mm to 1.18 mm).
[0073] In a possible embodiment, the downsampling of the multimodal medical image to be identified to obtain a shallow feature map and a deep feature map includes: extracting shallow features of the multimodal medical image to be identified through a dynamic kernel mixer and a dual-domain collaborative attention mechanism module in a first sampling layer to obtain the shallow feature map; and extracting deep features of the multimodal medical image to be identified through a dynamic kernel mixer and a dual-domain collaborative attention mechanism module in a second sampling layer to obtain the deep feature map.
[0074] In an embodiment of the present application, the preprocessed multimodal tensor X can be input to an encoder based on nnUNet (a network structure), which includes two downsampling levels, namely the first sampling layer and the second sampling layer mentioned above. Each layer is embedded with a dynamic kernel mixer and a dual-domain collaborative attention mechanism. Specifically, shallow feature extraction can be achieved for the first sampling layer. Specifically, the global feature vector g∈R can be generated by embedding the dynamic kernel mixer input X0=X, compressing the spatial dimension to 1×1×D′ through GAP (global average pooling) B×Cm ; Then through the two-layer fully connected network (FC1:C m →4C m , FC2:4C m →κ×C k , κ=3 is the number of nuclear groups, C k =3 is the number of single-core channels) to generate κ groups of convolution kernel parameters W k ∈R κ×Ck×3×3 ; Then generate dynamic weight π∈R through Softmax (normalization) B×κ Among them, a depth-wise separable convolution (e.g., 3×3 kernel) is performed on X0, and the output feature Figure X 1∈R B×Cm×H′×W′×D′ Then, local spatial features can be extracted through the dual-domain collaborative attention mechanism and deep separable convolution (e.g., 3×3 kernel): X1′′=DWConv(X1), where DWConv is a separable convolution; and then the channel mixer is used: the channel is expanded to 4C through 1×1 convolution. m , GeLU (Gaussian Error Unit) is activated and compressed back to Cm, the formula is: X1′′=Conv 1x1 (GeLU(Conv1x1(X1′))), Conv convolution; residual connection and batch normalization: X1=BN(X1+X1′′), BN means standardization. For the second sampling layer, downsampling (step 2) will reduce the size by half, repeat the DKM+DDC operation, and finally output low-resolution features Figure X 2∈R B×2Cm×H′ / 2×W′ / 2×D′ / 2 Data structure changes: the number of channels doubles with downsampling (C m →2C m), the spatial size is halved, forming a hierarchical feature pyramid. Specifically, hardware association can use GPU parallel computing depth separable convolution and fully connected layers, using CUDA core to accelerate matrix operations. For an example, see Figure 3 , Figure 3 An architecture diagram of LM-UNet provided in an embodiment of the present application. The ordinary convolution of the LM-UNet model (a medical image segmentation model based on a hybrid multi-layer perceptron) is replaced with a DKM module, which can adaptively complete feature extraction. On top of this, the DDC attention mechanism is entered after the DKM. The attention weights of space and channels are generated. The ACI-L module (a feature extraction module) is used on the jump connection to enhance feature extraction. The ECCM channel mixer is proposed at the bottleneck layer to provide an efficient and accurate feature modeling solution for medical image segmentation tasks. In an example, see Figure 4 , where the working principle of dynamic convolution is as follows: the input features complete basic feature extraction and normalization through triple operations, and then enter the parallel convolution layer after processing, and generate differentiated feature maps by adaptively using different numbers of convolution kernels. Use the Dropout layer (dropout layer) to randomly discard to prevent the model from overfitting. Finally, the features are obtained by weighted summation. In this application, DKM generates κ groups (κ=3-5) of dynamic convolution kernel parameters based on global average pooling and a two-layer fully connected network, and realizes anatomical structure-sensitive feature extraction (such as the tumor core area weight π=0.62) through Softmax weight allocation, replacing the traditional fixed convolution kernel, adapting to the heterogeneous characteristics of tumors and dynamic organ morphological changes, and reducing the number of parameters.
[0075] In this application, the DKM module can generate κ groups (κ=3-5) of candidate convolution kernel parameters based on global average pooling and a two-layer fully connected network, and dynamically assign weights through Softmax. This technical feature directly affects the feature extraction of different anatomical structures: in the brain tumor segmentation scenario, due to the significant differences in image features between tumor tissue and normal tissue, traditional fixed convolution kernels are difficult to specifically extract features of the tumor core area; while DKM can dynamically assign weights to the tumor core area up to 0.62, and can adaptively adjust the receptive field to accurately capture tumor boundary details. Experimental data show that on the BraTS2021 (a dataset) dataset, the Dice coefficient of tumor core area segmentation is improved from 0.848 of the traditional method to 0.924 after adopting this module, an increase of 0.076. The principle is that the dynamic kernel can selectively enhance the response of the target area and reduce background noise interference based on the differences in texture, grayscale and other features of different tissues. In this application, the DDC module utilizes a channel-space decoupled parallel attention architecture. Through a two-dimensional matrix reshaping operation (H×W×C → (H×W)×C) and a dynamic sparsification computation strategy, it overcomes the quadratic complexity bottleneck of the traditional self-attention mechanism. For the liver vessel segmentation task, the traditional Transformer (a network model) architecture works by separating the channel dimension from the spatial dimension. The channel attention branch rapidly filters valid channel information through matrix operations, while the spatial attention branch precisely locates vascular structures using global pooling. The synergistic effect of these two approaches reduces the HD95 error of liver vessel segmentation from 4.83mm in traditional methods to 3.35mm, while also improving inference speed, achieving both improved global modeling efficiency and accuracy. In this application, the dual-domain collaborative attention module utilizes a channel-space decoupled parallel attention architecture, utilizing a two-dimensional matrix reshaping operation (H×W×C → (H×W)×C) and dynamic sparsification computation to achieve linear complexity global modeling. The computational complexity is reduced by 40% compared to traditional self-attention, while suppressing noise channels (attenuation coefficient 0.3-0.5) and enhancing anatomical structure response (gain coefficient 1.2-1.5).
[0076] In a possible implementation, the extraction and fusion of multi-scale features from the deep feature map to obtain multi-scale features includes: extracting multi-scale context features from the deep feature map through an extended channel compression mixer; and splicing the extracted context features to obtain the multi-scale features. Specifically, the processing object can be the low-resolution feature output by the encoder. Figure X 2. Input to ECCM (Extended Channel Compression Mixer). The specific steps may include: 1. Dilated convolution group, using 3D dilated convolution with dilation rate of 1 / 3 / 5 through parallel branches to extract multi-scale context and output X 2a ,X 2b ,X 2c∈R B×2Cm×H′ / 2×W′ / 2×D′ / 2 ; 2. Feature splicing and compression, by splicing multi-scale features: X concat =Concat(X 2a ,X 2b ,X 2c ), Concat means concatenation; 1×1 convolution compresses the channel to 2C m :Xbottleneck= Conv1x1(X concat ), where Xbottleneck is the convolution compression result and Conv represents convolution. The data structure can expand the receptive field through the dilated convolution, keep the spatial size unchanged, and achieve the fusion of "local details and global context". In an example, see Figure 5 The ECCM works as follows: Input data X first passes through a depthwise conv layer (depthwise separable convolution) to extract spatial features before entering the channel mixer. In the channel mixer, the first conv2d (two-dimensional convolution) expands the number of channels, the GeLU activation function introduces nonlinearity, and the second conv2d compresses the channels. The original input X is then added to the processed features via a residual connection, and finally batch normalization is performed to produce the output Y. This structure effectively reduces computational effort, accelerates convergence, and improves generalization.
[0077] In this application, the ECCM module utilizes depthwise separable convolution combined with a channel bottleneck architecture. By reducing the number of cross-channel interaction parameters, the module significantly reduces the number of model parameters while maintaining feature expression capabilities. In low-dose CT image segmentation, traditional models are prone to overfitting due to severe noise interference, resulting in decreased segmentation accuracy. The ECCM module, however, enhances the stability of feature propagation through residual connections and batch normalization, while also compressing redundant channel information, improving the model's noise immunity on low-dose CT data. In principle, depthwise separable convolution reduces redundant convolution operations, while the channel compression mechanism effectively suppresses noise transfer between channels, ultimately reducing the total number of model parameters and enabling efficient operation even on resource-constrained devices. In this application, an extended channel compression mixer is implemented, combining depthwise separable convolution with a "channel expansion-compression" mechanism (with an expansion factor of 4-8) to reduce the number of cross-channel interaction parameters. This achieves lightweight cross-channel feature interaction, improves feature propagation stability, and supports low-power deployment on edge devices.
[0078] In the embodiment of the present application, step S14 is to upsample the multi-scale features and fuse them with the shallow feature map to obtain the features to be segmented. This can achieve feature recovery and attention calibration. Specifically, the processing object can be the neck layer feature X bottleneckThe encoder cross-layer features X1, X0 are input to the decoder, and each layer is embedded with DDC (dual-domain collaborative attention module). Specifically, the processing steps may include: Level 1: upsampling and cross-layer fusion. Among them, upsampling (interpolation + convolution) converts X bottleneck The size is restored to H′ / 2×W′ / 2×D′ / 2, concatenated with encoder X2, and output X3∈R B×4Cm ×H′ / 2×W′ / 2×D′ / 2; then through the DDC module, the channel attention branch is used to achieve feature reshaping: X3′∈R B×(H′W′D′ / 2)×4Cm ; Channel interaction: ∈R B×4Cm×4Cm ; Spatial gating: G c =Sigmoid(Conv 1x1 (X3))∈R B ×4Cm×H′ / 2×W′ / 2×D′ / 2 ; Channel enhancement: , reshaped back to a 3D tensor. Spatial attention branch: global pooling: g s =GlobalPool(X3)∈R B×4Cm , GlobalPool is global pooling; weight generation: W s =Softmax(FC(g s ))∈R B×4Cm ; Space Enhancement: Unsqueeze(-1) Unsqueeze(-1), Unsqueeze means expansion; dual domain fusion: X3=Y c + Y s Level 2: Output layer feature optimization. You can repeat upsampling and DDC operations, merge with encoder X1, and finally output segmented logits through 1×1 convolution: Logits∈R B×Cout×H′×W′×D′ (Cout is the number of categories, such as brain tumor segmentation Cout = 3). Data structure changes: The spatial dimensions can be restored through upsampling and cross-layer splicing, and the DDC module dynamically calibrates feature responses in the channel and spatial dimensions. For an example, see Figure 6, where the DDC attention mechanism is as follows: the input X is divided into two paths. In the Channel-only self-attention module, the channel attention features are obtained through Conv(1x1), Reshape, multiplication, softmax and other operations; in the Spatial-only self-attention module, the spatial attention features are obtained through Conv(1x1), Global pooling, Reshape, multiplication, softmax and other operations. After the two features are fused, they are added to X to obtain the output Y, realizing the weighted attention of the channel and spatial dimensions. In an example, see Figure 7 Among them, the ACI-L module can input X and then branch it through a 1×1 convolution layer. One branch undergoes adaptive average pooling to change the shape of the array and tensor, and then convolves it through a 1×1 convolution layer. Then, the weight is generated by the Sigmoid function to change the shape of the array and tensor; it is multiplied element-by-element with the other feature, and then convolved through a 1×1 convolution layer to obtain the output Y, realizing feature weighted adjustment.
[0079] In this application, a modality-specific kernel parameter generation mechanism and a cross-channel mixer are designed for multimodal medical imaging. In the CT-MRI-PET multimodal brain tumor segmentation scenario, the traditional method uses a simple channel splicing method, which cannot fully utilize the complementary information of each modality. The present invention generates exclusive convolution kernels for different modalities through DKM. For example, a higher weight (0.72) is assigned to the metabolically active area of the PET image. At the same time, ECCM realizes the deep fusion of cross-modal features, which increases the cross-modal segmentation Dice coefficient from 0.869 of the traditional method to 0.900, an increase of 3.1%. The principle is that the modality-specific kernel can extract the key features of each modality in a targeted manner, and the cross-channel mixer realizes feature complementarity through information interaction, ultimately improving the generalization ability of the model on multi-center, multimodal data and reducing the errors caused by different equipment and imaging parameters. It can be seen that the method of the embodiment of the present application has constructed a complete technical system of "adaptive feature extraction - efficient global modeling - multimodal generalization - hardware collaborative acceleration" in the field of medical image segmentation, providing innovative solutions for precision medicine and real-time diagnosis, and has significant clinical application value and technical foresight.
[0080] In a possible implementation, segmentation is performed according to the segmentation feature to obtain a segmentation result. Figure 8 ,include:
[0081] Step S81, generating a probability map according to the segmentation features through an activation function;
[0082] Step S82, generating a binary mask according to the probability map by presetting a threshold corresponding to the lesion category;
[0083] Step S83: Obtain the size and / or location information of the target lesion by dilation and / or corrosion according to the binary mask.
[0084] In this embodiment of the present application, the logits tensor output by the decoder can have a data structure of R B ×Cout×H′×W′×D′ Specifically, the processing steps may include: probability map generation, specifically, the category probability distribution can be obtained by the Softmax activation function: Prob=Softmax(Logits,dim=1)∈R B×Cout×H′×W′×D′ ; Threshold segmentation, which can be achieved by applying an adaptive threshold (e.g., Otsu algorithm) to each category to generate a binary mask Mask∈R B×Cout×H′×W′×D′ dim=1 indicates row-wise operation; morphological optimization removes isolated noise points through dilation / erosion operations, outputting the final segmentation result. Specifically, hardware integration allows GPU-accelerated Softmax calculations, while the CPU performs morphological post-processing (lightweight operations). This application implements a multimodal processing pipeline from preprocessing (bias field correction and rigid registration) to multimodal fusion (feature concatenation + attention weighting) through a global feature-guided dynamic convolution kernel generation method (including global average pooling, fully connected networks, and Softmax weight assignment) and a channel-space decoupled feature calibration method (including matrix reshaping, channel weight generation, and spatial gating). In one example, the corresponding hardware module combination in this application includes an encoder-decoder architecture consisting of DKM, DDC, and ECCM, supporting multimodal input and dynamic image processing. This allows for optimized inference on edge devices based on TensorRT quantization. This enables organ segmentation (e.g., liver vasculature and brain tumors) and microlesion detection for multimodal medical imaging: CT / MRI / PET. And dynamic sequence analysis: cardiac MRI motion compensation, low-dose CT noise suppression and other clinical scenarios.
[0085] In a second aspect of the present application, a lightweight segmentation system for medical images is provided. Figure 9 , the system comprising:
[0086] Image acquisition module 901, used to acquire multimodal medical images to be identified;
[0087] A downsampling module 902 is configured to downsample the multimodal medical image to be identified to obtain a shallow feature map and a deep feature map;
[0088] A feature fusion module 903 is used to extract and fuse multi-scale features from the deep feature map to obtain multi-scale features;
[0089] An upsampling module 904 is configured to upsample the multi-scale features and fuse them with the shallow feature map to obtain features to be segmented;
[0090] The feature segmentation module 905 is configured to perform segmentation according to the segmentation features to obtain a segmentation result, wherein the segmentation result includes size, shape and position information of the target lesion or organ.
[0091] In one possible embodiment, the image acquisition module is specifically configured to acquire an original multimodal image, wherein the original multimodal image includes at least one of a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, and a positron emission tomography (PET) image; spatially register the original multimodal image with a preset template to obtain a registered multimodal image; and normalize and resize the registered multimodal image to obtain the multimodal medical image to be identified.
[0092] In one possible embodiment, the downsampling module is specifically used to extract shallow features of the multimodal medical image to be identified through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the first sampling layer to obtain the shallow feature map; and to extract deep features of the multimodal medical image to be identified through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the second sampling layer to obtain the deep feature map.
[0093] In a possible implementation, the feature fusion module is specifically configured to extract multi-scale context features from the deep feature map by extending a channel compression mixer; and to splice the extracted context features to obtain the multi-scale features.
[0094] In one possible embodiment, the feature segmentation module is specifically used to generate a probability map based on the segmentation features through an activation function; generate a binary mask based on the probability map through a preset threshold corresponding to the lesion category; and obtain the size and / or location information of the target lesion through expansion and / or corrosion based on the binary mask.
[0095] It can be seen that through the system of the embodiment of the present application, after acquiring the multimodal medical image to be identified, deep feature maps and shallow feature maps can be obtained by downsampling, and then multi-scale features can be extracted and fused to obtain multi-scale features. Then, features to be segmented can be obtained through upsampling and feature fusion, and the size and / or location information of the target lesion can be obtained through feature segmentation. In this way, not only can the model be used to achieve automated segmentation, but also the size and / or location information of the target lesion can be obtained.
[0096] The present application also provides an electronic device, such as Figure 10 Shown, including:
[0097] Memory 1001, used for storing computer programs;
[0098] The processor 1002 is configured to execute the program stored in the memory 1001 by performing the following steps:
[0099] Acquire multimodal medical images to be identified;
[0100] Downsampling the multimodal medical image to be identified to obtain a shallow feature map and a deep feature map;
[0101] Extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features;
[0102] Upsampling the multi-scale features and fusing them with the shallow feature map to obtain features to be segmented;
[0103] Segmentation is performed according to the segmentation features to obtain a segmentation result, wherein the segmentation result includes size and / or location information of the target lesion.
[0104] The communication bus mentioned in the electronic devices mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into address buses, data buses, control buses, etc. For ease of illustration, only a single thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.
[0105] The communication interface is used for communication between the above electronic device and other devices.
[0106] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0107] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0108] In another embodiment provided in the present application, a computer-readable storage medium is further provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of any of the above-mentioned lightweight segmentation methods for medical images are implemented.
[0109] In another embodiment provided by the present application, a computer program product including instructions is further provided, which, when executed on a computer, enables the computer to execute any of the lightweight segmentation methods for medical images in the above embodiments.
[0110] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or solid-state drive (SSD).
[0111] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0112] Each embodiment in this specification is described in a related manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system, electronic device, and storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For related portions, refer to the descriptions of the method embodiments.
[0113] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application are included in the scope of protection of the present application.
Claims
1. A lightweight segmentation method for medical images, characterized in that: The method comprises: Acquire multimodal medical images to be identified; Downsampling the multimodal medical image to be identified to obtain a shallow feature map and a deep feature map; Extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features; Upsampling the multi-scale features and fusing them with the shallow feature map to obtain features to be segmented; Performing segmentation according to the segmentation features to obtain a segmentation result, wherein the segmentation result includes size, shape, and position information of the target lesion or organ; The acquiring of the multimodal medical image to be identified includes: acquiring an original multimodal image, wherein the original multimodal image includes at least one of a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, and a positron emission tomography (PET) image; spatially registering the original multimodal image with a preset template to obtain a registered multimodal image; and normalizing and resizing the registered multimodal image to obtain the multimodal medical image to be identified; The downsampling of the multimodal medical image to be identified to obtain a shallow feature map and a deep feature map includes: extracting shallow features of the multimodal medical image to be identified through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the first sampling layer to obtain the shallow feature map; and extracting deep features of the multimodal medical image to be identified through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the second sampling layer to obtain the deep feature map.
2. The method according to claim 1, characterized in that The extracting and fusing multi-scale features of the deep feature map to obtain multi-scale features includes: Extracting multi-scale context features from the deep feature map by extending the channel compression mixer; The extracted context features are spliced to obtain the multi-scale features.
3. The method according to claim 1, characterized in that The performing segmentation according to the segmentation feature to obtain a segmentation result includes: generating a probability map according to the segmentation features through an activation function; Generate a binary mask based on the probability map by presetting a threshold corresponding to the lesion category; According to the binary mask, the size and / or position information of the target lesion is obtained through dilation and / or erosion.
4. A lightweight segmentation system for medical images, characterized in that: The system comprises: An image acquisition module, used to acquire multimodal medical images to be identified; A downsampling module, configured to downsample the multimodal medical image to be identified to obtain a shallow feature map and a deep feature map; A feature fusion module is used to extract and fuse multi-scale features from the deep feature map to obtain multi-scale features; An upsampling module is used to upsample the multi-scale features and fuse them with the shallow feature map to obtain the features to be segmented; a feature segmentation module, configured to perform segmentation according to the segmentation features to obtain a segmentation result, wherein the segmentation result includes size, shape, and position information of the target lesion or organ; The image acquisition module is specifically configured to acquire an original multimodal image, wherein the original multimodal image includes at least one of a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, and a positron emission tomography (PET) image; spatially register the original multimodal image with a preset template to obtain a registered multimodal image; and normalize and resize the registered multimodal image to obtain the multimodal medical image to be identified; The downsampling module is specifically used to extract shallow features of the multimodal medical image to be identified through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the first sampling layer to obtain the shallow feature map; and to extract deep features of the multimodal medical image to be identified through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the second sampling layer to obtain the deep feature map.
5. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 3 when executing a program stored in a memory.
6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
DCE-MRI breast tumor image segmentation method based on deep neural network
CN113139981A
Multi-level lesion detection optimized pathological section image segmentation detection method
CN116596952A