Lightweight segmentation method and system for medical image

Through downsampling of multimodal medical images, multi-scale feature extraction and fusion, the problem of manual diagnosis error is solved, and the automated and precise segmentation of lesions such as tumors is realized.

CN120339306AActive Publication Date: 2025-07-18厦门工学院

Patent Information

Application Number
CN202510833286.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

In the existing medical imaging diagnosis, there are errors in manual diagnosis, which affects the recognition accuracy of information such as tumor location.

Method used

By obtaining multimodal medical images, downsampling is performed to obtain shallow and deep feature maps, multi-scale feature extraction and fusion is performed, combining the extended channel compression mixer and the dual-domain collaborative attention mechanism to achieve feature segmentation and generate size, shape and position information of the target lesion or organ.

Benefits of technology

It realizes automatic segmentation of medical images, improves the diagnostic accuracy of tumors and other lesions, and reduces artificial errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339306A_ABST
    Figure CN120339306A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a lightweight segmentation method and system for a medical image. The method comprises the following steps: acquiring a to-be-identified multi-modal medical image; performing down-sampling on the to-be-identified multi-modal medical image to obtain a shallow feature map and a deep feature map; performing multi-scale feature extraction and fusion on the deep feature map to obtain multi-scale features; performing up-sampling on the multi-scale features, and fusing the multi-scale features with the shallow feature map to obtain features to be segmented; and performing segmentation according to the segmentation features to obtain a segmentation result. According to the method, the deep feature map and the shallow feature map can be obtained through down-sampling, then the multi-scale features are obtained through extraction and fusion of the multi-scale features, then the to-be-segmented features are obtained through up-sampling and feature fusion, and therefore the size and / or position information of a target focus is obtained through feature segmentation. Therefore, automatic segmentation can be realized through the model, and the size and / or position information of the target focus can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of computer models, and particularly to a lightweight segmentation method and system for medical images. Background Art

[0002] Currently, with the rapid development of medical devices, doctors can diagnose and treat diseases through medical images. For example, during the diagnosis and treatment of tumors, doctors can use CT (Computed Tomography), MRI (Nuclear Magnetic Resonance Imaging), and PET (Positron Emission Tomography) for tumor diagnosis. However, although diseases can be diagnosed through various medical images, during diagnosis, the identification of information such as the location of tumors still relies on the experience of doctors, which may affect the accuracy of diagnosis due to human factors. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to provide a lightweight segmentation method and system for medical images to solve the problem of errors in manual diagnosis during the medical image diagnosis process. The specific technical solutions are as follows: In the first aspect of the embodiments of the present application, a lightweight segmentation method for medical images is first provided. The method includes: Obtain multi-modal medical images to be recognized; Perform downsampling on the multi-modal medical images to be recognized to obtain shallow feature maps and deep feature maps; Extract and fuse multi-scale features from the deep feature maps to obtain multi-scale features; Perform upsampling on the multi-scale features and fuse them with the shallow feature maps to obtain features to be segmented; Perform segmentation based on the segmentation features to obtain a segmentation result, where the segmentation result includes the size, shape, and location information of the target lesion or organ.

[0004] In a possible implementation manner, the obtaining of the multi-modal medical images to be recognized includes: Obtain original multi-modal images, where the original multi-modal images include at least one of computer tomography (CT) images, nuclear magnetic resonance imaging (MRI), and positron emission tomography (PET) images; Perform spatial registration on the original multi-modal images and a preset template to obtain registered multi-modal images; Perform normalization and size adjustment on the registered multi-modal images to obtain the multi-modal medical images to be recognized.

[0005] In a possible implementation, downsampling the multi-modal medical image to be recognized to obtain a shallow feature map and a deep feature map includes: Extracting shallow features of the multi-modal medical image to be recognized through a dynamic kernel mixer and a dual-domain collaborative attention mechanism module in a first sampling layer to obtain the shallow feature map; Extracting deep features of the multi-modal medical image to be recognized through a dynamic kernel mixer and a dual-domain collaborative attention mechanism module in a second sampling layer to obtain the deep feature map.

[0006] In a possible implementation, extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features includes: Extracting multi-scale context features from the deep feature map through an extended channel compression mixer; Concatenating the extracted context features to obtain the multi-scale features.

[0007] In a possible implementation, segmenting according to the segmentation features to obtain a segmentation result includes: Generating a probability map according to the segmentation features through an activation function; Generating a binary mask according to the probability map through a threshold corresponding to a preset lesion category; According to the binary mask, obtaining the size and / or position information of the target lesion through dilation and / or erosion.

[0008] In the second aspect of the embodiments of the present application, a lightweight segmentation system for medical images is provided. The system includes: An image acquisition module for acquiring a multi-modal medical image to be recognized; A downsampling module for downsampling the multi-modal medical image to be recognized to obtain a shallow feature map and a deep feature map; A feature fusion module for extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features; An upsampling module for upsampling the multi-scale features and fusing them with the shallow feature map to obtain features to be segmented; A feature segmentation module for segmenting according to the segmentation features to obtain a segmentation result, where the segmentation result includes the size, shape, and position information of the target lesion or organ.

[0009] In a possible implementation, the image acquisition module is specifically configured to acquire an original multi-modal image, where the original multi-modal image includes at least one of a computed tomography (CT) image, a magnetic resonance imaging (MRI), and a positron emission tomography (PET) image; perform spatial registration on the original multi-modal image and a preset template to obtain a registered multi-modal image; perform normalization and size adjustment on the registered multi-modal image to obtain the multi-modal medical image to be recognized.

[0010] In a possible implementation, the downsampling module is specifically configured to extract shallow features from the multi-modal medical image to be recognized through a dynamic kernel mixer and a dual-domain collaborative attention mechanism module in the first sampling layer to obtain the shallow feature map; extract deep features from the multi-modal medical image to be recognized through a dynamic kernel mixer and a dual-domain collaborative attention mechanism module in the second sampling layer to obtain the deep feature map.

[0011] In a possible implementation, the feature fusion module is specifically configured to extract multi-scale context features from the deep feature map through an extended channel compression mixer; splice the extracted context features to obtain the multi-scale features.

[0012] In a possible implementation, the feature segmentation module is specifically configured to generate a probability map according to the segmentation features through an activation function; generate a binary mask according to a threshold corresponding to a preset lesion category based on the probability map; obtain the size and / or position information of the target lesion according to the binary mask through dilation and / or erosion.

[0013] On the other hand, an embodiment of the present application further provides an electronic device, including: a memory for storing a computer program; a processor, when executing the program stored in the memory, implements any of the above lightweight segmentation methods for medical images.

[0014] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, any of the above lightweight segmentation methods for medical images is implemented.

[0015] On the other hand, an embodiment of the present application further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the above lightweight segmentation methods for medical images.

[0016] Advantages of the embodiments of the present application: An embodiment of the present application provides a lightweight segmentation method and system for medical images. The method includes: obtaining multi-modal medical images to be recognized; performing downsampling on the multi-modal medical images to be recognized to obtain a shallow feature map and a deep feature map; extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features; performing upsampling on the multi-scale features and fusing them with the shallow feature map to obtain features to be segmented; and performing segmentation based on the segmentation features to obtain a segmentation result, where the segmentation result includes the size, shape, and position information of the target lesion or organ. Through the solution of the embodiment of the present application, after obtaining the multi-modal medical images to be recognized, a deep feature map and a shallow feature map can be obtained through downsampling, then multi-scale features can be obtained through the extraction and fusion of multi-scale features, and then features to be segmented can be obtained through upsampling and feature fusion, so as to obtain the size and / or position information of the target lesion through feature segmentation. Thus, not only can automatic segmentation be achieved through the model to obtain the size and / or position information of the target lesion.

[0017] Of course, implementing any product or method of the present application does not necessarily require achieving all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.

[0019] Figure 1 It is a schematic flowchart of a lightweight segmentation method for medical images provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of obtaining multi-modal images provided by an embodiment of the present application; Figure 3 It is an architecture diagram of LM-UNet provided by an embodiment of the present application; Figure 4 It is a generation flowchart of a dynamic kernel mixer provided by an embodiment of the present application; Figure 5 It is a module architecture diagram of ECCM provided by an embodiment of the present application; Figure 6 It is an architecture diagram of DDC provided by an embodiment of the present application; Figure 7 It is an architecture diagram of ACI-L provided by an embodiment of the present application; Figure 8 It is a schematic flowchart of feature segmentation provided by an embodiment of the present application; Figure 9 It is a schematic structural diagram of a lightweight segmentation system for medical images provided by an embodiment of the present application; Figure 10 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0020] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.

[0021] In the first aspect of the embodiments of the present application, first, a lightweight segmentation method for medical images is provided. Refer to Figure 1 , Figure 1 It is a schematic flow diagram of a lightweight segmentation method for medical images provided by an embodiment of the present application. The method includes: Step S11: Obtain the multi-modal medical image to be recognized; Step S12: Downsample the multi-modal medical image to be recognized to obtain a shallow feature map and a deep feature map; Step S13: Extract and fuse multi-scale features from the deep feature map to obtain multi-scale features; Step S14: Upsample the multi-scale features and fuse them with the shallow feature map to obtain the feature to be segmented; Step S15: Segment according to the segmentation feature to obtain a segmentation result, where the segmentation result includes the size, shape, and position information of the target lesion or organ.

[0022] Corresponding to the above step S11, the multi-modal medical images in the embodiments of the present application may include various types of medical images. Specifically, it may include CT (Computed Tomography), MRI (Nuclear Magnetic Resonance Imaging), PET (Positron Emission Tomography), etc. And it should be noted that the multi-modal medical images in the embodiments of the present application are medical images for the same lesion location. In one example, when the solution in the embodiments of the present application is applied to the diagnosis of tumors, the multi-modal medical image is a multi-modal medical image for the location of the tumor.

[0023] Corresponding to the above step S12, in the embodiments of the present application, the multi-modal medical image to be recognized can be downsampled by multiple downsampling layers to obtain shallow feature maps and deep feature maps. Specifically, a sampling layer can be used to collect high-resolution image features, and then another sampling layer can be used to collect low-resolution image features. In one case, both of the two sampling layers in the embodiments of the present application can include a DKM (Dynamic Kernel Mixer) and a DDC (Dual-Domain Collaborative Attention Mechanism), and downsampling is performed through the dynamic kernel mixer and the dual-domain collaborative attention mechanism.

[0024] Corresponding to the above step S13, when extracting and fusing multi-scale features of the deep feature map, multi-scale features of the deep feature map can be first extracted. Specifically, an ECCM (Extended Channel Compression Mixer) can be used for extraction and fusion. In one case, when performing extraction and fusion of multi-scale features, dilated convolution can be first performed on the deep feature map, and then the convolution results are fused to obtain scale features. Among them, feature fusion can be performed through various methods, for example, feature splicing, summation, etc.

[0025] Corresponding to the above step S14, when upsampling the multi-scale features and fusing them with the shallow feature map, the multi-scale features can be first upsampled, and then the upsampling result is fused with the shallow feature map. Among them, during upsampling, the multi-scale features can be restored to a feature map of a preset size, and then spliced with the shallow feature map to obtain the feature to be segmented. In actual use, the feature to be segmented can refer to logits, which is the input of the segmentation layer and also the output of the layer before the segmentation layer. The segmentation layer can perform segmentation based on the logits to obtain a segmentation result.

[0026] Corresponding to the above step S15, the segmentation result includes the size, shape, and position information of the target lesion or organ. For example, the segmentation result can be the size, shape, and position information of an organ such as the liver in the image, or the size, shape, and position information of certain lesions such as tumors. When performing segmentation based on the segmentation feature, feature segmentation can be achieved by generating a probability map and performing binary masking through a threshold, etc., so as to obtain the size and / or position information of the target lesion.

[0027] It can be seen that through the solution of the embodiments of the present application, after obtaining the multi-modal medical images to be recognized, deep feature maps and shallow feature maps can be obtained through downsampling, and then multi-scale features can be obtained by extracting and fusing multi-scale features. Then, through upsampling and feature fusion, the features to be segmented are obtained, so that the size and / or position information of the target lesion can be obtained through feature segmentation. Thus, not only can automatic segmentation be achieved through the model to obtain the size and / or position information of the target lesion.

[0028] In a possible implementation manner, for the obtaining of the multi-modal medical images to be recognized, refer to Figure 2 , Figure 2 which is a schematic flowchart of a process for obtaining multi-modal images provided by the embodiments of the present application, and includes: Step S21: Obtain the original multi-modal images, where the original multi-modal images include at least one of computer tomography (CT) images, magnetic resonance imaging (MRI), and positron emission tomography (PET) images; Step S22: Perform spatial registration on the original multi-modal images and a preset template to obtain the registered multi-modal images; Step S23: Normalize and adjust the size of the registered multi-modal images to obtain the multi-modal medical images to be recognized.

[0029] Among them, the original multi-modal images in the embodiments of the present application include at least one of computer tomography (CT) images, magnetic resonance imaging (MRI), and positron emission tomography (PET) images. Specifically, the original multi-modal images, such as the HU value (Hounsfield unit) sequence of CT and the multi-contrast sequence of MRI, can be stored in the DICOM (Digital Imaging and Communications in Medicine) format. The corresponding data structure can be a three-dimensional array I ∈ R H×W×Cm , where H / W is the spatial size; C m is the number of modalities. For example, when MRI has four modalities, C m = 4; R is a real number.

[0030] In one case, when performing spatial registration on the original multi-modal images and a preset template to obtain the registered multi-modal images, the SyN (a typical denial-of-service attack) algorithm of the ANTs (a database) library can be used to align the multi-modal images to a standard template, such as the MNI152 (a brain atlas) brain atlas, and output the registered image I reg ∈ R H×W×Cm to ensure the spatial consistency of anatomical structures.

[0031] In one case, when normalizing and resizing the registered multimodal images, normalization can be performed according to a preset specification. In one example, for CT images, the HU value can be truncated to [-200, 300] and normalized to [0, 1]. The formula is: for the corresponding array I norm =clip(I reg , -200, 300) - (-200) / 500, where clip is a loss function; in another example, for MRI data, Z-Score (z-score) normalization can be used to eliminate field strength inhomogeneity. In the embodiments of the present application, when resizing, the image can be Resized (changed in size) to a fixed size H′×W′×D′, such as 256×256×64 for 3D (three-dimensional) images, through trilinear interpolation, thereby outputting a four-dimensional tensor X∈R B×Cm×H′×W′×D′ , where H′, W′, and D′ represent height, width, and depth respectively; B is the batch size. In actual use, the registration logic can be executed by the CPU (Central Processing Unit), and the GPU (Graphics Processing Unit) can be used to accelerate the interpolation and normalization calculations, such as PyTorch CUDA (a PyTorch version of a parallel computing platform and programming model) tensor operations. In the present application, for modal and dynamic image processing, a modality-specific kernel generation mechanism (such as a PET metabolic region weight of 0.72) and spatio-temporal feature decoupling (optical flow method motion compensation) are used. A cross-modal improvement of 3.1% is achieved, and the dynamic sequence error is reduced by 48.7% (for example, the myocardial wall segmentation error in cardiac MRI is reduced from 2.3 mm to 1.18 mm).

[0032] In a possible implementation manner, downsampling the to-be-recognized multimodal medical image to obtain a shallow feature map and a deep feature map includes: extracting shallow features of the to-be-recognized multimodal medical image through a dynamic kernel mixer and a dual-domain collaborative attention mechanism module in a first sampling layer to obtain the shallow feature map; extracting deep features of the to-be-recognized multimodal medical image through a dynamic kernel mixer and a dual-domain collaborative attention mechanism module in a second sampling layer to obtain the deep feature map.

[0033] In the embodiments of the present application, the preprocessed multimodal tensor X can be input into an encoder based on nnUNet (a network structure). This encoder contains 2 downsampling levels, which are also the above-mentioned first sampling layer and second sampling layer. Each layer embeds a dynamic kernel mixer and a dual-domain collaborative attention mechanism. Specifically, for the first sampling layer, shallow feature extraction can be achieved. Specifically, by embedding a dynamic kernel mixer, X0 = X can be input, and the spatial dimension can be compressed to 1×1×D′ through GAP (global average pooling) to generate a global feature vector g∈RB×Cm ; Then, through a double-layer fully connected network (FC1: C m →4C m , FC2: 4C m →κ×C k , where κ = 3 is the number of kernel groups, and C k =3 is the number of channels per single kernel), κ groups of convolution kernel parameters W k ∈R κ×Ck×3×3 are generated; then, through Softmax (normalization), dynamic weights π∈R B×κ are generated. Among them, depthwise separable convolution (e.g., 3×3 kernel) is performed on X0, and the output feature Figure X 1∈R B×Cm×H′×W′×D′ is obtained. Then, local spatial features can be extracted through the dual-domain collaborative attention mechanism and depthwise separable convolution (e.g., 3×3 kernel): X1′′ = DWConv(X1), where DWConv is depthwise separable convolution; then, a channel mixer is used: the number of channels is extended to 4C through 1×1 convolution m , and after activation by GeLU (Gaussian error unit), it is compressed back to Cm. The formula is: X1′′ = Conv 1x1 (GeLU(Conv1x1(X1′))), where Conv is convolution; residual connection and batch normalization: X1 = BN(X1 + X1′′), and BN represents batch normalization. For the second sampling layer, downsampling (stride 2) halves the size, and the DKM + DDC operations are repeated, and finally, low-resolution feature Figure X 2∈R B×2Cm×H′ / 2×W′ / 2×D′ / 2 is output. Data structure change: the number of channels doubles with downsampling (C m →2C m ), and the spatial size is halved, forming a hierarchical feature pyramid. Specifically, for hardware association, GPU parallel computing can be used for depthwise separable convolution and fully connected layers, and CUDA cores can be used to accelerate matrix operations. In one example, see Figure 3 , Figure 3 is an architecture diagram of the LM-UNet provided by an embodiment of this application. The ordinary convolution of the LM-UNet (a medical image segmentation model based on a hybrid multi-layer perceptron) model is replaced with a DKM module, which can adaptively complete feature extraction. On this basis, after the DKM, it enters the DDC attention mechanism. Spatial and channel attention weights are generated. The ACI-L module (a feature extraction module) is used on the skip connection to enhance feature extraction. An ECCM channel mixer is proposed in the bottleneck layer, providing an efficient and accurate feature modeling solution for medical image segmentation tasks. In one example, see Figure 4, where the working principle of dynamic convolution is as follows: The input features complete basic feature extraction and normalization through a triple operation, and then enter the parallel convolution layer. By adaptively using different numbers of convolution kernels, differential feature maps are generated. The Dropout layer is used to randomly discard to prevent the model from overfitting. Finally, features are obtained through weighted summation. In this application, DKM is based on global average pooling and a double-layer fully connected network, generates κ groups (κ = 3 - 5) of dynamic convolution kernel parameters, and realizes anatomy structure-sensitive feature extraction through Softmax weight allocation (such as the weight of the tumor core area π = 0.62), replacing the traditional fixed convolution kernel, adapting to tumor heterogeneous features and dynamic organ morphological changes, and reducing the number of parameters.

[0034] In this application, the DKM module can generate κ groups (κ = 3 - 5) of candidate convolution kernel parameters based on global average pooling and a double-layer fully connected network, and dynamically allocate weights through Softmax. This technical feature directly acts on the feature extraction of different anatomical structures: in the brain tumor segmentation scenario, due to the significant differences in the image features between tumor tissues and normal tissues, it is difficult for traditional fixed convolution kernels to specifically extract the features of the tumor core area; while DKM can reach a weight of 0.62 for the tumor core area through dynamic weight allocation, can adaptively adjust the receptive field, and accurately capture the details of the tumor boundary. Experimental data shows that on the BraTS2021 (a dataset) dataset, after adopting this module, the Dice coefficient of tumor core area segmentation is improved from 0.848 of the traditional method to 0.924, with an increase of 0.076. The principle is that the dynamic kernel can selectively enhance the response of the target area and reduce background noise interference according to the feature differences such as texture and gray scale of different tissues. In this application, the DDC module adopts a channel-spatial decoupled parallel attention architecture, and through a two-dimensional matrix reshaping operation (H×W×C → (H×W)×C) and a dynamic sparsification calculation strategy, breaks through the quadratic complexity bottleneck of the traditional self-attention mechanism. In the liver vessel segmentation task, the principle of the traditional Transformer (a network model) architecture is to separate the channel dimension and the spatial dimension for processing. The channel attention branch quickly screens out effective channel information through matrix operations, and the spatial attention branch accurately locates the vessel structure using global pooling. The combined effect of the two reduces the HD95 error of liver vessel segmentation from 4.83mm of the traditional method to 3.35mm, and at the same time improves the inference speed, achieving a double improvement in global modeling efficiency and accuracy. In this application, the dual-domain collaborative attention module uses a channel-spatial decoupled parallel attention architecture, and through two-dimensional matrix reshaping (H×W×C → (H×W)×C) and dynamic sparsification calculation, realizes linear complexity global modeling. The amount of calculation is reduced by 40% compared with the traditional self-attention, and at the same time suppresses noise channels (attenuation coefficient 0.3 - 0.5) and enhances the response of anatomical structures (gain coefficient 1.2 - 1.5).

[0035] In a possible implementation, extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features includes: extracting multi-scale context features from the deep feature map through an extended channel compression mixer; splicing the extracted context features to obtain the multi-scale features. Specifically, the processing object can be the low-resolution features output by the encoder. Figure X 2, which is input into the ECCM (Extended Channel Compression Mixer). The specific steps may include: 1. Atrous convolution group, using 3D atrous convolution with dilation rates of 1 / 3 / 5 through parallel branches to extract multi-scale context and output X 2a , X 2b , X 2c ∈R B×2Cm×H′ / 2×W′ / 2×D′ / 2 ; 2. Feature splicing and compression, by splicing multi-scale features: X concat =Concat(X 2a , X 2b , X 2c ), where Concat represents splicing; a 1×1 convolution compresses the channels to 2C m : Xbottleneck = Conv1x1(X concat ), where Xbottleneck is the result of convolution compression and Conv represents convolution. The data structure can expand the receptive field through atrous convolution, keep the spatial dimension unchanged, and achieve the fusion of "local details - global context". In one example, see Figure 5 , where the working principle of ECCM is: the input data X first passes through the Depthwise conv layer (depthwise separable convolution) to extract spatial features, and then enters the Channel mixer. In the Channel mixer, the first conv2d (two-dimensional convolution) expands the number of channels, the GeLU activation function introduces non-linearity, and the second conv2d compresses the channels. The original input X is added to the processed features through a residual connection, and finally Y is output through batchNorm (batch normalization). This structure effectively reduces the computational amount, accelerates convergence, and improves the generalization ability.

[0036] In this application, the ECCM module adopts depthwise separable convolution combined with a channel bottleneck structure. Through the design of reducing the cross-channel interaction parameter quantity, while maintaining the feature expression ability, it significantly reduces the model parameter quantity. In low-dose CT image segmentation, due to severe noise interference, traditional models are prone to overfitting, resulting in a decline in segmentation accuracy; while the ECCM module can enhance the stability of feature propagation through residual connection and batch normalization, and at the same time compress the redundant channel information, improving the noise resistance of the model on low-dose CT data. In principle, depthwise separable convolution reduces the redundant calculation of convolution operations, and the channel compression mechanism effectively suppresses the transmission of noise between channels, ultimately reducing the total model parameter quantity and enabling efficient operation on resource-constrained devices. In this application, the extended channel compression mixer, through depthwise separable convolution combined with the "channel expansion - compression" mechanism (expansion multiple of 4 - 8 times), reduces the cross-channel interaction parameter quantity. It realizes lightweight cross-channel feature interaction, enhances the stability of feature propagation, and supports low-power deployment of edge devices.

[0037] In the embodiment of this application, in step S14, the multi-scale features are upsampled and fused with the shallow feature map to obtain the features to be segmented, which can achieve feature recovery and attention calibration. Specifically, the processing object can be the neck layer feature X bottleneck The cross-layer features X1 and X0 of the encoder are input into the decoder, and each layer embeds a DDC (dual-domain collaborative attention module). Specifically, the processing steps can include: Level 1: Upsampling and cross-layer fusion. Among them, upsampling (interpolation + convolution) restores the size of X bottleneck to H′ / 2×W′ / 2×D′ / 2, and is concatenated with the encoder X2, and the output X3 ∈ R B×4Cm ×H′ / 2×W′ / 2×D′ / 2; then through the DDC module, feature reshaping is realized using the channel attention branch: X3′ ∈ R B×(H′W′D′ / 2)×4Cm ; Channel interaction: ∈ R B×4Cm×4Cm ; Spatial gating: G c = Sigmoid(Conv 1x1 (X3)) ∈ R B ×4Cm×H′ / 2×W′ / 2×D′ / 2 ; Channel enhancement: , reshaping back to a three-dimensional tensor. Spatial attention branch: Global pooling: g s = GlobalPool(X3) ∈ R B×4Cm , GlobalPool is global pooling; Weight generation: W s = Softmax(FC(g s )) ∈ R B×4Cm ; Spatial enhancement: Unsqueeze(-1) Unsqueeze(-1), where Unsqueeze means expansion; Dual-domain fusion: X3 = Y c + Y s Level 2: Optimization of output layer features. The upsampling and DDC operations can be repeated, fused with the encoder X1, and finally the segmentation logits are output through a 1×1 convolution: Logits ∈ R B×Cout×H′×W′×D′ (Cout is the number of categories. For example, in brain tumor segmentation, Cout = 3). Data structure change: The spatial dimension can be restored through upsampling and cross-layer concatenation, and the DDC module dynamically calibrates the feature responses in both the channel and spatial dimensions. In one example, see Figure 6 , where the DDC attention mechanism is as follows: The input X is split into two paths. In the Channel-only self Attention module, through operations such as Conv(1x1), Reshape, multiplication, and softmax, the channel attention features are obtained; in the Spatial - only selfAttention module, through operations such as Conv(1x1), Global pooling, Reshape, multiplication, and softmax, the spatial attention features are obtained. After the two paths of features are fused and added to X, the output Y is obtained, realizing attention weighting in both the channel and spatial dimensions. In one example, see Figure 7 , where the ACI-L module can input X, which is convolved through a 1×1 convolutional layer and then branched. One branch goes through adaptive average pooling, changes the shape of the array and tensor, then is convolved through a 1×1 convolutional layer, and then generates weights through the Sigmoid function and changes the shape of the array and tensor; it multiplies element-wise with the features of the other branch, and then is convolved through a 1×1 convolutional layer to obtain the output Y, realizing feature weighting adjustment.

[0038] In this application, for multimodal medical images, a modality-specific kernel parameter generation mechanism and a cross-channel mixer are designed. In the CT-MRI-PET multimodal brain tumor segmentation scenario, traditional methods use simple channel splicing and cannot fully utilize the complementary information of each modality. However, in the present invention, DKM generates specific convolutional kernels for different modalities. For example, a higher weight (0.72) is assigned to the metabolically active region of the PET image. At the same time, ECCM realizes the deep fusion of cross-modal features, improving the cross-modal segmentation Dice coefficient from 0.869 of the traditional method to 0.900, an increase of 3.1%. The principle lies in that the modality-specific kernel can specifically extract the key features of each modality, and the cross-channel mixer realizes feature complementarity through information interaction, ultimately improving the generalization ability of the model on multi-center and multimodal data and reducing the errors caused by different devices and imaging parameters. It can be seen that the method of the embodiment of this application constructs a complete technical system of "adaptive feature extraction - efficient global modeling - multimodal generalization - hardware co-acceleration" in the field of medical image segmentation, providing an innovative solution for precision medicine and real-time diagnosis, and having significant clinical application value and technological foresight.

[0039] In a possible implementation manner, segmenting according to the segmentation features to obtain a segmentation result, see Figure 8 , includes: Step S81, generating a probability map according to the segmentation features through an activation function; Step S82, generating a binary mask according to the probability map through a threshold corresponding to a preset lesion category; Step S83, obtaining the size and / or position information of the target lesion according to the binary mask through dilation and / or erosion.

[0040] Among them, in the embodiment of this application, the logits tensor output by the decoder, the data structure can be R B ×Cout×H′×W′×D′ . Specifically, the processing steps may include: probability map generation. Specifically, the class probability distribution can be obtained through the Softmax activation function: Prob = Softmax(Logits, dim = 1) ∈ R B×Cout×H′×W′×D′ ; threshold segmentation, which can apply an adaptive threshold (such as the Otsu algorithm) to each class to generate a binary mask Mask ∈ R B×Cout×H′×W′×D′; dim=1 indicates row-wise operation; for morphological optimization, isolated noise points can be removed through dilation / erosion operations to output the final segmentation result. Specifically, hardware association can accelerate Softmax calculation through GPU, and CPU performs post-morphological processing (lightweight operation). In this application, a method for generating dynamic convolution kernels guided by global features (including global average pooling, fully connected network, and Softmax weight allocation); and a feature calibration method for decoupling channels and space (including matrix reshaping, channel weight generation, and spatial gating) are used to implement a multi-modal processing pipeline: a full-process method from preprocessing (bias field correction, rigid registration) to multi-modal fusion (feature concatenation + attention weighting). In one example, the hardware module combination corresponding to this application: an encoder-decoder architecture including DKM, DDC, and ECCM, supporting multi-modal input and dynamic image processing. It can achieve edge device inference optimization based on TensorRT quantization. Implement multi-modal medical images: organ segmentation (such as liver blood vessels, brain tumors) and micro-lesion detection of CT / MRI / PET. And dynamic sequence analysis: clinical scenarios such as cardiac MRI motion compensation and low-dose CT noise suppression.

[0041] In the second aspect of the embodiments of the present application, a lightweight segmentation system for medical images is provided. Refer to Figure 9 , the system includes: An image acquisition module 901, configured to acquire multi-modal medical images to be recognized; A downsampling module 902, configured to downsample the multi-modal medical images to be recognized to obtain shallow feature maps and deep feature maps; A feature fusion module 903, configured to extract and fuse multi-scale features from the deep feature maps to obtain multi-scale features; An upsampling module 904, configured to upsample the multi-scale features and fuse them with the shallow feature maps to obtain features to be segmented; A feature segmentation module 905, configured to perform segmentation based on the segmentation features to obtain a segmentation result, where the segmentation result includes the size, shape, and position information of the target lesion or organ.

[0042] In a possible implementation manner, the image acquisition module is specifically configured to acquire original multi-modal images, where the original multi-modal images include at least one of computer tomography (CT) images, magnetic resonance imaging (MRI), and positron emission tomography (PET) images; perform spatial registration on the original multi-modal images and a preset template to obtain registered multi-modal images; and perform normalization and size adjustment on the registered multi-modal images to obtain the multi-modal medical images to be recognized.

[0043] In a possible implementation manner, the downsampling module is specifically configured to extract shallow features from the multi-modal medical image to be recognized through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the first sampling layer, so as to obtain the shallow feature map; and extract deep features from the multi-modal medical image to be recognized through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the second sampling layer, so as to obtain the deep feature map.

[0044] In a possible implementation manner, the feature fusion module is specifically configured to extract multi-scale context features from the deep feature map through an extended channel compression mixer; and splice the extracted context features to obtain the multi-scale features.

[0045] In a possible implementation manner, the feature segmentation module is specifically configured to generate a probability map according to the segmentation features through an activation function; generate a binary mask according to a threshold corresponding to a preset lesion category according to the probability map; and obtain the size and / or position information of the target lesion according to the binary mask through dilation and / or erosion.

[0046] It can be seen that through the system of the embodiments of the present application, after obtaining the multi-modal medical image to be recognized, deep feature maps and shallow feature maps can be obtained through downsampling, and then multi-scale features can be obtained through extraction and fusion of multi-scale features, and then upsampling and feature fusion are performed to obtain the features to be segmented, so that the size and / or position information of the target lesion can be obtained through feature segmentation. Therefore, not only can automatic segmentation be realized through the model to obtain the size and / or position information of the target lesion.

[0047] The embodiments of the present application also provide an electronic device, as Figure 10 shown, including: A memory 1001 for storing a computer program; A processor 1002, configured to implement the following steps when executing the program stored on the memory 1001: Obtain a multi-modal medical image to be recognized; Perform downsampling on the multi-modal medical image to be recognized to obtain a shallow feature map and a deep feature map; Extract and fuse multi-scale features from the deep feature map to obtain multi-scale features; Upsample the multi-scale features and fuse them with the shallow feature map to obtain features to be segmented; Perform segmentation according to the segmentation features to obtain a segmentation result, where the segmentation result includes the size and / or position information of the target lesion.

[0048] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0049] The communication interface is used for communication between the above electronic device and other devices.

[0050] The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.

[0051] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0052] In another embodiment provided by the present application, a computer-readable storage medium is also provided. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of any of the above-mentioned lightweight segmentation methods for medical images are implemented.

[0053] In another embodiment provided by the present application, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute any of the lightweight segmentation methods for medical images in the above embodiments.

[0054] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.

[0055] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.

[0056] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system, electronic device, and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0057] The foregoing are only the preferred embodiments of the present application and are not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included within the scope of protection of the present application.

Claims

1. A lightweight segmentation method for medical images, characterized in that The method includes: Obtaining a multi-modal medical image to be recognized; Downsampling the multi-modal medical image to be recognized to obtain a shallow feature map and a deep feature map; Extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features; Upsampling the multi-scale features and fusing them with the shallow feature map to obtain features to be segmented; Performing segmentation based on the segmentation features to obtain a segmentation result, where the segmentation result includes the size, shape, and location information of the target lesion or organ.

2. The method according to claim 1, wherein The obtaining of the multi-modal medical image to be recognized includes: Obtaining an original multi-modal image, where the original multi-modal image includes at least one of a computed tomography (CT) image, a magnetic resonance imaging (MRI), and a positron emission tomography (PET) image; Performing spatial registration on the original multi-modal image and a preset template to obtain a registered multi-modal image; Normalizing and adjusting the size of the registered multi-modal image to obtain the multi-modal medical image to be recognized.

3. The method according to claim 1, wherein The downsampling of the multi-modal medical image to be recognized to obtain a shallow feature map and a deep feature map includes: Extracting shallow features from the multi-modal medical image to be recognized through a dynamic kernel mixer and a dual-domain collaborative attention mechanism module in a first sampling layer to obtain the shallow feature map; Extracting deep features from the multi-modal medical image to be recognized through a dynamic kernel mixer and a dual-domain collaborative attention mechanism module in a second sampling layer to obtain the deep feature map.

4. The method according to claim 1, wherein The extracting and fusing of multi-scale features from the deep feature map to obtain multi-scale features includes: Extracting multi-scale context features from the deep feature map through an extended channel compression mixer; Concatenating the extracted context features to obtain the multi-scale features.

5. The method according to claim 1, characterized in that, The performing of segmentation based on the segmentation features to obtain a segmentation result includes: Generating a probability map according to the segmentation features through an activation function; Generating a binary mask according to the probability map through a threshold corresponding to a preset lesion category; Obtaining the size and / or location information of the target lesion according to the binary mask through dilation and / or erosion.

6. A lightweight segmentation system for medical images, characterized in that, The system includes: An image acquisition module for obtaining a multi-modal medical image to be recognized; A downsampling module for downsampling the multi-modal medical image to be recognized to obtain a shallow feature map and a deep feature map; A feature fusion module for extracting and fusing multi-scale features from the deep feature map to obtain multi-scale features; An upsampling module for upsampling the multi-scale features and fusing them with the shallow feature map to obtain features to be segmented; A feature segmentation module for performing segmentation based on the segmentation features to obtain a segmentation result, where the segmentation result includes the size, shape, and location information of the target lesion or organ.

7. The system according to claim 6, wherein The image acquisition module is specifically configured to acquire original multi-modal images, where the original multi-modal images include at least one of computer tomography (CT) images, magnetic resonance imaging (MRI), and positron emission tomography (PET) images; perform spatial registration on the original multi-modal images and a preset template to obtain registered multi-modal images; perform normalization and size adjustment on the registered multi-modal images to obtain the multi-modal medical images to be recognized.

8. The system according to claim 6, wherein the downsampling module is specifically configured to extract shallow features of the multi-modal medical images to be recognized through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the first sampling layer to obtain the shallow feature map; extract deep features of the multi-modal medical images to be recognized through the dynamic kernel mixer and the dual-domain collaborative attention mechanism module in the second sampling layer to obtain the deep feature map.

9. An electronic device, characterized in that, including: a memory for storing a computer program; a processor for implementing the method according to any one of claims 1-5 when executing the program stored on the memory.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • DCE-MRI breast tumor image segmentation method based on deep neural network

    CN113139981A

  • Multi-level lesion detection optimized pathological section image segmentation detection method

    CN116596952A

  • Image segmentation method and device based on multi-path fusion convolution, equipment and medium

    CN119068187A

  • Microscopic hyperspectral image segmentation method based on multi-scale attention fusion

    CN119810116A

Cited By

  • Method for carrying out lung lobe region segmentation on SPECTV / Q image by using deep learning model

    CN120953302A

  • Image segmentation method and system based on multi-scale feature fusion

    CN121095568A

  • Method and system for rapidly segmenting and identifying microbiological image under microscope

    CN121147917A

  • Multi-modal medical image segmentation method and device, electronic equipment and medium

    CN121354228A