Automatic Medical Image Segmentation Method and Device Based on Variational Mixture-of-Experts Model
Through the combination of the variational hybrid expert model and the adaptive dictionary enhancement module, the problem of insufficient feature expression in multimodal medical image segmentation is solved, high-precision segmentation of complex anatomical structures is achieved, and the robustness and adaptability of the segmentation task is improved.
Patent Information
- Application Number
- CN202510535439.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The prior art has problems in the segmentation of multimodal medical image, such as insufficient feature expression, poor robustness and limited adaptability to complex anatomical structures, especially in the segmentation accuracy of throat cancer and lymph node segmentation.
The medical image automatic segmentation method based on the variational hybrid expert model is adopted, and multi-level convolution and adaptive dictionary enhancement are performed through the U-Net encoder, and shared feature extraction and dynamic routing are combined with the variational hybrid expert module. The multi-head attention mechanism and jump connection are used to achieve feature fusion, and high-precision segmentation results are output.
Improve the accuracy and robustness of medical image segmentation, improve the adaptability to complex anatomical structures, especially in the tasks of throat cancer and lymph node segmentation, which significantly improves segmentation accuracy and adaptability under noise interference.
Smart Images

Figure CN120047459B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and medical image processing, and more particularly, to a method and apparatus for automatic medical image segmentation based on a variational mixture of experts model. Background Art
[0002] In the field of medical image analysis, accurate segmentation of three-dimensional medical images such as laryngeal cancer and lymph nodes is of great significance, and the results directly affect clinical diagnosis, treatment planning, and efficacy evaluation. The core of the medical image segmentation task is to accurately extract the target area from complex anatomical structures. However, traditional methods mainly rely on manual feature extraction and rule-based algorithms. When dealing with multi-modal medical images (such as T1, T1C, T2-weighted MRI magnetic resonance imaging), these methods often show problems of insufficient accuracy and robustness due to noise interference, insufficient feature expression, and poor adaptability to complex anatomical structures.
[0003] In recent years, deep learning techniques have been widely used in medical image segmentation. Among them, U-Net, as a classic network architecture, gradually downsamples through the encoder to extract high-level semantic features, and uses the decoder to upsample to restore the spatial resolution. Combining skip connections to fuse low-level and high-level features, it has achieved remarkable results in various segmentation tasks. However, U-Net still has limitations when dealing with multi-modal medical images. Especially in the segmentation of complex anatomical structures such as laryngeal cancer and lymph nodes, it is easily affected by noise interference or insufficient feature expression, resulting in a decrease in segmentation accuracy. In addition, variational autoencoders are introduced to enhance the generalization ability of the model, but their specific optimization strategies in medical image segmentation are not yet mature, and it is difficult to fully utilize the advantages of latent distribution learning. The mixture of experts model processes complex tasks through multiple expert networks collaborating, and each expert focuses on a specific data subset or task sub-problem, and has shown potential in various applications. However, there is a lack of a network architecture in the prior art that effectively combines U-Net, variational autoencoders, and the mixture of experts model, resulting in difficulty in fully exerting the synergistic potential of the three in high-precision segmentation tasks. At the same time, there is still room for improvement in feature enhancement and multi-modal feature fusion in the existing methods, which limits the performance of the model in complex segmentation tasks.
[0004] In view of this, the present application is proposed. Summary of the Invention
[0005] The present invention aims to provide a method and apparatus for automatic medical image segmentation based on a variational mixture of experts model to solve the problems of insufficient feature expression, poor robustness, and limited adaptability to complex anatomical structures shown by existing methods in multi-modal medical image processing, so as to improve the accuracy, robustness, and adaptability to multi-modal medical images of medical image segmentation such as laryngeal cancer and lymph nodes.
[0006] To solve the above technical problems, the present invention is achieved through the following technical solutions:
[0007] A method for automatic segmentation of medical images based on a variational mixture of experts model, comprising:
[0008] S1, obtaining three-dimensional medical images;
[0009] S2, inputting the three-dimensional medical images into a U-Net encoder for multi-level convolutional, downsampling feature extraction and adaptive dictionary enhancement operations to optimize the local structural expression of features and generate enhanced feature maps at each level;
[0010] S3, inputting the enhanced feature map of the last level into a variational mixture of experts module for shared feature extraction and fusing it with the output of the dynamic routing selection results of multiple expert models to generate fused features;
[0011] S4, inputting the fused features combined with the enhanced feature map of the last level into a U-Net decoder for upsampling operations, and sequentially splicing the output of the upsampling with the enhanced feature map of the current level to restore the spatial resolution of the image and obtain decoded fused features;
[0012] S5, inputting the decoded fused features into a single-channel convolutional layer and outputting the segmentation result.
[0013] Preferably, each level of the U-Net encoder includes a convolutional block, a max-pooling downsampling, and an adaptive dictionary enhancement module;
[0014] Performing three-dimensional convolution, activation, and normalization operations through the convolutional block;
[0015] Performing multi-scale feature extraction through the max-pooling downsampling operation to obtain feature maps at each level;
[0016] Performing enhancement operations based on a dynamic dictionary and an attention mechanism on each sample of the feature maps at each level through the adaptive dictionary enhancement module to optimize the local structural expression of features.
[0017] Preferably, the enhancement operation based on a dynamic dictionary and an attention mechanism performed by the adaptive dictionary enhancement module is specifically:
[0018] Performing three-dimensional block extraction on the input feature map to obtain block vectors;
[0019] Generating a structural dictionary for each sample of the feature map through a dynamic dictionary generator;
[0020] Inputting the block vectors and the structural dictionaries of all samples into a multi-head attention module for enhancement and reconstruction to obtain an enhanced feature map.
[0021] Preferably, the three-dimensional block extraction operation is as follows:
[0022] Set the size of the three-dimensional block according to the levels of the U-Net encoder. Divide the input feature map into non-overlapping blocks through a three-dimensional unfolding operation, and flatten each block into a one-dimensional vector to generate a set of block vectors. The expression is:
[0023] ;
[0024] Where, represents the i-th flattened block vector; represents the real number field; represents the current layer's number of channels; represents the spatial side length of the block; represents the total number of elements of the block, that is, the volume of the three-dimensional block;
[0025] Generate a structure dictionary for each sample of the feature map through a dynamic dictionary generator. Specifically:
[0026] Apply three-dimensional adaptive average pooling to the feature map to compress the spatial dimension of the feature map to the smallest unit and flatten it into a one-dimensional vector;
[0027] Expand the one-dimensional vector to the hidden dimension through the first fully connected network and apply the ReLU activation function to obtain the first fully connected vector;
[0028] Generate a dictionary vector by passing the first fully connected vector through the second fully connected network. The length of the dictionary vector is consistent with the number of channels to obtain a structure dictionary. The expression is:
[0029] ;
[0030] Where, represents the structure dictionary; represents the three-dimensional average pooling operation for compressing the spatial dimension; represents the flattening operation for converting a three-dimensional tensor into a one-dimensional vector;
[0031] represents the fully connected layer that expands the number of channels to 2 times; ReLU represents the activation function;
[0032] represents the fully connected layer that maps the dimension to the dictionary size; represents the input feature map of the current layer;
[0033] represents the dimension of the finally generated structure dictionary;
[0034] The operations of enhancing and reconstructing the multi-head attention module are specifically as follows:
[0035] Using the block vector as the query vector, and the structure dictionaries of all samples as the key vector and value vector, calculate the block vector after multi-head attention enhancement to focus on the structural expression, and obtain the enhanced block vector. The expression is:
[0036] ;
[0037] where, represents the set of enhanced block vectors; MultiHeadAttn represents the multi-head attention mechanism; P represents the set of block vectors; D represents the structure dictionary, which is used as the key vector and value vector in the multi-head attention mechanism at the same time;
[0038] Reshape the enhanced block vector into the original space dimension and perform a residual connection with the input feature map to obtain the enhanced feature map. The expression is:
[0039] ;
[0040] where, represents the enhanced feature map; represents the input feature map; represents the operation of reshaping into the same dimension as ; represents the residual connection.
[0041] Preferably, the variational mixture of experts module includes a shared feature extractor and an expert network composed of multiple expert models, which is used for dynamic routing selection and probability distribution modeling of input features;
[0042] wherein, the shared feature extractor is used to extract the general pattern of the input features to reduce the redundant calculations of the expert models;
[0043] Each of the expert models is composed of three variational U-Nets and independently processes the input features to generate variational parameters for feature extraction of different modality input features; the variational parameters include the mean and log variance of the expert models.
[0044] Preferably, the process of extracting the input features by the shared feature extractor is as follows:
[0045] Reduce the dimension of the input features through the first three-dimensional convolutional layer to extract low-order features;
[0046] Apply the ReLU activation function to enhance the non-linear expression of the low-order features;
[0047] The features after ReLU activation are passed through the second 3D convolutional layer to restore the number of channels to the dimension of the original input features, and the shared features are output. The expression is:
[0048] ;
[0049] Among them, represents the general features generated by the shared feature extractor; represents the feature map of the nth level, that is, the output feature map of the last bottleneck layer; represents the first 3D convolutional dimensionality reduction operation, that is, from channels to channels; represents the second 3D convolutional restoration operation, that is, from channels restored to channels; , both represent the number of channels, and .
[0050] Preferably, the process of each expert model processing the input features is as follows:
[0051] The input features are mapped to a high-dimensional feature space through the 3D convolutional layer of the U-Net encoder to extract high-dimensional features;
[0052] The ReLU activation function is applied to enhance the non-linear expression of the high-dimensional features;
[0053] The features after ReLU activation are respectively input into two parallel decoders to generate corresponding variational parameters through 3D convolution. The expression is:
[0054] ;
[0055] Among them, represents the mean generated by the jth expert model; represents the variance generated by the jth expert model; represents the mapping function of the jth expert model; represents the input features of the nth level, that is, the enhanced features of the last bottleneck layer; represents the activation function; represents the 3D convolutional expansion operation, that is, from channels expanded to channels; , both represent the number of channels, and .
[0056] Preferably, after the input features are processed by multiple expert models, a gating and routing mechanism is used for dynamic routing selection, and the routing output result is fused with the shared features. The specific process is as follows:
[0057] The input features pass through a weighting mechanism to calculate the selection weight vector of the expert models. The expression is:
[0058] ;
[0059] where, represents the selection weight vector of the expert models; represents the normalization function; represents a three-dimensional convolution operation, that is, mapping from channels to channels; b represents the learnable bias term;
[0060] According to the selected weight vector, calculate the mean weighted sum of each selected expert model as the routing output corresponding to each sample in the input features. The expression is:
[0061] ;
[0062] where, represents the routing output of the th sample of the input features; represents the normalized weight of the th sample in the th selected expert model; represents the number of the top K of the most relevant experts; represents the mean of the th most relevant expert, ; (·) represents the mean calculation function;
[0063] Fuse the routing outputs of all input feature samples with the shared features and then output.
[0064] Preferably, it also includes training optimization using a loss function;
[0065] where, the loss function includes a segmentation loss and a variational regularization loss, and a regularization weight is used to balance the segmentation loss and the variational regularization loss;
[0066] The segmentation loss is used to measure the difference between the predicted label and the true label;
[0067] The variational regularization loss introduces a variational regularization term and quantifies the difference between the distribution of the output of the variational mixture of experts module and the standard normal distribution through KL divergence;
[0068] The expression of the loss function is as follows:
[0069] ;
[0070] where: represents the loss function; represents the segmentation loss; represents the predicted label; represents the ground truth label;
[0071] represents the regularization weight; represents the top K of the most relevant expert models selected; KL represents the KL divergence; represents the variance generated by the j-th expert model; represents the mean generated by the j-th expert model; represents the normal distribution of the j-th expert; N(0, 1) represents the standard normal distribution.
[0072] The present invention also provides a medical image automatic segmentation device based on a variational mixture of experts model, including:
[0073] An acquisition unit, configured to acquire three-dimensional medical images;
[0074] A U-Net encoding unit, configured to input the three-dimensional medical image into a U-Net encoder for multi-level convolution, downsampling feature extraction, and adaptive dictionary enhancement operations to optimize the local structural expression of features and generate enhanced feature maps at each level;
[0075] A variational mixture of experts unit, configured to input the enhanced feature map of the last level into a variational mixture of experts module for shared feature extraction, and fuse it with the output of the dynamic routing selection of multiple expert models to generate fused features;
[0076] A decoding unit, configured to input the fused features in combination with the enhanced feature map of the last level into a U-Net decoder for upsampling operations, and concatenate the output of the upsampling with the enhanced feature map of the current level level by level to restore the spatial resolution of the image and obtain decoded fused features;
[0077] An output unit, configured to input the decoded fused features into a single-channel convolutional layer and output a segmentation result.
[0078] The present invention also provides a medical image automatic segmentation device based on a variational mixture of experts model, including a processor and a memory. The memory stores a computer program that can be executed by the processor to implement a medical image automatic segmentation method based on a variational mixture of experts model as described above.
[0079] The present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of the device where the computer-readable storage medium is located, the above-mentioned automatic medical image segmentation method based on a variational mixture of experts model is implemented.
[0080] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0081] By introducing a variational mixture of experts module and an adaptive dictionary enhancement module, the present invention optimizes the feature extraction and representation capabilities, thereby achieving more accurate and reliable segmentation results.
[0082] First, the variational mixture of experts module enhances the diversity and pertinence of feature representation through a dynamic routing mechanism, selects the optimal expert for specialization processing for different modalities or feature patterns, thereby improving the model's adaptability to complex anatomical structures. Second, the adaptive dictionary enhancement module optimizes the local structure expression through sample-specific dynamic dictionaries and attention mechanisms, improves the detail perception ability of the feature map, and thus improves the segmentation accuracy. Third, multi-level feature fusion realizes seamless fusion of low-level and high-level features through skip connections and decoders, ensures the integrity of spatial information, and further improves the model's segmentation performance for complex anatomical structures.
[0083] The present invention is applicable to various medical imaging modalities, including three-dimensional medical image data such as CT and MRI. In particular, it can effectively cope with the problems of noise interference and insufficient feature expression in the segmentation tasks of pharyngeal cancer and lymph nodes. By dynamically routing multi-modal data and modeling the uncertainty probability distribution, the present invention realizes the accurate extraction and fusion of different modal features, thereby improving the robustness and adaptability of the segmentation task.
[0084] In summary, by introducing a variational mixture of experts module and an adaptive dictionary enhancement module, the present invention solves the deficiencies of the prior art in the segmentation tasks of pharyngeal cancer and lymph nodes, and significantly improves the model's feature extraction ability and segmentation accuracy. This solution has a clear implementation path and broad application prospects, providing important technical support for the field of medical image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0086] Figure 1Flowchart of a medical image automatic segmentation method based on a variational mixture of experts model provided for Example 1.
[0087] Figure 2 Schematic diagram of the overall structure of a medical image automatic segmentation method based on a variational mixture of experts model provided for Example 1.
[0088] Figure 3 Schematic diagram of the structure of a medical image automatic segmentation device based on a variational mixture of experts model provided for Example 1.
[0089] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Specific embodiments
[0090] To make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention to be protected, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work fall within the scope of protection of the present invention.
[0091] Example 1
[0092] Example 1 of the present invention provides a medical image automatic segmentation method based on a variational mixture of experts model, which can be implemented by a medical image automatic segmentation device based on a variational mixture of experts model (hereinafter referred to as the segmentation device), and particularly, is executed by one or more processors in the segmentation device.
[0093] In this embodiment, the segmentation device can be an electronic device equipped with a processor, and the processor has a computer program of this medical image automatic segmentation method based on a variational mixture of experts model and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited here.
[0094] In this embodiment, T1 and T1C (enhanced T1) are two important imaging sequences in magnetic resonance imaging (MRI). T1 refers to T1-weighted imaging, which mainly reflects the differences in the longitudinal relaxation time of tissues. T1C is enhanced T1-weighted imaging, which is performed after intravenous injection of a gadolinium-based contrast agent (such as gadopentetate dimeglumine) on the basis of T1-weighted imaging.
[0095] The present invention constructs an efficient automatic medical image segmentation framework by combining a U-Net encoder-decoder structure, an adaptive dictionary enhancement operation, and a variational mixture of experts module. The core lies in using the variational mixture of experts model to process complex multi-modal medical image features, and restoring high-resolution segmentation results through progressive upsampling and feature fusion.
[0096] The method of the present invention is mainly applied to the accurate segmentation of three-dimensional images of pharyngeal cancer and lymph nodes. Of course, it can also be used for the segmentation of other three-dimensional images, which is not limited here.
[0097] As Figure 1 - Figure 2 shown, an automatic medical image segmentation method based on a variational mixture of experts model includes steps S1 to S5.
[0098] S1. Obtain three-dimensional medical images.
[0099] Input data: Three-dimensional medical images (such as CT, MRI, etc.) are the basis for the segmentation task. Three-dimensional images can provide richer spatial information, which helps to accurately segment complex anatomical structures and provide raw data for subsequent feature extraction and processing.
[0100] S2. Input the three-dimensional medical images into the U-Net encoder for multi-level convolution, downsampling feature extraction, and adaptive dictionary enhancement operations to optimize the local structural expression of features and generate enhanced feature maps at each level.
[0101] In this embodiment, a U-Net-based encoding-decoding framework is adopted, which is designed specifically for efficiently processing three-dimensional input data. Each level of the U-Net encoder includes a convolutional block, a max-pooling downsampling, and an adaptive dictionary enhancement module. Three-dimensional convolution, activation, and normalization operations are performed through the convolutional block; multi-scale feature extraction is performed through the max-pooling downsampling operation to obtain feature maps at each level; an enhancement operation based on a dynamic dictionary and an attention mechanism is performed on each sample of the feature maps at each level through the adaptive dictionary enhancement module to optimize the local structural expression of features.
[0102] Compared with the traditional U-Net, this model introduces a multi-scale adaptive dictionary enhancement (E-Dictionary) module in the encoding path to strengthen the feature expression ability, and embeds a variational mixture of experts module in the bottleneck layer to achieve dynamic feature routing and uncertainty modeling.
[0103] As Figure 2As shown, the input three-dimensional data enters the encoding path. First, after performing a convolution operation through the convolutional layer of the first level, four levels of max-pooling downsampling and convolution operations are successively performed to gradually compress the spatial resolution and expand the channel dimension, and multi-scale features of the image are gradually extracted. The feature maps generated by each level of convolution and downsampling operation have different receptive fields and resolutions, which can capture the local and global information of the image and provide multi-scale feature inputs for the subsequent adaptive dictionary enhancement and variational mixture of experts modules.
[0104] After generating each level of feature map, the E-Dictionary adaptive dictionary enhancement module based on the combination of dynamic dictionary and attention mechanism is applied to locally enhance the feature map to obtain the enhanced feature map of each level, so as to optimize the local structure expression of the feature and improve the detail expression ability.
[0105] Specifically, the steps for performing the adaptive dictionary enhancement operation on each level of feature map are as follows:
[0106] First, three-dimensional block extraction is performed on the input feature map to obtain block vectors, specifically:
[0107] According to the hierarchical setting of the U-Net encoder, the size of the three-dimensional block is set (for example, the first three layers are larger and the fourth layer is smaller). The input feature map is divided into non-overlapping blocks through three-dimensional unfolding operation, and each block is flattened into a one-dimensional vector to generate a set of block vectors. The expression is:
[0108] ;
[0109] where represents the i-th flattened block vector; represents the real number field; represents the current layer's number of channels; represents the spatial side length of the block; represents the total number of elements of the block, that is, the volume of the three-dimensional block.
[0110] For example, the first-level feature map is divided into a large number of blocks, and each block contains a specific number of elements.
[0111] Then, a structure dictionary is generated for each sample of the feature map through a dynamic dictionary generator, specifically:
[0112] Three-dimensional adaptive average pooling is applied to the feature map to compress the spatial dimension of the feature map to the minimum unit and flatten it into a one-dimensional vector; the one-dimensional vector is expanded to the hidden dimension through the first fully connected network, and the ReLU activation function is applied to obtain the first fully connected vector; the first fully connected vector is passed through the second fully connected network to generate a dictionary vector, and the length of the dictionary vector is consistent with the number of channels to obtain the structure dictionary. The expression is:
[0113] ;
[0114] Among them, represents the structure dictionary; represents the 3D average pooling operation, which is used to compress the spatial dimension; represents the flattening operation, which is used to convert the 3D tensor into a 1D vector;
[0115] represents the fully connected layer that expands the number of channels to 2 times, that is, expands to ; ReLU represents the activation function; represents the fully connected layer that maps the dimension to the dictionary size; represents the input feature map of the current layer; represents the dimension of the finally generated structure dictionary.
[0116] Dictionary learning is a sparse representation method that represents the input signal by learning a set of basis vectors (dictionary). The adaptive dictionary enhancement operation dynamically adjusts the dictionary according to the input feature map to better capture the local features of the image.
[0117] Finally, the block vector and the structure dictionaries of all samples are input into the multi-head attention module for enhancement and reconstruction to obtain the enhanced feature map. Specifically:
[0118] Using the block vector as the query vector, the structure dictionaries of all samples as the key vector and value vector, calculate the block vector enhanced by multi-head attention to focus on the structure expression, and obtain the enhanced block vector. The expression is:
[0119] ;
[0120] Among them, represents the set of enhanced block vectors; MultiHeadAttn represents the multi-head attention mechanism; P represents the set of block vectors; D represents the structure dictionary, which is also used as the key vector (key) and value vector (value) in the multi-head attention mechanism;
[0121] Reshape the enhanced block vector into the original spatial dimension and perform a residual connection with the input feature map to obtain the enhanced feature map. The expression is:
[0122] ;
[0123] Among them, represents the enhanced feature map; represents the input feature map; represents reshaping into the same shape as Operations in the same dimension; Indicates a residual connection.
[0124] S3, Input the enhanced feature map of the last level into the variational mixture of experts module for shared feature extraction, and fuse it with the output of the dynamic routing selection results of multiple expert models to generate fused features.
[0125] In the encoding path, the enhanced feature map of the fourth level enters the bottleneck layer, is processed by the variational mixture of experts module, and generates fused features and variational parameters.
[0126] The variational mixture of experts module is the core component of the bottleneck layer, aiming to perform dynamic routing and probability modeling on the input enhanced feature map through the cooperation of multiple expert model networks and a shared feature extractor. This module not only enriches the diversity of feature representations but also introduces uncertainty quantification through variational parameters, improving the adaptability and reliability of the model.
[0127] Among them, the shared feature extractor is used to extract the general pattern of the input features to reduce the redundant calculations of the expert models.
[0128] The process of extracting the input features by the shared feature extractor is as follows:
[0129] The input features are dimensionally reduced through the first three-dimensional convolutional layer to extract low-order features; the ReLU activation function is applied to enhance the non-linearity of the low-order features; the features after ReLU activation are passed through the second three-dimensional convolutional layer to restore the number of channels to the original input feature dimension, and the shared features are output. The expression is:
[0130] ;
[0131] Among them, represents the general features generated by the shared feature extractor; represents the feature map of the nth level, that is, the output feature map of the last bottleneck layer; represents the first three-dimensional convolutional dimensional reduction operation, that is, from (such as 64) channels to (such as 256) channels; represents the second three-dimensional convolutional restoration operation, that is, from (256) channels back to (64) channels; 、 both represent the number of channels, and .
[0132] Each expert model network consists of three variational U-Nets and independently processes the input features to generate variational parameters for feature extraction of input features in different modalities or feature patterns. The process of each expert model processing the input features is as follows:
[0133] The input features are mapped to a high-dimensional feature space through the three-dimensional convolutional layer of the U-Net encoder to extract high-dimensional features; the ReLU activation function is applied to enhance the non-linear expression of the high-dimensional features; the ReLU-activated features are respectively input into two parallel decoders to generate corresponding variational parameters through three-dimensional convolution; the variational parameters include the mean and log variance of the expert model, and the expression is:
[0134] ;
[0135] where, represents the mean generated by the j-th expert model; represents the variance generated by the j-th expert model; represents the mapping function of the j-th expert model; represents the input feature at the n-th level, that is, the enhanced feature of the last bottleneck layer; represents the activation function; represents the three-dimensional convolution expansion operation, that is, expanding from channels to channels; , both represent the number of channels, and , such as , .
[0136] After the input features are processed by multiple expert models, a gating and routing mechanism is used for dynamic routing selection, and the routing output result is fused with the shared features and output. The specific process is as follows:
[0137] First, an expert selection weight vector logits is extracted from the input features through a single-channel convolutional layer, and a learnable bias is superimposed to adjust the selection tendency; then, the function is applied to convert the selection weight vector logits into a weight distribution to select the top K (such as the top two) experts from the weight distribution and normalize their weights. The expression is:
[0138] ;
[0139] where, represents the selection weight vector of the expert model; represents the normalization function; represents the three-dimensional convolution operation, that is, mapping from channels (such as 256 channels) to a number (such as 3) of channels; b represents a learnable bias term;
[0140] Finally, according to the selected weight vector, calculate the mean weighted sum of each selected expert model as the routing output corresponding to each sample in the input features. The expression is:
[0141] ;
[0142] where, represents the routing output of the th sample of the input features; represents the normalized weight of the th selected expert model in the th sample; represents the number of the top K (such as the top 2) of the most relevant experts; represents the th mean of the most relevant experts, ; (·) represents the mean calculation function;
[0143] Fuse the routing output of all input feature samples with the shared features and then output. The expression is:
[0144] ;
[0145] where, o represents the fused features; s represents the general features generated by the shared feature extractor; r represents the routing output of all samples.
[0146] S4. Input the fused features combined with the enhanced feature map of the last level into the U-Net decoder for upsampling operations, and gradually splice the output of the upsampling with the enhanced feature map of the current level to restore the spatial resolution of the image, obtaining the decoded fused features.
[0147] In this embodiment, the decoder gradually restores the spatial resolution of the features through four levels of upsampling. After each level of upsampling and convolution, it combines the splicing and fusion processing output of the corresponding encoded features.
[0148] Specifically, after inputting the fused features combined with the enhanced feature map of the last level into the U-Net decoder for upsampling operations, the output of the upsampling and the enhanced features of the next level Figure 1 serve as the input for the next level of upsampling, and gradually splice the output of the upsampling to restore the spatial resolution of the image, obtaining the decoded fused features.
[0149] S5. Input the decoded fused features into a single-channel convolutional layer and output the segmentation result.
[0150] In this embodiment, after multiple levels ( Figure 2After the upsampling of level 4 and the convolutional fusion stitching, through the last single-channel convolution (such as using the Sigmoid activation), the segmentation result is output.
[0151] The segmentation result, such as the class label in the 3D image, realizes the automatic segmentation of medical images.
[0152] Specifically, this embodiment can use Python as the main programming language and implement it using the Pytorch deep learning framework.
[0153] In another preferred embodiment, it also includes training and optimizing the model corresponding to the method of the present invention using a loss function. The loss function combines the segmentation loss and the variational regularization loss, and uses a regularization weight to balance the segmentation loss and the variational regularization loss.
[0154] The segmentation loss is used to measure the difference between the predicted label and the true label; the variational regularization loss introduces a variational regularization term, and quantifies the difference between the distribution output by the variational mixture of experts module and the standard normal distribution through the KL divergence.
[0155] The expression of the loss function is as follows:
[0156] ;
[0157] Where: represents the loss function; represents the segmentation loss; represents the predicted label; represents the true label;
[0158] represents the regularization weight; represents the top K of the most relevant expert models selected; KL represents the KL divergence; represents the variance generated by the j-th expert model; represents the mean generated by the j-th expert model; represents the normal distribution of the j-th expert; N(0,1) represents the standard normal distribution.
[0159] A network structure integrating U-Net, variational autoencoder, and mixture-of-experts model proposed by the method of the present invention enhances the feature representation ability by introducing an adaptive dictionary enhancement (E-Dictionary) module in the encoder, fuses multi-modal data features using a variational mixture-of-experts module in the bottleneck layer, and achieves high-precision segmentation through skip connections and a decoder. The E-Dictionary module optimizes feature expressions through dictionary learning and sparse coding; the variational mixture-of-experts block combines multiple expert networks and a shared expert network to dynamically select the optimal feature representation; the encoder-decoder structure of U-Net ensures the integrity of spatial information. This multi-level and multi-modal collaborative design is particularly suitable for complex segmentation tasks of pharyngeal cancer and lymph nodes.
[0160] In practical applications, assume that a medical institution needs to accurately segment a patient's pharyngeal cancer and lymph nodes to assist in diagnosis and treatment planning. First, obtain the patient's CT or MRI three-dimensional medical image data as input data and import it into the network model corresponding to the method of the present invention. In the encoding path, the input data is input into the multi-level feature extraction and adaptive dictionary enhancement module as shown in Figure 2 to generate a multi-scale enhanced feature map. The last-level enhanced feature map enters the bottleneck layer, where the variational mixture-of-experts module performs dynamic feature routing and uncertainty modeling. The decoding path gradually restores the spatial resolution through a four-level upsampling module and fuses multi-level feature information through skip connections, finally generating a segmentation result. The segmentation result can be directly used to assist doctors in formulating surgical plans or evaluating treatment effects.
[0161] In addition, the present invention performs excellently in multi-modal medical image processing. For example, when processing multi-modal data such as T1, T1C, and T2-weighted MRI, the variational mixture-of-experts module selects the optimal expert for specialization processing according to different modalities or feature patterns through a dynamic routing mechanism, thereby enhancing the diversity of feature representation. The adaptive dictionary enhancement module optimizes the local structure expression through sample-specific dynamic dictionaries and attention mechanisms, improving the detail perception ability of the feature map. Multi-level feature fusion realizes the seamless fusion of low-level and high-level features through skip connections and the decoding path, ensuring the integrity of spatial information. These designs significantly improve the segmentation performance of the model for complex anatomical structures and solve the limitations of the prior art in terms of noise interference and insufficient feature expression.
[0162] In summary, by introducing a variational mixture of experts module and an adaptive dictionary enhancement module, the present invention solves the deficiencies of the prior art in 3D medical image segmentation tasks, particularly in pharyngeal cancer and lymph node segmentation tasks. From the preprocessing of the input data to the generation of the final segmentation result, each step of the present invention is carefully designed and optimized to ensure the high performance and robustness of the model. The present invention is not only applicable to multiple medical imaging modalities but also can effectively address the segmentation challenges of complex anatomical structures, providing important technical support for the field of medical image segmentation.
[0163] Embodiment 2
[0164] As Figure 3 shown, the second embodiment of the present invention also provides a medical image automatic segmentation device based on a variational mixture of experts model, including:
[0165] An acquisition unit for acquiring 3D medical images;
[0166] A U-Net encoding unit for inputting the 3D medical images into a U-Net encoder for multi-level convolution, downsampling feature extraction, and adaptive dictionary enhancement operations to optimize the local structural expression of features and generate enhanced feature maps at each level;
[0167] A variational mixture of experts unit for inputting the enhanced feature map of the last level into a variational mixture of experts module for shared feature extraction and fusing with the output of the dynamic routing selection of multiple expert models to generate fused features;
[0168] A decoding unit for inputting the fused features in combination with the enhanced feature map of the last level into a U-Net decoder for upsampling operations and successively splicing the output of the upsampling with the enhanced feature map of the current level to restore the spatial resolution of the image and obtain decoded fused features;
[0169] An output unit for inputting the decoded fused features into a single-channel convolutional layer and outputting a segmentation result.
[0170] Embodiment 3
[0171] The third embodiment of the present invention also provides a medical image automatic segmentation device based on a variational mixture of experts model, which includes a memory and a processor. A computer program is stored in the memory and can be executed by the processor to implement the medical image automatic segmentation method based on the variational mixture of experts model as described above.
[0172] Embodiment 4
[0173] The fourth embodiment of the present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of the device where the computer-readable storage medium is located, the automatic medical image segmentation method based on the variational mixture of experts model as described above is implemented.
[0174] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0175] In addition, each functional module in various embodiments of the present invention can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0176] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
[0177] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0178] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: the situation of A existing alone, A and B existing simultaneously, and B existing alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0179] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0180] The "first / second" mentioned in the embodiments is only used to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0181] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An automatic medical image segmentation method based on a variational mixture of experts model, characterized in that, Including: S1. Obtain a three-dimensional medical image; S2. Input the three-dimensional medical image into a U-Net encoder for multi-level convolutional, downsampling feature extraction, and adaptive dictionary enhancement operations to optimize the local structural expression of features and generate enhanced feature maps at each level. Among them, each level of the U-Net encoder includes a convolutional block, a max-pooling downsampling, and an adaptive dictionary enhancement module. Through the adaptive dictionary enhancement module, an enhancement operation based on a dynamic dictionary and an attention mechanism is performed on each sample of the feature map at each level to optimize the local structural expression of features; The specific operation of the adaptive dictionary enhancement module for performing the enhancement operation based on the dynamic dictionary and the attention mechanism is as follows: Perform three-dimensional block extraction on the input feature map to obtain block vectors; Generate a structural dictionary for each sample of the feature map through a dynamic dictionary generator; Input the block vectors and the structural dictionaries of all samples into a multi-head attention module for enhancement and reconstruction to obtain an enhanced feature map; S3. Input the enhanced feature map of the last level into a variational mixture-of-experts module for shared feature extraction, and fuse it with the dynamic routing selection results output by multiple expert models to generate a fused feature. Among them, the variational mixture-of-experts module includes a shared feature extractor and expert networks, and is used for dynamic routing selection and probability distribution modeling of the input features. The expert networks are composed of multiple expert models; Among them, the shared feature extractor is used to extract the general pattern of the input features to reduce the redundant calculations of the expert models; Each of the expert models is composed of three variational U-Nets and independently processes the input features to generate variational parameters for feature extraction of different modality input features. The variational parameters include the mean and log variance of the expert model. After the input features are processed by multiple expert models, a gating and routing mechanism is used for dynamic routing selection, and the routing output results are fused with the shared features and output to obtain a fused feature; S4. Input the fused feature combined with the enhanced feature map of the last level into a U-Net decoder for upsampling operations, and sequentially splice the output of the upsampling and the enhanced feature map of the current level to restore the spatial resolution of the image to obtain a decoded fused feature; S5. Input the decoded fused feature into a single-channel convolutional layer and output a segmentation result.
2. The automatic medical image segmentation method based on the variational mixture of experts model according to claim 1, wherein Perform three-dimensional convolution, activation, and normalization operations through the convolutional block of the U-Net encoder; Perform multi-scale feature extraction through the max-pooling downsampling operation of the U-Net encoder to obtain feature maps at each level.
3. The automatic medical image segmentation method based on a variational mixture of experts model according to claim 2, wherein , The three-dimensional block extraction operation is as follows: Set the size of the three-dimensional block according to the level of the U-Net encoder, divide the input feature map into non-overlapping blocks through a three-dimensional unfolding operation, and flatten each block into a one-dimensional vector to generate a set of block vectors. The expression is: ; Among them, represents the i-th flattened block vector; represents the real number field; represents the current number of channels in the layer; represents the spatial side length of the block; represents the total number of elements of the block, that is, the volume of the three-dimensional block; Generate a structural dictionary for each sample of the feature map through a dynamic dictionary generator, specifically: Apply three-dimensional adaptive average pooling to the feature map to compress the spatial dimension of the feature map to the smallest unit and flatten it into a one-dimensional vector; Expand the one-dimensional vector to the hidden dimension through the first fully connected network and apply the ReLU activation function to obtain the first fully connected vector; Generate a dictionary vector by passing the first fully connected vector through a second fully connected network. The length of the dictionary vector is consistent with the number of channels to obtain a structure dictionary. The expression is: ; Among them, represents a structure dictionary; represents a 3D average pooling operation for compressing the spatial dimension; represents a flattening operation for converting a 3D tensor into a 1D vector; The fully-connected layer that expands the number of channels to twice; ReLU represents the activation function; A fully connected layer that maps dimensions to the dictionary size; Indicates the input feature map of the current layer; Indicates the dimension of the finally generated structure dictionary; The specific operations of the multi-head attention module for enhancement and reconstruction are as follows: Use the block vector as the query vector, the structure dictionaries of all samples as the key vectors and value vectors, and calculate the block vector after multi-head attention enhancement to focus on the structure expression to obtain an enhanced block vector. The expression is: ; Among them, represents a set of enhanced block vectors; MultiHeadAttn represents the multi-head attention mechanism; P represents the set of block vectors; D represents the structure dictionary and serves as both the key vector and the value vector in the multi-head attention mechanism; Reshape the enhanced block vector into the original spatial dimension and perform a residual connection with the input feature map to obtain an enhanced feature map. The expression is: ; Among them, represents the enhanced feature map; represents the input feature map; represents reshaping into the same dimension as the operation of; represents the residual connection.
4. The automatic medical image segmentation method based on a variational mixture of experts model according to claim 2, wherein , The process of extracting input features through the shared feature extractor is: Reduce the dimension of the input features through the first three-dimensional convolutional layer to extract low-order features; Apply the ReLU activation function to enhance the non-linear expression of the low-order features; Restore the number of channels of the features after ReLU activation to the original input feature dimension through the second three-dimensional convolutional layer and output the shared features. The expression is: ; Among them, represents the general features generated by the shared feature extractor; represents the feature map of the nth level, that is, the output feature map of the last bottleneck layer; represents the first 3D convolutional dimensionality reduction operation, that is, from channels reduced to channels; represents the second 3D convolutional restoration operation, that is, from channels restored to channels; , both represent the number of channels, and .
5. The automatic medical image segmentation method based on a variational mixture of experts model according to claim 1, characterized in that , The process of each expert model processing the input features is: Map the input features to a high-dimensional feature space through the three-dimensional convolutional layer of the U-Net encoder to extract high-dimensional features; Apply the ReLU activation function to enhance the non-linear expression of the high-dimensional features; Input the features after ReLU activation into two parallel decoders respectively to generate corresponding variational parameters through three-dimensional convolution; The expression is: ; Among them, represents the mean generated by the j-th expert model; represents the variance generated by the j-th expert model; represents the mapping function of the j-th expert model; represents the input feature at the n-th level, that is, the enhanced feature of the last bottleneck layer; represents the activation function; represents the 3D convolution expansion operation, that is, from channels expanded to channels; , both represent the number of channels, and .
6. The automatic medical image segmentation method based on the variational mixture of experts model according to claim 5, wherein , After the input features are processed by multiple expert models, a gating and routing mechanism is used for dynamic routing selection, and the routing output result is fused with the shared features for output. The specific process is: The input features pass through a weighting mechanism to calculate the selection weight vector of the expert model, and the expression is: ; Among them, represents the selection weight vector of the expert model; represents the normalization function; represents a three-dimensional convolution operation, that is, from channels are mapped to channels; b represents a learnable bias term; According to the selected weight vector, calculate the mean weighted sum of each selected expert model as the routing output corresponding to each sample in the input features. The expression is: ; Among them, represents the routing output of the th sample of the input feature; represents the th sample, the th normalized weight of the selected expert model; represents the top K quantity of the most relevant experts; represents the th mean value of the most relevant experts, ; (·) represents the mean value calculation function; Fuse and output the routing outputs of all input feature samples with the shared features.
7. The automatic medical image segmentation method based on the variational mixture of experts model according to claim 1, characterized in that , It also includes training optimization using a loss function; Among them, the loss function includes a segmentation loss and a variational regularization loss, and a regularization weight is used to balance the segmentation loss and the variational regularization loss; The segmentation loss is used to measure the difference between the predicted label and the true label; The variational regularization loss introduces a variational regularization term, and quantifies the difference between the distribution output by the variational mixture of experts module and the standard normal distribution through the KL divergence; The expression of the loss function is as follows: ; Wherein: represents the loss function; represents the segmentation loss; represents the predicted label; represents the true label; Denote the regularization weights; Denote the top K of the most relevant expert models selected; KL denotes the KL divergence; Denote the variance generated by the j-th expert model; Denote the mean generated by the j-th expert model; Denote the normal distribution of the j-th expert; N(0, 1) denotes the standard normal distribution.
8. An automatic medical image segmentation device based on a variational mixture of experts model, which is used to implement an automatic medical image segmentation method based on a variational mixture of experts model as described in any one of claims 1-7, characterized in that, It includes: An acquisition unit for acquiring three-dimensional medical images; A U-Net encoding unit for inputting the three-dimensional medical images into the U-Net encoder for multi-level convolution, downsampling feature extraction, and adaptive dictionary enhancement operations to optimize the local structure expression of the features and generate enhanced feature maps at each level; A variational mixture of experts unit for inputting the enhanced feature map of the last level into the variational mixture of experts module for shared feature extraction, and fusing it with the output of the dynamic routing selection of multiple expert models to generate a fused feature; A decoding unit for inputting the fused feature combined with the enhanced feature map of the last level into the U-Net decoder for upsampling operations, and gradually splicing the output of the upsampling with the enhanced feature map of the current level to restore the spatial resolution of the image to obtain a decoded fused feature; An output unit, configured to input the decoded fusion features into a single-channel convolutional layer and output a segmentation result.
Citation Information
Patent Citations
Method for generating DWI image based on CT cerebral infarction image
CN117422788A
Single-source-domain generalization medical image segmentation method and device based on shape dictionary
CN119672346A