Automatic medical image segmentation method and device based on variational hybrid expert model

By introducing the variational mixing expert module and the adaptive dictionary enhancement module in the automatic segmentation method of medical images, the problems of insufficient feature expression and poor robustness in multimodal medical image processing are solved, and the medical image segmentation effect with higher accuracy and robustness are achieved.

CN120047459AActive Publication Date: 2025-05-27XIAMEN UNIV OF TECH

Patent Information

Application Number
CN202510535439.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-27
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The prior art has problems in the multimodal medical image processing, such as insufficient feature expression, poor robustness and limited adaptability to complex anatomical structures, especially in the segmentation tasks of throat cancer and lymph nodes, which are difficult to achieve high-precision and robust segmentation.

Method used

The medical image automatic segmentation method based on the variational hybrid expert model is adopted, and multi-level convolution and downsampled feature extraction is performed through the U-Net encoder, and feature expression is optimized in combination with the adaptive dictionary enhancement module. Then, the feature input variational hybrid expert module is fused to generate fusion features. Finally, upsampling and feature fusion are performed through the U-Net decoder to restore the spatial resolution of the image and output the segmentation result.

Benefits of technology

It significantly improves the accuracy and robustness of medical image segmentation, improves the adaptability to multimodal medical images, and can segment more accurately and reliably when dealing with complex anatomical structures such as throat cancer and lymph nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047459A_ABST
    Figure CN120047459A_ABST
Patent Text Reader

Abstract

The invention provides a medical image automatic segmentation method and device based on a variational hybrid expert model, and relates to the technical field of medical image processing. The method comprises the following steps: acquiring a three-dimensional medical image, and inputting the three-dimensional medical image into a U-Net encoder for multi-level convolution, down-sampling feature extraction and adaptive dictionary enhancement operation to generate each level of enhanced feature map; inputting the enhanced feature map of the last level into a variational hybrid expert module to perform shared feature extraction, and outputting and fusing with dynamic routing selection results of a plurality of expert models to generate fused features; inputting the fused features and the enhanced feature map of the last level into a U-Net decoder for up-sampling operation, and splicing the output of up-sampling and the enhanced feature map of the current level step by step to recover the spatial resolution of the image to obtain decoded fused features; and inputting the decoding fusion feature into a single-channel convolutional layer, and outputting a segmentation result. According to the method, the accuracy and robustness of medical image segmentation and the adaptive capacity to the multi-modal medical image are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning and medical image processing, and in particular, to a method and device for automatic medical image segmentation based on a variational mixture of experts model. Background Art

[0002] In the field of medical image analysis, the accurate segmentation of three-dimensional medical images such as laryngeal cancer and lymph nodes is of great significance, and the results directly affect clinical diagnosis, treatment planning, and efficacy evaluation. The core of the medical image segmentation task is to accurately extract the target area from complex anatomical structures. However, traditional methods mainly rely on manual feature extraction and rule-based algorithms. When dealing with multi-modal medical images (such as T1, T1C, T2-weighted MRI magnetic resonance imaging), these methods often show problems of insufficient accuracy and robustness due to noise interference, insufficient feature expression, and poor adaptability to complex anatomical structures.

[0003] In recent years, deep learning techniques have been widely used in medical image segmentation. Among them, U-Net, as a classic network architecture, gradually downsamples through an encoder to extract high-level semantic features, and uses a decoder to upsample to restore the spatial resolution. Combining skip connections to fuse low-level and high-level features, it has achieved remarkable results in various segmentation tasks. However, U-Net still has limitations when dealing with multi-modal medical images. Especially in the segmentation of complex anatomical structures such as laryngeal cancer and lymph nodes, it is easily affected by noise interference or insufficient feature expression, resulting in a decrease in segmentation accuracy. In addition, variational autoencoders have been introduced to enhance the generalization ability of the model, but their specific optimization strategies in medical image segmentation are not yet mature, and it is difficult to fully utilize the advantages of latent distribution learning. The mixture of experts model processes complex tasks through multiple expert networks collaborating, and each expert focuses on a specific data subset or task sub-problem, and has shown potential in various applications. However, there is a lack of a network architecture in the prior art that effectively combines U-Net, variational autoencoders, and the mixture of experts model, resulting in difficulty in fully exerting the synergistic potential of the three in high-precision segmentation tasks. At the same time, there is still room for improvement in feature enhancement and multi-modal feature fusion in the existing methods, which limits the performance of the model in complex segmentation tasks.

[0004] In view of this, the present application is proposed. Summary of the Invention

[0005] The present invention aims to provide a device and method for automatic medical image segmentation based on a variational mixture of experts model to solve the problems of insufficient feature expression, poor robustness, and limited adaptability to complex anatomical structures shown by existing methods in multi-modal medical image processing, so as to improve the accuracy, robustness, and adaptability to multi-modal medical images of medical image segmentation such as laryngeal cancer and lymph nodes.

[0006] To solve the above technical problems, the present invention is implemented through the following technical solutions: An automatic medical image segmentation method based on a variational mixture of experts model, comprising: S1, obtaining three-dimensional medical images; S2, inputting the three-dimensional medical images into a U-Net encoder for multi-level convolution, downsampling feature extraction, and adaptive dictionary enhancement operations to optimize the local structural expression of features and generate enhanced feature maps at each level; S3, inputting the enhanced feature map of the last level into a variational mixture of experts module for shared feature extraction, and fusing it with the output of the dynamic routing selection results of multiple expert models to generate fused features; S4, inputting the fused features combined with the enhanced feature map of the last level into a U-Net decoder for upsampling operations, and successively splicing the output of the upsampling with the enhanced feature map of the current level to restore the spatial resolution of the image, obtaining decoded fused features; S5, inputting the decoded fused features into a single-channel convolutional layer and outputting a segmentation result.

[0007] Preferably, each level of the U-Net encoder includes a convolutional block, a max-pooling downsampling, and an adaptive dictionary enhancement module; Performing three-dimensional convolution, activation, and normalization operations through the convolutional block; Performing multi-scale feature extraction through max-pooling downsampling operations to obtain feature maps at each level; Performing enhancement operations based on a dynamic dictionary and an attention mechanism on each sample of the feature maps at each level through the adaptive dictionary enhancement module to optimize the local structural expression of features.

[0008] Preferably, the enhancement operation based on a dynamic dictionary and an attention mechanism performed by the adaptive dictionary enhancement module is specifically: Performing three-dimensional block extraction on the input feature map to obtain block vectors; Generating a structural dictionary for each sample of the feature map through a dynamic dictionary generator; Inputting the block vectors and the structural dictionaries of all samples into a multi-head attention module for enhancement and reconstruction to obtain an enhanced feature map.

[0009] Preferably, the three-dimensional block extraction operation is: Setting the size of the three-dimensional block according to the level of the U-Net encoder, dividing the input feature map into non-overlapping blocks through three-dimensional unfolding operations, and flattening each block into a one-dimensional vector to generate a set of block vectors, with the expression: ; where represents the i-th flattened block vector; denotes the real number field; represents the current number of channels in the layer; represents the spatial side length of the block; represents the total number of elements in the block, i.e., the volume of the three-dimensional block; Generate a structure dictionary for each sample of the feature map through a dynamic dictionary generator, specifically: Apply three-dimensional adaptive average pooling to the feature map to compress the spatial dimension of the feature map to the smallest unit and flatten it into a one-dimensional vector; Expand the one-dimensional vector to the hidden dimension through the first fully connected network and apply the ReLU activation function to obtain the first fully connected vector; Generate a dictionary vector by passing the first fully connected vector through the second fully connected network. The length of the dictionary vector is consistent with the number of channels to obtain the structure dictionary, and the expression is: ; where, represents the structure dictionary; represents the three-dimensional average pooling operation for compressing the spatial dimension; represents the flattening operation for converting a three-dimensional tensor into a one-dimensional vector; represents the fully connected layer that expands the number of channels to 2 times; ReLU represents the activation function; represents the fully connected layer that maps the dimension to the dictionary size; represents the current input feature map of the layer; represents the dimension of the finally generated structure dictionary; The operation of the multi-head attention module for enhancement and reconstruction is specifically as follows: Take the block vector as the query vector, the structure dictionaries of all samples as the key vector and value vector, and calculate the block vector after multi-head attention enhancement to focus on the structure expression to obtain the enhanced block vector. The expression is: ; where, represents the set of enhanced block vectors; MultiHeadAttn represents the multi-head attention mechanism; P represents the set of block vectors; D represents the structure dictionary, which is used as both the key vector and value vector in the multi-head attention mechanism; Reshape the enhanced block vector to the original spatial dimension and perform a residual connection with the input feature map to obtain the enhanced feature map. The expression is: ; where, Denote the enhanced feature map; Denote the input feature map; Denote the operation of reshaping to the same dimension as the same dimension; Denote the residual connection.

[0010] Preferably, the variational mixture of experts module includes a shared feature extractor and an expert network composed of multiple expert models, which is used for dynamic routing selection and probability distribution modeling of input features; Among them, the shared feature extractor is used to extract the general pattern of input features to reduce the redundant calculation of expert models; Each of the expert models is composed of three variational U-Nets and independently processes the input features to generate variational parameters for feature extraction of input features in different modalities; the variational parameters include the mean and log variance of the expert model.

[0011] Preferably, the process of extracting input features by the shared feature extractor is as follows: The input features are processed by the first three-dimensional convolutional layer for dimensionality reduction to extract low-order features; Apply the ReLU activation function to enhance the non-linear expression of the low-order features; The features after ReLU activation are passed through the second three-dimensional convolutional layer to restore the number of channels to the original input feature dimension, and the shared features are output. The expression is: ; Among them, Denote the general features generated by the shared feature extractor; Denote the feature map at the nth level, that is, the output feature map of the last bottleneck layer; Denote the first three-dimensional convolutional dimensionality reduction operation, that is, from channels to channels; Denote the second three-dimensional convolutional restoration operation, that is, from channels to channels; 、 Both denote the number of channels, and .

[0012] Preferably, the process of each expert model processing the input features is as follows: The input features are mapped to a high-dimensional feature space through the three-dimensional convolutional layer of the U-Net encoder to extract high-dimensional features; Apply the ReLU activation function to enhance the non-linear expression of the high-dimensional features; The features after ReLU activation are respectively input into two parallel decoders to generate corresponding variational parameters through 3D convolution; the expression is: ; Among them, represents the mean generated by the j-th expert model; represents the variance generated by the j-th expert model; represents the mapping function of the j-th expert model; represents the input feature at the n-th level, that is, the enhanced feature of the last bottleneck layer; represents the activation function; represents the 3D convolution expansion operation, that is, expanding from channels to channels; , both represent the number of channels, and .

[0013] Preferably, after the input features are processed by multiple expert models, a gating and routing mechanism is used for dynamic routing selection, and the routing output result is fused with the shared features. The specific process is as follows: The input features pass through a weighting mechanism to calculate the selection weight vector of the expert model. The expression is: ; Among them, represents the selection weight vector of the expert model; represents the normalization function; represents the 3D convolution operation, that is, mapping from channels to channels; b represents the learnable bias term; According to the selected weight vector, calculate the weighted sum of the means of each selected expert model as the routing output corresponding to each sample in the input features. The expression is: ; Among them, represents the routing output of the -th sample of the input features; represents the -th sample and the -th selected expert model's normalized weight; represents the number of the top K of the most relevant experts; represents the -th most relevant expert's mean, ; (·) represents the mean calculation function; Fuse the routing outputs of all input feature samples with the shared features and then output.

[0014] Preferably, it also includes training optimization using a loss function; Among them, the loss function includes a segmentation loss and a variational regularization loss, and a regularization weight is used to balance the segmentation loss and the variational regularization loss; The segmentation loss is used to measure the difference between the predicted label and the true label; The variational regularization loss introduces a variational regularization term, and quantifies the difference between the distribution output by the variational mixture of experts module and the standard normal distribution through the KL divergence; The expression of the loss function is as follows: ; Where: represents the loss function; represents the segmentation loss; represents the predicted label; represents the true label; represents the regularization weight; represents the top K of the most relevant expert models selected; KL represents the KL divergence; represents the variance generated by the j-th expert model; represents the mean generated by the j-th expert model; represents the normal distribution of the j-th expert; N(0,1) represents the standard normal distribution.

[0015] The present invention also provides a medical image automatic segmentation device based on a variational mixture of experts model, including: An acquisition unit, configured to acquire three-dimensional medical images; A U-Net encoding unit, configured to input the three-dimensional medical image into a U-Net encoder for multi-level convolution, downsampling feature extraction, and adaptive dictionary enhancement operations to optimize the local structure expression of features and generate enhanced feature maps at each level; A variational mixture of experts unit, configured to input the enhanced feature map of the last level into a variational mixture of experts module for shared feature extraction, and fuse it with the dynamic routing selection results output by multiple expert models to generate fused features; A decoding unit, configured to input the fused features combined with the enhanced feature map of the last level into a U-Net decoder for upsampling operations, and gradually splice the output of the upsampling with the enhanced feature map of the current level to restore the spatial resolution of the image and obtain decoded fused features; An output unit, configured to input the decoded fused features into a single-channel convolutional layer and output a segmentation result.

[0016] The present invention also provides a medical image automatic segmentation device based on a variational mixture of experts model, including a processor and a memory. A computer program is stored in the memory and can be executed by the processor to implement a medical image automatic segmentation method based on a variational mixture of experts model as described above.

[0017] The present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, a medical image automatic segmentation method based on a variational mixture of experts model as described above is implemented.

[0018] In summary, compared with the prior art, the present invention has the following beneficial effects: By introducing a variational mixture of experts module and an adaptive dictionary enhancement module, the present invention optimizes the feature extraction and representation capabilities, thereby achieving more accurate and reliable segmentation results.

[0019] First, the variational mixture of experts module enhances the diversity and pertinence of feature representation through a dynamic routing mechanism, selects the optimal expert for specialization processing for different modalities or feature patterns, thereby improving the model's adaptability to complex anatomical structures. Second, the adaptive dictionary enhancement module optimizes the local structure expression through a sample-specific dynamic dictionary and an attention mechanism, improves the detail perception ability of the feature map, and thus improves the segmentation accuracy. Third, multi-level feature fusion realizes seamless fusion of low-level and high-level features through skip connections and a decoder, ensures the integrity of spatial information, and further improves the model's segmentation performance for complex anatomical structures.

[0020] The present invention is applicable to multiple medical imaging modalities, including three-dimensional medical image data such as CT and MRI. In particular, it can effectively cope with the problems of noise interference and insufficient feature expression in pharyngeal cancer and lymph node segmentation tasks. By dynamically routing and modeling the uncertainty probability distribution of multi-modal data, the present invention realizes the accurate extraction and fusion of different modal features, thereby improving the robustness and adaptability of the segmentation task.

[0021] In summary, the present invention solves the deficiencies of the prior art in pharyngeal cancer and lymph node segmentation tasks by introducing a variational mixture of experts module and an adaptive dictionary enhancement module, and significantly improves the model's feature extraction ability and segmentation accuracy. This solution has a clear implementation path and broad application prospects, providing important technical support for the field of medical image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0023] Figure 1 It is a flowchart of a medical image automatic segmentation method based on a variational mixture of experts model provided for the first embodiment.

[0024] Figure 2 It is a schematic diagram of the overall structure of a medical image automatic segmentation method based on a variational mixture of experts model provided for the first embodiment.

[0025] Figure 3 It is a schematic diagram of the structure of a medical image automatic segmentation device based on a variational mixture of experts model provided for the first embodiment.

[0026] The following will further elaborate on the present invention in conjunction with the drawings and specific embodiments. Specific Embodiments

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0028] First Embodiment The first embodiment of the present invention provides a medical image automatic segmentation method based on a variational mixture of experts model, which can be implemented by a medical image automatic segmentation device based on a variational mixture of experts model (hereinafter referred to as the segmentation device), and particularly, is executed by one or more processors in the segmentation device.

[0029] In this embodiment, the segmentation device may be an electronic device equipped with a processor, and the processor has a computer program of the automatic medical image segmentation method based on the variational mixture of experts model and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited herein.

[0030] In this embodiment, T1 and T1C (enhanced T1) are two important imaging sequences in magnetic resonance imaging (MRI). T1 refers to T1-weighted imaging, which mainly reflects the differences in tissue longitudinal relaxation times. T1C is enhanced T1-weighted imaging, which is performed after intravenous injection of a gadolinium contrast agent (such as gadopentetate dimeglumine) on the basis of T1-weighted imaging.

[0031] The present invention constructs an efficient automatic medical image segmentation framework by combining a U-Net encoder-decoder structure, an adaptive dictionary enhancement operation, and a variational mixture of experts module. The core lies in using the variational mixture of experts model to process complex multi-modal medical image features and restoring high-resolution segmentation results through progressive upsampling and feature fusion.

[0032] The method of the present invention is mainly applied to the accurate segmentation of three-dimensional images of laryngeal cancer and lymph nodes. Of course, it can also be used for the segmentation of other three-dimensional images, which is not limited herein.

[0033] As Figure 1 - Figure 2 shown, an automatic medical image segmentation method based on a variational mixture of experts model includes steps S1 to S5.

[0034] S1. Obtain a three-dimensional medical image.

[0035] Input data: Three-dimensional medical images (such as CT, MRI, etc.) are the basis for the segmentation task. Three-dimensional images can provide richer spatial information, which helps to accurately segment complex anatomical structures and provides raw data for subsequent feature extraction and processing.

[0036] S2. Input the three-dimensional medical image into the U-Net encoder for multi-level convolution, downsampling feature extraction, and adaptive dictionary enhancement operation to optimize the local structure expression of the features and generate each-level enhanced feature map.

[0037] This embodiment adopts a U-Net-based encoding-decoding framework, which is designed specifically for efficiently processing three-dimensional input data. Each level of the U-Net encoder includes a convolution block, a max-pooling downsampling, and an adaptive dictionary enhancement module. Through the convolution block, three-dimensional convolution, activation, and normalization operations are performed; through the max-pooling downsampling operation, multi-scale feature extraction is performed to obtain each-level feature map; through the adaptive dictionary enhancement module, an enhancement operation based on a dynamic dictionary and an attention mechanism is performed on each sample of each-level feature map to optimize the local structure expression of the features.

[0038] Compared with the traditional U-Net, this model introduces a multi-scale adaptive dictionary enhancement (E-Dictionary) module in the encoding path to strengthen the feature expression ability, and embeds a variational mixture of experts module in the bottleneck layer to achieve dynamic feature routing and uncertainty modeling.

[0039] As Figure 2 shown, the input 3D data enters the encoding path. First, after convolution operations through the convolutional layers of the first level, four-level max-pooling downsampling and convolutional operations are then performed in sequence to gradually compress the spatial resolution and expand the channel dimension, and gradually extract the multi-scale features of the image. The feature maps generated by each level of convolution and downsampling operation have different receptive fields and resolutions, and can capture the local and global information of the image, providing multi-scale feature inputs for the subsequent adaptive dictionary enhancement and variational mixture of experts modules.

[0040] After generating the feature maps of each level, an E-Dictionary adaptive dictionary enhancement module that combines a dynamic dictionary and an attention mechanism is applied to locally enhance the feature maps, obtaining the enhanced feature maps of each level to optimize the local structure expression of the features and improve the detail expression ability.

[0041] Specifically, the steps for performing adaptive dictionary enhancement operations on the feature maps of each level are as follows: First, three-dimensional block extraction is performed on the input feature maps to obtain block vectors, specifically: According to the layer settings of the U-Net encoder, the size of the three-dimensional blocks is set (for example, the first three layers are larger and the fourth layer is smaller). The input feature maps are divided into non-overlapping blocks through three-dimensional unfolding operations, and each block is flattened into a one-dimensional vector to generate a set of block vectors. The expression is: ; where represents the i-th flattened block vector; represents the real number field; represents the current layer's number of channels; represents the spatial side length of the block; represents the total number of elements in the block, that is, the volume of the three-dimensional block.

[0042] For example, the first-layer feature maps are divided into a large number of blocks, and each block contains a specific number of elements.

[0043] Then, a structure dictionary is generated for each sample of the feature maps through a dynamic dictionary generator, specifically: Apply three-dimensional adaptive average pooling to the feature map to compress the spatial dimension of the feature map to the smallest unit and flatten it into a one-dimensional vector; expand the one-dimensional vector to the hidden dimension through the first fully-connected network and apply the ReLU activation function to obtain the first fully-connected vector; generate a dictionary vector by passing the first fully-connected vector through the second fully-connected network, where the length of the dictionary vector is consistent with the number of channels, to obtain a structure dictionary. The expression is: ; Among them, represents the structure dictionary; represents the three-dimensional average pooling operation for compressing the spatial dimension; represents the flattening operation for converting a three-dimensional tensor into a one-dimensional vector; represents the fully-connected layer that expands the number of channels to 2 times, that is expanded to ; ReLU represents the activation function; represents the fully-connected layer that maps the dimension to the dictionary size; represents the current layer's input feature map; represents the dimension of the finally generated structure dictionary.

[0044] Dictionary learning is a sparse representation method that represents the input signal by learning a set of basis vectors (dictionary). The adaptive dictionary enhancement operation dynamically adjusts the dictionary according to the input feature map to better capture the local features of the image.

[0045] Finally, input the block vector and the structure dictionaries of all samples into the multi-head attention module for enhancement and reconstruction to obtain the enhanced feature map. Specifically: Use the block vector as the query vector, the structure dictionaries of all samples as the key vector and value vector, and calculate the block vector enhanced by multi-head attention to focus on the structure expression to obtain the enhanced block vector. The expression is: ; Among them, represents the set of enhanced block vectors; MultiHeadAttn represents the multi-head attention mechanism; P represents the set of block vectors; D represents the structure dictionary, which is also used as the key vector (key) and value vector (value) in the multi-head attention mechanism; Reshape the enhanced block vector into the original spatial dimension and perform a residual connection with the input feature map to obtain the enhanced feature map. The expression is: ; Among them, represents the enhanced feature map; Represents the input feature map; Represents the reshaping operation to the same dimension as ; Represents the residual connection.

[0046] S3. Input the enhanced feature map of the last level into the variational mixture of experts module for shared feature extraction, and fuse it with the output of the dynamic routing selection results of multiple expert models to generate the fused features.

[0047] In the encoding path, the enhanced feature map of the fourth level enters the bottleneck layer, is processed by the variational mixture of experts module, and generates the fused features and variational parameters.

[0048] The variational mixture of experts module is the core component of the bottleneck layer, aiming to perform dynamic routing and probability modeling on the input enhanced feature map through the cooperation of multiple expert model networks and the shared feature extractor. This module not only enriches the diversity of feature representations but also introduces uncertainty quantification through variational parameters, improving the adaptability and reliability of the model.

[0049] Among them, the shared feature extractor is used to extract the general pattern of the input features to reduce the redundant calculations of the expert models.

[0050] The process of extracting the input features by the shared feature extractor is as follows: The input features are dimensionally reduced by the first three-dimensional convolutional layer to extract low-order features; the ReLU activation function is applied to enhance the non-linear expression of the low-order features; the features after ReLU activation are passed through the second three-dimensional convolutional layer to restore the number of channels to the original input feature dimension, and the shared features are output. The expression is: ; Among them, represents the general features generated by the shared feature extractor; represents the feature map of the nth level, that is, the output feature map of the last bottleneck layer; represents the first three-dimensional convolutional dimension reduction operation, that is, from (such as 64) channels to (such as 256) channels; represents the second three-dimensional convolutional restoration operation, that is, from (256) channels restored to (64) channels; 、 both represent the number of channels, and .

[0051] Each expert model network consists of three variational U-Nets and independently processes the input features to generate variational parameters for feature extraction of input features in different modalities or feature patterns. The process of each expert model processing the input features is as follows: The input features are mapped to a high-dimensional feature space through the three-dimensional convolutional layer of the U-Net encoder to extract high-dimensional features; the ReLU activation function is applied to enhance the non-linear expression of the high-dimensional features; the ReLU-activated features are respectively input into two parallel decoders to generate corresponding variational parameters through three-dimensional convolution; the variational parameters include the mean and log variance of the expert model, and the expression is: ; Among them, represents the mean generated by the j-th expert model; represents the variance generated by the j-th expert model; represents the mapping function of the j-th expert model; represents the input features at the n-th level, that is, the enhanced features of the last bottleneck layer; represents the activation function; represents the three-dimensional convolution expansion operation, that is, expanding from channels to channels; , both represent the number of channels, and , such as , .

[0052] After the input features are processed by multiple expert models, a gating and routing mechanism is used for dynamic routing selection, and the routing output result is fused with the shared features and output. The specific process is as follows: First, the expert selection weight vector logits are extracted from the input features through a single-channel convolutional layer, and a learnable bias is superimposed to adjust the selection tendency; then, the function is applied to convert the selection weight vector logits into a weight distribution, so as to select the top K (such as the top two) experts from the weight distribution and normalize their weights. The expression is: ; Among them, represents the selection weight vector of the expert model; represents the normalization function; represents the three-dimensional convolution operation, that is, mapping from channels (such as 256 channels) to channels (such as 3 channels); b represents the learnable bias term; Finally, according to the selected weight vector, calculate the mean weighted sum of each selected expert model as the routing output corresponding to each sample in the input features, and the expression is: ; where, represents the routing output of the -th sample of the input features; represents the normalized weight of the -th selected expert model in the -th sample; represents the number of the top K (such as the top 2) of the most relevant experts; represents the mean of the -th most relevant expert, ; (·) represents the mean calculation function; After fusing the routing outputs of all input feature samples with the shared features, the output is: ; where, o represents the fused feature; s represents the general feature generated by the shared feature extractor; r represents the routing outputs of all samples.

[0053] S4. Input the fused feature combined with the enhanced feature map of the last level into the U-Net decoder for upsampling operation, and gradually splice the output of upsampling with the enhanced feature map of the current level to restore the spatial resolution of the image, and obtain the decoded fused feature.

[0054] In this embodiment, the decoder gradually restores the spatial resolution of the feature through four-level upsampling. After each level of upsampling and convolution, it outputs through the splicing and fusion processing of the corresponding encoded feature.

[0055] Specifically, after inputting the fused feature combined with the enhanced feature map of the last level into the U-Net decoder for upsampling operation, the output of upsampling and the enhanced feature of the next level Figure 1 are used as the input of the next level of upsampling, and the output of upsampling is gradually spliced to restore the spatial resolution of the image, and the decoded fused feature is obtained.

[0056] S5. Input the decoded fused feature into a single-channel convolutional layer and output the segmentation result.

[0057] In this embodiment, after multi-level ( Figure 2 here are four levels) upsampling and convolutional fusion splicing, the segmentation result is output through the last single-channel convolution (such as using Sigmoid activation).

[0058] The segmentation result, such as the class label in the three-dimensional image, realizes the automatic segmentation of medical images.

[0059] Specifically, in this embodiment, Python can be used as the main programming language and implemented using the Pytorch deep learning framework.

[0060] In another preferred embodiment, it further includes training and optimizing the model corresponding to the method of the present invention using a loss function. The loss function combines a segmentation loss and a variational regularization loss, and uses a regularization weight to balance the segmentation loss and the variational regularization loss.

[0061] The segmentation loss is used to measure the difference between the predicted label and the true label; the variational regularization loss introduces a variational regularization term and quantifies the difference between the distribution output by the variational mixture of experts module and the standard normal distribution through the KL divergence.

[0062] The expression of the loss function is as follows: ; Where: represents the loss function; represents the segmentation loss; represents the predicted label; represents the true label; represents the regularization weight; represents the top K of the most relevant expert models selected; KL represents the KL divergence; represents the variance generated by the j-th expert model; represents the mean generated by the j-th expert model; represents the normal distribution of the j-th expert; N(0,1) represents the standard normal distribution.

[0063] A network structure that combines a U-Net, a variational autoencoder, and a mixture of experts model proposed by the method of the present invention enhances the feature representation ability by introducing an adaptive dictionary enhancement (E-Dictionary) module in the encoder, fuses multi-modal data features using a variational mixture of experts module at the bottleneck layer, and achieves high-precision segmentation through skip connections and a decoder. The E-Dictionary module optimizes the feature expression through dictionary learning and sparse coding; the variational mixture of experts block combines multiple expert networks and a shared expert network to dynamically select the optimal feature representation; the encoder-decoder structure of the U-Net ensures the integrity of spatial information. This multi-level and multi-modal collaborative design is particularly suitable for the complex segmentation tasks of laryngeal cancer and lymph nodes.

[0064] In practical applications, assume that a certain medical institution needs to accurately segment a patient's pharyngeal cancer and lymph nodes to assist in diagnosis and treatment planning. First, obtain the patient's CT or MRI three-dimensional medical image data as input data and import it into the network model corresponding to the method of the present invention. In the encoding path, the input data is input into the multi-level feature extraction and adaptive dictionary enhancement module as shown in Figure 2 for processing to generate a multi-scale enhanced feature map. The last-level enhanced feature map enters the bottleneck layer, and the variational mixture of experts module performs dynamic feature routing and uncertainty modeling. The decoding path gradually restores the spatial resolution through a four-level upsampling module and fuses multi-level feature information through skip connections, and finally generates a segmentation result. The segmentation result can be directly used to assist doctors in formulating surgical plans or evaluating treatment effects.

[0065] In addition, the present invention performs excellently in multi-modal medical image processing. For example, when processing multi-modal data such as T1, T1C, and T2-weighted MRI, the variational mixture of experts module selects the optimal expert for specialization processing according to different modalities or feature patterns through a dynamic routing mechanism, thereby enhancing the diversity of feature representation. The adaptive dictionary enhancement module optimizes the local structure expression through sample-specific dynamic dictionaries and attention mechanisms, improving the detail perception ability of the feature map. The multi-level feature fusion realizes the seamless fusion of low-level and high-level features through skip connections and the decoding path, ensuring the integrity of spatial information. These designs significantly improve the segmentation performance of the model for complex anatomical structures and solve the limitations of the prior art in terms of noise interference and insufficient feature expression problems.

[0066] In summary, the present invention solves the deficiencies of the prior art in the three-dimensional medical image segmentation task, especially the pharyngeal cancer and lymph node segmentation tasks, by introducing the variational mixture of experts module and the adaptive dictionary enhancement module. From the preprocessing of the input data to the generation of the final segmentation result, each step of the present invention is carefully designed and optimized to ensure the high performance and robustness of the model. The present invention is not only applicable to multiple medical imaging modalities but also can effectively cope with the segmentation challenges of complex anatomical structures, providing important technical support for the field of medical image segmentation.

[0067] Embodiment 2 As shown in Figure 3 , the second embodiment of the present invention also provides a medical image automatic segmentation device based on a variational mixture of experts model, including: An acquisition unit for acquiring three-dimensional medical images; A U-Net encoding unit for inputting the three-dimensional medical image into a U-Net encoder to perform multi-level convolution, downsampling feature extraction, and adaptive dictionary enhancement operations to optimize the local structure expression of features and generate each-level enhanced feature map; A variational mixture of experts unit is used to input the enhanced feature map of the last stage into a variational mixture of experts module for shared feature extraction, and fuse it with the output of the dynamic routing selection of multiple expert models to generate a fused feature; A decoding unit is used to input the fused feature combined with the enhanced feature map of the last stage into a U-Net decoder for upsampling operations, and concatenate the output of the upsampling with the enhanced feature map of the current level step by step to restore the spatial resolution of the image, obtaining a decoded fused feature; An output unit is used to input the decoded fused feature into a single-channel convolutional layer and output a segmentation result.

[0068] Embodiment III The third embodiment of the present invention also provides a medical image automatic segmentation device based on a variational mixture of experts model, which includes a memory and a processor. A computer program is stored in the memory, and the computer program can be executed by the processor to implement the medical image automatic segmentation method based on the variational mixture of experts model as described above.

[0069] Embodiment IV The fourth embodiment of the present invention also provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, the medical image automatic segmentation method based on the variational mixture of experts model as described above is implemented.

[0070] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0071] In addition, in each embodiment of the present invention, each functional module may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.

[0072] If the above-mentioned function is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes. It should be noted that in this document, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such a process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0073] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0074] It should be understood that the term " / and" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after.

[0075] Depending on the context, as used herein, the word "if" can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0076] The "first / second" mentioned in the embodiments is only to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0077] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for automatic segmentation of medical images based on a variational mixture expert model, characterized in that: include: S1, obtain three-dimensional medical images; S2, inputting the three-dimensional medical image into a U-Net encoder to perform multi-level convolution, downsampling feature extraction and adaptive dictionary enhancement operations to optimize the local structural expression of the features and generate each level of enhanced feature map; S3, input the enhanced feature map of the last level into the variational hybrid expert module for shared feature extraction, and fuse it with the dynamic routing selection result output of multiple expert models to generate fused features; S4, inputting the fusion feature into the U-Net decoder for upsampling in combination with the enhanced feature map of the last level, and concatenating the upsampled output with the enhanced feature map of the current level step by step to restore the spatial resolution of the image, thereby obtaining a decoded fusion feature; S5, input the decoded fusion features into a single-channel convolutional layer and output the segmentation result.

2. The method for automatic segmentation of medical images based on a variational mixture expert model according to claim 1, characterized in that ,Each level of the U-Net encoder includes a convolution block, a maximum pooling downsampling and an adaptive dictionary enhancement module; Perform three-dimensional convolution, activation and normalization operations through convolution blocks; Multi-scale feature extraction is performed through the maximum pooling downsampling operation to obtain feature maps at each level; The adaptive dictionary enhancement module performs an enhancement operation based on a dynamic dictionary and an attention mechanism on each sample of each level of feature map to optimize the local structural expression of the feature.

3. The method for automatic segmentation of medical images based on a variational mixture expert model according to claim 2, characterized in that: The adaptive dictionary enhancement module performs the following enhancement operations based on the dynamic dictionary and attention mechanism: Perform three-dimensional block extraction on the input feature map to obtain a block vector; Generate a structure dictionary for each sample of the feature map through a dynamic dictionary generator; The block vector and the structure dictionary of all samples are input into the multi-head attention module for enhancement and reconstruction to obtain an enhanced feature map.

4. The method for automatic segmentation of medical images based on a variational mixture expert model according to claim 3, characterized in that ,The three-dimensional block extraction operation is: The size of the three-dimensional block is set according to the level of the U-Net encoder. The input feature map is divided into non-overlapping blocks through a three-dimensional expansion operation, and each block is flattened into a one-dimensional vector to generate a set of block vectors. The expression is: ; in, represents the i-th flattened block vector; represents the field of real numbers; Indicates the current The number of channels of the layer; Represents the spatial edge length of the block; Represents the total number of elements of the block, that is, the volume of the three-dimensional block; The dynamic dictionary generator generates a structure dictionary for each sample in the feature map, specifically: Apply 3D adaptive average pooling to the feature map to compress the spatial dimension of the feature map to the smallest unit and flatten it into a one-dimensional vector; Expanding the one-dimensional vector to a hidden dimension through a first layer of fully connected network and applying a ReLU activation function to obtain a first fully connected vector; The first fully connected vector is passed through the second layer of fully connected network to generate a dictionary vector, the length of which is consistent with the number of channels, to obtain a structural dictionary, which is expressed as: ; in, Represents a structure dictionary; Represents a three-dimensional average pooling operation, which is used to compress the spatial dimension; Represents the flattening operation, which is used to convert a three-dimensional tensor into a one-dimensional vector; ReLU represents a fully connected layer that expands the number of channels to 2 times; ReLU represents an activation function; represents a fully connected layer that maps dimensions to dictionary sizes; Indicates the current The input feature map of the layer; Represents the dimension of the final generated structure dictionary; The operations of enhancement and reconstruction of the multi-head attention module are as follows: The block vector is used as the query vector, and the structure dictionary of all samples is used as the key vector and value vector to calculate the block vector enhanced by multi-head attention to focus on the structural expression and obtain the enhanced block vector, which is expressed as: ; in, Represents a set of enhanced block vectors; MultiHeadAttn represents a multi-head attention mechanism; P represents a set of block vectors; D represents a structure dictionary, which serves as both a key vector and a value vector in the multi-head attention mechanism; The enhanced block vector is reshaped into the original spatial dimension and residually connected with the input feature map to obtain the enhanced feature map, which is expressed as: ; in, Represents the enhanced feature map; A feature map representing the input; Indicates that Reshape into Operations of the same dimension; represents a residual connection.

5. The method for automatic segmentation of medical images based on a variational mixture expert model according to claim 2, characterized in that ,The variational hybrid expert module includes a shared feature extractor and an expert ,network consisting of multiple expert models for dynamic routing selection and ,probability distribution modeling of input features; Wherein, the shared feature extractor is used to extract the common pattern of input features to reduce redundant calculations of the expert model; Each of the expert models is composed of three variational U-Nets, and independently processes input features to generate variational parameters to perform feature extraction of input features of different modalities; the variational parameters include the mean and logarithmic variance of the expert model.

6. The method for automatic segmentation of medical images based on a variational mixture expert model according to claim 5, characterized in that ,The process of extracting input features through the shared feature extractor is: The input features are reduced in dimension through the first 3D convolutional layer to extract low-order features; Applying the ReLU activation function to enhance the nonlinear expression of the low-order features; The features activated by ReLU are restored to the original input feature dimension through the second three-dimensional convolutional layer, and the shared features are output. The expression is: ; in, represents the common features generated by the shared feature extractor; Represents the feature map of the nth level, that is, the output feature map of the last bottleneck layer; Represents the first three-dimensional convolution dimensionality reduction operation, that is, Channels down to channels; Represents the second 3D convolution recovery operation, which is composed of Channels restored to channels; , Both represent the number of channels, and .

7. The method for automatic segmentation of medical images based on a variational mixture expert model according to claim 5, characterized in that ,The process of each expert model processing input features is: The input features are mapped to the high-dimensional feature space through the three-dimensional convolutional layer of the U-Net encoder to extract high-dimensional features; Applying the ReLU activation function to enhance the nonlinear expression of the high-dimensional features; The features after ReLU activation are input into two parallel decoders respectively, and the corresponding variational parameters are generated through three-dimensional convolution; the expression is: ; in, represents the mean value generated by the j-th expert model; represents the variance generated by the jth expert model; Represents the mapping function of the j-th expert model; Represents the input features of the nth level, that is, the enhanced features of the last bottleneck layer; represents the activation function; Represents the three-dimensional convolution expansion operation, that is, Channels expanded to channels; , Both represent the number of channels, and .

8. The method for automatic segmentation of medical images based on a variational mixture expert model according to claim 7, characterized in that After the input features are processed by multiple expert models, the gating and routing mechanism is used for dynamic routing selection, and the routing output results are fused with the shared features for output. The specific process is as follows: Input features through The weighted mechanism calculates the selection weight vector of the expert model, which is expressed as: ; in, Represents the selection weight vector of the expert model; represents the normalization function; Represents a three-dimensional convolution operation, that is, Channels are mapped to channels; b represents a learnable bias term; According to the selected weight vector, the mean weighted sum of each selected expert model is calculated as the routing output corresponding to each sample in the input feature, and the expression is: ; in, The first Routing output of samples; Indicates In the sample Normalized weights of selected expert models; represents the number of top K most relevant experts; Indicates The average of the most relevant experts, ; (·) represents the mean calculation function; The routing outputs of all input feature samples are fused with the shared features and then output.

9. The method for automatic segmentation of medical images based on a variational mixture expert model according to claim 1, characterized in that ,It also includes using loss function for training optimization; The loss function includes a segmentation loss and a variational regularization loss, and a regularization weight is used to balance the segmentation loss and the variational regularization loss; The segmentation loss is used to measure the difference between the predicted label and the true label; The variational regularization loss introduces a variational regularization term, and quantifies the difference between the distribution of the variational hybrid expert module output and the standard normal distribution through KL divergence; The expression of the loss function is as follows: ; in: represents the loss function; represents the segmentation loss; represents the predicted label; represents the true label; represents the regularization weight; represents the top K most relevant expert models selected; KL represents KL divergence; represents the variance generated by the j-th expert model; represents the mean value generated by the j-th expert model; represents the normal distribution of the jth expert; N(0,1) represents the standard normal distribution.

10. A medical image automatic segmentation device based on variational mixture expert model, characterized in that: include: An acquisition unit, used for acquiring a three-dimensional medical image; A U-Net encoding unit, used for inputting the three-dimensional medical image into a U-Net encoder to perform multi-level convolution, downsampling feature extraction and adaptive dictionary enhancement operations to optimize the local structural expression of the features and generate each level of enhanced feature map; The variational hybrid expert unit is used to input the enhanced feature map of the last level into the variational hybrid expert module for shared feature extraction, and fuse it with the dynamic routing selection result output of multiple expert models to generate fused features; A decoding unit, used to input the fusion feature combined with the enhanced feature map of the last level into the U-Net decoder for upsampling operation, and gradually splice the upsampled output with the enhanced feature map of the current level to restore the spatial resolution of the image, and obtain a decoded fusion feature; The output unit is used to input the decoded fusion features into a single-channel convolutional layer and output the segmentation result.

Citation Information

Patent Citations

  • Method for generating DWI image based on CT cerebral infarction image

    CN117422788A

  • Medical image segmentation model training method, segmentation method, equipment and program product

    CN118115838A

  • Eye fundus image optic cup and optic disc segmentation method based on multi-score expert perception reasoning

    CN118247503A

  • Single-source-domain generalization medical image segmentation method and device based on shape dictionary

    CN119672346A

  • Devices for generation of synthetic 3D representations of myelin content

    WO2025068000A1

Cited By

  • Communication segmentation learning system and method for adaptive channel compression, and medium

    CN120896677A

  • Automatic wiring cable segmentation method based on hybrid expert model

    CN121366288A

  • Prompt medical image segmentation method based on interlayer feature fusion and balance expert

    CN122175993A