Image Segmentation Method, System, Terminal and Storage Medium Based on Parallel Network Framework and Dynamic Fusion
Through the parallel network framework and dynamic fusion method, the problems of insufficient global dependence and high computational cost in brain tumor image segmentation are solved, and accurate segmentation in the absence of modality is achieved, which improves the robustness and performance of image segmentation.
Patent Information
- Application Number
- CN202510429723.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The prior art problems of insufficient global dependence, high computational cost and poor adaptability due to lack of modality in brain tumor image segmentation.
The parallel network framework and dynamic fusion method are adopted to obtain gating features through dual convolution operations, and the state space model is used for global enhancement processing, and a single-modal representation is constructed. The missing modal information is fused through the dynamic sharing model, and the network is trained in combination with the balanced loss terms, divergence loss terms and mutual information loss terms, and the final segmentation result is output.
It improves the accuracy and robustness of image segmentation, can still produce accurate segmentation results when the mode is missing, and fully extracts features when the multimodal is complete, enhancing the overall segmentation performance.
Smart Images

Figure CN119942130B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional image segmentation, and particularly relates to an image segmentation method, system, terminal and computer-readable storage medium based on a parallel network framework and dynamic fusion. Background Art
[0002] With the rapid development of imaging technology, magnetic resonance imaging (MRI) plays a crucial role in brain tumor image segmentation.
[0003] However, in actual scenarios, due to various reasons (such as differences in scanning equipment or protocols, patient contrast agent allergies, imaging time or cost limitations, image damage, etc.), it is not always possible to ensure the acquisition of complete data for all modalities. If the missing modalities are simply regarded as having no information or the relevant cases are directly deleted, the generalization ability of the deep learning model will be greatly weakened, thereby affecting the accurate segmentation of brain tumor images.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of the present invention is to provide an image segmentation method, system, terminal and computer-readable storage medium based on a parallel network framework and dynamic fusion, aiming to solve the problems of insufficient global dependence, high computational cost and lack of adaptability in the application process in the existing technology for the brain tumor segmentation modeling process with missing modalities.
[0006] To achieve the above object, the present invention provides an image segmentation method based on a parallel network framework and dynamic fusion. The image segmentation method based on a parallel network framework and dynamic fusion includes the following steps:
[0007] Obtain multiple original features of a target image set, input all the original features into a parallel network framework for dual-path convolution operation to obtain gating features corresponding to the original features of each sample, and construct spatial features corresponding to each sample according to each gating feature;
[0008] Perform global enhancement processing on each spatial feature by using a state space model, and obtain all single-modal representations of each sample;
[0009] Input all the single-modal representations into a dynamic sharing model to obtain multiple predicted segmentation results of the target image set;
[0010] Obtain the original segmentation result of the target image set, construct the original classification probability and the predicted classification probability, construct a balance loss term according to the original segmentation result and multiple predicted segmentation results, and construct a divergence loss term according to the original classification probability and the predicted classification probability;
[0011] Construct the dependency relationship between the distribution probability in the full modality and the distribution probability in the missing modality, extract the deepest latent features of the distribution probability in the full modality and the distribution probability in the missing modality according to the dependency relationship, and construct a mutual information loss term according to the deepest latent features;
[0012] Use the balance loss term, the divergence loss term and the mutual information loss term to train the parallel network framework to obtain the target parallel network framework, and output the final segmentation result of the target image set through the target parallel network framework.
[0013] Optionally, in the image segmentation method based on a parallel network framework and dynamic fusion, where obtaining multiple original features of a target image set, inputting all the original features into the parallel network framework for dual-path convolution operations to obtain gating features corresponding to the original features of each sample, and constructing spatial features corresponding to each sample according to each gating feature, specifically includes:
[0014] Obtain the original features corresponding to each sample in the target image set, and input all the original features into multiple encoders in the parallel network framework respectively;
[0015] Extract the local features corresponding to the original features of each sample through the gating convolution module of each encoder, perform dual-path convolution operations on the multiple original features according to each local feature, and output the corresponding gating features:
[0016] ;
[0017] where, represents the gating feature output by the -th layer encoder, represents the feature input into the gating convolution module of the -th layer encoder, represents the gating convolution module, represents the non-linear activation function, represents 's convolution kernel, represents 's convolution kernel, represents element-wise multiplication;
[0018] Input all the gating features into the state space model in the parallel network framework. The state space model performs residual connection on each pair of the gating features and the features output by the state space model to obtain corresponding spatial features:
[0019] ;
[0020] wherein, represents the spatial features output by the -th layer encoder, represents the state space model.
[0021] Optionally, in the image segmentation method based on the parallel network framework and dynamic fusion, wherein the state space model is used to perform global enhancement processing on each of the spatial features and obtain all single-modal representations of each of the samples, specifically including:
[0022] Input the samples corresponding to each of the spatial features into multiple state space models of the parallel network framework in parallel for analysis;
[0023] Each of the state space models analyzes the feature connection relationships between different levels of a single modality in each of the samples, and constructs the global dependence information corresponding to each modality in each sample according to all the feature connection relationships;
[0024] Obtain the global information enhancement channels respectively constructed by each of the state space models, and construct the feature sequences of each modality in each of the samples according to each of the spatial features;
[0025] Input all the feature sequences in each sample into the corresponding global information enhancement channels to update the corresponding global dependence information in real time, and perform global enhancement processing on each of the spatial features according to the globally dependent information updated in real time;
[0026] Obtain the missing modality information input by the user, and obtain the single-modal representation corresponding to each sample according to the missing modality information.
[0027] Optionally, in the image segmentation method based on the parallel network framework and dynamic fusion, wherein the step of inputting all the single-modal representations into the dynamic sharing model to obtain multiple predicted segmentation results of the target image set specifically includes:
[0028] Input all the single-modal representations and the missing modality information into the dynamic sharing model, and construct an available modality set in the dynamic sharing model according to all the single-modal representations and the missing modality information:
[0029] ;
[0030] ;
[0031] Among them, represents the th unimodal representation of the th sample, represents the available modality index, represents the number of modalities;
[0032] The dynamic sharing model determines multiple missing modalities for each sample according to the available modality set, and performs a fusion process on the available modality set to obtain a predicted segmentation result of the available modality set under each missing modality:
[0033] ;
[0034] Among them, represents the predicted segmentation result of the th sample under the available modality.
[0035] Optionally, for the image segmentation method based on a parallel network framework and dynamic fusion, wherein, obtaining the original segmentation result of the target image set, constructing the original classification probability and the predicted classification probability, and constructing a balance loss term according to the original segmentation result and multiple predicted segmentation results, and constructing a divergence loss term according to the original classification probability and the predicted classification probability, specifically includes:
[0036] Obtain the original segmentation result and the original classification result of the target image set, construct the original classification probability according to the original segmentation result and the original classification result, and construct the predicted classification probability according to all the predicted segmentation results and the original classification result:
[0037] ;
[0038] ;
[0039] ;
[0040] ;
[0041] ;
[0042] Among them, represents the original classification probability, represents the predicted classification probability, represents the gamma function, represents the number of classifications in the original classification result, represents the parameter of the original classification probability, , and respectively represent the parameters of the first, second, and the parameters of the original classification probability, represents the parameter of the predicted classification probability, , and respectively represent the parameters of the first, second, and the parameters of the predicted classification probability, represents the standard d-dimensional simplex, parameters of the predicted classification probability, represents the predicted classification result in the predicted classification probability, represents the predicted classification result of the first category in the predicted classification probability, represents the predicted classification result of the second category in the predicted classification probability, represents the predicted classification result of the category in the predicted classification probability, represents the predicted classification result of the category in the predicted classification probability, represents the original classification result in the original classification probability, represents the original classification result of the first category in the original classification probability, represents the original classification result of the second category in the original classification probability, represents the original classification result of the category in the original classification probability, represents the original classification result of the
[0043] Analyze the overlap between the original segmentation result and multiple predicted segmentation results to construct a balanced loss term:
[0044] ;
[0045] where represents the balanced loss term, represents the true label of the th pixel in the original segmentation result, represents the predicted probability of the th pixel in the predicted segmentation result;
[0046] Construct a divergence loss term based on the original classification probability and the predicted classification probability:
[0047] ;
[0048] ;
[0049] ;
[0050] Among them, and represent the divergence loss term, represents the degree of difference, and are in a conjugate exponential relationship with each other, , and all represent hyperparameters, represents the original classification probability, represents the predicted classification probability, represents the original segmentation result, represents the predicted segmentation result.
[0051] Optionally, for the image segmentation method based on the parallel network framework and dynamic fusion, wherein, constructing the dependency relationship between the distribution probability in the full modality and the distribution probability in the missing modality, extracting the deepest latent features of the distribution probability in the full modality and the distribution probability in the missing modality according to the dependency relationship, and constructing the mutual information loss term according to the deepest latent features, specifically including:
[0052] Construct the distribution probability in the full modality and the distribution probability in the missing modality according to the available modality set and all the missing modalities of each sample;
[0053] Extract the true entropy of the distribution probability in the full modality and the predicted entropy of the distribution probability in the missing modality from the parallel network framework, and quantify the dependency relationship between the distribution probability in the full modality and the distribution probability in the missing modality according to the true entropy and the predicted entropy:
[0054] ;
[0055] Among them, represents the dependency relationship, represents the true entropy, represents the predicted entropy, represents the distribution probability in the full modality, represents the distribution probability in the missing modality;
[0056] Extract the deepest latent features of the distribution probability in the full modality and the distribution probability in the missing modality according to the dependency relationship and the true entropy, obtain the joint probability distribution between the distribution probability in the full modality and the distribution probability in the missing modality, and construct the mutual information loss term according to the joint probability distribution:
[0057] ;
[0058] ;
[0059] Among them, represents the joint probability distribution of the random variables P and , and represents the mutual information loss term, represents the probability of P given Q occurs.
[0060] Optionally, for the image segmentation method based on the parallel network framework and dynamic fusion, where training the parallel network framework using the balance loss term, the divergence loss term, and the mutual information loss term to obtain a target parallel network framework, and outputting a final segmentation result of the target image set through the target parallel network framework specifically includes:
[0061] Construct a total loss function according to the balance loss term, the divergence loss term, and the mutual information loss term:
[0062] ;
[0063] Among them, represents the total loss function, represents the balance loss term, represents the mutual information loss term, represents the divergence loss term, represents the weight adjustment coefficient of represents the weight adjustment coefficient of;
[0064] Input the total loss function into the parallel network framework for training to obtain a target parallel network framework;
[0065] Perform segmentation processing on the target image set through the decoding network in the target parallel network framework, and output the final segmentation result.
[0066] In addition, to achieve the above object, the present invention also provides an image segmentation system based on a parallel network framework and dynamic fusion, where the image segmentation system based on the parallel network framework and dynamic fusion includes:
[0067] A spatial feature construction module, configured to obtain multiple original features of a target image set, input all the original features into a parallel network framework for dual-path convolution operations to obtain gated features corresponding to the original features of each sample respectively, and construct spatial features corresponding to each sample according to each gated feature;
[0068] A modality extraction module, configured to globally enhance each of the spatial features by using a state space model and obtain all unimodal representations of each sample;
[0069] A result prediction module, configured to input all the unimodal representations into a dynamic sharing model to obtain multiple predicted segmentation results of the target image set;
[0070] A first loss term construction module, configured to obtain the original segmentation result of the target image set, construct an original classification probability and a predicted classification probability, construct a balanced loss term according to the original segmentation result and the multiple predicted segmentation results, and construct a divergence loss term according to the original classification probability and the predicted classification probability;
[0071] A second loss term construction module, configured to construct a dependency relationship between the distribution probability in the full modality and the distribution probability in the missing modality, extract the deepest latent features of the distribution probability in the full modality and the distribution probability in the missing modality according to the dependency relationship, and construct a mutual information loss term according to the deepest latent features;
[0072] A model training module, configured to use the balanced loss term, the divergence loss term, and the mutual information loss term to train the parallel network framework to obtain a target parallel network framework, and output a final segmentation result of the target image set through the target parallel network framework.
[0073] In addition, to achieve the above object, the present invention further provides a terminal, where the terminal includes: a memory, a processor, and an image segmentation program based on a parallel network framework and dynamic fusion stored on the memory and executable on the processor. When the image segmentation program based on the parallel network framework and dynamic fusion is executed by the processor, the steps of the above-mentioned image segmentation method based on the parallel network framework and dynamic fusion are implemented.
[0074] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores an image segmentation program based on a parallel network framework and dynamic fusion. When the image segmentation program based on the parallel network framework and dynamic fusion is executed by a processor, the steps of the above-mentioned image segmentation method based on the parallel network framework and dynamic fusion are implemented.
[0075] In the present invention, multiple original features of a target image set are obtained, and all the original features are input into a parallel encoding network for dual-path convolution operations to obtain gated features corresponding to multiple samples. Spatial features corresponding to each sample are constructed based on all the gated features; a state space model is used to globally enhance each of the spatial features to obtain multiple unimodal representations of multiple samples; all the unimodal representations are input into a dynamic sharing model to obtain multiple predicted segmentation results of the target image set; the original segmentation result of the target image set is obtained, an original classification probability and a predicted classification probability are constructed, and a balance loss term is constructed based on the original segmentation result and the predicted segmentation result, and a divergence loss term is constructed based on the original classification probability and the predicted classification probability; a dependency relationship between the original classification probability and the predicted classification probability is constructed, and the true posterior of the predicted classification probability is estimated according to the dependency relationship to obtain a mutual information loss term; the parallel encoding network is trained using the balance loss term, the divergence loss term, and the mutual information loss term to obtain a target parallel encoding network, and the final segmentation result of the target image set is output through the target parallel encoding network. In the present invention, multiple parallel encoder branches are provided to fully learn their own features, which can effectively avoid the drawback of insufficient fusion information caused by missing modalities, generate accurate image segmentation results when modalities are missing, and when multiple modalities are complete, each branch can fully extract the features of the modalities to enhance the overall image segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 is a flowchart of a preferred embodiment of the image segmentation method based on a parallel network framework and dynamic fusion according to the present invention;
[0077] Figure 2 is a schematic diagram of a parallel network framework of a preferred embodiment of the image segmentation method based on a parallel network framework and dynamic fusion according to the present invention;
[0078] Figure 3 is a schematic diagram of an image segmentation result of a preferred embodiment of the image segmentation method based on a parallel network framework and dynamic fusion according to the present invention;
[0079] Figure 4 is a schematic diagram of a whole tumor region segmentation result of a preferred embodiment of the image segmentation method based on a parallel network framework and dynamic fusion according to the present invention;
[0080] Figure 5 is a schematic diagram of a tumor core region segmentation result of a preferred embodiment of the image segmentation method based on a parallel network framework and dynamic fusion according to the present invention;
[0081] Figure 6 is a schematic diagram of an enhanced tumor segmentation result of a preferred embodiment of the image segmentation method based on a parallel network framework and dynamic fusion according to the present invention;
[0082] Figure 7 It is the segmentation result diagram of the cardiac image of the preferred embodiment of the image segmentation method based on the parallel network framework and dynamic fusion of the present invention;
[0083] Figure 8 It is the segmentation result diagram of the conventional scan of the preferred embodiment of the image segmentation method based on the parallel network framework and dynamic fusion of the present invention;
[0084] Figure 9 It is the segmentation result diagram of the abnormal state of the preferred embodiment of the image segmentation method based on the parallel network framework and dynamic fusion of the present invention;
[0085] Figure 10 It is the segmentation result diagram of the specificity of the preferred embodiment of the image segmentation method based on the parallel network framework and dynamic fusion of the present invention;
[0086] Figure 11 It is the structural diagram of the preferred embodiment of the image segmentation system based on the parallel network framework and dynamic fusion of the present invention;
[0087] Figure 12 It is the structural diagram of the preferred embodiment of the terminal of the present invention. Specific embodiments
[0088] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further elaborates on the present invention with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0089] The image segmentation method based on the parallel network framework and dynamic fusion described in the preferred embodiment of the present invention, as Figure 1 shown, the image segmentation method based on the parallel network framework and dynamic fusion includes the following steps:
[0090] Step S10: Obtain multiple original features of the target image set, input all the original features into the parallel network framework for dual-path convolution operation to obtain the gated features corresponding to the original features of each sample, and construct the spatial features corresponding to each sample according to each gated feature.
[0091] Among them, the structure of the robustness missing modality MRI segmentation method framework (parallel network framework) based on Mamba (selective state space model) state space modeling and information theory criteria proposed in this embodiment is specifically as Figure 2 shown (wherein, Figure 2 in represents the th layer encoder, represents the Layer encoder), where the Mamba-based encoder extracts features of multiple modalities ( respectively represent the features of the modalities, where is the mean of these features), which can be used to construct the subsequent mutual information loss term (that is , and the specific process is as described in the following steps). The Mamba-based decoder contains two activation functions, softmax and softplus. After decoding the features through these two activation functions, it outputs the segmentation features (that is (that is ) for constructing the balance loss term ) and the classification features (that is (that is ) for constructing the divergence loss term ). Then, the DS sharing mechanism (DynamicSharing, DS) is used to perform weighted average fusion on the segmentation results of available modalities to construct the predicted labels ( ), the ground truth labels ( ), and the fusion features ( ) for constructing and the fusion features ( ) for constructing .
[0092] Among them, the target image set (such as Figure 2 in ) is input into the encoder of the parallel network framework for convolution operations. The encoder extracts features from each modality through independent sub-encoders. Each encoder extracts features according to different modalities of each sample to generate preliminary feature representations. The parallel network framework can handle combinations of different modality data to cope with the challenge of modality missing and ensure that each modality information is independently and effectively utilized.
[0093] Specifically, the original features corresponding to each sample in the target image set are obtained, and all the original features are respectively input into multiple encoders in the parallel network framework; the local features corresponding to the original features of each sample are extracted through the gated convolution module of each encoder, and two-way convolution operations are performed on the multiple original features according to each local feature, and the corresponding gated features are output:
[0094] ;
[0095] Among them, represents the gated feature output by the -th layer encoder, and represents the feature input into the gated convolution module of the -th layer encoder, represents a gated convolution module, represents a non-linear activation function, represents the convolution kernel of represents the convolution kernel of represents element-wise multiplication; input all the gated features into the state space model in the parallel network framework, and the state space model performs residual connection on each pair of the gated features and the features output by the state space model to obtain corresponding spatial features:
[0096] ;
[0097] wherein, represents the spatial feature output by the -th layer encoder, represents the state space model.
[0098] Wherein, in each layer of the encoder, a GSC (Gated Spatial Convolution) module is introduced to enhance the local feature extraction and spatial information fusion capabilities of the model. The GSC module combines a 3×3×3 convolution kernel and a 1×1×1 convolution kernel through a dual-path convolution operation to obtain feature representations at different scales. These features are fused through a gating mechanism, effectively capturing spatial information while extracting local features. Among them, the gating mechanism can generate gated features by performing element-wise multiplication on the features output by the two-path convolutions, and combine these gated features with the original features input to the GSC through a residual connection. This design effectively improves the expression ability of local features, especially the ability to capture spatial features at different scales.
[0099] Step S20: Use the state space model to globally enhance each of the spatial features and obtain all single-modal representations of each sample.
[0100] Among them, in image segmentation tasks, traditional convolutional networks often rely on local receptive fields to obtain features. The limited size of the convolution kernel causes the model to pay more attention to local texture, boundaries and other information. Although the receptive field can be expanded to a certain extent by stacking multiple layers of convolutions or using pyramid pooling, etc., the modeling ability for large-scale and long-range dependencies is still insufficient. For example, in medical images, the context information of some anatomical structures or lesions often spans a large spatial range, and it is difficult to learn the global correlation of the entire image simply relying on local convolutions. Therefore, in the embodiments of the present application, an SSM (State Space Model) is introduced between the encoder and the decoder for global feature modeling and capturing long-range dependencies.
[0101] Specifically, the samples corresponding to each of the spatial features are respectively and parallelly input into a plurality of state space models of the parallel network framework for analysis; each state space model respectively analyzes the feature connection relationships between different levels of a single modality in each of the samples, and constructs global dependency information corresponding to each modality in each sample according to all the feature connection relationships; obtains a globally information-enhanced channel respectively constructed by each state space model, and constructs a feature sequence of each modality in each sample according to each of the spatial features; inputs all the feature sequences in each sample into the corresponding globally information-enhanced channel to perform real-time update on the corresponding global dependency information, and performs global enhancement processing on each of the spatial features according to the real-time updated global dependency information; obtains the missing modality information input by the user, and obtains a single-modal representation corresponding to each sample according to the missing modality information.
[0102] Among them, the state space model maintains a "state vector" or "state representation", which can be continuously updated during the feature flow of the network. Traditional convolution only performs convolution operations locally, while the state space model establishes a global or long-distance dependency between feature maps (or feature sequences), and updates features distributed at different positions or different resolution levels into a shared global "state". When the convolutional segmentation network is embedded with the SSM module, the corresponding position can obtain a representation of "feature aggregation within the entire image range" after one or multiple information interactions. In this way, it can better distinguish the relationships between different anatomical parts and enhance the capture of long-distance information (such as organ boundaries, lesion context). By introducing the SSM, the parallel network framework in this embodiment no longer only depends on the "small-range" features seen by the local convolution kernel, but can rely on the "global perspective" provided by the SSM to perform richer and more global reasoning, thereby improving the image segmentation accuracy and significantly enhancing the accuracy and robustness of the segmentation.
[0103] Step S30: Input all the single-modal representations into the dynamic sharing model to obtain multiple predicted segmentation results of the target image set.
[0104] Among them, in the medical image segmentation task, there is a situation of modality missing, and the decoder needs to dynamically fuse the segmentation results of each modality according to the available modality combinations. Therefore, the embodiment of the present application introduces a DS (Dynamic Sharing) mechanism for dynamically adjusting the fusion strategy according to the number of available modalities.
[0105] Specifically, input all the single-modal representations and the missing modality information into the dynamic sharing model, and construct an available modality set according to all the single-modal representations and the missing modality information in the dynamic sharing model:
[0106] ;
[0107] ;
[0108] wherein, represents the th single-modal representation of the th sample, represents the available modality index, represents the number of modalities; the dynamic sharing model determines multiple missing modalities for each sample according to the available modality set, and performs a fusion process on the available modality set to obtain a predicted segmentation result of the available modality set under each missing modality:
[0109] ;
[0110] wherein, represents the predicted segmentation result of the th sample under the available modalities.
[0111] Specifically, for each sample, the decoder selects a weighted average method to fuse the segmentation results of each modality according to the different available modalities, so as to ensure that the model can output an effective segmentation result even when some modalities are missing. Since the input modalities may be missing, a dynamic fusion method is required to merge the segmentation results of the available modalities. A set of segmentation results of the available modalities is set, and the segmentation results of the available modalities are fused by using a weighted sum and averaging method. The fused segmentation result is used as the final predicted output, so that the model can flexibly adapt to different combinations of available modalities, thereby enhancing the robustness in the case of missing modalities.
[0112] Furthermore, when performing fusion, if only the first and second modalities are available and the remaining modalities are missing, the single-modal output results of the first channel and the second channel are weighted and averaged, and the obtained result is considered as the missing modality segmentation image obtained when only the first and second modalities are available. By flexibly adapting to different combinations of modalities, the robustness of the model in the case of missing modalities is enhanced, and the accuracy of the segmentation result is ensured.
[0113] Step S40: Obtain the original segmentation result of the target image set, construct the original classification probability and the predicted classification probability, and construct a balance loss term according to the original segmentation result and the multiple predicted segmentation results, and construct a divergence loss term according to the original classification probability and the predicted classification probability.
[0114] Among them, in order to further optimize the parallel network framework in this application, a loss function is designed by combining the Dice loss (an index for measuring the similarity between two sets), the Hölder divergence (Kullback-Leibler divergence, KL divergence, a method for measuring the difference between two probability distributions) loss based on the Dirichlet distribution, and the mutual information loss.
[0115] Specifically, obtain the original segmentation result and the original classification result of the target image set, construct the original classification probability according to the original segmentation result and the original classification result, and construct the predicted classification probability according to all the predicted segmentation results and the original classification result:
[0116] ;
[0117] ;
[0118] ;
[0119] ;
[0120] ;
[0121] Among them, represents the original classification probability, represents the predicted classification probability, represents the gamma function, represents the number of classifications in the original classification result, represents the parameter of the original classification probability, 、 and respectively represent the parameters of the 1st, 2nd, and th original classification probabilities, represents the parameter of the th original classification probability, represents the parameter of the predicted classification probability, 、 and respectively represent the parameters of the 1st, 2nd, and th predicted classification probabilities, represents the standard -dimensional simplex, represents the parameter of the th predicted classification probability, represents the predicted classification result in the predicted classification probability, represents the predicted classification result of the 1st category in the predicted classification probability, represents the predicted classification result of the 2nd category in the predicted classification probability, represents the predicted classification result of the category in the predicted classification probability, represents the predicted classification result of the category in the predicted classification probability, represents the original classification result in the original classification probability, represents the original classification result of the 1st category in the original classification probability, represents the original classification result of the 2nd category in the original classification probability, represents the original classification result of the category in the original classification probability, represents the original classification result of the category in the original classification probability; analyze the overlap between the original segmentation result and the multiple predicted segmentation results to construct a balanced loss term:
[0122] ;
[0123] where, represents the balanced loss term, represents the true label of the th pixel in the original segmentation result, represents the predicted probability of the th pixel in the predicted segmentation result; construct a divergence loss term according to the original classification probability and the predicted classification probability:
[0124] ;
[0125] ;
[0126] ;
[0127] where, and represent the divergence loss term, represents the degree of difference, and have a conjugate exponential relationship with each other, , and all represent hyperparameters, represents the original classification probability, represents the predicted classification probability, represents the original segmentation result, represents the predicted segmentation result.
[0128] Among them, by introducing the Dirichlet distribution, the parallel network framework can accurately capture and represent the uncertainty between multiple modalities, effectively improving the robustness of the network.
[0129] Further, in this embodiment, the evidence value is obtained by applying the softplus activation function (a mathematical function commonly used in the field of deep learning) to the logits (logical units):
[0130] ;
[0131] where represents the index of the classification in the original classification result.
[0132] Further, the Hölder divergence loss based on the Dirichlet distribution can flexibly and robustly quantify the differences between probability distributions, thereby calculating the difference between the true Dirichlet distribution and the predicted Dirichlet distribution in this embodiment, and thus calculating the divergence-based loss.
[0133] Step S50: Construct the dependency relationship between the distribution probability in the full modality and the distribution probability in the missing modality, extract the deepest latent features of the distribution probability in the full modality and the distribution probability in the missing modality according to the dependency relationship, and construct a mutual information loss term according to the deepest latent features.
[0134] Among them, the mutual information loss function can enhance the ability of the model to extract relevant information from available modalities, especially in the case where some modalities are missing.
[0135] Specifically, according to the available modality set and all the missing modalities of each sample, construct the distribution probability in the full modality and the distribution probability in the missing modality; extract the true entropy of the distribution probability in the full modality and the predicted entropy of the distribution probability in the missing modality from the parallel network framework, and quantify the dependency relationship between the distribution probability in the full modality and the distribution probability in the missing modality according to the true entropy and the predicted entropy:
[0136] ;
[0137] where represents the dependency relationship, represents the true entropy, represents the predicted entropy, represents the distribution probability in the full modality, represents the distribution probability in the missing modality; extract the deepest latent features of the distribution probability in the full modality and the distribution probability in the missing modality according to the dependency relationship and the true entropy, obtain the joint probability distribution between the distribution probability in the full modality and the distribution probability in the missing modality, and construct a mutual information loss term according to the joint probability distribution:
[0138] ;
[0139] ;
[0140] Among them, represents the joint probability distribution of the random variables P and and the joint probability distribution, represents the mutual information loss term, represents the probability of P given Q.
[0141] Among them, in this embodiment, by random quantization of the dependence relationship between the distribution probabilities in the full modality and the distribution probabilities in the missing modality, the correlation between different modalities or different feature branches is emphasized, and the network is encouraged to still be able to extract relevant and effective information from the remaining modalities in the case of missing modalities, that is, to extract the deepest latent features in the two distribution probabilities.
[0142] Step S60, use the balance loss term, the divergence loss term and the mutual information loss term to train the parallel network framework to obtain a target parallel network framework, and output the final segmentation result of the target image set through the target parallel network framework.
[0143] Specifically, construct a total loss function according to the balance loss term, the divergence loss term and the mutual information loss term:
[0144] ;
[0145] Among them, represents the total loss function, represents the balance loss term, represents the mutual information loss term, represents the divergence loss term, represents the weight adjustment coefficient of represents the weight adjustment coefficient of ; input the total loss function into the parallel network framework for training to obtain a target parallel network framework; perform segmentation processing on the target image set through the decoding network in the target parallel network framework, and output the final segmentation result.
[0146] Among them, in addition to the Dice loss, the Hölder divergence (HD) loss based on the Dirichlet distribution and the mutual information (MI) loss are introduced. This combination is not common in multi-modal medical image segmentation. In particular, the Hölder divergence based on Dirichlet is used to characterize the uncertainty of the prediction classification, and the combination with the mutual information term can better balance the accuracy and uncertainty mining in the training process.
[0147] Furthermore, asFigure 3 As shown, it presents a comparison of image segmentation performance under different missing modality conditions. Figure 3 In (a), it represents the segmentation results based on the CHAOS dataset (Combined (CT-MR) Healthy Abdominal Organ Segmentation, a challenge mainly for abdominal organ segmentation). Each row (from top to bottom) shows the results under IN (input), OUT (output), and full-modal input. Each column (from left to right) compares the segmentation outputs of ACN (the first method, Automatic Control Network), SMU-Net (the second method, Style Matching U-Net), M³AE (the third method, Multimodal Masked Autoencoders, where Multimodal means multi-modal, Masked means masked, and Autoencoders means autoencoders), MMCFormer (the fourth method, Missing Modality Compensation Transformer), MML-MM-SSF (the fifth method, Multimodal Machine Learning - Multimodal Semantic Fusion Model), GGMD (the sixth method, Gradient-Guided Modality Decoupling for Missing-Modality Robustness), and the method used in this embodiment. The last column shows the ground truth image. Figure 3 In (b), it is the segmentation result based on the BraTS 2018 dataset (an authoritative open-source dataset for brain tumor segmentation in the field of medical imaging). Each row (from top to bottom) shows the results of T1 (the first modality), T2 (the second modality), T1ce (the third modality), Flair (the fourth modality), and full-modal input. Each column compares the segmentation outputs of different methods under various missing modalities. The last column shows the ground truth image, which highlights the robustness and effectiveness of the method proposed in this embodiment compared with the prior art in various missing modality scenarios.
[0148] Specifically, the effectiveness of the method used in this embodiment can be judged by the data in Table 1, Table 2, and Table 3 below:
[0149] Table 1: Evaluation Table for WT (Whole Tumor) Segmentation Results
[0150]
[0151] Table 2: Evaluation Table for ET (Enhancing Tumor) Segmentation Results
[0152]
[0153] Table 3: Evaluation Table for TC (Tumor Core) Segmentation Results
[0154]
[0155] Among them, the horizontal elements in Table 1, Table 2, and Table 3 all represent different unimodals, combinations of different unimodals, and full-modal inputs. The last column represents the mean of all modal segmentation results. In the first column, multiple methods (models) are divided. U-HVED represents the Unruh-Hybrid Variable Energy Detection Model, RFNet represents the Region-aware Fusion Network, mmFormer represents a deep learning network (Multimodal Medical Transformer) designed specifically for medical image segmentation tasks (especially brain tumor segmentation), EMRM represents the Enhanced Mobile Radio Module, SSFM represents the method of "Multi-modal Learning with Missing Modality via Shared-Specific Feature Modelling", QuMo represents the Quantum Mobility, GSS represents the method of "Scratch Each Other’s Back: Incomplete Multi-modal Brain Tumor Segmentation Via Category Aware Group Self-Support Learning", MFTrans represents a multi-functional neural network model based on the Transformer architecture, MTI represents the method of "Multimodal Transformer of Incomplete MRI Data for Brain Tumor Segmentation", GGDM represents the method of "Gradient-Guided Modality Decoupling for Missing-Modality Robustness", and OUR represents the parallel network framework used in the present invention.
[0156] Furthermore, as Figure 4 , Figure 5 andFigure 6 As shown, in this embodiment, a quantitative comparison of the segmentation results of the BraTS 2018 and BraTS 2020 datasets (authoritative open-source datasets for brain tumor segmentation in the field of medical imaging) under different missing modalities was also investigated. Specifically, the segmentation results of the whole tumor region, the tumor core region, and the enhanced tumor were compared. According to Figures 4 - 6 the results in, the effectiveness of the method proposed in this application in dealing with incomplete data can be clarified, showing a superior reconstruction quality compared to the baseline method.
[0157] Furthermore, as Figure 7 , Figure 8 , Figure 9 and Figure 10 shown, in this embodiment, a quantitative comparison of the segmentation results of the CHAOS dataset (Combined (CT-MR) Healthy Abdominal Organ Segmentation, a challenge mainly for abdominal organ segmentation) under different missing modalities was investigated. Specifically, the quantitative comparison results of different methods for cardiac imaging (as Figure 7 shown, LV-dice, Left Ventricle dice coefficient), conventional scans (as Figure 8 shown, Routine K-space, medical imaging scan), abnormal states (as Figure 9 shown, LK, Lucas-Kanade method), and specificity (as Figure 10 shown, SP, Standardized Patients, standardized patients) were compared; Figures 7 - 10 The reconstruction performance of the CHAOS dataset in various missing modality cases was demonstrated, proving that the method used in this embodiment can achieve high-quality segmentation even when the input data is incomplete, and the performance exceeds the baseline technology.
[0158] The present invention sets multiple parallel encoder branches to fully learn its own features, which can effectively avoid the disadvantage of insufficient fusion information caused by missing modalities and generate accurate image segmentation results when modalities are missing. When multiple modalities are complete, each branch can fully extract the features of the modality, enhancing the overall image segmentation performance.
[0159] Furthermore, as Figure 11 shown, based on the above image segmentation method based on a parallel network framework and dynamic fusion, the present invention also correspondingly provides an image segmentation system based on a parallel network framework and dynamic fusion. Among them, the image segmentation system based on a parallel network framework and dynamic fusion includes:
[0160] A spatial feature construction module 51, configured to obtain multiple original features of a target image set, input all the original features into a parallel encoding network for dual-path convolution operations to obtain gated features corresponding to multiple samples, and construct spatial features corresponding to each sample according to all the gated features;
[0161] A modality extraction module 52, configured to globally enhance each of the spatial features by using a state space model and obtain all single-modal representations of each sample;
[0162] A result prediction module 53, configured to input all the single-modal representations into a dynamic sharing model to obtain multiple predicted segmentation results of the target image set;
[0163] A first loss term construction module 54, configured to obtain the original segmentation result of the target image set, construct an original classification probability and a predicted classification probability, construct a balanced loss term according to the original segmentation result and the predicted segmentation result, and construct a divergence loss term according to the original classification probability and the predicted classification probability;
[0164] A second loss term construction module 55, configured to construct a dependency relationship between the original classification probability and the predicted classification probability, estimate the true posterior of the predicted classification probability according to the dependency relationship to obtain a mutual information loss term;
[0165] A model training module 56, configured to train the parallel encoding network by using the balanced loss term, the divergence loss term, and the mutual information loss term to obtain a target parallel encoding network, and output a final segmentation result of the target image set through the target parallel encoding network.
[0166] Further, as Figure 12 shown, based on the above image segmentation method and system based on a parallel network framework and dynamic fusion, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 12 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0167] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In some other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store the application software installed on the terminal and various types of data, such as the program code of the installed terminal, etc. The memory 20 may also be used to temporarily store the data that has been output or will be output. In one embodiment, an image segmentation program 40 based on a parallel network framework and dynamic fusion is stored on the memory 20, and the image segmentation program 40 based on the parallel network framework and dynamic fusion can be executed by the processor 10, so as to implement the image segmentation method based on the parallel network framework and dynamic fusion in the present application.
[0168] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chips, and is used to run the program code stored in the memory 20 or process data, such as executing the image segmentation method based on the parallel network framework and dynamic fusion, etc.
[0169] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. The display 30 is used to display the information on the terminal and to display a visual user interface. The components of the terminal communicate with each other through a system bus.
[0170] In one embodiment, when the processor 10 executes the image segmentation program 40 based on the parallel network framework and dynamic fusion in the memory 20, the steps of the image segmentation method based on the parallel network framework and dynamic fusion as described above are implemented.
[0171] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an image segmentation program based on a parallel network framework and dynamic fusion, and when the image segmentation program based on the parallel network framework and dynamic fusion is executed by a processor, the steps of the image segmentation method based on the parallel network framework and dynamic fusion as described above are implemented.
[0172] In summary, the present invention provides an image segmentation method and related devices based on a parallel network framework and dynamic fusion. The method includes: obtaining multiple original features of a target image set, inputting all the original features into a parallel encoding network for dual-path convolution operations to obtain gated features corresponding to multiple samples, and constructing spatial features corresponding to each sample according to all the gated features; using a state space model to perform global enhancement processing on each spatial feature to obtain multiple single-modal representations of multiple samples; inputting all the single-modal representations into a dynamic sharing model to obtain multiple predicted segmentation results of the target image set; obtaining the original segmentation results of the target image set, constructing original classification probabilities and predicted classification probabilities, and constructing a balance loss term according to the original segmentation results and the predicted segmentation results, and constructing a divergence loss term according to the original classification probabilities and the predicted classification probabilities; constructing a dependency relationship between the original classification probabilities and the predicted classification probabilities, estimating the true posterior of the predicted classification probabilities according to the dependency relationship to obtain a mutual information loss term; training the parallel encoding network using the balance loss term, the divergence loss term, and the mutual information loss term to obtain a target parallel encoding network, and outputting the final segmentation results of the target image set through the target parallel encoding network. The present invention sets multiple parallel encoder branches to fully learn their own features, which can effectively avoid the drawback of insufficient fusion information caused by missing modalities, generate accurate image segmentation results when modalities are missing, and when multiple modalities are complete, each branch can fully extract the features of the modalities to enhance the overall image segmentation performance.
[0173] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal including that element.
[0174] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer, and when the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0175] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or modifications can be made according to the above description, and all such improvements and modifications shall fall within the protection scope of the appended claims of the present invention.
Claims
1. An image segmentation method based on a parallel network framework and dynamic fusion, characterized in that The image segmentation method based on a parallel network framework and dynamic fusion includes: Obtain multiple original features of a target image set, input all the original features into a parallel network framework for dual-path convolution operation to obtain gating features corresponding to the original features of each sample respectively, and construct spatial features corresponding to each sample according to each gating feature; Perform global enhancement processing on each spatial feature by using a state space model, and obtain all unimodal representations of each sample; Input all the unimodal representations into a dynamic sharing model to obtain multiple predicted segmentation results of the target image set; Obtain the original segmentation result of the target image set, construct an original classification probability and a predicted classification probability, construct a balance loss term according to the original segmentation result and the multiple predicted segmentation results, and construct a divergence loss term according to the original classification probability and the predicted classification probability; Construct a dependency relationship between the distribution probability in the full modality and the distribution probability in the missing modality, extract the deepest-layer latent features of the distribution probability in the full modality and the distribution probability in the missing modality according to the dependency relationship, and construct a mutual information loss term according to the deepest-layer latent features; Use the balance loss term, the divergence loss term, and the mutual information loss term to train the parallel network framework to obtain a target parallel network framework, and output the final segmentation result of the target image set through the target parallel network framework.
2. The image segmentation method based on a parallel network framework and dynamic fusion according to claim 1, wherein, The obtaining of multiple original features of a target image set, inputting all the original features into a parallel network framework for dual-path convolution operation to obtain gating features corresponding to the original features of each sample respectively, and constructing spatial features corresponding to each sample according to each gating feature specifically includes: Obtain the original features corresponding to each sample in the target image set, and input all the original features into multiple encoders in the parallel network framework respectively; Extract local features corresponding to the original features of each sample through the gating convolution module of each encoder, perform dual-path convolution operation on the multiple original features according to each local feature, and output corresponding gating features: ; Among them, represents the gated feature output by the -th layer encoder, represents the feature input to the gated convolution module in the -th layer encoder, represents the gated convolution module, represents the non-linear activation function, represents 's convolution kernel, represents 's convolution kernel, represents element-wise multiplication; Input all the gating features into the state space model in the parallel network framework, and the state space model performs residual connection on each pair of the gating features and the features output by the state space model to obtain corresponding spatial features: ; Among them, represents the spatial features output by the layer encoder, represents the state space model.
3. The image segmentation method based on a parallel network framework and dynamic fusion according to claim 1, wherein The performing of global enhancement processing on each spatial feature by using a state space model, and obtaining all unimodal representations of each sample specifically includes: Input the samples corresponding to each spatial feature into multiple state space models of the parallel network framework in parallel for analysis; Each state space model respectively analyzes the feature connection relationship between different levels of a single modality in each sample, and constructs global dependency information corresponding to each modality in each sample according to all the feature connection relationships; Obtain the global information enhancement channels respectively constructed by each state space model, and construct a feature sequence of each modality in each sample according to each spatial feature; Input all the feature sequences in each sample into the corresponding global information enhancement channel to update the corresponding global dependency information in real time, and perform global enhancement processing on each spatial feature according to the real-time updated global dependency information; Obtain the missing modality information input by the user, and obtain the unimodal representation corresponding to each sample according to the missing modality information.
4. The image segmentation method based on a parallel network framework and dynamic fusion according to claim 3, wherein, The step of inputting all the unimodal representations into the dynamic sharing model to obtain multiple predicted segmentation results of the target image set specifically includes: Input all the unimodal representations and the missing modality information into the dynamic sharing model, and construct an available modality set according to all the unimodal representations and the missing modality information in the dynamic sharing model: ; ; Among them, represents the th single-modal representation of the th sample, represents the available modality index, represents the number of modalities; The dynamic sharing model determines multiple missing modalities of each sample according to the available modality set, and performs fusion processing on the available modality set to obtain the predicted segmentation results of the available modality set under each missing modality: ; Among them, represents the predicted segmentation result of the th sample under the available modalities.
5. The image segmentation method based on a parallel network framework and dynamic fusion according to claim 1, characterized in that The step of obtaining the original segmentation result of the target image set, constructing the original classification probability and the predicted classification probability, and constructing a balance loss term according to the original segmentation result and multiple predicted segmentation results, and constructing a divergence loss term according to the original classification probability and the predicted classification probability specifically includes: Obtain the original segmentation result and the original classification result of the target image set, construct the original classification probability according to the original segmentation result and the original classification result, and construct the predicted classification probability according to all the predicted segmentation results and the original classification result: ; ; ; ; ; Among them, represents the original classification probability, represents the predicted classification probability, represents the gamma function, represents the number of classifications in the original classification result, represents the parameter of the original classification probability, 、 and respectively represent the parameters of the first, second, and Kth original classification probabilities, represents the th parameter of the original classification probability, represents the parameter of the predicted classification probability, 、 and respectively represent the parameters of the first, second, and th predicted classification probabilities, represents the standard dimensional simplex, represents the th parameter of the predicted classification probability, represents the predicted classification result in the predicted classification probability, represents the predicted classification result of the first category in the predicted classification probability, represents the predicted classification result of the second category in the predicted classification probability, represents the predicted classification result of the th category in the predicted classification probability, represents the predicted classification result of the th category in the predicted classification probability, represents the original classification result in the original classification probability, represents the original classification result of the first category in the original classification probability, represents the original classification result of the second category in the original classification probability, represents the original classification result of the th category in the original classification probability, represents the original classification result of the th category in the original classification probability; Analyze the overlapping situation between the original segmentation result and multiple predicted segmentation results to construct a balance loss term: ; Among them, represents the balance loss term, represents the true label of the -th pixel in the original segmentation result, represents the predicted probability of the -th pixel in the predicted segmentation result; Construct a divergence loss term according to the original classification probability and the predicted classification probability: ; ; ; Among them, and represent the divergence loss term, represents the difference degree, and have a conjugate exponent relationship with each other, , and all represent hyperparameters, represents the original classification probability, represents the predicted classification probability, represents the original segmentation result, represents the predicted segmentation result.
6. The image segmentation method based on a parallel network framework and dynamic fusion according to claim 4, characterized in that The step of constructing the dependency relationship between the distribution probability under the full modality and the distribution probability under the missing modality, extracting the deepest latent features of the distribution probability under the full modality and the distribution probability under the missing modality according to the dependency relationship, and constructing a mutual information loss term according to the deepest latent features specifically includes: Construct the distribution probability under the full modality and the distribution probability under the missing modality according to the available modality set and all the missing modalities of each sample; Extract the true entropy of the distribution probability under the full modality and the predicted entropy of the distribution probability under the missing modality from the parallel network framework, and quantify the dependency relationship between the distribution probability under the full modality and the distribution probability under the missing modality according to the true entropy and the predicted entropy: ; Among them, represents a dependency relationship, represents the true entropy, represents the predicted entropy, represents the distribution probability in the full-modal state, represents the distribution probability in the missing-modal state; Extract the deepest latent features of the distribution probability under the full modality and the distribution probability under the missing modality according to the dependency relationship and the true entropy, obtain the joint probability distribution between the distribution probability under the full modality and the distribution probability under the missing modality, and construct a mutual information loss term according to the joint probability distribution: ; ; Among them, represents the joint probability distribution of the random variables P and , represents the mutual information loss term, represents the probability of P given Q occurs.
7. The image segmentation method based on a parallel network framework and dynamic fusion according to claim 1, wherein The step of training the parallel network framework using the balance loss term, the divergence loss term and the mutual information loss term to obtain a target parallel network framework, and outputting the final segmentation result of the target image set through the target parallel network framework specifically includes: Construct a total loss function based on the balance loss term, the divergence loss term, and the mutual information loss term: ; Among them, represents the total loss function, represents the balance loss term, represents the mutual information loss term, represents the divergence loss term, represents the weight adjustment coefficient of represents the weight adjustment coefficient of Input the total loss function into the parallel network framework for training to obtain a target parallel network framework; Perform segmentation processing on the target image set through the decoding network in the target parallel network framework, and output the final segmentation result.
8. An image segmentation system based on a parallel network framework and dynamic fusion, characterized in that, The image segmentation system based on the parallel network framework and dynamic fusion is applied to the image segmentation method based on the parallel network framework and dynamic fusion according to any one of claims 1-7. The image segmentation system based on the parallel network framework and dynamic fusion includes: A spatial feature construction module, configured to obtain multiple original features of a target image set, input all the original features into a parallel network framework for dual-path convolution operations to obtain gating features corresponding to the respective original features of each sample, and construct spatial features corresponding to each sample according to each gating feature; A modality extraction module, configured to perform global enhancement processing on each spatial feature by using a state space model and obtain all single-modal representations of each sample; A result prediction module, configured to input all the single-modal representations into a dynamic sharing model to obtain multiple predicted segmentation results of the target image set; A first loss term construction module, configured to obtain the original segmentation result of the target image set, construct an original classification probability and a predicted classification probability, construct a balance loss term according to the original segmentation result and the multiple predicted segmentation results, and construct a divergence loss term according to the original classification probability and the predicted classification probability; A second loss term construction module, configured to construct a dependency relationship between the distribution probability in the full modality and the distribution probability in the missing modality, extract the deepest latent features of the distribution probability in the full modality and the distribution probability in the missing modality according to the dependency relationship, and construct a mutual information loss term according to the deepest latent features; A model training module, configured to use the balance loss term, the divergence loss term, and the mutual information loss term to train the parallel network framework to obtain a target parallel network framework, and output the final segmentation result of the target image set through the target parallel network framework.
9. A terminal, characterized in that, The terminal includes: a memory, a processor, and an image segmentation program based on the parallel network framework and dynamic fusion stored on the memory and executable on the processor. When the image segmentation program based on the parallel network framework and dynamic fusion is executed by the processor, the steps of the image segmentation method based on the parallel network framework and dynamic fusion according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an image segmentation program based on the parallel network framework and dynamic fusion. When the image segmentation program based on the parallel network framework and dynamic fusion is executed by a processor, the steps of the image segmentation method based on the parallel network framework and dynamic fusion according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Parallel fractural network evolution image segmentation method
CN102831613A
MRI (Magnetic Resonance Imaging) image segmentation method with adaptive mode
CN119478407A