Multi-modal image segmentation method and system based on missing modality, terminal and storage medium

By constructing similarity, divergence and mutual information loss terms in a three-dimensional convolutional neural network, and processing multimodal image segmentation method, the problem of model performance degradation in the absence of multiple modes is solved, efficient multimodal image segmentation is achieved, and the stability and accuracy of the model in the case of incomplete data is enhanced.

CN119963845APending Publication Date: 2025-05-09SHENZHEN MSU-BIT UNIVERSITY

Patent Information

Application Number
CN202510423507.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art cannot maintain model performance when multiple modes are missing, cannot improve segmentation accuracy when taking into account multimodals, and the model performance declines in few modes, making it difficult to cover the needs of various scenarios.

Method used

A multimodal image segmentation method based on a three-dimensional convolutional neural network is adopted. By acquiring multiple modal data of the target image set, a single-modal representation is output, and a similarity loss term, divergence loss term and mutual information loss term are constructed. The mutual information is processed through the variational information maximization method, and the three-dimensional convolutional neural network is trained to obtain the final segmentation result.

Benefits of technology

Maintain model performance in the absence of modality, improve multimodal segmentation accuracy, enhance the stability and accuracy of the model in the case of incomplete data, and is suitable for image segmentation tasks in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963845A_ABST
    Figure CN119963845A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal image segmentation method and system based on missing modalities, a terminal and a storage medium, and the method comprises the steps: obtaining a target image set, carrying out the independent parallel processing of each modal of each sample, and combining divergence loss, mutual information transmission and the dynamic fusion of all single-modal representations of each sample, thereby obtaining a target image set; and constructing a total loss function to train a model, and finally outputting a segmentation result of the target image set through the trained model. According to the method, the condition of multi-modal data incompleteness is comprehensively considered, feature alignment between different modals is effectively promoted through divergence loss and a high mutual information knowledge transfer method, the cross-modal information transfer precision is improved, the difference between a prediction result and a label is effectively measured, missing modals can be more effectively processed, and the accuracy of cross-modal information transfer is improved. And the stability and the accuracy of the model under the condition of incomplete data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image segmentation technology, and in particular to a missing modality-based multimodal image segmentation method, system, terminal and computer-readable storage medium. Background Art

[0002] With the continuous development of imaging technology, different imaging modalities (such as ultrasound) can provide multi-angle and multi-level information for the same target area. Multimodal image segmentation or recognition can comprehensively utilize the complementary information between different image modalities to obtain more accurate analysis results.

[0003] However, in practice, some modality data are often missing due to objective conditions (object tolerance, cost, time constraints, equipment limitations, etc.), making it difficult for algorithms originally designed for multimodality to achieve ideal performance in actual applications.

[0004] However, existing modality-specific methods require separate training of corresponding networks for different combinations of missing modalities, resulting in the number of models growing exponentially with the number of modality combinations. The computational and storage overhead is extremely considerable, which poses a major obstacle to application. As for the unified single model method, when the number of available modalities is extremely limited (for example, there is only a single modality or a very small number of modalities), the performance of the model usually drops significantly, making it difficult to cover the needs of various scenarios.

[0005] Therefore, the prior art still needs to be improved and developed. Summary of the invention

[0006] The main purpose of the present invention is to provide a multimodal image segmentation method, system, terminal and computer-readable storage medium based on missing modalities, aiming to solve the problems in the prior art that the model performance cannot be maintained on the basis of existing modality information under the condition of missing multiple modalities, and the segmentation accuracy cannot be improved while taking into account multiple modalities without significantly increasing the number of models and parameter scale, and the model effectiveness cannot be improved in the case of few modalities.

[0007] To achieve the above object, the present invention provides a multimodal image segmentation method based on missing modality, and the multimodal image segmentation method based on missing modality comprises the following steps: Acquire a target image set, extract multiple modal data of each sample according to the target image set, input all the modal data into a three-dimensional convolutional neural network, and output all single modal representations of each modal data; Obtaining an initial segmentation result of the target image set, generating a segmentation prediction result according to all the unimodal representations, and constructing a similarity loss term according to the initial segmentation result and the segmentation prediction result; Constructing a divergence loss term according to the segmentation prediction result and the initial segmentation result; Extracting multiple feature vector pairs in the three-dimensional convolutional neural network, constructing mutual information between the complete mode and the missing mode according to the multiple feature vector pairs, and processing the mutual information by a variational information maximization method to obtain a mutual information loss term; The three-dimensional convolutional neural network is trained using the similarity loss term, the divergence loss term and the mutual information loss term to obtain a target three-dimensional convolutional neural network, and the current image set is segmented by the target three-dimensional convolutional neural network to obtain a final segmentation result.

[0008] Optionally, the multimodal image segmentation method based on missing modality, wherein the step of acquiring a target image set, extracting multiple modality data of each sample according to the target image set, inputting all the modality data into a three-dimensional convolutional neural network, and outputting all single modality representations of each modality data, specifically includes: Acquire a target image set, and extract multiple modality data of each sample in the target image set; Input all the modal data of each sample into the constructed three-dimensional convolutional neural network, and output multiple single-modal features of multiple modal data respectively through multiple channel encoders in the three-dimensional convolutional neural network: ; in, Representation sample The The unimodal features of the modes, Representation sample The Channel encoders, Indicates samples; Each of the unimodal features is input into the backbone network of the three-dimensional convolutional neural network, and the corresponding unimodal representation is output: ; in, Representation sample No. A unimodal representation of the modalities, represents the backbone network, Represents the parameters of the backbone network.

[0009] Optionally, the multimodal image segmentation method based on missing modality, wherein the initial segmentation result of the target image set is obtained, a segmentation prediction result is generated according to all the unimodal representations, and all unimodal representations of each modality data of the similarity loss term are constructed according to the initial segmentation result and the segmentation prediction result, further comprising: Obtaining missing modality information input by a user, and screening valid modalities and missing modalities from all the single-modality representations according to the missing modality information; When inputting modality data into the dynamic fusion model in the three-dimensional convolutional neural network, all the missing modalities are respectively left vacant.

[0010] Optionally, the multimodal image segmentation method based on missing modality, wherein the step of obtaining the initial segmentation result of the target image set, generating a segmentation prediction result according to all the single modality representations, and constructing a similarity loss term according to the initial segmentation result and the segmentation prediction result, specifically includes: Obtaining an initial segmentation result of the target image set; Based on all the single modal representations, multiple valid modal sets are constructed: ; in, Representation sample The effective mode set of Representation sample The unimodal representation of the first mode of Representation sample The unimodal representation of the second mode of Representation sample No. Unimodal representation of modalities; All the valid modality sets are input into the dynamic fusion model, and the segmentation prediction results of the target image set are output: ; in, represents the segmentation prediction result, represents the fusion operator of the dynamic fusion model; Construct a similarity loss term based on the initial segmentation result and the segmentation prediction result ; in, Represents the initial segmentation result.

[0011] Optionally, in the multimodal image segmentation method based on missing modality, the step of constructing a divergence loss term according to the segmentation prediction result and the initial segmentation result specifically includes: Constructing an initial probability density function of the initial segmentation result and a predicted probability density function of the segmentation prediction result; A divergence metric function is constructed according to the initial probability density function and the predicted probability density function, and a divergence loss term is determined: ; in, represents the divergence measure function, represents the difference between the initial probability density function and the predicted probability density function, and Indicates adjustment and The hyperparameters of and are conjugate indices of each other, represents the initial probability density, represents the predicted probability density, represents the initial distribution of the initial segmentation results, represents the predicted distribution of the segmentation prediction results, Indicates the range of integration; ; in, represents the divergence loss term, represents the segmentation prediction result, represents the initial segmentation result, , and Respectively represent the dimensions of the sample’s data volume in depth, height, and width, represents the mapping function, Indicated in The predicted probability at Indicated in The initial probability at Indicated in The difference between .

[0012] Optionally, the multimodal image segmentation method based on missing modality, wherein the step of extracting multiple feature vector pairs in the three-dimensional convolutional neural network, constructing the mutual information between the complete modality and the missing modality according to the multiple feature vector pairs, and processing the mutual information by a variational information maximization method to obtain a mutual information loss term, specifically includes: A plurality of feature vector pairs are extracted from all channel encoders in the three-dimensional convolutional neural network, and the mutual information between the complete modality and the missing modality is constructed based on all the feature vector pairs: ; in, represents mutual information, express The entropy of Indicates a given under conditions The conditional entropy of Represents the deep features under all modalities, represents the deep features in the missing mode, Indicates that in the known In the case specific characteristics of Calculate the mean and variance of all the unimodal representations, calculate the negative logarithm of the conditional probability of the missing modality based on the mean and the variance, and construct the mutual information loss term by minimizing the negative logarithm: ; in, represents the mutual information loss term, Indicates the total number of channel encoders, Indicates Channel encoders, represents the layer-by-layer weighting coefficient, Indicates pairs of feature vectors, Indicates the extraction of The expected value of the eigenvector pair is represents the predicted probability distribution.

[0013] Optionally, the multimodal image segmentation method based on missing modality, wherein the similarity loss term, the divergence loss term and the mutual information loss term are used to train the three-dimensional convolutional neural network to obtain a target three-dimensional convolutional neural network, and the current image set is segmented by the target three-dimensional convolutional neural network to obtain a final segmentation result, specifically includes: A total loss function is constructed according to the similarity loss term, the divergence loss term and the mutual information loss term: ; in, represents the total loss function, represents the similarity loss term, represents the mutual information loss term, represents the divergence loss term, express The weight adjustment coefficient, express The weight adjustment coefficient of Training the three-dimensional convolutional neural network using the total loss function to obtain a target three-dimensional convolutional neural network; The current image set input by the user is obtained, the current image set is input into the target three-dimensional convolutional neural network for segmentation processing, and the final segmentation result is output.

[0014] In addition, to achieve the above-mentioned object, the present invention further provides a multimodal image segmentation system based on missing modality, wherein the multimodal image segmentation system based on missing modality comprises: A feature extraction module, used to obtain a target image set, extract multiple modal data of each sample according to the target image set, input all the modal data into a three-dimensional convolutional neural network, and output all single modal representations of each modal data; A similarity loss term construction module, used to obtain an initial segmentation result of the target image set, generate a segmentation prediction result according to all the unimodal representations, and construct a similarity loss term according to the initial segmentation result and the segmentation prediction result; A divergence loss term construction module, used to construct a divergence loss term according to the segmentation prediction result and the initial segmentation result; A mutual information loss term construction module, used to extract multiple feature vector pairs in the three-dimensional convolutional neural network, construct the mutual information between the complete mode and the missing mode according to the multiple feature vector pairs, and process the mutual information through a variational information maximization method to obtain a mutual information loss term; The model training module is used to train the three-dimensional convolutional neural network using the similarity loss term, the divergence loss term and the mutual information loss term to obtain a target three-dimensional convolutional neural network, and to segment the current image set through the target three-dimensional convolutional neural network to obtain a final segmentation result.

[0015] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a multimodal image segmentation program based on missing modality stored in the memory and executable on the processor, wherein the multimodal image segmentation program based on missing modality, when executed by the processor, implements the steps of the multimodal image segmentation method based on missing modality as described above.

[0016] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multimodal image segmentation program based on a missing modality, and when the multimodal image segmentation program based on a missing modality is executed by a processor, the steps of the multimodal image segmentation method based on a missing modality as described above are implemented.

[0017] In the present invention, a target image set is obtained, multiple modal data of each sample are extracted according to the target image set, all the modal data are input into a three-dimensional convolutional neural network, and all single-modal representations of each modal data are output; an initial segmentation result of the target image set is obtained, a segmentation prediction result is generated according to all the single-modal representations, and a similarity loss term is constructed according to the initial segmentation result and the segmentation prediction result; a divergence loss term is constructed according to the segmentation prediction result and the initial segmentation result; a plurality of feature vector pairs are extracted from the three-dimensional convolutional neural network, mutual information between complete modalities and missing modalities is constructed according to the plurality of feature vector pairs, and the mutual information is processed by a variational information maximization method to obtain a mutual information loss term; the three-dimensional convolutional neural network is trained using the similarity loss term, the divergence loss term and the mutual information loss term to obtain a target three-dimensional convolutional neural network, and the current image set is segmented by the target three-dimensional convolutional neural network to obtain a final segmentation result. The present invention comprehensively considers the situation of incomplete multimodal data. Through divergence loss and high mutual information knowledge transfer methods, it effectively promotes the feature alignment between different modalities and improves the accuracy of cross-modal information transfer. The difference between the prediction result and the label is effectively measured, and the missing modality can be handled more effectively, which improves the stability and accuracy of the model in the case of incomplete data. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flow chart of a preferred embodiment of the multimodal image segmentation method based on missing modality of the present invention; Figure 2 It is a segmentation framework diagram of a preferred embodiment of the multimodal image segmentation method based on missing modality of the present invention; Figure 3 is a comparison diagram of segmentation results of a preferred embodiment of the multimodal image segmentation method based on missing modality of the present invention; Figure 4 It is a structural diagram of a preferred embodiment of the multimodal image segmentation system based on missing modality of the present invention; Figure 5 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0020] The multimodal image segmentation method based on missing modality described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the multimodal image segmentation method based on missing modality includes the following steps: Step S10: obtain a target image set, extract multiple modal data of each sample according to the target image set, input all the modal data into a three-dimensional convolutional neural network, and output all single modal representations of each modal data.

[0021] Among them, Figure 2 As shown, by converting the acquired target image set (e.g. Figure 2 In , ... Representing samples in the target image set) are input into a three-dimensional convolutional neural network to encode different modal data in parallel. For example, the encoder processes multiple modal data in parallel to obtain deep features of different modalities (such as Figure 2 In , ... The decoder is then used to obtain the corresponding single-modal representation to retain the feature information of each modality.

[0022] Furthermore, after decoding the features, the decoder obtains the corresponding valid features and generates the predicted labels ( Figure 2 In and the true label ( Figure 2 In ), which are used to construct the subsequent similarity loss term and divergence loss term respectively.

[0023] Specifically, a target image set is obtained, and multiple modal data of each sample in the target image set are extracted; all the modal data of each sample are input into the constructed three-dimensional convolutional neural network, and multiple single-modal features of multiple modal data are output respectively through multiple channel encoders in the three-dimensional convolutional neural network: ; in, Representation sample The The unimodal features of the modes, Representation sample The Channel encoders, Indicates samples; input each of the unimodal features into the backbone network of the three-dimensional convolutional neural network, and output the corresponding unimodal representation: ; in, Representation sample No. A unimodal representation of the modalities, represents the backbone network, Represents the parameters of the backbone network.

[0024] Among them, the multiple modal data in each sample are first input into different channel encoders to realize parallel encoding and obtain the corresponding single modal features. This process mainly involves preliminary feature conversion and dimensionality reduction of the modal data, which can prepare for the subsequent shared backbone network to reduce the accuracy loss of the entire process.

[0025] Among them, this embodiment adopts a parallel network processing framework, which can independently activate the corresponding feature channels for each available modality, significantly reducing the computational burden and memory requirements. Compared with the traditional model that relies on full-modal input, this method can still run efficiently when only some modalities are available, which not only optimizes the utilization of computing resources, but also expands its applicability in resource-constrained environments.

[0026] Furthermore, the missing modality information input by the user is obtained, and the valid modalities and missing modalities are screened from all the single modality representations according to the missing modality information; when the modality data is input into the dynamic fusion model in the three-dimensional convolutional neural network, all the missing modalities are respectively left vacant.

[0027] Among them, all modal data of each sample are sequentially input into different channel encoders, and multiple single-modal representations of each sample can be obtained through multiple parallel processing channel encoders. In the subsequent fusion stage of the single-modal representation input into the dynamic fusion model, if there is missing modal data, the corresponding path will be flexibly "skipped" in the dynamic fusion model, and the corresponding channel will be treated as vacant, and only the available modality will be used for segmentation. Through independent processing and optimization of each modal feature, the model can effectively reduce the computational complexity and memory usage of the reasoning process while maintaining high performance. This feature makes the present invention practical in real-time or high-load medical environments, and can complete accurate brain tumor segmentation in a relatively short time.

[0028] Step S20: obtaining an initial segmentation result of the target image set, generating a segmentation prediction result according to all the unimodal representations, and constructing a similarity loss term according to the initial segmentation result and the segmentation prediction result.

[0029] Among them, if there is a sample with missing modal data, the corresponding unimodal representation cannot be generated. Therefore, in this embodiment, the path of the missing modality is automatically judged and skipped in the fusion stage, and only the unimodal representations of the remaining available modalities are fused.

[0030] Specifically, an initial segmentation result of the target image set is obtained; and multiple valid modality sets are constructed according to all the single modality representations: ; in, Representation sample The effective mode set of Representation sample The unimodal representation of the first mode of Representation sample The unimodal representation of the second mode of Representation sample No. All the valid modality sets are input into the dynamic fusion model, and the segmentation prediction results of the target image set are output: ; in, represents the segmentation prediction result, A fusion operator representing a dynamic fusion model; constructing a similarity loss term based on the initial segmentation result and the segmentation prediction result ; in, Represents the initial segmentation result.

[0031] Among them, according to the constructed single modality representation, all valid modalities in the target image set can be determined, thereby constructing a valid modality set. The dynamic combination (DC) module in the three-dimensional convolutional neural network is used to process the valid modality set, and the segmentation prediction results of the target image set are generated through the learnable or preset fusion operator of the DC module. Accordingly, in the stage of model reasoning, no matter how many modalities are missing in the target image set, the model can adaptively infer based on the remaining modalities (valid modality set), so that the model can still work effectively when some modalities are missing, which improves the applicability and enhances the robustness of the model in actual image segmentation application scenarios.

[0032] Furthermore, the original segmentation results (i.e., initial segmentation results) of the target image set are obtained, and the similarity between the original segmentation results and the segmentation prediction results is constructed to construct a similarity loss term. This can ensure the consistency between the predicted segmentation of the target image and the true label at the pixel or voxel level, which is suitable for scenarios with strongly unbalanced image distribution.

[0033] Step S30: construct a divergence loss term according to the segmentation prediction result and the initial segmentation result.

[0034] Among them, traditional divergences (for example, KL divergence (Kullback-Leibler divergence) and Jensen-Shannon divergence (‌Jensen-Shannon divergence‌)) are prone to instability when dealing with asymmetric distribution or noisy data. The Hölder divergence used in this embodiment has better stability in dealing with noise and asymmetric distribution, and can significantly enhance the robustness and accuracy of the model.

[0035] Specifically, an initial probability density function of the initial segmentation result and a predicted probability density function of the segmentation prediction result are constructed; a divergence metric function is constructed according to the initial probability density function and the predicted probability density function, and a divergence loss term is determined: ; in, represents the divergence measure function, represents the difference between the initial probability density function and the predicted probability density function, and Indicates adjustment and The hyperparameters of and are conjugate indices of each other, represents the initial probability density, represents the predicted probability density, represents the initial distribution of the initial segmentation results, represents the predicted distribution of the segmentation prediction results, Indicates the range of integration; ; in, represents the divergence loss term, represents the segmentation prediction result, represents the initial segmentation result, , and Respectively represent the dimensions of the sample’s data volume in depth, height, and width, represents the mapping function, Indicated in The predicted probability at Indicated in The initial probability at Indicated in The difference between .

[0036] There are different conjugate exponents in the divergence metric function of this embodiment, and the outliers and asymmetric distribution can be finely regulated by taking different values. and The product of is greater than 0.

[0037] Furthermore, in addition to this embodiment, for different segmentation scenarios, a piecewise Hölder divergence can also be used (for example, different value) or combining the Hölder divergence with Dice / Cross-Entropy etc. into a hybrid loss to achieve the same effect as in this embodiment.

[0038] Furthermore, in order to measure the difference between the predicted probability and the true label, a divergence loss term can be defined to deal with asymmetric distribution and noise, thereby ensuring that the model is more robust in dealing with severe imbalance or noise that may exist in the segmentation annotation data.

[0039] Step S40: extract multiple feature vector pairs in the three-dimensional convolutional neural network, construct the mutual information between the complete mode and the missing mode according to the multiple feature vector pairs, and process the mutual information through the variational information maximization method to obtain the mutual information loss term.

[0040] Among them, in the image of the missing modality, it is difficult for the model to obtain complete multimodal features, resulting in incomplete information and limited generalization ability. Therefore, it is necessary to use the complete modality (or a richer set of modalities) to perform knowledge distillation on the missing modality so that the missing modality can also obtain certain cross-modal information compensation. When some modalities are missing, the available modalities can help the network learn important feature associations through mutual information measurement, thereby reducing the sensitivity to missing information.

[0041] Specifically, multiple feature vector pairs are extracted from all channel encoders in the three-dimensional convolutional neural network, and the mutual information between the complete modality and the missing modality is constructed based on all the feature vector pairs: ; in, represents mutual information, express The entropy of Indicates a given under conditions The conditional entropy of Represents the deep features under all modalities, represents the deep features in the missing mode, Indicates that in the known In the case Specific features; calculating the mean and variance of all the single modal representations, calculating the negative logarithm of the conditional probability of the missing modality based on the mean and the variance, and constructing the mutual information loss term by minimizing the negative logarithm: ; in, represents the mutual information loss term, Indicates the total number of channel encoders, Indicates Channel encoders, Represents the layer-by-layer weighted coefficient, emphasizing the key role of high-level semantic features in knowledge distillation. Indicates pairs of feature vectors, Indicates the extraction of The expected value of the eigenvector pair is represents the predicted probability distribution.

[0042] Among them, in order to measure the feature correlation between the complete modality and the missing modality, in this embodiment, feature vector pairs are extracted from each channel encoder of the three-dimensional convolutional neural network, and the mutual information between the complete modality and the missing modality is estimated by calculating entropy and conditional entropy, thereby guiding the information transmission between network levels; the mutual information is processed by the variational information maximization method, and the variational information maximization estimation is performed directly within the network, which can effectively transfer the information of the complete modality to the missing modality path, significantly improving the model accuracy.

[0043] Furthermore, in addition to this embodiment, the features of the complete mode and the missing mode can be compared and lost to achieve the purpose of "similar feature distribution" to replace the step of "mutual information distillation".

[0044] Among them, through the knowledge distillation mechanism proposed in this embodiment, there is no need to pre-complete or interpolate the missing modalities, thereby reducing the additional model complexity; by introducing the Hölder divergence and mutual information knowledge transfer based on modality missing, the problem of incomplete multimodal data is effectively solved, and the accuracy and robustness of image segmentation (especially brain tumor segmentation) are improved, which not only improves the segmentation accuracy, but also enhances the adaptability of the model, which is of great significance to the common clinical modality data missing problem.

[0045] Step S50: Use the similarity loss term, the divergence loss term and the mutual information loss term to train the three-dimensional convolutional neural network to obtain a target three-dimensional convolutional neural network, and segment the current image set through the target three-dimensional convolutional neural network to obtain a final segmentation result.

[0046] Among them, integrating the constructed losses into a unified total loss framework and combining them with GPU (Graphics Processing Unit) acceleration can achieve considerable training efficiency.

[0047] Specifically, a total loss function is constructed according to the similarity loss term, the divergence loss term and the mutual information loss term: ; in, represents the total loss function, represents the similarity loss term, represents the mutual information loss term, represents the divergence loss term, express The weight adjustment coefficient, express The weight adjustment coefficient is as follows; the three-dimensional convolutional neural network is trained by the total loss function to obtain a target three-dimensional convolutional neural network; a current image set input by a user is obtained, the current image set is input into the target three-dimensional convolutional neural network for segmentation processing, and a final segmentation result is output.

[0048] Among them, after constructing the total loss function, the parameters of the shared backbone network, channel encoder and DC module are updated by gradient descent or other optimization methods, and the training is stopped according to the performance of the validation set or the number of training rounds to obtain a more robust target three-dimensional convolutional neural network. The target image set is segmented by the target three-dimensional convolutional neural network to obtain a final segmentation result with higher accuracy.

[0049] Furthermore, if Figure 3 As shown, the segmentation results of four models after using different modalities to input the BraTS 2018 dataset (BrainTumor Segmentation Challenge 2018) are shown in the second row; the reproduction result of the second model (MMCFormer, Missing Modality Compensation Transformer) is shown in the third row; the reproduction result of the third modality (MML-MM-SSF, Multimodal Machine Learning - Multimodal SemanticFusion Model) is shown in the fourth row; the reproduction result of the fourth model, that is, the model used in this embodiment, is shown in the fifth row; Figure 3Each column in represents a different input. The first four columns show the results of different single-modal inputs, the fifth column shows the results of using all four modalities as input at the same time, and the last column shows the corresponding true labels.

[0050] Furthermore, the following table can clearly show the accuracy of the segmentation results output by the method used in this application compared with the existing conventional methods: Table 1: WT (Whole Tumor, whole tumor area) segmentation result evaluation table

[0051] Table 2: ET (Enhancing Tumor) segmentation result evaluation table

[0052] Table 3: TC (Tumor Core) segmentation result evaluation table

[0053] Among them, the horizontal elements in Tables 1, 2 and 3 represent different single modalities, combinations of different single modalities and full modal inputs. The last column represents the mean of all modal segmentation results. In the first column, multiple methods (models) are divided. U-HVED represents Unruh-Hybrid Variable Energy Detection Model, ACN represents Automatic Control Network, RFNet represents Region-aware Fusion Network, SMU-Net represents Style Matching U-Net, D2-Net represents Dense Detection and Description Network, mmFormer represents a deep learning network designed for medical image segmentation tasks (especially brain tumor segmentation), EMRM represents Enhanced Mobile Radio Module, and SSFM represents Multi-modal Learning with Missing Modality. via Shared-Specific Feature Modelling), QuMo stands for Quantum Mobility, GSS stands for Scratch Each Other's Back: Incomplete Multi-modal Brain Tumor Segmentation Via Category Aware Group Self-Support Learning, MFTrans stands for Multi-Function Transformer Network based on the Transformer architecture, and MTI stands for Multimodal Transformer of Incomplete MRI Data for Brain Tumor Segmentation.GGDM stands for Gradient-Guided Modality Decoupling for Missing-Modality Robustness, and OUR stands for the parallel-based three-dimensional convolutional neural network used in the present invention.

[0054] Furthermore, Tables 1, 2 and 3 all show the quantitative results of segmentation performance measured based on the Dice Similarity Coefficient (DSC) on the BraTS 2018 dataset. By comparing the DSC values ​​of different segmentation methods, their overall effect can be evaluated. The higher the DSC value, the better the segmentation accuracy. The method used in this application has a higher segmentation accuracy than the segmentation results obtained by other models.

[0055] Specifically, combined with the Hölder divergence hyperparameter ( ) and mutual information knowledge transfer strategy, the present invention effectively enhances the performance of the model under different modal data. In a single modality input scenario (taking the Brats 2018 dataset as an example), the segmentation performance can be improved by 1.4% to 15.3% compared with the existing multimodal segmentation technology. Therefore, even in the case of missing modalities or incomplete information, the model can still maintain high accuracy and robustness, and has broad clinical application potential.

[0056] Compared with conventional KL divergence and other measurement methods, the present invention further enhances the model's adaptability to asymmetric distribution by introducing Hölder divergence. On the Brats2018 dataset, the Dice coefficient increased by 6.1% on average. This feature overcomes the limitation of traditional methods in handling abnormal or unbalanced data to a certain extent, and provides higher stability and accuracy for dealing with complex clinical data.

[0057] The present invention comprehensively considers the situation of incomplete multimodal data. Through divergence loss and high mutual information knowledge transfer methods, it effectively promotes the feature alignment between different modalities and improves the accuracy of cross-modal information transfer. The difference between the prediction result and the label is effectively measured, and the missing modality can be handled more effectively, which improves the stability and accuracy of the model in the case of incomplete data.

[0058] Furthermore, if Figure 4 As shown, based on the above-mentioned multimodal image segmentation method based on missing modality, the present invention also provides a multimodal image segmentation system based on missing modality, wherein the multimodal image segmentation system based on missing modality includes: A feature extraction module 51 is used to obtain a target image set, extract multiple modal data of each sample according to the target image set, input all the modal data into a three-dimensional convolutional neural network, and output all single modal representations of each modal data; A similarity loss term construction module 52 is used to obtain an initial segmentation result of the target image set, generate a segmentation prediction result according to all the unimodal representations, and construct a similarity loss term according to the initial segmentation result and the segmentation prediction result; A divergence loss term construction module 53, used to construct a divergence loss term according to the segmentation prediction result and the initial segmentation result; A mutual information loss term construction module 54 is used to extract multiple feature vector pairs in the three-dimensional convolutional neural network, construct the mutual information between the complete mode and the missing mode according to the multiple feature vector pairs, and process the mutual information by a variational information maximization method to obtain a mutual information loss term; The model training module 55 is used to train the three-dimensional convolutional neural network using the similarity loss term, the divergence loss term and the mutual information loss term to obtain a target three-dimensional convolutional neural network, and to segment the current image set through the target three-dimensional convolutional neural network to obtain a final segmentation result.

[0059] Furthermore, if Figure 5 As shown, based on the above-mentioned missing modality-based multimodal image segmentation method and system, the present invention also provides a terminal accordingly, and the terminal includes a processor 10, a memory 20 and a display 30. Figure 5 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0060] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Further, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a multimodal image segmentation program 40 based on missing modality is stored on the memory 20, and the multimodal image segmentation program 40 based on missing modality can be executed by the processor 10, thereby realizing the multimodal image segmentation method based on missing modality in the present application.

[0061] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the multimodal image segmentation method based on missing modality.

[0062] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.

[0063] In one embodiment, when the processor 10 executes the missing modality-based multimodal image segmentation program 40 in the memory 20 , the steps of the missing modality-based multimodal image segmentation method described above are implemented.

[0064] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multimodal image segmentation program based on a missing modality, and when the multimodal image segmentation program based on a missing modality is executed by a processor, the steps of the multimodal image segmentation method based on a missing modality as described above are implemented.

[0065] In summary, the present invention provides a multimodal image segmentation method based on missing modality and related equipment, the method comprising: obtaining a target image set, extracting multiple modal data of each sample according to the target image set, inputting all the modal data into a three-dimensional convolutional neural network, and outputting all single-modal representations of each modal data; obtaining an initial segmentation result of the target image set, generating a segmentation prediction result according to all the single-modal representations, and constructing a similarity loss term according to the initial segmentation result and the segmentation prediction result; constructing a divergence loss term according to the segmentation prediction result and the initial segmentation result; extracting multiple feature vector pairs in the three-dimensional convolutional neural network, constructing the mutual information between the complete modality and the missing modality according to the multiple feature vector pairs, and processing the mutual information through a variational information maximization method to obtain a mutual information loss term; using the similarity loss term, the divergence loss term and the mutual information loss term to train the three-dimensional convolutional neural network to obtain a target three-dimensional convolutional neural network, and segmenting the current image set through the target three-dimensional convolutional neural network to obtain a final segmentation result. The present invention comprehensively considers the situation of incomplete multimodal data. Through divergence loss and high mutual information knowledge transfer methods, it effectively promotes the feature alignment between different modalities and improves the accuracy of cross-modal information transfer. The difference between the prediction result and the label is effectively measured, and the missing modality can be handled more effectively, which improves the stability and accuracy of the model in the case of incomplete data.

[0066] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or terminal including the element.

[0067] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.

[0068] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A multimodal image segmentation method based on missing modality, characterized in that: The multimodal image segmentation method based on missing modality includes: Acquire a target image set, extract multiple modal data of each sample according to the target image set, input all the modal data into a three-dimensional convolutional neural network, and output all single modal representations of each modal data; Obtaining an initial segmentation result of the target image set, generating a segmentation prediction result according to all the unimodal representations, and constructing a similarity loss term according to the initial segmentation result and the segmentation prediction result; Constructing a divergence loss term according to the segmentation prediction result and the initial segmentation result; Extracting multiple feature vector pairs in the three-dimensional convolutional neural network, constructing mutual information between the complete mode and the missing mode according to the multiple feature vector pairs, and processing the mutual information by a variational information maximization method to obtain a mutual information loss term; The three-dimensional convolutional neural network is trained using the similarity loss term, the divergence loss term and the mutual information loss term to obtain a target three-dimensional convolutional neural network, and the current image set is segmented by the target three-dimensional convolutional neural network to obtain a final segmentation result.

2. The multimodal image segmentation method based on missing modality according to claim 1, characterized in that: The step of acquiring a target image set, extracting multiple modal data of each sample according to the target image set, inputting all the modal data into a three-dimensional convolutional neural network, and outputting all single modal representations of each modal data specifically includes: Acquire a target image set, and extract multiple modality data of each sample in the target image set; Input all the modal data of each sample into the constructed three-dimensional convolutional neural network, and output multiple single-modal features of multiple modal data respectively through multiple channel encoders in the three-dimensional convolutional neural network: ; in, Representation sample The The unimodal features of the modes, Representation sample The Channel encoders, Indicates samples; Each of the unimodal features is input into the backbone network of the three-dimensional convolutional neural network, and the corresponding unimodal representation is output: ; in, Representation sample No. A unimodal representation of the modalities, represents the backbone network, Represents the parameters of the backbone network.

3. The multimodal image segmentation method based on missing modality according to claim 2, characterized in that: The step of obtaining the initial segmentation result of the target image set, generating a segmentation prediction result according to all the unimodal representations, and constructing all unimodal representations of each modality data of the similarity loss term according to the initial segmentation result and the segmentation prediction result, further includes: Obtaining missing modality information input by a user, and screening valid modalities and missing modalities from all the single-modality representations according to the missing modality information; When inputting modality data into the dynamic fusion model in the three-dimensional convolutional neural network, all the missing modalities are respectively left vacant.

4. The multimodal image segmentation method based on missing modality according to claim 1, characterized in that: The obtaining of the initial segmentation result of the target image set, generating a segmentation prediction result according to all the unimodal representations, and constructing a similarity loss term according to the initial segmentation result and the segmentation prediction result specifically includes: Obtaining an initial segmentation result of the target image set; Based on all the single modal representations, multiple valid modal sets are constructed: ; in, Representation sample The effective mode set of Representation sample The unimodal representation of the first mode of Representation sample The unimodal representation of the second mode of Representation sample No. Unimodal representation of modalities; All the valid modality sets are input into the dynamic fusion model, and the segmentation prediction results of the target image set are output: ; in, represents the segmentation prediction result, represents the fusion operator of the dynamic fusion model; Construct a similarity loss term based on the initial segmentation result and the segmentation prediction result ; in, Represents the initial segmentation result.

5. The multimodal image segmentation method based on missing modality according to claim 1, characterized in that: The step of constructing a divergence loss term according to the segmentation prediction result and the initial segmentation result specifically includes: Constructing an initial probability density function of the initial segmentation result and a predicted probability density function of the segmentation prediction result; A divergence metric function is constructed according to the initial probability density function and the predicted probability density function, and a divergence loss term is determined: ; in, represents the divergence measure function, represents the difference between the initial probability density function and the predicted probability density function, and Indicates adjustment and The hyperparameters of and are conjugate indices of each other, represents the initial probability density, represents the predicted probability density, represents the initial distribution of the initial segmentation results, represents the predicted distribution of the segmentation prediction results, Indicates the range of integration; ; in, represents the divergence loss term, represents the segmentation prediction result, represents the initial segmentation result, , and Respectively represent the dimensions of the sample’s data volume in depth, height, and width, represents the mapping function, Indicated in The predicted probability at Indicated in The initial probability at Indicated in The difference between .

6. The missing modality-based multimodal image segmentation method according to claim 1, characterized in that: The step of extracting multiple feature vector pairs from the three-dimensional convolutional neural network, constructing mutual information between the complete mode and the missing mode according to the multiple feature vector pairs, and processing the mutual information by a variational information maximization method to obtain a mutual information loss term specifically includes: A plurality of feature vector pairs are extracted from all channel encoders in the three-dimensional convolutional neural network, and the mutual information between the complete modality and the missing modality is constructed based on all the feature vector pairs: ; in, represents mutual information, express The entropy of Indicates a given under conditions The conditional entropy of Represents the deep features under all modalities, represents the deep features in the missing mode, Indicates that in the known In the case specific characteristics of Calculate the mean and variance of all the unimodal representations, calculate the negative logarithm of the conditional probability of the missing modality based on the mean and the variance, and construct the mutual information loss term by minimizing the negative logarithm: ; in, represents the mutual information loss term, Indicates the total number of channel encoders, Indicates Channel encoders, represents the layer-by-layer weighting coefficient, Indicates pairs of feature vectors, Indicates the extraction of The expected value of the eigenvector pair is represents the predicted probability distribution.

7. The multimodal image segmentation method based on missing modality according to claim 1, characterized in that: The method of training the three-dimensional convolutional neural network using the similarity loss term, the divergence loss term, and the mutual information loss term to obtain a target three-dimensional convolutional neural network, and segmenting the current image set by using the target three-dimensional convolutional neural network to obtain a final segmentation result specifically includes: A total loss function is constructed according to the similarity loss term, the divergence loss term and the mutual information loss term: ; in, represents the total loss function, represents the similarity loss term, represents the mutual information loss term, represents the divergence loss term, express The weight adjustment coefficient, express The weight adjustment coefficient of Training the three-dimensional convolutional neural network using the total loss function to obtain a target three-dimensional convolutional neural network; The current image set input by the user is obtained, the current image set is input into the target three-dimensional convolutional neural network for segmentation processing, and the final segmentation result is output.

8. A multimodal image segmentation system based on missing modality, characterized in that: The multimodal image segmentation system based on missing modality is applied to the multimodal image segmentation method based on missing modality according to any one of claims 1 to 7, and the multimodal image segmentation system based on missing modality includes: A feature extraction module, used to obtain a target image set, extract multiple modal data of each sample according to the target image set, input all the modal data into a three-dimensional convolutional neural network, and output all single modal representations of each modal data; A similarity loss term construction module, used to obtain an initial segmentation result of the target image set, generate a segmentation prediction result according to all the unimodal representations, and construct a similarity loss term according to the initial segmentation result and the segmentation prediction result; A divergence loss term construction module, used to construct a divergence loss term according to the segmentation prediction result and the initial segmentation result; A mutual information loss term construction module, used to extract multiple feature vector pairs in the three-dimensional convolutional neural network, construct the mutual information between the complete mode and the missing mode according to the multiple feature vector pairs, and process the mutual information through a variational information maximization method to obtain a mutual information loss term; The model training module is used to train the three-dimensional convolutional neural network using the similarity loss term, the divergence loss term and the mutual information loss term to obtain a target three-dimensional convolutional neural network, and to segment the current image set through the target three-dimensional convolutional neural network to obtain a final segmentation result.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a multimodal image segmentation program based on missing modality stored in the memory and executable on the processor. When the multimodal image segmentation program based on missing modality is executed by the processor, the steps of the multimodal image segmentation method based on missing modality as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a multimodal image segmentation program based on missing modality, and when the multimodal image segmentation program based on missing modality is executed by a processor, the steps of the multimodal image segmentation method based on missing modality as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Missing modal data processing method based on mutual information constraint and adversarial learning

    CN119415848A

Cited By

  • Multimodal emotion recognition method and system based on hypergraph diffusion and evidence fusion, terminal and storage medium

    CN121009512A

  • Multimodal emotion recognition method and system based on hypergraph diffusion and evidence fusion, terminal and storage medium

    CN121009512B

  • Image classification method and device, electronic equipment and storage medium

    CN121962717A