Hybrid uncertainty MRI semi-supervised segmentation method based on context information
By introducing a context extraction module and an uncertainty hybrid module in the semi-supervised medical image segmentation method, combining 3D CNN and Mean-Teacher architecture, the shortcomings of existing methods in image volume context learning are solved, achieving higher segmentation accuracy and lower risk of overfitting.
Patent Information
- Application Number
- CN202411947687.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
The existing semi-supervised medical image segmentation method has shortcomings in image volume context learning, resulting in the model underestimating the context details in the dataset during training, thereby reducing the accuracy of segmentation of complex regions.
A hybrid uncertainty MRI semi-supervised segmentation method based on context information is designed. By introducing a context extraction module and an uncertainty hybrid module into the network architecture, using 3D CNN and Mean-Teacher architectures, the volume context information of the image is learned, and information fusion is combined with local and global uncertainties to improve the segmentation accuracy of the model.
This method can effectively learn the volume context information of the image, improve the accuracy of the segmentation of the model for the lesion area, perform better than other semi-supervised segmentation methods, and significantly reduce the model parameters and reduce the possibility of overfitting.
Smart Images

Figure CN119992081A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision of artificial intelligence, and in particular to a hybrid uncertainty MRI semi-supervised segmentation method based on context information. Background Art
[0002] Medical image segmentation plays an important role in medical diagnosis and treatment. The goal of medical image segmentation is to accurately measure the patient's lesion area based on medical images. It effectively promotes doctors' treatment plans by providing quantitative analysis of relevant lesion volumes, positioning of pathologically modified tissues, and visualization of anatomical structures. At present, medical image segmentation is mainly divided into: deep learning supervised methods, deep learning semi-supervised methods, and deep learning unsupervised methods. These different methods can be distinguished by different categories, different technologies, and whether or not to use label standards.
[0003] Semi-supervised segmentation is currently a very hot and important topic in the field of image processing. This technology aims to use a small number of labeled images and a large number of unlabeled images for training, which greatly reduces the model's demand for labeled data. The basic idea of style transfer is to generate corresponding labels for unlabeled images after learning labeled data so that they can match the corresponding reference images. In the early research on semi-supervised medical image segmentation, neural network-based methods gradually dominated. By using deep neural networks trained with labeled data, high-level feature representations of images can be captured, thereby more accurately helping unlabeled images generate corresponding labels. A key component of this method is the accuracy of label generation. Only correctly generated labels can effectively help model training.
[0004] In recent years, deep learning has developed rapidly and has shown strong performance in the field of medical image segmentation. With the ability of deep learning models to perform automatic feature learning, boundary operations and efficient performance, it is possible to obtain high-precision measurement results while taking into account both time and space efficiency. Since convolutional neural networks (CNNs) have powerful feature extraction capabilities, they can obtain rich pixel information from medical images, which can then be used for accurate lesion area segmentation. Therefore, a large number of medical image segmentation methods based on deep learning have sprung up, such as: V-Net based on 3D CNN, TransUNet combining Transformer and U-Net, AC-MT based on consistency learning, and UA-MT based on Mean-Teacher structure. Although these methods have achieved good results to a certain extent, they are all supervised methods and require a large number of medical lesion images corresponding to real labels. However, due to the low contrast and noise interference of medical images, they usually show poor visual effects. This makes it very laborious and difficult to annotate such pixel-level annotations. In addition, medical image annotation requires more expertise and time than natural images, so it is almost impossible to construct a large number of medical image datasets with high-precision labels, which greatly slows down the development of medical image segmentation. In comparison, unlabeled data is easier to obtain. Therefore, semi-supervised medical image segmentation methods that do not require a large number of real labels but only require a small amount of labeled data and a large amount of unlabeled data have emerged. In 2017, Tarvainen et al. first utilized a pseudo-label-based student-teacher semi-supervised learning paradigm, making semi-supervised methods an emerging wave in medical image segmentation. The current main semi-supervised methods include consistency learning, adversarial learning, contrastive learning, and pseudo-labeling. Although these methods have greatly promoted the development of current semi-supervised segmentation and performed well in segmentation with a small amount of labeled data. However, the current semi-supervised segmentation still has some problems.
[0005] (1) Existing uncertainty regularization focuses on seeking more complex data-level or model-level perturbations for unsupervised consistency learning to improve the training effect of the model, thereby ignoring the learning of image volume context, causing the model to underestimate the context details in the dataset during training, resulting in a decrease in the accuracy of the model's segmentation of these complex areas.
[0006] Existing uncertainty calculations mainly estimate image uncertainty at the pixel level. They do not consider the global information of the image, causing the model to easily ignore the contextual and structural information of the image. This will lead to the model's inaccurate understanding of the shape, size, and spatial relationship of the image, thereby reducing the model's accuracy in segmenting the lesion area.
[0007] Therefore, a new solution to the above problems needs to be proposed. Summary of the invention
[0008] The purpose of the present invention is to provide a hybrid uncertainty MRI semi-supervised segmentation method based on contextual information, design and train a model through a deep learning method to achieve fast and accurate lesion segmentation, so as to solve the technical problems raised in the background technology.
[0009] To achieve the above object, the present invention provides the following technical solution: a hybrid uncertainty MRI semi-supervised segmentation method based on context information, comprising at least the following steps:
[0010] S1: Preprocess the medical images in the public dataset;
[0011] S2: Divide the original data set into two parts: labeled images and unlabeled images, and divide the ratio;
[0012] S3: Design a network architecture for learning context information, i.e. build a context extraction module;
[0013] S4: Design a loss function for mixed uncertainty calculation, that is, build an uncertainty mixing module;
[0014] S5: Build the CUAMT model based on the context extraction module and the uncertainty mixing module, use the Mean-Teacher architecture to train the model and perform measurement verification, and the model training is completed;
[0015] S6: Use the trained model for supervised segmentation.
[0016] Furthermore, the S1 at least comprises the following steps:
[0017] All volumes of the original medical images were normalized to zero mean and unit variance;
[0018] Each 3D MRI volume is then cropped with its enlarged edges according to the target;
[0019] Each 3D image is cropped to a size of 128×128×80.
[0020] Furthermore, the S2 at least includes the following steps:
[0021] 80 images of the original data set are used for training, 20 images are used for validation, and 54 images are used for testing;
[0022] Among the 80 medical images used for training, only 10% of the images have labels, and the labels of the remaining images are removed.
[0023] Furthermore, the design idea of the context extraction module in S3 is:
[0024] In order to obtain semantic feature information of different dimensions in medical images, a context information learning network is designed based on 3D CNN;
[0025] By extracting features from images at different scales and upsampling them to the same scale for information fusion, the model can better learn the volume context information of the image, making the model training more accurate and robust.
[0026] Furthermore, the context extraction module includes two parts: feature extraction and 3D attention mechanism;
[0027] The feature extraction part uses a hybrid pyramid structure to extract context features;
[0028] In the 3D attention mechanism section, the CA attention mechanism is improved to a 3D attention mechanism.
[0029] In the 3D attention mechanism section, we improve the CA attention mechanism into a 3D attention mechanism.
[0030] Furthermore, the execution process of the feature extraction part includes at least the following steps:
[0031] First, the input image passes through a 1x1x1 3D convolution operation to reduce the number of computational parameters and channel dimensions of the feature map;
[0032] Next, it is divided into two parts;
[0033] The first part uses 3x3x3 3D convolution to extract the original features of the image, and the original features include at least texture, shape and edge;
[0034] The second part uses 3D maximum pooling with strides of 2, 4, and 6 to perform pooling operations with different kernel sizes, followed by a 3x3x3 3D convolution operation to extract more advanced features and achieve the extraction of features of different scales;
[0035] After that, the above three convolution operations are followed by an upsampling operation to restore the pooled feature map to the same input size as the original image for subsequent feature fusion to obtain a semantically rich feature map, thereby enhancing the network’s understanding of the spatial structure and contextual information of the input data.
[0036] Furthermore, the execution process of the 3D attention mechanism part includes at least the following steps:
[0037] First, we replace the convolution and pooling operations inside the network with 3D convolution and 3D pooling, and add BatchNorm3d as a normalization operation, thereby improving it from a 2D attention mechanism to a 3D attention mechanism.
[0038] The 3D attention mechanism combines the channel information and position information of the input data. The channel information provides rich feature expressions, while the position information provides contextual information between different elements. The feature expressions include at least texture, shape and edge.
[0039] The combination of channel information and position information enables the model to better understand the semantic features and contextual information of the input image, thereby improving the performance and expression ability of the model.
[0040] Furthermore, the uncertainty mixing module adopts the existing uncertainty calculation method in the calculation of local uncertainty, and calculates the local uncertainty of each sample by performing multiple forward propagations during the training process and calculating the average value of multiple prediction results. The formula is as follows:
[0041]
[0042] Where T represents the number of predictions, Represents the prediction result of input x using weights at time step t, P c is the average prediction result, and H(p)1 represents our local uncertainty.
[0043] In the global uncertainty calculation, a denoising autoencoder (DAE) is introduced to encode the available segmentation masks in a nonlinear latent space to learn anatomically aware prior representations.
[0044] This learned representation can capture the global information from the segmentation mask, thereby mapping inaccurate predictions to credible segmentations to learn the global information captured from the segmentation mask, thereby calculating a global uncertainty, which is calculated as follows:
[0045] H(p)2=∥f d (x i )-f t (x i )∥ 2
[0046] where f d (x i ) is the DAE for the input x i The prediction result, f t (x i ) is the teacher model for input x iThe prediction results, H(p)2 represents our global uncertainty.
[0047] Then, the calculated local uncertainty and global uncertainty are combined to obtain the final mixed uncertainty, as shown in the following formula:
[0048] C1=1-H(p)1
[0049]
[0050] Cmix=αC1+(1-αC2)
[0051] Among them, C1 and C2 are reliable regions calculated by the information compression method in information theory using local and global uncertainties, α is an adjustable weight control factor set according to experience, e is a mathematical constant, which is approximately equal to 2.71828, and Cmix is the mixed uncertainty.
[0052] This mixed uncertainty can then be used to calculate the final consistency loss, as shown below:
[0053]
[0054] where v represents a voxel, Cmix is the calculated mixing uncertainty, γ is the uncertainty weighting factor, and when γ = 0, the consistency loss Lc will be equivalent to the standard average teacher method.
[0055] Compared with the prior art, the present invention has the following beneficial effects:
[0056] Compared with the traditional supervised segmentation method, the present invention does not require a large amount of artificial prior knowledge. Relying on the powerful learning ability of the neural network, the model has strong performance and more accurate segmentation results in the task of segmenting lesion images.
[0057] Compared with other deep learning semi-supervised segmentation methods, the present invention is reasonably designed and expanded based on 3D CNN, so that 3D CNN can fuse and learn semantic features of different scales, thereby learning the volume context information of the image and using it for training, so as to strengthen the model's application of context information, breaking the limitation of insufficient utilization of image volume context in previous consistency semi-supervised methods. At the same time, the calculation method of capturing global information by anatomical perception representation prior is used to calculate global uncertainty and the traditional calculation uncertainty calculation method is used to calculate local uncertainty, and the local uncertainty information is combined with the global uncertainty information to obtain a mixed uncertainty to help the model better determine the uncertainty area of the image, thereby improving the segmentation accuracy. Compared with other excellent semi-supervised segmentation methods, it has better performance and generalization ability; and it significantly reduces the parameters of the model, thereby reducing the possibility of overfitting and meeting the lightweight requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0059] Figure 1 A flow chart of the overall remote physiological measurement method provided by the present invention;
[0060] Figure 2 A schematic diagram of the CUAMT model provided by the present invention;
[0061] Figure 3 A schematic diagram of a context information extraction module provided by the present invention;
[0062] Figure 4 A schematic diagram of a hybrid uncertainty module provided by the present invention;
[0063] Figure 5 A visual comparison diagram of the segmentation results of different models provided by the present invention on dataset A;
[0064] Figure 6 This is a visual comparison chart of the segmentation results of different models provided by the present invention on data set B. DETAILED DESCRIPTION
[0065] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0066] The present invention focuses on seeking more complex data-level or model-level disturbances for unsupervised consistency learning to improve the training effect of the model in view of the existing uncertainty regularization, thereby ignoring the learning of the image volume context and the existing uncertainty calculation. The image uncertainty is mainly estimated at the pixel level. It does not consider the global information of the image and other issues, resulting in segmentation difficulties. This patent proposes a new fast, efficient and accurate hybrid uncertainty semi-supervised segmentation model CUAMT based on context information. A context extraction module is designed in CUAMT. By extracting the semantic features of the input image at different scales, the volume context information of the image is learned and used for training to enhance the model's application of context information. At the same time, an uncertainty mixing module is also designed to ensure that the model better selects a certain area to guide the segmentation task by mixing global uncertainty information and local uncertainty information. The problem of poor robustness and low segmentation accuracy in the current consistency semi-supervised segmentation is solved. In order to reflect the performance of the proposed semi-supervised medical image segmentation method in different situations, multiple data sets are used for training verification to expand the scope of application of the method.
[0067] Embodiment 1:
[0068] See also Figure 1 , a hybrid uncertainty MRI semi-supervised segmentation method based on context information, comprising at least the following steps:
[0069] S1: Preprocess the medical images in the public dataset;
[0070] S2: Divide the original data set into two parts: labeled images and unlabeled images, and divide the ratio;
[0071] S3: Design a network architecture for learning context information, i.e. build a context extraction module;
[0072] S4: Design a loss function for mixed uncertainty calculation, that is, build an uncertainty mixing module;
[0073] S5: Build the CUAMT model based on the context extraction module and the uncertainty mixing module, use the Mean-Teacher architecture to train the model and perform measurement verification, and the model training is completed;
[0074] S6: Use the trained model for supervised segmentation.
[0075] CUAMT mainly consists of two modules: context extraction module and uncertainty mixing module. The context extraction module extracts the semantic features of the input image at different scales, thereby learning the volume context information of the image and using it for training to enhance the model's application of context information. The uncertainty mixing module is used to mix global uncertainty information and local uncertainty information to ensure that the model can better select certain areas to guide the segmentation task. The combination of the two improves the performance and generalization ability of the model. The overall architecture of the model can be seen in Figure 2 .
[0076] The context extraction module is as follows Figure 3 As shown, it mainly includes two parts: feature extraction and 3D attention mechanism.
[0077] In the feature extraction part, we use a hybrid pyramid structure to extract contextual features. Specifically, the execution process of this module is as follows: First, the input image passes through a 1x1x1 3D convolution operation to reduce the number of computational parameters and channel dimensions of the feature map. Then, the module is divided into two parts: the first part uses a 3x3x3 3D convolution to extract the original features of the image (such as texture, shape, and edges, etc.); the second part uses 3D maximum pooling with strides of 2, 4, and 6 to perform pooling operations with different kernel sizes, followed by a 3x3x3 3D convolution operation to extract higher-level features and achieve the extraction of features of different scales. After that, the three convolution operations are upsampled to restore the pooled feature map to the same input size as the original image for subsequent feature fusion to obtain a semantically rich feature map, thereby enhancing the network's understanding of the spatial structure and contextual information of the input data.
[0078] In the 3D attention mechanism part, we improved the CA attention mechanism into a 3D attention mechanism. First, we replaced its internal convolution and pooling operations with 3D convolution and 3D pooling, and added BatchNorm3d as a normalization operation, thereby improving it from a 2D attention mechanism to a 3D attention mechanism. The 3D attention mechanism combines the channel information and position information of the input data. The channel information provides rich feature expressions (such as texture, shape, and edges, etc.), while the position information provides contextual information between different elements. The combination of the two enables the model to better understand the semantic features and contextual information of the input image, thereby improving the performance and expression ability of the model.
[0079] When designing the uncertainty hybrid module, in order to address the problem that the current uncertainty estimation method estimates uncertainty at the pixel level, it does not consider the global information in the dataset, which may cause the model to produce misleading uncertainty estimates and limit the propagation of uncertainty. Consider introducing a denoising autoencoder (DAE), which can encode the available segmentation mask in a nonlinear latent space to learn anatomically perceived prior representations. This learned representation can capture the global information from the segmentation mask, thereby mapping inaccurate predictions to credible segmentations. To learn the global information captured from the segmentation mask, a global uncertainty is calculated. Then, by fusing local and global uncertainties, the model is guaranteed to better select certain areas to guide the segmentation task.
[0080] Uncertainty hybrid modules such as Figure 4 As shown, in the calculation of local uncertainty, the existing uncertainty calculation method is used to calculate the local uncertainty of each sample by performing multiple forward propagations during the training process and calculating the average value of multiple prediction results. The formula is as follows:
[0081]
[0082] Where T represents the number of predictions, Represents the prediction result of input x using weights at time step t, P c is the average prediction result, and H(p)1 represents our local uncertainty.
[0083] In the global uncertainty calculation, a denoising autoencoder (DAE) is introduced to encode the available segmentation masks in a nonlinear latent space to learn anatomically aware prior representations.
[0084] This learned representation can capture the global information from the segmentation mask, thereby mapping inaccurate predictions to credible segmentations to learn the global information captured from the segmentation mask, thereby calculating a global uncertainty, which is calculated as follows:
[0085] H(p)2=∥f d (x i )-f t (x i )∥ 2
[0086] where f d (x i ) is the DAE for the input x i The prediction result, f t (x i ) is the teacher model for input x iThe prediction results, H(p)2 represents our global uncertainty.
[0087] Then, the calculated local uncertainty and global uncertainty are combined to obtain the final mixed uncertainty, as shown in the following formula:
[0088] C1=1-H(p)1
[0089]
[0090] Cmix=αC1+(1-αC2)
[0091] Among them, C1 and C2 are reliable regions calculated by the information compression method in information theory using local and global uncertainties, α is an adjustable weight control factor set according to experience, e is a mathematical constant, which is approximately equal to 2.71828, and mix is the mixed uncertainty.
[0092] This mixed uncertainty can then be used to calculate the final consistency loss, as shown below:
[0093]
[0094] where v represents a voxel, Cmix is the calculated mixing uncertainty, γ is the uncertainty weighting factor, and when γ = 0, the consistency loss Lc will be equivalent to the standard average teacher method.
[0095] The model was validated and measured using at least the Dice similarity coefficient (Dice), Jaccard index (Jaccard), 95% Hausdorff distance (95HD) and average surface distance (ASD) as evaluation metrics for semi-supervised segmentation.
[0096]
[0097] In the formula, TP, TN, FP and FN represent the number of true positives, true negatives, false positives and false negatives, respectively. A and B represent the ground truth and segmentation result, respectively. SA and SB are the surface voxel sets corresponding to A and B, respectively, and d(sB,S(A)) represents the shortest Euclidean distance from voxel sB to set sA. Similarly, d(sA,S(B)) is the shortest Euclidean distance from voxel sA to set S(B). In addition, 95HD is defined as the 95th quantile of Hausdorff distance (95HD), not the maximum value.
[0098] Embodiment 2:
[0099] Based on the first embodiment, the present embodiment takes the cardiac image as an application example, and other image applications may also be processed in different embodiments.
[0100] A hybrid uncertainty MRI semi-supervised segmentation method based on context information includes at least the following steps:
[0101] S1: Preprocess the cardiac images in the public dataset;
[0102] S2: Divide the original data set into two parts: labeled images and unlabeled images, and divide the ratio;
[0103] S3: Design a network architecture for learning context information, i.e. build a context extraction module;
[0104] S4: Design a loss function for mixed uncertainty calculation, that is, build an uncertainty mixing module;
[0105] S5: Build the CUAMT model based on the context extraction module and the uncertainty mixing module, use the Mean-Teacher architecture to train the model and perform measurement verification, and the model training is completed;
[0106] S6: Use the trained model for supervised segmentation.
[0107] S1 includes at least the following steps:
[0108] All volumes of the original cardiac images were normalized to zero mean and unit variance;
[0109] Each 3D MRI volume is then cropped with its enlarged edges according to the target;
[0110] Each 3D image is cropped to a size of 128×128×80.
[0111] S2 includes at least the following steps:
[0112] 80 images of the original data set are used for training, 20 images are used for validation, and 54 images are used for testing;
[0113] Only 10% of the 80 heart images used for training have labels, and the labels of the remaining images are removed.
[0114] The design idea of the context extraction module in S3 is:
[0115] In order to obtain semantic feature information of different dimensions in cardiac images, a context information learning network is designed based on 3D CNN;
[0116] By extracting features from images at different scales and upsampling them to the same scale for information fusion, the model can better learn the volume context information of the image, making the model training more accurate and robust.
[0117] In summary:
[0118] The present invention takes the accuracy of lesion segmentation as an example. Two publicly recognized data sets are used for processing to verify the effectiveness of the method proposed in the present invention. Among them, after preprocessing, 80 images of data set A are used for training, 20 images are used for verification, and 54 images are used for testing. Only 8 images have labels in the training stage; after preprocessing, 250 images of data set B are used for training, 25 images are used for verification, and 60 images are used for testing. Only 25 images have labels in the training stage. When measuring the performance of the test method, the following indicators are used to evaluate the measurement model: Dice similarity coefficient (Dice), Jaccard index (Jaccard), 95% Hausdorff distance (95HD) and average surface distance (ASD).
[0119] The method based on the CUAMT model is compared with a variety of excellent semi-supervised segmentation methods. The segmentation accuracy on dataset A is shown in Table 1, and the visual comparison effect is shown in Figure 5 ; The segmentation accuracy on dataset B is shown in Table 2, and the visual comparison effect is shown in Figure 6 ; The segmentation results of two organs in different parts show that CUAMT has strong generalization ability and robustness. All experiments show that Style-rPPG has the best measurement performance among all unsupervised methods, and has strong cross-dataset capabilities. At the same time, the measurement speed is fast and meets the lightweight requirements.
[0120] Table 1 Segmentation comparison results on dataset A (the best one in each indicator is shown in bold)
[0121]
[0122]
[0123] Table 2 Segmentation comparison results on dataset B (the best in each indicator is shown in bold)
[0124]
[0125]
[0126] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
Claims
1. A hybrid uncertainty MRI semi-supervised segmentation method based on contextual information, characterized by: At least the following steps are included: S1: Preprocess the medical images in the public dataset; S2: Divide the original data set into two parts: labeled images and unlabeled images, and divide the ratio; S3: Design a network architecture for learning context information, i.e. build a context extraction module; S4: Design a loss function for mixed uncertainty calculation, that is, build an uncertainty mixing module; S5: Build the CUAMT model based on the context extraction module and the uncertainty mixing module, use the Mean-Teacher architecture to train the model and perform measurement verification, and the model training is completed; S6: Use the trained model for supervised segmentation.
2. The hybrid uncertainty MRI semi-supervised segmentation method based on contextual information according to claim 1 is characterized by: The S1 at least comprises the following steps: All volumes of the original medical images were normalized to zero mean and unit variance; Each 3D MRI volume is then cropped with its enlarged edges according to the target; Each 3D image is cropped to a size of 128×128×80.
3. The hybrid uncertainty MRI semi-supervised segmentation method based on contextual information according to claim 1, characterized in that: The S2 at least comprises the following steps: 80 images of the original data set are used for training, 20 images are used for validation, and 54 images are used for testing; Among the 80 cardiac medical images used for training, only 10% of the images have labels, and the labels of the remaining images are removed.
4. The hybrid uncertainty MRI semi-supervised segmentation method based on contextual information according to claim 1, characterized in that: The design idea of the context extraction module in S3 is: In order to obtain semantic feature information of different dimensions in medical images, a context information learning network is designed based on 3D CNN; By extracting features from images at different scales and upsampling them to the same scale for information fusion, the model can better learn the volume context information of the image, making the model training more accurate and robust.
5. The hybrid uncertainty MRI semi-supervised segmentation method based on contextual information according to claim 4, characterized in that: The context extraction module includes two parts: feature extraction and 3D attention mechanism; The feature extraction part uses a hybrid pyramid structure to extract context features; In the 3D attention mechanism section, the CA attention mechanism is improved to a 3D attention mechanism. In the 3D attention mechanism section, we improve the CA attention mechanism into a 3D attention mechanism.
6. The hybrid uncertainty MRI semi-supervised segmentation method based on contextual information according to claim 5, characterized in that: The execution process of the feature extraction part includes at least the following steps: First, the input image passes through a 1x1x1 3D convolution operation to reduce the number of computational parameters and channel dimensions of the feature map; Next, it is divided into two parts; The first part uses 3x3x3 3D convolution to extract the original features of the image, and the original features include at least texture, shape and edge; The second part uses 3D maximum pooling with strides of 2, 4, and 6 to perform pooling operations with different kernel sizes, followed by a 3x3x3 3D convolution operation to extract more advanced features and achieve the extraction of features of different scales; After that, the above three convolution operations are followed by an upsampling operation to restore the pooled feature map to the same input size as the original image for subsequent feature fusion to obtain a semantically rich feature map, thereby enhancing the network’s understanding of the spatial structure and contextual information of the input data.
7. The hybrid uncertainty MRI semi-supervised segmentation method based on contextual information according to claim 5, characterized in that: The execution process of the 3D attention mechanism part includes at least the following steps: First, we replace the convolution and pooling operations inside the network with 3D convolution and 3D pooling, and add BatchNorm3d as a normalization operation, thereby improving it from a 2D attention mechanism to a 3D attention mechanism. The 3D attention mechanism combines the channel information and position information of the input data. The channel information provides rich feature expressions, while the position information provides contextual information between different elements. The feature expressions include at least texture, shape and edge. The combination of channel information and position information enables the model to better understand the semantic features and contextual information of the input image, thereby improving the performance and expression ability of the model.
8. The hybrid uncertainty MRI semi-supervised segmentation method based on contextual information according to claim 1, characterized in that: The uncertainty mixing module uses the existing uncertainty calculation method in the calculation of local uncertainty. It calculates the local uncertainty of each sample by performing multiple forward propagations during the training process and calculating the average value of multiple prediction results. The formula is as follows: Where T represents the number of predictions, represents the prediction result of input x using weights at time step t, P c is the average prediction result, H(p)1 represents the local uncertainty, In the global uncertainty calculation, a denoising autoencoder (DAE) is introduced to encode the available segmentation masks in a nonlinear latent space to learn anatomically aware prior representations. This learned representation can capture the global information from the segmentation mask, thereby mapping inaccurate predictions to credible segmentations to learn the global information captured from the segmentation mask, thereby calculating a global uncertainty, which is calculated as follows: H(p)2=∥f d (x i )-f t (x i )∥ 2 where f d (x i ) is the DAE for the input x i The prediction result, f t (x i ) is the teacher model for input x i The prediction results, H(p)2 represents the global uncertainty, Then, the calculated local uncertainty and global uncertainty are combined to obtain the final mixed uncertainty, as shown in the following formula: C1=1-H(p)1 Cmix=αC1+(1-αC2) Among them, C1 and C2 are reliable regions calculated by information compression method in information theory using local and global uncertainties, α is an adjustable weight control factor set according to experience, γ is the uncertainty weighting factor, e is a mathematical constant, and Cmix is the mixed uncertainty; This mixed uncertainty can then be used to calculate the final consistency loss, as shown below: where v represents a voxel, γ is the uncertainty weighting factor, and when γ = 0, the consistency loss Lc will be equivalent to the standard average teacher method.