Semi-supervised heart and brain image segmentation method and system based on dynamic prototype learning
Through the dynamic prototype learning method, combined with the advantages of labeling and unlabeled data, the problem that prototype learning in the existing technology is difficult to capture data connections, and better semantic fusion and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510314176.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-01
AI Technical Summary
The existing semi-supervised brain image segmentation method based on prototype learning is difficult to effectively capture the connection between labeled data and unlabeled data, making it difficult for the learned prototype to provide a comprehensive understanding of the data, especially when the boundaries between categories are blurred or labeled data are scarce.
The semi-supervised heart-brain image segmentation method based on dynamic prototype learning is adopted. By obtaining labeled images and unlabeled images, the teacher network generates pseudo-labels, and mixed prototype learning is performed through the student network, and the prototype learning strategy is dynamically adjusted to combine the accuracy of labeled data and the richness of unlabeled data.
The semantic fusion of labeled and unlabeled data is realized, ensuring that prototype generation can combine the accuracy of labeled data and the richness of unlabeled data, improve the representativeness and generalization ability of prototype learning, and adapt to different data distribution and feature changes.
Smart Images

Figure CN120236079A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cardiac and cerebral image segmentation, and particularly to a semi-supervised cardiac and cerebral image segmentation method and system based on dynamic prototype learning. Background Art
[0002] In the field of medical imaging, image segmentation technology plays a crucial role in the diagnosis, treatment planning, and effect evaluation of cardiac and cerebral diseases. The goal of image segmentation is to accurately extract the target regions (such as the heart, brain, etc.) from complex medical images, so that doctors can more clearly analyze the lesion sites and make decisions.
[0003] In traditional supervised learning, the model relies on a large amount of labeled data for training. However, the acquisition of labeled data is usually expensive and time-consuming. Especially in the field of medical images, the labeling workload of professional doctors is huge. Therefore, the acquisition cost of labeled data is very high. Unsupervised learning, although it can use a large amount of unlabeled data for training, its performance is often affected by data noise and the lack of clear labels, making it difficult to ensure high accuracy. Semi-supervised learning is precisely proposed to make up for this contradiction. It combines a small amount of labeled data with a large amount of unlabeled data to improve the performance of the learning model. In the task of medical image segmentation, semi-supervised learning can effectively utilize the potential of unlabeled data, reduce the workload of manual labeling, and at the same time improve the generalization ability and robustness of the model. This makes semi-supervised learning an important research direction in cardiac and cerebral image segmentation, especially in the case where the cost of data labeling is high and the acquisition is difficult, it has great application value.
[0004] Traditional semi-supervised methods usually rely on the consistency assumption or self-training methods to generate pseudo-labels. These methods are easily disturbed by noise and incorrect labels, and lack an effective mechanism to capture the potential structure of unlabeled data, thus affecting the learning effect. Different from traditional semi-supervised methods, semi-supervised methods based on prototype learning can effectively capture the potential structure of data by constructing prototypes or class centers for each category, and optimize the embedding distribution of different category features, thus overcoming the deficiencies of traditional methods. In existing patents, such as Patent Application Publication No.: CN119169292A and CN118864486A, both belong to semi-supervised methods based on prototypes in the field of medical image segmentation. However, in the research, it is found that the existing semi-supervised methods based on prototypes have the following problems:
[0005] 1. The labeled data is located at the tail of the overall distribution, and relying solely on the labeled data cannot reflect the overall distribution. Most existing methods construct prototypes using labeled data and unlabeled data respectively. Since the connection between the two cannot be effectively captured and utilized, the learned prototypes are difficult to provide a comprehensive understanding of the data, especially in the case where the boundaries between categories are blurred or the labeled data is scarce.
[0006] 2. Existing prototype fusion methods usually adopt a fixed growth mode to calculate the consistency weight between unlabeled data and labeled data, which has certain limitations when dealing with diverse data sets. This fixed growth mode lacks adaptability to the characteristics of the data set. Due to the difficulty in flexibly coping with changes in different data distributions and features, the effect of prototype fusion may be limited. Summary of the Invention
[0007] In view of the deficiencies of the existing technology, the present invention provides a semi-supervised heart and brain image segmentation method and system based on dynamic prototype learning, which solves the technical problems existing in the above-mentioned existing technology.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] In the first aspect, the present invention provides a semi-supervised heart and brain image segmentation method based on dynamic prototype learning, including:
[0010] Obtain labeled images and unlabeled images, and preprocess the labeled images and unlabeled images; input the unlabeled images into a trained teacher network to output pseudo-labels of the unlabeled images;
[0011] Select labeled images and unlabeled images, mix them by exchanging regions of the labeled images and unlabeled images at different spatial positions, and input the mixed images into a trained student network to output segmented images and intermediate features;
[0012] The student network is optimized through the intermediate features, calculates the confidence of the unlabeled regions according to the unlabeled regions in the intermediate features, obtains the confidence weight, and obtains the unlabeled prototype according to the confidence weight; randomly samples the labeled regions according to the labeled regions in the intermediate features to obtain labeled prototypes; fuse the unlabeled prototypes and the labeled prototypes to generate mixed prototypes.
[0013] As a further technical solution, the training method of the teacher network is: first randomly select two groups of images in the labeled images, randomly select two regions on the two groups of images, cut out the selected regions and paste them at the corresponding positions of the other images to generate two groups of first mixed images; then input the generated first mixed images into the first U-Net network for training, and the trained first U-Net network is the teacher network.
[0014] As a further technical solution, the training method of the student network is as follows: First, randomly select labeled images and unlabeled images, randomly select cropping regions on the labeled images and unlabeled images respectively, crop the cropping regions and paste them to the corresponding positions of the other images to generate second mixed images; then input the generated second mixed images into the second U-Net network for training, and the trained second U-Net network is the student network.
[0015] As a further technical solution, mix the pseudo-labels with the true labels of the labeled images to generate mixed labels and use them to supervise the student network. Specifically: Obtain the pseudo-labels output by the teacher network and the true labels in the labeled images, randomly determine cropping regions on the true labels and pseudo-labels respectively, crop the cropping regions and paste them to the corresponding positions of the other images to obtain mixed labels.
[0016] As a further technical solution, for calculating the confidence of the unlabeled regions in the intermediate features and obtaining the confidence weights, the specific method is: By obtaining the class probability distribution output by the student network, calculate the difference between the maximum probability and the second probability of each voxel point to obtain the class probability difference, and perform normalization processing on the class probability difference, and use the normalized class probability difference as the confidence weights.
[0017] As a further technical solution, the specific method for obtaining the unlabeled prototypes according to the confidence weights is: Determine whether to be sampled from the unlabeled regions in the intermediate features according to the confidence weights of each voxel point, perform average aggregation on the sampled features, and calculate prototypes with the same number of channels as the intermediate features, which are the unlabeled prototypes.
[0018] As a further technical solution, the specific method for obtaining the labeled prototypes by random sampling is: There are valid voxels in the labeled regions of the intermediate features, randomly select voxels, and perform average aggregation on the features obtained from the selected voxels to obtain the labeled prototypes.
[0019] In a second aspect, the present invention provides a semi-supervised heart and brain image segmentation system based on dynamic prototype learning, including the following modules:
[0020] A pre-training module, configured to: Obtain labeled images and unlabeled images, and preprocess the labeled images and unlabeled images; Input the unlabeled images into the trained teacher network to output the pseudo-labels of the unlabeled images;
[0021] A self-training module, configured to: Select labeled images and unlabeled images, mix them by exchanging regions of the labeled images and unlabeled images at different spatial positions, and input the mixed images into the trained student network to output segmentation images and intermediate features;
[0022] The prototype learning module is configured to: optimize the student network through the intermediate features, calculate the confidence of the unlabeled area based on the unlabeled area in the intermediate features, obtain the confidence weight, and obtain the unlabeled prototype according to the confidence weight; obtain the labeled prototype by random sampling according to the labeled area in the intermediate features; and fuse the unlabeled prototype and the labeled prototype to generate a hybrid prototype.
[0023] One or more technical solutions of the present invention have the following beneficial effects:
[0024] The present invention maps labeled and unlabeled data to the same latent space, enabling semantic fusion of the two, ensuring that subsequent prototype generation can combine the accuracy of labeled data and the richness of unlabeled data; two different sampling methods are used for labeled features and unlabeled features, avoiding the model relying too much on certain specific voxel points, thereby improving the representativeness and generalization ability of prototype learning.
[0025] The present invention uses a dynamic adjustment function to control the prototype learning strategy. In the initial stage of training, it relies on labeled data. As the training progresses, the weight of unlabeled data is gradually increased, and the fusion strategy can be flexibly adjusted according to task requirements, thereby enhancing the adaptability and generalization ability of the method to different data distributions and feature changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0027] Figure 1 It is a flowchart of the semi-supervised cardiac and brain image segmentation method based on dynamic prototype learning of the present invention;
[0028] Figure 2 It is a flowchart of the pre-training part of the present invention;
[0029] Figure 3 It is a flowchart of the self-training part of the present invention;
[0030] Figure 4 It is a flowchart of the prototype learning module of the present invention;
[0031] Figure 5 It is a graph of the dynamic adjustment function of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0033] Example 1
[0034] In this embodiment, a semi-supervised heart and brain image segmentation method based on dynamic prototype learning is provided, as Figure 1 shown, specifically including the following methods:
[0035] 1. Pretraining part:
[0036] Obtain labeled images and unlabeled images, and preprocess the labeled images and unlabeled images. Among them, as Figure 1 shown, the given dataset D = {D l , D i}, which consists of N labeled images and M unlabeled images (N << M). The labeled image dataset Unlabeled image dataset The method for preprocessing the labeled images and unlabeled images is: smooth the images and reduce noise, and perform data augmentation operations such as random flipping and cropping.
[0037] As Figure 2 shown, pre-train the teacher network in advance. The specific training method is: first randomly select two groups of images from the labeled images, randomly select two regions on the two groups of images, cut out the selected regions and paste them to the corresponding positions of the other images to generate two groups of first mixed images; then input the generated first mixed images into the first U-Net network for training. The trained first U-Net network is the teacher network, which is used to generate pseudo-labels of unlabeled images for the student network. Specifically, input the unlabeled images into the trained teacher network to output the pseudo-labels of the unlabeled images.
[0038] 2. Self-training part:
[0039] As Figure 3 shown, select labeled images and unlabeled images, mix them by exchanging regions of the labeled images and unlabeled images at different spatial positions, and input the mixed images into the second U-Net network for training as the student network. The student network is optimized by stochastic gradient descent, and the teacher network is optimized by the exponential moving average of the student network.
[0040] Specifically: first randomly select labeled images and unlabeled images, randomly select cropping regions on the labeled images and unlabeled images, cut out the cropping regions and paste them to the corresponding positions of the other images, which can achieve semantic fusion of the two, and generate two second mixed images I a , I b ; then input the generated second mixed images I a , I b into the student network to output two segmentation images O a , O b and intermediate features fa , f b . To supervise the student network, the pseudo - labels generated by the teacher network and the true labels (Y a , Y b ) of the annotated images are subjected to the same mixing operation to obtain the mixed label M a and M b . Specifically, the method for obtaining the mixed label is as follows: Mix the pseudo - labels with the true labels of the annotated images to generate a mixed label, which is used to supervise the student network. Specifically: Obtain the pseudo - labels output by the teacher network and the true labels in the annotated images, randomly determine the cropping regions on the true labels and the pseudo - labels respectively, crop the cropping regions and paste them to the corresponding positions of the other image to obtain the mixed label.
[0041] M a and M b will be used to supervise the student network, and the loss function L a , O b for segmenting the image O mix can be calculated by the following formula:
[0042] where L mix is a linear combination of the Dice loss and the cross - entropy loss.
[0043] 3. Prototype learning part:
[0044] As Figure 4 shown, this part explains that the student network is optimized through the above - mentioned intermediate features. First, according to the unannotated regions in the intermediate features, calculate the confidence of the unannotated regions and obtain the confidence weights, and obtain the unannotated prototypes according to the confidence weights; then, according to the annotated regions in the intermediate features, use random sampling to obtain the annotated prototypes; fuse the unannotated prototypes and the annotated prototypes to generate the mixed prototypes.
[0045] Specifically: For the unannotated regions a in the intermediate feature f , a confidence - based probability sampling method is adopted to enhance the reliability of learning prototypes from unannotated images. Through confidence - based probability sampling, when selecting voxel points, it is possible to avoid relying too much on high - confidence regions, ensure that low - confidence regions also have the opportunity to be considered, and thus improve the representativeness and generalization ability of prototype learning. For each voxel point, first obtain the class probability distribution output by the student network, calculate the difference between the maximum probability and the second - maximum probability of each voxel point to obtain the class probability difference ΔP i , ΔP i =P(x i , c1)-P(x i , c2); where P(xi ,c1) represents the i-th voxel point x i The probability of belonging to category c1, P(x i , c2) represents the i-th voxel point x i The probability of belonging to category c2. ΔP i The smaller it is, the higher the uncertainty of the student network in predicting the voxel category. The category probability difference is normalized and represented as ΔP i norm ,The normalization process standardizes the category probability difference to a uniform range, and uses the normalized category probability difference as the confidence weight w i , w i =ΔP i norm .
[0046] The specific method of obtaining the unlabeled prototype according to the confidence weight is as follows: a According to the confidence weight w of each voxel point in the unlabeled area i Decide whether to be sampled, average and aggregate the features composed after sampling, and calculate the prototype with the same number of intermediate feature channels, which is the unlabeled prototype. The features composed after sampling are expressed as: in, Represents the features after sampling, Cat(f a ,*) indicates f a Each voxel point in the sample is selected according to the confidence weight. The features composed after sampling are averaged and aggregated to calculate multiple prototypes with the same number of intermediate feature channels. It is the unlabeled prototype, which represents the potential semantic information of the unlabeled part. Each prototype corresponds to an intermediate feature channel, reflecting the comprehensive semantic features of the channel.
[0047] For the intermediate feature f a The marked area in A random sampling strategy is used to increase the diversity of generated prototypes. Assume that the labeled area Total T l valid voxels, randomly select T l ×(1-w d (t)) The voxel gets the feature Perform average aggregation to obtain the labeled prototype Among them, w d (t) is a dynamic adjustment function, which indicates the dynamic sampling ratio in the (t)th round of training. <w d <1, here it is used to adjust the number of random samples during the training process. Through this dynamic adjustment function w d Weighted fusion of unlabeled and labeled prototypes to obtain the hybrid prototype qa , Finally, the probability of each class voxel is approximated by calculating the cosine similarity between the feature and the prototype, and the prototype loss is obtained:
[0048] where L pc represents the prototype loss, and s a is the cosine similarity between q a and f a . The total loss function in the self-training stage (student network) is defined as:
[0049] L total = L mix + L pc .
[0050] 4. Dynamic adjustment function part:
[0051] The dynamic adjustment function w d is defined in the above description, which is used to weighted fuse the unlabeled and labeled prototypes. By gradually reducing the number of samples from the labeled area and increasing the importance of the prototypes generated by the unlabeled area, dynamic fusion is achieved. Specifically, in this embodiment, an exponential growth weight update strategy is adopted. A smaller weight is used at the beginning of training to maintain the dependence on the labeled data. As the training progresses, the weight of the unlabeled data in the prototype fusion process is gradually increased. The growth rate α and the growth threshold h are introduced into the dynamic adjustment function. α determines the rate of increase in the contribution of the unlabeled data to the prototype fusion during training, and the growth threshold sets the maximum influence degree of the unlabeled data on the prototype fusion process. By reasonably setting these two parameters, the dynamic adjustment function can flexibly adjust the fusion strategy according to the task requirements, thereby enhancing the adaptability and generalization ability of the method. The formula of the dynamic adjustment function is as follows:
[0052] where t_max is the total number of training epochs.
[0053] 5. Evaluation index part:
[0054] MSE is used to measure the difference between the predicted segmentation result and the true label. The smaller its value, the more accurate the segmentation result. The formula is as follows:
[0055] The Dice coefficient is an index to measure the overlap degree between the predicted and true regions. The closer its value is to 1, the better the segmentation effect. It is usually used to evaluate the similarity of image segmentation. The formula is: The cross-entropy loss guides the model optimization by calculating the gap between the predicted class and the true label. The formula is:
[0056] Example Two
[0057] In this embodiment, a weather prediction system for multi-scale spatio-temporal feature extraction and fusion is provided, including the following modules:
[0058] A pre-training module, configured to: obtain labeled images and unlabeled images, and preprocess the labeled images and unlabeled images; input the unlabeled images into a trained teacher network to output pseudo-labels of the unlabeled images;
[0059] A self-training module, configured to: select labeled images and unlabeled images, mix them by exchanging regions of the labeled images and unlabeled images at different spatial positions, and input the mixed images into a trained student network to output segmented images and intermediate features;
[0060] A prototype learning module, configured to: optimize the student network through the intermediate features, calculate the confidence of the unlabeled regions according to the unlabeled regions in the intermediate features, and obtain confidence weights, and obtain unlabeled prototypes according to the confidence weights; obtain labeled prototypes by random sampling according to the labeled regions in the intermediate features; fuse the unlabeled prototypes and the labeled prototypes to generate mixed prototypes.
[0061] For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A semi-supervised heart-brain image segmentation method based on dynamic prototype learning, characterized in that: include: Obtain annotated images and unannotated images, and preprocess the annotated images and unannotated images; Input the unlabeled image into the trained teacher network and output the pseudo label of the unlabeled image; Select annotated images and unannotated images, mix them by exchanging the regions of the annotated images and unannotated images at different spatial locations, input the mixed images into the trained student network, and output the segmented images and intermediate features; The student network is optimized through the intermediate features, and the confidence of the unlabeled area is calculated according to the unlabeled area in the intermediate features, and the confidence weight is obtained, and the unlabeled prototype is obtained according to the confidence weight; according to the labeled area in the intermediate features, a labeled prototype is obtained by random sampling; the unlabeled prototype is fused with the labeled prototype to generate a hybrid prototype.
2. The semi-supervised heart-brain image segmentation method based on dynamic prototype learning as claimed in claim 1, characterized in that: The training method of the teacher network is as follows: first, randomly select two groups of images from the annotated images, randomly select two areas on the two groups of images respectively, crop the selected areas and paste them to the corresponding positions of the other image to generate two groups of first mixed images; then input the generated first mixed images into the first U-Net network for training, and the trained first U-Net network is the teacher network.
3. The semi-supervised heart-brain image segmentation method based on dynamic prototype learning as claimed in claim 1, characterized in that: The training method of the student network is as follows: first, randomly select a labeled image and an unlabeled image, randomly select a cropping area on the labeled image and the unlabeled image respectively, crop the cropping area and paste it to the corresponding position of the other image to generate a second mixed image; then input the generated second mixed image into the second U-Net network for training, and the trained second U-Net network is the student network.
4. The semi-supervised heart-brain image segmentation method based on dynamic prototype learning as claimed in claim 1, characterized in that: The pseudo-label is mixed with the real label of the annotated image to generate a mixed label, and is used to supervise the student network. Specifically, the pseudo-label output by the teacher network and the real label in the annotated image are obtained, and the cropping area is randomly determined on the real label and the pseudo-label respectively, and the cropping area is cropped and pasted to the corresponding position of the other image to obtain the mixed label.
5. The semi-supervised heart-brain image segmentation method based on dynamic prototype learning as claimed in claim 1, characterized in that: The method of calculating the confidence of the unlabeled area according to the unlabeled area in the intermediate feature and obtaining the confidence weight is as follows: by obtaining the category probability distribution output by the student network, calculating the difference between the maximum probability and the second probability of each voxel point, obtaining the category probability difference, and normalizing the category probability difference, and using the normalized category probability difference as the confidence weight.
6. The semi-supervised heart-brain image segmentation method based on dynamic prototype learning as claimed in claim 1, characterized in that: The specific method to obtain the unlabeled prototype based on the confidence weight is as follows: from the unlabeled area in the intermediate feature, decide whether to be sampled according to the confidence weight of each voxel point, average and aggregate the features composed after sampling, and calculate the prototype with the same number of channels as the intermediate feature, which is the unlabeled prototype.
7. The semi-supervised heart-brain image segmentation method based on dynamic prototype learning as claimed in claim 1, characterized in that: The specific method of obtaining the labeled prototype by random sampling is: there are valid voxels in the labeled area of the intermediate features, voxels are randomly selected, and the features obtained by the selected voxels are averaged and aggregated to obtain the labeled prototype.
8. A semi-supervised heart-brain image segmentation system based on dynamic prototype learning, characterized in that: Includes the following modules: The pre-training module is configured to: obtain annotated images and unannotated images, and pre-process the annotated images and unannotated images; input the unannotated images into the trained teacher network, and output pseudo labels of the unannotated images; The self-training module is configured to: select an annotated image and an unannotated image, mix the annotated image and the unannotated image by exchanging regions of the annotated image and the unannotated image at different spatial locations, input the mixed images into the trained student network, and output the segmented image and the intermediate features; The prototype learning module is configured as follows: the student network is optimized through the intermediate features, the confidence of the unlabeled area is calculated according to the unlabeled area in the intermediate features, and the confidence weight is obtained, and the unlabeled prototype is obtained according to the confidence weight; according to the labeled area in the intermediate features, a labeled prototype is obtained by random sampling; and the unlabeled prototype is fused with the labeled prototype to generate a hybrid prototype.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the semi-supervised heart-brain image segmentation method based on dynamic prototype learning as described in any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the semi-supervised heart and brain image segmentation method based on dynamic prototype learning as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Semi-supervised medical image segmentation method based on model self-distillation and prototype learning
CN118864486A
Semi-supervised skin image automatic segmentation method based on boundary perception and hybrid prototype consistency learning
CN119169292A