Consistency-based semi-supervised left atrium segmentation method based on teacher-student model

By constructing a teacher-student model for consistent semi-supervised learning, and combining adversarial perturbation and consistency loss, the problems of time-consuming and labor-intensive traditional medical image segmentation methods and the dependence of deep learning models on a large amount of labeled data are solved. This achieves efficient and accurate left atrial segmentation, which is applicable to the diagnosis and treatment of atrial fibrillation patients.

CN119904642BActive Publication Date: 2025-10-24GUILIN UNIV OF ELECTRONIC TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510085136.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-10-24
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Traditional medical image segmentation methods are time-consuming and labor-intensive, relying on a large amount of labeled data, which makes it difficult to meet the needs of efficient and accurate diagnosis of atrial fibrillation patients. Furthermore, deep learning models do not perform well when trained with limited labeled data.

Method used

We construct teacher-student models, including teacher image segmentation networks and student image segmentation networks. Through consistency semi-supervised learning, we utilize a small amount of labeled data and a large amount of unlabeled data, combined with adversarial perturbations and consistency loss from different decoders, to optimize model parameters and improve segmentation accuracy and robustness.

Benefits of technology

Under limited annotation resources, this method significantly improves the efficiency and accuracy of medical image segmentation tasks, reduces data preparation costs, enhances the model's generalization ability and robustness, and is suitable for segmenting complex left atrial structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904642B_ABST
    Figure CN119904642B_ABST
Patent Text Reader

Abstract

The application discloses a consistency semi-supervised left atrium segmentation method based on a teacher-student model, and constructs a teacher-student model, wherein the teacher image segmentation network and the student image segmentation network each comprise one encoder and three decoders, the input of the encoder forms the input of the teacher image segmentation network or the student image segmentation network, and the output of the three decoders forms three outputs of the teacher image segmentation network or the student image segmentation network; the encoder and each decoder form a sub-segmentation network of a U-Net structure, and the intensity of the added adversarial disturbance in the feature transmission channels of different sub-segmentation networks is different. The application utilizes the strategy of effective segmentation of a small amount of labeled data and a large amount of unlabeled data, and through mutual consistency learning of adversarial disturbance and the teacher-student image segmentation network, the segmentation accuracy can be improved while reducing the labeling demand, so that the performance of the semi-supervised medical image segmentation task is effectively improved, and the training cost is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of nuclear magnetic resonance image segmentation, and particularly relates to a consistent semi-supervised left atrium segmentation method based on a teacher-student model. BACKGROUND

[0002] Atrial fibrillation is one of the most common persistent arrhythmias, affecting the quality of life and health of millions of people worldwide. Atrial fibrillation is characterized by a chaotic rhythm of atrial electrical signals and extremely fast frequency, which leads to ineffective pumping of the heart, thereby increasing the risk of heart failure and stroke. Blood flow abnormalities cause blood to stagnate in the atrium, forming blood clots that can enter the brain with the blood flow, causing a stroke. Therefore, atrial fibrillation not only poses a serious threat to the health of patients, but also places a huge burden on the medical system.

[0003] For patients with atrial fibrillation, accurate segmentation of atrial structure helps to assess the degree of atrial fibrosis and develop catheter ablation strategies. Due to the anatomical structure and functional characteristics of the left atrium, atrial fibrillation is more common in the left atrium, so it is of great significance to establish a left atrial segmentation model for patients to assist doctors in treating atrial fibrillation. Traditional medical image segmentation methods rely on manual operation, and professional radiologists or technicians need to manually label the atrial region in the image. This process not only consumes time and effort, but also the accuracy and repeatability of the results are limited by the operator's experience and professional level. In addition, manual segmentation methods are inefficient when dealing with large amounts of data, making it difficult to meet the needs of rapid clinical diagnosis. With the development of medical imaging technology, the amount of image data obtained has increased dramatically, making manual segmentation methods even more impractical. At the same time, manual segmentation methods are difficult to fully utilize the complex information contained in medical image data, which may lead to inaccurate diagnosis and treatment planning.

[0004] In recent years, deep learning technology has made significant progress in the field of medical image segmentation. By automatically learning feature representations in image data, deep learning models can achieve high-accuracy image segmentation, greatly improving the efficiency and accuracy of the segmentation task. In particular, in the segmentation of left atrial structure in patients with atrial fibrillation, deep learning methods can automatically identify and segment the atrial region by learning the morphological features of the atrium from a large amount of image data, providing important information for the diagnosis and treatment of atrial fibrillation. However, the training of deep learning models usually requires a large amount of labeled data, and in the field of medical images, such data is costly to obtain, requiring a lot of manpower and time. In addition, the quality of data labeling directly affects the performance of the model, and high-quality labeling often requires highly specialized knowledge. Therefore, how to effectively train deep learning models with limited labeled data and reduce the dependence on large amounts of labeled data has become a research hotspot.

[0005] In summary, the accurate diagnosis and treatment of atrial fibrillation have an urgent need for accurate atrial structure segmentation, and the advantages and disadvantages of traditional segmentation methods and deep learning techniques indicate the direction of future research, that is, to study a segmentation method that combines the powerful ability of deep learning and a small amount of labeled data to achieve efficient and accurate atrial structure segmentation. SUMMARY

[0006] The present application aims at the challenge of medical image segmentation under the condition of limited labeled data, and provides a consistency semi-supervised left atrial segmentation method based on teacher-student model.

[0007] To solve the above problems, the present application is realized by the following technical solutions:

[0008] The consistency semi-supervised left atrial segmentation method based on teacher-student model comprises the following steps:

[0009] Step 1, a teacher-student model composed of a teacher image segmentation network and a student image segmentation network is constructed, wherein the teacher image segmentation network and the student image segmentation network have the same structure and each comprise one encoder and three decoders, and the three decoders share one encoder, the input of the encoder forms the input of the teacher image segmentation network or the student image segmentation network, and the output of the three decoders forms three outputs of the teacher image segmentation network or the student image segmentation network; the encoder and each decoder form a sub-segmentation network of U-Net structure, that is, the input of the encoder forms the input of the sub-segmentation network, the output of the encoder is connected to the input of the decoder, the output of the decoder forms the output of the sub-segmentation network, and a feature transmission channel adding adversarial disturbance is arranged between the same resolution ports of the encoder and the decoder; the strength of the adversarial disturbance added by the feature transmission channels of different sub-segmentation networks is different, wherein the adversarial disturbance added by the feature transmission channel between the encoder and the first decoder is less than that added by the feature transmission channel between the encoder and the second decoder, and the adversarial disturbance added by the feature transmission channel between the encoder and the second decoder is less than that added by the feature transmission channel between the encoder and the third decoder;

[0010] Step 2, a left atrial image dataset is obtained, and after processing the left atrial image dataset, a labeled data sample set and an unlabeled data sample set are obtained;

[0011] Step 3, first, initialize the teacher-student model by using the labeled data sample set; then, input the labeled data sample set and the unlabeled data sample set into the initialized teacher-student model for consistency learning, calculate the consistency loss by comparing the output difference of the teacher image segmentation network and the student image segmentation network, and update the parameters of the teacher-student model by minimizing the consistency loss and the exponential moving average strategy; finally, use part of the labeled data samples in the labeled data sample set to further refine the parameters of the teacher-student model.

[0012] Step 4, use the independent left atrium image test set to evaluate the teacher-student model obtained in step 3, and use the sub-segmentation network with the best performance in the student image segmentation network as the final left atrium segmentation model.

[0013] Step 5, input the left atrium image to be segmented into the left atrium segmentation model obtained in step 4 to obtain the segmentation result.

[0014] In the above scheme, the encoder is composed of 8 convolution and activation modules, 4 down-sampling modules and a pyramid pooling module; the input of the first convolution and activation module forms the input of the encoder, the output of the first convolution and activation module is connected to the input of the second convolution and activation module, and the output of the second convolution and activation module is connected to the input of the first down-sampling module; the output of the first down-sampling module is connected to the input of the third convolution and activation module, the output of the third convolution and activation module is connected to the input of the fourth convolution and activation module, and the output of the fourth convolution and activation module is connected to the input of the second down-sampling module; the output of the second down-sampling module is connected to the input of the fifth convolution and activation module, the output of the fifth convolution and activation module is connected to the input of the sixth convolution and activation module, and the output of the sixth convolution and activation module is connected to the input of the third down-sampling module; the output of the third down-sampling module is connected to the input of the seventh convolution and activation module, the output of the seventh convolution and activation module is connected to the input of the eighth convolution and activation module, and the output of the eighth convolution and activation module is connected to the input of the fourth down-sampling module; the output of the fourth down-sampling module is connected to the input of the pyramid pooling module, and the output of the pyramid pooling module forms the output of the encoder; the input of the first convolution and activation module is used as the input of the first feature transmission channel, the input of the third convolution and activation module is used as the input of the second feature transmission channel, the input of the fifth convolution and activation module is used as the input of the third feature transmission channel, and the input of the seventh convolution and activation module is used as the input of the fourth feature transmission channel.

[0015] In the scheme, the decoder is composed of 8 convolution and activation modules, 4 up-sampling modules and 4 fusion modules; the input of the first up-sampling module forms the input of the decoder, the output of the first up-sampling module is connected to the input of the first convolution and activation module, the output of the first convolution and activation module is connected to the input of the second convolution and activation module, the output of the second convolution and activation module is connected to one input of the first fusion module, and the output of the first fusion module is connected to the input of the second up-sampling module; the output of the second up-sampling module is connected to the input of the third convolution and activation module, the output of the third convolution and activation module is connected to the input of the fourth convolution and activation module, the output of the fourth convolution and activation module is connected to one input of the second fusion module, and the output of the second fusion module is connected to the input of the third up-sampling module; the output of the third up-sampling module is connected to the input of the fifth convolution and activation module, the output of the fifth convolution and activation module is connected to the input of the sixth convolution and activation module, the output of the sixth convolution and activation module is connected to one input of the third fusion module, and the output of the third fusion module is connected to the input of the fourth up-sampling module; the output of the fourth up-sampling module is connected to the input of the seventh convolution and activation module, the output of the seventh convolution and activation module is connected to the input of the eighth convolution and activation module, the output of the eighth convolution and activation module is connected to one input of the fourth fusion module, and the output of the fourth fusion module forms the output of the decoder; the other input of the first fusion module is the output of the first feature transmission channel, the other input of the second fusion module is the output of the first feature transmission channel, the other input of the third fusion module is the output of the first feature transmission channel, and the other input of the fourth fusion module is the output of the first feature transmission channel.

[0016] In the scheme, the up-sampling module of the first decoder adopts bilinear interpolation, the up-sampling module of the second decoder adopts transpose convolution, and the up-sampling module of the third decoder adopts sub-pixel convolution.

[0017] In step 3, when the teacher-student model is subjected to consistency learning:

[0018] For the student image segmentation network, the parameters thereof are updated by adopting the minimum consistency loss back propagation, that is:

[0019]

[0020] For the teacher image segmentation network, the parameters thereof are updated by adopting the exponential moving average strategy, that is:

[0021] θ t ←α·θ t +(1-α)·θ s

[0022] wherein θ s represents the parameters of the student image segmentation network, and θ tdenotes the parameters of the teacher image segmentation network, denotes the learning rate, and denotes the weight of the historical parameters, denotes the gradient of the parameters of the student image segmentation network, and L denotes the consistency loss,

[0023]

[0024] where i denotes the current sample index, and j denotes the current output port index, denotes the output probability distribution of the student image segmentation network at the j output ports for the i-th sample, denotes the output probability distribution of the teacher image segmentation network at the j output ports for the i-th sample, and ‖*‖ 2 denotes the square of the Euclidean distance.

[0025] Compared with the prior art, the present application has the following characteristics:

[0026] (1) A teacher-student model composed of a teacher image segmentation network and a student image segmentation network is constructed, wherein the teacher image segmentation network and the student image segmentation network each include 1 encoder and 3 decoders, the 3 decoders share 1 encoder, and the encoder and each decoder form a sub-segmentation network of U-Net structure. The connection channels between the shared encoder and the 3 decoders are independent of each other, and there is no direct connection between the decoders. In the training process, although the decoders are independent, they can jointly act on the encoder through the back propagation algorithm to optimize the model parameters.

[0027] (2) The decoders of the 3 U-Net structure sub-segmentation networks of the teacher image segmentation network and the student image segmentation network adopt different up-sampling strategy combinations (transpose convolution, bilinear interpolation, sub-pixel convolution), which can combine the advantages of their respective decoders while weakening their respective shortcomings, so as to facilitate the consistency learning of the subsequent model and improve the model training effect.

[0028] (3) Different intensity of adversarial perturbations are added to the respective feature transmission channels of the 3 U-Net structure sub-segmentation networks of the teacher image segmentation network and the student image segmentation network. The introduced adversarial perturbations enhance the effectiveness of the training data by directly challenging the weaknesses of the model, forcing the model to learn more generalized feature representations. Different intensities of adversarial perturbations are added to different types of decoders, which can generate output differences containing more rich abstract information, thereby improving the generalization ability of the model and significantly improving the robustness of the model.

[0029] (4) The constructed teacher-student model can be trained by mixing a small amount of labeled data and a large amount of unlabeled data, thereby significantly reducing the dependence on a large amount of manually annotated data, reducing the cost of data preparation and obtaining data samples, and effectively improving the performance of the medical image segmentation task under the condition of limited annotation resources. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 A flowchart of the consistency semi-supervised left atrium segmentation method based on the teacher-student model.

[0031] Figure 2 The teacher-student model.

[0032] Figure 3 A schematic diagram of a sub-segmentation network of the U-Net structure. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application is further described in detail below with reference to specific examples and the accompanying drawings.

[0034] The consistency semi-supervised left atrium segmentation method based on the teacher-student model, as shown in Figure 1 , includes the following steps:

[0035] Step 1, constructing a teacher-student model.

[0036] The teacher-student model of the present application is composed of a teacher image segmentation network and a student image segmentation network, as shown in Figure 2 , wherein the structures of the teacher image segmentation network and the student image segmentation network are the same, and each includes one encoder and three decoders, and the three decoders share the same encoder, the input of the encoder forms the input of the teacher image segmentation network or the student image segmentation network, and the outputs of the three decoders form three outputs of the teacher image segmentation network or the student image segmentation network. The shared encoder is connected to the three decoders through independent connection channels, and the decoders are not directly connected to each other. In the training process, although each decoder is independent, it can act on the shared encoder through the back propagation algorithm, thereby optimizing the model parameters. The shared encoder and each decoder together constitute a sub-segmentation network based on the U-Net structure, and the feature transmission channels of each sub-segmentation network add different intensities of adversarial perturbations. The decoding network based on multi-feature map fusion can more effectively integrate feature information from different levels and improve the detail restoration ability of segmentation. The adversarial perturbation finds the direction most sensitive to the input change of the model output by calculating the gradient of the loss function of the model to the input image, that is, the direction that misleads the model to make a wrong judgment, and generates a small but meaningful adversarial perturbation according to the direction.

[0037] The sub-segmentation network of the present application draws lessons from the classic U-Net structure, the structure of which is shown in Figure 3 , that is, the input of the encoder forms the input of the sub-segmentation network, the output of the encoder is connected to the input of the decoder, the output of the decoder forms the output of the sub-segmentation network, and the feature transmission channels adding adversarial perturbations are arranged between the same resolution ports of the encoder and the decoder. The encoder part extracts features of different resolutions through multiple down-sampling layers, while the decoder part restores the resolution of the image by using multiple up-sampling layers, and fuses the features by using the skip connection of the feature transmission channel, so that the model can understand the data more deeply.

[0038] The above-mentioned encoder is composed of 8 convolution and activation modules, 4 down-sampling modules and a pyramid pooling module. The input of the first convolution and activation module forms the input of the encoder, the output of the first convolution and activation module is connected to the input of the second convolution and activation module, and the output of the second convolution and activation module is connected to the input of the first down-sampling module; the output of the first down-sampling module is connected to the input of the third convolution and activation module, the output of the third convolution and activation module is connected to the input of the fourth convolution and activation module, and the output of the fourth convolution and activation module is connected to the input of the second down-sampling module; the output of the second down-sampling module is connected to the input of the fifth convolution and activation module, the output of the fifth convolution and activation module is connected to the input of the sixth convolution and activation module, and the output of the sixth convolution and activation module is connected to the input of the third down-sampling module; the output of the third down-sampling module is connected to the input of the seventh convolution and activation module, the output of the seventh convolution and activation module is connected to the input of the eighth convolution and activation module, and the output of the eighth convolution and activation module is connected to the input of the fourth down-sampling module; the output of the fourth down-sampling module is connected to the input of the pyramid pooling module, and the output of the pyramid pooling module forms the output of the encoder. The input of the first convolution and activation module is taken as the input of the first feature transmission channel, the input of the third convolution and activation module is taken as the input of the second feature transmission channel, the input of the fifth convolution and activation module is taken as the input of the third feature transmission channel, and the input of the seventh convolution and activation module is taken as the input of the fourth feature transmission channel.

[0039] The decoder is composed of 8 convolution and activation modules, 4 up-sampling modules and 4 fusion modules. The input of the first up-sampling module forms the input of the decoder, the output of the first up-sampling module is connected to the input of the first convolution and activation module, the output of the first convolution and activation module is connected to the input of the second convolution and activation module, the output of the second convolution and activation module is connected to one input of the first fusion module, and the output of the first fusion module is connected to the input of the second up-sampling module; the output of the second up-sampling module is connected to the input of the third convolution and activation module, the output of the third convolution and activation module is connected to the input of the fourth convolution and activation module, the output of the fourth convolution and activation module is connected to one input of the second fusion module, and the output of the second fusion module is connected to the input of the third up-sampling module; the output of the third up-sampling module is connected to the input of the fifth convolution and activation module, the output of the fifth convolution and activation module is connected to the input of the sixth convolution and activation module, the output of the sixth convolution and activation module is connected to one input of the third fusion module, and the output of the third fusion module is connected to the input of the fourth up-sampling module; the output of the fourth up-sampling module is connected to the input of the seventh convolution and activation module, the output of the seventh convolution and activation module is connected to the input of the eighth convolution and activation module, the output of the eighth convolution and activation module is connected to one input of the fourth fusion module, and the output of the fourth fusion module forms the output of the decoder. Another input of the first fusion module is the output of the first feature transmission channel, another input of the second fusion module is the output of the first feature transmission channel, another input of the third fusion module is the output of the first feature transmission channel, and another input of the fourth fusion module is the output of the first feature transmission channel.

[0040] In view of the two decoders with different up-sampling strategies added on the basis of the classic U-Net in the present application, it is necessary to explain the design considerations. Although the transpose convolution used in the classic U-Net decoder can learn the effective parameters in the up-sampling process and is beneficial to the recovery of image details and textures, its inherent pixel arrangement non-uniformity may lead to checkerboard artifacts, and the parameter amount is relatively large and the calculation cost is high. Therefore, considering that subsequent consistency learning can fuse the advantages of different decoders and complement their shortcomings, the present application additionally adds two decoders which respectively use bilinear interpolation and sub-pixel convolution as the up-sampling strategy.

[0041] For bilinear interpolation: 1) Bilinear interpolation is a non-learning upsampling method, which is fast and low in computational cost. This makes the model more efficient in real-time or resource-limited applications. 2) Bilinear interpolation mainly estimates new pixel values by weighted average of adjacent pixel values. This method can naturally smooth the image and help provide better visual effects in some medical image processing tasks that require smooth edges and transitions. For the left atrium, which has a delicate boundary and structure, bilinear interpolation can provide a smoother image transition and avoid abrupt changes caused by upsampling, which is particularly important for accurate identification of cardiac structures. 3) Transposed convolution may cause a checkerboard effect when upsampling, which is caused by uneven pixel arrangement due to the convolution kernel and step size. In left atrium segmentation, this effect may cause false positives or false negatives in the segmentation results, affecting the accuracy of diagnosis. Bilinear interpolation does not have this problem, which ensures the continuity and naturalness of the image, which is very important for medical applications that require high-quality edge details. 4) Bilinear interpolation provides a non-learning upsampling method, and the consistency and stability of the output are higher than those of transposed convolution. In transposed convolution, insufficient model training or minor changes in data may result in significantly different output results, while bilinear interpolation can ensure consistent output under the same input conditions. At the same time, because bilinear interpolation does not depend on the parameters learned by the model, it can reduce the risk of overfitting to some extent. In specific applications, this can help the model better generalize to new, unseen data. Overall, in left atrium segmentation applications that require extremely high image quality, bilinear interpolation can make up for the corresponding shortcomings of transposed convolution in consistent learning with its smooth image quality, efficient computing performance, and excellent stability. This helps improve segmentation accuracy, reduces potential diagnostic errors, and speeds up processing, so the invention chooses bilinear interpolation as the upsampling method for one of the decoders.

[0042] For sub-pixel convolution: 1) Sub-pixel convolution increases the number of channels of the feature map, and then realizes up-sampling through a pixel rearrangement process. This method allows the model to increase the details and complexity of the data without directly increasing the spatial dimension, so that the image details and clarity can be better maintained after up-sampling. For left atrium segmentation, this means that the edges of the heart structure can be more accurately depicted, improving the accuracy of segmentation, especially in the key areas of the heart edge. 2) Sub-pixel convolution avoids the checkerboard effect through pixel rearrangement, providing smoother and continuous image features, which is beneficial for generating more accurate and natural segmentation results. 3) In sub-pixel convolution, up-sampling is realized by first performing deep convolution to increase the number of channels, and then performing pixel rearrangement. Such a process allows the model to make more full use of network parameters when performing feature extraction, enhancing the ability of feature extraction. This is particularly beneficial for the segmentation of complex areas such as the left atrium, as the model needs to accurately understand and analyze the complex structure of the heart. 4) Because sub-pixel convolution can more naturally restore image details, it helps the model to show better robustness when processing various medical images. Especially when faced with the complex and varied image features of the left atrium, it can more reliably produce high-quality segmentation results. In summary, the use of sub-pixel convolution in left atrium segmentation can bring higher image segmentation accuracy compared to traditional transpose convolution, reduce common problems such as checkerboard effect in image processing, and improve the computational efficiency and robustness of the model. It can make up for the shortcomings of transpose convolution in consistency learning. This helps to improve segmentation accuracy, reduce potential diagnostic errors, and speed up processing, so the invention chooses sub-pixel convolution as the up-sampling method of the other decoder.

[0043] Accordingly, the three decoders of the three sub-segmentation networks of the three U-Net structures of the teacher image segmentation network and the student image segmentation network of the present application respectively adopt different up-sampling strategies, wherein the up-sampling module of the first decoder adopts bilinear interpolation, the up-sampling module of the second decoder adopts transpose convolution, and the up-sampling module of the third decoder adopts sub-pixel convolution.

[0044] Since the accurate segmentation of left atrium is crucial for the diagnosis and treatment planning of heart disease, it is necessary to ensure that the segmentation model can perform highly reliable under various conditions. Adversarial perturbation can be regarded as a carefully designed input perturbation, which is designed to challenge the decision boundary of the model. It can effectively reveal the potential defects and deficiencies of the model by directly challenging the weaknesses of the model, so as to enhance the effectiveness of the training data. Adversarial perturbation can significantly improve the robustness of the model by revealing the vulnerability of the model when facing carefully designed input perturbations, compared with ordinary perturbations. Adversarial perturbation can force the model to learn more general feature representation by simulating input changes under edge data distribution or boundary conditions where the model is prone to make mistakes, so as to perform better when facing unknown data. These characteristics provide important clues for improving the model structure and training strategy. Therefore, in the specific medical image processing task of left atrium segmentation, the use of adversarial perturbation can bring a series of advantages and significance, especially in improving the robustness, accuracy and generalization ability of the segmentation model. For improving the robustness of the model, in the real world, medical images are often affected by various noises, such as different scanning devices, changes in operating conditions, etc. Adversarial perturbation can train the model to maintain its performance when facing real-world data, especially in the case of data quality. And because the left atrium may show significant morphological differences between different patients, through adversarial training, the model can learn more general feature representation, so as to better handle the variation in anatomical structure. For enhancing the generalization ability of the segmentation model, adversarial perturbation can simulate those cases that are not covered in the training data, helping the model to make accurate predictions when encountering new or rare cases, which is crucial for ensuring the wide applicability of the model in clinical applications. For improving the accuracy of the segmentation model, in the training process containing adversarial samples, the model needs to learn how to correctly segment the left atrium from complex or subtle changes in images. This high-difficulty training can make the model learn more deep and robust features, ultimately improving its accuracy in standard and challenging tasks.

[0045] Although composite perturbations can simulate a variety of real-world changes, as shown in the Chinese invention patent application CN116958544A, "Left atrium image segmentation method based on consistency learning", composite perturbations can improve the robustness of the model to natural disturbances, but adversarial perturbations focus on the most vulnerable parts of the model, providing some unique advantages. First, adversarial perturbations are more targeted than composite perturbations, which are specifically designed to mislead the model. They are calculated to precisely target the weaknesses of the model. This approach can effectively reveal the most unstable aspects of the model, helping developers to better understand and improve the model. Second, adversarial perturbations are also more efficient. Adversarial perturbations are usually generated through gradient optimization, which makes them more effective in testing and enhancing the model. In contrast, composite perturbations may include a series of non-targeted changes that, while increasing the general robustness of the model, may not be enough to touch the key weaknesses of the model. Third, adversarial perturbations can better improve the generalization ability of the model. Since adversarial perturbations are systematic and targeted, they can help the model learn deeper and more abstract feature representations during training, thereby improving the model's generalization ability on unseen data. This training approach forces the model to learn not only how to handle specific and obvious perturbations, but also to maintain stability in a wider range of adversarial scenarios. Finally, adversarial perturbations can provide a clearer optimization target than composite perturbations. When training a model using adversarial perturbations, the optimization target is very clear - to reduce the error on adversarial samples. This clear goal helps the direction of model training, while composite perturbations may involve multiple different directions of optimization, sometimes leading to scattered training focus. In summary, for the left atrium segmentation scenario, adversarial perturbations can better improve the robustness, accuracy, and generalization ability of the segmentation model compared to general random perturbations or even composite perturbations (a method that includes multiple different types of perturbations, such as random noise, blurring, brightness changes, etc.).

[0046] Accordingly, the feature transmission channel between the encoder and the decoder of the 3 U-Net structure sub-segmentation network of the teacher image segmentation network and the student image segmentation network of the present application adopts adversarial perturbations.

[0047] Since the teacher image segmentation network and the student image segmentation network increase two decoders with different upsampling methods based on the classic U-net, different intensity of adversarial perturbations is designed in the respective feature transmission channels for different upsampling structures, aiming to generate more diverse feature maps, so as to promote the model to learn more robust and generalizable feature representations. For the decoder using transpose convolution as upsampling: since the weights of transpose convolution are learned, it has a certain adaptability to small changes in input, but it will still be affected by large perturbations, especially when the perturbation affects the learned convolution kernel. Applying moderate intensity of adversarial perturbations to this decoder can help test and enhance its ability to handle more complex image structures, while avoiding performance degradation caused by excessive perturbations. Therefore, moderate intensity of adversarial perturbations is applied to the decoder using transpose convolution as upsampling. For the decoder using bilinear interpolation as upsampling: since bilinear interpolation directly depends on the values of adjacent pixels, small perturbations to these pixel values can directly affect the interpolation results, especially at the edges or details of the image. Applying small intensity of adversarial perturbations to this decoder can prevent excessive impact on the accuracy of interpolation, while still testing its robustness to minor changes. Therefore, small intensity of adversarial perturbations is applied to the decoder using bilinear interpolation as upsampling. For the decoder using subpixel convolution as upsampling: since subpixel convolution generally has high tolerance to input perturbations due to its operation method, it achieves high-resolution output by rearranging pixels, rather than directly relying on the precise values of surrounding pixels, which has high tolerance. Applying large perturbations to this decoder can help further strengthen the model's ability to handle complex input changes, ensuring that it can maintain good performance even under large input perturbations. Therefore, large intensity of adversarial perturbations is applied to the decoder using subpixel convolution as upsampling.

[0048] The shared encoder captures the core features and information of the input data and transmits these features and information to different decoders respectively, and adds adversarial perturbations of different intensities in the respective feature transmission channels, and the three decoders will receive features with different disturbances and output segmentation results with differences. This can test and enhance the robustness of each decoder, especially when they handle different types of image features (such as details of different scales). Each decoder can be trained to handle its specific task, such as one handling finer texture details and the other handling more macroscopic structures.

[0049] Compared to the approach adopted by the Chinese invention patent application "Left Atrial Image Segmentation Method Based on Consistency Learning," published under Publication No. CN116958544A, which applies perturbations to different layers of different decoders, the present invention applies perturbations to the entire decoder structure, enabling a more systematic assessment and improvement of the robustness of the decoder as a whole, avoiding the local optimization issues that can arise from focusing solely on inter-layer robustness. Furthermore, by applying perturbations of varying strengths to each decoder, the entire network can be made more robust when integrating information obtained from different decoders. This approach can better optimize the model's understanding and processing of complex image structures than simply applying perturbations to different layers. Furthermore, because consistency constraints are subsequently required, varying perturbations of varying strengths across different decoders strengthen the model's ability to maintain output consistency, further improving its accuracy and robustness, which is particularly important for left atrial segmentation. Furthermore, applying adversarial perturbations of varying strengths to different decoders provides greater flexibility, allowing the perturbation strength to be adjusted according to specific application requirements to maximize the model's performance in specific scenarios. This approach is highly scalable, facilitating future optimization and adaptation to a wider range of task requirements. In summary, applying adversarial perturbations of varying strength to different decoders can more effectively leverage the advantages of a multi-decoder architecture, specifically enhancing model robustness, optimizing the effects of consistency loss, and facilitating more refined model performance tuning. This strategy is particularly advantageous in left atrial segmentation tasks, which have complex structures and requirements.

[0050] Accordingly, the feature transmission channels of the three U-Net structured sub-segmentation networks of the teacher image segmentation network and the student image segmentation network of the present invention adopt adversarial perturbations of different intensities, wherein the adversarial perturbation added to the feature transmission channel between the encoder and the first decoder is smaller than the adversarial perturbation added to the feature transmission channel between the encoder and the second decoder, and the adversarial perturbation added to the feature transmission channel between the encoder and the second decoder is smaller than the adversarial perturbation added to the feature transmission channel between the encoder and the third decoder.

[0051] Step 2: Build sample data.

[0052] First, a 3D magnetic resonance imaging (MRI) cardiac dataset is acquired. The images in the left atrial image dataset can be sourced from existing data on the internet or acquired using MRI equipment. After processing the left atrial image dataset, i.e., through image preprocessing (e.g., normalization) and data augmentation (e.g., flipping, scaling, and / or random cropping), a portion of the images that have not been calibrated by a professional physician are directly used to form an unlabeled data sample set, while the remaining portion of the images that have been calibrated by a professional physician form a labeled data sample set.

[0053] Step 3, teacher-student model learning with sample data.

[0054] Step 3.1 initialization: initialize teacher-student model with labeled data sample set.

[0055] The encoder part of the teacher-student model is initialized with pre-trained weights, so that the encoder has good feature extraction capability. The decoder part adopts Xavier initialization, which can automatically adjust the initial distribution of weights according to the number of input and output neurons, thereby effectively alleviating the problems of gradient vanishing and gradient explosion, and helping the stability and efficiency of neural network training. Finally, the entire network is fine-tuned using the labeled left atrial image dataset.

[0056] Step 3.2 consistency learning: mix the labeled data sample set and the unlabeled data sample set as input to the initialized teacher-student model for consistency learning, calculate the consistency loss by comparing the output difference of the teacher image segmentation network and the student image segmentation network, and update the teacher-student model parameters using the minimum loss consistency loss backpropagation and EMA strategy.

[0057] Mix the labeled data sample set and the unlabeled data sample set as input data, input to the teacher-student model for consistency learning. In the early stage of training, prefer to use more labeled data samples for model training, and gradually increase the proportion of unlabeled data samples as the training progresses, to increase the weight of the impact of unlabeled data on performance. During consistency learning, the encoders of the teacher image segmentation network and the student image segmentation network of the teacher-student model capture the core features and information of the input data, and transmit these features and information to different decoders respectively, and add corresponding different intensity of adversarial perturbations in the respective feature transmission channels, to obtain multiple segmentation results. Calculate the mean square error of the output probability distribution of each student image segmentation network and all teacher image segmentation networks between the 3 results of the teacher image segmentation network and the student image segmentation network, i.e. the consistency loss. Calculate the consistency loss by comparing the output difference of the student image segmentation network and the teacher image segmentation network, first update the student image segmentation network parameters by minimizing the consistency loss backpropagation, and then update the teacher image segmentation network parameters using the EMA strategy. Repeat this process until all data training is completed.

[0058] For the student image segmentation network, the minimum consistency loss backpropagation is used to update its parameters. Let P t be the output probability distribution of the student image segmentation network, and N be the total number of data samples, then the mean square error consistency loss (L) can be defined as:

[0059]

[0060] where L denotes the consistency loss value calculated to measure the difference between the outputs of the student image segmentation network and the teacher image segmentation network. N denotes the total number of data samples, and i denotes the current sample index, which traverses all samples from 1 to N. j denotes the current output port index, which traverses all output ports 1, 2, 3. denotes the square of the Euclidean distance between the output probability distributions of the student image segmentation network and the teacher image segmentation network at the jth output port for the ith sample, which measures the difference between the output of the student image segmentation network at the port and the output of the teacher image segmentation network at the port.

[0061] During the training process, our goal is to minimize this loss L. Specifically, the gradient of the loss function L with respect to the parameters of the student image segmentation network is calculated, and the parameters of the student image segmentation network are updated:

[0062]

[0063] where θ s denotes the parameters of the student image segmentation network. η is the learning rate, which is a positive scalar that determines the step size of each parameter update. denotes the gradient of the parameters of the student image segmentation network, which indicates the direction in which the loss function grows the fastest, and the negative direction is the direction in which the loss function decreases the fastest. is the gradient of the loss function L with respect to the parameters θ s of the student image segmentation network.

[0064] For the teacher image segmentation network, the EMA (Exponential Moving Average) strategy is used to update its parameters, ensuring that the update of the teacher image segmentation network parameters is more smooth:

[0065] θ t ← α · θ t + (1 - α) · θ s

[0066] where θ t is the parameter of the teacher image segmentation network. α is a coefficient close to 1, representing the weight of the historical parameters. It controls the weight of the historical parameters when updating the teacher network parameters.

[0067] In this way, the model ensures that the student image segmentation network can learn stable and accurate knowledge from the teacher image segmentation network by minimizing consistency loss, thereby improving the performance of the student image segmentation network. Compared with traditional consistency learning, the teacher-student model consistency learning increases the stability of the model output through the exponential moving average (EMA) mechanism. The EMA strategy helps to smooth the parameter update process of the teacher image segmentation network, thereby suppressing potential fluctuations in the training process and improving the consistency and stability of the teacher network output. This smoothed output is more reliable when generating pseudo-labels and can reduce noise interference during the learning process of the student network. This smoothed output is particularly important in left atrium segmentation, as it can ensure that the model maintains good performance even in the face of individual differences and technical biases in the dataset. Traditional consistency learning, on the other hand, usually relies on multiple outputs of a single model (or multiple models handling the same task), which are directly compared through a consistency loss to encourage the model output to remain consistent when subjected to slight perturbations. While this also helps to improve robustness, the model output can be more susceptible to random fluctuations in the training process without the help of the EMA mechanism. Meanwhile, in terms of the use of unlabeled data, in the teacher-student model, the output of the teacher image segmentation network (based on the EMA smoothed parameters of the student image segmentation network) is used to generate pseudo-labels, which are fed back to the student image segmentation network as consistency targets. This not only enhances the guiding role of the teacher image segmentation network, but also enables the student image segmentation network to learn effectively without explicit labeling. For left atrium segmentation, this means that the model can better learn from a large number of unlabeled medical images in the absence of sufficient labeled data, improving the accuracy and consistency of its predictions. Traditional consistency learning, on the other hand, usually enforces consistency between different outputs of a model or between inputs generated through data augmentation. This approach can also utilize unlabeled data, but lacks a teacher image segmentation network that smooths and stabilizes the output to guide the learning process. Meanwhile, the combination of smoothing and consistency learning provided by EMA enables the teacher-student model to better generalize to new, unseen data. In left atrium segmentation, this helps the model to handle image data from different patients, different devices, or obtained under different conditions, ensuring the accuracy and reliability of the segmentation results.

[0068] Step 3.3, fully supervised training: using part of the labeled data samples in the labeled data sample set to conduct fully supervised training on the consistency-learned teacher-student model, further refining the teacher-student model parameters, applying semantic constraints, and using early stopping to monitor the training process and save the optimal network parameters to improve the model segmentation accuracy;

[0069] Step 4, the teacher-student model obtained in step 3 is evaluated using an independent left atrium image test set to verify its performance, and the sub-segmentation network with the best performance in the student image segmentation network is taken as the final left atrium segmentation model.

[0070] In order to test the generalization ability of the obtained network, the test set is used for verification, and the performance indicators of the segmentation results are Dice coefficient, Jaccard coefficient, 95% Hausdorff distance (95% HD) and average surface distance ASD.

[0071] The Dice coefficient mainly evaluates the accuracy of segmentation by measuring the set similarity, and the formula of the Dice coefficient is as follows:

[0072]

[0073] Wherein, A and B represent the real data label of left atrium and the predicted left atrium region respectively.

[0074] The Jaccard coefficient mainly evaluates the accuracy of segmentation by measuring the degree of overlap of the region, and the Jaccard coefficient is defined as follows:

[0075]

[0076] Wherein, A and B represent the real data label of left atrium and the predicted left atrium region respectively.

[0077] 95% HD is evaluated by calculating the distance of the segmentation boundary, and the definition of 95% HD is as follows:

[0078] HD 95 (A,B)=max(h 95 (A,B),h 95 (B,A))

[0079] Wherein, A represents the point set of the left atrium boundary in the predicted segmentation result, and B represents the point set of the left atrium boundary in the real segmentation label.

[0080] The average surface distance ASD is an index for measuring the difference between two surfaces, and the definition of ASD is as follows:

[0081]

[0082] Wherein, A represents the point set of the left atrium boundary predicted by the segmentation model, a∈A: indicates a specific point in the point set A, which represents a point on the left atrium boundary predicted by the model. B represents the point set of the real left atrium boundary, b∈B: indicates a specific point in the point set B, which represents a point on the real left atrium boundary.

[0083] Step 5, input the left atrium image to be segmented into the left atrium segmentation model obtained in step 4 to obtain a segmentation result.

[0084] The present application designs an average teacher segmentation network composed of a shared encoder and three slightly different independent decoders, and sets different intensities of adversarial noise according to the sensitivity of different decoders. During the training of the average teacher segmentation network, the teacher-student model is first initialized with a labeled data sample set; then the mixed labeled data sample set and unlabeled data sample set are input into the teacher-student model for consistency learning; then part of the labeled data samples in the labeled data sample set are used to perform full supervision training on the teacher-student model to further refine the parameters of the teacher-student model. Then, the trained teacher-student model is evaluated using an independent test set, and the sub-segmentation network with the best performance is used as the final left atrium segmentation model to segment the left atrium image to be segmented. The present application uses the strategy of effective segmentation with a small amount of labeled data and a large amount of unlabeled data, and through mutual consistency learning of adversarial disturbance and teacher-student image segmentation network, it can reduce the labeling demand while improving the segmentation accuracy, thereby effectively improving the performance of the semi-supervised medical image segmentation task and greatly reducing the training cost.

[0085] It should be noted that although the above embodiments of the present application are illustrative, this is not a limitation of the present application, therefore the present application is not limited to the above specific embodiments. Any other embodiments obtained by those skilled in the art under the inspiration of the present application without departing from the principles of the present application are considered to be within the protection scope of the present application.

Claims

1. A consistency-based semi-supervised left atrium segmentation method based on teacher-student model, characterized in that, The steps comprise the following: Step 1, constructing a teacher-student model composed of a teacher image segmentation network and a student image segmentation network, wherein the teacher image segmentation network and the student image segmentation network have the same structure and each comprise one encoder and three decoders, the upsampling module of the first decoder adopts bilinear interpolation, the upsampling module of the second decoder adopts transpose convolution, the upsampling module of the third decoder adopts sub-pixel convolution, and the three decoders share one encoder, the input of the encoder forms the input of the teacher image segmentation network or the student image segmentation network, and the outputs of the three decoders form three outputs of the teacher image segmentation network or the student image segmentation network; the encoder and each decoder form a sub-segmentation network of a U-Net structure, that is, the input of the encoder forms the input of the sub-segmentation network, the output of the encoder is connected to the input of the decoder, the output of the decoder forms the output of the sub-segmentation network, and a feature transmission channel adding adversarial disturbance is arranged between the ports of the same resolution of the encoder and the decoder; the strengths of the adversarial disturbances added by the feature transmission channels of different sub-segmentation networks are different, wherein the adversarial disturbance added by the feature transmission channel between the encoder and the first decoder is smaller than the adversarial disturbance added by the feature transmission channel between the encoder and the second decoder, and the adversarial disturbance added by the feature transmission channel between the encoder and the second decoder is smaller than the adversarial disturbance added by the feature transmission channel between the encoder and the third decoder; Step 2, obtaining a left atrium image dataset, and obtaining a labeled data sample set and an unlabeled data sample set after processing the left atrium image dataset; Step 3, first initializing the teacher-student model using the labeled data sample set; then inputting the labeled data sample set and the unlabeled data sample set into the initialized teacher-student model for consistency learning, calculating a consistency loss by comparing the output differences of the teacher image segmentation network and the student image segmentation network, and updating the parameters of the teacher-student model by using the minimum loss consistency loss back propagation and the exponential moving average strategy; then performing full supervision training on the consistency-learned teacher-student model using part of the labeled data samples in the labeled data sample set to further refine the parameters of the teacher-student model; Step 4, evaluating the teacher-student model obtained in step 3 using an independent left atrium image test set, and taking the sub-segmentation network with the best performance in the student image segmentation network as the final left atrium segmentation model; Step 5, inputting a left atrium image to be segmented into the left atrium segmentation model obtained in step 4 to obtain a segmentation result.

2. The teacher-student model based consensus semi-supervised left atrium segmentation method according to claim 1, characterized in that, The encoder is composed of 8 convolution and activation modules, 4 down-sampling modules and a pyramid pooling module. The input of the first convolution and activation module forms the input of the encoder, the output of the first convolution and activation module is connected to the input of the second convolution and activation module, and the output of the second convolution and activation module is connected to the input of the first downsampling module; the output of the first downsampling module is connected to the input of the third convolution and activation module, the output of the third convolution and activation module is connected to the input of the fourth convolution and activation module, and the output of the fourth convolution and activation module is connected to the input of the second downsampling module; the output of the second downsampling module is connected to the input of the fifth convolution and activation module, the output of the fifth convolution and activation module is connected to the input of the sixth convolution and activation module, and the output of the sixth convolution and activation module is connected to the input of the third downsampling module; the output of the third downsampling module is connected to the input of the seventh convolution and activation module, the output of the seventh convolution and activation module is connected to the input of the eighth convolution and activation module, and the output of the eighth convolution and activation module is connected to the input of the fourth downsampling module; the output of the fourth downsampling module is connected to the input of the pyramid pooling module, and the output of the pyramid pooling module forms the output of the encoder; The input of the first convolution and activation module is used as the input of the first feature transmission channel, the input of the third convolution and activation module is used as the input of the second feature transmission channel, the input of the fifth convolution and activation module is used as the input of the third feature transmission channel, and the input of the seventh convolution and activation module is used as the input of the fourth feature transmission channel.

3. The teacher-student model based consensus semi-supervised left atrium segmentation method according to claim 1, characterized in that, The decoder consists of 8 convolution and activation modules, 4 upsampling modules and 4 fusion modules; The input of the first upsampling module forms the input of the decoder, the output of the first upsampling module is connected to the input of the first convolution and activation module, the output of the first convolution and activation module is connected to the input of the second convolution and activation module, the output of the second convolution and activation module is connected to an input of the first fusion module, and the output of the first fusion module is connected to the input of the second upsampling module; the output of the second upsampling module is connected to the input of the third convolution and activation module, the output of the third convolution and activation module is connected to the input of the fourth convolution and activation module, the output of the fourth convolution and activation module is connected to an input of the second fusion module, and the output of the second fusion module is connected to the input of the third upsampling module; the output of the third upsampling module is connected to the input of the fifth convolution and activation module, the output of the fifth convolution and activation module is connected to the input of the sixth convolution and activation module, the output of the sixth convolution and activation module is connected to an input of the third fusion module, and the output of the third fusion module is connected to the input of the fourth upsampling module; the output of the fourth upsampling module is connected to the input of the seventh convolution and activation module, the output of the seventh convolution and activation module is connected to the input of the eighth convolution and activation module, the output of the eighth convolution and activation module is connected to an input of the fourth fusion module, and the output of the fourth fusion module forms the output of the decoder; Another input of the first fusion module is used as the output of the first feature transmission channel, another input of the second fusion module is used as the output of the first feature transmission channel, another input of the third fusion module is used as the output of the first feature transmission channel, and another input of the fourth fusion module is used as the output of the first feature transmission channel.

4. The teacher-student model based consensus semi-supervised left atrium segmentation method according to claim 1, characterized in that, In step 3, when the teacher-student model is consistent learning: For the student image segmentation network, the parameters are updated by minimizing the consistency loss backpropagation, that is: For the teacher image segmentation network, the parameters are updated by using the exponential moving average strategy, that is: θ t ← α · θ t + (1 - α) · θ s where θ s denotes the parameters of the student image segmentation network, θ t denotes the parameters of the teacher image segmentation network, η denotes the learning rate, and α denotes the weight of the historical parameters, denotes the gradient of the parameters of the student image segmentation network, and L denotes the consistency loss, where i represents a current sample index, j represents a current output port index, denotes an output probability distribution of the student image segmentation network at the j output ports for the i-th sample, denotes an output probability distribution of the teacher image segmentation network at the j output ports for the i-th sample, 2 denotes a square of the Euclidean distance, and N represents a total number of data samples.

Citation Information

Patent Citations

  • Left atrium image segmentation method based on consistency learning

    CN116958544A

  • Semi-supervised medical image segmentation method, system, equipment and medium

    CN117095014A

  • Image steganography utilizing adversarial perturbations

    WO2022241307A1