A bias-variance decomposition segmentation method of medical images

CN117351210BActive Publication Date: 2026-09-29JIANGXI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311467546.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-07
Publication Date
2026-09-29
Estimated Expiration
2043-11-07

AI Technical Summary

Technical Problem

然而,异方差噪声的估计常常不稳定,导致数据不确定性的评估不可靠

Benefits of technology

[0035](1)本发明致力于通过偏差-方差分解理论计算方差对教师预测的数据不确定性进行建模来进行有效的知识迁移。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117351210B_ABST
    Figure CN117351210B_ABST
Patent Text Reader

Abstract

The application discloses a bias-variance decomposition segmentation method of a medical image, and comprises the following steps: acquiring a medical image; inputting the medical image into an image segmentation model to acquire a segmented medical image, wherein the image segmentation model adopts a teacher network and a student network, and bias and variance in a distillation process of the teacher network and the student network are decoupled by using a bias-variance theory. On the basis of finding that unreliable modeling of data uncertainty is caused by bias-variance coupling in a knowledge distillation process, the application is committed to decoupling bias and variance in the knowledge distillation, so that a more optimal student network segmentation model is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image semantic segmentation technology based on knowledge distillation, and particularly relates to a bias-variance decomposition segmentation method for medical images. Background Technology

[0002] While deep learning has demonstrated excellent performance in medical image segmentation, its high computational cost makes deploying most deep neural networks on portable embedded devices a significant challenge. Knowledge distillation, with its potential for model compression, has attracted increasing attention in the field of medical image segmentation. Maximizing the mutual information between teacher and student networks is central to knowledge distillation, aiming to achieve knowledge transfer by having students estimate the teacher's distribution. However, teacher and student distributions are not always consistent, and the knowledge provided by teachers can sometimes be unreliable, for example, containing spurious labels. Noise introduced by teachers can introduce data uncertainty, which, if transferred indiscriminately to students, may interfere with their learning. Knowledge distillation requires active student participation in exploring the information provided by the teacher. Furthermore, students need the ability to identify any knowledge interference in the teacher's network.

[0003] Directly maximizing the mutual information between teachers and students can be challenging. To address this, Variational Information Distillation (VID) introduces a variational distribution aimed at maximizing the variational lower bound. VID typically assumes the use of heteroscedastic mean and homoscedastic noise for the variational distribution. However, experiments show that considering heteroscedastic noise leads to unstable training results. Existing research utilizes heteroscedastic mean and heteroscedastic noise to construct the variational distribution, where both teacher and student networks use the same decoder, and the weights of the student decoder are initialized by the teacher network. The final results show that this strategy performs well in super-resolution tasks. A strategy for modeling data uncertainty in unsupervised semantic segmentation domain adaptation has been proposed. To improve pseudo-label learning, it is suggested to combine an auxiliary classifier that helps predict variance. Using the predicted data uncertainty, the Auxiliary Classifier Model (ACM) can achieve satisfactory results. However, the estimation of heteroscedastic noise is often unstable, leading to unreliable assessments of data uncertainty. Researching reliable modeling of data uncertainty and its application in knowledge distillation will be an important and highly sought-after research area. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a bias-variance decomposition segmentation method for medical images, which can avoid unreliable evaluations caused by data uncertainty.

[0005] To achieve the above objectives, the present invention provides a bias-variance decomposition segmentation method for medical images, comprising:

[0006] Acquiring medical images;

[0007] The medical image is input into the image segmentation model to obtain the segmented medical image. The image segmentation model uses a teacher network and a student network, and the bias and variance in the distillation process of the teacher network and the student network are decoupled using bias-variance theory.

[0008] Optionally, the process of obtaining the image segmentation model includes:

[0009] The corresponding first predicted image is obtained using the teacher network;

[0010] Obtain the predicted expectation image of the student network;

[0011] Based on the first predicted image and the predicted expected image, the bias-variance decomposition distillation loss is obtained;

[0012] The bias-variance decomposition distillation loss is optimized and trained using a balance factor until it converges, thus obtaining the image segmentation model.

[0013] Optionally, before obtaining the corresponding first predicted image using the teacher network, the process includes: pre-training the teacher network.

[0014] Optionally, pre-training the teacher network includes:

[0015] The initial medical image is deformed using a random mask to obtain the deformed medical image.

[0016] The deformed medical image is used to perform feature learning using an encoder, and the medical image after feature learning is used to perform image restoration using a decoder to obtain the restored medical image;

[0017] Based on the initial medical image and the restored medical image, the reconstruction loss is calculated;

[0018] Based on the reconstruction loss, after obtaining the initial medical image and the restored medical image, the teacher network is trained until the initial medical image and the restored medical image are consistent, thus completing the pre-training of the teacher network.

[0019] Optionally, obtaining the predicted expectation image of the student network includes: obtaining a second predicted image using the student network, combining it with the expectation estimation module to predict the student's expectation, and obtaining the predicted expectation image of the student network.

[0020] Optionally, combining the expectation estimation module to predict students' expectations, obtaining the predicted expectation image of the student network includes:

[0021] The second predicted image is masked using a random mask to obtain mask features;

[0022] The mask features are reconstructed using 3×3 convolution and activation functions to obtain the reconstructed features;

[0023] Based on the reconstructed features, the predicted expected image of the student network is obtained.

[0024] Optionally, obtaining the bias-variance decomposition distillation loss includes:

[0025] The deviation is obtained based on the first predicted image and the second predicted image;

[0026] Based on the second predicted image and the expected predicted image, the variance is obtained;

[0027] Based on the aforementioned bias and variance, the bias-variance decomposition distillation loss is obtained.

[0028] Optionally, the bias-variance decomposition distillation loss is optimized and trained using a balance factor until it converges, including:

[0029] The bias-variance decomposition distillation loss is optimized by using a balance factor to obtain the optimized bias-variance decomposition distillation loss.

[0030] By introducing preset coefficients, the bias-variance decomposition distillation loss during optimization training converges, and the optimal bias-variance decomposition distillation loss is obtained to complete the training of the image segmentation model.

[0031] Optionally, the optimal bias-variance decomposition distillation loss is:

[0032]

[0033] Among them, L kd The deviation-variance decomposition distillation loss value is given, and α is a preset coefficient. For the deviation, exp(-var kd ) is the balance factor, var kd Let V be the variance and E be the expected value.

[0034] Compared with the prior art, the present invention has the following advantages and technical effects:

[0035] (1) This invention aims to effectively transfer knowledge by modeling the uncertainty of teachers’ predicted data through variance calculation using the bias-variance decomposition theory.

[0036] (2) Based on the discovery that the unreliable modeling of data uncertainty is attributed to the bias-variance coupling in the knowledge distillation process, this invention is committed to decoupling the bias and variance in knowledge distillation, thereby obtaining a better student network segmentation model.

[0037] (3) When faced with a dilemma in minimizing bias and variance, this invention proposes to incorporate a balancing factor to control the overall loss consisting of bias and variance, while making full use of teachers’ information and reducing the transmission of incorrect information to students. Attached Figure Description

[0038] The accompanying drawings, which constitute a part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0039] Figure 1 This is a flowchart illustrating a medical image segmentation method according to an embodiment of the present invention;

[0040] Figure 2 This is a flowchart illustrating the expected estimation module according to an embodiment of the present invention;

[0041] Figure 3 This is a schematic diagram of the teacher pre-training process according to an embodiment of the present invention;

[0042] Figure 4 This is a schematic diagram of the bias-variance decomposition knowledge distillation network in an embodiment of the present invention. Detailed Implementation

[0043] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0044] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0045] In maximizing mutual information, "pseudo-labels" present in the teacher network can cause data perturbation. To mitigate the impact of this perturbation, it is necessary to model the inherent data uncertainty in the teacher network. This invention proposes that the unreliable modeling of data uncertainty can be attributed to bias-variance coupling in the knowledge distillation process. On the one hand, it is desirable to make the student distribution network as similar as possible to the teacher network, aiming to maintain consistency between the student and teacher networks. This process is related to bias learning and can be viewed as the ability of students to fit teacher predictions. On the other hand, identifying the differences between the teacher and student networks is crucial when maximizing mutual information, with the aim of identifying ambiguities in the information provided by the teacher. This process, related to variance learning, can be viewed as perceiving the information perturbation inherent in the teacher network. Bias-variance decomposition theory can decouple bias and variance in knowledge distillation, thereby enabling the modeling of reliable data uncertainty in the teacher network. Defining expectation is key to obtaining bias and variance; this invention proposes using an expectation estimator to simulate expectation calculation.

[0046] Once biased learning reaches a certain level, simultaneously reducing both bias and variance becomes problematic. That is, minimizing bias typically increases variance, and vice versa. To achieve a balance in the trade-off between minimizing bias and variance, this invention incorporates a balancing factor to weigh biased learning against variance learning. The aim is to control the overall loss comprised of bias and variance, while fully utilizing information in the teacher network and reducing the likelihood of transmitting misleading information to the student network.

[0047] This invention provides a bias-variance decomposition segmentation method for medical images, the specific steps of which are as follows: Figure 1 As shown, including

[0048] Acquiring medical images;

[0049] The medical image is input into the image segmentation model to obtain the segmented medical image. The image segmentation model uses a teacher network and a student network. The bias and variance in the distillation process of the teacher network and the student network are decoupled using bias-variance theory.

[0050] The BENIN core of the bias-variance decomposition knowledge distillation network proposed in this invention, such as... Figure 4 As shown, it includes:

[0051] 1) First, obtain the pre-trained teacher model.

[0052] The key to constructing the proxy task lies in designing a reasonable medical image transformation, which is performed using a method of applying masks to organs. This strategy transforms the medical image by applying two forms of random masks: an outer mask and an inner mask. If a random mask is applied outside the selected window, it is called an outer mask. The inner mask is applied inside the selected window. Subsequently, an encoder-decoder architecture is used to reconstruct the original input, such as... Figure 3 As shown. Finally, when the teacher completes the pre-training, the learned representational knowledge with anatomical consistency can be transferred to the student network.

[0053] 2) Secondly, design an expectation estimator to simulate expectation calculation.

[0054] Specifically, suppose the student's prediction is f s ∈R C×H×W The teacher's output is y t ∈R C×H×W The expectation is represented as E(f) s First, predict f for the students. s Apply a random mask to obtain mask features Then, the expectation estimator is used to obtain the reconstructed features E(f). s The expected value predicted by the student is then represented by the reconstructed features. The EEM consists of two 3×3 convolutions and a ReLU activation function between them.

[0055] Based on bias-variance decomposition theory, variance learning is used to model the uncertainty of reliable data. In this study, an expectation estimator is designed to simulate expectation calculation, with the aim of understanding the variance var. kd Modeling is performed.

[0056] 3) Then, obtain the final bias-variance decomposition distillation loss.

[0057] This invention will use bias learning This is represented as the process by which the student's output matches the teacher's prediction, and the variance learning var is used. kd This is represented as a measure of the inherent data uncertainty for teachers in medical image segmentation. The process of maximizing mutual information involves... and var kd The dilemma of not being able to minimize both simultaneously. To address this issue, a balance factor exp(-var) is introduced. kd To regulate the training process and var kd The overall loss. Introduce another coefficient α∈[0,1], which is a monotonically decreasing function (for simplicity, a linear function is used), gradually decreasing as the training process progresses. Loss transition to loss

[0058] Bias-variance decomposition theory enables this invention to independently measure bias and variance in the knowledge distillation process. By balancing the trade-off between bias and variance in medical image segmentation, a balance factor exp(-var) is introduced. kd This causes distillation losses. Gradual convergence aims to address the challenges of bias and variance learning during the distillation process.

[0059] 4) By applying the above theory to classical medical image semantic segmentation and knowledge distillation networks, the bias-variance decomposition knowledge distillation network BENIN proposed in this invention can be obtained.

[0060] Bias-variance decomposition and data uncertainty in learning algorithms: Assuming the sample set is X and the target set is Y, the learning algorithm of this invention is represented by f, Y = f(X) + ∈(X), where ∈(·) represents the error. In the deep learning paradigm (taking regression as an example), it is often necessary to minimize the error between the sample x∈X and the target y∈Y, as shown in formula (1).

[0061] ∈(x)=E[(yf(x)) 2 (1)

[0062] Where ∈(·) represents the error, y is the learning objective, and f(x) is the learning algorithm used.

[0063] Further decomposing formula (1) yields formula (2).

[0064] ∈(x)=E[(yE(f(x))+E(f(x))-f(x))] 2

[0065] =E[yE(f(x))] 2 +E[(f(x)-E[f(x)]) 2 ]

[0066] =bias 2 +variance (2)

[0067] Where E(f(x)) is the expectation of the learning algorithm, and bias 2 Here, is the bias form of bias learning, and variance is the variance formula for variance learning.

[0068] From formula (2), the expectation E(f(x)) of the learning algorithm is expressed. The difference between the target and the expectation is called the bias. The bias is obtained through bias learning, as shown in formula (3).

[0069] bias 2=E[yE(f(x))] 2 (3)

[0070] The goal of bias learning is to drive the learning algorithm to make its predictions as close as possible to the target y. In other words, it evaluates the learning algorithm's ability to fit its target.

[0071] Similarly, the variance formula (4) for variance learning can be derived from formula (2).

[0072] variance=E[(f(x)-E[f(x)]) 2 (4)

[0073] Variance can be used to assess target interference, thus representing the data uncertainty implicit in annotated labels. In knowledge distillation, the predictions of teacher networks are often influenced by pseudo-labels; therefore, modeling the data uncertainty of teacher predictions by calculating variance is crucial for effective knowledge transfer. Estimation of data uncertainty can be achieved by building a network module that directly predicts noise. It is suggested to use an auxiliary classifier to predict variance and estimate the data uncertainty derived from pseudo-labels. However, it is worth noting that estimates of data uncertainty are generally considered unreliable. In this invention, based on bias-variance decomposition theory, variance learning is employed to model reliable data uncertainty. How to define expectation is key to obtaining bias and variance. This invention aims to design an expectation estimator to simulate expectation calculation.

[0074] Knowledge distillation, and various deep neural network models, have been widely applied to medical image segmentation. However, these methods inevitably increase the cost of introducing various components and the amount of storage space required. Therefore, a large number of lightweight networks have been proposed and applied to real-time semantic segmentation. However, when the network is simplified to speed up inference, its performance sometimes drops significantly, thus greatly affecting the practical application of this method.

[0075] Knowledge distillation, due to its powerful model compression capabilities, has attracted significant academic interest in the field of medical image segmentation. It attempts to transfer knowledge from a trained teacher network to a lightweight student network. Logit-based distillation transforms the logit activations of each channel into a probabilistic shape, allowing the use of probability distribution metrics such as KL divergence to measure the difference between the student and teacher distributions. Feature-based knowledge distillation distills features from the teacher's intermediate layers to the student. MGD proposes extracting knowledge by reconstructing the teacher's features. MGD first uses the MAE principle to randomly mask a certain number of regions in the student's output, then develops a simple generative module to encourage the student to recover the teacher's feature information, thus significantly improving the representational power of the student network. Research indicates that one reason for the effectiveness of knowledge distillation is the regularization provided by soft labels. Based on this approach, this paper rethinks soft labels and investigates the trade-off between supervised tasks and knowledge distillation tasks from the student. The proposed method uses weighted soft labels to handle the bias-variance game in the distillation process.

[0076] Existing methods cannot effectively address the problem of unreliable data uncertainty caused by misleading information contained in the teacher's input. The purpose of this invention is to represent bias learning as the process by which the student's output aligns with the teacher's prediction, and to represent variance learning as a measure of the inherent data uncertainty of the teacher in medical image segmentation. The aim is to distill consistent segmentation knowledge for the student while eliminating interference from unreliable segmentation information. Furthermore, to address the dilemma of bias and variance learning during the distillation process, a trade-off is imposed between bias and variance in medical image segmentation, and a balance factor is introduced to gradually bring the distillation loss to a convergent state.

[0077] Medical image semantic segmentation based on self-supervised learning aims to classify the pixels of an image using models such as FCN, U-Net, and DeepLab. Because these tasks require high accuracy, supervised learning is the most popular paradigm for medical image segmentation. However, training medical image diagnostic models using supervised learning often requires manual annotation of medical data by experts. This process can be costly and may lead to inconsistencies in annotation methods due to different experts. Self-supervised learning is a method that learns feature representations by leveraging effective and challenging proxy tasks, without relying on expert-annotated data. This approach is particularly well-suited for performing medical image diagnostic tasks.

[0078] The general approach of self-supervised learning based on classification, contrastive learning, or reconstruction is to transform each original medical image and then initiate an agent task that can recover each image transformation. It is proposed to first learn contextual information from the medical image by applying external or internal masks to organs, thereby generating meaningful image transformations. Then, an encoder-decoder architecture is used for image recovery. Finally, a pre-trained model with better representations is employed to transfer superior knowledge to the downstream task (i.e., the student).

[0079] One of the most significant challenges in applying knowledge distillation to medical image segmentation is the lack of pre-trained teachers. This challenge is exacerbated by the scarcity of medical image data and the time-consuming and laborious process of lesion labeling. Fortunately, different tasks in medical image diagnosis and the various modalities involved often exhibit anatomical consistency, enabling the development of effective self-supervised proxy tasks. These proxy tasks can explore the teacher's generalized representations, thereby improving the level of knowledge transfer from students. This invention combines self-supervised learning with knowledge distillation to enhance the quality of knowledge transfer from pre-trained teachers.

[0080] 1. In this invention, based on the bias-variance decomposition theory, variance learning is used to model reliable data uncertainty. How to define expectation is key to obtaining bias and variance. This embodiment focuses on designing an expectation estimator to simulate expectation calculation.

[0081] Maximizing mutual information can be formulated as maximizing formula (1).

[0082] I(t;s)=E t [-logp(t)]-E t,s [-logp(t|s)] (5)

[0083] Where p(t) represents the teacher distribution, and p(t|s) is the conditional distribution of teachers predicted given students. In E t With [-logp(t)] constant, maximizing I(t; s) is equivalent to minimizing E. t,s [-logp(t|s)], because the teacher's prediction is known. Typically, to minimize E... t,s[-logp(t|s)] often requires minimizing the ∈(x) error of a sample x∈X with a target y∈Y (as shown in Equation (1)). However, teacher predictions often contain noise that is considered as spurious labels. When maximizing mutual information, the estimation of heteroscedastic noise in the variational distribution can sometimes be unstable, leading to incorrect measurement of data uncertainty. This invention utilizes bias-variance decomposition theory to obtain the bias formula (as shown in Equation (3)) and the variance formula (as shown in Equation (4)). Since expectation calculation acts as a bridge between student and teacher predictions, obtaining expectation is crucial for obtaining bias and variance. In this work, this invention develops an expectation estimation module to predict student expectations.

[0084] The goal of reconstruction-based self-supervised models (such as MAE) is to recover the original input. For example, minimizing the mean squared error essentially involves regressing the expectation of the input samples. Based on the results of this analysis, this invention develops a lightweight reconstruction proxy task to achieve the desired estimation results. Figure 2 As shown.

[0085] A random mask is generated according to formula (6), and then appended to the student's prediction f. s Then, the expectation estimation module is used to reconstruct the masked pixels from the remaining pixels.

[0086]

[0087] Among them, M H,W For the constructed mask, R H,W ∈(0,1) represents a random number, and λ represents a hyperparameter that can adjust the mask ratio.

[0088] Finally, the variance in the knowledge distillation process can be measured using the variance formula and the provided expectation estimation module. Variance is used to model the data uncertainty of teacher predictions, as shown below:

[0089] var kd =E[(f s -E(f s )) 2 (7)

[0090] Among them, f s For student predictions, var kd Let E(f) be the variance. s ( ) represents the expected predictions of the students.

[0091] 2. The purpose of this invention is to represent bias learning as the process by which the student's output and the teacher's prediction reach consistency, and to represent variance learning as a measure of the inherent data uncertainty of the teacher in medical image segmentation. Furthermore, to address the dilemma of bias and variance learning during the distillation process, a balance factor is introduced to balance the trade-off between bias and variance in medical image segmentation, allowing the distillation loss to gradually converge.

[0092] This invention proposes that the unreliable estimation of data uncertainty in teacher predictions is attributed to bias-variance coupling in knowledge distillation. Bias-variance decomposition theory can independently measure bias and variance in the knowledge distillation process. Bias (as shown in Equation (8)) represents the degree of matching between student and teacher predictions, while variance (as shown in Equation (7)) measures the difference between student predictions and their expectations.

[0093]

[0094] in, For deviation, y t For teacher prediction, as shown in formula (9), the initial bias-variance decomposition distillation loss can be obtained.

[0095]

[0096] in, The initial bias variance decomposition distillation loss.

[0097] According to formula (9), the process of maximizing mutual information has and var kd The dilemma of not being able to minimize simultaneously. To address this issue, this invention introduces a balance factor to regulate the overall loss for both sides during the game. Students can not only perceive the data uncertainty caused by the variance of the teacher's predictions throughout the game, but also actively extract appropriate knowledge during the distillation process, thereby avoiding interference from noise or redundant information in the teacher's presentation. A feasible paradigm is to use the balance factor exp(-var kd The deviation-variance decomposition distillation loss is incorporated into formula (9) to obtain the final deviation-variance decomposition distillation loss (as shown in formula (10)), which can then be used for startup. and var kd The game between bias and variance is explored. Students can actively uncover consistent information while eliminating interference from inconsistent information. Ultimately, a bias-variance game in medical image segmentation is used to optimize the distillation loss.

[0098]

[0099] in, This represents the optimized bias-variance decomposition distillation loss.

[0100] Several metrics are available for calculating bias and variance in knowledge distillation, including KL divergence, MSE, and cross-entropy. For example... Figure 3 As shown, in the early stages of the knowledge distillation process, bias is primarily optimized, while variance is gradually incorporated into the optimization scope in the later stages. To adhere to this rule, a coefficient α∈[0,1] is introduced, which is a monotonically decreasing function (a linear function is used for simplicity), gradually decreasing from α to 1 as the training process progresses. The initial bias-variance decomposition distillation loss transition (i.e., Equation (9)) to The final deviation-variance decomposition distillation loss is calculated as follows (i.e., Equation (10)). Therefore, the final deviation-variance decomposition distillation loss is as follows:

[0101]

[0102] Among them, L kd This represents the final deviation-variance decomposition distillation loss value, where α is a preset coefficient. For the deviation, exp(-var kd ) is the balance factor, var kd Let V be the variance and E be the expected value.

[0103] 3. This invention combines self-supervised learning with knowledge distillation to improve the quality of pre-trained teachers. Different tasks in medical image diagnosis and the various patterns involved often exhibit consistency in anatomical structures, enabling the development of reasonable self-supervised pretext tasks. Pretext tasks can explore the teacher's generalized representations, thereby improving knowledge transfer to students.

[0104] Due to the scarcity of medical image data and the time-consuming and laborious process of lesion location labeling, pre-training teachers using supervised learning is extremely challenging. Fortunately, medical images from different tasks or various modalities often share anatomical structural consistency, thus teacher networks trained on related or similar tasks can effectively transfer knowledge to student networks. This invention employs self-supervised learning for teacher pre-training. The goal of self-supervised learning is to develop an effective agent task that can be used to mine generalizable feature representations.

[0105] For example, this invention uses the classic ResNet-101 network as the teacher and the ResNet34 network backbone as the student, employing DeeplapV3 as the segmentation head and FCN as the auxiliary segmentation head to obtain the bias-variance decomposition knowledge distillation network BENIN designed in this invention. The specific structure of the student network is shown in Table 1.

[0106] Table 1

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

Claims

1. A bias-variance decomposition segmentation method for medical images, characterized in that, include: Acquiring medical images; The medical image is input into the image segmentation model to obtain the segmented medical image. The image segmentation model uses a teacher network and a student network. The bias and variance in the distillation process of the teacher network and the student network are decoupled using bias-variance theory. The process of obtaining the image segmentation model includes: The corresponding first predicted image is obtained using the teacher network; Obtain the predicted expectation image of the student network; Based on the first predicted image and the predicted expected image, the bias-variance decomposition distillation loss is obtained; The bias-variance decomposition distillation loss is optimized and trained using a balance factor until it converges, thus obtaining the image segmentation model. The bias-variance decomposition distillation loss is optimized and trained using a balance factor until it converges, including: The bias-variance decomposition distillation loss is optimized by using a balance factor to obtain the optimized bias-variance decomposition distillation loss. By introducing preset coefficients, the bias-variance decomposition and distillation loss during optimization training converges, and the optimal bias-variance decomposition and distillation loss is obtained to complete the training of the image segmentation model. The optimal bias-variance decomposition distillation loss is: in, This represents the deviation-variance decomposition distillation loss value. For preset coefficients, For deviation, As a balance factor, For variance, This is the expected value.

2. The bias-variance decomposition and segmentation method for medical images according to claim 1, characterized in that, Before obtaining the corresponding first predicted image using the teacher network, the process includes: pre-training the teacher network.

3. The bias-variance decomposition and segmentation method for medical images according to claim 2, characterized in that, The pre-training process for the teacher network includes: The initial medical image is deformed using a random mask to obtain the deformed medical image. The deformed medical image is used to perform feature learning using an encoder, and the medical image after feature learning is used to perform image restoration using a decoder to obtain the restored medical image; Based on the initial medical image and the restored medical image, the reconstruction loss is calculated; Based on the reconstruction loss, after obtaining the initial medical image and the restored medical image, the teacher network is trained until the initial medical image and the restored medical image are consistent, thus completing the pre-training of the teacher network.

4. The bias-variance decomposition and segmentation method for medical images according to claim 1, characterized in that, Obtaining the predicted expectation image of the student network includes: obtaining a second predicted image using the student network, combining it with the expectation estimation module to predict the student's expectation, and obtaining the predicted expectation image of the student network.

5. The bias-variance decomposition and segmentation method for medical images according to claim 4, characterized in that, Combining the expectation estimation module to predict students' expectations, the predicted expectation image of the student network is obtained by: The second predicted image is masked using a random mask to obtain mask features; The mask features are reconstructed using 3×3 convolution and activation functions to obtain the reconstructed features; Based on the reconstructed features, the predicted expected image of the student network is obtained.

6. The bias-variance decomposition and segmentation method for medical images according to claim 5, characterized in that, Obtaining the bias-variance decomposition distillation loss includes: The deviation is obtained based on the first predicted image and the second predicted image; Based on the second predicted image and the expected predicted image, the variance is obtained; Based on the aforementioned bias and variance, the bias-variance decomposition distillation loss is obtained.