Semi-supervised stomatological image segmentation method based on cross-frequency cooperative training
Through cross-frequency collaborative training and wavelet transformation technology, combined with the CFC-Net model, the existing semi-supervised learning method lacks self-learning ability in oral image segmentation, ignoring the destruction of frequency domain information and structural integrity, achieving more efficient and accurate image segmentation.
Patent Information
- Application Number
- CN202510475199.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In the existing semi-supervised learning method, there is a problem that teacher models lack self-learning ability, ignore frequency domain information, and destroy the integrity of training image structure in oral imaging segmentation.
The semi-supervised oral image segmentation method with cross-frequency collaborative training is adopted to obtain low-frequency, high-frequency and full-frequency images through wavelet transformation, and the first and second phases of training are combined with the CFC-Net model to generate mixed training samples to ensure the consistency of replacement positions between images of different frequency and preserve the integrity of the image structure.
Improve the accuracy of image segmentation, effectively utilize unlabeled data, ensure the integrity of the image structure, and provide more reliable segmentation results.
Smart Images

Figure CN119991691A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to a cross-frequency collaborative training semi-supervised stomatological image segmentation method. Background Art
[0002] Periodontitis is an inflammatory disease that affects the apical tissue, with a global prevalence of over 50%. Root canal therapy is the main treatment for periodontitis and can preserve the basic functions of affected teeth in the oral cavity. However, the root canal system itself is complex, and there are considerable differences in root canal morphology between different individuals, affected by factors such as age and geographical region. This complexity makes the success of root canal treatment largely dependent on the expertise and subjective judgment of clinicians, and even minor negligence can lead to unpredictable results or treatment failure. Therefore, there is an urgent need for an intelligent analysis method to assess the health of a patient's root canal system before root canal therapy. In clinical practice, large-scale annotated medical data is often difficult to obtain, which makes the application of semi-supervised segmentation algorithms even more important.
[0003] Although the semi-supervised learning deep learning network image segmentation methods that have been made public have shown good segmentation results, they still have some limitations. First, the teacher model usually lacks self-learning ability. The teacher model parameters are the exponential moving average of the student model parameters. The teacher model guides the student model by generating pseudo labels of unlabeled data and introduces consistency regularization. Some previous methods do not enable the gradient update of the teacher model, but only update the parameters of the student model through training. This may lead to the problem that the teacher model can only be the EMA of the student model parameters. If the student model deviates, the teacher cannot correct itself, resulting in errors accumulating over time. Secondly, frequency domain information is crucial for medical image segmentation, especially for targets such as tooth root canals, which are small in size and have rich high-frequency details. However, most existing semi-supervised learning methods for medical image segmentation focus mainly on feature learning in the spatial domain, while ignoring the valuable information contained in the frequency domain; in addition, some previous methods adopt a hybrid mechanism in which image patches with lower uncertainty replace image patches with higher uncertainty to generate new training samples and produce more reliable pseudo labels. Although this method has been proven to be effective, it has a serious disadvantage. For example, Adaptive Bidirectional Displacement for Semi-Supervised Medical Image Segmentation (arXiv:2405.00378v1 [cs.CV] 1 May 2024) proposes an anti-confidence bidirectional displacement operation to mark images to generate more unreliable information samples, thereby promoting model learning; however, this replacement scheme still does not consider the position information of the replaced blocks, which will directly damage the structural integrity of the newly generated training images, which may damage the performance of the model and even cause a certain degree of output errors.
[0004] Therefore, in the application of dental image segmentation and recognition, there is an urgent need for a method that can overcome the limitations of existing methods and provide more efficient and accurate segmentation and recognition output. Summary of the invention
[0005] The purpose of the present invention is to provide a semi-supervised dental image segmentation method with cross-frequency collaborative training to solve the technical problems that the teacher model usually lacks self-learning ability, the analysis model ignores the frequency domain information in the image, and the structural integrity of the lesions in the training image is destroyed.
[0006] A cross-frequency collaborative training semi-supervised dental image segmentation method comprises the following steps: S1. Data preprocessing: Acquire images in a data set, and obtain pre-training samples by wavelet transform, wherein the pre-training samples include low-frequency images, high-frequency images, and full-frequency images; S2. First stage training: The pre-trained samples are input into the CFC-Net model for the first stage of training. Output the loss function of the first stage of training; S3. Mixed sample generation: Get the output of the CFC-Net model network, calculate the joint probability distribution and uncertainty map, Based on the low-frequency image, high-frequency image, full-frequency image and uncertainty map, high-confidence image blocks are selected for bidirectional cyclic mixing to generate mixed training samples; S4, second stage training: Receive mixed training samples and input them into the CFC-Net model for supervised and unsupervised training. Output the loss function of the second stage training; S5. Model output: The total loss function is calculated by the loss function of the first stage training and the loss function of the second stage training. The final model is obtained by iterative convergence of the total loss function. The image is segmented using the final model.
[0007] Furthermore, the CFC-Net model includes an expertise student network SL network for processing low-frequency images, an expertise student network SH network for processing high-frequency images, and a comprehensive teacher network CT network for processing full-frequency images. The three networks work together to complete semi-supervised learning training, and the two expertise student networks and the comprehensive teacher network have the same structure.
[0008] Furthermore, in S1, data preprocessing, Wavelet transform is applied to each image using formula (1) to extract low-frequency image, high-frequency image and full-frequency image. in represents the wavelet coefficients, (n, m) represents the pixels in the EF image, and (x, y) represents The coordinates in the domain, LL is the low-frequency component of the image, the low-frequency image is represented by LF, (LH, HL, HH) are three groups of high-frequency components of the image, the high-frequency image composed of high-frequency components is represented by HF, and the original image is a full-frequency image represented by EF.
[0009] Furthermore, the image data in the dataset includes labeled data and unlabeled data; The cross-frequency consistency loss is calculated by the pseudo-labels generated by the two expert student networks. , The pseudo labels generated by the comprehensive teacher network supervise the two expert student networks to calculate the full-frequency consistency loss , The unsupervised loss function L in the first training stage unsup1 It is the sum of full-frequency consistency loss and cross-frequency consistency loss; Supervised training of labeled data, calculating the cross entropy loss and DICE loss of SL network, SH network and CT network; The first training stage has a supervised loss function L sup1 is the sum of the segmentation losses of the three networks, calculated using formula (4), with the input LF, HF, and EF components sharing the same label, in represents the segmentation loss, , and Respectively represent the outputs of SL network, SH network and CT network, and Represents the corresponding label; i is the serial number of the annotated image, and m is the total number of annotated images.
[0010] Furthermore, unsupervised training is performed on unlabeled data: In the unsupervised path, the parameters of the CT network are updated using the combined EMA of the SL network and the SH network, calculated by formula (5): Where α and β are smoothing coefficients, which control the EMA update rate and the relative contribution of the SL network and the SH network to the EMA, respectively. represents the weight of the corresponding network at the t-th iteration, sl is low frequency, and sh is high frequency; θ with both superscript and subscript represents the parameters of the corresponding sl network or sh network at the t-th iteration; The CT network is used to generate pseudo labels, and the output of the two expert student networks is constrained by the full-frequency consistency loss. The mutual supervision of the two expert student networks is achieved through cross-frequency consistency loss to avoid error accumulation. The full-frequency consistency loss constrains the expert student network by synthesizing the pseudo-labels generated by the teacher network; The CT network generates pseudo labels, supervises the SS network, and uses formulas (7, 8) for optimization. is the loss function of the CT network relative to the SL network, where represents the pseudo label generated by the CT network, is the loss function of the CT network relative to the SH network, where represents the pseudo label generated by the CT network, j is the serial number of the unlabeled image, and n is the total number of unlabeled images; The cross-frequency consistency loss is achieved by generating pseudo labels between two specialized student networks. The SL network and the SH network generate pseudo labels for collaborative learning, and are calculated using formulas (9, 10). is the loss function of the SL network relative to the SH network, where represents the pseudo-label generated by the SL network, is the loss function of the SH network relative to the SL network, where represents the pseudo-label generated by the SH network, and j is the number of unlabeled images.
[0011] Furthermore, the specific steps in S3, generating a mixed sample, include: Get the output of the CFC-Net model network, Compute the joint probability distribution P and uncertainty map U, Divide the low-frequency image, high-frequency image, full-frequency image and uncertainty map into blocks. Get the location information of high confidence tiles to get replacement location information. The high confidence block acquisition step is to sort the blocks segmented from the uncertainty map U from high to low according to the joint probability distribution, and select the top 25% to 35% of the blocks. According to the replacement position information, a bidirectional cyclic mixing is performed between the low-frequency, high-frequency and full-frequency image blocks at the same position to generate mixed training samples. The blending operation only replaces image blocks at the same position, preserving the integrity of the image structure.
[0012] Furthermore, S4, the steps of the second stage training include: The mixed training samples are input into the CFC-Net network, and supervised and unsupervised training processes are performed to obtain the supervised loss function L of the second stage training. sup2 and the unsupervised loss function L in the second training stage unsup2 .
[0013] Furthermore, S5, the model output step includes: The total loss function is composed of the total supervised loss Lsup and the total unsupervised loss L unsup Weighted composition, the weights are adjusted by Gaussian warm-up function, Total loss function Calculated by formula (2), Where λ is the regularization term used to control the unsupervised loss weight, defined by the Gaussian warm-up function and calculated by formula (3), Where t represents the current iteration, represents the λ value at the t-th iteration, is the total number of iterations, is a hyperparameter representing the maximum value of λ, The total supervised loss L sup is the supervised loss function L for the first stage of training sup1 And the supervised loss function L of the second stage training sup2 the sum of The total unsupervised loss L unsup is the unsupervised loss function in the first training stage Lunsup1 , and the unsupervised loss function in the second training stage Lunsup2 the sum of The supervised loss function L of the first stage training sup1 and the unsupervised loss function in the first training stage Lunsup1 Output from the first stage of training, The supervised loss function L for the second stage training sup2 and the unsupervised loss function in the second training phase Lunsup2 Output from the second stage training, The model is iteratively trained until the total loss function is minimized or the number of iterations is reached; the trained CFC-Net model is output for segmenting uncertainty images and mining quantity and information.
[0014] Compared with the prior art, the technical solution provided by this application has beneficial effects.
[0015] The CFC-Net in this application provides a semi-supervised CBCT image tooth root canal segmentation method, which integrates multi-frequency information to ensure the accuracy of root canal segmentation. At the same time, this method can effectively utilize a large amount of clinical unlabeled data to further improve the model performance.
[0016] The hybrid mechanism of this method ensures that the replacement positions between images of different frequencies are consistent, and the replaced training samples still contain complete image information, and the original structure of the image will not be destroyed due to the replacement of low-certainty image blocks with high-certainty image blocks.
[0017] This method will provide important insights into tract number and morphology, allowing clinicians to develop precise surgical plans while minimizing labor and financial costs.
[0018] We also conducted detailed experimental tests on the proposed CFC-Net on three other public dental segmentation tasks. The results showed that our method can also show a certain degree of robustness in other dental segmentation tasks and is superior to the existing image segmentation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] in: Figure 1 A flowchart of the semi-supervised dental image segmentation method with cross-frequency collaborative training in this application; Figure 2 This is the network structure diagram of the CFC-Net model in this application; Figure 3 A visual comparison of CFC-Net and other semi-supervised segmentation methods in this application on the root canal segmentation task; Figure 4 This is a visual comparison of CFC-Net and other semi-supervised segmentation methods on the TDD dataset in this application; Figure 5 This is a visual comparison of CFC-Net and other semi-supervised segmentation methods on the CTooth dataset in this application; Figure 6 This is a visual comparison of CFC-Net and other semi-supervised segmentation methods on the NKUT dataset in this application; Figure 7 This is a quantitative comparison chart of CFC-Net and other semi-supervised segmentation methods in the root canal segmentation task in this application; Figure 8 This is a quantitative comparison chart of CFC-Net and other semi-supervised segmentation methods on the TDD dataset in this application; Fig. 9 This is a quantitative comparison chart of CFC-Net and other semi-supervised segmentation methods on the CTooth dataset in this application; Fig.10This is a quantitative comparison chart of CFC-Net and other semi-supervised segmentation methods on the NKUT dataset in this application. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of this application.
[0022] like Figure 1 As shown, the process of the semi-supervised dental image segmentation method with cross-frequency collaborative training in the present application is shown.
[0023] like Figure 2 As shown, the two main structures of CFC-Net of the present application are shown, namely the cross-frequency coordination module (CFC-MT) shown on the left and the uncertainty cross-frequency mixing mechanism module (UCF-Mix) shown on the right.
[0024] In CFC-MT, a collaborative training architecture of two expert students (SS networks) and one comprehensive teacher (CT network) is used to complete semi-supervised learning training, and the structures of the three networks are the same.
[0025] In the SS network, the one that processes low-frequency images is marked as the SL network, and the one that processes high-frequency images is marked as the SH network.
[0026] In UCF-Mix, low-frequency images, high-frequency images, and full-frequency images are obtained by wavelet transforming the original image, and high-determinacy image blocks at the same position of the low-frequency image, high-frequency image, and full-frequency image are bidirectionally cyclically mixed to improve the network robustness while retaining the integrity of the target structure.
[0027] The specific steps of model training are described in detail below: For the clinical dental dataset, CFC-Net was first used to perform the first round of network training to obtain the segmentation map.
[0028] During training, for labeled images, cross entropy loss and DICE loss are used jointly for supervision; For unlabeled images, two sets of loss functions are used for supervision, namely, the cross-frequency consistency loss between the two expert students and the full-frequency consistency loss performed by the comprehensive teacher on the two expert students respectively; like Figure 2As shown in the left part, given an input image, the first step is to apply the wavelet transform described in formula (1) to extract its frequency components: LL, HL, LH and HH, in represents the wavelet coefficients, (n, m) represents the pixels in the EF image, and (x, y) represents Coordinates in the domain.
[0029] LL is the low-frequency component of the image, and the low-frequency image is represented by LF. (LH, HL, HH) are three groups of high-frequency components of the image. The high-frequency image composed of high-frequency components is represented by HF. The original image is a full-frequency image represented by EF.
[0030] For labeled data, all networks are trained using labels for supervision.
[0031] For unlabeled data, we use two loss functions: and to optimize unsupervised training.
[0032] It is the full-frequency consistency loss function of the comprehensive teacher on the two expertise students respectively; is the cross-frequency consistency loss function between two expert students; The loss function of the entire training process It consists of two parts, namely and , as shown in Formula 2.
[0033] Where λ is the regularization term used to control the unsupervised loss weight, defined by the Gaussian warm-up function as shown in Formula 3: Where t represents the current iteration, represents the λ value at the t-th iteration, is the total number of iterations, is a hyperparameter representing the maximum value of λ.
[0034] For training with labeled data: The SS network and the CT network are trained using labeled data; for pairs of LF, EF, and HF components, all components are trained with the same labels.
[0035] The first training stage has a supervised loss function L sup1 It is defined in formula (4).
[0036] in represents the segmentation loss. , and Represent the outputs of the low-frequency SS network, high-frequency SS network, and CT network, respectively, and Indicates the corresponding label. i is the serial number of the annotated image, and m is the total number of annotated images.
[0037] For training with unlabeled data: In the unsupervised path, the parameters of the CT network are updated using the combined EMA of the two SS networks, as defined in Equation (5), where α and β are smoothing coefficients that control the EMA update rate and the relative contribution of the high- and low-frequency SS networks to the EMA, respectively. It represents the weight of the corresponding network at the tth iteration, sl is low frequency and sh is high frequency. Theta with both superscript and subscript represents the parameters of the corresponding sl network or sh network at the tth iteration.
[0038] The unsupervised loss function L in the first training stage unsup1 It consists of two parts, as defined in formula (6).
[0039] The pseudo labels generated by the CT network are used to supervise the outputs of the two SS networks simultaneously.
[0040] The main goal is to enable the CT network to guide and constrain the training process of the SS network from a full-frequency perspective, thereby preventing errors and accumulation of errors during the training process.
[0041] As shown in Equations 7 and 8, this gives The formula is: exist In , two SS networks generate pseudo labels for each other to achieve mutual supervision.
[0042] is the loss function of the CT network relative to the low-frequency SS network, where represents the pseudo labels generated by the CT network.
[0043] is the loss function of the CT network relative to the high-frequency SS network, where represents the pseudo labels generated by the CT network.
[0044] j is the serial number of the unlabeled image, and n is the total number of unlabeled images.
[0045] This consistency mechanism allows SS networks to learn collaboratively, leverage their respective strengths to make up for weaknesses, correct errors, and avoid isolated learning.
[0046] As shown in Equations 9 and 10, this gives Formula; is the loss function of the low-frequency SS network relative to the high-frequency SS network, where represents the pseudo labels generated by the low-frequency SS network; is the loss function of the high-frequency SS network relative to the low-frequency SS network, where represents the pseudo-label generated by the high-frequency SS network; j is the number of unlabeled images.
[0047] The specific steps of UCF-Mix to generate mixed training samples are described as follows.
[0048] Generate supervised mixed samples using labeled data; Generate unsupervised samples using unlabeled data.
[0049] The generation process of UCF-Mix is illustrated by taking supervised mixed samples as an example; the process of generating new unsupervised samples is similar. Receive output from SS and CT networks , and , calculate the joint probability distribution P and uncertainty map U.
[0050] The joint probability distribution P is defined as formula (11).
[0051] The uncertainty map U is calculated using formula (12), where c represents the number of segmentation categories and ϵ is a small parameter introduced to prevent the logarithm calculation from approaching zero.
[0052] In the formula, , and The output obtained for the SS and CT networks; where c represents the number of segmentation categories and ϵ is a small parameter introduced.
[0053] The specific implementation method is as follows Figure 2 As shown in the right part, the low-frequency component image, the full-frequency component image, the high-frequency component image, and the uncertainty map U are all segmented into multiple blocks according to the same segmentation principle; Obtaining the position information of the high confidence block to obtain the replacement position information; The steps for obtaining high-confidence blocks are to sort the blocks segmented by the uncertainty map U from high to low according to the joint probability distribution, and select the top 25% to 35% of the blocks.
[0054] Perform bilateral mixing of the patches corresponding to the replacement patch positions of the uncertainty map U in the low-frequency component image, the full-frequency component image, and the high-frequency component image to obtain mixed training samples; the specific mixing operation is: In the first round, the selected patches are mixed in the forward order of sl → t → sh → sl to obtain mixed training sample 1: EF1, HF1, LF1.
[0055] In the second round, the same patches are mixed in the reverse order of sh → t → sl → sh to obtain mixed training sample 2: EF2, LF2, HF2.
[0056] Taking EF2 as an example, the HF image corresponding to sh is replaced with the image of the corresponding patch in EF according to the replacement patch position information to generate EF2. The acquisition of similar mixed training samples will not be repeated.
[0057] The mixed samples are input into CFC-MT for the second stage of training. like Figure 2 As shown in the middle part; the training process is the same as the first stage training process, only the obtained training samples are different. The supervised and unsupervised loss functions of this stage are respectively expressed as the supervised loss function L of the second stage training sup2 and the unsupervised loss function in the second training phase Lunsup2 , which is in the same form as L sup1 and L unsup1 The calculation process is the same, except that the training samples used are different, which are distinguished by marking, and will not be described here.
[0058] It should be noted that in the second stage of training, mixed training sample 1 and mixed training sample 2 are used for one training iteration respectively.
[0059] The total loss of the entire training process, the total supervised loss Lsup and the total unsupervised loss L unsup is the sum of the losses of the two stages, as defined in Equations (13) and (14).
[0060] It is worth noting that UCF-Mix only mixes the patches at the same position in different frequency components. Therefore, it does not require any modification to the training labels, thus preserving the structural integrity of the target.
[0061] The optimization target is the loss function of the entire training process When the iteration converges to the set value or the number of iterations, the model parameters are obtained. The model is used to segment the input dental images and mine the medical information therein.
[0062] The module of the finally trained CFC-Net model for image segmentation and information mining is the trained teacher network, and the student network no longer participates in image segmentation and information mining.
[0063] The parameters and data flows are summarized in the following table:
[0064] It should be noted that the three network architectures of BCP, UCMT and CFC-Net of this application are different. Although all three are MT architectures and all use the mix-up mechanism, the difference lies in the different mixing methods and purposes of the three. BCP mixes random patches of labeled and unlabeled data, emphasizing the consistency of the output distribution of labeled and unlabeled data; UCMT replaces low-uncertainty blocks with high-certainty blocks in an image, emphasizing high confidence, but may destroy the structure and integrity of the image. CFC-Net uses the uncertainty graph generated by the mixture of three networks, selects the K blocks with the highest comprehensive confidence in the results of the three networks, and rotates them in the same position among the three, emphasizing the guarantee of high confidence while enhancing the performance of cross-frequency learning and representation.
[0065] It is further explained that, compared with the previous two mixed methods, the mixed method selected by this application has complete labels, because the blocks in the same position are exchanged and mixed with each other, which will not cause structural confusion in the original image. For example, after the high-frequency block No. 1 and the full-frequency block No. 1 are exchanged with each other, there is full-frequency information in the high-frequency image and prominent high-frequency information in the full-frequency image, but because of the same position, the labels output by the two remain consistent.
[0066] UCMT hybrid replaces low-uncertainty blocks with high-certainty blocks in the same image, which will lose relevant information and destroy the original structure of the original image. For example, the network has high certainty in identifying the center of a tumor, but low certainty in learning the edge of the tumor. UCMT directly cuts the center of the tumor and sticks it to the edge of the tumor. Intuitively, the uncertainty in the image is reduced, but at the cost of destroying the integrity of the lesion and losing the image content.
[0067] In the extreme case, the algorithm uses the block in the middle to identify the most certainty, and replaces the entire image with the same block. In this way, the uncertainty will be very low when the image is output, because each of its blocks is a block that performs well in the entire image. However, it is precisely because of this replacement that the integrity of the image is destroyed. It turns out that this image may contain non-tumor information such as skin and hair. Although image recognition is low certainty, we still need the information of the block when analyzing image information, otherwise this mixture will lose some necessary analysis information, resulting in incomplete information or even misjudgment or wrong judgment.
[0068] CFC does not emphasize improving the high certainty of the image, but mainly aims to obtain more detailed information of the image from the perspective of cross-frequency without reducing the certainty of the image. The information of the three images of low frequency, high frequency and full frequency is comprehensively calculated to obtain the joint probability map. On this uncertainty map, good blocks with high certainty are selected, and forward and reverse cycles are performed between the three images to perform cross-information training.
[0069] Using different data sets, the segmentation method proposed in this application is compared with the existing segmentation method. The comparison effect is shown in Figure 3-Figure 6 Demonstration, quantitative analysis of the comparison through Figure 7-10 Show, from Figure 3-Figure 10 It can be seen that the image information after segmentation is fully retained without any missing or abnormally excessive information. At the same time, the overall performance of quantitative information is at the forefront, and most indicators rank first.
[0070] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0071] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. A cross-frequency collaborative training method for semi-supervised dental image segmentation, characterized in that: The steps include: S1. Data preprocessing: Acquire images in a data set, and obtain pre-training samples by wavelet transform, wherein the pre-training samples include low-frequency images, high-frequency images, and full-frequency images; S2. First stage training: The pre-trained samples are input into the CFC-Net model for the first stage of training. Output the loss function of the first stage of training; S3. Mixed sample generation: Get the output of the CFC-Net model network, calculate the joint probability distribution and uncertainty map, Based on the low-frequency image, high-frequency image, full-frequency image and uncertainty map, high-confidence image blocks are selected for bidirectional cyclic mixing to generate mixed training samples; S4, second stage training: Receive mixed training samples and input them into the CFC-Net model for supervised and unsupervised training. Output the loss function of the second stage training; S5. Model output: The total loss function is calculated by the loss function of the first stage training and the loss function of the second stage training. The final model is obtained by iterative convergence of the total loss function. The image is segmented using the final model.
2. The cross-frequency collaborative training semi-supervised dental image segmentation method according to claim 1, characterized in that: The CFC-Net model includes an expert student network SL network for processing low-frequency images, an expert student network SH network for processing high-frequency images, and a comprehensive teacher network CT network for processing full-frequency images. The three networks work together to complete semi-supervised learning training, and the two expert student networks and the comprehensive teacher network have the same structure.
3. The semi-supervised dental image segmentation method of cross-frequency collaborative training according to claim 1, characterized in that: In S1, data preprocessing, Wavelet transform is applied to each image using formula (1) to extract low-frequency image, high-frequency image and full-frequency image. in represents the wavelet coefficients, (n, m) represents the pixels in the EF image, and (x, y) represents The coordinates in the domain, LL is the low-frequency component of the image, the low-frequency image is represented by LF, (LH, HL, HH) are the three groups of high-frequency components of the image, the high-frequency image composed of high-frequency components is represented by HF, and the original image is a full-frequency image represented by EF.
4. The semi-supervised dental image segmentation method of cross-frequency collaborative training according to claim 2, characterized in that: The image data in the dataset includes labeled data and unlabeled data; The cross-frequency consistency loss is calculated by the pseudo-labels generated by the two expert student networks. , The pseudo labels generated by the comprehensive teacher network supervise the two expert student networks to calculate the full-frequency consistency loss , The unsupervised loss function L in the first training stage unsup1 It is the sum of full-frequency consistency loss and cross-frequency consistency loss; Supervised training of labeled data, calculating the cross entropy loss and DICE loss of SL network, SH network and CT network; The first training stage has a supervised loss function L sup1 is the sum of the segmentation losses of the three networks, calculated using formula (4), with the input LF, HF, and EF components sharing the same label, in represents the segmentation loss, , and Respectively represent the outputs of SL network, SH network and CT network, and Represents the corresponding label; i is the serial number of the annotated image, and m is the total number of annotated images.
5. The semi-supervised dental image segmentation method of cross-frequency collaborative training according to claim 4, characterized in that: Unsupervised training for unlabeled data: In the unsupervised path, the parameters of the CT network are updated using the combined EMA of the SL network and the SH network, calculated by formula (5), Where α and β are smoothing coefficients, which control the EMA update rate and the relative contribution of the SL network and the SH network to the EMA, respectively. represents the weight of the corresponding network at the t-th iteration, sl is low frequency, and sh is high frequency; θ with both superscript and subscript represents the parameters of the corresponding sl network or sh network at the t-th iteration; The CT network is used to generate pseudo labels, and the output of the two expert student networks is constrained by the full-frequency consistency loss. The mutual supervision of the two expert student networks is achieved through cross-frequency consistency loss to avoid error accumulation. The full-frequency consistency loss constrains the expert student network by synthesizing the pseudo-labels generated by the teacher network; The CT network generates pseudo labels, supervises the SS network, and optimizes using formulas (7, 8). is the loss function of the CT network relative to the SL network, where represents the pseudo label generated by the CT network, is the loss function of the CT network relative to the SH network, where represents the pseudo label generated by the CT network, j is the serial number of the unlabeled image, and n is the total number of unlabeled images; The cross-frequency consistency loss is achieved by generating pseudo labels between two specialized student networks. The SL network and the SH network generate pseudo labels for collaborative learning, and are calculated using formulas (9, 10). is the loss function of the SL network relative to the SH network, where represents the pseudo-label generated by the SL network, is the loss function of the SH network relative to the SL network, where represents the pseudo-label generated by the SH network, and j is the number of unlabeled images.
6. The cross-frequency collaborative training semi-supervised dental image segmentation method according to claim 5, characterized in that: S3. The specific steps of generating mixed samples include: Get the output of the CFC-Net model network, Compute the joint probability distribution P and uncertainty map U, Divide the low-frequency image, high-frequency image, full-frequency image and uncertainty map into blocks. Get the location information of high confidence tiles to get replacement location information. The high confidence block acquisition step is to sort the blocks segmented from the uncertainty map U from high to low according to the joint probability distribution, and select the top 25% to 35% of the blocks. According to the replacement position information, a bidirectional cyclic mixing is performed between the low-frequency, high-frequency and full-frequency image blocks at the same position to generate mixed training samples. The blending operation only replaces image blocks at the same position, preserving the integrity of the image structure.
7. The semi-supervised dental image segmentation method of cross-frequency collaborative training according to claim 6, characterized in that: S4. The steps of the second stage training include: The mixed training samples are input into the CFC-Net network, and supervised and unsupervised training processes are performed to obtain the supervised loss function L of the second stage training. sup2 and the unsupervised loss function L in the second training stage unsup2 .
8. The semi-supervised dental image segmentation method of cross-frequency collaborative training according to claim 7, characterized in that: S5. The model output step includes: The total loss function is composed of the total supervised loss L sup and the total unsupervised loss L unsup Weighted composition, the weights are adjusted by Gaussian warm-up function, Total loss function Calculated by formula (2), Where λ is the regularization term used to control the unsupervised loss weight, defined by the Gaussian warm-up function and calculated by formula (3), Where t represents the current iteration, represents the λ value at the t-th iteration, is the total number of iterations, is a hyperparameter representing the maximum value of λ, The total supervised loss L sup is the supervised loss function L for the first stage of training sup1 And the supervised loss function L of the second stage training sup2 the sum of The total unsupervised loss L unsup is the unsupervised loss function in the first training stage Lunsup1 , and the unsupervised loss function in the second training stage Lunsup2 the sum of The supervised loss function L of the first stage training sup1 and the unsupervised loss function in the first training stage Lunsup1 Output from the first stage of training, The supervised loss function L for the second stage training sup2 and the unsupervised loss function in the second training phase Lunsup2 Output from the second stage training, The model is iteratively trained until the total loss function is minimized or the number of iterations is reached; the final trained CFC-Net model is output, and the image is segmented and information mined using the final trained CFC-Net model.
Citation Information
Patent Citations
Semi-supervised semantic segmentation method based on scale perception attention
CN115661463A
Semi-supervised medical image segmentation method based on consistency loss function
CN116258730A
Semi-supervised medical image segmentation method, system, equipment and medium
CN117095014A
Semi-supervised domain generalization medical image segmentation method and system
CN118657790A
Semi-supervised medical image segmentation method and device
CN119693392A