A semi-supervised oral cavity image segmentation method for cross-frequency collaborative training

Through the semi-supervised oral image segmentation method of cross-frequency collaborative training, the CFC-Net model is used to fuse multi-frequency information, which solves the problem of lack of self-learning and ignoring frequency domain information of the teacher model, and achieves efficient accuracy and structural integrity of tooth root canal segmentation, which is suitable for the treatment of root canal in periodontitis.

CN119991691BActive Publication Date: 2025-07-22NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510475199.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-22
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing semi-supervised learning methods have problems in the treatment of root canals of periodontitis that the teacher model lacks self-learning ability, ignores frequency domain information, and damages in the structural integrity of the training image, resulting in inaccurate segmentation recognition.

Method used

The semi-supervised oral image segmentation method with cross-frequency collaborative training is adopted, and multi-frequency information fusion is performed through the CFC-Net model, and low-frequency, full-frequency and high-frequency images are obtained using wavelet transform. Combined with the combined probability distribution and uncertainty map, high confidence tiles are generated for mixed training, so as to realize collaborative learning of teachers and students' networks and preserve the integrity of the image structure.

Benefits of technology

It improves the accuracy and model performance of tooth root canal segmentation, can effectively utilize unmarked data, reduce clinical costs, and provide accurate surgical plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991691B_ABST
    Figure CN119991691B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of medical image processing, and specifically relates to a semi-supervised oral cavity image segmentation method for cross-frequency collaborative training, aiming to solve the problems of insufficient self-learning ability of the teacher model, neglect of domain information, and damage to the lesion structure by the mixing mechanism in the existing semi-supervised segmentation method in medical image segmentation; decomposing the image into low-frequency, high-frequency, and full-frequency images through wavelet transform, using two specialized student networks and a comprehensive teacher network for collaborative training, using cross-entropy and Dice loss to supervise the labeled data, and combining cross-frequency consistency and full-frequency consistency losses to mine the unlabeled data; through an uncertainty cross-frequency mixing mechanism, bidirectionally cyclically mixing between high-confidence image patches to generate new samples and retaining the integrity of the target structure. An efficient and accurate segmentation scheme is provided, which provides reliable support for the preoperative evaluation of oral cavity treatment, reduces the clinical labor cost, and improves the treatment success rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and particularly relates to a semi-supervised oral cavity image segmentation method for cross-frequency collaborative training. Background Art

[0002] Periodontitis is an inflammatory disease that affects the apical tissues, with a global prevalence rate exceeding 50%. Root canal treatment is the main treatment method for periodontitis and can preserve the basic functions of the affected teeth in the oral cavity. However, the root canal system itself is very complex, and affected by factors such as age and geographical region, there are quite large differences in the root canal morphology of different individuals. This complexity makes the success of root canal treatment largely dependent on the professional knowledge and subjective judgment of clinicians, and even a slight oversight can lead to unpredictable results or treatment failure. Therefore, there is an urgent need for an intelligent analysis method to evaluate the health status of the patient's root canal system before root canal treatment. In clinical practice, it is often difficult to obtain large-scale labeled medical data, which makes the application of semi-supervised segmentation algorithms more important.

[0003] Although the publicly available semi-supervised learning deep learning network image segmentation methods have shown good segmentation effects, they still have some limitations. First of all, the teacher model usually lacks the ability of self-learning. The parameters of the teacher model are the exponential moving average of the parameters of the student model. The teacher model guides the student model by generating pseudo-labels for unlabeled data and introducing consistency regularization. Some previous methods do not turn on the gradient update of the teacher model and only update the parameters of the student model through training. The problem that may arise is that the teacher model can only be the EMA of the parameters of the student model. If the student model has a deviation, the teacher cannot correct itself, resulting in the accumulation of errors over time. Secondly, frequency domain information is crucial for medical image segmentation, especially for targets such as dental root canals, which are small in size and have rich high-frequency details. However, most of the existing semi-supervised learning methods for medical image segmentation mainly focus on feature learning in the spatial domain and ignore the valuable information contained in the frequency domain. Moreover, some previous methods adopt a hybrid mechanism in which image patches with lower uncertainty replace those with higher uncertainty to generate new training samples and produce more reliable pseudo-labels. Although this method has been proven to be effective, it has a serious drawback. For example, "Adaptive Bidirectional Displacement for Semi-Supervised Medical Image Segmentation" (arXiv:2405.00378v1 [cs.CV] 1 May 2024) proposes a bidirectional displacement operation of anti-confidence to label images to generate more samples with unreliable information, thus promoting model learning. However, this replacement scheme still does not consider the position information of the replaced patches, which will directly damage the structural integrity of the newly generated training images, thereby possibly damaging the performance of the model and even causing a certain degree of output errors.

[0004] Therefore, in the application of oral cavity imaging segmentation recognition, there is an urgent need for a method that can overcome the limitations of existing methods and provide more efficient and accurate segmentation recognition output. Summary of the Invention

[0005] The purpose of the present invention is to provide a semi-supervised oral cavity imaging segmentation method for cross-frequency collaborative training to solve the technical problems that the teacher model usually lacks the ability of self-learning, the analysis model ignores the frequency domain information in the image, and the structural integrity of the lesions in the training images is damaged.

[0006] A semi-supervised oral cavity imaging segmentation method for cross-frequency collaborative training includes the following steps:

[0007] S1. Data preprocessing:

[0008] Obtain the images in the dataset and obtain pre-training samples through wavelet transform. The pre-training samples include low-frequency images, high-frequency images, and full-frequency images;

[0009] S2. First-stage training:

[0010] Input the pre-training samples into the CFC-Net model for the first-stage training,

[0011] Output the loss function of the first-stage training;

[0012] S3. Generation of mixed samples:

[0013] Obtain the output of the CFC-Net model network, calculate the joint probability distribution and uncertainty map,

[0014] Based on the low-frequency images, high-frequency images, full-frequency images, and uncertainty map, select high-confidence image patches for two-way cyclic mixing to generate mixed training samples;

[0015] S4. Second-stage training:

[0016] Receive the mixed training samples and input them into the CFC-Net model for supervised training and unsupervised training,

[0017] Output the loss function of the second-stage training;

[0018] S5. Model output:

[0019] Calculate the total loss function through the loss function of the first-stage training and the loss function of the second-stage training,

[0020] Obtain the final model through the iterative convergence of the total loss function,

[0021] Use the final model to segment the images.

[0022] Furthermore, the CFC-Net model includes a specialized student network SL for processing low-frequency images, a specialized student network SH for processing high-frequency images, and a comprehensive teacher network CT for processing full-frequency images. The three networks cooperate to complete semi-supervised learning training, and the structures of the two specialized student networks and the comprehensive teacher network are the same.

[0023] Furthermore, in S1. Data preprocessing,

[0024] Apply wavelet transform to each image using formula (1) to extract low-frequency images, high-frequency images, and full-frequency images,

[0025]

[0026] where denotes the wavelet coefficient, (n, m) denotes the pixel in the EF image, and (x, y) denotes the coordinates in the domain, LL is the low-frequency component of the image, the low-frequency image is denoted by LF, (LH, HL, HH) are the three high-frequency components of the image, and the high-frequency image composed of the high-frequency components is denoted by HF, and the original image is the full-frequency image denoted by EF.

[0027] Furthermore, the image data in the dataset includes labeled data and unlabeled data;

[0028] Calculate the cross-frequency consistency loss from the pseudo-labels generated mutually between two specialized student networks ,

[0029] Supervise the two specialized student networks by the pseudo-labels generated by the comprehensive teacher network to calculate the full-frequency consistency loss ,

[0030] The unsupervised loss function L unsup1 in the first training stage is the sum of the full-frequency consistency loss and the cross-frequency consistency loss;

[0031] Perform supervised training on the labeled data, and calculate the cross-entropy loss and DICE loss of the SL network, SH network, and CT network;

[0032] The supervised loss function L sup1 in the first training stage is the sum of the segmentation losses of the three networks, calculated using formula (4), inputting the LF, HF, and EF components, sharing the same label,

[0033]

[0034] where denotes the segmentation loss, and the terms , and respectively denote the outputs of the SL network, SH network, and CT network, while denotes the corresponding label; i is the serial number of the labeled image, and m is the total number of labeled images.

[0035] Furthermore, perform unsupervised training on the unlabeled data:

[0036] In the unsupervised path, the parameters of the CT network are updated using the combined EMA of the SL network and SH network, calculated by formula (5),

[0037]

[0038] where α and β are smoothing coefficients, controlling the EMA update rate and the relative contributions of the SL network and SH network to the EMA respectively, denotes the weights of the corresponding network at the t-th iteration, where sl is the low frequency and sh is the high frequency; θ with both superscript and subscript represents the parameters of the corresponding sl network or sh network at the t-th iteration;

[0039] The CT network is used to generate pseudo-labels, and the outputs of the two specialized student networks are constrained by the full-frequency consistency loss.

[0040] Mutual supervision of the two specialized student networks is achieved through the cross-frequency consistency loss, avoiding error accumulation.

[0041] The full-frequency consistency loss constrains the specialized student network by integrating the pseudo-labels generated by the teacher network.

[0042] The CT network generates pseudo-labels to supervise the SS network and is optimized using equations (7, 8).

[0043]

[0044]

[0045] is the loss function of the CT network relative to the SL network, where denotes the pseudo-labels generated by the CT network.

[0046] is the loss function of the CT network relative to the SH network, where denotes the pseudo-labels generated by the CT network.

[0047] j is the serial number of the unlabeled image, and n is the total number of unlabeled images.

[0048] The cross-frequency consistency loss is achieved by the two specialized student networks mutually generating pseudo-labels. The SL network and the SH network mutually generate pseudo-labels for collaborative learning and are calculated using equations (9, 10).

[0049]

[0050]

[0051] is the loss function of the SL network relative to the SH network, where denotes the pseudo-labels generated by the SL network.

[0052] is the loss function of the SH network relative to the SL network, where denotes the pseudo-labels generated by the SH network, and j is the number of unlabeled images.

[0053] Furthermore, the specific steps in generating the mixed samples include:

[0054] Obtain the output of the CFC-Net model network,

[0055] Calculate the joint probability distribution P and the uncertainty map U,

[0056] Perform image block division on the low-frequency image, high-frequency image, full-frequency image, and uncertainty map,

[0057] Obtain the position information of the high-confidence patches to get the replacement position information,

[0058] The steps for obtaining the high-confidence patches are as follows: Sort the patches after dividing the uncertainty map U in descending order of the joint probability distribution, and select 25% - 35% of the patches with the top rankings.

[0059] According to the replacement position information, perform bidirectional cyclic mixing among the low-frequency, high-frequency, and full-frequency image patches at the same position to generate mixed training samples.

[0060] The mixing operation only replaces the image patches at the same position, retaining the integrity of the image structure.

[0061] Furthermore, the steps in the second-stage training include:

[0062] Input the mixed training samples into the CFC-Net network, perform the supervised and unsupervised training processes, and obtain the supervised loss function L in the second-stage training sup2 and the unsupervised loss function L in the second training stage unsup2 .

[0063] Furthermore, the steps in the model output include:

[0064] The total loss function is composed of the total supervised loss L sup and the total unsupervised loss L unsup weighted, and the weights are adjusted by the Gaussian warm-up function.

[0065] The total loss function is obtained by calculating through formula (2).

[0066]

[0067] where λ is a regularization term used to control the weight of the unsupervised loss, defined by the Gaussian warm-up function and calculated through formula (3).

[0068]

[0069] where t represents the current iteration, represents the value of λ at the t-th iteration. is the total number of iterations, is a hyperparameter representing the maximum value of λ,

[0070] The total supervised loss L sup is the supervised loss function L for the first-stage training sup1 and the supervised loss function L for the second-stage training sup2 sum,

[0071] The total unsupervised loss L unsup is the unsupervised loss function for the first training stage Lunsup1 , and the unsupervised loss function for the second training stage Lunsup2 sum,

[0072] The supervised loss function L for the first-stage training sup1 and the unsupervised loss function for the first training stage Lunsup1 are output by the first-stage training,

[0073] The supervised loss function L for the second-stage training sup2 and the unsupervised loss function for the second training stage Lunsup2 are output by the second-stage training,

[0074] The model is iteratively trained until the total loss function is minimized or the number of iterations is reached; the trained CFC-Net model is output for segmenting the uncertainty image and mining quantity and information.

[0075] Compared with the prior art, the beneficial effects of the technical solution provided by the present application.

[0076] The CFC-Net in the present application provides a semi-supervised method for segmenting dental root canals in CBCT images. This method fuses multi-frequency information to ensure the accuracy of root canal segmentation. At the same time, this method can effectively utilize a large amount of clinically unlabeled data to further improve the model performance.

[0077] The hybrid mechanism of this method ensures that the replacement positions between different frequency images are consistent. The training samples after replacement still contain complete image information and will not damage the original structure of the image due to the replacement of low-certainty image blocks by high-certainty image blocks.

[0078] This method will provide important insights into the number and morphology of the canals, enabling clinicians to formulate precise surgical plans while minimizing labor and financial costs.

[0079] We also conducted detailed experimental tests on the proposed CFC-Net for other three public oral cavity segmentation tasks. The results prove that our method can also demonstrate certain robustness in other oral cavity segmentation tasks, outperforming the existing image segmentation methods. Description of the Drawings

[0080] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0081] Among them:

[0082] Figure 1 is the flowchart of the semi-supervised oral cavity image segmentation method for cross-frequency collaborative training in this application;

[0083] Figure 2 is the network structure diagram of the CFC-Net model in this application;

[0084] Figure 3 is the visual comparison diagram of the CFC-Net in the root canal segmentation task and other semi-supervised segmentation methods in this application;

[0085] Figure 4 is the visual comparison diagram of the CFC-Net on the TDD dataset and other semi-supervised segmentation methods in this application;

[0086] Figure 5 is the visual comparison diagram of the CFC-Net on the CTooth dataset and other semi-supervised segmentation methods in this application;

[0087] Figure 6 is the visual comparison diagram of the CFC-Net on the NKUT dataset and other semi-supervised segmentation methods in this application;

[0088] Figure 7 is the quantitative comparison diagram of the CFC-Net in the root canal segmentation task and other semi-supervised segmentation methods in this application;

[0089] Figure 8 is the quantitative comparison diagram of the CFC-Net on the TDD dataset and other semi-supervised segmentation methods in this application;

[0090] Figure 9 is the quantitative comparison diagram of the CFC-Net on the CTooth dataset and other semi-supervised segmentation methods in this application;

[0091] Figure 10This is a quantitative comparison graph of CFC-Net with other semi-supervised segmentation methods on the NKUT dataset in this application. Detailed implementation manners

[0092] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0093] As Figure 1 shown, the flow of the semi-supervised oral cavity image segmentation method with cross-frequency collaborative training in this application is shown.

[0094] As Figure 2 shown, two main structures of CFC-Net in this application are shown, which are the cross-frequency collaborative module (CFC-MT) shown on the left and the uncertainty cross-frequency mixing mechanism module (UCF-Mix) shown on the right, respectively.

[0095] In CFC-MT, a collaborative training architecture using two specialized students (SS networks) and one comprehensive teacher (CT network) is used to complete semi-supervised learning training, and the structures of the three networks are the same.

[0096] In the SS network, the one for processing low-frequency images is marked as the SL network, and the one for processing high-frequency images is marked as the SH network.

[0097] In UCF-Mix, the low-frequency image, high-frequency image, and full-frequency image obtained by wavelet transform of the original image are acquired, and the high-certainty image patches at the same positions in the low-frequency image, high-frequency image, and full-frequency image are circularly mixed bidirectionally, while enhancing the network robustness, the integrity of the target structure is retained.

[0098] The following elaborates in detail the specific steps of model training:

[0099] For the clinical dental dataset, first use CFC-Net for the first round of network training; a segmentation map is obtained.

[0100] During training, for the labeled images, cross-entropy loss and DICE loss are jointly used for supervision;

[0101] For the unlabeled images, two sets of loss functions are used for supervision, which are the cross-frequency consistency loss between the two specialized students and the full-frequency consistency loss of the comprehensive teacher for the two specialized students respectively;

[0102] As Figure 2As shown in the left part, given an input image, the first step is to apply the wavelet transform described in formula (1) to extract its frequency components: LL, HL, LH, and HH,

[0103]

[0104] where represents the wavelet coefficients, (n, m) represents the pixels in the EF image, and (x, y) represents the coordinates in the domain.

[0105] LL is the low-frequency component of the image, the low-frequency image is represented by LF, (LH, HL, HH) are the three groups of high-frequency components of the image, the high-frequency image composed of high-frequency components is represented by HF, and the original image is the full-frequency image represented by EF.

[0106] For labeled data, all networks are supervised and trained using the labels.

[0107] For unlabeled data, we use two loss functions: and to optimize the unsupervised training.

[0108] is the full-frequency consistency loss function for the comprehensive teacher to two expert students respectively; is the cross-frequency consistency loss function between two expert students;

[0109] The loss function for the entire training process consists of two parts, namely and , as shown in formula 2.

[0110]

[0111] where λ is a regularization term used to control the weight of the unsupervised loss, defined by a Gaussian warm-up function, as shown in formula 3:

[0112]

[0113] where t represents the current iteration, represents the value of λ at the t-th iteration, is the total number of iterations, is a hyperparameter representing the maximum value of λ.

[0114] For training with labeled data:

[0115] The SS network and the CT network are trained using the labeled data; for the paired LF, EF, and HF components, all components are trained using the same label.

[0116] The supervised loss function L in the first training stage sup1 is defined in Equation (4).

[0117]

[0118] where represents the segmentation loss. The term , and represent the outputs of the low-frequency SS network, the high-frequency SS network, and the CT network respectively, while represents the corresponding label. i is the serial number of the labeled image, and m is the total number of labeled images.

[0119] For the training of unlabeled data:

[0120] In the unsupervised path, the parameters of the CT network are updated using the combined EMA of the two SS networks, as defined in Equation (5),

[0121]

[0122] where α and β are smoothing coefficients that control the EMA update rate and the relative contributions of the high-frequency and low-frequency SS networks to the EMA respectively. represents the weights of the corresponding network at the t-th iteration, sl for low-frequency and sh for high-frequency. θ with both superscript and subscript represents the parameters of the corresponding sl network or sh network at the t-th iteration.

[0123] The unsupervised loss function L in the first training stage unsup1 consists of two parts, as defined in Equation (6).

[0124]

[0125] The pseudo-labels generated by the CT network are used to supervise the outputs of the two SS networks simultaneously.

[0126] The main objective of

[0127] is to enable the CT network to guide and constrain the training process of the SS network from the full-frequency perspective, thereby preventing the accumulation of errors and mistakes during the training process. As shown in Equations 7 and 8, the equations of

[0128]

[0129]

[0130] In two SS networks generate pseudo-labels for each other to achieve mutual supervision.

[0131] is the loss function of the CT network with respect to the low-frequency SS network, where represents the pseudo-labels generated by the CT network.

[0132] is the loss function of the CT network with respect to the high-frequency SS network, where represents the pseudo-labels generated by the CT network.

[0133] j is the serial number of the unlabeled image, and n is the total number of unlabeled images.

[0134] This consistency mechanism allows the SS networks to collaborate in learning, leveraging their respective advantages to make up for weaknesses, correct mistakes, and avoid isolated learning.

[0135] As shown in Equations 9 and 10, the formula for is given;

[0136]

[0137]

[0138] is the loss function of the low-frequency SS network with respect to the high-frequency SS network, where represents the pseudo-labels generated by the low-frequency SS network;

[0139] is the loss function of the high-frequency SS network with respect to the low-frequency SS network, where represents the pseudo-labels generated by the high-frequency SS network;

[0140] j is the number of unlabeled images.

[0141] The specific steps for UCF-Mix to generate mixed training samples are described as follows.

[0142] Use labeled data to generate supervised mixed samples;

[0143] Use unlabeled data to generate unsupervised samples.

[0144] Taking the supervised mixed samples as an example to illustrate the generation process of UCF-Mix; the process of generating new unsupervised samples is similar. That is

[0145] Receive the outputs obtained from the SS and CT networks of 、 and , calculate the joint probability distribution P and the uncertainty map U.

[0146] The joint probability distribution P is defined as in Equation (11).

[0147] The uncertainty map U is calculated using Equation (12), where c represents the number of segmentation categories, and ϵ is a small parameter introduced to prevent the logarithm calculation from approaching zero.

[0148]

[0149]

[0150] In the formula, , and are the outputs obtained from the SS and CT networks; where c represents the number of segmentation categories, and ϵ is a small parameter introduced.

[0151] The specific implementation method is as shown in the Figure 2 right part. The low-frequency component image, full-frequency component image, high-frequency component image, and uncertainty map U are all segmented into multiple patches according to the same segmentation principle;

[0152] Obtain the position information of the high-confidence patches to get the replacement position information;

[0153] The steps for obtaining high-confidence patches are as follows: Sort the patches after segmenting the uncertainty map U according to the joint probability distribution from high to low, and select 25% - 35% of the patches with the top rankings.

[0154] Perform bilateral mixing on the patches in the low-frequency component image, full-frequency component image, and high-frequency component image corresponding to the replacement patch positions of the uncertainty map U to obtain mixed training samples; the specific mixing operation is as follows:

[0155] In the first round, mix the selected patches in the forward order of sl → t → sh → sl to obtain the first mixed training samples: EF1, HF1, LF1.

[0156] In the second round, mix the same patches in the reverse order of sh → t → sl → sh to obtain the second mixed training samples: EF2, LF2, HF2.

[0157] Taking EF2 as an example, replace the corresponding patch image in EF with the patch corresponding to the HF image of sh according to the replacement patch position information to generate EF2. The acquisition of similar mixed training samples will not be elaborated here.

[0158] Input the mixed samples into CFC-MT for the second-stage training,

[0159] as Figure 2 shown in the middle part; the training process is the same as that of the first-stage training, only the training samples obtained are different. The supervised and unsupervised loss functions in this stage are respectively represented as the supervised loss function L sup2 in the second-stage training and the unsupervised loss function Lunsup2 in the second training stage. Their forms are the same as the calculation processes of L sup1 and L unsup1 , only the training samples used are different, which are distinguished by marking and will not be elaborated here.

[0160] It should be noted that in the second-stage training, use the mixed training sample one and the mixed training sample two to conduct one training iteration respectively.

[0161] The total loss of the entire training process, the total supervised loss L sup and the total unsupervised loss L unsup are the sum of the losses in the two stages, as defined in Formulas (13) and (14).

[0162]

[0163]

[0164] It is worth noting that UCF-Mix only mixes the blocks at the same positions in different frequency components. Therefore, it does not need to modify any training labels, thus retaining the structural integrity of the target.

[0165] The optimization objective is the loss function of the entire training process Iteratively converge to the set value or reach the number of iterations to obtain the model parameters. The model is used to segment the input dental images and mine the medical information therein.

[0166] The module for segmenting and information mining of the image by the finally trained CFC-Net model is the trained teacher network, and the student network no longer participates in the image segmentation and information mining.

[0167] The parameters and data flow summary are listed in the following table:

[0168]

[0169] It should be noted that the three network architectures of BCP, UCMT, and the CFC-Net of the present application are different. Although all three are MT architectures and all use the mix-up mechanism, the difference lies in their mixing methods and purposes. BCP mixes random patches of labeled and unlabeled data, emphasizing the consistency of the output distributions of labeled and unlabeled data; UCMT replaces low-uncertainty blocks with high-certainty blocks in a single image, emphasizing high confidence, but it may damage the structure and integrity of the image. CFC-Net uses the uncertainty maps generated by mixing three networks, selects the K blocks with the highest comprehensive confidence among the results of the three networks, and rotates them at the same positions among the three, emphasizing enhancing the performance of cross-frequency learning and representation while ensuring high confidence.

[0170] Furthermore, it needs to be explained that compared with the mixing methods of the former two, the mixing method selected in the present application has the integrity of labels because the blocks at the same position are exchanged and mixed with each other, without causing the structural chaos of the original image. For example, after the high-frequency block No. 1 and the full-frequency block No. 1 are exchanged with each other, the high-frequency image contains full-frequency information, and the full-frequency image contains prominent high-frequency information. However, due to the same position, the labels output by the two still remain consistent.

[0171] The UCMT mixing replaces low-uncertainty blocks with high-certainty blocks in the same image. On the one hand, this will lose relevant information, and on the other hand, it will also damage the original structure of the original image. For example, the network has high certainty in identifying the center of a tumor and low certainty in learning the edge of the tumor. UCMT directly cuts out the tumor center block and pastes it to the position of the tumor edge. Intuitively, the uncertainty in the figure is reduced in this way, but the price is that the integrity of the lesion is damaged and the image content is lost.

[0172] In the extreme case, the algorithm uses the part in the middle with the highest recognition certainty and replaces the entire image with the same block. In this way, when the image is output, the uncertainty will be very low because each block of it is a block that performs well in the entire image. However, precisely due to such replacement, the integrity of the image is damaged. There may be non-tumor information such as skin and hair in the original image. Although the image recognition has low certainty, we still need the information of the block when analyzing the image information. Otherwise, this kind of mixing will lose some necessary analysis information, resulting in incomplete information and even misjudgment or wrong judgment.

[0173] In CFC, improving the high certainty of the image is not emphasized. Instead, from the perspective of cross-frequency, more detailed information of the image is obtained without reducing the certainty of the image. The information of low-frequency, high-frequency, and full-frequency images is combined to calculate the joint probability map to obtain the joint uncertainty map. On this uncertainty map, blocks with high certainty and good recognition are selected, and forward and reverse cycles are performed among the three images for cross-information training.

[0174] Using different data sets, the segmentation method proposed in this application and the existing segmentation methods are compared, and the comparison results are shown by Figures 3 - 6 displayed, and the quantitative analysis comparison is shown by Figures 7 - 10 displayed. It can be seen from Figures 3 - 10 that the image information after segmentation is completely retained without loss or abnormal excess, and at the same time, it is at the forefront in terms of overall performance in the quantitative information, and most of the indicators rank first.

[0175] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0176] The above-described embodiments merely represent several implementation manners of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application.

Claims

1. A semi-supervised oral cavity image segmentation method for cross-frequency collaborative training, characterized in that, It includes the following steps: S1. Data preprocessing: Obtain the images in the dataset, and obtain pre-training samples through wavelet transform. The pre-training samples include low-frequency images, high-frequency images, and full-frequency images; S2. First-stage training: Input the pre-training samples into the CFC-Net model for the first-stage training, Output the loss function of the first-stage training; S3. Generation of mixed samples: Obtain the output of the CFC-Net model network, Calculate the joint probability distribution P and the uncertainty map U, Perform image block division on the low-frequency images, high-frequency images, full-frequency images, and uncertainty map, Obtain the position information of the high-confidence patches to get the replacement position information, The step of obtaining the high-confidence patches is to sort the patches after dividing the uncertainty map U according to the joint probability distribution from high to low, and select the top 25% - 35% of the patches, According to the replacement position information, perform bidirectional cyclic mixing among the low-frequency, high-frequency, and full-frequency image patches at the same position to generate mixed training samples, The mixing operation only replaces the image patches at the same position, and retains the integrity of the image structure; S4. Second-stage training: Receive the mixed training samples and input them into the CFC-Net model for supervised training and unsupervised training, Output the loss function of the second-stage training; S5. Model output: Calculate the total loss function through the loss function of the first-stage training and the loss function of the second-stage training, Obtain the final model through the iterative convergence of the total loss function, Use the final model to segment the images.

2. The semi-supervised oral cavity image segmentation method with cross-frequency collaborative training according to claim 1, characterized in that The CFC-Net model includes a specialized student network SL network for processing low-frequency images, a specialized student network SH network for processing high-frequency images, and a comprehensive teacher network CT network for processing full-frequency images. The three networks cooperate to complete semi-supervised learning training, and the two specialized student networks and the comprehensive teacher network have the same structure.

3. The semi-supervised oral cavity image segmentation method with cross-frequency collaborative training according to claim 2, characterized in that In S1. Data preprocessing, Apply wavelet transform to each image using formula (1) to extract low-frequency images, high-frequency images, and full-frequency images, Among them represents the wavelet coefficient, (n, m) represents the pixel in the EF image, and (x, y) represents the coordinates in the domain, LL is the low-frequency component of the image, the low-frequency image is represented by LF, (LH, HL, HH) are the three groups of high-frequency components of the image, the high-frequency image composed of high-frequency components is represented by HF, and the original image is a full-frequency image represented by EF.

4. The semi-supervised oral cavity image segmentation method with cross-frequency collaborative training according to claim 3, characterized in that The image data in the dataset includes labeled data and unlabeled data; Calculate the cross-frequency consistency loss using the pseudo-labels generated mutually between two specialized student networks , The pseudo-labels generated by the comprehensive teacher network supervise the calculation of the full-frequency consistency loss by two expert student networks , The unsupervised loss function L in the first training stage unsup1 is the sum of the full-frequency consistency loss and the cross-frequency consistency loss; Perform supervised training on the labeled data, and calculate the cross-entropy loss and DICE loss of the SL network, SH network, and CT network; The supervised loss function L in the first training stage sup1 is the sum of the segmentation losses of the three networks respectively, calculated using formula (4), with the LF, HF, and EF components as inputs and sharing the same label Among them represents the segmentation loss, and the terms , and represent the outputs of the SL network, the SH network, and the CT network respectively, while represents the corresponding label; i is the serial number of the labeled image, and m is the total number of labeled images.

5. The semi-supervised oral cavity image segmentation method with cross-frequency collaborative training according to claim 4, characterized in that Perform unsupervised training on the unlabeled data: In the unsupervised path, the parameters of the CT network are updated using the combined EMA of the SL network and the SH network, and are calculated by formula (5), where α and β are smoothing coefficients that control the EMA update rate and the relative contributions of the SL network and the SH network to the EMA, respectively, represents the weights of the corresponding network at the t-th iteration, sl is low frequency, and sh is high frequency; θ with both superscripts and subscripts represents the parameters of the corresponding SL network or SH network at the t-th iteration; Use the CT network to generate pseudo-labels, and constrain the outputs of the two specialized student networks through the full-frequency consistency loss, Realize the mutual supervision of the two specialized student networks through the cross-frequency consistency loss to avoid error accumulation, The full-frequency consistency loss constrains the expertise student network through the pseudo-labels generated by the integrated teacher network; The CT network generates pseudo-labels to supervise the SS network and is optimized using formulas (7, 8). is the loss function of the CT network relative to the SL network, where represents the pseudo-labels generated by the CT network, is the loss function of the CT network relative to the SH network, where represents the pseudo-labels generated by the CT network j is the serial number of the unannotated image, and n is the total number of unannotated images; The cross-frequency consistency loss is achieved by the two expertise student networks generating pseudo-labels for each other. The SL network and the SH network generate pseudo-labels for each other for collaborative learning and are calculated using formulas (9, 10). is the loss function of the SL network relative to the SH network, where represents the pseudo-labels generated by the SL network, is the loss function of the SH network relative to the SL network, where represents the pseudo-labels generated by the SH network, and j is the number of unlabeled images.

6. The semi-supervised oral cavity image segmentation method with cross-frequency collaborative training according to claim 1, characterized in that S4. The steps of the second-stage training include: Input the mixed training samples into the CFC-Net network to perform the supervised and unsupervised training processes, and obtain the supervised loss function \(L\) for the second-stage training sup2 and the unsupervised loss function \(L\) for the second training stage unsup2 .

7. The semi-supervised oral cavity image segmentation method with cross-frequency collaborative training according to claim 6, characterized in that S5. The model output steps include: The total loss function is composed of the total supervised loss L sup and the total unsupervised loss L unsup weighted, and the weights are adjusted by a Gaussian warm-up function Total loss function is obtained by calculation through formula (2). where λ is a regularization term used to control the weight of the unsupervised loss, defined by a Gaussian warm-up function and calculated through formula (3). where \(t\) represents the current iteration, represents the value of \(\lambda\) at the \(t\)-th iteration, is the total number of iterations, is a hyperparameter representing the maximum value of \(\lambda\), The total supervised loss L sup is the sum of the supervised loss function L sup1 for the first-stage training and the supervised loss function L sup2 for the second-stage training. The total unsupervised loss L unsup is the unsupervised loss function in the first training stage Lunsup1 , and the unsupervised loss function in the second training stage Lunsup2 , and their sum The supervised loss function L of the first-stage training sup1 and the unsupervised loss function of the first training stage Lunsup1 are output by the first-stage training The supervised loss function L for the second-stage training sup2 and the unsupervised loss function for the second training stage Lunsup2 are output by the second-stage training The model is iteratively trained until the total loss function is minimized or the number of iterations is reached; the finally trained CFC-Net model is output, and the finally trained CFC-Net model is used to segment the image and mine information.

Citation Information

Patent Citations

  • Semi-supervised semantic segmentation method based on scale perception attention

    CN115661463A

  • Semi-supervised medical image segmentation method based on consistency loss function

    CN116258730A