Airway image segmentation model training method, airway image segmentation method and device

Through MDC and SMPS regularization and selective cross-supervised learning of heterogeneous subnetwork, confirmation bias and pseudo-label quality problems in semi-supervised segmentation of airway images are solved, and segmentation accuracy and model generalization capabilities are improved.

CN120259335APending Publication Date: 2025-07-04SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510216360.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the existing semi-supervised segmentation method of airway images, subnetworks with the same structure limit the prediction diversity, resulting in confirmation bias problems, pseudo-label quality has a significant impact on segmentation performance, and manual labeling costs are high.

Method used

Using heterogeneous first and second subnets, the missegmented areas of labeled images are identified through MDC regularization, and SMPS regularization recognizes reliable pseudo-labels of unlabeled images. Combined with selective cross-supervised learning, the loss function is optimized to update network parameters.

Benefits of technology

The prediction accuracy and segmentation accuracy of the airway image segmentation model are improved, the generalization ability of the model is enhanced, the confirmation bias is alleviated, and the pseudo-label is effectively utilized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259335A_ABST
    Figure CN120259335A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of airway image segmentation, and provides an airway image segmentation model training method and an airway image segmentation method and device.The method comprises the steps that batch airway images are input into two heterogeneous sub-networks respectively, and prediction results are obtained; determining an MDC regular term of the two sub-networks in the prediction inconsistent area according to the hard tag in the prediction result, and calculating the supervised loss of the hard tag according to the MDC regular term; determining SMPS masks of the two sub-networks according to the confidence coefficient of the pseudo tag in the prediction result, calculating the similarity between voxel features learned by the target layer of the two sub-networks and the prototype, and determining the weight of the pseudo tag according to the similarity; calculating unsupervised loss of the pseudo label according to the SMPS mask and the weight; determining the total loss of the prediction result according to the supervised loss and the unsupervised loss; and updating network parameters of the two sub-networks according to the total loss. According to the embodiment of the invention, the prediction precision of the airway image segmentation model can be improved, and the segmentation precision of airway image segmentation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the technical field of airway image segmentation, and in particular, to a method for training an airway image segmentation model, an airway image segmentation method, and an apparatus therefor. Background Art

[0002] An airway image refers to an image of an airway structure (including a series of tubular structures in the upper and lower airways) obtained through medical imaging techniques (such as CT, ultrasound, etc.). In clinical applications such as computer-aided diagnosis, image-guided intervention, and radiotherapy, automatically and accurately segmenting airway images is of great significance.

[0003] In recent years, techniques for segmenting airway images using pre-trained neural network models have emerged. For example, an airway image segmentation model obtained by training a convolutional neural network using a semi-supervised learning method. The semi-supervised learning method has shown good performance in three-dimensional airway image segmentation by simultaneously using a limited amount of labeled data and abundant unlabeled data.

[0004] Since the semi-supervised learning method results in unlabeled images lacking true label values, and manually labeling a large-scale three-dimensional airway image in its entirety would consume a large amount of human resources and time costs, the semi-supervised learning method based on pseudo-labels has gradually become an important strategy for semi-supervised segmentation of airway images. Such methods usually use sub-networks with the same structure, and by introducing input perturbations, feature perturbations, or network perturbations, a consistency learning strategy is adopted between the sub-networks to encourage them to produce consistent predictions (pseudo-labels) for unlabeled images. However, the sub-networks with the same structure limit the diversity of predictions for labeled and unlabeled images, making it difficult for the sub-networks to correct each other, and causing the existing semi-supervised segmentation of airway images to possibly have a confirmation bias problem. On the other hand, the quality of pseudo-labels has a significant impact on the performance of semi-supervised segmentation of airway images. Summary of the Invention

[0005] The purpose of the embodiments of this specification is to provide a method for training an airway image segmentation model, an airway image segmentation method, and an apparatus therefor, so as to improve the prediction accuracy of the airway image segmentation model and the segmentation accuracy of airway image segmentation.

[0006] To achieve the above purpose, on the one hand, the embodiments of this specification provide a method for training an airway image segmentation model, including:

[0007] Inputting batch images extracted from an airway image dataset into heterogeneous first and second sub-networks respectively to obtain prediction results; the batch images include labeled images and unlabeled images, and the prediction results include hard labels corresponding to the labeled images and pseudo-labels corresponding to the unlabeled images;

[0008] Determine the MDC regularization terms corresponding to the first sub-network and the second sub-network in the prediction inconsistent region;

[0009] Calculate the supervised loss of the hard labels according to the corresponding MDC regularization terms;

[0010] Determine the SMPS masks of the first sub-network and the second sub-network according to the confidence of the pseudo-labels;

[0011] Calculate the similarity between the voxel features learned by the target layers of the first sub-network and the second sub-network and the prototypes according to the SMPS masks, and determine the weights of the pseudo-labels according to the similarity;

[0012] Calculate the unsupervised loss of the pseudo-labels according to the SMPS masks and the weights;

[0013] Determine the total loss of the prediction results according to the supervised loss and the unsupervised loss;

[0014] Update the network parameters of the first sub-network and the second sub-network according to the total loss;

[0015] Iteratively execute the above steps to obtain the first sub-network and the second sub-network that meet the preset conditions, and determine them as the airway image segmentation model.

[0016] In the airway image segmentation model training method of the embodiments of this specification, before respectively inputting the batch of images extracted from the airway image dataset into the heterogeneous first sub-network and second sub-network, it further includes:

[0017] Perform truncated normalization processing on the images in the airway image dataset.

[0018] In the airway image segmentation model training method of the embodiments of this specification, determining the mutual difference correction MDC regularization terms corresponding to the first sub-network and the second sub-network in the prediction inconsistent region according to the hard labels includes:

[0019] Determine the masks corresponding to the first sub-network and the second sub-network in the prediction inconsistent region according to the hard labels;

[0020] Calculate the segmentation results obtained by the first sub-network and the second sub-network in the prediction inconsistent region according to the masks, and the true labels corresponding to the prediction inconsistent region;

[0021] Calculate the mean square errors between the segmentation results of the first sub-network and the second sub-network and the true labels respectively, and correspondingly use them as the MDC regularization terms of the first sub-network and the second sub-network in the prediction inconsistent region.

[0022] In the airway image segmentation model training method according to the embodiments of this specification, calculating the supervised loss of the hard label corresponding to the MDC regularization term includes:

[0023] Calculating the supervised loss of the hard label predicted by the first sub-network in the prediction inconsistent region according to the formula ; and,

[0024] Calculating the supervised loss of the hard label predicted by the second sub-network in the prediction inconsistent region according to the formula ;

[0025] wherein, is the supervised loss of the hard label predicted by the first sub-network in the prediction inconsistent region, is the hard label predicted by the first sub-network in the prediction inconsistent region, Y l is the true label, is the Dice loss with respect to Y l , is the Focal loss with respect to Y l , is the MDC regularization term of the first sub-network in the prediction inconsistent region, and β is the balance parameter, is the supervised loss of the hard label predicted by the second sub-network in the prediction inconsistent region, is the hard label predicted by the second sub-network in the prediction inconsistent region, is the Dice loss with respect to Y l , is the Focal loss with respect to Y l , is the MDC regularization term of the second sub-network in the prediction inconsistent region.

[0026] In the airway image segmentation model training method according to the embodiments of this specification, determining the selective mutual pseudo-supervision SMPS masks of the first sub-network and the second sub-network according to the confidence of the pseudo-labels includes:

[0027] Calculating the prediction entropy value of the voxel features of the unlabeled image by the first sub-network according to the formula and using it as the prediction confidence of the voxel features of the unlabeled image by the first sub-network;

[0028] Calculating the prediction entropy value of the voxel features of the unlabeled image by the second sub-network according to the formula and using it as the prediction confidence of the voxel features of the unlabeled image by the second sub-network;

[0029] When C A <C B , according to the formula determine the SMPS masks of the first sub-network and the second sub-network; or,

[0030] When C A >C B , according to the formula determine the SMPS masks of the first sub-network and the second sub-network;

[0031] wherein, C A is the predicted entropy value of the voxel features of the unlabeled image by the first sub-network, is the predicted output of the voxel features of the unlabeled image by the first sub-network, T is the matrix transpose, C B is the predicted entropy value of the voxel features of the unlabeled image by the second sub-network, is the predicted output of the voxel features of the unlabeled image by the second sub-network, is the unlabeled image, M A is the SMPS mask of the first sub-network, I represents the indicator function, H, W, and L are the height, width, and length of the voxel respectively, M B is the SMPS mask of the second sub-network, and I is a matrix with all elements being 1.

[0032] In the airway image segmentation model training method of the embodiments of this specification, calculating the similarity between the voxel features learned by the target layers of the first sub-network and the second sub-network and the prototype according to the SMPS mask includes:

[0033] According to the formula calculate the prototype to which the voxel features learned by the penultimate layer of the first sub-network belong to the c-th class;

[0034] According to the formula calculate the prototype to which the voxel features learned by the penultimate layer of the second sub-network belong to the c-th class;

[0035] According to the formula calculate the cosine similarity between the voxel features learned by the penultimate layer of the first sub-network belonging to the c-th class and the prototype;

[0036] According to the formula calculate the cosine similarity between the voxel features learned by the penultimate layer of the second sub-network belonging to the c-th class and the prototype;

[0037] wherein, is the prototype to which the voxel features learned by the penultimate layer of the first sub-network belong to the c-th class, The SMPS mask corresponding to the prototype where the voxel features learned in the penultimate layer of the first sub-network belong to the c-th class is the number of non-zero elements in The voxel features learned in the penultimate layer of the first sub-network that belong to the c-th class, e represents element-wise multiplication, h, w, and l are the height, width, and length of the voxel The prototype where the voxel features learned in the penultimate layer of the second sub-network belong to the c-th class The SMPS mask corresponding to the prototype where the voxel features learned in the penultimate layer of the second sub-network belong to the c-th class is the number of non-zero elements in The cosine similarity between the voxel features learned in the penultimate layer of the first sub-network that belong to the c-th class and the prototype, cos represents cosine similarity The cosine similarity between the voxel features learned in the penultimate layer of the second sub-network that belong to the c-th class and the prototype The voxel features learned in the penultimate layer of the second sub-network that belong to the c-th class

[0038] In the method for training the airway image segmentation model according to the embodiments of this specification, determining the weight of the pseudo-label according to the similarity includes:

[0039] According to the formula Calculate the weight of the pseudo-label predicted by the first sub-network;

[0040] According to the formula Calculate the weight of the pseudo-label predicted by the second sub-network;

[0041] where G A is the weight of the pseudo-label predicted by the first sub-network, C is the C-th class of the pseudo-label, sim represents similarity is the cosine similarity between the voxel features learned in the penultimate layer of the first sub-network that belong to the c-th class and the prototype is the predicted output of the voxel features of the unlabeled image by the first sub-network, G B is the weight of the pseudo-label predicted by the second sub-network is the cosine similarity between the voxel features learned in the penultimate layer of the second sub-network that belong to the c-th class and the prototype is the predicted output of the voxel features of the unlabeled image by the second sub-network

[0042] In the method for training the airway image segmentation model according to the embodiments of this specification, calculating the unsupervised loss of the pseudo-label according to the SMPS mask and the weight includes:

[0043] According to the formula Calculate the unsupervised loss of the prediction result of the first sub-network; and,

[0044] According to the formula Calculate the unsupervised loss of the prediction result of the first sub-network;

[0045] Wherein, is the unsupervised loss of the prediction result of the first sub-network, M B is the SMPS mask of the second sub-network, e represents element-wise multiplication, G A is the weight of the pseudo-label predicted by the first sub-network, is the predicted output of the voxel features of the unlabeled image by the first sub-network, is the pseudo-label predicted by the second sub-network for the voxel features, CE is the cross-entropy loss, is the unsupervised loss of the prediction result of the first sub-network, M A is the SMPS mask of the first sub-network, G B is the weight of the pseudo-label predicted by the second sub-network, is the predicted output of the voxel features of the unlabeled image by the second sub-network, is the pseudo-label predicted by the first sub-network for the voxel features.

[0046] In the method for training an airway image segmentation model according to an embodiment of this specification, determining the total loss of the prediction result according to the supervised loss and the unsupervised loss includes:

[0047] According to the formula Calculate the total loss of the prediction result of the first sub-network; and,

[0048] According to the formula Calculate the total loss of the prediction result of the first sub-network;

[0049] Wherein, L A is the total loss of the prediction result of the first sub-network, is the supervised loss of the prediction result of the first sub-network, is the unsupervised loss of the prediction result of the first sub-network, λ is the weight, L B is the total loss of the prediction result of the second sub-network, is the supervised loss of the prediction result of the second sub-network, is the unsupervised loss of the prediction result of the second sub-network.

[0050] In the method for training an airway image segmentation model according to an embodiment of this specification, the first sub-network and the second sub-network include a three-dimensional Unet network and a VNet network.

[0051] On the other hand, an embodiment of this specification also provides an airway image segmentation method, including:

[0052] Obtain the airway image to be segmented;

[0053] Input the airway image into a pre-trained airway image segmentation model to obtain the segmentation result of the airway image; wherein, the airway image segmentation model is trained according to the above model training method.

[0054] On the other hand, an embodiment of this specification also provides an airway image segmentation model training device, including:

[0055] A data input module, configured to respectively input a batch of images extracted from an airway image dataset into heterogeneous first and second sub-networks to obtain prediction results; the batch of images includes labeled images and unlabeled images, and the prediction results include hard labels corresponding to the labeled images and pseudo-labels corresponding to the unlabeled images;

[0056] An MDC regularization module, configured to determine the MDC regularization terms corresponding to the first and second sub-networks in the prediction inconsistent regions;

[0057] A supervised loss determination module, configured to calculate the supervised loss of the hard labels according to the corresponding MDC regularization terms;

[0058] An SMPS regularization module, configured to determine the SMPS masks of the first and second sub-networks according to the confidence levels of the pseudo-labels;

[0059] A weight determination module, configured to calculate the similarity between the voxel features learned by the target layers of the first and second sub-networks and the prototypes according to the SMPS masks, and determine the weights of the pseudo-labels according to the similarity;

[0060] An unsupervised loss determination module, configured to calculate the unsupervised loss of the pseudo-labels according to the SMPS masks and the weights;

[0061] A total loss determination module, configured to determine the total loss of the prediction results according to the supervised loss and the unsupervised loss;

[0062] A network parameter update module, configured to update the network parameters of the first and second sub-networks according to the total loss;

[0063] An iteration control module, configured to iteratively execute the above modules to obtain the first and second sub-networks that meet the preset conditions, and determine them as the airway image segmentation model.

[0064] On the other hand, an embodiment of this specification also provides an airway image segmentation device, including:

[0065] An image acquisition module for acquiring airway images to be segmented;

[0066] An image segmentation module for inputting the airway image into a pre-trained airway image segmentation model to obtain a segmentation result of the airway image; wherein, the airway image segmentation model is trained according to the above model training method.

[0067] On the other hand, an embodiment of this specification also provides a computer device, including a memory, a processor, and a computer program stored on the memory. When the computer program is run by the processor, it executes the instructions of the above method.

[0068] On the other hand, an embodiment of this specification also provides a computer storage medium, on which a computer program is stored. When the computer program is run by the processor of a computer device, it executes the instructions of the above method.

[0069] On the other hand, an embodiment of this specification also provides a computer program product, which includes a computer program. When the computer program is run by the processor of a computer device, it executes the instructions of the above method.

[0070] As can be seen from the technical solutions provided by the embodiments of this specification above, the semi-supervised learning framework constructed by the embodiments of this specification adopts two segmentation networks with different structures (i.e., heterogeneous first and second sub-networks), which is beneficial to enhancing the generalization ability of model prediction. By introducing MDC regularization to guide the model to identify and re-learn the regions in the labeled airway images that are prone to mis-segmentation, and by introducing SMPS regularization to identify more reliable pseudo-labels in the unlabeled airway images and select corresponding unlabeled image voxels for cross-supervision; thus, through the mutual correction and selective cross-supervised learning between the heterogeneous segmentation networks, the segmentation accuracy of the model for airway images is greatly enhanced. Description of the Drawings

[0071] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0072] Figure 1 Schematic diagrams of the application environment for airway image segmentation in some embodiments of this specification are shown;

[0073] Figure 2Shows a flowchart of a method for training an airway image segmentation model in some embodiments of this specification;

[0074] Figure 3 Shows Figure 2 A flowchart of determining the mutual difference correction MDC regularization term corresponding to the first sub-network and the second sub-network in the predicted inconsistent region in the method shown;

[0075] Figure 4 Shows a schematic diagram of the training principle of the first sub-network in an exemplary embodiment of this specification;

[0076] Figure 5 Shows a schematic diagram of the training principle of the second sub-network in an exemplary embodiment of this specification;

[0077] Figure 6 Shows a flowchart of an airway image segmentation method in some embodiments of this specification;

[0078] Figure 7 Shows a structural block diagram of an airway image segmentation model training device in some embodiments of this specification;

[0079] Figure 8 Shows a structural block diagram of an airway image segmentation device in some embodiments of this specification;

[0080] Figure 9 Shows a structural block diagram of a computer device in some embodiments of this specification.

[0081]

Explanation of Reference Numerals

[0082] 10, Client;

[0083] 20, Server;

[0084] 71, Data Input Module;

[0085] 72, MDC Regularization Module;

[0086] 73, Supervised Loss Determination Module;

[0087] 74, SMPS Regularization Module;

[0088] 75, Weight Determination Module;

[0089] 76, Unsupervised Loss Determination Module;

[0090] 77, Total Loss Determination Module;

[0091] 78, Network Parameter Update Module;

[0092] 79, Iteration Control Module;

[0093] 81. Image acquisition module;

[0094] 82. Image segmentation module;

[0095] 902. Computer device;

[0096] 904. Processor;

[0097] 906. Memory;

[0098] 908. Driving mechanism;

[0099] 910. Input / output interface;

[0100] 912. Input device;

[0101] 914. Output device;

[0102] 916. Presentation device;

[0103] 918. Graphical user interface;

[0104] 920. Network interface;

[0105] 922. Communication link;

[0106] 924. Communication bus. Detailed implementation

[0107] In order to enable those skilled in the art of this technology to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0108] It should be noted that in the embodiments of this specification, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved are all information and data that have been authorized and agreed by the user and fully authorized by all parties. That is, the acquisition, transmission, storage, use, processing, etc. of data in the technical solutions of this application all comply with the relevant provisions of national laws and regulations.

[0109] Figure 1The schematic diagram of the application environment in some embodiments of this specification is shown; the application environment includes a client 10 and a server 20. The server 20 can obtain the airway image to be segmented from the client 10; input the airway image into a pre-trained airway image segmentation model to obtain the segmentation result of the airway image. In some embodiments of this specification, the client 10 may include an image acquisition module integrated in a CT device, an MRI device, an ultrasonic detection device, an optical coherence tomography device (OCT), an electrical impedance tomography device (EIT), an infrared thermal imaging device, a 3D scanner, etc. In some embodiments of this specification, the server 20 may include image analysis and processing software integrated in a CT device, an MRI device, an ultrasonic detection device, an optical coherence tomography device (OCT), an electrical impedance tomography device (EIT), an infrared thermal imaging device, a 3D scanner, etc. In some other embodiments of this specification, the server 20 may also be image analysis and processing software independent of a CT device, an MRI device, an ultrasonic detection device, an optical coherence tomography device (OCT), an electrical impedance tomography device (EIT), an infrared thermal imaging device, a 3D scanner, etc.

[0110] An embodiment of this specification provides a method for training an airway image segmentation model, referring to Figure 2 As shown, in some embodiments of this specification, the method for training an airway image segmentation model may include the following steps:

[0111] Step 201: Input the batch of images extracted from the airway image dataset into heterogeneous first and second sub-networks respectively to obtain prediction results; the batch of images includes labeled images and unlabeled images, and the prediction results include hard labels corresponding to the labeled images and pseudo-labels corresponding to the unlabeled images.

[0112] Step 202: Determine the mutual discrepancy correction (MDC) regularization terms corresponding to the first and second sub-networks in the prediction inconsistent regions according to the hard labels.

[0113] Step 203: Calculate the supervised loss of the hard labels according to the corresponding MDC regularization terms.

[0114] Step 204: Determine the selective mutual pseudo-supervision (SMPS) masks of the first and second sub-networks according to the confidence of the pseudo-labels.

[0115] Step 205: Calculate the similarity between the voxel features learned by the target layers of the first sub-network and the second sub-network and the prototypes according to the SMPS mask, and determine the weights of the pseudo-labels according to the similarity.

[0116] Step 206: Calculate the unsupervised loss of the pseudo-labels according to the SMPS mask and the weights.

[0117] Step 207: Determine the total loss of the prediction results according to the supervised loss and the unsupervised loss.

[0118] Step 208: Update the network parameters of the first sub-network and the second sub-network according to the total loss.

[0119] Step 209: Iteratively execute the above steps 201-208 to obtain the first sub-network and the second sub-network that meet the preset conditions, and determine them as the airway image segmentation model.

[0120] The semi-supervised learning framework constructed in the embodiments of this specification adopts two segmentation networks with different structures (i.e., heterogeneous first sub-network and second sub-network), which is beneficial to enhancing the generalization ability of model prediction. By introducing MDC regularization to guide the model to identify the regions that are prone to mis-segmentation in the labeled airway images and re-learn, and by introducing SMPS regularization to identify more reliable pseudo-labels in the unlabeled airway images and select the corresponding unlabeled image voxels for cross-supervision; thus, through the mutual correction and selective cross-supervised learning between the heterogeneous segmentation networks, the segmentation accuracy of the model for airway images is greatly enhanced.

[0121] In some embodiments of this specification, the airway image generally refers to a three-dimensional airway image; the airway image dataset refers to an image data set containing a small number of labeled images (i.e., labeled airway images) and a large number of unlabeled images (i.e., unlabeled airway images). For example, in some embodiments of this specification, the airway image dataset may contain N labeled data sets and M unlabeled data sets where the i-th labeled data is denoted as and its corresponding true label is denoted as Here, H, W, and L respectively represent the height, width, and depth (number of slices) of the airway image, C represents the category to which the image voxel belongs, is an unlabeled image.

[0122] In some embodiments of the present specification, before respectively inputting the batch of images extracted from the airway image dataset into the heterogeneous first sub-network and second sub-network, it may further include performing truncated normalization processing on the images in the airway image dataset. In this way, all data can be retained as much as possible, and the influence of outliers on the normalization result can be reduced. Among them, the truncated normalization processing indicates: first perform truncation processing, and then perform normalization processing on the processing result of the truncation processing. For example, in an exemplary embodiment of the present specification, taking the airway CT image as an example, each airway CT image's CT value can be adjusted to be limited within a specified range based on functions such as the clip function, and then normalized to [0, 1]. Among them, the clip function will limit the CT value of each airway CT image between a given minimum value (for example, -1000 HU) and a maximum value (for example, 400 HU), and the values outside the range will be correspondingly mapped to the minimum value or the maximum value.

[0123] In some embodiments of the present specification, during each training, a part of the labeled images X l can be extracted from the labeled data set D l of the airway image dataset by means of hierarchical random sampling, etc., and a part of the unlabeled images X u can be extracted from the unlabeled data set D u to form a batch for the current input. Then, the airway image data of this batch is the batch of images.

[0124] In some embodiments of the present specification, the heterogeneous first sub-network and second sub-network refer to: the first sub-network and the second sub-network have different network structures. Compared with the traditional semi-supervised learning that uses homogeneous sub-networks, the heterogeneous first sub-network and second sub-network can generate more diverse prediction results for the labeled and unlabeled airway images, thereby alleviating the confirmation bias problem of the semi-supervised learning method. For example, in an exemplary embodiment of the present specification, it may include a three-dimensional Unet network and a VNet network, that is, the first sub-network is a three-dimensional Unet network and the second sub-network is a VNet network; or, the second sub-network is a three-dimensional Unet network and the first sub-network is a VNet network.

[0125] In some embodiments of the present specification, since the data input in each batch includes both labeled images and unlabeled images, the prediction results may include hard labels corresponding to the labeled images and pseudo-labels corresponding to the unlabeled images. Among them, hard labels (Hard Labels) are also called discrete labels. In a classification task, hard labels can be represented by one-hot encoding for each sample's category. Pseudo-label is a term in semi-supervised learning technology. The trained model is used to generate prediction labels (pseudo-labels) for unlabeled data, and they are combined with the labeled data to retrain the model, so as to utilize the unlabeled data to improve the model performance when the training data is insufficient.

[0126] For example, in an exemplary embodiment of the present specification, taking the first sub-network (denoted as sub-network A) as a 3D Unet network (denoted as f A (·)), and the second sub-network (denoted as sub-network B) as a VNet network (denoted as f B (·)) as an example, then the batch data X B (i.e., X l +X u ) is input into sub-network A and sub-network B respectively. θ A and θ B represent the model parameters of sub-network A and B respectively. Then the network prediction outputs are respectively expressed as:

[0127]

[0128] Among them, is the prediction result obtained by inputting X l into sub-network A, is the prediction result obtained by inputting X u into sub-network A, is the prediction result obtained by inputting X l into sub-network B, is the prediction result obtained by inputting X u into sub-network B.

[0129] Referring to Figure 3 as shown, in some embodiments of the present specification, determining the mutual difference correction MDC regularization term corresponding to the first sub-network and the second sub-network in the prediction inconsistent region according to the hard label may include the following steps:

[0130] Step 301: Determine the masks corresponding to the first sub-network and the second sub-network in the prediction inconsistent region according to the hard label.

[0131] In some embodiments of the present specification, a mask is a binary (0 or 1) or multi-channel matrix in computer graphics, image processing, or computer vision, used to mark specific regions in an image that need to be processed or retained; in terms of effect, a mask is an image operation used to partially or completely cover (or hide) an object or element. In model training, by hiding part of the input data with a mask, the model is forced to rely on context information to predict the masked part, thereby reducing the model's over-focus on specific labels and improving the model's generalization ability.

[0132] In some embodiments of the present specification, the input airway image has a mask. If the first sub-network and the second sub-network have inconsistent predictions for the same image region of the same labeled airway image, that is, the class labels output by the two sub-networks for the same image region of the same labeled airway image are different, then the corresponding mask can be located.

[0133] For example, for the labeled airway image data, by comparing the predicted hard labels of sub-network A and sub-network B, that is, by comparing and the mask corresponding to the region where the two predictions are inconsistent can be determined. where is the predicted hard label of sub-network A, is the predicted hard label of sub-network B, and I represents the indicator function.

[0134] Step 302: Calculate the segmentation results obtained by the first sub-network and the second sub-network in the prediction inconsistent region according to the mask, and the ground truth label corresponding to the prediction inconsistent region.

[0135] In some embodiments of the present specification, calculating the segmentation results obtained by the first sub-network and the second sub-network in the prediction inconsistent region according to the mask may include:

[0136] According to the formula calculate the segmentation result obtained by the first sub-network in the prediction inconsistent region.

[0137] According to the formula calculate the segmentation result obtained by the second sub-network in the prediction inconsistent region.

[0138] where R is the part of the mask M region taken from the first sub-network and the second sub-network.

[0139] Step 303: Calculate the mean squared error between the segmentation results of the first sub-network and the second sub-network and the ground truth label respectively, and use them as the MDC regularization terms of the first sub-network and the second sub-network in the prediction inconsistent region.

[0140] For the prediction inconsistent regions of the labeled images, by introducing the MDC regularization terms of subnet A and subnet B and by calculating the predicted output and the ground truth label of the mean square error (MSE), the wrong predictions of the two networks can be corrected, thereby encouraging the model to identify potential mis-segmented regions and re-learn them. Therefore, the MDC regularization of subnet A and subnet B is specifically expressed as:

[0141]

[0142] where is the MDC regularization term of subnet A, MSE is the mean square error, is the MDC regularization term of subnet B. In some other embodiments of this specification, the mean square error can also be replaced by other algorithms for quantifying the degree of difference.

[0143] In some embodiments of this specification, calculating the supervised loss of the hard label corresponding to the MDC regularization term may include:

[0144] Calculating the supervised loss of the hard label predicted by the first subnet in the prediction inconsistent region according to the formula ;

[0145] Calculating the supervised loss of the hard label predicted by the second subnet in the prediction inconsistent region according to the formula ;

[0146] where is the supervised loss of the hard label predicted by the first subnet in the prediction inconsistent region, is the hard label predicted by the first subnet in the prediction inconsistent region, Y l is the ground truth label, is the Dice loss with respect to Y l , is the Focal loss with respect to Y l , is the MDC regularization term of the first subnet in the prediction inconsistent region, β is the balance parameter, is the supervised loss of the hard label predicted by the second subnet in the prediction inconsistent region, is the hard label predicted by the second subnet in the prediction inconsistent region, is the Dice loss with respect to Yl The Dice loss is relative to Y l the Focal loss, and is the MDC regularization term of the second sub-network in the prediction inconsistent region.

[0147] In some embodiments of the present specification, determining the selective mutual pseudo-supervision SMPS mask of the first sub-network and the second sub-network according to the confidence of the pseudo-label may include:

[0148] According to the formula calculate the prediction entropy value of the voxel features of the unlabeled image by the first sub-network, and use it as the prediction confidence of the voxel features of the unlabeled image by the first sub-network;

[0149] According to the formula calculate the prediction entropy value of the voxel features of the unlabeled image by the second sub-network, and use it as the prediction confidence of the voxel features of the unlabeled image by the second sub-network.

[0150] Wherein, a voxel is short for a volume pixel, and the three-dimensional object containing voxels can be represented by volume rendering or extracting the polygon isosurface of a given threshold contour. A voxel is the smallest unit of digital data segmentation in three-dimensional space; the voxel feature is the feature corresponding to the voxel.

[0151] When the prediction confidence of a sub-network for a certain voxel is higher, the sub-network needs to use the prediction of another sub-network for supervised learning. For example, if C A > C B , it indicates that sub-network A uses the prediction result of sub-network B for supervised learning. Conversely, if C A < C B , it indicates that sub-network B uses the prediction result of sub-network A for supervised learning.

[0152] Therefore, when C A < C B , the SMPS mask of the first sub-network and the second sub-network can be determined according to the formula ;

[0153] And when C A > C B , the SMPS mask of the first sub-network and the second sub-network can be determined according to the formula . Wherein, the SMPS mask is the corresponding mask after introducing SMPS regularization.

[0154] Wherein, C A is the prediction entropy value of the voxel features of the unlabeled image by the first sub-network, is the prediction output of the voxel features of the unlabeled image by the first sub-network, T is the matrix transpose, and C B is the predicted entropy value of the voxel features of the unlabeled image by the second sub-network, is the prediction output of the voxel features of the unlabeled image by the second sub-network, is the unlabeled image, M A is the SMPS mask of the first sub-network, I represents the indicator function, H, W, and L are the height, width, and length of the voxel respectively, M B is the SMPS mask of the second sub-network, and I is a matrix with all elements being 1.

[0155] That is, through M A and M B it is possible to determine which pseudo-labels are reliable, thereby alleviating the impact of unreliable pseudo-labels on model learning and helping to prevent the two sub-networks from falling into learning collapse.

[0156] In some embodiments of this specification, calculating the similarity between the voxel features learned by the target layers of the first sub-network and the second sub-network and the prototypes according to the SMPS mask includes:

[0157] According to the formula calculate the prototype to which the voxel features learned by the penultimate layer of the first sub-network belong to the c-th class;

[0158] According to the formula calculate the prototype to which the voxel features learned by the penultimate layer of the second sub-network belong to the c-th class;

[0159] According to the formula calculate the cosine similarity between the voxel features learned by the penultimate layer of the first sub-network belonging to the c-th class and the prototype; where the prototype is the center point of the clustering, and each class corresponds to a prototype.

[0160] According to the formula calculate the cosine similarity between the voxel features learned by the penultimate layer of the second sub-network belonging to the c-th class and the prototype;

[0161] Among them, is the prototype to which the voxel features learned by the penultimate layer of the first sub-network belong to the c-th class, is the SMPS mask corresponding to the prototype to which the voxel features learned by the penultimate layer of the first sub-network belong to the c-th class, is the number of non-zero elements in, is the voxel feature learned by the penultimate layer of the first sub-network belonging to the c-th class, e represents the element-wise product, h, w, and l are the height, width, and length of the voxel, The prototype of the voxel features learned by the penultimate layer of the second sub-network belonging to the c-th class, The SMPS mask corresponding to the prototype of the voxel features learned by the penultimate layer of the second sub-network belonging to the c-th class, is the number of non-zero elements in The cosine similarity between the voxel features belonging to the c-th class learned by the penultimate layer of the first sub-network and the prototype, where cos represents the cosine similarity, The cosine similarity between the voxel features belonging to the c-th class learned by the penultimate layer of the second sub-network and the prototype, is the voxel features belonging to the c-th class learned by the penultimate layer of the second sub-network. In some other embodiments of this specification, the cosine similarity can also be replaced by other similarity measurement algorithms.

[0162] In some embodiments of this specification, the penultimate layer of the sub-network refers to the last extracted feature, and the prediction output (i.e., the prediction result) can be obtained by inputting the last feature into the last fully connected layer.

[0163] In some embodiments of this specification, determining the weight of the pseudo-label according to the similarity may include:

[0164] Calculating the weight of the pseudo-label predicted by the first sub-network according to the formula ;

[0165] Calculating the weight of the pseudo-label predicted by the second sub-network according to the formula ;

[0166] where G A is the weight of the pseudo-label predicted by the first sub-network, C is the C-th class of the pseudo-label, and sim represents the similarity, is the cosine similarity between the voxel features belonging to the c-th class learned by the penultimate layer of the first sub-network and the prototype, is the prediction output of the voxel features of the unlabeled image by the first sub-network, G B is the weight of the pseudo-label predicted by the second sub-network, is the cosine similarity between the voxel features belonging to the c-th class learned by the penultimate layer of the second sub-network and the prototype, is the prediction output of the voxel features of the unlabeled image by the second sub-network.

[0167] For example, if the prediction probabilities and of predicting the c-th class are high, and at the same time the similarity between the voxel features and the prototype is high (i.e., and If it is larger, it is considered that the reliability of the pseudo-label is relatively high, and these pseudo-labels can be given a relatively large weight; if the predicted probability of the c-th category and is relatively high, but the similarity between the voxel feature and the prototype is relatively low (that is, and is relatively high), it is considered that the reliability of the pseudo-label is relatively low, and these pseudo-labels can be given a relatively small weight; thus, selectively using reliable pseudo-labels for cross-supervision is realized.

[0168] Due to the complexity of airway images, such as irregular segmentation shapes, adjacent boundaries, etc., the reliability of pseudo-labels may vary depending on the position of the voxels. Therefore, using the geometric structure information of airway image voxels as the basis for evaluating the reliability of pseudo-labels and assigning reliability weights to them is beneficial to further improving the prediction accuracy of the model.

[0169] In some embodiments of the present specification, calculating the unsupervised loss of the pseudo-label according to the SMPS mask and the weight may include:

[0170] Calculating the unsupervised loss of the prediction result of the first sub-network according to the formula ; and,

[0171] Calculating the unsupervised loss of the prediction result of the first sub-network according to the formula ;

[0172] wherein, is the unsupervised loss of the prediction result of the first sub-network, M B is the SMPS mask of the second sub-network, e represents element-wise multiplication, G A is the weight of the pseudo-label predicted by the first sub-network, is the prediction output of the first sub-network for the voxel feature of the unlabeled image, is the pseudo-label predicted by the second sub-network for the voxel feature, CE is the cross-entropy loss, is the unsupervised loss of the prediction result of the first sub-network, M A is the SMPS mask of the first sub-network, G B is the weight of the pseudo-label predicted by the second sub-network, is the prediction output of the second sub-network for the voxel feature of the unlabeled image, is the pseudo-label predicted by the first sub-network for the voxel feature.

[0173] In some embodiments of the present specification, determining the total loss of the prediction result according to the supervised loss and the unsupervised loss may include:

[0174] According to the formula Calculate the total loss of the prediction results of the first sub-network; and,

[0175] According to the formula Calculate the total loss of the prediction results of the first sub-network;

[0176] where, L A is the total loss of the prediction results of the first sub-network, is the supervised loss of the prediction results of the first sub-network, is the unsupervised loss of the prediction results of the first sub-network, λ is the weight, L B is the total loss of the prediction results of the second sub-network, is the supervised loss of the prediction results of the second sub-network, is the unsupervised loss of the prediction results of the second sub-network.

[0177] In some embodiments of the present specification, updating the network parameters of the first sub-network and the second sub-network according to the total loss may include updating the network parameters of the first sub-network and the second sub-network by methods such as backpropagation and gradient descent with the goal of minimizing the total loss; by iteratively executing the above steps 201-208, finally, a first sub-network and a second sub-network that meet the preset conditions can be obtained, and at least one of them is used as the airway image segmentation model. Among them, meeting the preset conditions means: meeting one or more preset model evaluation index values.

[0178] In some embodiments of the present specification, for the training principle of training the heterogeneous sub-network A and sub-network B using the airway image segmentation model training method, reference can also be made to Figure 4 and Figure 5 as shown. Among them, Figure 4 the marked image in is the labeled image; Figure 4 and Figure 5 the symbols in can refer to the explanations in the above text and will not be elaborated here.

[0179] In summary, the airway image segmentation model training method in the embodiments of the present specification has the following advantages:

[0180] (1) It can generate more diverse prediction results for the labeled and unlabeled airway images, thereby alleviating the confirmation bias problem of the semi-supervised learning method and improving the generalization ability of the model.

[0181] (2) By identifying the regions in the labeled images that are prone to segmentation errors and selectively using the reliable pseudo-labels among them for cross-supervision, the segmentation accuracy of the model can be improved.

[0182] (3) The mutual correction between sub-networks and the selective cross-supervised learning further improve the segmentation accuracy of the model.

[0183] Using the airway image segmentation model training method of the embodiments of this specification, experiments were conducted on the publicly available airway image dataset. Four evaluation metrics, DSC, TPR, BD, and TD, were used for the airway image dataset. The DSC coefficient is a statistic that measures the Dice similarity between two samples. TPR represents the ability of the model to correctly detect the true segmentation result. BD represents the proportion of the number of correctly detected airway branches to the total number of true branches. TD represents the proportion of the correctly detected tree length to the true airway tree length. The comparison of the segmentation results of the airway image segmentation model training method of the embodiments of this specification (shown in the last row of Table 1) with the segmentation results of the comparative algorithms is shown in Table 1.

[0184] Table 1

[0185]

[0186] As shown in Table 1, in the fully supervised learning setting, 3DUNet and VNet were respectively used to train the model on 100% (63) labeled images, and the obtained segmentation results were considered the upper bound performance of the model. Only 20% (12) labeled images were used to train the model, and the obtained segmentation results were used as the baseline performance. In the semi-supervised learning setting, 20% (12) labeled data and 80% (51) unlabeled data in the training set were used to train the model, and the segmentation performances of the method of the present invention and other semi-supervised learning methods (such as the MT model and the MCF model) were compared. In terms of the DSC, TPR, BD, and TD metrics, compared with the most competitive semi-supervised learning method, MCF, the segmentation results of the embodiments of this specification are 1.37%, 1.54%, 4.52%, and 2.92% higher, verifying the effectiveness of the semi-supervised learning framework proposed in the embodiments of this specification in the airway image segmentation task.

[0187] Based on obtaining the airway image segmentation model by using the above airway image segmentation model training method, the embodiments of this specification also provide an airway image segmentation method. Refer to Figure 6 As shown, in some embodiments of this specification, the airway image segmentation method may include the following steps:

[0188] Step 601, obtain the airway image to be segmented.

[0189] Step 602, input the airway image into the pre-trained airway image segmentation model to obtain the segmentation result of the airway image; wherein, the airway image segmentation model is trained by the above airway image segmentation model training method.

[0190] In some other embodiments of this specification, before inputting the airway image into the pre-trained airway image segmentation model, it may further include performing truncated normalization processing on the airway image to be segmented to further improve the segmentation accuracy of the airway image.

[0191] Although the process flows described above include multiple operations that occur in a specific order, it should be clearly understood that these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel (e.g., using a parallel processor or a multi-threaded environment).

[0192] Corresponding to the above-described method for training an airway image segmentation model, an embodiment of this specification also provides an airway image segmentation model training apparatus. Referring to Figure 7 as shown, in some embodiments of this specification, the airway image segmentation model training apparatus may include:

[0193] A data input module 71, configured to respectively input a batch of images extracted from an airway image dataset into a heterogeneous first sub-network and a second sub-network to obtain prediction results; the batch of images includes labeled images and unlabeled images, and the prediction results include hard labels corresponding to the labeled images and pseudo-labels corresponding to the unlabeled images;

[0194] An MDC regularization module 72, configured to determine a mutual difference correction MDC regularization term corresponding to the first sub-network and the second sub-network in a prediction inconsistent region according to the hard labels;

[0195] A supervised loss determination module 73, configured to calculate a supervised loss of the hard labels according to the MDC regularization term;

[0196] An SMPS regularization module 74, configured to determine a selective mutual pseudo-supervision SMPS mask of the first sub-network and the second sub-network according to the confidence of the pseudo-labels;

[0197] A weight determination module 75, configured to calculate a similarity between voxel features learned by a target layer of the first sub-network and the second sub-network and a prototype according to the SMPS mask, and determine a weight of the pseudo-labels according to the similarity;

[0198] An unsupervised loss determination module 76, configured to calculate an unsupervised loss of the pseudo-labels according to the SMPS mask and the weight;

[0199] A total loss determination module 77, configured to determine a total loss of the prediction results according to the supervised loss and the unsupervised loss;

[0200] A network parameter update module 78, configured to update network parameters of the first sub-network and the second sub-network according to the total loss;

[0201] An iteration control module 79, configured to iteratively execute the above modules to obtain a first sub-network and a second sub-network that meet preset conditions, and determine them as an airway image segmentation model.

[0202] Corresponding to the above airway image segmentation method, an embodiment of this specification also provides an airway image segmentation device. Referring to Figure 8 as shown, in some embodiments of this specification, the airway image segmentation device may include:

[0203] An image acquisition module 81, configured to acquire an airway image to be segmented;

[0204] An image segmentation module 82, configured to input the airway image into a pre-trained airway image segmentation model to obtain a segmentation result of the airway image; wherein, the airway image segmentation model is trained by the above airway image segmentation model training method.

[0205] For the convenience of description, when describing the above device, various units are described separately according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware.

[0206] An embodiment of this specification also provides a computer device. As Figure 9 shown, in some embodiments of this specification, the computer device 902 may include one or more processors 904, such as one or more central processing units (CPUs) or graphics processing units (GPUs), and each processing unit may implement one or more hardware threads. The computer device 902 may also include any memory 906, which is used to store any kind of information such as code, settings, data, etc. In a specific embodiment, a computer program stored on the memory 906 and executable on the processor 904, when the computer program is run by the processor 904, may execute the instructions of the airway image segmentation model training method or the airway image segmentation method described in any of the above embodiments. Non-limitingly, for example, the memory 906 may include any one or more combinations of the following: any type of RAM, any type of ROM, flash memory devices, hard disks, optical discs, etc. More generally, any memory may use any technology to store information. Further, any memory may provide volatile or non-volatile retention of information. Further, any memory may represent a fixed or removable component of the computer device 902. In one case, when the processor 904 executes the associated instructions stored in any memory or combination of memories, the computer device 902 may perform any operation of the associated instructions. The computer device 902 also includes one or more drive mechanisms 908 for interacting with any memory, such as a hard disk drive mechanism, an optical disc drive mechanism, etc.

[0207] The computer device 902 may also include an input / output interface 910 (I / O) for receiving various inputs (via the input device 912) and for providing various outputs (via the output device 914). A specific output mechanism may include a presentation device 916 and an associated graphical user interface 918 (GUI). In other embodiments, the input / output interface 910 (I / O), the input device 912, and the output device 914 may not be included, and it may only be a computer device in the network. The computer device 902 may also include one or more network interfaces 920 for exchanging data with other devices via one or more communication links 922. One or more communication buses 924 couple the components described above together.

[0208] The communication link 922 may be implemented in any manner, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 922 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.

[0209] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), computer-readable storage media, and computer program products according to some embodiments of the present specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processors to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processors generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0210] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processors to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0211] These computer program instructions can also be loaded onto a computer or other programmable data processors, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocksFigure 1 Steps of the functions specified in one or more boxes.

[0212] In a typical configuration, a computer device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0213] Memory may include non-permanent memory in computer-readable media, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0214] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information accessible by a computer device. As defined in this specification, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0215] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0216] The embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The embodiments of this specification can also be practiced in a distributed computing environment where tasks are performed by remote processors connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0217] It should also be understood that in the embodiments of this specification, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, both A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally indicates that the associated objects before and after are in an "or" relationship.

[0218] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content.

[0219] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0220] The above description is only for the embodiments of this application and is not used to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. A method for training an airway image segmentation model, characterized in that Including: Inputting the batch of images extracted from the airway image dataset into heterogeneous first and second sub-networks respectively to obtain prediction results; the batch of images includes labeled images and unlabeled images, and the prediction results include hard labels corresponding to the labeled images and pseudo-labels corresponding to the unlabeled images; Determining the mutual difference correction (MDC) regularization terms corresponding to the first and second sub-networks in the prediction inconsistent regions according to the hard labels; Calculating the supervised loss of the hard labels according to the MDC regularization terms; Determining the selective mutual pseudo-supervision (SMPS) masks of the first and second sub-networks according to the confidence levels of the pseudo-labels; Calculating the similarity between the voxel features learned by the target layers of the first and second sub-networks and the prototypes according to the SMPS masks, and determining the weights of the pseudo-labels according to the similarity; Calculating the unsupervised loss of the pseudo-labels according to the SMPS masks and the weights; Determining the total loss of the prediction results according to the supervised loss and the unsupervised loss; Updating the network parameters of the first and second sub-networks according to the total loss; Iteratively executing the above steps to obtain the first and second sub-networks that meet the preset conditions, and determining them as the airway image segmentation model.

2. The airway image segmentation model training method according to claim 1, wherein Before inputting the batch of images extracted from the airway image dataset into heterogeneous first and second sub-networks respectively, it further includes: Performing truncated normalization processing on the images in the airway image dataset.

3. The airway image segmentation model training method according to claim 1, wherein Determining the mutual difference correction (MDC) regularization terms corresponding to the first and second sub-networks in the prediction inconsistent regions according to the hard labels, including: Determining the masks corresponding to the first and second sub-networks in the prediction inconsistent regions according to the hard labels; Calculating the segmentation results obtained by the first and second sub-networks in the prediction inconsistent regions according to the masks, and the true labels corresponding to the prediction inconsistent regions; Respectively calculating the mean square errors between the segmentation results of the first and second sub-networks and the true labels, and correspondingly using them as the MDC regularization terms of the first and second sub-networks in the prediction inconsistent regions.

4. The airway image segmentation model training method according to claim 1, characterized in that Calculating the supervised loss of the hard labels according to the MDC regularization terms, including: According to the formula calculate the supervised loss of the hard labels predicted by the first sub-network in the prediction inconsistent region; and According to the formula calculate the supervised loss of the hard labels predicted by the second sub-network in the prediction inconsistent region; Among them, is the supervised loss of the hard labels predicted by the first sub-network in the prediction inconsistent region, is the hard label predicted by the first sub-network in the prediction inconsistent region, Y l is the ground truth label, is the Dice loss with respect to Y l and is the Focal loss with respect to Y l ; is the MDC regularization term of the first sub-network in the prediction inconsistent region, and β is the balance parameter. is the supervised loss of the hard labels predicted by the second sub-network in the prediction inconsistent region, is the hard label predicted by the second sub-network in the prediction inconsistent region, is the Dice loss with respect to Y l and is the Focal loss with respect to Y l ; is the MDC regularization term of the second sub-network in the prediction inconsistent region.

5. The airway image segmentation model training method according to claim 1, wherein Determining the selective mutual pseudo-supervision (SMPS) masks of the first and second sub-networks according to the confidence levels of the pseudo-labels, including: According to the formula Calculate the prediction entropy value of the voxel features of the unlabeled image by the first sub-network, and use it as the prediction confidence of the voxel features of the unlabeled image by the first sub-network; According to the formula Calculate the prediction entropy value of the voxel features of the unlabeled image by the second sub-network, and use it as the prediction confidence of the voxel features of the unlabeled image by the second sub-network; When C A <C B At this time, according to the formula determine the SMPS masks of the first sub-network and the second sub-network; or When C A > C B , according to the formula determine the SMPS masks of the first sub-network and the second sub-network; Among them, C A is the prediction entropy value of the voxel features of the unlabeled image by the first sub-network, is the prediction output of the voxel features of the unlabeled image by the first sub-network, T is the matrix transpose, C B is the prediction entropy value of the voxel features of the unlabeled image by the second sub-network, is the prediction output of the voxel features of the unlabeled image by the second sub-network, is the unlabeled image, M A is the SMPS mask of the first sub-network, I represents the indicator function, H, W, and L are the height, width, and length of the voxel respectively, M B is the SMPS mask of the second sub-network, and I is a matrix with all elements being 1.

6. The airway image segmentation model training method according to claim 1, wherein Calculating the similarity between the voxel features learned by the target layers of the first and second sub-networks and the prototypes according to the SMPS masks, including: According to the formula calculate the prototype of the voxel features learned by the penultimate layer of the first sub-network belonging to the c-th class; According to the formula calculate the prototype of the voxel features learned by the penultimate layer of the second sub-network belonging to the c-th class; According to the formula Calculate the cosine similarity between the voxel features learned by the penultimate layer of the first sub-network that belong to the c-th class and the prototype; According to the formula Calculate the cosine similarity between the voxel features belonging to the c-th class learned by the penultimate layer of the second sub-network and the prototype; Among them, is the prototype of the voxel features learned by the penultimate layer of the first sub-network belonging to the c-th class, is the SMPS mask corresponding to the prototype of the voxel features learned by the penultimate layer of the first sub-network belonging to the c-th class, is the number of non-zero elements in, is the voxel feature belonging to the c-th class learned by the penultimate layer of the first sub-network, e represents the element-wise product, h w, and l are the height, width, and length of the voxel, is the prototype of the voxel features learned by the penultimate layer of the second sub-network belonging to the c-th class, is the SMPS mask corresponding to the prototype of the voxel features learned by the penultimate layer of the second sub-network belonging to the c-th class, is the number of non-zero elements in, is the cosine similarity between the voxel feature belonging to the c-th class learned by the penultimate layer of the first sub-network and the prototype, cos represents the cosine similarity, is the cosine similarity between the voxel feature belonging to the c-th class learned by the penultimate layer of the second sub-network and the prototype, is the voxel feature belonging to the c-th class learned by the penultimate layer of the second sub-network.

7. The method for training an airway image segmentation model according to claim 1, wherein Determining the weights of the pseudo-labels according to the similarity, including: According to the formula Calculate the weights of the pseudo-labels predicted by the first sub-network; According to the formula Calculate the weights of the pseudo-labels predicted by the second sub-network; Among them, G A is the weight of the pseudo-label predicted by the first sub-network, C is the C-th class of the pseudo-label, and sim represents the similarity. is the cosine similarity between the voxel features belonging to the c-th class learned by the penultimate layer of the first sub-network and the prototype. is the prediction output of the voxel features of the unlabeled image by the first sub-network, and G B is the weight of the pseudo-label predicted by the second sub-network. is the cosine similarity between the voxel features belonging to the c-th class learned by the penultimate layer of the second sub-network and the prototype. is the prediction output of the voxel features of the unlabeled image by the second sub-network.

8. The airway image segmentation model training method according to claim 1, wherein Calculating the unsupervised loss of the pseudo-labels according to the SMPS masks and the weights, including: According to the formula calculate the unsupervised loss of the prediction result of the first sub-network; and, According to the formula Calculate the unsupervised loss of the prediction result of the first sub-network; Among them, is the unsupervised loss of the prediction result of the first sub-network, M B is the SMPS mask of the second sub-network, e represents the element-wise product, G A is the weight of the pseudo-label predicted by the first sub-network, is the prediction output of the voxel features of the unlabeled image by the first sub-network, is the pseudo-label predicted by the second sub-network for the voxel features, CE is the cross-entropy loss, is the unsupervised loss of the prediction result of the first sub-network, M A is the SMPS mask of the first sub-network, G B is the weight of the pseudo-label predicted by the second sub-network, is the prediction output of the voxel features of the unlabeled image by the second sub-network, is the pseudo-label predicted by the first sub-network for the voxel features.

9. The airway image segmentation model training method according to claim 1, wherein, Determining the total loss of the prediction results according to the supervised loss and the unsupervised loss, including: According to the formula calculate the total loss of the prediction result of the first sub-network; and, According to the formula calculate the total loss of the prediction result of the first sub-network; Among them, L A is the total loss of the prediction result of the first sub-network, is the supervised loss of the prediction result of the first sub-network, is the unsupervised loss of the prediction result of the first sub-network, λ is the weight, L B is the total loss of the prediction result of the second sub-network, is the supervised loss of the prediction result of the second sub-network, is the unsupervised loss of the prediction result of the second sub-network.

10. The airway image segmentation model training method according to claim 1, wherein, The first and second sub-networks include 3D Unet network and VNet network.

11. An airway image segmentation method, characterized in that, Including: Obtaining the airway image to be segmented; Input the airway image into a pre-trained airway image segmentation model to obtain the segmentation result of the airway image; wherein, the airway image segmentation model is trained according to the method described in any one of claims 1-10.

12. An airway image segmentation model training device, characterized in that, It includes: A data input module, configured to input a batch of images extracted from an airway image dataset into heterogeneous first and second sub-networks respectively to obtain prediction results; the batch of images includes labeled images and unlabeled images, and the prediction results include hard labels corresponding to the labeled images and pseudo-labels corresponding to the unlabeled images; An MDC regularization module, configured to determine the mutual difference correction MDC regularization terms corresponding to the first and second sub-networks in the prediction inconsistent regions according to the hard labels; A supervised loss determination module, configured to calculate the supervised loss of the hard labels according to the MDC regularization terms; An SMPS regularization module, configured to determine the selective mutual pseudo-supervision SMPS masks of the first and second sub-networks according to the confidence levels of the pseudo-labels; A weight determination module, configured to calculate the similarities between the voxel features learned by the target layers of the first and second sub-networks and the prototypes according to the SMPS masks, and determine the weights of the pseudo-labels according to the similarities; An unsupervised loss determination module, configured to calculate the unsupervised loss of the pseudo-labels according to the SMPS masks and the weights; A total loss determination module, configured to determine the total loss of the prediction results according to the supervised loss and the unsupervised loss; A network parameter update module, configured to update the network parameters of the first and second sub-networks according to the total loss; An iteration control module, configured to iteratively execute the above modules to obtain the first and second sub-networks that meet the preset conditions, and determine them as the airway image segmentation model.

13. An airway image segmentation device, characterized in that, It includes: An image acquisition module, configured to acquire the airway image to be segmented; An image segmentation module, configured to input the airway image into a pre-trained airway image segmentation model to obtain the segmentation result of the airway image; wherein, the airway image segmentation model is trained according to the method described in any one of claims 1-10.

14. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, When the computer program is run by the processor, it executes the instructions of the method described in any one of claims 1-11.

15. A computer storage medium, on which a computer program is stored, characterized in that, When the computer program is run by the processor of the computer device, it executes the instructions of the method described in any one of claims 1-11.

16. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is run by the processor of the computer device, it executes the instructions of the method described in any one of claims 1-11.