Three-dimensional medical image segmentation method

Through the dual-network collaborative training framework and data enhancement technology, the problem of low segmentation accuracy in three-dimensional medical image segmentation is solved, and stronger generalization ability and higher segmentation accuracy are achieved.

CN120219754AActive Publication Date: 2025-06-27SUZHOU UNIV

Patent Information

Application Number
CN202510696653.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing three-dimensional medical image segmentation method faces sample edge blur and category imbalance, and the segmentation accuracy is low, making it difficult to effectively learn under finite labeled data.

Method used

Using the dual network collaborative training framework, the supervision loss of the labeled sample and the consistency loss of the unlabeled sample is calculated through the predicted segmentation probability graph, soft pseudo-label and predicted symbol distance field output by the two subnets, and the joint optimization of the labeled and unlabeled data is achieved. At the same time, a simple attention module and an edge enhancement module are introduced to perform data enhancement to alleviate the problem of category imbalance.

Benefits of technology

The generalization ability of the model under limited annotation data is improved, the segmentation performance of three-dimensional medical images is optimized, and the segmentation accuracy and edge stability of small-category organs are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219754A_ABST
    Figure CN120219754A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a three-dimensional medical image segmentation method, which comprises the following steps: acquiring two prediction segmentation probability graphs of each sample based on two parallel sub-networks of a dual-network segmentation model, and calculating a soft pseudo label of an unlabeled sample; utilizing a distance regression head and a hyperbolic tangent function to obtain two prediction symbol distance fields of each sample based on the decoding feature map; calculating segmentation loss and regression loss of each labeled sample, and adding the segmentation loss and the regression loss to obtain supervision loss; calculating the pseudo label consistency loss and the symbol distance consistency loss of each unlabeled sample, and adding the pseudo label consistency loss and the symbol distance consistency loss to obtain consistency loss; optimizing the dual-network segmentation model under the joint constraint of supervision loss and consistency loss; and obtaining an input segmentation result of the to-be-segmented three-dimensional medical image by using the trained dual-network segmentation model. According to the method, the problem of class imbalance in the three-dimensional medical image is effectively relieved, and the segmentation precision of the weak edge region is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a three-dimensional medical image segmentation method. Background Art

[0002] Image segmentation is an important research direction in the fields of computer vision and image processing. Its goal is to divide an image into several non-overlapping connected regions and extract the region of interest (ROI). In the field of medical image segmentation, the segmentation of three-dimensional medical images plays a crucial role in clinical diagnosis and treatment planning. However, deep learning models usually rely on a large amount of high-quality labeled data, and the labeling process of medical images is costly and time-consuming, which severely restricts its popularization in practical applications.

[0003] In recent years, semi-supervised learning methods have become a research hotspot. The core idea is to improve the generalization ability of the model by effectively using unlabeled data and alleviate the problem of scarce labeled data. However, in the case of blurred edges and complex structures, pseudo-label noise is prone to concentrate in the edge regions in existing semi-supervised segmentation methods, resulting in local missegmentation or even overall morphological distortion, affecting the reliability of the model in clinical applications. In addition, another important challenge faced by medical image segmentation is the class imbalance problem. Since the number of samples in small classes in medical images is usually small, the model tends to predict large classes, leading to a decline in performance. These small-class organs are often small in volume and scarce in samples, further increasing the difficulty of model training.

[0004] In summary, since the model depends on sample data for training, and in the case of blurred sample edges, complex results, and a small number of small-class samples, there are problems of blurred segmentation edges and low segmentation accuracy in existing three-dimensional medical image segmentation. Summary of the Invention

[0005] Therefore, the technical problem to be solved by the present invention is to overcome the problem of low segmentation accuracy in the prior art when facing blurred edges and class imbalance of three-dimensional medical image samples.

[0006] To solve the above technical problem, the present invention provides a three-dimensional medical image segmentation method, including: Obtain a three-dimensional medical image sample set including labeled samples and unlabeled samples; Input all samples into two sub-networks in parallel of a dual-network segmentation model respectively, obtain two predicted segmentation probability maps of each sample, and calculate the soft pseudo-labels of the unlabeled samples based on the unnormalized classification scores of the unlabeled samples; the sub-networks are constructed based on the V-Net network; Using a distance regression head and the hyperbolic tangent function, two predicted signed distance fields for each sample are obtained based on the decoded feature maps output by the decoder of each sub-network for all samples. For each labeled sample: Calculate the cross-entropy loss function between each predicted segmentation probability map and the ground truth label to obtain the segmentation loss; Calculate the signed distance function loss between each predicted signed distance field and the ground truth signed distance function to obtain the regression loss; Add the segmentation loss and the regression loss to obtain the supervised loss of the labeled sample. For each unlabeled sample: Calculate the average of the cross-entropy loss and the Dice loss between each predicted segmentation probability map and the soft pseudo-label to obtain the pseudo-label consistency loss; Calculate the signed distance function loss between the two predicted signed distance fields to obtain the signed distance consistency loss; Add the pseudo-label consistency loss and the signed distance consistency loss to obtain the consistency loss of the unlabeled sample. Based on the supervised loss of the labeled samples and the consistency loss of the unlabeled samples, a total loss function is obtained, and each sub-network in the dual-network segmentation model is trained. Use the trained dual-network segmentation model to obtain the segmentation result of the input 3D medical image to be segmented.

[0007] Preferably, the predicted signed distance field of a sample is expressed as: ; where represents the predicted signed distance field of class samples after passing through sub-network ; The sub-network identifier , where A and B respectively represent two parallel sub-networks of the dual-network segmentation model; The sample annotation status identifier , when it is a labeled sample, when it is an unlabeled sample; is the hyperbolic tangent function, represents the distance regression head with an output channel of 1; represents

[0008] the decoded feature map output by the decoder of class samples after passing through sub-network Preferably, calculating the supervised loss of the labeled samples includes: ; Calculate the cross-entropy loss function between the predicted segmentation probability map of the labeled sample in each sub-network and its corresponding ground truth label to obtain the segmentation loss which is expressed as: ; Add the segmentation loss and the regression loss to obtain the supervised loss of the labeled sample , expressed as: ; wherein, and respectively represent the learnable weighting factors of the first sub-network A and the second sub-network B; represents the cross-entropy loss, and Y represents the true label of the labeled sample; and respectively represent the predicted segmentation probability maps output by the labeled sample passing through the first sub-network A and the second sub-network B, and the calculation formula is ; is the decoded feature map obtained by the labeled sample passing through the sub-network All unnormalized classification scores of the voxels in are expressed as: , represents a 1×1×1 three-dimensional convolutional layer, represents the classification score corresponding to the th category in the labeled sample, , represents the total number of organ categories in the sample; represents the Softmax function; represents the exponential function; and respectively represent the predicted signed distance fields output by the labeled sample passing through the first sub-network A and the second sub-network B; represents the signed distance function loss, which is the average of the L1 loss and the mean square error loss; represents the true signed distance function of the labeled sample, and the expression is: , represents the infimum, represents the voxel point in the true label, represents the voxel point on the surface of the true label, represents the voxel point and the squared Euclidean distance between, , and respectively represent the external region, surface and internal region of the true label.

[0009] Preferably, calculate the consistency loss of the unlabeled sample, including: Calculate the average of the cross-entropy loss and the Dice loss between the predicted segmentation probability map of the unlabeled sample in each sub-network and the soft pseudo-label to obtain the pseudo-label consistency loss , expressed as: ; Calculate the signed distance function loss between the predicted signed distance fields of unlabeled samples in each sub-network to obtain the signed distance consistency loss , expressed as: ; Add the pseudo-label consistency loss and the signed distance consistency loss to obtain the consistency loss of unlabeled samples , expressed as: ; Among them, and respectively represent the predicted segmentation probability maps output by the unlabeled samples passing through the first sub-network A and the second sub-network B; represents the segmentation loss, which is the average of the cross-entropy loss and the Dice loss; and respectively represent the soft pseudo-labels corresponding to the unnormalized classification scores output by the unlabeled samples passing through the first sub-network A and the second sub-network B; in the sub-network the soft pseudo-label ; represents the preset soft pseudo-label smoothing degree parameter, represents the decoded feature map obtained by the unlabeled samples passing through the sub-network ; represents the unnormalized classification scores of all voxels in ; represents the -th category corresponding unnormalized classification score in the unlabeled samples, , represents the total number of organ categories in the samples; and respectively represent the predicted signed distance fields output by the unlabeled samples passing through the first sub-network A and the second sub-network B.

[0010] Preferably, based on the supervised loss of the labeled samples and the consistency loss of the unlabeled samples, the total loss function is obtained, expressed as: ; Among them, represents the consistency loss weight, and the expression is , represents the current training epoch, represents the preset maximum training epoch.

[0011] Preferably, the two predicted segmentation probability maps output by the two sub-networks of the trained dual-network segmentation model are averaged and fused to obtain a fused segmentation probability map; an argmax operation is performed on the fused segmentation probability map to obtain a predicted segmentation label; based on the predicted segmentation label, the segmentation result of the three-dimensional medical image to be segmented is obtained.

[0012] Preferably, the three-dimensional medical image is input into the sub-network based on the V-Net network to obtain the predicted segmentation probability map of the sample, including: Taking the three-dimensional medical image sample as the input of the encoder of the sub-network ; where the sub-network identifier , A and B respectively represent the two parallel sub-networks of the dual-network segmentation model; the sample annotation status identifier , when , it is an annotated sample, , when ; after passing through four cascaded encoding layers, an encoded feature map is output; in each encoding layer, the input feature map passes through a cascaded three-dimensional convolutional block and a simple attention module in sequence and then outputs an attention convolutional feature map, and after passing through a downsampling unit, the output feature map of each encoding layer is output, expressed as: ; represents the output feature map of the th encoding layer of the sub-network ; represents the downsampling block, including a three-dimensional convolutional block with a stride of 2; represents the sub-network th attention convolutional feature map output after passing through the cascaded three-dimensional convolutional block and the simple attention module in the encoding layer; represents the simple attention module; represents the three-dimensional convolutional block in the encoder, including a plurality of convolutional layers, a batch normalization layer, and a ReLU activation layer connected in sequence; The encoded features are input into the bottleneck module, and after passing through the cascaded three-dimensional convolutional block and the simple attention module, a global feature map is output; The global feature map is input into the decoder, and after passing through four cascaded decoding layers, a decoded feature map is output; The decoded feature map is input into the segmentation prediction unit of the output layer, and the corresponding predicted segmentation probability map is output.

[0013] Preferably, the global feature map is input into the decoder, and after passing through four cascaded decoding layers, a decoded feature map is output, including: In each decoding layer, the input feature map is upsampled and then skip-connected with the attention convolution feature map of the corresponding encoding layer. After passing through a 3D convolution block, the output feature map of each decoding layer is obtained, which is expressed as: ; represents the sub-network the output feature map of the -th decoding layer of the sub-network ; represents the 3D convolution block in the decoder, including a plurality of convolution layers, batch normalization layers, and ReLU activation layers connected in series in sequence; represents pixel-wise addition, represents the sub-network the attention convolution feature map of the -th encoding layer of the sub-network; After passing the output feature map of the second decoding layer through a dimension adjustment block and a distance regression head, a predicted symbol distance map is obtained, and then it is adjusted using a preset edge sensitivity parameter to obtain an edge-enhanced feature map;

[0014] Preferably, the obtaining of the decoded feature map includes: Calculating the predicted symbol distance map of the output feature map of the second decoding layer , which is expressed as: ; Based on the preset edge sensitivity parameter and the predicted symbol distance map, an edge-enhanced feature map is obtained, which is expressed as: ; Based on the highest-layer output feature map of the decoder and the edge-enhanced feature map , the decoded feature map is obtained, which is expressed as: ; wherein, represents a distance regression head with an output channel of 1, represents a dimension adjustment block.

[0015] Preferably, before inputting a 3D medical image sample into the dual-network segmentation model, data augmentation of the 3D medical image sample is also included, including: Performing random cropping and random flipping on both the labeled samples and the unlabeled samples to obtain corresponding labeled 3D input sample data and unlabeled 3D input sample data ; Based on the regional volume of the small-category organ regions contained in each labeled three-dimensional input sample data and the volume of the category with the smallest volume in the small-category organ set , calculate the volume weight of the small-category organ region ; Based on the positions of the voxel points in the small-category organ region, calculate the signed distance function of the small-category organ region ; Let \(\inf\) denote the infimum, let \(x_i\) denote the voxel point at the \(i\)-th layer slice of the three-dimensional input sample data at the position \((x,y,z)\), let \(x_s\) denote the voxel point on the surface of the small-category organ region, let \(d^2(x_i,x_s)\) denote the squared Euclidean distance between the voxel point \(x_i\) and \(x_s\); let \(x_i\in\Omega_{out}\), \(x_i\in\Omega_{s}\), and \(x_i\in\Omega_{in}\) respectively denote the external region, surface, and internal region of the small-category organ region \(\Omega\); Based on the volume weight and signed distance function of the small-category organ region, and the fitted Dirac function, calculate the active contour deformation field of the target organ region , expressed as: ; Let \(\alpha\) denote the deformation amplitude control parameter, let \(\delta(x)\) denote the fitted Dirac function, with the expression , where \(c\) is a positive constant, and \(x\) is the calculation variable of the fitted Dirac function; Let \(\nabla\phi\) denote the gradient of the level set, expressed as ; Based on the three-dimensional Gaussian kernel function and the active contour deformation field, calculate the smoothed deformation field ; where \(G_{\sigma}\) is the three-dimensional Gaussian kernel function, \(\sigma\) is the standard deviation, and \(\ast\) is the convolution operation symbol; Based on the smoothed deformation field, perform smoothed deformation on the labeled three-dimensional input sample data and its corresponding ground truth label to obtain the input sample corresponding to the labeled sample and the enhanced ground truth label , where \(\odot\) is the element-wise multiplication operation symbol.

[0016] ​​​​​​​​​The above technical solution of the present invention has the following beneficial effects compared with the prior art: In the three-dimensional medical image segmentation method of the present invention, when training the dual-network segmentation model, a dual-network collaborative training framework is adopted. Based on the predicted segmentation probability maps, soft pseudo-labels, and predicted signed distance fields output by the two sub-networks, the supervised loss of the labeled samples and the consistency loss of the unlabeled samples are calculated for joint optimization, realizing effective learning of the labeled samples and the unlabeled samples, thereby obtaining stronger generalization ability in the limited-label scenario and further optimizing the segmentation performance of the model for three-dimensional medical images. For the labeled samples, cross-entropy loss is used to learn the classification of each voxel category of the organ, and the signed distance function loss is used to accurately learn the geometric shape of the target organ to ensure the basic segmentation performance. For the unlabeled data, through cross-collaborative training of the soft pseudo-labels and predicted segmentation probability maps output by the two sub-networks, mutual supervision is carried out to avoid the prediction bias of a single model; and through the signed distance function consistency, the collaboration of geometric prediction is forced to improve the stability of the segmentation edge.

[0017] At the same time, the present invention introduces a simple attention module in the encoder part and integrates edge enhancement module operations in the decoder part. The present invention performs attention weighting on the voxel-level features in the three-dimensional medical image through the simple attention module, enhancing the perception of key regions. And the simple attention module calculates the attention weights through an energy function and does not rely on additional parameters, and can improve the network's perception ability of key regions in three-dimensional medical images at a low computational cost. The present invention enables the network to pay more attention to the edge regions through edge enhancement operations, enhancing the perception ability of weak edge regions, thereby improving the segmentation performance of weak edge medical images and further improving the segmentation accuracy.

[0018] The present invention performs data augmentation on three-dimensional medical images, proposes an active contour deformation data augmentation strategy, and performs shape transformation on the small organ regions in the labeled samples and unlabeled samples through the level set method to enhance the diversity of training samples, thereby alleviating the class imbalance problem, enhancing the feature representation ability of small category organs, enhancing the generalization ability and robustness of the model, and enabling the trained dual-network segmentation model to have higher segmentation accuracy. Description of the Drawings

[0019] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to the specific embodiments of the present invention in combination with the drawings, where: Figure 1 is the step flowchart of the three-dimensional medical image segmentation method provided by the present invention; Figure 2 is the data augmentation step flowchart; Figure 3It is a flowchart of the steps of a 3D medical image segmentation method based on active contour deformation and edge enhancement; Figure 4 It is a comparison schematic diagram of the 2D cross-section and 3D segmentation results of the present invention and different semi-supervised methods at a 5% annotation ratio. Specific embodiments

[0020] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited are not intended to limit the present invention.

[0021] Referring to Figure 1 as shown, the flowchart of the steps of the 3D medical image segmentation method provided by the present invention specifically includes: S101: Obtain a 3D medical image sample set including labeled samples and unlabeled samples; S102: Input all samples into two sub-networks in parallel of the dual-network segmentation model respectively, obtain two predicted segmentation probability maps for each sample, and calculate the soft pseudo-labels of the unlabeled samples based on the unnormalized classification scores of the unlabeled samples; the sub-network is constructed based on the V-Net network; S103: Use the distance regression head and the hyperbolic tangent function to obtain two predicted signed distance fields for each sample based on the decoded feature maps output by the decoder of each sub-network for all samples; The predicted signed distance field of the class of samples after passing through the sub-network is expressed as: ; Among them, the sub-network identifier , A and B respectively represent two sub-networks in parallel of the dual-network segmentation model; the sample annotation status identifier , when it is a labeled sample, when it is an unlabeled sample; is the hyperbolic tangent function, represents the distance regression head with an output channel of 1; represents the decoded feature map output by the decoder of the sub-network for the S104: For each labeled sample: Calculate the cross-entropy loss function of each predicted segmentation probability map and the true label to obtain the segmentation loss; calculate the signed distance function loss of each predicted signed distance field and the true signed distance function to obtain the regression loss; add the segmentation loss and the regression loss to obtain the supervised loss of the labeled sample; S105: For each unlabeled sample: Calculate the average of the cross-entropy loss and the Dice loss between each predicted segmentation probability map and the soft pseudo-label to obtain the pseudo-label consistency loss; Calculate the signed distance function loss between two predicted signed distance fields to obtain the signed distance consistency loss; Add the pseudo-label consistency loss and the signed distance consistency loss to obtain the consistency loss of the unlabeled sample; S106: Based on the supervision loss of the labeled samples and the consistency loss of the unlabeled samples, obtain the total loss function, and train each sub-network in the dual-network segmentation model; S107: Use the trained dual-network segmentation model to obtain the segmentation result of the input three-dimensional medical image to be segmented.

[0022] Specifically, in step S104, calculating the supervision loss of the labeled samples includes: S104-1: Calculate the cross-entropy loss function between the predicted segmentation probability map of the labeled sample in each sub-network and its corresponding ground truth label to obtain the segmentation loss , expressed as: ; S104-2: Calculate the signed distance function loss between the predicted signed distance field of the labeled sample in each sub-network and the true signed distance function to obtain the regression loss , expressed as: ; S104-3: Add the segmentation loss and the regression loss to obtain the supervision loss of the labeled sample , expressed as: ; Among them, and respectively represent the learnable weighting factors of the first sub-network A and the second sub-network B; represents the cross-entropy loss, and Y represents the ground truth label of the labeled sample; and respectively represent the predicted segmentation probability maps output by the labeled sample passing through the first sub-network A and the second sub-network B, and the calculation formula is ; is the decoded feature map obtained by the labeled sample passing through the sub-network The unnormalized classification scores of all voxels in are expressed as: , represents a 1×1×1 three-dimensional convolutional layer, represents the classification score corresponding to the th category in the labeled sample, , represents the total number of organ categories in the sample; denotes the Softmax function; denotes the exponential function; and respectively denote the predicted signed distance fields output by the first sub-network A and the second sub-network B for the labeled samples; denotes the signed distance function loss, which is the average of the L1 loss and the mean squared error loss; denotes the true signed distance function of the labeled sample, and the expression is: , denotes the infimum, denotes the voxel points in the true label, denotes the voxel points on the surface of the true label, denotes the voxel point and denotes the squared Euclidean distance between , and respectively denote the external region, surface and internal region of the true label.

[0023] Specifically, in step S105, the consistency loss of the unlabeled samples is calculated, including: S105-1: Calculate the average of the cross-entropy loss and the Dice loss between the predicted segmentation probability map of the unlabeled sample in each sub-network and the soft pseudo-label to obtain the pseudo-label consistency loss , which is expressed as: ; S105-2: Calculate the signed distance function loss between the predicted signed distance fields of the unlabeled sample in each sub-network to obtain the signed distance consistency loss , which is expressed as: ; S105-3: Add the pseudo-label consistency loss and the signed distance consistency loss to obtain the consistency loss of the unlabeled sample , which is expressed as: ; where and respectively denote the predicted segmentation probability maps output by the first sub-network A and the second sub-network B for the unlabeled sample; denotes the segmentation loss, which is the average of the cross-entropy loss and the Dice loss; and respectively denote the soft pseudo-labels corresponding to the unnormalized classification scores output by the first sub-network A and the second sub-network B for the unlabeled sample; in the sub-network the soft pseudo-label ; denotes the preset soft pseudo-label smoothing degree parameter, Indicates the decoded feature map obtained by the unlabeled samples passing through the sub-network All the unnormalized classification scores of the voxels in are expressed as: ; Indicates the unnormalized classification score corresponding to the th category in the unlabeled samples, , represents the total number of organ categories in the samples; and respectively represent the predicted signed distance fields output by the unlabeled samples passing through the first sub-network A and the second sub-network B.

[0024] Therefore, based on the supervised loss of the labeled samples and the consistency loss of the unlabeled samples, the total loss function is obtained and expressed as: ; Among them, represents the consistency loss weight, and the expression is , represents the current training epoch, represents the preset maximum number of training epochs.

[0025] In this embodiment, after training for the preset maximum number of training epochs and obtaining the trained dual-network segmentation model, the three-dimensional medical image to be segmented is input for image segmentation. Specifically, the two predicted segmentation probability maps output by the two sub-networks of the trained dual-network segmentation model are averaged and fused to obtain the fused segmentation probability map; an argmax operation is performed on the fused segmentation probability map to obtain the predicted segmentation label; based on the predicted segmentation label, the segmentation result of the three-dimensional medical image to be segmented is obtained.

[0026] In the whole training process of the embodiment of the present invention, the two sub-networks are jointly optimized through the supervised loss and the consistency loss, and effective learning of the labeled and unlabeled data is realized based on the self-supervised and mutual-supervised mechanisms. Therefore, stronger generalization ability can be obtained in the limited-label scenario, and the segmentation accuracy of small-category targets can be improved.

[0027] Based on the above embodiment, in this embodiment, an improvement and extension are made based on the classic V-Net. A simple attention module is introduced in the encoder part, and an edge enhancement module is fused in the decoder part to obtain EEVNet. Based on EEVNet, the three-dimensional medical image is input into the sub-network (EEVNet) based on the V-Net network to obtain the predicted segmentation probability map of the sample, including: S201: Use the three-dimensional medical image sample as the input of the encoder of the sub-network ; among them, the sub-network identifier , A and B respectively represent two sub-networks in parallel of the dual-network segmentation model; the sample annotation situation identifier , when , it is a labeled sample, , when it is an unlabeled sample; after passing through four cascaded encoding layers, an encoded feature map is output; In each encoding layer, the input feature map passes through a cascaded three-dimensional convolutional block and a simple attention module in sequence and then outputs an attention convolutional feature map, and after passing through a downsampling unit, the output feature map of each encoding layer is output, which is expressed as: ; ; Among them, represents the output feature map of the th encoding layer of the sub-network ; represents a downsampling block, including a three-dimensional convolutional block with a stride of 2; represents the th attention convolutional feature map output after passing through the cascaded three-dimensional convolutional block and the simple attention module in the encoding layer of the sub-network represents a simple attention module; represents a three-dimensional convolutional block in the encoder, including a plurality of convolutional layers, a batch normalization layer, and a ReLU activation layer connected in sequence; S202: Input the encoded features into the bottleneck module, pass through the cascaded three-dimensional convolutional block and the simple attention module, and output a global feature map; S203: Input the global feature map into the decoder, and after passing through four cascaded decoding layers, output a decoded feature map; S204: Input the decoded feature map into the segmentation prediction unit of the output layer, and output the corresponding predicted segmentation probability map.

[0028] Specifically, in step S203, input the global feature map into the decoder, and after passing through four cascaded decoding layers, output a decoded feature map, including: S203-1: In each decoding layer, the input feature map is upsampled and then undergoes a skip connection with the attention convolutional feature map of the corresponding encoding layer, and then passes through a three-dimensional convolutional block, and the output feature map of each decoding layer is output, which is expressed as: ; Among them, represents the output feature map of the th decoding layer of the sub-network ; represents a three-dimensional convolutional block in the decoder, including a plurality of convolutional layers, a batch normalization layer, and a ReLU activation layer connected in sequence; Denotes an upsampling block, including a 3D transposed convolution with a stride of 2, a batch normalization layer, and a ReLU activation layer connected in series in sequence; Denotes pixel-by-pixel addition, Denotes a subnetwork The attention convolution feature map of the S203-2: After passing the output feature map of the second decoding layer through a dimension adjustment block and a distance regression head to obtain a predicted symbol distance map, it is adjusted using a preset edge sensitivity parameter to obtain an edge-enhanced feature map, including: Calculating the predicted symbol distance map of the output feature map of the second decoding layer which is expressed as: ; ; Based on the preset edge sensitivity parameter and the predicted symbol distance map, obtaining the edge-enhanced feature map which is expressed as: ; Among them, denotes a distance regression head with an output channel of 1, denotes a dimension adjustment block; S203-3: Fusing the edge-enhanced feature map and the output feature map of the highest layer, and outputting a decoded feature map, which is expressed as: ; where is the decoded feature map, is the output feature map of the highest layer of the decoder, is the edge-enhanced feature map.

[0029] The role of adding a simple attention module to each layer of the decoder in this embodiment is to perform attention weighting on voxel-level features and enhance the perception of key regions. The simple attention module calculates attention weights through an energy function and does not rely on additional parameters. Compared with traditional channel attention, it can improve the network's perception ability of key regions at a lower computational cost. The targets in medical images usually have the characteristics of unclear edges and few pixels in the edge regions, resulting in difficult model training. Through the edge enhancement module, the network can pay more attention to the edge regions during the training process, enhance the perception ability of weak edge regions, and thus improve the segmentation performance of weak edge medical images.

[0030] Based on the above embodiments, in the embodiments of the present invention, before inputting a 3D medical image sample into the dual-network segmentation model, it further includes performing data augmentation on the 3D medical image sample. Referring to Figure 2 as shown, it is a flowchart of the data augmentation steps, and the specific steps include: S301: Randomly crop and randomly flip both the labeled samples and the unlabeled samples to obtain the corresponding labeled three-dimensional input sample data and the unlabeled three-dimensional input sample data ; S302: Calculate the volume weight of the small-category organ region based on the regional volume of each small-category organ region included in the labeled three-dimensional input sample data , and the volume of the category with the smallest volume in the small-category organ set ; ; S303: Calculate the signed distance function of the small-category organ region based on the positions of the voxel points in the small-category organ region ; denotes the infimum, denotes the voxel point at the -th slice of the three-dimensional input sample data at the position, denotes the voxel point on the surface of the small-category organ region, denotes the voxel point and the squared Euclidean distance between them; , and respectively denote the external region, the surface, and the internal region of the small-category organ region ; S304: Calculate the active contour deformation field of the target organ region based on the volume weight and the signed distance function of the small-category organ region, and the fitted Dirac function, expressed as: ; ; denotes the deformation amplitude control parameter, denotes the fitted Dirac function, with the expression , is a positive constant, is the calculation variable of the fitted Dirac function; denotes the gradient of the level set, expressed as ; S305: Calculate the smoothed deformation field based on the three-dimensional Gaussian kernel function and the active contour deformation field; is the three-dimensional Gaussian kernel function, is the standard deviation, is the convolution operation symbol; S306: Based on the smoothed deformation field, the labeled three-dimensional input sample data, and its corresponding true label Perform smooth deformation to obtain the input sample corresponding to the labeled sample and the enhanced true label , is the symbol for element-wise multiplication operation.

[0031] Based on the data augmentation strategy of active contour deformation and the curve evolution theory of the level set method, the embodiments of the present invention perform shape transformation on the small organ region to enhance the diversity of training samples. This transformation can effectively adjust the shape and volume of small-category organs, thereby enhancing the generalization ability and robustness of the model.

[0032] Based on the above embodiments, in the embodiments of the present invention, the image enhancement method, the improved sub-network based on the V-Net network, and the model training method provided by the present invention are all applied to this embodiment to segment the three-dimensional medical image to be segmented. Referring to Figure 3 shown in the figure, it is the step flow chart of the three-dimensional medical image segmentation method based on active contour deformation and edge enhancement, which specifically includes: S401: Divide and preprocess the medical image dataset, and apply data augmentation operations to the training set composed of labeled data and unlabeled data to generate input samples; S401-1: Divide and preprocess the medical image dataset, and construct a training set containing labeled data and unlabeled data; S401-2: Randomly crop and flip the training set to obtain the corresponding labeled three-dimensional sample data and unlabeled three-dimensional input sample data ; S401-3: Calculate the volume weight of the small-category organs contained in each labeled three-dimensional input sample data , where represents the small-category organ region, represents the volume of the small-category organ, is the minimum symbol, represents the label set of the small category, is the volume size of the category with the smallest volume in the label set of the small category; S401-4: Calculate the signed distance function of the small-category organ region , where , and respectively represent the surface, external region, and internal region of the organ , is the th layer slice in the three-dimensional sample data The voxel points at the position, the voxel points representing the surface of the small-category organ region, and the squared Euclidean distance between the voxel points; is the voxel point on the organ surface, is the voxel and the squared Euclidean distance between, represents the infimum; S401-5: Calculate the active contour deformation field , where is the parameter controlling the deformation amplitude, is the fitted Dirac function, whose role is to limit the deformation to occur only near the zero level set; is a positive constant, is the calculation variable for the fitted Dirac function, is the gradient of the level set. In this embodiment, , ; S401-6: Calculate the smoothed deformation field , where is the three-dimensional Gaussian kernel function, is the standard deviation, is the convolution operation symbol. In this embodiment, ; S401-7: Calculate and obtain the input sample after the active contour deformation of the labeled three-dimensional medical image and the enhanced true label , where is the true label of, represents the element-wise multiplication operation symbol.

[0033] S402: Construct a dual-network collaborative training framework for the edge-enhanced EEVNet, initialize two sub-networks EEVNetA and EEVNetB, and input the input samples into the two sub-networks for training respectively; the EEVNet includes an encoder, a bottleneck layer, a decoder, and a dual-output layer; S402-1: Configure the training parameters, including: In this embodiment, the batch size is set to 2; the maximum number of training epochs is set to ; the optimizer uses the stochastic gradient descent method with momentum, the momentum coefficient is 0.95, and the weight decay coefficient is ; the learning rate is dynamically adjusted using the polynomial decay strategy, and is updated to at the th epoch, where is the initial learning rate, is set to 0.001; S402-2: Input the input samples into the sub-networks EEVNetA and EEVNetB with the same network structure respectively; S402-3: The input samples pass through the encoders in the sub-networks to extract multi-level features. The encoder has a total of four encoding layers, and each encoding layer is stacked by a three-dimensional convolutional block, a simple attention module, and a downsampling block. The layer encoding process can be expressed as: ; ; Among them, represents the sub-network identifier, where sub-network A and sub-network B are two parallel sub-networks in the method; represents the sample annotation situation. When , it is a labeled sample, and when , it is an unlabeled sample; , represents the feature map input at the layer. When , represents the three-dimensional medical image sample input into the encoder of sub-network represents the feature map after being processed by the three-dimensional convolutional block and the simple attention module, and represents the feature map after one layer of encoding; is the operation of the simple attention module; represents the three-dimensional convolutional block in the encoder, which is composed of several convolutional layers, batch normalization layers, and ReLU activation layers; is the downsampling operation, which downsamples the feature map through a three-dimensional convolution with a stride of 2; S402-4: After four layers of encoding, the deepest feature map enters the bottleneck layer for further processing, which is expressed as ; S402-5: The low-level features in the encoder are passed into the decoder through skip connections. The decoder has a total of four decoding layers, and each decoding layer is stacked by a three-dimensional convolutional block and an upsampling block. The layer decoding process can be expressed as , where , represents the output feature of the layer of the decoder, and is the decoding feature of the layer. When , represents the feature map output by the bottleneck layer; is the attention convolutional feature map of the layer of the encoding layer, Denotes the element-wise addition symbol, is an upsampling block, consisting of a 3D transposed convolution with a stride of 2, a batch normalization layer, and a ReLU activation layer; Denotes the 3D convolutional block in the decoder, which is composed of multiple convolutional layers, batch normalization layers, and ReLU activation layers connected in series in sequence; S402-6: Calculate the predicted signed distance function of the output features of the second layer , where, is a dimension adjustment block that adjusts the feature dimension through convolution and transposed convolution to align it with the highest-level features align; is a distance regression head with an output channel of 1; Calculate the features after edge enhancement , where, is a hyperparameter that controls the edge sensitivity. In this embodiment, ; The edge enhancement operation is performed on the output features of the second layer and fused with the highest-level features, which can be expressed as , where, is the highest-level fused feature, is the highest-level decoded feature; S402-7: Each sub-network outputs predicted segmentation predictions and predicted signed distance functions, including: For the segmentation prediction, first generate unnormalized classification scores for each voxel, then perform a Softmax operation on the classification scores to obtain segmentation probabilities, and perform a temperature-scaled Softmax operation on the classification scores of unlabeled data to obtain soft pseudo-labels. Calculate the unnormalized classification scores , where, is the unnormalized classification scores of all voxels in the decoded feature map obtained by the class samples through the sub-network is a 1×1×1 3D convolutional layer; Calculate the segmentation probability , where, is the predicted segmentation probability of the class samples through the sub-network denotes the Softmax function, is the exponential function. Denotes the classification score corresponding to the th category in the class samples, denotes the total number of organ categories in the sample; The temperature-scaled Softmax operation on the classification scores of unlabeled data to obtain soft pseudo-labels , where, is the temperature parameter for controlling the smoothness of the soft pseudo-label. In this embodiment, .

[0034] The symbol distance function prediction is implemented through a distance regression head and a hyperbolic tangent function, and can be expressed as , where, is the hyperbolic tangent function, represents the decoded feature map output by the decoder of the class sample through the sub-network , represents the predicted symbol distance field of the class sample after passing through the sub-network .

[0035] S403: The dual networks are optimized under the joint constraints of the supervised loss and the consistency loss until the total loss function converges, and the trained EEVNetA and EEVNetB networks are obtained; S403-1: Calculate the segmentation loss , where, is the cross-entropy loss, and are the learnable weighting factors of the two networks respectively; Calculate the symbol distance function regression loss , where, is the true symbol distance function, is the symbol distance function loss, and this loss is calculated as the average of the L1 loss and the mean squared error loss; The supervised loss of the labeled data; S403-2: Calculate the pseudo-label consistency loss , where, represents the segmentation loss of the unlabeled data, defined as the average of the cross-entropy loss and the Dice loss; Calculate the symbol distance function consistency loss ; Calculate the consistency loss of the unlabeled data; S403-3: Calculate the consistency loss weight , where, is the current training epoch, represents the preset maximum training epoch; Calculate the total loss function, expressed as: ; S403-4: In each training epoch, the two sub-networks respectively perform forward propagation to obtain outputs, calculate their respective supervised losses and the consistency losses with each other. Subsequently, the total loss is backpropagated, and their respective parameters are updated independently; S403-5: Until the total loss function converges, obtain the trained EEVNetA and EEVNetB networks.

[0036] S404: Input the medical image to be segmented into the trained dual network, and output the final segmentation result; S404-1: Input the 3D medical image to be segmented into the trained EEVNetA and EEVNetB sub-networks, and respectively obtain the predicted segmentation probability maps and ; S404-2: Integrate the prediction results of the two sub-networks, and adopt the average fusion strategy to calculate the final segmentation probability map ; S404-3: Perform the argmax operation on to output the predicted segmentation label , and obtain the final segmentation result.

[0037] The three-dimensional medical image segmentation method based on active contour deformation and edge enhancement described in the present invention is based on the curve evolution theory of the level set method. Through the data augmentation strategy of active contour deformation, the shape of the small organ region is transformed to enhance the diversity of training samples. This transformation can effectively adjust the shape of small-category organs and expand their volume, thereby enhancing the generalization ability and robustness of the model. The present invention is improved and extended on the basis of the classical V-Net. A simple attention module is introduced in the encoder part, and an edge enhancement module is fused in the decoder part to construct the edge-enhanced EEVNet. Specifically, adding a simple attention module to each layer of the decoder is to perform attention weighting on voxel-level features and enhance the perception of key regions. The simple attention module calculates the attention weight through an energy function and does not rely on additional parameters. Compared with traditional channel attention, it can improve the network's perception ability of key regions at a lower computational cost. The targets in medical images usually have the characteristics of unclear edges and few pixels in the edge region, resulting in difficult model training. The edge enhancement module can make the network pay more attention to the edge region during the training process, enhance the perception ability of weak edge regions, and thus improve the segmentation performance of weak edge medical images. In addition, during the entire training process of the present invention, the two sub-networks are jointly optimized through the supervised loss and the consistency loss, and effective learning of labeled and unlabeled data is achieved based on the self-supervised and mutual-supervised mechanisms. Thus, stronger generalization ability can be obtained in the limited label scenario.

[0038] Based on the above embodiments, in order to further illustrate the effect of the 3D medical image segmentation method based on active contour deformation and edge enhancement provided by the embodiments of the present invention, the 3D medical image segmentation method based on active contour deformation and edge enhancement provided by this embodiment and the existing MT (Mean Teacher) method, UA-MT (Uncertainty-Aware Mean Teacher) method, ICT (Interpolation Consistency Training) method, DHC (Dual-debiased Heterogeneous Co-training) method, and STAC (Shape Transformation Driven by Active Contour) method are respectively used for image segmentation simulation experiments, and the experimental results are compared.

[0039] Specifically, both sub-networks of this embodiment adopt the Pytorch 2.4.1 deep learning framework and are implemented in the Ubuntu20.04 operating system environment. The training and testing of the model are both carried out on an NVIDIA GeForce RTX 4090 GPU with a single video memory capacity of 24GB. The parameter settings of the network are specifically as follows: the batch size is set to 2; the maximum number of training epochs is set to ; the optimizer uses the stochastic gradient descent method with momentum, the momentum coefficient is 0.95, and the weight decay coefficient is ; the initial learning rate is set to 0.001; the hyperparameters are set as: , , , , .

[0040] All the training set images, test set images, and 3D medical images to be segmented used in the experiments are from the AMOS database. This database contains 500 CTs and 100 MR abdominal scans, covering a total of 15 abdominal organs. Then the total number of categories including the background 。Two semi-supervised training settings were adopted in the experiment: the proportion of labeled data was 2% and 5% respectively, and the remaining samples were used as unlabeled data. In all experiments, DSC (Dice Similariy Coefficient), ASD (Average Surface Distance), and 95%HD (95% Hausdorff Distance) were used as evaluation metrics. The definition of DSC is , where is the ground truth label region of the -th class, is the predicted segmentation region of the -th class; ASD is defined as , where and represent the contour point sets of the model segmentation prediction region and the ground truth annotation region respectively, represents the contour point and is the Euclidean distance between them, represents the number of contour points of the predicted segmentation region; 95%HD is defined as , where represents the 95%-quantile function, and are the voxel point sets of the segmentation prediction region and the ground truth annotation region respectively. The larger the DSC, the higher the overlap between the prediction result and the ground truth annotation, and the higher the segmentation accuracy; the smaller the ASD value, the closer the predicted contour is to the ground truth annotation, and the smoother and more accurate the segmentation edge; the smaller the 95%HD value, the smaller the edge error of the model in the worst case, and the closer the segmentation edge is to the ground truth annotation.

[0041] Table 1 shows the comparison results of the segmentation performance of the MT method, UA-MT method, ICT method, DHC method, STAC method, and the image segmentation method in the present invention under two annotation ratio settings (2% and 5%) in the AMOS image library. The experimental results show that at the 2% annotation ratio, compared with other methods, the image segmentation method of the present invention achieves the highest DSC value, indicating that the image segmentation method of the present invention has a high segmentation accuracy. In addition, the two indicators of ASD and 95%HD of the image segmentation method of the present invention are also the lowest, showing its good performance in edge segmentation. When the annotation ratio is 5%, the DSC of the image segmentation method of the present invention reaches 53.79%, and the ASD and 95%HD are 4.86 and 10.47 respectively. The overall segmentation performance is also more superior compared with other methods, which further verifies the effectiveness of the proposed active contour deformation data augmentation strategy and edge enhancement module in improving the segmentation performance. In addition, the adopted dual-network collaborative training framework effectively promotes the network to fully learn the labeled and unlabeled data, thereby improving the segmentation performance of the image segmentation method of the present invention under the condition of a very small labeled data ratio.

[0042] Table 1 Comparison of Segmentation Performance of Different Semi-Supervised Models under Different Annotation Ratios

[0043] Furthermore, in order to visually compare the performance differences of each model in the multi-organ segmentation task, 3 3D abdominal CT scan images are selected from the AMOS image library, referring to Figure 4 As shown, it is a comparison schematic diagram of the two-dimensional cross-section and three-dimensional segmentation results of the present invention and different semi-supervised methods under the 5% annotation ratio; specifically, it shows the two-dimensional cross-section segmentation results and the corresponding three-dimensional segmentation visualization results of the MT method, UA-MT method, ICT method, DHC method, STAC method, and the image segmentation method in the present invention for these 3 images (numbered 1-3) under the 5% annotation ratio. Among them, the first and second rows are the two-dimensional cross-section segmentation results and the three-dimensional segmentation visualization results of image No. 1, the third and fourth rows are the two-dimensional cross-section segmentation results and the three-dimensional segmentation visualization results of image No. 2, and the fifth and sixth rows are the two-dimensional cross-section segmentation results and the three-dimensional segmentation visualization results of image No. 3. The first column is the ground truth result, and the second to seventh columns are the segmentation results of the MT method, UA-MT method, ICT method, DHC method, STAC method, and the image segmentation method in the present invention respectively. From Figure 4It can be observed that for small-category organ targets, the image segmentation method of the present invention significantly reduces the phenomenon of mis-segmentation and under-segmentation, and shows better overall segmentation integrity than other methods. In addition, from the three-dimensional volume segmentation results, it can be seen that the image segmentation method of the present invention can maintain the structural integrity of small-category organs, and rarely has the problem of local fracture in other models.

[0044] The three-dimensional medical image segmentation method described in the present invention adopts a dual-network collaborative training framework when training a dual-network segmentation model. Based on the predicted segmentation probability map, soft pseudo-label and predicted signed distance field output by the two sub-networks, the supervision loss of the labeled samples and the consistency loss of the unlabeled samples are calculated for joint optimization, thereby achieving effective learning of labeled samples and unlabeled samples, thereby obtaining stronger generalization capabilities in limited label scenarios, and further optimizing the segmentation performance of the model for three-dimensional medical images. For labeled samples, the cross entropy loss is used to deal with the problem of category imbalance, and the signed distance function loss is used to accurately learn the geometric shape of the target organ to ensure basic segmentation performance. For unlabeled data, the soft pseudo-labels and predicted segmentation probability maps output by the two sub-networks are cross-coordinated and mutually supervised to avoid the prediction deviation of a single model; and the signed distance function consistency is used to force the synergy of geometric prediction and improve the stability of segmentation edges. At the same time, the present invention introduces a simple attention module in the encoder part and integrates the edge enhancement module operation in the decoder part. The present invention uses a simple attention module to perform weighted attention on voxel-level features in three-dimensional medical images, thereby enhancing the perception of key areas. The simple attention module calculates the attention weights through an energy function and does not rely on additional parameters, which can improve the network's perception of key areas in three-dimensional medical images at a lower computational cost. The present invention uses edge enhancement operations to make the network pay more attention to edge areas and enhance the perception of weak edge areas, thereby improving the segmentation performance of weak edge medical images and further improving the segmentation accuracy. The present invention performs data enhancement on three-dimensional medical images, proposes an active contour deformation data enhancement strategy, and uses a level set method to transform the shape of small organ areas in labeled samples and unlabeled samples to enhance the diversity of training samples, thereby alleviating the problem of category imbalance, enhancing the feature representation ability of small category organs, and enhancing the generalization and robustness of the model, so that the trained dual-network segmentation model can have higher segmentation accuracy.

[0045] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0046] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0047] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0048] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0049] Obviously, the above embodiments are only examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to exhaustively list all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.

Claims

1. A three-dimensional medical image segmentation method, characterized in that, Including: Obtain a three-dimensional medical image sample set including labeled samples and unlabeled samples; Input all samples into two sub-networks in parallel of a dual-network segmentation model respectively, obtain two predicted segmentation probability maps for each sample, and calculate the soft pseudo-labels of the unlabeled samples based on the unnormalized classification scores of the unlabeled samples; The sub-network is constructed based on the V-Net network; Using a distance regression head and a hyperbolic tangent function, obtain two predicted signed distance fields for each sample based on the decoded feature maps output by each sub-network for all samples; For each labeled sample: calculate the cross-entropy loss function between each predicted segmentation probability map and the ground truth label to obtain the segmentation loss; calculate the signed distance function loss between each predicted signed distance field and the ground truth signed distance function to obtain the regression loss; add the segmentation loss and the regression loss to obtain the supervised loss of the labeled sample; For each unlabeled sample: calculate the average of the cross-entropy loss and the Dice loss between each predicted segmentation probability map and the soft pseudo-labels to obtain the pseudo-label consistency loss; calculate the signed distance function loss between the two predicted signed distance fields to obtain the signed distance consistency loss; add the pseudo-label consistency loss and the signed distance consistency loss to obtain the consistency loss of the unlabeled sample; Based on the supervised loss of the labeled samples and the consistency loss of the unlabeled samples, obtain the total loss function and train each sub-network in the dual-network segmentation model; Use the trained dual-network segmentation model to obtain the segmentation result of the input three-dimensional medical image to be segmented.

2. The three-dimensional medical image segmentation method according to claim 1, wherein The predicted signed distance field of the sample is expressed as: ; Among them, denotes the predicted signed distance field of class samples after passing through the sub-network ; the sub-network identifier , where A and B respectively represent two sub-networks in parallel of the dual-network segmentation model; the sample annotation status identifier , when , it is an annotated sample, , when , it is an unannotated sample; is the hyperbolic tangent function, denotes the decoded feature map output by the decoder of class samples after passing through the sub-network .

3. The three-dimensional medical image segmentation method according to claim 2, characterized in that Calculating the supervised loss of the labeled samples includes: Calculate the cross-entropy loss function of the predicted segmentation probability map of the labeled samples in each sub-network and its corresponding ground truth label to obtain the segmentation loss , which is expressed as: ; Calculate the signed distance function loss between the predicted signed distance field of the annotated samples in each sub-network and the ground-truth signed distance function to obtain the regression loss , which is expressed as: ; Add the segmentation loss to the regression loss to obtain the supervised loss of the labeled samples , which is expressed as: ; Among them, and respectively represent the learnable weighting factors of the first sub-network A and the second sub-network B; represents the cross-entropy loss, and Y represents the true label of the labeled sample; and respectively represent the predicted segmentation probability maps output by the labeled sample passing through the first sub-network A and the second sub-network B. The calculation formula is ; is the decoded feature map obtained by the labeled sample passing through the sub-network The unnormalized classification scores of all voxels in are expressed as: , represents a 1×1×1 three-dimensional convolutional layer, represents the classification score corresponding to the th category in the labeled sample, , represents the total number of organ categories in the sample; represents the Softmax function; represents the exponential function; and respectively represent the predicted signed distance fields output by the labeled sample passing through the first sub-network A and the second sub-network B; represents the signed distance function loss, which is the average of the L1 loss and the mean square error loss; represents the true signed distance function of the labeled sample, and the expression is: , represents the infimum, represents the voxel point in the true label, represents the voxel point on the surface of the true label, represents the voxel point and The squared Euclidean distance between , and respectively represent the external region, surface and internal region of the true label.

4. The three-dimensional medical image segmentation method according to claim 3, wherein, Calculating the consistency loss of the unlabeled samples includes: Calculate the average of the cross-entropy loss and Dice loss between the predicted segmentation probability map of the unlabeled samples in each sub-network and the soft pseudo-label to obtain the pseudo-label consistency loss , denoted as: ; Calculate the signed distance function loss between the predicted signed distance fields of the unlabeled samples in each sub-network to obtain the signed distance consistency loss , denoted as: ; Add the pseudo-label consistency loss and the signed-distance consistency loss to obtain the consistency loss of unlabeled samples , which is expressed as: ; Among them, and respectively represent the predicted segmentation probability maps output by the first sub-network A and the second sub-network B for the unlabeled samples; represents the segmentation loss, which is the average of the cross-entropy loss and the Dice loss; and respectively represent the soft pseudo-labels corresponding to the unnormalized classification scores output by the first sub-network A and the second sub-network B for the unlabeled samples; in the sub-network the soft pseudo-label ; represents the preset soft pseudo-label smoothing degree parameter, represents the decoded feature map obtained by the unlabeled samples passing through the sub-network The unnormalized classification scores of all voxels in are expressed as: ; represents the unnormalized classification score corresponding to the th category in the unlabeled samples, , represents the total number of organ categories in the samples; and respectively represent the predicted signed distance fields output by the first sub-network A and the second sub-network B for the unlabeled samples.

5. The three-dimensional medical image segmentation method according to claim 4, wherein Supervision loss based on labeled samples Consistency loss with unlabeled samples , to obtain the total loss function , expressed as: ; Among them, represents the consistency loss weight, and the expression is , represents the current training round, represents the preset maximum number of training rounds.

6. The three-dimensional medical image segmentation method according to claim 1, wherein Average and fuse the two predicted segmentation probability maps output by the two sub-networks of the trained dual-network segmentation model to obtain a fused segmentation probability map; perform an argmax operation on the fused segmentation probability map to obtain a predicted segmentation label; based on the predicted segmentation label, obtain the segmentation result of the three-dimensional medical image to be segmented.

7. The three-dimensional medical image segmentation method according to claim 1, wherein Input the three-dimensional medical image into the sub-network based on the V-Net network to obtain the predicted segmentation probability map of the sample, including: Take the three-dimensional medical image sample as the input of the encoder of the sub-network ; where the sub-network identifier , A and B respectively represent two sub-networks in parallel of the dual-network segmentation model; the sample annotation status identifier , when , it is an annotated sample, , it is an unannotated sample; after passing through four cascaded encoding layers, an encoded feature map is output; in each encoding layer, the input feature map passes through a cascaded three-dimensional convolution block and a simple attention module in sequence and then outputs an attention convolution feature map, and after passing through a downsampling unit, the output feature map of each encoding layer is output, expressed as: ; ; represents the output feature map of the -th encoding layer of the sub-network , ; represents a downsampling block, including a three-dimensional convolution block with a stride of 2; represents the attention convolution feature map output after passing through the cascaded three-dimensional convolution block and the simple attention module for times in the -th encoding layer of the sub-network represents a simple attention module; represents a three-dimensional convolution block in the encoder, including a plurality of convolution layers, a batch normalization layer, and a ReLU activation layer cascaded in sequence Input the encoded features into the bottleneck module, and after passing through a series of three-dimensional convolutional blocks and a simple attention module, output a global feature map; Input the global feature map into the decoder, and after passing through four series-connected decoding layers, output a decoded feature map; Input the decoded feature map into the segmentation prediction unit of the output layer to output the corresponding predicted segmentation probability map.

8. The three-dimensional medical image segmentation method according to claim 7, wherein Input the global feature map into the decoder, and after passing through four series-connected decoding layers, output a decoded feature map, including: In each decoding layer, the input feature map is upsampled and then skip-connected to the attention convolution feature map of the corresponding encoding layer. After passing through a 3D convolution block, the output feature map of each decoding layer is obtained, which is expressed as: ; represents the sub-network the output feature map of the -th decoding layer of the sub-network ; represents the 3D convolution block in the decoder, including a plurality of convolution layers, batch normalization layers, and ReLU activation layers connected in series in sequence; represents element-wise addition represents the sub-network the attention convolution feature map of the -th encoding layer of the sub-network; After passing the output feature map of the second decoding layer through a dimension adjustment block and a distance regression head, obtain a predicted signed distance map, and then adjust it using a preset edge sensitivity parameter to obtain an edge-enhanced feature map; Fuse the edge-enhanced feature map and the output feature map of the highest layer to output a decoded feature map.

9. The three-dimensional medical image segmentation method according to claim 8, wherein Obtaining the decoded feature map includes: Calculate the output feature map of the second decoding layer of the predicted symbol distance map , expressed as: ; Based on a preset edge sensitivity parameter and a predicted symbol distance map, an edge-enhanced feature map is obtained , which is expressed as: ; Based on the output feature map of the highest layer of the decoder and the edge enhancement feature map , the decoded feature map is obtained , which is expressed as: ; Among them, represents the distance regression head with output channel 1, represents the dimension adjustment block.

10. The three-dimensional medical image segmentation method according to claim 1, wherein Before inputting the three-dimensional medical image sample into the dual-network segmentation model, it also includes data augmentation for the three-dimensional medical image sample, including: Randomly crop and randomly flip both the labeled samples and the unlabeled samples to obtain the corresponding labeled three-dimensional input sample data and the unlabeled three-dimensional input sample data ; Based on the regional volume of the small-category organ regions contained in each labeled three-dimensional input sample data and the volume of the category with the smallest volume in the small-category organ set calculate the volume weight of the small-category organ region ; ;​ Calculate the signed distance function of the small-class organ region based on the positions of the voxel points in the small-class organ region ; Denote the infimum, Denote the voxel point at the position in the layer slice of the three-dimensional input sample data, Denote the voxel point on the surface of the small-class organ region, Denote the voxel point and the squared Euclidean distance between them; , and respectively denote the external region, surface and internal region in the small-class organ region ; Calculate the active contour deformation field of the target organ region based on the volume weight and signed distance function of the small-category organ region, as well as the fitted Dirac function, which is expressed as: , which is expressed as: ; represents the deformation amplitude control parameter, represents the fitted Dirac function, and its expression is , is a positive constant, is the calculation variable of the fitted Dirac function; represents the gradient of the level set, which is expressed as ; Based on a three-dimensional Gaussian kernel function and the active contour deformation field, calculate the smoothed deformation field ; is the three-dimensional Gaussian kernel function, is the standard deviation, is the convolution operation symbol; Based on the smooth deformation field, the labeled three-dimensional input sample data and its corresponding ground truth label are smoothly deformed to obtain the input samples corresponding to the labeled samples and the enhanced ground truth label , where ⊙ is the element-wise multiplication operator.

Citation Information

Patent Citations

  • Cross-modal medical image registration method based on symbolic distance function collaborative segmentation

    CN116452645A

  • Semi-supervised medical image segmentation method of mutual pseudo supervised edge perception double CNN

    CN118297976A

  • Semi-supervised medical image segmentation method and system based on dual-network adaptive pseudo label generation

    CN119887805A

Cited By

  • Medical image sequence semi-supervised segmentation model construction method and application

    CN120598982A

  • CT image identity determination method, model training method, medium and equipment

    CN121236009A