Semi-Supervised Medical 3D Image Segmentation Method and System Based on Stage Similarity Constraint

By introducing staged similarity constraints in semi-supervised medical 3D image segmentation, using the parallel subnet and similarity constraint loss function, the problems of low accuracy and low computational efficiency in the prior art are solved, and more efficient and accurate medical 3D image segmentation is achieved.

CN119863625BActive Publication Date: 2025-05-27SHANDONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510344328.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-05-27
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The existing semi-supervised medical 3D image segmentation technology has problems of low accuracy and low computational efficiency when processing medical 3D images, especially the lack of effective similarity constraints in the stages before and after feature extraction.

Method used

A semi-supervised medical 3D image segmentation method based on stage similarity constraints is proposed. By constructing parallel subnets A and subnet B, and adding an early feature similarity minimization module in the early stage of feature extraction, adding a global similarity maximization module in the later stage of feature extraction, and optimizing model training using the similarity constraint loss function.

Benefits of technology

The prediction accuracy and computing efficiency of the model on labelless data are improved, the diversity of feature learning and the regularization effect of the model are enhanced, and the segmentation effect and operation accuracy are improved compared with traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863625B_ABST
    Figure CN119863625B_ABST
Patent Text Reader

Abstract

The present invention discloses a semi-supervised medical 3D image segmentation method and system based on stage similarity constraints, which trains an image segmentation model to segment 3D images to obtain segmentation results. When the segmentation model is trained, a sample group is simultaneously input into subnet A and subnet B for feature extraction, respectively obtaining low-dimensional features extracted in the early stage and prediction results. The early feature similarity minimization module calculates the similarity based on the low-dimensional features and minimizes the similarity. The global similarity maximization module calculates the similarity matrices of the feature embeddings of subnet A and subnet B and the class prototypes respectively, and then calculates the global similarity maximization loss to maximize the similarity of the similarity matrices of subnet A and subnet B. Finally, the final model loss is calculated for gradient backpropagation. From the perspective of similarity constraints, similarity-based loss functions are added respectively in the early and late stages of model feature extraction, enabling the model to learn knowledge that is also applicable to unlabeled data from labeled data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semi-supervised medical image processing, and particularly to a semi-supervised medical 3D image segmentation method and system based on staging similarity constraint. Background Technique

[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] The automatic segmentation of medical images is the basis and key component for developing effective computer-aided diagnosis (CAD) systems. Accurate segmentation results can quantitatively analyze the morphological attributes of organs and tissues, providing valuable disease diagnosis information for clinicians. With the development of deep learning, convolutional neural networks have achieved impressive improvements in various automatic image segmentation applications. Although significant progress has been made in automatic medical image segmentation using deep learning, the scarcity of accurately labeled training data remains a major obstacle to the widespread adoption of such technologies in clinical settings.

[0004] As a solution, the concept of semi-supervised segmentation has been proposed to enable the model to be trained using less labeled but abundant unlabeled data.

[0005] The dataset for semi-supervised medical image segmentation includes a small portion of labeled images and a large portion of unlabeled images. The task is to segment the unlabeled data and obtain good segmentation results by learning knowledge from a dataset containing a small portion of labeled images and a large portion of unlabeled images.

[0006] Medical 3D images not only have the characteristic that pixel-by-pixel annotation is extremely labor-intensive, but also the blurred edges and low regional contrast have become obstacles to segmenting precise regions. In recent years, semi-supervised learning methods have proposed a large number of methods for utilizing unlabeled data to improve segmentation accuracy. Generally speaking, they can be roughly divided into two categories: consistency regularization and pseudo-label methods. Methods based on consistency regularization encourage the model to produce invariant predictions for the input image under small perturbations at the data, feature, and model levels. Specifically, they can add small perturbations to unlabeled samples and enforce the consistency of model predictions on the original data and the perturbed data, or directly use adversarial regularization on the entire unlabeled dataset to enforce a similar prediction distribution. While methods based on pseudo-labels mainly follow the self-training pipeline. The key to this method lies in assigning high-quality pseudo-labels to unlabeled data, and then merging the labeled dataset with the pseudo-labeled dataset to retrain the model. In this case, low-quality pseudo-labeling may have higher uncertainty, may contain more noise, and may potentially mislead the model during training. In addition, some recent methods have attempted to combine the two to achieve good performance, that is, the model should produce similar pseudo-labels for the same unlabeled data when receiving different degrees of interference. Although such methods have achieved some results, they still do not comprehensively consider the similarity differences between the features of labeled data and unlabeled data at the pre- and post-feature extraction stages. To sum up, the existing technologies for solving the existing problems of semi-supervised medical 3D image segmentation are not perfect, lacking highly accurate and efficient solutions. Summary of the Invention

[0007] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a semi-supervised medical 3D image segmentation method and system based on staged similarity constraints. From the perspective of similarity constraints, an image segmentation model is constructed, and similarity-based loss functions are added respectively in the early and late stages of the model to extract features, so that the model can learn knowledge that is also applicable to unlabeled data from labeled data.

[0008] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0009] In a first aspect, the present invention provides a semi-supervised medical 3D image segmentation method based on staged similarity constraints, including:

[0010] Obtain labeled samples and unlabeled samples of medical 3D images, and construct a sample group;

[0011] Build an image segmentation model, input the sample group into the image segmentation model for training to obtain a trained image segmentation model; the image segmentation model includes subnet A and subnet B in parallel, and a pre - feature similarity minimization module and a global similarity maximization module are set; when the image segmentation model is trained, the sample group is input into subnet A and subnet B simultaneously for feature extraction, and the pre - extracted low - dimensional features and prediction results are obtained respectively; the pre - feature similarity minimization module calculates the similarity based on the pre - extracted low - dimensional features and makes the similarity minimum; the prediction results are embedded to obtain high - dimensional features in the later stage of feature extraction, and class prototypes and feature embeddings are obtained through screening; next, the global similarity maximization module calculates the similarity matrices of the feature embeddings and class prototypes of subnet A and subnet B respectively, and then calculates the global similarity maximization loss to maximize the similarity of the similarity matrices of subnet A and subnet B; finally, calculate the final model loss and perform gradient backpropagation;

[0012] Based on the trained image segmentation model, segment the medical 3D image to be segmented to obtain an image segmentation result.

[0013] In a further technical solution, the pre - feature similarity minimization module calculates the cosine similarity of the pre - extracted low - dimensional features of subnet A and subnet B, and uses a loss function to make the similarity of the low - dimensional features of subnet A and subnet B minimum.

[0014] In a further technical solution, sample self - attention and cross - sample attention are sequentially added between the encoder and decoder of subnet B.

[0015] In a further technical solution, the image segmentation model adopts a dynamic pseudo - label generation strategy, selects the subnet for generating pseudo - labels based on the size of the supervised loss, and then obtains the pseudo - labels corresponding to the unlabeled samples.

[0016] In a further technical solution, the supervised loss is expressed as:

[0017]

[0018] where, represents the supervised loss, represents the loss weight, represents the Dice loss, represents the true label of the labeled sample, represents the prediction result of subnet A on the labeled sample, represents the prediction result of subnet B on the labeled sample.

[0019] In a further technical solution, the global similarity maximization module calculates the similarity of the feature embeddings and class prototypes of subnet A and subnet B specifically as:

[0020] On the unlabeled samples, compare the per-pixel predicted classes of subnet A and subnet B for the same sample, and select the pixels where subnet A and subnet B have the same prediction;

[0021] For each predicted class, sort the pixels with the same prediction according to the confidence score;

[0022] Select the features of the top set number of pixels as unlabeled features;

[0023] Calculate the cosine similarity between the unlabeled features and the selected class prototypes respectively to obtain a similarity matrix.

[0024] In a further technical solution, the total loss function of the image segmentation model includes a supervised loss and an unsupervised loss, and the unsupervised loss is expressed as:

[0025]

[0026] where represents the similarity minimization loss, represents the unlabeled loss, represents the global similarity maximization loss, , , respectively represent the loss weights of different losses.

[0027] In a second aspect, the present invention provides a semi-supervised medical 3D image segmentation system based on staged similarity constraints, including:

[0028] A data acquisition module, which is configured to: acquire labeled samples and unlabeled samples of medical 3D images and construct a sample group;

[0029] A model training module, which is configured to: construct an image segmentation model, input the sample group into the image segmentation model for training to obtain a trained image segmentation model; the image segmentation model includes subnet A and subnet B arranged in parallel, and a pre-stage feature similarity minimization module and a global similarity maximization module are set; when the image segmentation model is trained, the sample group is simultaneously input into subnet A and subnet B for feature extraction to respectively obtain pre-stage extracted low-dimensional features and prediction results; the pre-stage feature similarity minimization module calculates the similarity based on the pre-stage extracted low-dimensional features and minimizes the similarity; the prediction results are embedded to obtain high-dimensional features in the later stage of feature extraction, and class prototypes and feature embeddings are obtained through screening; next, the global similarity maximization module calculates the similarity matrices of the feature embeddings and class prototypes of subnet A and subnet B respectively, and further calculates the global similarity maximization loss to maximize the similarity of the similarity matrices of subnet A and subnet B; finally, calculate the final model loss and perform gradient backpropagation;

[0030] A model segmentation module, configured to: segment a medical 3D image to be segmented based on the trained image segmentation model to obtain an image segmentation result.

[0031] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the semi-supervised medical 3D image segmentation method based on staging similarity constraints as described in the first aspect are implemented.

[0032] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the semi-supervised medical 3D image segmentation method based on staging similarity constraints as described in the first aspect are implemented.

[0033] The above one or more technical solutions have the following beneficial effects:

[0034] In terms of segmentation effect, the present invention firstly proposes a semi-supervised medical 3D image segmentation model based on staging similarity constraints. The two subnets make predictions for the same sample, which increases the diversity of model predictions. The early feature similarity minimization module further increases the diversity of model feature learning. The cross-sample attention establishes dense cross-sample correlations among a group of samples, realizing the transfer of label prior knowledge to unlabeled data. The global similarity maximization module constrains the similarity matrices of the two subnets to be consistent, regularizing the model. Compared with traditional MT (Mean Teacher)-based methods, the accuracy of model operation is improved.

[0035] In terms of practicality and scalability, the method is based on the pseudo-label method and consistency regularization, combined with the CNN network and the attention mechanism. In the semi-supervised medical 3D image segmentation method based on staging similarity constraints of the present invention, similarity-based loss functions are added respectively in the early and late stages of model feature extraction from the perspective of similarity. A feature similarity minimization constraint is added in the early stage of feature extraction, and a global feature similarity maximization constraint is added in the late stage of feature extraction, enabling the model to learn knowledge applicable to unlabeled data from labeled data.

[0036] In terms of computational efficiency, for the sake of computational efficiency, full correlation calculations are not performed on all labeled and unlabeled pixels. On labeled data, class prototypes are generated using the features where the prediction results of the two subnets are the same as the true labels, and a single memory bank is established for storage. When selecting unlabeled pixels and the memory bank for full correlation calculations, only the pixels where the two subnets make the same prediction are selected, which greatly improves the computational efficiency. Description of the Drawings

[0037] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not unduly limit the invention.

[0038] Figure 1 It is a structural diagram of the image segmentation model of the embodiment of the present invention;

[0039] Figure 2 It is a schematic diagram of cross-sample attention of the embodiment of the present invention;

[0040] Figure 3 It is a schematic diagram of the global similarity maximization module of the embodiment of the present invention;

[0041] Figure 4 It is an example diagram of the segmentation result of the medical 3D image of the embodiment of the present invention;

[0042] Figure 5 It is an example diagram of the ground truth of the semi-supervised medical 3D image based on the staging similarity constraint of the embodiment of the present invention. Detailed implementation manners

[0043] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0044] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0045] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0046] Embodiment 1

[0047] As Figure 1 shown, this embodiment discloses a semi-supervised medical 3D image segmentation method based on staging similarity constraint, and the method includes the following steps:

[0048] S1: Obtain the labeled samples and unlabeled samples of the medical 3D image, and construct a sample group;

[0049] In this embodiment, publicly available datasets on the network are used for experiments, namely The Left Atrial Dataset and The NIH pancreas dataset.

[0050] S2: Build an image segmentation model, input the sample group into the image segmentation model for training to obtain a trained image segmentation model; the image segmentation model includes subnet A and subnet B in parallel, and a pre - feature similarity minimization module and a global similarity maximization module are set; when the image segmentation model is trained, the sample group is input into subnet A and subnet B simultaneously for feature extraction, and the pre - extracted low - dimensional features and prediction results are obtained respectively; the pre - feature similarity minimization module calculates the similarity based on the pre - extracted low - dimensional features and makes the similarity minimum; the prediction result is embedded to obtain the high - dimensional features in the later stage of feature extraction, and class prototypes and feature embeddings are obtained through screening; next, the global similarity maximization module calculates the similarity matrices of the feature embeddings and class prototypes of subnet A and subnet B respectively, and then calculates the global similarity maximization loss to maximize the similarity of the similarity matrices of subnet A and subnet B; finally, calculate the final model loss and perform gradient backpropagation;

[0051] In this embodiment, the image segmentation model includes subnet A and subnet B in parallel. Both subnet A and subnet B are encoder - decoder structures, and both use V - net as the backbone but have different up - sampling methods. On the basis of the backbone network, a pre - feature similarity minimization module and a global similarity maximization module with complementary performance are innovatively designed, which are located in the early and later stages of feature extraction respectively, and a loss function is constructed from the perspective of similarity to constrain the training of the image segmentation model. In addition, the present invention selects a dynamic pseudo - label selection strategy for the selection of pseudo - labels, that is, instead of setting a complex screening strategy, a dynamic pseudo - label generation network is selected, which reduces the computational overhead while achieving more reliable label propagation.

[0052] The pre - feature similarity minimization module is located in the early stage of feature extraction. The low - dimensional features of the encoder in the two sub - networks after the second block are extracted as pre - features. By calculating the cosine similarity of the pre - features of the two sub - networks and using a loss function to minimize the similarity of the low - dimensional features of the two sub - networks, different features are extracted by the two sub - networks, thereby preventing the two sub - networks (subnet A and subnet B) from collapsing with each other and increasing the diversity of the model's learned features and subsequent predictions.

[0053] Specifically, as Figure 1 shown, the sample group is input into subnet B and subnet A respectively. For the labeled samples , extract their pre - features from subnet B , extract their pre - features from subnet A , calculate the minimum similarity between the two in the pre - feature similarity minimization module, and obtain the loss ; for the unlabeled samples , extract their pre - features from subnet B , the preliminary features are extracted from subnet A , and the calculated loss is . Next, and The average value of is used as the loss for minimizing the similarity of preliminary features to participate in the final loss calculation.

[0054] The similarity loss calculation formula is expressed as:

[0055]

[0056] Among them, represents the loss for minimizing the similarity of preliminary features, represents the preliminary features of unlabeled samples or labeled samples extracted by subnet A, represents the preliminary features of unlabeled samples or labeled samples extracted by subnet B.

[0057] Furthermore, in order to enable the image segmentation model to capture global and local information in the input data in different ways and improve the feature expression ability of the model, the present invention sequentially adds sample self-attention and cross-sample attention between the encoder and decoder of subnet B to model the relationships within samples and between samples, denoted as T1 and T2 respectively, Figure 2 is a schematic diagram of cross-sample attention. This mechanism consists of two consecutive Transformer encoder layers, each layer including a multi-head attention and an MLP block with layer normalization after each block. The first Transformer encoder layer performs in-sample self-attention on the spatial dimension of each sample, while the second Transformer encoder layer performs inter-sample self-attention along the batch dimension for further information propagation between different samples. That is, along the sample dimension, pixels at the same spatial position in the samples are fed into the self-attention module to construct cross-sample relationships. In this way, by using two self-attention modules installed sequentially along different dimensions, attention between samples in a batch (batch) is realized on the basis of realizing single-sample attention, and thus efficient mutual attention calculation between all pixels is realized. It should be noted that since the implementation of cross-sample attention requires more than one sample, the subnet with cross-sample attention is not suitable for model verification and testing, so cross-sample attention is only added to subnet B here.

[0058] However, minimizing the similarity of early-stage features may introduce too strong perturbations to the model, making the features extracted by the sub-networks may contain less meaningful information for prediction, resulting in inconsistent and unreliable predictions of the two sub-networks. Therefore, it is necessary to ensure that the sub-networks make meaningful predictions. Specifically, for the labeled data, the true labels are used to supervise the training of the two sub-networks to generate semantically meaningful predictions. For the labeled data (samples), the Dice losses are calculated respectively for the true labels and the predicted labels of the two sub-networks to obtain and , represents the Dice loss of the true label, represents the Dice loss of the predicted labels of the two sub-networks, and the supervised loss composed of the two is as follows:

[0059]

[0060] wherein, represents the true label of the labeled data, represents the prediction result of sub-network A on the labeled data, represents the prediction result of sub-network B on the labeled data, represents the Dice loss, represents the loss weight, which is used to balance the supervised loss.

[0061] For the unlabeled data, each of the two sub-networks generates sub-network pseudo-labels. The present invention uses a dynamic pseudo-label generation strategy to use the accuracy of the labeled data on the two sub-networks as an index of the segmentation accuracy of the two sub-networks, and selects the pseudo-labels generated by the sub-network with higher segmentation accuracy as the final pseudo-labels, and constrains the predictions and the final pseudo-labels generated by the sub-network with poorer accuracy. The corresponding loss function is constructed as follows:

[0062]

[0063]

[0064]

[0065] wherein, represents the pseudo-label supervision loss, represents the pseudo-label supervision loss of sub-network A, represents the pseudo-label supervision loss of sub-network B, represents the mean square error loss, represents the final pseudo-label of the unlabeled data, represents the prediction result of sub-network A on the unlabeled data, It represents the prediction result of subnet B on unlabeled data. The dynamic pseudo-label generation strategy selects the subnet for generating pseudo-labels based on the magnitude of the supervised loss, thereby obtaining the pseudo-labels corresponding to the unlabeled samples. Compared with the fixed setting of using a certain subnet as the network for generating pseudo-labels, such a dynamic selection of the subnet to generate the final pseudo-labels further avoids the problem of low-quality pseudo-labels and further enables the model to fall into cognitive bias.

[0066] Furthermore, the global similarity maximization module is as Figure 3 shown. The global similarity maximization module enables the model to learn the features that both subnets make correct predictions in the labeled data, and further learn the information on the labeled data to improve the overall model segmentation effect.

[0067] The global similarity is calculated from the features obtained by the unlabeled data passing through the subnets and the class prototypes representing the correct prediction features. Attach a projection head to each of the two subnets. After the unlabeled data passes through the subnets and then through the projection heads, two groups of embeddings sampled from the projection features extracted from the unlabeled samples can be obtained. It should be noted that the embeddings are sampled from the embeddings corresponding to the same set of positions in the two subnets.

[0068] Different from other methods that use clustering methods to determine class prototypes, since medical image segmentation is a binary classification problem, the present invention establishes a unified and shared class prototype library for the two subnets to store class prototypes. When selecting class prototypes, it is further fully utilized the labeled data, that is, compare the predicted classes of the two subnets for the same data, and select the features in the labeled data where the predictions of the two subnets are both the same as the true class of the image as class prototypes. However, as the number of training iteration cycles increases and the segmentation accuracy of the training set becomes higher and higher, it is impossible for the present invention to use each prototype that meets the above conditions as a class prototype to calculate the global similarity. Therefore, the present invention sets a fixed number of feature slots for the class prototype library. For each training sample, the class prototype library updates the feature slots corresponding to the labeled samples in the current mini-batch in a query-like manner. In the class prototype library, the present invention randomly samples a certain number of embeddings from each class to calculate the global similarity to increase its diversity.

[0069] To reduce the computational overhead of calculating the global similarity, it is also not possible to calculate the similarity between the embeddings of all unlabeled samples and the class prototypes of the corresponding classes. Here, first compare the per-pixel predicted classes of the two subnets for the same data on the unlabeled samples, and select the pixels where the two subnets have the same prediction. Next, for each class, sort the confidence scores of these pixels, and then select the features of the top i pixels as the sampled unlabeled features. Next, calculate the cosine similarity between the top i pixels selected by the two subnets and the selected class prototypes respectively. and Respectively represent the similarity matrices between subnet A, subnet B, and class prototypes, and the calculation method is as follows:

[0070]

[0071] Among them, represents the number of sampled prototype features, represents the number of class prototypes, represents the temperature coefficient, represents the unlabeled features, represents the high-quality class prototypes. Next, through the CE loss constraint and , while encouraging the two subnets to learn useful semantic information from the class prototypes, it also encourages the two subnets to produce consistent predictions. represents the global similarity maximization loss, which is expressed as follows:

[0072]

[0073] Among them, represents the number of sampled features, represents the CE loss. The CE loss is used to minimize the difference between the two similarity matrices to further align the predictions between the two subnets and reduce the noise in the pseudo-labels, finally obtaining a more accurate segmentation result.

[0074] Furthermore, the total loss function of the model is expressed as follows:

[0075]

[0076] Among them, represents the supervised loss, represents the unsupervised loss.

[0077] The total loss function consists of two parts: supervised loss and unsupervised loss. Among them, the supervised loss is composed of three parts, which is expressed as follows:

[0078]

[0079] Among them, represents the similarity minimization loss, represents the unlabeled loss, represents the global similarity maximization loss, , , respectively represent the loss weights of different losses.

[0080] In this embodiment, a training of an image segmentation model is achieved by using a data path, a model storage path, etc., as well as training parameters such as initialization, deviation, regularization, an initial learning rate, a learning rate reduction method, an optimization algorithm, the number of iterations, a data augmentation method, etc.

[0081] It should be noted that, in addition to only one subnet adding a cross-sample attention mechanism, the upsampling methods of the two subnets are also different, which further ensures the diversity of the predictions of the two subnets.

[0082] During the training process, N labeled samples are randomly selected from the training set of the dataset in each epoch, and N is determined by the proportion of the labeled samples in the overall training set. In each iteration, two samples are selected from the selected labeled data, and two samples are randomly selected from the remaining unlabeled data to form a new set of samples and are respectively fed into the two subnets. After a set of samples are respectively fed into the two subnets, the early-stage features of the two subnets for a set of samples are extracted for the calculation of the early-stage feature minimization module. Next, the model selects the final pseudo-label based on the dice loss of the subnet on the labeled data and narrows the prediction result and the result of the final pseudo-label. The prediction results generated by the two subnets pass through the embedding module to obtain the late-stage features, and then the class prototypes and feature embeddings are obtained. The similarities between the embedding features of the two subnets and the class prototypes under the corresponding classifications are respectively calculated and . After the final loss is calculated, the gradient is backpropagated.

[0083] The model testing process is similar to the training process, such as setting the input image and using the model. In one embodiment, the input includes a test folder path, a test model path, a test model, the number of test images, and a test result output path. The finally trained model is tested on the entire test set, and the test result returns the average result of the test set. Finally, the visualization of the test result is performed to display the 3D medical image segmentation result generated by the model.

[0084] S3: Segment the medical 3D image to be segmented based on the trained segmentation model to obtain an image segmentation result.

[0085] In this embodiment, Figure 4 is an example diagram of the segmentation result of a certain medical 3D image to be segmented, Figure 5 is an example diagram of the ground truth of the medical 3D image to be segmented. By comparing the two, it can be seen that the method of this embodiment improves the segmentation result of the medical 3D image.

[0086] Embodiment 2

[0087] This embodiment discloses a semi-supervised medical 3D image segmentation system based on staging similarity constraints, including:

[0088] A data acquisition module, which is configured to: acquire labeled samples and unlabeled samples of medical 3D images, and construct a sample group;

[0089] A model training module, which is configured to: construct an image segmentation model, input the sample group into the image segmentation model for training, and obtain a trained image segmentation model; the image segmentation model includes subnet A and subnet B arranged in parallel, and a pre - feature similarity minimization module and a global similarity maximization module are set; when the image segmentation model is trained, the sample group is input into subnet A and subnet B simultaneously for feature extraction, and a pre - extracted low - dimensional feature and a prediction result are obtained respectively; the pre - feature similarity minimization module calculates the similarity based on the pre - extracted low - dimensional feature and makes the similarity minimum; the prediction result is embedded to obtain a high - dimensional feature in the later stage of feature extraction, and a class prototype and a feature embedding are obtained through screening; next, the global similarity maximization module calculates the similarity matrices of the feature embeddings of subnet A and subnet B and the class prototype respectively, and then calculates the global similarity maximization loss to maximize the similarity of the similarity matrices of subnet A and subnet B; finally, calculate the final model loss and perform gradient backpropagation;

[0090] A model segmentation module, which is configured to: segment the medical 3D image to be segmented based on the trained image segmentation model, and obtain an image segmentation result.

[0091] Embodiment III

[0092] The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method in Embodiment I are implemented.

[0093] Embodiment IV

[0094] The purpose of this embodiment is to provide a computer - readable storage medium. A computer - readable storage medium stores a computer program, and when the program is executed by a processor, the steps of the method in Embodiment I are executed.

[0095] The steps involved in the devices in Embodiments III and IV above correspond to those in Method Embodiment I. For specific implementation manners, reference can be made to the relevant description part of Embodiment I. The term "computer - readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0096] Those skilled in the art should understand that each module or step of the present invention described above can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0097] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0098] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts by those skilled in the art on the basis of the technical solution of the present invention are still within the protection scope of the present invention.

Claims

1. A semi-supervised medical 3D image segmentation method based on stage similarity constraint, characterized in that: include: Obtain labeled samples and unlabeled samples of medical 3D images and construct a sample group; An image segmentation model is constructed, and the sample group is input into the image segmentation model for training to obtain a trained image segmentation model; the image segmentation model includes a parallel subnet A and a subnet B, and a pre-feature similarity minimization module and a global similarity maximization module are set; when the image segmentation model is trained, the sample group is simultaneously input into subnet A and subnet B for feature extraction, and low-dimensional features extracted in the pre-feature and prediction results are obtained respectively; the pre-feature similarity minimization module calculates similarity based on the low-dimensional features extracted in the pre-feature and minimizes the similarity; the prediction result is embedded to obtain high-dimensional features in the post-feature extraction stage, and a class prototype and feature embedding are obtained by screening; Next, the global similarity maximization module calculates the feature embedding and similarity matrix of the class prototype of subnet A and subnet B respectively, and then calculates the global similarity maximization loss to maximize the similarity of the similarity matrix of subnet A and subnet B; finally, the final model loss is calculated and the gradient is back-propagated; The medical 3D image to be segmented is segmented based on the trained image segmentation model to obtain an image segmentation result.

2. The semi-supervised medical 3D image segmentation method based on stage similarity constraint according to claim 1, characterized in that: The early feature similarity minimization module calculates the cosine similarity of the low-dimensional features extracted in the early stage of subnet A and subnet B, and uses the loss function to minimize the low-dimensional feature similarity of subnet A and subnet B.

3. The semi-supervised medical 3D image segmentation method based on stage similarity constraint according to claim 1, characterized in that: Sample self-attention and cross-sample attention are added between the encoder and decoder of the subnetwork B in sequence.

4. The semi-supervised medical 3D image segmentation method based on stage similarity constraint according to claim 1, characterized in that: The image segmentation model adopts a dynamic pseudo-label generation strategy, selects a subnet to generate pseudo-labels based on the size of the supervised loss, and then obtains the pseudo-labels corresponding to the unlabeled samples.

5. The semi-supervised medical 3D image segmentation method based on stage similarity constraint according to claim 4, characterized in that: The supervised loss is expressed as: in, indicates that there is supervised loss, represents the loss weight, represents DIce loss, represents the true label of the labeled sample, represents the prediction result of subnet A on labeled samples, Represents the prediction results of subnet B on labeled samples.

6. The semi-supervised medical 3D image segmentation method based on stage similarity constraint according to claim 1, characterized in that: The global similarity maximization module calculates the similarity of feature embedding and class prototype of subnet A and subnet B respectively as follows: Compare the pixel-by-pixel prediction categories of subnet A and subnet B for the same sample on unlabeled samples, and select pixels with the same predictions by subnet A and subnet B; For each predicted class, pixels with the same prediction are sorted according to the confidence score; Select the features of the top set number of pixels as unlabeled features; The cosine similarity between the unlabeled features and the selected class prototypes is calculated to obtain a similarity matrix.

7. The semi-supervised medical 3D image segmentation method based on stage similarity constraint according to claim 1, characterized in that: The total loss function of the image segmentation model includes supervised loss and unsupervised loss, and the unsupervised loss is expressed as: in, represents the similarity minimization loss, represents the unlabeled loss, represents the global similarity maximization loss, , , They represent the loss weights of different losses respectively.

8. A semi-supervised medical 3D image segmentation system based on stage similarity constraints, characterized in that: include: A data acquisition module is configured to: acquire labeled samples and unlabeled samples of medical 3D images and construct a sample group; The model training module is configured to: construct an image segmentation model, input the sample group into the image segmentation model for training, and obtain a trained image segmentation model; the image segmentation model includes parallel subnets A and B, and sets a pre-feature similarity minimization module and a global similarity maximization module; when the image segmentation model is trained, the sample group is simultaneously input into subnets A and B for feature extraction, and low-dimensional features extracted in the pre-feature and prediction results are obtained respectively; the pre-feature similarity minimization module calculates similarity based on the low-dimensional features extracted in the pre-feature and minimizes the similarity; the prediction results are embedded to obtain high-dimensional features in the post-feature extraction stage, and the class prototype and feature embedding are obtained by screening; Next, the global similarity maximization module calculates the feature embedding and similarity matrix of the class prototype of subnet A and subnet B respectively, and then calculates the global similarity maximization loss to maximize the similarity of the similarity matrix of subnet A and subnet B; finally, the final model loss is calculated and the gradient is back-propagated; The model segmentation module is configured to segment the medical 3D image to be segmented based on the trained image segmentation model to obtain an image segmentation result.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the semi-supervised medical 3D image segmentation method based on stage similarity constraint as described in any one of claims 1 to 7 are implemented.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the semi-supervised medical 3D image segmentation method based on stage similarity constraint as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • OCT retina image field adaptive segmentation method and system

    CN113096137A

  • Multi-tissue-component image segmentation method and system based on semi-supervised prototype feature alignment

    CN118071765A