Heart MRI (Magnetic Resonance Imaging) image segmentation method, system, equipment and medium
Through implicit neural network and information migration enhancement methods, the information distribution mismatch between labeled data and unlabeled data in cardiac MRI image segmentation is solved, efficient feature alignment and detail retention are achieved, and segmentation accuracy is improved.
Patent Information
- Application Number
- CN202510368047.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has a problem of mismatch in information distribution between labeled data and unlabeled data in cardiac MRI image segmentation, which leads to the inability to effectively transfer knowledge and feature alignment, especially when high-resolution image processing is severely lost.
Implicit neural network is used to align features, and through the information migration enhancement method, label data and pseudo-label data are mixed, pseudo-label data is generated using the teacher network, and pseudo-label network is optimized to realize the alignment and aggregation of features at different levels.
Effectively reduce the distribution gap between labeled and unlabeled data, retain the learned knowledge, improve the segmentation accuracy of heart and boundary details, and perform well especially in low labeled data.
Smart Images

Figure CN120298430A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to image segmentation, and particularly relates to a method, a system, a device and a medium for segmenting cardiac MRI images. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Segmenting the internal structure of magnetic resonance imaging (MRI) to divide structures or tissues, so that doctors can diagnose diseases, formulate treatment plans and perform surgical planning more accurately, is very important for many clinical applications. At present, a variety of medical image segmentation techniques based on supervised learning have been proposed, but this usually requires a large amount of labeled data. Given that the acquisition of medical image annotation is very complex and expensive, and the fine processing of contours also requires a great deal of manual effort, this has promoted the attention of semi-supervised methods that use a small amount of labeled data and a large amount of unlabeled data. In recent years, semi-supervised segmentation has attracted more and more attention and has been widely used in the field of medical image analysis. Generally, in semi-supervised medical image segmentation, the problem of information distribution mismatch between a large amount of labeled data and a small amount of unlabeled data is widespread.
[0004] Existing models mostly adopt CNN or Transformer architectures for MRI images. There are some problems in their processing of MRI image features. For example, when dealing with high-resolution images, convolutional operations may not be able to effectively capture the fine-grained structures in the images, especially in the extraction of fine features in edge and small-scale regions. In addition, although Transformer can handle long-range dependencies, it often performs worse than CNN in learning local features of images. Both may be affected by resolution limitations, resulting in detail loss in segmentation tasks, especially for complex organ or tumor morphologies. However, implicit neural networks (INNs) represent the target structure implicitly by mapping coordinates to density fields or segmentation labels. This method can capture detailed geometric features while avoiding the detail loss caused by resolution limitations in traditional explicit segmentation methods. Implicit neural networks represent through a continuous coordinate space and do not rely on a fixed grid structure, thus having stronger representation capabilities and being able to accurately depict complex boundaries and tiny structures in MRI images. In addition, implicit neural networks can maintain efficient feature representations at different resolutions, avoiding information loss in traditional methods, and having better generalization capabilities, especially performing prominently in few-shot learning and low-contrast image processing. And to effectively integrate information at different levels in medical images, the most widely used methods, such as U-Net and its variants and transformers, adopt bilinear upsampling and convolutional operations to align low-resolution deep features with high-resolution shallow features. However, bilinear upsampling has the problem of blurring the precise context learned in deep features, and when convolutional operations are used to handle high-resolution segmentation tasks, the resolution difference between high-level context and low-level details may be quite significant.
[0005] To solve the above problems, most semi-supervised methods used by researchers are basically centered around symmetrically training the consistency relationship between labeled data and unlabeled data, which leads to the knowledge learned from the labeled data part being discarded, and the knowledge learned from the unlabeled data being difficult to apply to the labeled data. The resulting empirical distribution mismatch, as well as the inability of traditional methods to perform efficient and accurate feature alignment, requires further improvement. Summary of the Invention
[0006] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a method, system, device and medium for cardiac MRI image segmentation. Through the data augmentation method of information migration, the distribution gap between labeled data / pseudo-labeled data and unlabeled data can be effectively reduced, and the learned knowledge can be retained to the greatest extent.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a method for segmenting cardiac MRI images, comprising:
[0009] Obtaining a cardiac MRI image to be processed;
[0010] Inputting the cardiac MRI image to be processed into a trained teacher network to obtain a segmentation result predicted based on the features continuously aligned by the feature alignment implicit function;
[0011] Wherein, the training process of the teacher network is as follows:
[0012] Obtaining a cardiac MRI image training dataset, which includes labeled data and unlabeled data;
[0013] Performing information transfer enhancement on the pseudo-labeled data and the labeled data through a randomly generated zero-centered mask to obtain mixed input data;
[0014] Training a student network using the mixed input data, and updating the network parameters of the teacher network according to the network parameters of the trained student network;
[0015] Generating pseudo-labels for the unlabeled data using the teacher network, mixing the labeled data and the pseudo-labels through a randomly generated zero-centered mask to generate a supervision signal, and guiding and optimizing the training of the student network based on the supervision signal.
[0016] In a second aspect, the present invention provides a cardiac MRI image segmentation system, comprising:
[0017] An acquisition module configured to: obtain a cardiac MRI image to be processed;
[0018] A segmentation prediction module configured to: input the cardiac MRI image to be processed into a trained teacher network to obtain a segmentation result predicted based on the features continuously aligned by the feature alignment implicit function;
[0019] Wherein, the training process of the teacher network is as follows:
[0020] Obtaining a cardiac MRI image training dataset, which includes labeled data and unlabeled data;
[0021] Performing information transfer enhancement on the pseudo-labeled data and the labeled data through a randomly generated zero-centered mask to obtain mixed input data;
[0022] Training a student network using the mixed input data, and updating the network parameters of the teacher network according to the network parameters of the trained student network;
[0023] Generate pseudo - labels for the unlabeled data using the teacher network, mix the labeled data and the pseudo - labels through a randomly generated zero - centered mask to generate a supervision signal, and guide and optimize the training of the student network based on the supervision signal.
[0024] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.
[0025] In a fourth aspect, the present invention provides a computer - readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.
[0026] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.
[0027] The above - mentioned one or more technical solutions have the following beneficial effects:
[0028] In the present invention, information migration is carried out between labeled / pseudo - labeled data and unlabeled data. Through this data augmentation method, the distribution gap between labeled / pseudo - labeled data and unlabeled data can be effectively reduced, and the learned knowledge is retained to the greatest extent. Using an implicit neural network for alignment can map each feature map at different levels to a continuous feature map, so as to query and align its features at any coordinate position, and conveniently and quickly aggregate features at different levels and of different shapes, achieving great results in the segmentation of the heart and the processing of boundary details.
[0029] Advantages of additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0031] Figure 1 It is the overall architecture diagram of the INIT network and information migration proposed in the first embodiment of the present invention;
[0032] Figure 2 It is the specific structure diagram of each part of the INIT network in the first embodiment of the present invention;
[0033] Figure 3This is the qualitative result graph of various image segmentation methods in Embodiment 1 of the present invention on the ACDC cardiac medical MRI dataset using 10% labeled data;
[0034] Figure 4 This is the qualitative result graph of various image segmentation methods in Embodiment 1 of the present invention on the SAX cardiac medical MRI dataset using 10% labeled data;
[0035] Figure 5 This is the generalization verification result on the LAX dataset;
[0036] Figure 6 This is the ablation result graph of different information migration directions in Embodiment 1 of the present invention on the ACDC cardiac medical MRI dataset using 10% labeled data;
[0037] Figure 7 This is the ablation result graph of different information migration directions in Embodiment 1 of the present invention on the SAX cardiac medical MRI dataset using 10% labeled data;
[0038] Figure 8 This is the quantitative metric result of training in Embodiment 1 of the present invention on the ACDC dataset using 5%, 10%, and 20% labeled data;
[0039] Figure 9 This is the quantitative metric result of training in Embodiment 1 of the present invention on the internal SAX dataset using 5%, 10%, and 20% labeled data;
[0040] Figure 10 This is the ablation study result of different information migration directions in Embodiment 1 of the present invention;
[0041] Figure 11 This is the ablation study result of selecting the hyperparameter α on the ACDC dataset in Embodiment 1 of the present invention;
[0042] Figure 12 This is the ablation study result of selecting the hyperparameter β on the ACDC dataset in Embodiment 1 of the present invention;
[0043] Figure 13 This is the ablation study result of comparing the functional effects of each implicit network component on the ACDC dataset in Embodiment 1 of the present invention;
[0044] Figure 14 This is the ablation study result of determining the number of hidden layers in the multi-layer perceptron (MLP) and comparing on the ACDC dataset in Embodiment 1 of the present invention;
[0045] Figure 15This is a double - blind reader study in the ACDC dataset in the first embodiment of the present invention to compare the method of this embodiment with other state - of - the - art (SOTA) methods. Detailed implementation manners
[0046] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0047] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention.
[0048] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0049] Embodiment 1
[0050] This embodiment discloses a method for segmenting cardiac MRI images, including:
[0051] Obtain a cardiac MRI image to be processed;
[0052] Input the cardiac MRI image to be processed into a trained teacher network to obtain a segmentation result predicted based on the continuously aligned features of the feature - aligned implicit function.
[0053] Among them, the training process of the teacher network is as follows:
[0054] Obtain a cardiac MRI image training dataset, where the cardiac MRI image training dataset includes labeled data and unlabeled data;
[0055] Perform information transfer enhancement on the pseudo - labeled data and the labeled data through a randomly generated zero - centered mask to obtain mixed input data;
[0056] Use the mixed input data to train a student network, and update the network parameters of the teacher network according to the network parameters of the trained student network;
[0057] Generate pseudo - labels for the unlabeled data using the teacher network, mix the labeled data and the pseudo - labels through a randomly generated zero - centered mask to generate a supervision signal, and guide and optimize the training of the student network based on the supervision signal.
[0058] This embodiment proposes a method for information transfer semi - supervised cardiac MRI segmentation based on an implicit neural network. By training a student network, the weights are transferred to the teacher network in the EMA (Exponential Moving Average) manner to generate pseudo - labels, which form the supervision signal. For the dataset, information transfer is performed between the labeled data and the unlabeled data. Through this data augmentation method, the distribution gap between the labeled data and the unlabeled data can be effectively reduced, and the learned knowledge is retained to the greatest extent.
[0059] The backbone network of this embodiment, based on the traditional decoder - encoder, uses an implicit neural network to transform the input image into a continuous feature space representation, retaining the most feature information. Moreover, a novel alignment method (FAI) is adopted to map each feature map at different levels into a continuous feature map, so as to achieve querying and alignment of its features at any coordinate position, facilitating the aggregation of features at different levels and different shapes quickly and conveniently, and achieving great results in the segmentation of the heart and the processing of boundary details.
[0060] The following is a detailed description of a cardiac MRI image segmentation method proposed in this embodiment, including:
[0061] Define the three - dimensional volume of the medical image as \(T\in R\) W×H×L , and the goal of semi - supervised medical image segmentation is to predict each voxel label map \(Q\in\{0,1,\ldots,K - 1\}\) W×H×L , which represents the positions of the background and the target in \(T\), \(K\) is the number of categories, and \(W\), \(H\), \(L\) represent the width, height, and depth of the image respectively.
[0062] The training set \(D\) consists of \(N\) labeled data and \(M\) unlabeled data (\(N\ll M\)), denoted as two subsets: \(D = D\) l \(\cup D\) u , and the purpose is to segment the key regions in the input image (right ventricle, left ventricle, and myocardium).
[0063] In the framework of this embodiment, there is a teacher network and a student network \(N\) s \((T_1,T_2;\Delta\) s ), where and are unlabeled data, \(T_1\), \(T_2\) are the data after information transfer, and \(\Delta\) t and \(\Delta\) s are the network parameters in the teacher network and the student network respectively. The student network is optimized using stochastic gradient descent, and the teacher network is optimized using the exponential moving average (EMA) of the student network.
[0064] The student network training strategy is divided into three steps: First, pre-train the student network only using the labeled data, and then use the pre-trained student network to update the network parameters Δ of the teacher network. t Generate pseudo-labels for unlabeled images. In each iteration, first optimize the network parameters Δ of the student network using stochastic gradient descent. s Finally, use the exponential moving average (EMA) of the network parameters Δ of the student network to update the network parameters Δ of the teacher network. s t
[0065] When training the student network, first process the input data and use information transfer as data augmentation. The specific process is as Figure 2 shown. First, randomly generate a zero-center mask M ∈ {0, 1} W×H×L . Random generation means that the center may be close to the boundary. In this way, different feature parts are randomly transferred to enhance generalization, indicating whether this part belongs to labeled data (1) or unlabeled data (0). The size of the central region m of the zero-center mask is αW × αH × αL, where α is a hyperparameter ∈ (0, 1), and the value of α is determined by ablation experiments.
[0066] Both labeled and unlabeled data are used. By different data combinations, the quantity and diversity of the input dataset are maintained. Different input data T1 and T2 are formed through different combination methods, that is, different input data.
[0067]
[0068] Among them, i ≠ j, p ≠ q, 1 ∈ {1} W×H×L , ⊙ represents element-wise multiplication, D l and D u represent the labeled and unlabeled datasets respectively, and T1 and T2 represent different input data.
[0069] After processing the input data, in the overall network, the mixed input data T1 and T2 extract different-scale features through an encoder, restore different-scale features through a decoder, and send the features output by the decoder into a feature alignment implicit function (FAI) and then into a decoding function f θ (MLP).
[0070] Figure 1Shows the overall network architecture. The mixed input data extracts features of different scales through the encoder, restores features of different scales through the decoder, and sends the features output by the decoder into the Feature Alignment Implicit function (FAI) and then into the decoding function f θ (MLP), and finally obtains the segmentation result.
[0071] Specifically, the encoder part passes the input through two loops of 3×3 convolution, batch normalization, and Leaky activation function in sequence to obtain the output result. The decoder part passes the output of the encoder through 1×1 convolution and the upsampling part, and then splices the input features of the decoder through skip connection, and then passes through two loops of 3×3 convolution, batch normalization, and Leaky activation function to obtain the output result of the decoder.
[0072] Feature Alignment Implicit function: The function mainly depends on the implicit neural network; defines a decoding function f θ (usually a multi-layer perceptron MLP) to obtain a continuous feature map F. Given a discrete feature map, the feature vectors are regarded as latent codes evenly distributed in the two-dimensional space. For features of different levels, after dynamic hierarchical feature weighting, each feature vector is assigned a two-dimensional coordinate.
[0073] The eigenvalue of F at x q is defined as F(x q ):
[0074] F(x q ) = f θ (S′, x q - x′)
[0075] where S′ is the latent code closest to the query coordinate x q , and x′ is the coordinate of the latent code S′.
[0076] As Figure 2 (d) shows, regard the features in the feature map as latent codes evenly distributed in the two-dimensional space. Given a query coordinate x q , find the latent code closest to it for each feature map.
[0077] Using the decoding function f θ , a continuous feature map F can be defined on a discrete feature map. In actual operation, f θ is jointly learned with the encoder, so that the learned features can accurately represent the continuous information domain.
[0078] To better extract heart features, a feature weight map W is dynamically assigned to each feature at different levels output by the decoder. It is a learnable parameter map related to the spatial coordinate x′. Each element of the feature weight map W represents the weight or importance of the corresponding position for feature alignment. Taking the aligned feature map as an example, the implicit feature function is extended to a feature-aligned implicit function (FAI) (Feature-aligned Implicit Function), which directly defines the continuous feature map F on multi-level discrete feature maps with different resolutions. The FAI function realizes the process of mapping feature maps with different resolutions to the continuous feature map, and F is the feature map processed by the FAI function.
[0079] Specifically, the value of F at x q is defined as:
[0080]
[0081] where F(x q ) refers to the feature after the query coordinate x q is processed by the feature-aligned implicit function.
[0082] For the current coordinate, the nearest latent code is obtained from the features at each level, denoted as The feature weight maps at different levels are The relative coordinates and the corresponding encoded values are denoted as (·) is element-wise multiplication; S′ is the latent code closest to the query coordinate x q , and x′ is the coordinate of the latent code S′.
[0083] Among them, the weight feature maps at different levels are randomly generated. For weight maps of different sizes, the loss function is calculated based on the obtained results and GT during the training process, and they are dynamically updated.
[0084] After assigning weights to the latent codes at different levels and concatenating them with the relative coordinates, they are input into the MLP of the decoding function f θ . Intuitively, each latent code still represents a feature field that can be decoded through the relative coordinates, and f θ can decode the features at each level and simultaneously model the interactions between different levels.
[0085] The same method is also used for the alignment between the features generated in different stages. The features at different levels use different coordinate upper limits.
[0086] The feature vector finally obtained and input into the decoder is input into the feature-aligned implicit function (FAI) to obtain the updated feature.
[0087] The student network and the teacher network have completed the network update through the feature optimization of the FAI function:
[0088] N′ t = FAI(N t (Δ t ))
[0089] N′ s = FAI(N s (Δ s ))
[0090] Among them, Δ t , Δ s are the teacher network parameters and the student network parameters respectively, N t , N s are the teacher network and the student network respectively, N t (Δ t ) is the output of the teacher network with the parameter Δ t , and N s (Δ s ) is the output of the student network with the parameter Δ s .
[0091] Supervisory Signals / Pseudo Label: For the pseudo labels of semi-supervised training, only the labeled data is used to pre-train the backbone network to initialize the network weights.
[0092] After pre-training, the unlabeled data is fed into the teacher network to generate pseudo labels The initial determination of the pseudo label is obtained through the argmax of the pre-trained student network. The true labels of the labeled data and and the pseudo labels of the unlabeled data and undergo the same information transfer to obtain the outputs Y1 and Y2, and the supervisory signals are used to supervise the predictions of the student network for T1 and T2.
[0093]
[0094] Among them, and represent the pseudo labels of the unlabeled data generated by the teacher network. Thus, a method corresponding to the data augmentation part is proposed, and the true labels of the labeled data and and the pseudo labels of the unlabeled data and undergo the same information transfer:
[0095]
[0096] The resulting Y1 and Y2 will be used as supervision signals to supervise the predictions of the student network for T1 and T2.
[0097] Each input image of the student network consists of labeled images and unlabeled images. Intuitively, the GT mask of the labeled image is usually more accurate than the pseudo-label of the unlabeled image. The hyperparameter β is used to control the contribution of the unlabeled image pixels to the loss function. The input T1 and T2 are calculated through the student network to obtain the predicted Q1 and Q2. Through the linear combination L of the Dice loss and the cross-entropy loss seg Calculate the losses of Q1 and Q2 with Y1 and Y2. Use this to update the parameters of the student network and update the parameters of the teacher network by the EMA method.
[0098] The loss functions of T1 and T2 are calculated respectively as:
[0099] L1 = L seg (Q1, Y1) ⊙ M + βL seg (Q1, Y1) ⊙ (1 - M)
[0100] L2 = L seg (Q2, Y2) ⊙ (1 - M) + βL seg (Q2, Y2) ⊙ M
[0101] where L seg is the linear combination of the Dice loss and the cross-entropy loss.
[0102] Q1 and Q2 are calculated by the student network:
[0103] Q1 = N s ′(T1; Δ S ), Q2 = N s ′(T2; Δ S )
[0104] In each iteration, use the loss function L all to update the parameters Δ in the student network by the stochastic gradient descent method S :
[0105] L all = L1 + L2
[0106] Then update the parameters of the teacher network at the (k + 1)-th iteration
[0107]
[0108] Among them, ρ is the smoothing coefficient parameter.
[0109] Based on this, the overall framework of deep learning is completed. According to L all Backpropagation is performed to update the network parameters to obtain the final result.
[0110] To verify the practicality of the present invention, two datasets are used to evaluate the effectiveness of the INIT method proposed in this embodiment. One is the public Automatic Cardiac Diagnosis Challenge (ACDC) dataset; the other is the internal cardiac SAX dataset. The ACDC dataset is a four-class dataset for background, right ventricle, left ventricle, and myocardial segmentation, containing scans of 100 patients. The data split is fixed at 70, 10, and 20 patient scans for training, validation, and testing. The internal dataset SAX consists of 120 cases, with 8 - 11 slices per case. A total of 1194 pairs of Image-GT images are obtained. The data split in the dataset is fixed at a 7:1:2 ratio and is used for the training / validation / testing subsets.
[0111] All experiments were implemented with INIT using a fixed random seed on an NVIDIA Tesla A100 GPU. The pre-training and self-training processes were set to 10000 and 30000 epochs respectively. Data augmentation was performed using information transfer operations, and the model of this embodiment was trained with an SGD optimizer with an initial learning rate of 0.01, decaying by 10% every 2.5K iterations. In addition, α and β were set to 0.5 in this experiment. The batch size was 4. Four of the most common and general evaluation metrics were selected: dice score (%), Jaccard score (%), 95% Hausdorff distance (95HD), and average surface distance (ASD). For two object regions, Dice and Jaccard mainly calculate the overlap percentage between them, ASD calculates the average distance between their boundaries, and 95HD measures the distance between their closest points. In addition, the focus was placed on 10% of the labeled data of all comparison methods.
[0112] Qualitative Results: INIT was compared with several state-of-the-art methods, and the network of this embodiment was compared with DTC, URPC, SS-Net, and BCP. To compare the results obtained by different methods, all these algorithms were implemented, trained, and tested on the same dataset, and their default settings were provided in the public source code. The same proportion of labeled data was used in terms of data volume: 5%, 10%, 20%. When processing data, following the previous techniques such as SS-Net and BCP, two-dimensional slices were used to train the network of this embodiment. It should not be ignored that a 3D data can be sliced into multiple 2D slices, so more combinations can be generated from labeled and unlabeled slices than using 3D data. Therefore, during the training process, the knowledge of labeled data can be more fully transferred to unlabeled data, especially when the number of labels is very small.
[0113] The results of the comparative experiments are as Figure 4 , Figure 5 shown. To demonstrate the effectiveness of the method proposed in this embodiment, the first row shows the prediction results on a single slice after data segmentation, and the second row shows its overall 3D visualization. Among them, the red, green, and blue colors represent the right ventricle, myocardium, and left ventricle respectively. For the SAX dataset, to show the differences in its dataset, the colors of these three parts are swapped, that is, red: left ventricle, green: myocardium, blue: right ventricle. In these examples, it can be clearly seen that other methods have problems such as incorrect predictions and missing edges. In contrast, the method of this embodiment shows superior performance in overall segmentation and edge retention. At the same time, a new private dataset LAX was used to test the weights trained on the SAX dataset to verify the generalization ability of the network of this embodiment. As Figure 7 shown, the network of this embodiment shows the best performance in four metrics, successfully verifying the high generalization ability of the proposed network.
[0114] Quantitative Results: The comparison metrics between the method of this embodiment and other state-of-the-art (SOTA) methods are as Figure 8 shown. The quantitative measurement results of the ACDC and internal datasets in terms of Dice, Jaccard, 95HD, and ASD at 5%, 10%, and 20% are as Figure 8 , Figure 9 shown, showing the best performance at each proportion, superior to other competitors. The average performance of the four types of segmentation results on the ACDC dataset shows that in the case of a labeling rate of 5%, the solution of this embodiment achieved a 1.72% performance improvement, in the case of a labeling rate of 10%, the solution of this embodiment achieved a 1.35% improvement, and in the case of a labeling rate of 20%, the solution of this embodiment achieved a 1.21% improvement. The results show that the method of this embodiment shows better advantages when the amount of labeled data is small.
[0115] To measure the effectiveness of the method in this embodiment on other datasets, the internal dataset SAX was used. Similarly, under the labeled data volumes of 5%, 10%, and 20%, the network structure of this embodiment outperforms other state-of-the-art experiments. As can be seen from Figure 9 , the improvement is still the largest under the 5% annotation volume, once again demonstrating the powerful segmentation performance of the network in this embodiment under low labeled data volumes in semi-supervised methods.
[0116] Ablation experiments: Ablation studies were conducted to show the impact of each component in the network. This includes the information transfer direction, two hyperparameters α and β, each component of the implicit network, and the number of layers of the MLP hidden layer in the implicit network.
[0117] In the first ablation study, three different ablation experiments were conducted for the information transfer direction: forward transfer, backward transfer, and self-transfer. They respectively represent the transfer from labeled to unlabeled data, from unlabeled to labeled data, and the transfer between data of their respective types. Figure 10 The results of the ablation experiments are shown, Figure 5 , Figure 6 showing the segmentation result diagrams, which prove that the data augmentation part of this embodiment can improve the generalization ability of the model, retain learned knowledge to a greater extent, and enable the model to make full use of potential beneficial information.
[0118] In the second ablation study, the impact of the block size during information transfer on the ACDC dataset was studied. For the size α of the central region m of the mask M, it was set to {0.33, 0.5, 0.66, 0.75}. The effect is best when α is equal to 0.5. Either being too large or too small will affect the effect, which means that an inappropriate block size either cannot learn semantic information from the graph and is difficult to transfer its common semantics, or blocks the original image, losing too much information of the current image.
[0119] In the third ablation study, the weight β in the loss function was default set to 0.5, and then the results after changing it were tested. After changing β to {0.5, 1.0, 1.5, 2.0}, the performance all decreased, and the decrease was the most when the weight was changed to 2.0.
[0120] In the fourth ablation study, it was desired to know which component in the improved implicit network played the most important role. In this embodiment, the extracted feature vectors were used as latent codes, concatenated with relative coordinates, and then input into the MLP. Ablation experiments were conducted on each part, as Figure 13 shown. The two have complementary effects, and they jointly provide more comprehensive and richer information to support the training of the model. The third row also demonstrates the MLP's ability to understand and capture semantic information in the implicit neural network.
[0121] In the fifth ablation study, since the number of hidden layers in the multi-layer perceptron (MLP) has an important impact on network performance. Reducing the number of hidden layers may lead to a weakening of the network's representation ability, while increasing the number of layers will increase computational resources, training time, and the number of parameters of the network. Therefore, a comprehensive ablation experiment was also conducted to determine the optimal number of layers. On the ACDC dataset, the impact of the number of layers {2, 4, 8, 16, 32} on the network was tested when using 5%, 10%, and 20% of the data. There is Figure 14 shown: The larger the number of layers, the better the results. The best results were obtained when the number of layers was 16, but when it was 32, there was a decline instead. This may be because too many layers are not suitable for this task and may lead to a decrease in the accuracy of the extracted feature maps.
[0122] Reader study results: A double-blind reader study was conducted to evaluate and compare the performance of various methods in semi-supervised medical image segmentation. The study involved two radiology experts with more than 5 years of clinical experience, and a total of 20 cardiac segmentation results were used for evaluation, of which 14 groups were from the ACDC dataset and 6 groups were from an internal dataset. The readers evaluated the segmentation results in each group. Segmentation effect and boundary details were considered in the evaluation criteria. The evaluation of the segmentation effect included the overall image quality and general aspects. The results were evaluated using a five-point scale, with 1 point indicating that the segmentation map was difficult to be successfully distinguished, and 5 points indicating successful segmentation, which could be applied to clinical assistance for doctors' diagnosis.
[0123] Figure 15 The reader study results shown indicate that when the network of this embodiment was applied to the ACDC and internal SAX datasets, the highest reader scores were obtained among all methods. This shows that the method of this embodiment has achieved great results in successful cardiac segmentation and boundary detail processing. Figure 15 A double-blind reader study (mean±SD) comparing the method of this embodiment with other SOTA methods on the ACDC dataset is shown, where "mean" and "SD" represent the mean and standard deviation respectively. The best results are marked in bold.
[0124] The above-mentioned multiple experimental studies show that the model of this embodiment has successfully reduced the distribution gap between the labeled and unlabeled data, and maximally retained the learned knowledge. The input image is transformed into a continuous feature space representation using an implicit neural network, retaining the most feature information. Through a large number of experiments, the effectiveness of each part is demonstrated, and competitive results are achieved on two datasets, ACDC and internal SAX. The network architecture proposed in this embodiment has superiority. The results show that the segmentation network of this embodiment is superior to the state-of-the-art methods and generates high-quality segmentation results.
[0125] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
[0126] Embodiment 2
[0127] The purpose of this embodiment is to provide a cardiac MRI image segmentation system, including:
[0128] An acquisition module, which is configured to: acquire a cardiac MRI image to be processed;
[0129] A segmentation prediction module, which is configured to: input the cardiac MRI image to be processed into a trained teacher network to obtain a segmentation result predicted based on the features continuously aligned by a feature alignment implicit function;
[0130] Wherein, the training process of the teacher network is:
[0131] Acquire a cardiac MRI image training dataset, which includes labeled data and unlabeled data;
[0132] Perform information transfer enhancement on the pseudo-labeled data and the labeled data through a randomly generated zero-centered mask to obtain mixed input data;
[0133] Use the mixed input data to train a student network, and update the network parameters of the teacher network according to the network parameters of the trained student network;
[0134] Generate pseudo-labels for the unlabeled data using the teacher network, mix the labeled data and the pseudo-labels through a randomly generated zero-centered mask to generate a supervision signal, and guide and optimize the training of the student network based on the supervision signal.
[0135] In more embodiments, there is also provided:
[0136] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.
[0137] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0138] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0139] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.
[0140] The method in Embodiment 1 can be directly embodied as being executed and completed by a hardware processor, or executed and completed by a combination of hardware and software modules in the processor. The software modules may be located in mature storage media in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0141] A computer program product includes a computer program. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented and completed.
[0142] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to execute the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules may be combined or divided as needed. The machine-executable instructions for program modules may be executed locally or within a distributed device. In a distributed device, program modules may be located in local and remote storage media.
[0143] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.
[0144] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.
[0145] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0146] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A method for segmenting cardiac MRI images, characterized in that, Including: Obtain a cardiac MRI image to be processed; Input the cardiac MRI image to be processed into a trained teacher network to obtain a segmentation result predicted based on the features continuously aligned by the feature alignment implicit function; Wherein, the training process of the teacher network is: obtain a cardiac MRI image training dataset, and the cardiac MRI image training dataset includes labeled data and unlabeled data; Perform information transfer enhancement on the pseudo-labeled data and the labeled data through a randomly generated zero-centered mask to obtain mixed input data; Use the mixed input data to train a student network, and update the network parameters of the teacher network according to the network parameters of the trained student network; Generate pseudo-labels for the unlabeled data using the teacher network, mix the labeled data and the pseudo-labels through a randomly generated zero-centered mask to generate a supervision signal, and guide and optimize the training of the student network based on the supervision signal.
2. The method for segmenting cardiac MRI images according to claim 1, wherein, The processing process of the student network for the mixed input data is specifically: Use an encoder-decoder to extract features from the mixed input data; Perform feature weight map assignment on the extracted features; wherein, each element of the feature weight map represents the weight of the corresponding position for feature alignment; Use the feature weight map to perform implicit alignment on the extracted features, and generate a segmentation prediction result based on the aligned features.
3. The method for segmenting cardiac MRI images according to claim 2, characterized in that, Using an encoder-decoder to extract features from the mixed input data is specifically: Input the mixed input data into the encoder, and extract features at different scales through a group of cyclic 3×3 convolutions, batch normalization, and Leaky activation functions; Input the features at different scales extracted by the encoder into the decoder, splice the input of the decoder through 1×1 convolution and upsampling followed by skip connection, and then restore the features at different scales through a group of cyclic 3×3 convolutions, batch normalization, and Leaky activation functions to obtain the extracted features.
4. The method for segmenting cardiac MRI images according to claim 2, wherein The process of mapping feature maps with different resolutions to a continuous feature map using the feature alignment implicit function is specifically: Among them, F(x q ) refers to the feature after the query coordinate x q is processed by the feature alignment implicit function; x q is the given query coordinate, is the nearest latent code obtained from the features at each level; is the feature weight map of different levels; is the relative coordinate and the corresponding encoded value; (·) is the element-wise multiplication, and f θ is the decoding function.
5. A method for cardiac MRI image segmentation according to claim 1, characterized in that, Construct a loss function based on the supervision signal to guide and optimize the training of the student network, and the loss function is specifically: L1 = L seg (Q1, Y1) ⊙ M + βL seg (Q1, Y1) ⊙ (1 - M) K2 = L seg (Q2, Y2) ⊙ (1 - M) + βL seg (Q2, Y2) ⊙ M L all = L1 + L2 where Q1 and Q2 are the calculation results of the mixed input data generated by the student network; L seg is a linear combination of the Dice loss and the cross-entropy loss; β is a hyperparameter that controls the contribution of unlabeled data to the loss function; M is a zero-centered mask, and L all is the final loss function.
6. A method for segmenting cardiac MRI images according to any one of claims 1-5, characterized in that, The student network uses stochastic gradient descent for optimization during training; based on the network parameters of the student network, the network parameters of the teacher network are optimized using exponential moving average.
7. A cardiac MRI image segmentation system, characterized in that, Including: An acquisition module, which is configured to: obtain a cardiac MRI image to be processed; A segmentation prediction module, which is configured to: input the cardiac MRI image to be processed into a trained teacher network to obtain a segmentation result predicted based on the features continuously aligned by the feature alignment implicit function; Wherein, the training process of the teacher network is: obtain a cardiac MRI image training dataset, and the cardiac MRI image training dataset includes labeled data and unlabeled data; Perform information transfer enhancement on the pseudo-labeled data and the labeled data through a randomly generated zero-centered mask to obtain mixed input data; Use the mixed input data to train a student network, and update the network parameters of the teacher network according to the network parameters of the trained student network; Generate the pseudo-labels of the unlabeled data by using the teacher network, mix the labeled data and the pseudo-labels through a randomly generated zero-centered mask to generate a supervision signal, and guide and optimize the training of the student network based on the supervision signal.
8. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in any one of claims 1-6 is completed.
9. A computer-readable storage medium, characterized in that, It is used to store computer instructions. When the computer instructions are executed by the processor, the method described in any one of claims 1-6 is completed.
10. A computer program product, characterized in that, It includes a computer program. When the computer program is executed by the processor, the method described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Semi-supervised medical image segmentation method based on hybrid-decoupling training
CN116935054A
Semi-supervised medical image segmentation method, system, equipment and medium
CN117095014A
Semi-supervised MRI image segmentation method
CN118154880A
Heart image segmentation method and device, medium and equipment
CN118587233A
Medical image segmentation multi-model aggregation method based on learnable token
CN119206228A