MRI-TRUS Image Registration Method and Device Based on Weakly Supervised Learning
Through the segmentation sub-network and registration sub-network based on weakly supervised learning, combined with dense transform field and self-attention mechanism, the fast and accurate time registration of MRI and TRUS images is achieved, solving the problems of manual labeling dependence and low registration efficiency in the prior art.
Patent Information
- Application Number
- CN202510146999.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The prior art relies on manual labeling data during real-time registration of MRI and TRUS images, resulting in high time and cost, and it is difficult to quickly and accurately complete registration in a high-load clinical environment.
The MRI-TRUS image registration method based on weakly supervised learning is adopted to obtain the mask to be registered through the segmented subnet, and the dense transform field and self-attention mechanism are used for precise registration in the registration subnet to reduce the dependence on the labeled data.
It realizes fast and accurate registration of MRI and TRUS images, reduces the dependence of manual annotation, improves the stability and efficiency of registration, and reduces the waste of medical resources.
Smart Images

Figure CN119625038B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of medical image processing, and in particular, to an MRI-TRUS image registration method and device based on weakly supervised learning. Background Art
[0002] Prostate cancer is one of the most common malignant tumors in men worldwide. Its early and accurate diagnosis is crucial for reducing mortality and improving the prognosis of patients. Currently, biopsy remains the gold standard for confirming prostate cancer, and transrectal ultrasound (TRUS) provides real-time image guidance during the biopsy process. However, due to the limitations of TRUS images in tissue resolution, it is difficult to accurately distinguish normal and cancerous tissues, resulting in limited diagnostic accuracy. Multi-parametric magnetic resonance imaging (mpMRI) has been widely used in the detection and localization of prostate cancer because of its superior soft tissue contrast. However, MRI is mainly used for preoperative image evaluation and cannot be applied in real-time during the operation, which limits its role in biopsy guidance.
[0003] Therefore, in order to better distinguish normal and cancerous tissues accurately during the operation, MRI and TRUS are often registered. Currently, the registration of MRI and TRUS is usually performed manually, which requires a large amount of time and effort and is prone to waste of medical resources, especially in a high-load clinical environment. If automatic registration is performed through an artificial intelligence model, a large amount of labeled data is required, and these labeled data often also need to be manually labeled by doctors, which is time-consuming and costly.
[0004] In summary, how to reduce the dependence on labeled data and thus quickly and accurately register MRI images and TRUS images in real-time is an urgent problem to be solved in the prior art. Summary of the Invention
[0005] The embodiments of the present application provide an MRI-TRUS image registration method and device based on weakly supervised learning. The method obtains a to-be-registered MRI mask and a to-be-registered TRUS mask through a segmentation sub-network, and then performs precise registration in a registration sub-network guided by the to-be-registered MRI mask and the to-be-registered TRUS mask.
[0006] In a first aspect, the embodiments of the present application provide an MRI-TRUS image registration method based on weakly supervised learning. The method includes:
[0007] Obtain a to-be-registered MRI image and a to-be-registered TRUS image, and input the to-be-registered MRI image and the to-be-registered TRUS image into a pre-trained segmentation sub-network to obtain a to-be-registered MRI mask and a to-be-registered TRUS mask;
[0008] Input the to-be-registered MRI mask, to-be-registered TRUS mask, to-be-registered MRI image, and to-be-registered TRUS image into a pre-trained registration sub-network to obtain a registration result. Among them, the registration sub-network consists of a transformation field unit and a registration unit. The transformation field unit obtains a dense transformation field based on the to-be-registered MRI mask, to-be-registered TRUS mask, to-be-registered MRI image, and to-be-registered TRUS image. In the registration unit, use the to-be-registered MRI mask as the moving label, the to-be-registered MRI image as the moving image, the to-be-registered TRUS mask as the fixed label, and the to-be-registered TRUS image as the fixed image, and deform the moving label and the moving image based on the dense transformation field to register the moving label and the moving image into the fixed label and the fixed image.
[0009] In a second aspect, an embodiment of the present application provides an MRI-TRUS image registration device based on weakly supervised learning, including:
[0010] An acquisition module, configured to acquire a to-be-registered MRI image and a to-be-registered TRUS image, and input the to-be-registered MRI image and the to-be-registered TRUS image into a pre-trained segmentation sub-network to obtain a to-be-registered MRI mask and a to-be-registered TRUS mask;
[0011] A registration module, configured to input the to-be-registered MRI mask, to-be-registered TRUS mask, to-be-registered MRI image, and to-be-registered TRUS image into a pre-trained registration sub-network to obtain a registration result. Among them, the registration sub-network consists of a transformation field unit and a registration unit. The transformation field unit obtains a dense transformation field based on the to-be-registered MRI mask, to-be-registered TRUS mask, to-be-registered MRI image, and to-be-registered TRUS image. In the registration unit, use the to-be-registered MRI mask as the moving label, the to-be-registered MRI image as the moving image, the to-be-registered TRUS mask as the fixed label, and the to-be-registered TRUS image as the fixed image, and deform the moving label and the moving image based on the dense transformation field to register the moving label and the moving image into the fixed label and the fixed image.
[0012] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute an MRI-TRUS image registration method based on weakly supervised learning.
[0013] In a fourth aspect, an embodiment of the present application provides a readable storage medium. A computer program is stored in the readable storage medium. The computer program includes program codes for controlling a process to execute the process, and the process includes an MRI-TRUS image registration method based on weakly supervised learning.
[0014] The main contributions and innovations of the present invention are as follows:
[0015] In the embodiment of the present application, the sub-network is segmented to obtain the MRI mask to be registered and the TRUS mask to be registered. In the segmentation sub-network, the SE module is used to adaptively adjust the key channel features of the feature map, suppress unimportant channels, and enhance the attention to the prostate ROI; the self-attention mechanism is used to capture the long-range dependencies of the image, reduce noise interference, improve the segmentation accuracy of fine regions, and fuse the feature information of different levels through skip connections to avoid feature loss; the transformation field unit in the registration sub-network is connected in series by a downsampling module and an upsampling module, and the number of downsampling layers is equal to that of the upsampling layers and skip connections are made. Residual connections are used for downsampling and trilinear interpolation is used for upsampling to accurately capture the key features at different scales and generate a dense transformation field. Through the downsampling layer and the upsampling layer, the key features at different scales can be accurately captured under the guidance of the MRI mask to be registered and the TRUS mask to be registered, thereby improving the registration accuracy, helping the registration sub-network to focus on the region of interest, and further improving the accuracy and efficiency of the registration.
[0016] The details of one or more embodiments of the present application are set forth in the following drawings and description, so as to make the other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0018] Figure 1 is a flowchart of a method for MRI-TRUS image registration based on weakly supervised learning according to an embodiment of the present application;
[0019] Figure 2 is a schematic structural diagram of a segmentation sub-network according to an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of the performance results of a pre-trained segmentation sub-network on a clinical dataset according to an embodiment of the present application;
[0021] Figure 4 is a schematic structural diagram of a transformation field unit according to an embodiment of the present application;
[0022] Figure 5 is an overall structural diagram of a registration sub-network according to an embodiment of the present application;
[0023] Figure 6Schematic diagram of the clinical performance results of a registration sub-network according to an embodiment of the present application;
[0024] Figure 7 Block diagram of the structure of an MRI-TRUS image registration device based on weak supervision learning according to an embodiment of the present application;
[0025] Figure 8 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0026] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0027] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0028] Embodiment 1
[0029] The embodiment of the present application provides an MRI-TRUS image registration method based on weak supervision learning. The method obtains a to-be-registered MRI mask and a to-be-registered TRUS mask through a segmentation sub-network, and then performs precise registration in the registration sub-network guided by the to-be-registered MRI mask and the to-be-registered TRUS mask. Specifically, referring to Figure 1 , the method includes:
[0030] Obtain a to-be-registered MRI image and a to-be-registered TRUS image, and input the to-be-registered MRI image and the to-be-registered TRUS image into a pre-trained segmentation sub-network to obtain a to-be-registered MRI mask and a to-be-registered TRUS mask;
[0031] Input the to-be-registered MRI mask, to-be-registered TRUS mask, to-be-registered MRI image, and to-be-registered TRUS image into the pre-trained registration sub-network to obtain a registration result. Among them, the registration sub-network consists of a transformation field unit and a registration unit. The transformation field unit obtains a dense transformation field based on the to-be-registered MRI mask, to-be-registered TRUS mask, to-be-registered MRI image, and to-be-registered TRUS image. In the registration unit, use the to-be-registered MRI mask as the moving label, the to-be-registered MRI image as the moving image, the to-be-registered TRUS mask as the fixed label, and the to-be-registered TRUS image as the fixed image, and deform the moving label and the moving image based on the dense transformation field to register the moving label and the moving image into the fixed label and the fixed image.
[0032] In some specific embodiments, obtain multiple MRI images and paired TRUS images from a public dataset and various clinical image data collected from a hospital as the training dataset. First, preprocess the training dataset, and then use the preprocessed training dataset to train the segmentation sub-network to obtain a pre-trained segmentation sub-network.
[0033] Specifically, in this solution, first, prostate image data is collected from a public dataset, including 48 MRI images of the MSDProstate dataset, 50 MRI images of the Promise12 dataset, and 73 pairs of MRI-TRUS paired images of the µ-RegPro dataset. Also, 82 pairs of MRI-TRUS clinical image data are collected from a certain hospital to ensure the reliability of the model in practical applications. When screening the clinical data, 50 cases of non-compliant data are excluded. The reasons for excluding these data include the lack of T2 sequence, artifacts in the MRI or TRUS images, and the missing or mismatched segmentation labels drawn by doctors. Finally, 32 pairs of high-quality clinical paired images are retained for subsequent experiments. In terms of the division of the dataset, first, the training data for the automatic segmentation model of the prostate region includes 136 MRI images and 65 TRUS images selected from the public dataset, and 35 MRI images and 8 TRUS images are used for testing. Then, in the training and validation of the MRI-TRUS registration model, 65 pairs of paired images are selected from the µ-RegPro dataset for training, and 8 pairs of paired images and 32 pairs of clinical images are used for testing. Finally, in the experiment to evaluate the improvement effect of MRI-TRUS registration on clinical diagnosis, 32 pairs of registered clinical images are used, and the diagnostic accuracy and confidence of doctors are compared and analyzed through the PI-RADS scoring system.
[0034] Specifically, the training dataset is preprocessed by means of format conversion, size cropping, normalization, and data augmentation. To reduce redundant information and improve the training efficiency of the segmentation sub-network, the top and bottom 10% slices of the MRI images and the first 25% and last 12.5% frames of the TRUS images are cropped. To enhance the generalization ability of the segmentation sub-network, data augmentation is applied by means of rotation, balancing, scaling, etc.
[0035] In some embodiments, the structure of the segmentation sub-network is as Figure 2 shown. The segmentation sub-network includes an encoding region and a decoding region with a symmetric structure. The encoding region is formed by cascading a first encoding unit and multiple second encoding units. The decoding region is formed by cascading multiple first decoding units and a second decoding unit. Moreover, each second encoding unit is skip-connected to the first decoding unit of the corresponding scale size.
[0036] Specifically, the segmentation sub-network is trained using the training samples in the training dataset. The training samples are MRI images or TRUS images labeled with the prostate region, and a segmentation loss is constructed. When the segmentation loss meets the set conditions or reaches the number of iteration rounds, the training is completed to obtain a pre-trained segmentation sub-network.
[0037] Specifically, in this solution, the segmentation sub-network enhances the model's ability to capture long-range dependent features through the SE module and the self-attention mechanism, and fuses feature information at different levels through skip connections, combining shallow detail features with deep semantic features, avoiding the problem of feature loss caused by the increase in network depth, helping to improve the segmentation accuracy, making the segmentation results more accurate in terms of details such as edges and the distinction of overall objects, so as to achieve efficient prostate ROI extraction for guiding subsequent image registration.
[0038] Furthermore, the first encoding unit is a separate encoder, the second encoding unit is an encoder cascaded with an SE module, the first decoding unit is a self-attention module cascaded with a decoder, and the second decoding unit is a separate decoder.
[0039] Specifically, compared with manual delineation by doctors, obtaining the MRI mask to be registered and the TRUS mask to be registered through the pre-trained segmentation sub-network has stronger robustness and consistency, reduces the time and cost of manual annotation, and avoids problems of human error and inconsistent annotation, thereby improving the stability and accuracy of subsequent image registration.
[0040] Specifically, the SE module extracts the global information of each channel through global average pooling operation and generates adaptive weights for each channel, thereby weighting important channels and enabling the segmentation sub-network to pay more attention to the prostate ROI during the learning process, avoiding interference from irrelevant features.
[0041] The SE module first performs global average pooling on each channel of the output feature map through a squeezing operation to obtain the global descriptor of each channel , representing the global information of the corresponding channel, and obtaining The formula for is as follows:
[0042]
[0043] where, is the eigenvalue of channel at position in the input feature map, and are the height and width of the feature map respectively. Then, in the Excitation operation, the SE module generates the attention weight for each channel, and this weight controls the contribution of each channel to the final feature. The formula for generating is as follows:
[0044]
[0045] where, and are the weight matrices of the fully connected layers, is the ReLU activation function, is the Sigmoid activation function, which is used to ensure that the weight of each channel is between 0 and 1. Finally, after obtaining the weight of each channel, the SE module applies these weights to the features of each channel in the original feature map through a recalibration operation to adjust the importance of the channels. The specific operation is as follows:
[0046]
[0047] where, is the channel before adjustment, is the importance of each channel, is the channel weight after adjusting the importance.
[0048] This solution adaptively adjusts the features of the key channels in the feature map through the SE module, while suppressing unimportant channels, thereby enhancing the feature expression ability. The SE module not only enhances the robustness of the network when processing MRI images and TRUS images, but also effectively improves the segmentation accuracy of complex regions such as the prostate and tumors, providing more accurate segmentation labels for subsequent registration tasks, thus ensuring the accuracy of the registration process.
[0049] Specifically, the self-attention mechanism in this solution enables the segmentation sub-network to capture long-range dependencies in MRI images and TRUS images. Especially when dealing with regions with blurred boundaries or complex backgrounds, it can dynamically adjust the attention area of the segmentation sub-network. Through the self-attention mechanism, the model can effectively reduce noise interference during the segmentation process and improve the segmentation accuracy of fine regions such as tumors and the prostate. In the self-attention mechanism, first, for the input feature map the queries, keys, and values are calculated, where N is the number of features and d is the dimension of each feature. The calculation formulas for the queries, keys, and values are as follows:
[0050]
[0051] where, are weight matrices, which are used to generate the query Q, key K, and value V respectively.
[0052] Then, the attention weights are obtained by calculating the similarity between the query and the key. That is to say, the dot product is used to calculate the inner product of the query and the key, and the weight distribution is obtained through the Softmax function:
[0053]
[0054] Finally, the obtained attention weights are used to perform a weighted average on V to obtain the weighted output:
[0055]
[0056] Through the self-attention mechanism, the segmentation sub-network can adaptively allocate attention according to the relationship between different input features, focusing on regions with strong correlation, that is, the difference between the prostate and tumors, while ignoring irrelevant parts. In practical applications, the self-attention mechanism can dynamically adjust the weights of features in different parts to effectively capture long-range dependencies and global context information. For example, when processing MRI images, it can automatically adjust the attention according to the lesion regions in the image to ensure that information on key regions such as the prostate and tumors receives more attention. When dealing with images with blurred boundaries or complex details, it can improve the recognition ability of these regions, thus better completing image segmentation.
[0057] Exemplarily, the performance results of the pre-trained segmentation sub-network in the clinical dataset are as follows Figure 3 shown Figure 3 In Figure 3 , 1-4 represent four cases of MRI image data, 5-8 represent four cases of TRUS image data. (a) is the original image input into the segmentation sub-network, (b) is the segmentation mask outlined by professional radiologists, and (c) is the segmentation mask automatically output by the segmentation sub-model
[0058] The pre-trained segmentation sub-network was tested using the test set, and the test results are shown in Table 1
[0059] Table 1 Test results of the segmentation sub-network
[0060]
[0061] Table 1 shows the performance evaluation results of the segmentation sub-network on different datasets, demonstrating the stable performance of the model in each dataset. In the public MRI dataset, the DSC of the model reached 0.9154, and the accuracy (Acc) was 0.9914, indicating its high precision in the MRI image segmentation of prostate structures. Although the performance on the clinical dataset decreased slightly, with DSC being 0.9010 and Acc being 0.9753, the segmentation performance of the model in actual clinical applications was still excellent. In the segmentation task of TRUS images, the model also performed well. The DSC in the public dataset was 0.9384, and the Acc was 0.9828, fully demonstrating its segmentation ability in TRUS images. In the clinical dataset, although the accuracy decreased, with DSC being 0.9173 and Acc being 0.9522, the model still showed strong adaptability
[0062] In terms of the HD95 index for boundary segmentation, no significant statistical difference was found between the public MRI dataset and the clinical MRI dataset (public MRI vs. clinical MRI, p = ns). However, in TRUS images, the performance of the public dataset was significantly better than that of the clinical dataset (public TRUS vs. clinical TRUS, p = 0.007). Similarly, in terms of the HD95 index, there was no significant difference between the MRI images and TRUS images in the public dataset (public MRI vs. public TRUS, p = ns), but in the clinical dataset, the performance of MRI images was significantly better than that of TRUS images (clinical MRI vs. clinical TRUS, p<0.001). Overall, the model demonstrated good robustness and cross-scene adaptability under different datasets and imaging modalities
[0063] In some embodiments, the pre-trained segmentation sub-network output MRI mask and TRUS mask are used as guiding information to combine the corresponding MRI images and TRUS images in the training dataset to perform weakly supervised training on the registration sub-network.
[0064] Specifically, due to the small amount of registration training sample data for MRI images and TRUS images, and the individual differences between different patients, this solution uses a weakly supervised training method to train the registration sub-network. Weak supervision can reduce the requirement for annotation accuracy. During registration, only relatively less annotation information is needed, saving time and labor costs. At the same time, it can also improve the generalization ability of the registration sub-network, so that the registration sub-network does not overly rely on the precise annotation mode, and can learn more general registration features and rules. When facing individual differences between different patients and differences in device imaging parameters, it is easier to find a suitable registration strategy.
[0065] In some embodiments, the structural schematic diagram of the transformation field unit is as Figure 4 shown. The transformation field unit is connected in series by a downsampling module and an upsampling module. The downsampling module includes multiple downsampling layers, and the upsampling module includes multiple upsampling layers, and the number of downsampling layers is equal to the number of upsampling layers.
[0066] Specifically, the number of downsampling layers and upsampling layers in this solution is four. In the registration network, through the downsampling layer and upsampling layer, key features at different scales can be accurately captured under the guidance of the MRI mask to be registered and the TRUS mask to be registered, thereby improving the registration accuracy, helping the registration sub-network to focus on the region of interest, and further improving the accuracy and efficiency of registration.
[0067] Furthermore, each downsampling layer is skip-connected to the upsampling layer of the corresponding scale size, and the downsampling layer uses a residual connection method for downsampling, and the upsampling layer uses a trilinear interpolation method for upsampling.
[0068] Specifically, the skip connection can effectively transfer the detailed information lost during the downsampling process to the upsampling stage, enabling the upsampling to better restore the details of the image using this information, and helping to retain the key features in the original image to avoid feature loss or blurring after multiple downsamplings and upsamplings.
[0069] Specifically, the downsampling layer uses a residual connection for downsampling. The residual connection can effectively alleviate the problem of gradient disappearance or gradient explosion caused by downsampling. By adding the input information and the downsampled information, it makes it easier for the network to learn the residual part after downsampling, improving the training stability and convergence speed of the network.
[0070] Specifically, the upsampling layer uses trilinear interpolation for upsampling to reasonably estimate the values of newly inserted pixels based on the values of surrounding pixels. While ensuring computational efficiency, it can better magnify the image, so as to combine with the detailed information transmitted by the skip connection and more accurately restore the features and structure of the original image.
[0071] In some specific embodiments, the overall structure of the registration sub-network is as Figure 5 shown. In the registration sub-network, the dense transformation field output by the transformation field unit is the transformation matrix for transforming the moving image and the moving label to the fixed image and the fixed label in the image space.
[0072] Specifically, in the registration sub-network, the to-be-registered MRI mask, the to-be-registered TRUS mask are combined with the corresponding image itself to better capture the key features in the image, thereby improving the registration accuracy. That is to say, in the registration sub-network, the to-be-registered MRI mask and the to-be-registered TRUS mask are used as guiding information to help the registration sub-network focus on the region of interest, further improving the accuracy and efficiency of registration.
[0073] Furthermore, the dense transformation field in this solution is used to optimize in multiple scales, effectively handle the anatomical differences in multi-modal images, and ensure the accurate registration of MRI images and TRUS images.
[0074] Specifically, the dense transformation field is used to deform the moving label and the moving image so that they are aligned with the corresponding fixed label and fixed image to complete the registration.
[0075] In some embodiments, when weakly training the registration sub-network, the segmentation consistency loss, the boundary error loss, and the target registration error loss are used as the loss functions. Among them, the segmentation consistency loss is used to evaluate the overlap degree between the deformed moving label and the fixed label, the boundary error loss is used to measure the maximum deviation of the boundary points between the deformed moving label and the fixed label, and the target registration error loss is used to measure the spatial alignment error of the key points between the moving image and the fixed image.
[0076] Specifically, the formula of the segmentation consistency loss is expressed as follows:
[0077]
[0078] where A is the deformed moving label, B is the fixed label, is the segmentation consistency loss, and the segmentation consistency loss is used to evaluate the overlap degree between the deformed moving label and the fixed label.
[0079] Specifically, the formula of the boundary error loss is expressed as follows:
[0080]
[0081] Among them, A is the deformed moving label, and B is the fixed label is the boundary error loss, and HD95 represents the 95th percentile of the Hausdorff distance. The boundary error loss is used to measure the maximum deviation of the boundary points between the deformed moving label and the fixed label
[0082] Specifically, the formula of the target registration error loss is expressed as follows
[0083]
[0084] Among them, N is the total number of key points represents the key points in the moving image represents the key points in the fixed image represents the target registration error loss, and the target registration error loss is used to measure the spatial alignment error of the key points between the moving image and the fixed image
[0085] Furthermore, in this solution, the segmentation consistency loss, the boundary error loss, and the target registration error loss are weighted and summed to obtain the multi-scale loss, and the parameters of the registration sub-network are adjusted based on the multi-scale loss
[0086] Specifically, the formula of the multi-scale loss is expressed as
[0087]
[0088] Among them is the multi-scale loss is the weight of the segmentation consistency loss is the weight of the boundary error loss is the weight of the target registration error loss
[0089] Exemplarily, in this solution is 0.5 is 0.25 is 0.25
[0090] Specifically, the registration sub-network in this solution uses the to-be-registered MRI mask and the to-be-registered TRUS mask as segmentation labels to provide more structural information, enabling the registration sub-network to better focus on and align the regions of interest such as the prostate or tumor, improving the registration accuracy. Especially in the registration tasks of complex anatomical structures or lesion regions, the possibility of misregistration can be significantly reduced
[0091] Specifically, Table 2 shows the evaluation indicators of the registration sub-model in the test dataset
[0092] Table 2 Evaluation Metrics of the Registration Sub-Model in the Test Dataset
[0093]
[0094] There are six key metrics involved in Table 2: DSC (Dice coefficient), HD95 (95% Hausdorff distance), Acc (accuracy), Pre (precision), Rec (recall), and TRE (target registration error). On the public dataset, the DSC of the model is 0.7481, HD95 is 10.1839, Acc is 0.9313, Pre is 0.7444, Rec is 0.7583, and TRE is 7.8562, showing that the model has a relatively high registration accuracy on this dataset, especially performing excellently in terms of the DSC and Acc metrics, fully demonstrating its ability to accurately segment and register the prostate region. In contrast, on the clinical dataset, the DSC of the model is 0.7295, HD95 is 11.1766, Acc is 0.8760, Pre is 0.8813, Rec is 0.6412, and TRE is 8.6307. Although there is no statistically significant difference between the public dataset and the clinical dataset in terms of these three key metrics of DSC, HD95, and TRE, in terms of the Pre metric, the clinical dataset is significantly better than the public dataset (p < 0.001), reflecting the robustness and higher registration accuracy of the model in the clinical environment. These results further verify the good generalization ability and adaptability of the model in complex clinical scenarios, indicating that it can continuously provide a high level of segmentation and registration accuracy in actual clinical applications.
[0095] Exemplarily, the performance results of the registration sub-network in clinical settings are as Figure 6 shown. In Figure 6 , (a) is the original MRI image, (b) is the original TRUS image, (c) is the standardized MRI image, and 1, 2 represent two different cases.
[0096] Example 2
[0097] Based on the same concept, referring to Figure 7 , this application also proposes an MRI-TRUS image registration device based on weakly supervised learning, including:
[0098] An acquisition module, configured to acquire the MRI image to be registered and the TRUS image to be registered, and input the MRI image to be registered and the TRUS image to be registered into a pre-trained segmentation sub-network to obtain the MRI mask to be registered and the TRUS mask to be registered;
[0099] A registration module is configured to input the MRI mask to be registered, the TRUS mask to be registered, the MRI image to be registered, and the TRUS image to be registered into a pre-trained registration sub-network to obtain a registration result. The registration sub-network consists of a transformation field unit and a registration unit. The transformation field unit obtains a dense transformation field based on the MRI mask to be registered, the TRUS mask to be registered, the MRI image to be registered, and the TRUS image to be registered. In the registration unit, the MRI mask to be registered is used as the moving label, the MRI image to be registered is used as the moving image, the TRUS mask to be registered is used as the fixed label, and the TRUS image to be registered is used as the fixed image. The moving label and the moving image are deformed based on the dense transformation field so that the moving label and the moving image are registered into the fixed label and the fixed image.
[0100] Embodiment III
[0101] This embodiment also provides an electronic device. Refer to Figure 8 , which includes a memory 404 and a processor 402. A computer program is stored in the memory 404, and the processor 402 is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0102] Specifically, the above processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0103] Among them, the memory 404 may include a mass storage 404 for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In suitable cases, the memory 404 may include removable or non-removable (or fixed) media. In suitable cases, the memory 404 may be internal or external to the data processing device. In a particular embodiment, the memory 404 is a non-volatile memory. In a particular embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). In suitable cases, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. In suitable cases, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data out dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0104] The memory 404 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.
[0105] By reading and executing the computer program instructions stored in the memory 404, the processor 402 implements any one of the MRI-TRUS image registration methods based on weakly supervised learning in the above embodiments.
[0106] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.
[0107] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above network may include wired or wireless networks provided by a communication provider of the electronic device. In one example, the transmission device includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0108] The input / output device 408 is used to input or output information. In this embodiment, the input information may be the MRI image to be registered and the TRUS image to be registered, etc., and the output information may be the registration result, etc.
[0109] Optionally, in this embodiment, the above processor 402 may be configured to execute the following steps by a computer program:
[0110] Obtain the MRI image to be registered and the TRUS image to be registered, and input the MRI image to be registered and the TRUS image to be registered into the pre-trained segmentation sub-network to obtain the MRI mask to be registered and the TRUS mask to be registered;
[0111] Input the to-be-registered MRI mask, to-be-registered TRUS mask, to-be-registered MRI image, and to-be-registered TRUS image into the pre-trained registration sub-network to obtain a registration result. Among them, the registration sub-network consists of a transformation field unit and a registration unit. The transformation field unit obtains a dense transformation field based on the to-be-registered MRI mask, to-be-registered TRUS mask, to-be-registered MRI image, and to-be-registered TRUS image. In the registration unit, use the to-be-registered MRI mask as the moving label, the to-be-registered MRI image as the moving image, the to-be-registered TRUS mask as the fixed label, and the to-be-registered TRUS image as the fixed image, and deform the moving label and the moving image based on the dense transformation field to register the moving label and the moving image into the fixed label and the fixed image.
[0112] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.
[0113] Generally, various embodiments can be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although the various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, as a non-limiting example, the blocks, devices, systems, technologies, or methods described herein can be implemented in hardware, software, firmware, dedicated circuits or logic, general hardware or a controller or other computing devices, or some combination thereof.
[0114] Embodiments of the present invention can be implemented by computer software, which can be executed by a data processor of a mobile device, such as in a processor entity, or can be implemented by hardware, or can be implemented by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components configured to execute the embodiments when the program runs. The one or more computer-executable components can be at least one software code or a part thereof. Additionally, in this regard, it should be noted that any box in the logical flow, as Figure 8 shown, can represent a program step, or interconnected logical circuits, boxes, and functions, or a combination of program steps and logical circuits, boxes, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media is a non-transitory medium.
[0115] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0116] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A MRI-TRUS image registration method based on weakly supervised learning, characterized in that: The following steps are involved: Acquire an MRI image to be registered and a TRUS image to be registered, input the MRI image to be registered and the TRUS image to be registered into a pre-trained segmentation subnetwork to obtain an MRI mask to be registered and a TRUS mask to be registered, wherein the segmentation subnetwork includes an encoding region and a decoding region of a symmetrical structure, wherein the encoding region is connected in series by a first encoding unit and a plurality of second encoding units, wherein the decoding region is connected in series by a plurality of first decoding units and a second decoding unit, and each second encoding unit is jump-connected to a first decoding unit of a corresponding scale, wherein the first encoding unit is a separate encoder, the second encoding unit is a series-connected encoder and SE module, the first decoding unit is a series-connected self-attention module and decoder, and the second decoding unit is a separate decoder; The MRI mask to be registered, the TRUS mask to be registered, the MRI image to be registered, and the TRUS image to be registered are input into a pre-trained registration sub-network to obtain a registration result, wherein the MRI mask and the TRUS mask output by the pre-trained segmentation sub-network are used as guiding information to combine the corresponding MRI image and TRUS image in the training data set to perform weak supervision training on the registration sub-network. When the registration sub-network is weakly trained, the segmentation consistency loss, the boundary error loss, and the target registration error loss are used as loss functions, wherein the segmentation consistency loss is used to evaluate the degree of overlap between the deformed moving label and the fixed label, the boundary error loss is used to measure the maximum deviation of the boundary point between the deformed moving label and the fixed label, and the target registration error loss is used to measure The key point is the spatial alignment error between the moving image and the fixed image. The registration subnetwork is composed of a transformation field unit and a registration unit. The transformation field unit obtains a dense transformation field based on the MRI mask to be registered, the TRUS mask to be registered, the MRI image to be registered and the TRUS image to be registered. The dense transformation field is a transformation matrix that transforms the moving image and the mobile label to the fixed image and the fixed label in the image space. In the registration unit, the MRI mask to be registered is used as the moving label, the MRI image to be registered is used as the moving image, the TRUS mask to be registered is used as the fixed label, and the TRUS image to be registered is used as the fixed image. The mobile label and the moving image are deformed based on the dense transformation field so that the mobile label and the moving image are registered to the fixed label and the fixed image.
2. The MRI-TRUS image registration method based on weakly supervised learning according to claim 1, characterized in that: The segmentation subnetwork is trained using training samples in the training data set, where the training samples are MRI images or TRUS images with prostate regions labeled, and segmentation loss is constructed. When the segmentation loss meets the set conditions or reaches the iteration round, the training is completed to obtain a pre-trained segmentation subnetwork.
3. The MRI-TRUS image registration method based on weakly supervised learning according to claim 1, characterized in that: The transform field unit is connected in series by a down-sampling module and an up-sampling module, wherein the down-sampling module includes a plurality of down-sampling layers, and the up-sampling module includes a plurality of up-sampling layers, and the number of the down-sampling layers is equal to the number of the up-sampling layers.
4. The MRI-TRUS image registration method based on weakly supervised learning according to claim 3, characterized in that: Each of the down-sampling layers is jump-connected to an up-sampling layer of a corresponding scale, and the down-sampling layer is down-sampled by a residual connection, and the up-sampling layer is up-sampled by a trilinear interpolation.
5. An MRI-TRUS image registration device based on weakly supervised learning, characterized in that: include: An acquisition module is used to acquire an MRI image to be registered and a TRUS image to be registered, and input the MRI image to be registered and the TRUS image to be registered into a pre-trained segmentation subnetwork to obtain an MRI mask to be registered and a TRUS mask to be registered, wherein the segmentation subnetwork includes an encoding region and a decoding region of a symmetrical structure, wherein the encoding region is connected in series by a first encoding unit and a plurality of second encoding units, and the decoding region is connected in series by a plurality of first decoding units and a second decoding unit, and each second encoding unit is jump-connected to a first decoding unit of a corresponding scale, wherein the first encoding unit is a separate encoder, the second encoding unit is a series-connected encoder and SE module, the first decoding unit is a series-connected self-attention module and decoder, and the second decoding unit is a separate decoder; A registration module is used to input the MRI mask to be registered, the TRUS mask to be registered, the MRI image to be registered, and the TRUS image to be registered into a pre-trained registration sub-network to obtain a registration result, wherein the MRI mask and the TRUS mask output by the pre-trained segmentation sub-network are used as guiding information to combine the corresponding MRI image and TRUS image in the training data set to perform weak supervision training on the registration sub-network. When the registration sub-network is weakly trained, the segmentation consistency loss, the boundary error loss, and the target registration error loss are used as loss functions, wherein the segmentation consistency loss is used to evaluate the overlap degree of the deformed moving label and the fixed label, the boundary error loss is used to measure the maximum deviation of the boundary point between the deformed moving label and the fixed label, and the target registration error loss is used to measure the maximum deviation of the boundary point between the deformed moving label and the fixed label. Used to measure the spatial alignment error of key points between a moving image and a fixed image, the registration subnetwork consists of a transformation field unit and a registration unit, the transformation field unit obtains a dense transformation field based on an MRI mask to be registered, a TRUS mask to be registered, an MRI image to be registered, and a TRUS image to be registered, the dense transformation field is a transformation matrix for transforming a moving image and a mobile label to a fixed image and a fixed label in the image space, in the registration unit, the MRI mask to be registered is used as a moving label, the MRI image to be registered is used as a moving image, the TRUS mask to be registered is used as a fixed label, and the TRUS image to be registered is used as a fixed image, and the mobile label and the moving image are deformed based on the dense transformation field so that the mobile label and the moving image are registered to the fixed label and the fixed image.
6. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to run the computer program to execute the MRI-TRUS image registration method based on weakly supervised learning according to any one of claims 1 to 4.
7. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes a program code for controlling a process to execute a process. When the program code is executed by a processor, the MRI-TRUS image registration method based on weakly supervised learning as described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Image registration method, and training method and device of image registration model
CN115908515A
Multi-task abdominal organ registration method based on mutual attention and semantic sharing
CN117036428A
Lung X-ray image registration method based on weak supervised learning
CN117593180A