Multi-modal prostate image registration method based on convolutional neural network
Through the multimodal prostate image registration method based on convolutional neural network, the problems of low ultrasound imaging quality and lack of spatial dimension information in the prior art are solved, and high-precision and fast image registration are achieved, which significantly improves the accuracy and efficiency of prostate cancer diagnosis and treatment.
Patent Information
- Application Number
- CN202510141440.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-27
AI Technical Summary
In the diagnosis and treatment of prostate cancer, the ultrasound imaging quality is low and the image is unclear, resulting in unclear lesions contours, affecting the accuracy of puncture positioning. In addition, the two-dimensional image guidance during operation lacks spatial dimension information, making it difficult for the puncture needle to accurately reach the target area.
The multimodal prostate image registration method based on convolutional neural network is adopted. By acquiring preoperative nuclear magnetic images and intraoperative ultrasound images for image preprocessing, the target organs are automatically segmented, and rigid registration is performed. The CNN-Transformer hybrid network model is used for weakly supervised non-rigid registration, achieving accurate registration of multimodal images.
Automatic image registration is realized, which significantly improves registration accuracy and speed, and the registration accuracy can approach 90%, reduces the generation of folded pixels, improves the quality and reliability of image registration, and meets the time efficiency requirements of clinical applications.
Smart Images

Figure CN120047497A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more particularly to a multi-modal prostate image registration method based on a convolutional neural network. Background Art
[0002] The incidence and mortality rates of prostate cancer are relatively high globally, and its proportion among male malignant tumors shows an upward trend. Prostate biopsy, as one of the core means for diagnosing prostate cancer, is widely used in clinical practice. Among them, transrectal ultrasound (TRUS)-guided prostate biopsy is a common operation method. However, this method exposes many problems that cannot be ignored in practical applications:
[0003] First, the quality of ultrasound imaging is low, and the images are not clear, resulting in unclear lesion contours, that is, "unable to see clearly". This makes it difficult for doctors to accurately judge the boundaries and shapes of tumors, affecting the accuracy of subsequent puncture positioning. Second, although preoperative CT / MRI images can clearly show tumors, during the operation, ultrasound guidance is used, and the lesion positions in CT / MRI images can only be combined by the doctor's spatial imagination, that is, "unable to find". Doctors need to integrate and convert the information of different images in their minds, which undoubtedly increases the difficulty and uncertainty of positioning. Third, intraoperative two-dimensional image guidance lacks spatial dimension information, making the puncture needle unable to accurately reach the target area, that is, "unable to puncture accurately". When the puncture needle enters the human tissue, due to the lack of accurate three-dimensional spatial guidance, it is easy to deviate from the target lesion, resulting in puncture failure or accidental puncture of surrounding normal tissues. Eventually, it is difficult to effectively detect malignant tumors, and even if detected, it is difficult to implement precise and effective treatment measures, seriously affecting the diagnosis and treatment effect of prostate cancer and the prognosis of patients.
[0004] Chinese Patent CN110363802B discloses a prostate image registration system and method based on automatic segmentation and pelvic alignment. The prostate CT image sequence and the prostate MRI image sequence are respectively jointly segmented by using a trained prostate multi-modal image segmentation network U-net to obtain a prostate CT image segmentation result image set C and a prostate MRI image segmentation result image set M. Then, two-stage registration is performed on these two image sets C and M to obtain a final registration image set L. Finally, the final registration image set L and the prostate CT image segmentation result image set C are fused and displayed. Although this method reduces the difference in prostate multi-modal images, there are still deficiencies in the registration accuracy. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a multi-modal prostate image registration method based on a convolutional neural network that can automatically achieve registration, has high registration accuracy, and short time.
[0006] The object of the present invention can be achieved by the following technical solutions:
[0007] A multimodal prostate image registration method based on a convolutional neural network, comprising the following steps:
[0008] Obtain preoperative magnetic resonance (MR) images and intraoperative ultrasound (US) images, and perform image preprocessing on the preoperative MR images and intraoperative US images respectively to obtain the preprocessed MR images and US images;
[0009] Perform an automatic segmentation operation on the target organ on the preprocessed MR images and US images to obtain segmentation information;
[0010] According to the segmentation information, pair and rigidly register the preprocessed MR images and US images;
[0011] Apply a CNN-Transformer hybrid network model based on the Unet architecture to perform weakly supervised non-rigid registration on the rigidly registered MR images and US images to achieve multimodal registration.
[0012] Further, the image preprocessing includes target organ region extraction, image enhancement, resampling, and two-dimensional slice amplification.
[0013] The specific operation of the target organ region extraction is as follows:
[0014] Convert the image into a grayscale image, highlight the difference between the target and the background through adaptive binarization, mark a rectangular box for the largest area region of the binary image, and crop the region outside the rectangular box. The cropped region is the target organ region.
[0015] Further, the automatic segmentation operation on the target organ is implemented through a U-Net segmentation network architecture to obtain a target organ segmentation mask in the multimodal mode.
[0016] Further, the pairing is specifically as follows:
[0017] According to the segmentation information, respectively obtain the area, perimeter, and centroid of the prostate region from the preprocessed MR images and US images, perform matching based on the area and perimeter information, and find the image groups with the same position information.
[0018] Further, the rigid registration includes:
[0019] For the image groups with the same position information, taking the centroid as the reference, correct the in-plane position deviation between the MR image and the US image through translation operations, and adjust the image angle difference by rotation transformation to obtain the rigidly registered MR images and US images.
[0020] Further, in the CNN-Transformer hybrid network model based on the Unet architecture, features are extracted by a feature extraction module based on U-Net with equal weights. CNN is used to perform encoding and decoding operations in the shallow layer of the network, and Transformer is applied for decoding and encoding in the deep layer of the network. At the same time, residual blocks are embedded at the input end of each encoder and the output end of the decoder.
[0021] Further, the weak supervision non-rigid registration specifically includes:
[0022] Taking the nuclear magnetic resonance image as the fixed image and the ultrasonic image as the moving image, and synchronously inputting them into the CNN-Transformer hybrid network model based on the Unet architecture to generate a corresponding deformation field;
[0023] Taking the ultrasonic label as the moving label, deforming it according to the generated deformation field, calculating the loss between the deformed label and the nuclear magnetic resonance label, and feeding it back to the network, and continuously iterating until an accurate deformation field is generated;
[0024] Deforming the ultrasonic image according to the finally generated accurate deformation field to obtain a registered image.
[0025] Further, the CNN-Transformer hybrid network model is trained with a hybrid edge loss. The acquisition of the hybrid edge loss specifically includes:
[0026] For the ultrasonic label and the nuclear magnetic resonance label respectively, perform Gaussian blur downsampling, Gaussian blur upsampling and Gaussian blur processing in sequence, subtract the processed label from the original label to obtain an edge optimized by gradient, and calculate the edge loss;
[0027] Combining the edge loss, multi-scale Dice, and bending energy loss in a set ratio to form a hybrid edge loss.
[0028] Further, the method further includes:
[0029] Obtaining control signals for each step through a human-computer interaction interface and realizing the display of the multi-modal registration effect.
[0030] The present invention also provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the multi-modal prostate image registration method based on a convolutional neural network as described above.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] 1. Automatic registration: The present invention does not require manual marking and delineation of images. It only needs to automatically match the intraoperative ultrasound images collected in real time with the preoperative MRI images at the corresponding positions, and input them into the registration model together, then the registration deformation of the intraoperative ultrasound images can be automatically achieved, with fast and efficient operation.
[0033] 2. High registration accuracy: The present invention adopts a method combining rigid registration and weakly supervised non-rigid registration to achieve multi-modal registration. In the weakly supervised non-rigid registration, the unique characteristics of ultrasound and MRI images are deeply considered, and the concept of neural network modularity is integrated to carefully construct the architecture of the proposed network. On the basis of following the main framework of the existing MambaMorph network, the same-weight feature extraction module and the shallow CNN encoder-decoder are cleverly used, and the core principles of the deep decoder and encoder of the network are replaced with the advanced attention mechanism based on Transformer. This innovative design can not only give full play to the inductive bias advantage of CNN, but also utilize the long attention characteristics of Transformer to effectively improve the network performance. The attention mechanism used therein downsamples the key matrix and the value matrix to a low-dimensional space, which not only greatly reduces the consumption of computing resources, but also enables the model to focus on the singular value region of the image, that is, the edge part of the image, and accurately capture key information. In addition, residual blocks Residual Block are cleverly added at the input end of the encoder and the output end of the decoder, which effectively enhances the long-distance memory efficiency of the model. Finally, a hybrid edge loss function customized for weakly supervised learning is introduced to help the model achieve high-precision registration of images based on the label edge information. Using the method of the present invention, the registration accuracy rate can approach 90%, and the generation of folded pixels can be significantly reduced, effectively improving the quality and reliability of image registration.
[0034] 3. Short registration time: During the surgical process, ultrasound images have a high refresh rate to timely respond to various changes occurring in real time during the operation. Therefore, the timeliness of the registration operation faces extremely strict requirements. The present invention innovatively implements a preprocessing process on the original data set, effectively screening out redundant information that is of no help to the registration network, and significantly improving the image quality. At the same time, with the pre-training of the model, only by pre-inputting the preoperative MRI image of the patient and combining it with the intraoperative ultrasound image updated in real time, the registration task can be successfully completed. During the registration process, each image only takes 0.3 seconds to complete non-rigid registration, fully meeting the strict requirements of clinical applications for time efficiency.
[0035] IV. Certain generalization ability: The method of the present invention aims to enhance the registration ability and optimize the registration accuracy of the model. Three datasets (covering one private dataset and two public datasets) are used for training. Through this way of using multiple datasets, the feature information of the images can be effectively enriched, making it sufficient to handle the image differences from different devices, and thus achieving a more ideal registration generalization effect.
[0036] V. Better interactivity and transferability: The present invention is controlled by designing an interactive interface, which is intelligent, fast, simple and clear. The operation instructions are obvious at a glance. While having a certain degree of interactivity, it has good transferability and less restrictions in use. Only one computer is needed to complete feature extraction and achieve preliminary classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.
[0039] As Figure 1 shown, this embodiment provides a multi-modal prostate image registration method based on a convolutional neural network. This method can realize the multi-modal image registration of preoperative magnetic resonance (MR) and intraoperative ultrasound (US) through a convolutional neural network, and specifically includes the following steps:
[0040] Step S101: Obtain the corresponding images of preoperative MR and intraoperative US.
[0041] In order to achieve more accurate registration results for US and MR images and have a certain generalization ability, this embodiment uses three datasets from different sources. These datasets are from clinical cases and two public datasets, covering a total of 7,745 pairs of MR and US images of 704 patients, so as to comprehensively and deeply verify the proposed method. Specifically, the clinical cases include 49 patients and 332 pairs of corresponding images; in the images from The Cancer Imaging Archive dataset, due to the poor imaging quality of some cases, only the maximum prostate sagittal plane section was selected by clinicians when processing their cases. This dataset involves 582 patients and a total of 3,374 images; while the μ-ProReg challenge dataset contains 4,039 images of 73 patients. In these datasets, the MR images all use T2-weighted sequences, and the US images are all collected by transrectal ultrasound equipment.
[0042] Step S102: To accelerate the network training speed, reduce the computational amount, improve the data quality, unify the data format, and meet the requirements of clinical applications, in this embodiment, a preprocessing operation is performed on the prostate dataset based on the contour search method, including extracting the target organ area, image enhancement, resampling, and two-dimensional slice amplification.
[0043] When extracting the target organ area, the color picture is first grayscaled. After adaptive binarization, contour detection is performed on the binary image to determine the area with the largest area and mark it with a rectangular box. The pixels outside the box are set to zero and cropped to reduce invalid information, enabling feature extraction to focus on the lesion and improving efficiency. In the image enhancement step, considering the low quality of ultrasound images, median filtering is used to remove noise, and then adaptive histogram equalization is used to highlight details, strengthen the image information, and weaken the interference of noise on the network. The resampling operation resamples the MRI and ultrasound images to the same voxel specification to ensure the matching of their information. The two-dimensional slice amplification cuts the resampled image along the sagittal plane to meet the conditions of intraoperative real-time two-dimensional image registration, providing strong support for accurate registration.
[0044] Step S103: To achieve the goal of subsequent weakly supervised registration, mask segmentation of the target organ needs to be carried out on the MRI image and the ultrasound image.
[0045] Specifically, the ultrasound image and the MRI image are respectively input into the U-Net segmentation network architecture. With the powerful image segmentation ability of the U-Net network, the segmentation masks of the same target organ in the multi-modal context are accurately obtained, and then the position information and shape information of the target organ are extracted therefrom, providing key data basis and image feature support for the subsequent registration process.
[0046] Step S104: According to the segmentation mask, the preprocessed MRI-ultrasound images are paired, and rigid registration is carried out within the successfully paired groups to shorten the network training time.
[0047] With the help of the segmentation mask, core data information such as the area, perimeter, and centroid of the prostate label on each image can be accurately measured. The area and perimeter information of the MRI image and the ultrasound image are compared and matched with each other to screen out the image groups with the same position information.
[0048] Subsequently, rigid registration operations are performed in these image groups. Taking the determined centroid as the reference point, translation means are used to correct the position offset of the MRI image and the ultrasound image in the plane, and rotation transformation is used to adjust the angle difference of the images, so that the images are initially consistent in terms of position and angle, laying a foundation for more accurate subsequent registration work and improving the overall efficiency.
[0049] Step S105: To achieve precise multi-modal non-rigid weakly supervised registration, a CNN-Transformer hybrid network model based on the Unet architecture is used for weakly supervised non-rigid registration of the augmented dataset, thereby achieving multi-modal registration.
[0050] Generally speaking, the non-rigid registration process in this step is as follows: taking the nuclear magnetic resonance image as the fixed image and the ultrasound image as the moving image, synchronously inputting them into the CNN-Transformer hybrid network model to prompt the network to generate the corresponding deformation field; using the ultrasound label as the moving label, deforming it according to the generated deformation field, calculating the loss between the deformed label and the nuclear magnetic resonance label (i.e., the fixed label), and feeding it back to the network, continuously iterating until an accurate deformation field is generated; finally, deforming the ultrasound image according to the finally generated deformation field to obtain the registered image; in addition, in the inference stage of the model, no label is involved.
[0051] Specifically, this embodiment uses a weight-based U-Net feature extraction module to extract features, accurately positioning the dimension of feature extraction at the pixel level. In the shallow part of the network, the CNN is used to complete the encoding and decoding operations, giving full play to its advantages in local feature extraction and processing; while in the deep region of the network, the Transformer is used to perform the decoding and encoding work, using its long attention mechanism to better capture the long-range dependencies in the image. The entire network architecture is constructed based on the U-Net architecture, which includes upsampling and downsampling operations to achieve effective fusion and transmission of features at different scales. At the same time, in order to enhance the long-term memory ability of the network, Residual Blocks are embedded at the input end of each encoder and the output end of the decoder.
[0052] In the loss calculation stage, this embodiment adopts a multi-scale loss mixing processing method. First, Gaussian blur is applied to the label and downsampled, then Gaussian blur is applied again and upsampled, and finally Gaussian blur is applied once more. The difference between the processed label and the original label is calculated to obtain the edge information for gradient optimization. The same operation is applied to the ultrasound and nuclear magnetic resonance labels, and the L1 loss between the two is calculated to obtain the edge loss. The edge loss, multi-scale Dice, and bending energy loss are combined in a ratio of 10:1:1 to form a mixed edge loss, and the model is trained with the mixed edge loss until convergence. With this carefully designed loss function, the network can more accurately complete the image registration work from multiple dimensions such as the position information, shape, and contour of the target organ, comprehensively improving the accuracy and reliability of registration.
[0053] In a preferred embodiment, the above method further includes:
[0054] Step S106: Construct a human-computer interaction interface for displaying the ultrasound-MRI registration and fusion images, and at the same time, control signals for the above steps can be obtained. This design has important value in clinical surgery, can provide doctors with key intraoperative prostate position and feature information, and strongly assist the smooth progress of the surgery. This interaction interface greatly improves the usability of this design. Its interface layout is clear and concise, and the operation instructions are clear and intuitive, enabling users to quickly get started and easily execute various operations, with a good interaction experience. At the same time, this interface shows excellent portability, has low requirements for the use environment, and only needs to be equipped with an ordinary computer to complete complex feature extraction and preliminary classification tasks, greatly expanding its application scope. In actual development, the python language can be preferentially selected to design this human-computer interaction interface, and a series of key operations such as image preprocessing, segmentation, pairing, and registration are realized through programming.
[0055] The above method adheres to the concept of precision medicine, abandons irrelevant imaging data, focuses on the prostate lesion area, automatically and accurately extracts key features from complex image information, relies on a modular architecture, while expanding the complexity of the network model, focuses on optimizing the feature capture ability of different levels, improves the feature extraction efficiency of the network, introduces the innovative thinking of a hybrid architecture, integrates the advantages of CNN and Transformer, achieves more excellent feature extraction and registration effects, uses automatically segmented mask labels to implement the strategy of weak supervision learning, continuously strengthens the registration accuracy of the network, and creates a human-computer interaction interface to intuitively present the multi-modal registration results to assist surgical operations, with the remarkable advantages of simple process, high efficiency, and excellent registration performance. Compared with the existing technology, it has important clinical application value and broad market development prospects.
[0056] If the above method is implemented in the form of software function units and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0057] In other embodiments, an electronic device is further provided, including one or more processors, a memory, and one or more programs stored in the memory, where the one or more programs include instructions for performing the method as described above.
[0058] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0060] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field according to the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.
Claims
1. A multimodal prostate image registration method based on convolutional neural network, characterized in that: The following steps are involved: Acquire a preoperative nuclear magnetic resonance image and an intraoperative ultrasound image, and perform image preprocessing on the preoperative nuclear magnetic resonance image and the intraoperative ultrasound image, respectively, to obtain a preprocessed nuclear magnetic resonance image and an ultrasound image; Performing an automatic segmentation operation on the target organ on the preprocessed nuclear magnetic resonance image and ultrasound image to obtain segmentation information; According to the segmentation information, the preprocessed nuclear magnetic resonance image and the ultrasound image are paired and rigidly registered; The CNN-Transformer hybrid network model based on the Unet architecture is used to perform weakly supervised non-rigid registration on the rigidly registered MRI images and ultrasound images to achieve multimodal registration.
2. The multimodal prostate image registration method based on convolutional neural network according to claim 1, characterized in that: The image preprocessing includes target organ region extraction, image enhancement, resampling and two-dimensional slice amplification.
3. The multimodal prostate image registration method based on convolutional neural network according to claim 1, characterized in that: The automatic segmentation operation of the target organ is implemented through the U-Net segmentation network architecture to obtain the target organ segmentation mask under multi-modality.
4. The multimodal prostate image registration method based on convolutional neural network according to claim 1, characterized in that: The pairing is specifically: According to the segmentation information, the area, perimeter and centroid of the prostate region are respectively obtained from the preprocessed nuclear magnetic resonance image and ultrasound image, and matching is performed based on the area and perimeter information to find an image group with the same position information.
5. The multimodal prostate image registration method based on convolutional neural network according to claim 4, characterized in that: The rigid registration includes: For the image group with the same position information, the centroid is used as a reference, the position deviation of the nuclear magnetic resonance image and the ultrasound image in the plane is corrected by translation operation, and the image angle difference is adjusted by rotation transformation to obtain the nuclear magnetic resonance image and ultrasound image after rigid registration.
6. The multimodal prostate image registration method based on convolutional neural network according to claim 1, characterized in that: In the CNN-Transformer hybrid network model based on the Unet architecture, an equally weighted U-Net-based feature extraction module is used to extract features, CNN is used to perform encoding and decoding operations in the shallow layer of the network, and Transformer is used to perform decoding and encoding in the deep layer of the network, and residual blocks are embedded in each encoder input and decoder output.
7. The multimodal prostate image registration method based on convolutional neural network according to claim 1, characterized in that: The weakly supervised non-rigid registration specifically includes: The nuclear magnetic resonance image is used as a fixed image and the ultrasound image is used as a moving image, and they are synchronously input into the CNN-Transformer hybrid network model based on the Unet architecture to generate a corresponding deformation field; The ultrasonic tag is used as a mobile tag, and is deformed according to the generated deformation field. The deformed tag and the nuclear magnetic resonance tag are subjected to loss calculation, and the loss is fed back to the network. The iteration is continued until an accurate deformation field is generated. The ultrasonic image is deformed according to the final generated precise deformation field to obtain a registered image.
8. The multimodal prostate image registration method based on convolutional neural network according to claim 1, characterized in that: The CNN-Transformer hybrid network model is trained using a hybrid edge loss, and the acquisition of the hybrid edge loss specifically includes: For ultrasound labels and nuclear magnetic resonance labels, Gaussian blur downsampling, Gaussian blur upsampling and Gaussian blur processing are performed in sequence, and the difference between the processed labels and the original labels is calculated to obtain the gradient optimized edge and calculate the edge loss; The edge loss, multi-scale Dice, and bending energy loss are combined in a set ratio to form a hybrid edge loss.
9. The multimodal prostate image registration method based on convolutional neural network according to claim 1, characterized in that: The method further includes: The control signals for each step are obtained through the human-computer interaction interface, and the multimodal registration effect is displayed.
10. A computer-readable storage medium, characterized in that: It includes one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the multimodal prostate image registration method based on convolutional neural network as described in any one of claims 1-9.
Citation Information
Patent Citations
A Prostate Image Registration System and Method Based on Automatic Segmentation and Pelvic Alignment
CN110363802B
Cited By
CT and ultrasonic image cross-modal registration method based on three-stage process and related device
CN122391315A