Multi-modal medical image registration method and system based on large segmentation model

By employing a multimodal medical image registration method based on a large segmentation model, and utilizing segmentation mask information fusion and a deformable registration network, the non-rigid deformation and noise robustness issues in image registration during prostate surgery are addressed. This method achieves high-precision image alignment, reduces the need for labeled data, and improves the reliability of clinical applications.

CN121074101AActive Publication Date: 2025-12-05SHANDONG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511604243.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2025-12-05
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

Existing multimodal medical image registration methods suffer from insufficient robustness to non-rigid deformation and noise in prostate surgery, and are highly dependent on labeled data, resulting in insufficient registration accuracy and difficulty in meeting clinical needs.

Method used

A multimodal medical image registration method based on a large segmentation model is adopted. By fusing segmentation mask information and a deformable registration network, the three-dimensional spatial deformation field is predicted to achieve high-precision image alignment and reduce dependence on annotation data.

Benefits of technology

It improves the accuracy and robustness of image registration, reduces the need for manual annotation, provides an accurate foundation for image fusion, and offers reliable image alignment support for medical applications such as prostate surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074101A_ABST
    Figure CN121074101A_ABST
Patent Text Reader

Abstract

The invention belongs to the related technical field of image registration, and provides a multi-modal medical image registration method and system based on a segmentation large model in order to solve the problems of poor non-rigid deformation robustness, inaccurate registration and the like in existing medical image registration. Taking the intraoperative prostate three-dimensional image as a target image; respectively extracting a moving mask and a target mask from the moving image and the target image by utilizing the large segmentation model; fusing the moving image, the target image, the moving mask and the target mask to obtain a fusion information tensor; based on the fusion information tensor, through the deformable registration network, a three-dimensional space deformation field is predicted, spatial transformation is carried out on the moving image, the moving image is aligned with the target image after deformation, the robustness of non-rigid deformation is improved, high-precision registration is achieved, and the requirement for manual annotation is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field related to image registration, and particularly relates to a multi-modal medical image registration method and system based on a segmentation large model. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] In the preoperative planning and intraoperative navigation process of prostate surgery, accurately registering different modalities of medical images such as magnetic resonance imaging (MRI) and transrectal ultrasound imaging (TRUS) is a key task. Multi-modal medical image registration aims to spatially align images from different imaging devices so that doctors can integrate the information provided by different modalities of images to more accurately locate lesions, plan surgical paths, and assess surgical risks.

[0004] However, this process faces many challenges. First, MRI and TRUS images have significant differences in imaging principles, resolution, and contrast. MRI can provide high-resolution soft tissue contrast, which helps to clearly show the anatomical structure of the prostate and the location of the lesion; while TRUS images have the advantage of real-time imaging, but have lower resolution and are greatly affected by patient physiological conditions and operator skill levels. Second, the prostate may undergo non-rigid deformation before and during surgery, which may be caused by factors such as bladder filling level, intestinal gas, patient position change, and surgical operation. In addition, there may be noise and artifacts in TRUS images, further increasing the difficulty of registration.

[0005] Traditional multi-modal medical image registration methods mainly rely on feature-based registration algorithms or intensity-based registration algorithms. Feature-based registration algorithms extract feature points, edges or contours, etc. in the image, and then find the correspondence between these features to achieve registration. However, these methods require high accuracy of feature extraction and uniformity of feature point distribution, and are prone to matching errors when dealing with non-rigid deformation and noise. Intensity-based registration algorithms minimize the intensity difference between images to find the best registration parameters, but this type of method usually has large computational load and is prone to local optimal solution, especially in the case of large deformation and noise between images.

[0006] In recent years, with the development of deep learning technology, some deep learning-based registration methods have been proposed. These methods can better handle the nonlinear deformation and noise between images by learning the feature representation of the images, and have higher registration accuracy and robustness than traditional registration methods. However, most existing deep learning registration methods require a large amount of labeled data for training, which is often difficult to obtain in the medical image field because the labeling of medical images requires professional medical knowledge and a large amount of time cost. In addition, these methods still have certain limitations in dealing with specific non-rigid deformation and multi-modal image registration in prostate surgery, such as insufficient robustness to image noise and the need to improve registration accuracy.

[0007] Therefore, how to develop a method that can efficiently and accurately handle multi-modal medical image registration in prostate surgery while reducing the dependence on labeled data and improving the robustness to noise and non-rigid deformation is a technical problem that needs to be solved at present. SUMMARY

[0008] To overcome the shortcomings of the prior art, the present application provides a multi-modal medical image registration method and system based on a segmentation large model, which improves the robustness to non-rigid deformation, achieves high-precision registration, and significantly reduces the need for manual labeling.

[0009] To achieve the above purpose, the present application adopts the following technical solutions: In a first aspect, the present application provides a multi-modal medical image registration method based on a segmentation large model, comprising: obtaining a preoperative prostate three-dimensional image and an intraoperative prostate three-dimensional image as a pair of images to be registered; inputting the pair of images to be registered into a trained registration model to obtain a registration result of the pair of images to be registered; wherein the processing process of the registration model for the pair of images to be registered is: taking the preoperative prostate three-dimensional image as a moving image and the intraoperative prostate three-dimensional image as a target image; extracting a moving mask and a target mask from the moving image and the target image respectively using a segmentation large model; fusing the moving image, the target image, the moving mask and the target mask to obtain a fusion information tensor; based on the fusion information tensor, predicting a three-dimensional space deformation field through a deformable registration network to perform spatial transformation on the moving image, so that the moving image is aligned with the target image after deformation.

[0010] In a second aspect, the present application provides a multi-modal medical image registration system based on a segmentation large model, comprising: an acquisition module configured to obtain a preoperative prostate three-dimensional image and an intraoperative prostate three-dimensional image as a pair of images to be registered; The registration module is configured to input the image pair to be registered into the trained registration model to obtain a registration result of the image pair to be registered. The registration model processes the image pair to be registered as follows: The preoperative prostate three-dimensional image is taken as a moving image, and the intraoperative prostate three-dimensional image is taken as a target image. The segmentation large model is used to extract a moving mask and a target mask from the moving image and the target image, respectively. The moving image, the target image, the moving mask and the target mask are fused to obtain a fusion information tensor. Based on the fusion information tensor, a three-dimensional space deformation field is predicted through a deformable registration network to perform a space transformation on the moving image, so that the moving image is aligned with the target image after deformation.

[0011] In a third aspect, the present application provides an electronic device, comprising a memory and a processor, and computer instructions stored in the memory and running on the processor, when the computer instructions are run by the processor, the method of the first aspect is completed.

[0012] In a fourth aspect, the present application provides a computer readable storage medium for storing computer instructions, when the computer instructions are executed by a processor, the method of the first aspect is completed.

[0013] The above one or more technical solutions have the following beneficial effects: The present application fuses the original image with the segmentation mask information, integrates the segmentation details into the image features, effectively combines the image intensity and the anatomical boundary information, and improves the registration accuracy; based on the fusion information tensor, a three-dimensional space deformation field is predicted to realize the space transformation of the moving image, and the moving image after deformation is highly consistent with the target image in the anatomical structure position and shape, while ensuring the smoothness and continuity of the displacement field, meeting the accuracy and reliability requirements of clinical multi-modal image registration, and providing accurate image fusion basis for medical applications such as prostate surgery.

[0014] The present application integrates the pre-training ability of SAM through the segmentation module, and only a small amount of sample fine-tuning is needed to realize high-precision segmentation of the prostate region. This weakly supervised strategy based on the basic model greatly reduces the demand for manual annotation, reduces the data acquisition and preprocessing cost in clinical application, and avoids registration deviation caused by annotation error.

[0015] The advantages of the additional aspects of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0016] The accompanying drawings, which form a part of this specification, are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification. The embodiments of the application, and their

[0017] Figure 1 A flow chart of a multi-modal medical image registration method based on a segmentation large model in an embodiment of the present application; Figure 2 A three-dimensional voxel segmentation large model network structure based on SAM in an embodiment of the present application; Figure 3 An information fusion layer network structure in an embodiment of the present application; Figure 4 A registration network layer network structure in an embodiment of the present application. DETAILED DESCRIPTION

[0018] It should be noted that the following detailed description is merely exemplary and is intended to provide further description of the application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the application belongs.

[0019] It should be noted that the terms used herein are merely intended to describe specific embodiments and are not intended to limit the exemplary embodiments according to the present application.

[0020] In the case of no conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0021] Embodiment one The present embodiment discloses a multi-modal medical image registration method based on a segmentation large model, comprising: Obtaining a preoperative prostate three-dimensional image and an intraoperative prostate three-dimensional image as a pair of images to be registered; Inputting the pair of images to be registered into the trained registration model to obtain a registration result of the pair of images to be registered; The processing procedure of the registration model for the pair of images to be registered is: Taking the preoperative prostate three-dimensional image as a moving image and the intraoperative prostate three-dimensional image as a target image; Extracting a moving mask and a target mask from the moving image and the target image respectively by using the segmentation large model; Fusing the moving image, the target image, the moving mask and the target mask to obtain a fusion information tensor; Based on the fusion information tensor, predicting a three-dimensional space deformation field by a deformable registration network to perform a space transformation on the moving image, so that the moving image is aligned with the target image after deformation.

[0022] The embodiment fuses the original image with the segmentation mask information, integrates the segmentation details into the image features, effectively combines the image intensity and the anatomical boundary information, and improves the registration accuracy; based on the fusion information tensor, a three-dimensional space deformation field is predicted, the spatial transformation of the moving image is realized, the moving image after deformation is highly consistent with the target image in the anatomical structure position and shape, while ensuring the smoothness and continuity of the displacement field, meeting the accuracy and reliability requirements of clinical multi-modal image registration, and providing accurate image fusion basis for medical applications such as prostate surgery.

[0023] The following will be combined Figure 1 A multi-modal medical image registration method based on a segmentation large model is described in detail: The embodiment proposes a registration model of a segmentation-fusion-registration integrated architecture for medical image registration, and the registration model comprises a SAM-based three-dimensional voxel segmentation large model, an information fusion layer, and a registration network layer.

[0024] S1: Obtain a preoperative prostate three-dimensional image as a moving image , obtain an intraoperative prostate three-dimensional image as a target image , and construct a pair of images to be registered.

[0025] In S1, the person skilled in the art can select a nuclear magnetic resonance imaging instrument, an ultrasonic imaging instrument, etc. to obtain the preoperative and intraoperative prostate three-dimensional images of the patient.

[0026] S2: Use the SAM-based three-dimensional voxel segmentation large model to extract a moving mask and a target mask from the moving image and the target image respectively, and accurately define important segmentation structures such as lesion areas.

[0027] As shown in Figure 2 , the SAM-based three-dimensional voxel segmentation large model takes the moving image and the target image as input and predicts the segmentation mask as output, specifically including: S201: input the moving image and the target image , and send them into the normalization layer respectively to standardize the pixel intensity of the image and eliminate the brightness difference between different modalities; S202: capture the long-distance spatial dependence relationship of the moving image and the target image across the slices through the memory attention layer, enhance the modeling ability of the context information on the prostate region, and suppress irrelevant background interference; S203: perform spatial dimension convolution operation on the features output by the memory attention layer through the three-dimensional convolution layer, extract the three-dimensional features containing spatial and depth information, and capture the three-dimensional structure details of the prostate; S204: After three-dimensional convolution, the high-dimensional features are compressed by a linear adjustment layer to reduce the channel dimension, retain key information, reduce computational complexity, control the number of new parameters, and improve model efficiency; S205: The compressed features are input into the SAM module, and the pre-trained SAM model is used to complete the semantic segmentation of important structures and generate a coarse-grained mask based on the generalization ability of the medical image with only a small amount of labeled samples; S206: The feature map is upsampled by the memory decoding layer to restore the input image resolution, and a high-precision moving mask and target mask are output .

[0028] S3: Through the information fusion layer, the moving image , the target image and their corresponding mask information are fused and processed to output the information fusion result.

[0029] As shown in Figure 3 , the present embodiment provides two different information fusion strategies, namely direct information splicing and feature fusion splicing. The following will be described respectively: S301: Direct information splicing, the technical principle is to directly splice the moving image, the target image and their segmentation masks along the channel dimension to form a combined tensor containing the original pixel intensity and anatomical boundary information.

[0030] The moving image , the target image , the moving mask , the target mask are directly spliced along the channel dimension to generate a combined tensor , wherein represents directly splicing elements along the first dimension of the vector. This strategy directly retains the original form of image pixel intensity and mask binary region, and is suitable for scenarios where the difference between image modalities is large and the intensity information and anatomical boundary information need to be used equally. For example, when the gray scale distribution of prostate MR images and ultrasound images is significantly different due to different imaging principles, the original features of each modality can be retained by direct splicing to provide comprehensive input information for the registration network.

[0031] S302: Feature fusion splicing, the technical principle is to selectively enhance the image features of the prostate region by feature extraction and weighted modulation, while retaining key anatomical structure information and reducing computational complexity.

[0032] Using a U-Net encoder, the moving mask and the target mask are subjected to hierarchical feature extraction to obtain corresponding image feature tensors and , the confidence values of the mask feature vector are used to weight and modulate the image features, and the enhanced feature representation is generated by element-wise multiplication and information combination operations , wherein represents element-wise multiplication, and the image features in the prostate region are selectively enhanced by the spatial distribution of the mask (such as the binary mask of the prostate region), which reduces the computational complexity while retaining key anatomical structure information, and is suitable for registration tasks that need to highlight local anatomical features.

[0033] This strategy is suitable for registration tasks that need to highlight local anatomical features, such as precise alignment focusing on the prostate tumor region. By using mask feature weighting, image details of the tumor boundary can be highlighted, and computational redundancy in non-key regions can be reduced.

[0034] S4: Use the registration network layer input fusion result to calculate a three-dimensional spatial deformation field, so that the moving image is spatially transformed by the deformation field to obtain a transformed moving image.

[0035] As shown in Figure 4 , the specific steps are as follows: S401: The registration network layer input is the fusion information tensor of the information fusion layer output , which is reduced to 1 / 2 of the original size by downsampling operation, reducing the amount of calculation while expanding the receptive field of the neural network to capture global structural features in the image; S402: Perform three-dimensional convolution operation on the down-sampled feature tensor, use 3x3x3 convolution kernel to extract local structure information in space and channel dimension, and gradually build deep feature representation by multi-layer convolution stacking to capture anatomical structure relationship of different scales; S403: Restore the deep feature representation to near original resolution by upsampling operation (such as tri-linear interpolation), and combine the shallow detail features of the down-sampling stage by jump connection to form multi-scale displacement field , , , ; taking into account the global deformation trend and local structure alignment accuracy; S404: Element-wise average the multi-scale displacement field, use Gaussian smoothing filter to eliminate noise and abnormal displacement values, generate smooth and continuous three-dimensional spatial deformation field ; S405: Use the final three-dimensional spatial deformation field to perform spatial transformation on the moving image , calculate the new position of each pixel in the target image coordinate system by nearest neighbor interpolation algorithm, generate the transformed moving image , represents the transformation operation.

[0036] S5: Calculate the similarity loss of the transformed moving image and the target image, the label overlap loss, and the smoothness loss based on the deformation field, for training the registration model.

[0037] Steps S1-S5 are executed in a loop until the registration model converges to complete training.

[0038] The specific process is as follows: S501: Similarity loss The normalized cross correlation (NCC) loss is calculated, which eliminates the inter-modal gray difference through normalization processing. The formula is:

[0039] wherein, is the image pixel index, and represent the pixel mean of the transformed target image and the target image, respectively; represents the moving image, represents the target image; represents the three-dimensional space deformation field.

[0040] S502: Label overlap loss The Dice coefficient loss is calculated using the transformed moving mask and the target mask The formula is:

[0041] wherein, and represent the intersection volume and the total volume of the two masks, respectively, both of which are scalars; is the moving mask; is the target mask.

[0042] S503: Deformation field smoothness loss To avoid abnormal displacement of the deformation field, a gradient norm penalty term is used to constrain its spatial smoothness. The formula is:

[0043] wherein, represents the total number of pixels in the deformation field, represents the spatial gradient of the deformation field at pixel , which suppresses the displacement mutation between adjacent pixels through L2 norm, ensuring the continuity of deformation in accordance with human anatomy; represents the square of the L2 norm.

[0044] S504: Weighted sum of triple loss to get total loss :

[0045] wherein, 、 、 are weight coefficients. The parameters of the registration network are updated by a back propagation algorithm, and the deformation field prediction ability is iteratively adjusted.

[0046] S505: Set the convergence threshold and the maximum number of iterations to 1000, and terminate the training when one of the following conditions is met: 1. The Dice coefficient on the validation set no longer improves within 100 consecutive iterations, and the loss value tends to be stable; 2. The maximum preset training round is reached to avoid overfitting.

[0047] S6: Apply the trained registration network to testing, repeat the steps S1-S4 to obtain the transformed image, and map it to the real physical environment to provide visualization assistance for prostate surgery.

[0048] According to the steps S1-S4, the registration of the preoperative prostate MR image and the intraoperative three-dimensional ultrasound image is implemented. Through the registration process, the preoperative MR reconstructed three-dimensional structure of the prostate is superimposed on the intraoperative ultrasound image, and key anatomical structures such as tumors and urethra are labeled. At the same time, with the help of the registration result, the displacement error of the prostate caused by factors such as respiration and compression is compensated, thereby providing visualization assistance for prostate surgery.

[0049] In this embodiment, the preoperative three-dimensional image and the intraoperative three-dimensional image are obtained to construct a pair of images to be registered, and after normalization, high-precision masks are extracted through a segmentation module containing a memory attention layer, a three-dimensional convolution layer and a SAM model; the information fusion layer processes image and mask information by direct information splicing or feature fusion splicing strategy; the registration network layer generates a three-dimensional deformation field through downsampling, three-dimensional convolution, upsampling and Gaussian smoothing, and transforms the moving image space; the network is trained to convergence by using similarity loss, label overlap loss and deformation field smoothness loss, and the registration result is mapped to the physical environment to assist surgery during testing. The computer device obtains images through the DICOM interface, and after segmentation, fusion and registration, it realizes the alignment of images and surgical space in combination with the optical positioning system, and provides a fusion view in a semi-transparent superposition manner to assist surgical navigation.

[0050] The embodiment constructs a full-process automation framework from image acquisition to physical space mapping. Preoperatively, multi-modal images are acquired by MRI and ultrasound, and a deformation field is generated by segmentation, fusion and registration. Intraoperatively, the registered virtual structure is mapped to the surgical robot coordinate system using an optical positioning sensor, realizing real-time fusion display of preoperative MR and intraoperative ultrasound. The framework supports docking with mainstream surgical navigation systems, provides dynamic anatomical guidance for prostate biopsy, tumor ablation and other surgeries, compensates for displacement errors caused by physiological movements such as respiration and compression, and improves surgical accuracy.

[0051] Embodiment two The purpose of the embodiment is to provide a multi-modal medical image registration system based on a segmentation large model, comprising: The acquisition module is configured to acquire a preoperative prostate three-dimensional image and an intraoperative prostate three-dimensional image as a pair of images to be registered; The registration module is configured to input the pair of images to be registered into the trained registration model to obtain a registration result of the pair of images to be registered; The processing process of the registration model on the pair of images to be registered is: The preoperative prostate three-dimensional image is taken as a moving image, and the intraoperative prostate three-dimensional image is taken as a target image; The segmentation large model is used to extract a moving mask and a target mask from the moving image and the target image respectively; The moving image, the target image, the moving mask and the target mask are fused to obtain a fusion information tensor; Based on the fusion information tensor, a three-dimensional space deformation field is predicted by a deformable registration network to perform spatial transformation on the moving image, so that the moving image is aligned with the target image after deformation.

[0052] In more embodiments, there are also provided: An electronic device includes a memory and a processor, and computer instructions stored on the memory and running on the processor, when the computer instructions are run by the processor, the method described in embodiment one is completed. For brevity, this will not be repeated here.

[0053] It should be understood that in the embodiment, the processor can be a central processing unit CPU, and the processor can also be other general-purpose processors, digital signal processors DSPs, application-specific integrated circuits ASICs, ready-to-program gate arrays FPGA or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0054] The memory can include read-only memory and random access memory, and provide instructions and data to the processor, a portion of the memory can also include non-volatile random access memory. For example, the memory can also store device type information.

[0055] A computer readable storage medium for storing computer instructions, which are executed by a processor to complete the method described in embodiment one.

[0056] The method in embodiment one can be directly embodied as a hardware processor to complete, or be completed by a combination of hardware and software modules in the processor. The software modules can be located in the storage medium mature in the art such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory to complete the steps of the above method in combination with the hardware. To avoid repetition, it will not be described in detail here.

[0057] A computer program product comprising a computer program, which, when executed by a processor, implements the method described in embodiment one.

[0058] The present application also provides at least one computer program product tangibly stored on a non-transitory computer readable storage medium. The computer program product includes computer executable instructions, such as those included in program modules, executed by devices at a destination real or virtual processor to perform processes / methods as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform particular tasks or implement particular abstract data types. In various embodiments, the functionality of program modules can be combined or split between program modules as desired. Machine executable instructions for program modules can be executed within a local or distributed device. In a distributed device, program modules can be located in local and remote memory storage devices.

[0059] The computer program code for carrying out the methods of the present application can be written in one or more programming languages. These computer program codes can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, so that the program codes cause the functions / operations specified in the flowcharts and / or block diagrams to be performed when the computer or other programmable data processing apparatus executes the program codes. The program codes can be executed entirely on the computer, partially on the computer, as a standalone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0060] In the context of the present application, the computer program code or related data can be carried by any suitable carrier to enable the device, apparatus or processor to perform the various processes and operations described above. Examples of carriers include signals, computer readable media, and the like. Examples of signals can include electrical, optical, radio, sound or other forms of propagated signals, such as carrier waves, infrared signals, and the like.

[0061] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the present embodiment can be realized in electronic hardware or in combination of computer software and electronic hardware. Whether the functions are realized in hardware or software manner depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0062] Although the specific embodiments of the present application are described above in combination with the drawings, it is not a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications or variations made by those skilled in the art on the basis of the technical solutions of the present application without creative labor are still within the scope of protection of the present application.

Claims

1. A multi-modal medical image registration method based on a segmentation large model, characterized in that, The method comprises the following steps: obtaining a preoperative prostate three-dimensional image and an intraoperative prostate three-dimensional image as a pair of images to be registered; inputting the pair of images to be registered into a trained registration model to obtain a registration result of the pair of images to be registered; wherein the processing procedure of the registration model for the pair of images to be registered is as follows: taking the preoperative prostate three-dimensional image as a moving image and taking the intraoperative prostate three-dimensional image as a target image; extracting a moving mask and a target mask from the moving image and the target image respectively by using a segmentation large model; fusing the moving image, the target image, the moving mask and the target mask to obtain a fusion information tensor; based on the fusion information tensor, predicting a three-dimensional spatial deformation field through a deformable registration network to perform spatial transformation on the moving image, so that the moving image is aligned with the target image after deformation.

2. The multi-modal medical image registration method based on a segmentation large model according to claim 1, wherein, extracting the moving mask and the target mask from the moving image and the target image respectively by using the segmentation large model, specifically as follows: performing standardization processing on the pixel intensity of the moving image and the target image; capturing long-distance spatial dependency of the moving image or the target image across slices through a memory attention layer to enhance the modeling capability of the context information of the prostate region; performing spatial dimension convolution operation on the output features of the memory attention layer through a three-dimensional convolution layer to extract stereo features containing spatial and depth information and capture three-dimensional structural details of the prostate; compressing the high-dimensional features after three-dimensional convolution through a linear adjustment layer; inputting the compressed features into a pre-trained SAM convolution-deconvolution module to obtain the moving mask corresponding to the moving image or the target mask corresponding to the target image.

3. The multi-modal medical image registration method based on a segmentation large model according to claim 1, wherein, fusing the moving image, the target image, the moving mask and the target mask to obtain the fusion information tensor, specifically as follows: splicing the moving image, the target image, the moving mask and the target mask along the channel dimension to generate the fusion information tensor.

4. The multi-modal medical image registration method based on a segmentation large model according to claim 1, wherein, fusing the moving image, the target image, the moving mask and the target mask to obtain the fusion information tensor, specifically as follows: extracting hierarchical features of the moving mask and the target mask respectively to obtain image feature tensors corresponding to the moving mask and the target mask; performing element multiplication and combination operation on the moving mask and the target mask and the corresponding image feature tensors respectively to obtain the fusion information tensor.

5. The multi-modal medical image registration method based on a segmentation large model according to claim 1, wherein, based on the fusion information tensor, predicting a three-dimensional spatial deformation field through a deformable registration network to perform spatial transformation on the moving image, specifically as follows: reducing the spatial resolution of the fusion information tensor to 1 / 2 of the original size through downsampling operation; performing three-dimensional convolution operation on the feature tensor after downsampling to extract local structural information in the spatial and channel dimensions, gradually constructing deep feature representation through multi-layer convolution stacking, and capturing anatomical structure relationship of different scales; restoring the deep feature representation to the original resolution through upsampling operation, combining the shallow detail features in the downsampling stage through jump connection to form a multi-scale displacement field; performing element-wise averaging on the multi-scale displacement field, eliminating noise and abnormal displacement values by using Gaussian smoothing filter to generate a smooth and continuous three-dimensional spatial deformation field; performing spatial transformation on the moving image by using the three-dimensional spatial deformation field, calculating the new position of each pixel in the coordinate system of the target image to generate the transformed moving image.

6. The multi-modal medical image registration method based on a segmentation large model according to claim 1, wherein, The loss function of the registration model training includes a similarity loss, a label overlap loss and a deformation field smoothness loss.

7. The multi-modal medical image registration method based on a segmentation large model according to claim 1 or 6, characterized in that, The loss function of the registration model training is specifically: ; ; ; ; where, is the image pixel index, is the target image pixel index, respectively represent the transformed target image and the target image pixel mean; denotes the moving image, denotes the target image; denotes the three-dimensional space deformation field; is the moving mask; is the target mask; and respectively represent the intersection volume and the total volume of the transformed moving mask and the target mask; denotes the spatial gradient of the deformation field at pixel ; is the loss function of the registration model; is the similarity loss; is the label overlap loss; is the deformation field smoothness loss; , , is the weight coefficient, and N represents the total number of pixels of the deformation field; denotes the square of the L2 norm.

8. A multi-modal medical image registration system based on a segmentation large model, characterized by, comprises: An acquisition module configured to acquire a preoperative prostate three-dimensional image and an intraoperative prostate three-dimensional image as a to-be-registered image pair; A registration module configured to input the to-be-registered image pair into the trained registration model to obtain a registration result of the to-be-registered image pair; The processing procedure of the registration model on the to-be-registered image pair is: taking the preoperative prostate three-dimensional image as a moving image and taking the intraoperative prostate three-dimensional image as a target image; extracting a moving mask and a target mask from the moving image and the target image respectively by using a segmentation large model; fusing the moving image, the target image, the moving mask and the target mask to obtain a fusion information tensor; based on the fusion information tensor, predicting a three-dimensional space deformation field by using a deformable registration network to perform a space transformation on the moving image, so that the moving image is aligned with the target image after deformation.

9. An electronic device, comprising: A computer device comprising a memory and a processor and computer instructions stored on the memory and running on the processor, when the computer instructions are run by the processor, the method of any one of claims 1-7 is completed.

10. A computer-readable storage medium, characterized in that, A computer device for storing computer instructions, when the computer instructions are executed by the processor, the method of any one of claims 1-7 is completed.

Citation Information

Patent Citations

  • High-precision three-dimensional medical image semantic segmentation system and method based on improved SAM model

    CN118864855A

  • Cross-scene multi-domain fusion small sample remote sensing target robust identification method

    CN118918476A

  • Percutaneous lumbar intervertebral disc puncture surgery navigation method and device based on multi-modal image

    CN119745507A

  • Medical image registration method based on large model robust features

    CN120782830A

  • Multi-modal medical segmentation method and framework based on transfer learning

    CN120852337A