Multimodal image-based navigation method and device for percutaneous lumbar disc puncture surgery

Through multimodal image segmentation and registration technology, the radiation exposure and insufficient precision problems of percutaneous lumbar disc puncture surgery in existing technologies are solved, and higher-precision navigation path planning and safety improvement are achieved.

CN119745507BActive Publication Date: 2025-10-03THE FIFTH AFFILIATED HOSPITAL SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411605645.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-10-03
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

In existing technologies, percutaneous lumbar disc puncture surgery performed under two-dimensional image guidance is subject to large radiation exposure, high risk of nerve root, dura mater and blood vessel damage, and insufficient surgical navigation accuracy and feasibility. In particular, it is difficult to accurately distinguish between bones, flesh and other anatomical structures in complex spinal structures.

Method used

A multimodal image-based percutaneous lumbar disc puncture surgical navigation method was adopted. The multi-contrast multimodal image segmentation model and the multimodal image rigid-elastic hybrid registration model were used, combined with preoperative MR and CT images, to perform image segmentation and registration, construct a three-dimensional visualization model of the spine, and generate a puncture navigation path.

Benefits of technology

It improves the accuracy and feasibility of surgical navigation, reduces radiation exposure and tissue damage risks by clearly displaying multiple structures of the spine, and achieves more accurate image guidance and navigation path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119745507B_ABST
    Figure CN119745507B_ABST
Patent Text Reader

Abstract

The present invention proposes a percutaneous lumbar disc puncture surgical navigation method and device based on multimodal images. The method includes: obtaining preoperative MR images and CT images, sequentially inputting a multi-contrast multimodal image segmentation model and a multimodal image rigid-elastic hybrid registration model to obtain a first registration image; based on the intraoperative three-dimensional CBCT image, calling the multi-contrast multimodal image segmentation model and the multimodal image rigid-elastic hybrid registration model again to obtain a second registration image; and constructing a three-dimensional visualization model of the spine based on the second registration image and generating a puncture navigation path. According to the technical solution of the embodiment of the present invention, it is possible to automatically align the surgical MR image, CT image, and intraoperative three-dimensional CBCT image using a mask, introduce two sequences, fat suppression and non-fat suppression, into the preoperative MR image, so that the second registration image can distinguish between bone, flesh, and spinal structure, thereby improving the accuracy of the three-dimensional visualization model of the spine, and improving the accuracy of the navigation path and the feasibility of the surgery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of surgical navigation, and in particular to a multimodal image-based percutaneous lumbar disc puncture surgical navigation method and device. Background Art

[0002] Currently, percutaneous lumbar discectomy is often performed under the guidance of two-dimensional C-arm CT (CBCT) images during the transforaminal puncture procedure. Multiple fluoroscopy sessions and repeated adjustments of the needle to the location of the disc lesion may be required. This process not only increases radiation exposure but also increases the risk of nerve root, dura mater, and vascular damage, resulting in disadvantages such as low radiation exposure, safety, and accuracy. Furthermore, the lumbar intervertebral foramen is complex, containing important anatomical structures such as fat, nerve roots, arteries and veins, foraminal ligaments, and sinus nerves. The medial aspect of the foramina is the dura mater and nerve roots. Degeneration-induced thickening of the ligamentum flavum and hyperplasia of the facet joints can further alter the morphology of the intervertebral foramen. These factors pose difficulties and risks to percutaneous lumbar discectomy.

[0003] In scenarios with complex spinal structures, percutaneous lumbar disc puncture surgery guided by two-dimensional CBCT images, magnetic resonance (MR) images, and computed tomography (CT) images currently suffers from difficulties in distinguishing between bone and flesh, lacks intuitive image guidance, relies heavily on the surgeon's experience, and can lead to intraoperative damage to tissues such as nerve roots, dura mater, vertebral endplates, and blood vessels. Consequently, surgical navigation lacks accuracy and feasibility. Summary of the Invention

[0004] The present invention aims to address at least one of the technical problems existing in the prior art. To this end, the present invention proposes a multimodal image-based percutaneous lumbar disc puncture surgical navigation method and device. This method can clearly display multiple spinal structures in all directions during surgical navigation, enabling intuitive and accurate image guidance and improving the accuracy and feasibility of surgical navigation.

[0005] In a first aspect, an embodiment of the present invention provides a multimodal image-based percutaneous lumbar disc puncture surgical navigation method, which is applied to a surgical navigation system. The surgical navigation system is pre-configured with a multi-contrast multimodal image segmentation model and a multimodal image rigid-elastic hybrid registration model. The multimodal image-based percutaneous lumbar disc puncture surgical navigation method includes:

[0006] Acquiring a preoperative spinal image set, wherein the preoperative spinal image set includes a preoperative thin-layer fat-suppressed MR image, a preoperative thin-layer non-fat-suppressed MR image, and a preoperative CT image, wherein the preoperative thin-layer fat-suppressed MR image is acquired based on a preset fat-suppressed sequence, and the preoperative thin-layer non-fat-suppressed MR image is acquired based on a preset non-fat-suppressed sequence;

[0007] Inputting the preoperative spinal image set into the multi-contrast multimodal image segmentation model for image segmentation, and fusing the image segmentation results to obtain a first mask, wherein the first mask is used to indicate the spinal structure or preoperative lesion target;

[0008] Inputting the preoperative spinal image set and the first mask into a multimodal image rigid-elastic hybrid registration model, performing a registration and alignment operation, and obtaining a first registered image, wherein the first registered image includes the preoperative thin-layer fat-suppressed MR image, the preoperative thin-layer non-fat-suppressed MR image, and the preoperative CT image that are registered with each other;

[0009] Acquiring an intraoperative three-dimensional CBCT image, and inputting the intraoperative three-dimensional CBCT image into the multi-contrast multimodal image segmentation model for image segmentation to obtain a second mask, wherein the second mask is used to indicate a vertebral structure;

[0010] Inputting the second mask, the intraoperative 3D CBCT image, and the first registered image into the multimodal image rigid-elastic hybrid registration model, performing the registration and alignment operation, and obtaining a second registered image, wherein the second registered image includes the first registered image and the intraoperative 3D CBCT image that are registered with each other;

[0011] A three-dimensional visualization model of the spine is constructed based on the second registered image, and when a target lesion target marked on the three-dimensional visualization model of the spine is acquired, a puncture navigation path is generated based on the target lesion target.

[0012] According to some embodiments of the present invention, the multi-contrast multimodal image segmentation model includes a first segmentation branch and a second segmentation branch, the first segmentation branch includes a first encoder and a first decoder, the second segmentation branch includes a second encoder and a second decoder, and the preoperative spinal image set is input into the multi-contrast multimodal image segmentation model for image segmentation, and the image segmentation results are fused to obtain a first mask, including:

[0013] Inputting the preoperative thin-layer fat-suppressed MR image into the first segmentation branch, extracting first regional features of the preoperative thin-layer fat-suppressed MR image using the first encoder, extracting first semantic information of the first regional features using the first decoder, and determining a first segmentation result based on the first semantic information, wherein the first regional features are used to indicate the spinal structure or preoperative lesion target;

[0014] Inputting the preoperative thin-layer non-fat-suppressed MR image into the first segmentation branch, and inputting the CT image into the second segmentation branch;

[0015] extracting a second regional feature of the preoperative thin-layer non-fat-suppressed MR image by the first encoder, extracting second semantic information of the second regional feature by the first decoder, and determining a second segmentation result based on the second semantic information, wherein the second regional feature is used to indicate the spinal structure or the preoperative lesion target;

[0016] inputting the second semantic information into the second segmentation branch, and extracting a third region feature of the CT image through the second encoder;

[0017] Performing comparative learning based on the second semantic information and the third region feature to obtain a third segmentation result;

[0018] The first segmentation result, the second segmentation result, and the third segmentation result are fused to obtain the first mask.

[0019] According to some embodiments of the present invention, the second decoder includes multiple attention perception fusion modules, and the performing comparative learning based on the second semantic information and the third region feature to obtain the third segmentation result includes:

[0020] Inputting the second semantic information into the attention perception fusion module to obtain an attention feature, and upsampling the attention feature and fusing it into the third region feature;

[0021] extracting third semantic information of the third region feature, and determining a third segmentation result based on the third semantic information and a preset contrast loss function, wherein the third region feature is used to indicate a vertebral structure;

[0022] The contrast loss function satisfies the following conditions: , is the cross entropy loss function of the first segmentation branch, is the cross entropy loss function of the second segmentation branch, and is the preset magnification factor, is the contrast loss function, is the total loss function of the multi-contrast multimodal image segmentation model.

[0023] According to some embodiments of the present invention, the first encoder and the second encoder perform encoding N times, the first decoder and the second decoder perform decoding N times, the number of the attention-aware fusion modules is N, and N is a positive integer greater than 2. Inputting the second semantic information into the attention-aware fusion module to obtain an attention feature, and upsampling the attention feature and fusing it into the third region feature includes:

[0024] Inputting the second semantic information obtained by the first decoder for the first time into the first attention perception fusion module, performing an attention extraction operation, and obtaining the first attention feature;

[0025] Upsampling the first attention feature, fusing it with the third region feature obtained by the Nth encoding of the second encoder, and then inputting it into the second attention perception fusion module;

[0026] Inputting the second semantic information obtained by the second decoding of the first decoder into the second attention perception fusion module, performing the attention extraction operation, and obtaining the second attention feature;

[0027] When the Nth attention feature is obtained, it is fused with the third region feature obtained by the first encoding of the second encoder to obtain the attention feature.

[0028] According to some embodiments of the present invention, the attention perception fusion module includes a channel attention module and a spatial attention module, the channel attention module includes a first convolution block, a first BN block and a first ReLu layer, a first maximum pooling layer, a first average pooling layer, a multi-layer perceptron and a first Sigmoid function, the spatial attention module includes a second convolution block, a second BN block, a third convolution block and a third BN block, and the attention extraction operation includes:

[0029] Determining a segmented image pair, the segmented image pair including the preoperative thin-slice non-fat-suppressed MR image and the preoperative CT image, or including the first registered image and the intraoperative three-dimensional CBCT image;

[0030] Obtaining a first transformed feature based on the second semantic information by sequentially processing the first convolution block, the first BN block, and the first ReLu layer;

[0031] Inputting the first transformed features into the first maximum pooling layer and the first average pooling layer respectively, fusing the outputs of the first maximum pooling layer and the first average pooling layer, and then sequentially passing the outputs through the multi-layer perceptron and the first sigmoid function to obtain a channel attention map;

[0032] Obtain a second transformed feature based on the third region feature passing through the second convolution block and the second BN block in sequence;

[0033] Inputting the first transformed feature into the spatial attention module, fusing the first transformed feature with the second transformed feature, and then sequentially passing through the third convolution block and the third BN block to obtain a spatial attention map;

[0034] The attention feature is obtained based on the spatial attention map and the channel attention map.

[0035] According to some embodiments of the present invention, the multimodal image rigid-elastic hybrid registration model includes a multiple affine matrix estimation module, an affine elastic fusion module, and a local rigid constraint module, and the registration and alignment operation includes:

[0036] Inputting a plurality of images to be registered into the multiple affine matrix estimation module to obtain multi-scale image features and initial deformation fields of each of the images to be registered, wherein the images to be registered include the preoperative thin-layer fat-suppressed MR image, the preoperative thin-layer non-fat-suppressed MR image, and the preoperative CT image, or the images to be registered include the first registration image and the intraoperative three-dimensional CBCT image;

[0037] Based on the same image to be registered, the corresponding multi-scale image features and the initial deformation field are input into the affine elastic fusion module, deformation field splicing is performed based on the initial deformation field to obtain a rigid elastic deformation field, and all the rigid elastic deformation fields are fused to output an initial registered image;

[0038] The initial registration image and the registration mask are input into the local rigid constraint module, and the rigid-elastic deformation field of the initial registration image is corrected by the local rigid constraint module to obtain a target registration image, wherein, when the image to be registered includes the preoperative spinal image set, the registration mask is the first mask, or when the image to be registered includes the intraoperative three-dimensional CBCT image, the registration mask is the second mask.

[0039] According to some embodiments of the present invention, the multiple affine matrix estimation module includes four consecutively arranged fourth convolution blocks, a flat layer, and a rigid transformation estimation module, the rigid transformation estimation module includes multiple dense layers and multiple affine matrices, the dense layers uniquely correspond to the affine matrices, and the multi-scale image features and initial deformation fields of each of the images to be registered are obtained, including:

[0040] Subjecting the image to be registered to convolution processing of each of the fourth convolution blocks in sequence to obtain the multi-scale image features;

[0041] After the multi-scale image features are expanded based on each scale through the flattening layer, the image features of each scale are input into one of the dense layers of the rigid transformation estimation module, and a rigid deformation field is calculated based on preset rigid transformation estimation parameters;

[0042] The initial deformation field is obtained by performing weighted averaging on all the rigid deformation fields.

[0043] According to some embodiments of the present invention, the affine elastic fusion module includes four combination units, three sixth convolution blocks and an integrated block, the combination unit includes an upsampling layer and a fifth convolution block, the deformation field is spliced ​​based on the initial deformation field to obtain a rigid elastic deformation field, and all the rigid elastic deformation fields are fused to output an initial registration image, including:

[0044] The image features of each scale are sequentially processed by the four combination units to obtain the structural image features of each scale;

[0045] Splicing the structural image features of the corresponding scale with the initial deformation field to obtain a rigid elastic deformation field;

[0046] Each rigid-elastic deformation field is input into the sixth convolution block for convolution processing, and then input into the integrated block for fusion to obtain the initial registration image.

[0047] According to some embodiments of the present invention, the step of correcting the rigid-elastic deformation field of the initial registration image by the local rigid constraint module to obtain a target registration image includes:

[0048] performing least squares regression calculation on the rigid-elastic deformation field based on the registration mask to obtain a reference deformation field;

[0049] constructing a rigid constraint cost function based on an error between the rigid elastic deformation field and the reference deformation field;

[0050] The initial registration image is modified based on the rigid constraint cost function to obtain the target registration image.

[0051] In a second aspect, an embodiment of the present invention provides a percutaneous lumbar disc puncture surgical navigation device based on multimodal images, comprising at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the percutaneous lumbar disc puncture surgical navigation method based on multimodal images as described in the first aspect above.

[0052] In a third aspect, an embodiment of the present invention provides an electronic device comprising the multimodal image-based percutaneous lumbar disc puncture surgical navigation device as described in the second aspect above.

[0053] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the multimodal image-based percutaneous lumbar disc puncture surgical navigation method as described in the first aspect above.

[0054] According to an embodiment of the present invention, the percutaneous lumbar disc puncture surgical navigation method based on multimodal images has at least the following beneficial effects: obtaining a preoperative spinal image set, wherein the preoperative spinal image set includes a preoperative thin-layer fat-suppressed MR image, a preoperative thin-layer non-fat-suppressed MR image and a preoperative CT image, the preoperative thin-layer fat-suppressed MR image is obtained based on a preset fat-suppressed sequence, and the preoperative thin-layer non-fat-suppressed MR image is obtained based on a preset non-fat-suppressed sequence; inputting the preoperative spinal image set into the multi-contrast multimodal image segmentation model for image segmentation, and fusing the image segmentation results to obtain a first mask, wherein the first mask is used to indicate the spinal structure or the preoperative lesion target; inputting the preoperative spinal image set and the first mask into the multimodal image rigid-elastic hybrid registration model, performing a registration and alignment operation, and obtaining a first registered image, which includes the corresponding The pre-operative thin-layer fat-suppressed MR image, the pre-operative thin-layer non-fat-suppressed MR image and the pre-operative CT image are mutually registered; an intra-operative three-dimensional CBCT image is acquired, and the intra-operative three-dimensional CBCT image is input into the multi-contrast multi-modal image segmentation model for image segmentation to obtain a second mask, wherein the second mask is used to indicate the vertebral structure; the second mask, the intra-operative three-dimensional CBCT image and the first registration image are input into the multi-modal image rigid-elastic hybrid registration model, and the registration alignment operation is performed to obtain a second registration image, wherein the second registration image includes the mutually registered first registration image and the intra-operative three-dimensional CBCT image; a three-dimensional visualization model of the spine is constructed based on the second registration image, and when the target lesion target marked on the three-dimensional visualization model of the spine is acquired, a puncture navigation path is generated based on the target lesion target. According to the technical solution of the embodiment of the present invention, image segmentation can be used to determine a mask representing the spinal structure or vertebral structure, and the mask can be used to automatically align the preoperative thin-layer MR image, CT image and intraoperative three-dimensional CBCT image, and the preoperative MR image introduces two sequences, fat suppression and non-fat suppression, so that the second registered image can clearly display the lumbar vertebral structure at the MR, CT and CBCT levels, improve the distinction between bones, flesh and various structures in the second registered image, improve the accuracy of the three-dimensional visualization model of the spine, and thus improve the accuracy of the navigation path and the feasibility of the operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 This is a flowchart of a multimodal image-based percutaneous lumbar disc puncture surgical navigation method provided by one embodiment of the present invention;

[0056] Figure 2 is a complete flow chart of a multimodal image-based percutaneous lumbar disc puncture surgical navigation method provided by another embodiment of the present invention;

[0057] Figure 3 is a structural diagram of a multi-contrast multimodal image segmentation model provided by another embodiment of the present invention;

[0058] Figure 4 is a structural diagram of an attention perception fusion module provided by another embodiment of the present invention;

[0059] Figure 5 is a structural diagram of a multimodal image rigid-elastic hybrid registration model provided by another embodiment of the present invention;

[0060] Figure 6 It is a structural diagram of a multimodal image-based percutaneous lumbar disc puncture surgical navigation device provided by another embodiment of the present invention. DETAILED DESCRIPTION

[0061] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0062] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the present invention.

[0063] In the description of the present invention, "several" means one or more, "many" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The use of "first" and "second" in the description is solely for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, implicitly specifying the number of the indicated technical features, or implicitly specifying the order of the indicated technical features.

[0064] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, and connecting should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0065] An embodiment of the present invention provides a percutaneous lumbar disc puncture surgical navigation method and device based on multimodal images, wherein the percutaneous lumbar disc puncture surgical navigation method based on multimodal images includes: acquiring a preoperative spinal image set, wherein the preoperative spinal image set includes a preoperative thin-layer fat-suppressed MR image, a preoperative thin-layer non-fat-suppressed MR image and a preoperative CT image, the preoperative thin-layer fat-suppressed MR image is acquired based on a preset fat-suppressed sequence, and the preoperative thin-layer non-fat-suppressed MR image is acquired based on a preset non-fat-suppressed sequence; inputting the preoperative spinal image set into the multi-contrast multimodal image segmentation model for image segmentation, fusing the image segmentation results to obtain a first mask, wherein the first mask is used to indicate a spinal structure or a preoperative lesion target; inputting the preoperative spinal image set and the first mask into a multimodal image rigid-elastic hybrid registration model, performing a registration and alignment operation, and obtaining a first registered image. The first registration image includes the mutually registered anterior thin-layer fat-suppressed MR image, the preoperative thin-layer non-fat-suppressed MR image and the preoperative CT image; an intraoperative three-dimensional CBCT image is acquired, and the intraoperative three-dimensional CBCT image is input into the multi-contrast multi-modal image segmentation model for image segmentation to obtain a second mask, wherein the second mask is used to indicate the vertebral structure; the second mask, the intraoperative three-dimensional CBCT image and the first registration image are input into the multi-modal image rigid-elastic hybrid registration model, and the registration alignment operation is performed to obtain a second registration image, the second registration image includes the mutually registered first registration image and the intraoperative three-dimensional CBCT image; a three-dimensional visualization model of the spine is constructed based on the second registration image, and when the target lesion target marked on the three-dimensional visualization model of the spine is acquired, a puncture navigation path is generated based on the target lesion target. According to the technical solution of the embodiment of the present invention, image segmentation can be used to determine a mask representing the spinal structure or vertebral structure, and the mask can be used to automatically align the preoperative thin-layer MR image, CT image and intraoperative three-dimensional CBCT image, and the preoperative MR image introduces two sequences, fat suppression and non-fat suppression, so that the second registered image can clearly display the lumbar vertebral structure at the MR, CT and CBCT levels, improve the distinction between bones, flesh and various structures in the second registered image, improve the accuracy of the three-dimensional visualization model of the spine, and thus improve the accuracy of the navigation path and the feasibility of the operation.

[0066] The technical solutions of the embodiments of the present invention are further described below based on the accompanying drawings.

[0067] Reference Figure 1 , Figure 1A flowchart of a multimodal image-based percutaneous lumbar disc puncture surgical navigation method is provided in an embodiment of the present invention. The multimodal image-based percutaneous lumbar disc puncture surgical navigation method includes but is not limited to the following steps:

[0068] S10, obtaining a preoperative spinal image set, wherein the preoperative spinal image set includes a preoperative thin-layer fat-suppressed MR image, a preoperative thin-layer non-fat-suppressed MR image, and a preoperative CT image, wherein the preoperative thin-layer fat-suppressed MR image is obtained based on a preset fat-suppressed sequence, and the preoperative thin-layer non-fat-suppressed MR image is obtained based on a preset non-fat-suppressed sequence;

[0069] S20, inputting the preoperative spinal image set into a multi-contrast multimodal image segmentation model for image segmentation, fusing the image segmentation results to obtain a first mask, wherein the first mask is used to indicate the spinal structure or preoperative lesion target;

[0070] S30, inputting the preoperative spinal image set and the first mask into a multimodal image rigid-elastic hybrid registration model, performing a registration and alignment operation, and obtaining a first registered image, the first registered image including the preoperative thin-layer fat-suppressed MR image, the preoperative thin-layer non-fat-suppressed MR image, and the preoperative CT image that are registered with each other;

[0071] S40, acquiring an intraoperative three-dimensional CBCT image, and inputting the intraoperative three-dimensional CBCT image into a multi-contrast multi-modal image segmentation model for image segmentation to obtain a second mask, wherein the second mask is used to indicate a vertebral structure;

[0072] S50, inputting the second mask, the intraoperative 3D CBCT image, and the first registered image into a multimodal image rigid-elastic hybrid registration model, performing a registration and alignment operation, and obtaining a second registered image, where the second registered image includes the first registered image and the intraoperative 3D CBCT image that are registered with each other;

[0073] S60, constructing a three-dimensional visualization model of the spine based on the second registered image, and when a target lesion target marked on the three-dimensional visualization model of the spine is obtained, generating a puncture navigation path based on the target lesion target.

[0074] It should be noted that fat suppression refers to medical images captured using fat suppression technology, a technique well known to those skilled in the art. Preoperative thin-section fat-suppressed MR images are captured using a fat-suppressed sequence, while preoperative thin-section non-fat-suppressed MR images are captured using a non-fat-suppressed sequence. Both preoperative thin-section fat-suppressed and non-fat-suppressed MR images are captured with the patient in the same position to ensure spatial alignment of the images.

[0075] For example, 3D-SPACE non-fat suppression and fat suppression sequences can be used to obtain transverse imaging data from two different sequences. The specific parameters for the 3D-SPACE non-fat suppression sequence are: repetition time / echo time of 2800.0ms / 189.0ms, flip angle of 45°, field of view of 240×240mm, matrix of 320×320, slice thickness of 0.8mm, bandwidth of 579kHz, and final image resolution of 0.8×0.8×0.8mm; the specific parameters for the 3D-SPACE fat suppression sequence are: repetition time / echo time of 2800.0ms / 189.0ms, flip angle of 45°, field of view of 240×240mm, matrix of 320×320, slice thickness of 0.8mm, bandwidth of 579kHz, and final image resolution of 0.8×0.8×0.8mm. After acquisition, the data can be exported in DICOM format.

[0076] It should be noted that the preoperative CT image can use the direction perpendicular to the patient's long axis as the scanning baseline, and scan the patient layer by layer without intervals. The scanning field of view is 180mm, the layer thickness is 1mm, and the patient's transverse CT imaging data is collected and exported in DICOM format.

[0077] It should be noted that the intraoperative three-dimensional CBCT images include multiple spinal structures, and three-dimensional scanning can be performed using a three-dimensional C-arm to obtain three-dimensional CBCT imaging data of the L4 / L5 or L5 / S1 lumbar segments, and the data can be exported in DICOM format.

[0078] It should be noted that after obtaining the preoperative spinal image set, the spinal structure or preoperative lesion target in each image can be marked first. The spinal structure can be L4 vertebra, L4 / L5 intervertebral disc, L5 vertebra, L4 nerve root, L5 nerve root, dura mater, and skin. In this embodiment, the preoperative lesion target is the lesion tissue of intervertebral disc herniation.

[0079] It should be noted that the multi-contrast multimodal image segmentation model includes multiple segmentation branches, which simultaneously segment two images and fuse the results obtained after the segmentation, so that the first mask can represent the features respectively possessed by the two images, thereby representing the spinal structure or preoperative lesion target through the first mask. For example, after the image is annotated according to the above-mentioned spinal structure, the preoperative thin-layer non-fat-suppressed MR image and the preoperative CT image are simultaneously input into the multi-contrast multimodal image segmentation model, and the semantic representation of the spinal structure before fat suppression is obtained by segmenting the preoperative thin-layer non-fat-suppressed MR image, the semantic representation of the spinal structure under fat suppression is obtained by segmenting the preoperative thin-layer fat-suppressed MR image, and the semantic representation of the vertebral structure is obtained by segmenting the preoperative CT image. The first mask obtained after fusion can semantically represent the spinal structure or preoperative lesion target at the fat-suppressed, non-fat-suppressed and CT levels, thereby providing structural reference data for image registration and alignment.

[0080] It should be noted that, by aligning the MR image and the CT image under the first mask, the same spinal structure or lesion target in the MR image and the CT image can be aligned on the image, and the first mask of this embodiment introduces fat-suppressed MR images and non-fat-suppressed MR images, which can use the difference in images under fat-suppressed and non-fat-suppressed conditions to distinguish the flesh boundary, so that the first registered image obtained by the alignment can effectively distinguish between bones and flesh, thereby realizing bone and flesh visualization in the subsequently established three-dimensional visualization model of the spine, effectively improving the accuracy and feasibility of surgical navigation.

[0081] It should be noted that, after acquiring the intraoperative three-dimensional CBCT image during the operation, this embodiment again calls the multi-contrast multimodal image segmentation model to perform image segmentation, so that the second mask corresponding to the semantic information obtained by the segmentation can represent the vertebral structure in the plane dimension, and then calls the multimodal image rigid-elastic hybrid registration model to perform secondary registration with the first registration image, so as to achieve accurate registration and alignment of the plane vertebral structure and the three-dimensional spinal structure, thereby obtaining a clear MR / CT / three-dimensional CBCT image and improving the accuracy of the constructed three-dimensional visualization model of the spine.

[0082] It should be noted that after obtaining the clear registered MR / CT / 3D CBCT images, those skilled in the art are familiar with how to construct a 3D visualization model of the spine, and this embodiment does not elaborate on the model construction technology.

[0083] It should be noted that after the target lesion target is selected in the three-dimensional visualization model of the spine, the registration relationship based on the second registration image can locate the target lesion target in the plane and space, and automatically generate multiple available puncture navigation paths, thereby providing multiple path selections and completing surgical navigation based on the final selected target puncture path. This embodiment does not make any improvements to the specific path planning and surgical navigation process, and will not be elaborated here.

[0084] The technical solution of this embodiment enables the registration of preoperative high-resolution MR images, CT images, and intraoperative three-dimensional CBCT images using a multimodal image rigid-elastic hybrid registration model. The registered images can distinguish between bone and flesh based on fat-suppressed and non-fat-suppressed MR images. A multi-contrast, multimodal image segmentation model is used to fuse information from preoperative CT images, thin-layer fat-suppressed MR images, and thin-layer non-fat-suppressed MR images from the same source to extract high-contrast spinal structural features. This allows for precise segmentation of multiple spinal structures and lesion targets. Finally, a complete three-dimensional spinal model is constructed using multimodal registration methods, improving the accuracy and feasibility of surgical navigation.

[0085] In addition, in one embodiment, referring to Figure 3 The multi-contrast multimodal image segmentation model includes a first segmentation branch and a second segmentation branch, the first segmentation branch includes a first encoder and a first decoder, and the second segmentation branch includes a second encoder and a second decoder. Figure 2 , Figure 1 Step S20 shown specifically includes but is not limited to the following steps:

[0086] S21, inputting the preoperative thin-layer fat-suppressed MR image into a first segmentation branch, extracting first regional features of the preoperative thin-layer fat-suppressed MR image using a first encoder, extracting first semantic information of the first regional features using a first decoder, and determining a first segmentation result based on the first semantic information, wherein the first regional features are used to indicate a spinal structure or a preoperative lesion target;

[0087] S22, inputting the preoperative thin-slice non-fat-suppressed MR image into the first segmentation branch, and inputting the CT image into the second segmentation branch;

[0088] S23, extracting a second regional feature from the preoperative thin-slice non-fat-suppressed MR image using the first encoder, extracting second semantic information of the second regional feature using the first decoder, and determining a second segmentation result based on the second semantic information, wherein the second regional feature is used to indicate a spinal structure or a preoperative lesion target;

[0089] S24, inputting the second semantic information into the second segmentation branch, and extracting third regional features of the CT image through the second encoder;

[0090] S25, performing comparative learning based on the second semantic information and the third region feature to obtain a third segmentation result;

[0091] S26 , fusing the first segmentation result, the second segmentation result, and the third segmentation result to obtain a first mask.

[0092] It should be noted that the multi-contrast multimodal image segmentation model is Figure 3As shown, in this embodiment, the preoperative thin-layer fat-suppressed MR image is first input into the first segmentation branch. In the first decoder, the preoperative thin-layer fat-suppressed MR image is first subjected to a first encoding operation to realize image characterization. The first encoding operation includes convolution (conv), batch normalization (BN), and activation (ReLu) in sequence. Then, based on the first region feature, a second encoding operation, three third encoding operations, a fourth encoding operation, and a fifth encoding operation are performed in sequence. The output feature is used as the first region feature. The second encoding operation includes maximum pooling (Maxpool), conv+BN+ReLu in sequence, the third encoding operation includes Maxpool and (conv+BN+ReLu)×2 in sequence, the fourth encoding operation includes Maxpool, and the fifth encoding operation includes upsampling (Upsample).

[0093] It should be noted that after encoding, the first region features are input into the first decoder. The first decoding operation obtains semantic information. After the first decoded information is fused with the features of the corresponding encoding layer, a second decoding operation is performed to obtain further optimized semantic information. This operation is repeated again, and the final output semantic information is subjected to a third decoding operation to obtain the first semantic information. The first decoding operation includes upsampling, the second decoding operation includes (conv+BN+ReLu)×2, and the third decoding operation includes conv.

[0094] For example, Figure 3 As shown, encoding is performed five times in the first encoder. After the first decoding operation is performed for the first time, the fused encoding layer is the image feature of the fifth encoding operation, and the second decoding information is obtained. After the first decoding operation is performed on the second decoding information, the fused encoding layer is the image feature of the fourth encoding operation, and so on, until the image feature of the first encoding operation is fused. Through the image segmentation operation of this embodiment, the encoding features of the same layer can be fused during decoding, and features with the same semantics can be continuously extracted, so that the first segmentation result can more clearly reflect the marked spinal structure or preoperative lesion target.

[0095] It should be noted that the process of extracting the second segmentation result of the preoperative thin-layer non-fat-suppressed MR image through the first segmentation branch can refer to the extraction process of the first segmentation result described above, and will not be repeated here.

[0096] It should be noted that after completing the segmentation of the preoperative thin-layer fat-suppressed MR image, the first segmentation result is saved as reference information for distinguishing bone and flesh boundaries. The preoperative thin-layer non-fat-suppressed MR image and the preoperative CT image are then simultaneously input into the multi-contrast multimodal image segmentation model for contrastive learning and segmentation. The second semantic information of the preoperative thin-layer non-fat-suppressed MR image is extracted by the first decoder and then input into the Attention-aware Fusion Module (AFM). This guides the second segmentation branch to focus on the vertebral structure when segmenting the preoperative CT image, emphasizing the extraction of regional features of the CT vertebral structure, so that the third segmentation result can clearly represent the vertebral structure.

[0097] It should be noted that after obtaining the first segmentation result, the second segmentation result and the third segmentation result, the three segmentation results can be combined with a preset loss function and then fused. The specific fusion process will not be described in detail here.

[0098] In addition, in one embodiment, referring to Figure 3 , the second decoder includes multiple attention perception fusion modules, referring to Figure 2 Step S25 specifically includes but is not limited to the following steps:

[0099] S251, inputting the second semantic information into the attention perception fusion module to obtain an attention feature, upsampling the attention feature and fusing it into the third region feature;

[0100] S252, extracting third semantic information of the third region feature, and determining a third segmentation result based on the third semantic information and a preset contrast loss function, wherein the third region feature is used to indicate a vertebral structure;

[0101] Among them, the contrast loss function satisfies the following conditions: , is the cross entropy loss function of the first segmentation branch, is the cross entropy loss function of the second segmentation branch, and is the preset magnification factor, is the contrast loss function, is the total loss function of the multi-contrast multimodal image segmentation model.

[0102] It should be noted that after obtaining the second semantic information, refer to Figure 3 As shown, the second semantic information of the corresponding scale is input into the AFM of the same scale. The number of AFMs is the same as the number of encoding times, so that when the features of the same scale are decoded, the attention features can be extracted through the attention mechanism, and the attention features are combined with the third region features, so that the extracted semantic information can be more focused on the vertebral structure.

[0103] It should be noted that the contrast loss function makes points of the same category near the vertebral structure similar in feature space, enhancing the features of the vertebral structure region and thus improving the segmentation accuracy of the vertebral structure. According to the above method, the thin-layer non-fat-suppressed MR image and thin-layer fat-suppressed MR image segmentation group data are input into the multi-contrast multimodal image segmentation model for automatic segmentation to enhance the features of the nerve root structure region and improve the segmentation accuracy of the nerve root structure. During testing, the segmentation results output by the thin-layer fat-suppressed MR image segmentation branch are saved. Finally, the segmentation results output by the thin-layer fat-suppressed MR image and CT image are automatically spliced, and the spliced ​​segmentation result is used as the final segmentation result.

[0104] It should be noted that during the multi-structure segmentation of multimodal images, the boundaries between the vertebral structure and muscles and ligaments, and between the nerve roots and the dura mater, are unclear, resulting in poor discrimination and discontinuous segmentation. To address these issues, this example employs a contrast loss function to make points of the same category near the vertebral structure and nerve roots similar in feature space, aiming to improve the discrimination between the vertebral structure and nerve root regions.

[0105] In addition, in one embodiment, referring to Figure 3 , the encoding times of the first encoder and the second encoder are N times, the decoding times of the first decoder and the second decoder are N times, the number of attention perception fusion modules is N, N is a positive integer greater than 2, refer to Figure 2 Step S251 specifically includes but is not limited to the following steps:

[0106] S2511, inputting the second semantic information obtained by the first decoding of the first decoder into the first attention perception fusion module, performing an attention extraction operation, and obtaining a first attention feature;

[0107] S2512, upsampling the first attention feature, fusing it with the third region feature obtained by the Nth encoding by the second encoder, and then inputting it into the second attention perception fusion module;

[0108] S2513, inputting the second semantic information obtained by the second decoding of the first decoder into the second attention perception fusion module, performing an attention extraction operation, and obtaining a second attention feature;

[0109] S2514, when the Nth attention feature is obtained, it is fused with the third region feature obtained by the first encoding of the second encoder to obtain the attention feature.

[0110] It should be noted that the first segmentation channel of this embodiment is based on the preoperative thin-layer non-fat-suppressed MR image, which records the spinal structure. The second segmentation channel is based on the preoperative CT image. Since the CT image is obtained based on a planar scan, this embodiment uses the attention mechanism to enhance the spinal structure features of the MR image in the three-dimensional space to the planar vertebral structure, so that the third semantic information can multiple times enhance the vertebral structure of the same scale during the decoding process, thereby improving the accuracy of the third segmentation result.

[0111] like Figure 3 As shown, N=5, the second semantic information obtained by the first decoding is input into the first AFM, the first attention feature is extracted, and the first attention feature is upsampled and fused with the third region feature of the fifth encoding, so that the features of the vertebral structure are enhanced in the third region feature, and multiple iterations are performed in this way.

[0112] In addition, in one embodiment, referring to Figure 4 The attention perception fusion module includes a channel attention module and a spatial attention module. The channel attention module includes the first convolution block, the first BN block and the first ReLu layer, the first maximum pooling layer, the first average pooling layer, the multi-layer perceptron and the first Sigmoid function. The spatial attention module includes the second convolution block, the second BN block, the third convolution block and the third BN block. Figure 2 The attention extraction operation in steps S2511 and S2513 includes:

[0113] S71, determining a segmented image pair, where the segmented image pair includes a preoperative thin-slice non-fat-suppressed MR image and a preoperative CT image, or includes a first registered image and an intraoperative three-dimensional CBCT image;

[0114] S72, obtaining a first transformed feature by sequentially processing the second semantic information through a first convolution block, a first BN block, and a first ReLu layer;

[0115] S73, inputting the first transformed features into the first maximum pooling layer and the first average pooling layer respectively, fusing the outputs of the first maximum pooling layer and the first average pooling layer, and then passing them through a multi-layer perceptron and a first sigmoid function in sequence to obtain a channel attention map;

[0116] S74, sequentially performing the third region feature through a second convolution block and a second BN block to obtain a second transformed feature;

[0117] S75: Input the first transformed feature into the spatial attention module, fuse the first transformed feature with the second transformed feature, and then pass it through the third convolution block and the third BN block in sequence to obtain a spatial attention map;

[0118] S76, obtain attention features based on the spatial attention map and the channel attention map.

[0119] It should be noted that this embodiment uses the segmented image pair of a preoperative thin-layer non-fat-suppressed MR image and a preoperative CT image as an example to illustrate the principle. The segmented image pair of the first registration image and the intraoperative three-dimensional CBCT image is similar, and will not be repeated later.

[0120] It should be noted that AFM includes a channel attention module and a spatial attention module for extracting regional feature information of feature structures. The second semantic information of this embodiment is the image feature of the thin layer non-fat-suppressed MR image, which is input into the first convolution block, the first BN block and the first ReLu layer. Figure 4 The Conv+BN+ReLu operation shown in the figure can perform a convolution operation on the second semantic information to achieve transformation, obtain the first transformed feature, and then input it into the first maximum pooling layer and the first average pooling layer respectively. After performing the corresponding pooling operation, the two pooling results are fused and then passed through the multi-layer perceptron and the first Sigmoid function to obtain the channel attention map.

[0121] It should be noted that the third regional feature of this embodiment is the vertebral structure feature obtained after CT image encoding, which is input into the second convolution block and the second BN block to obtain the second transformation feature. The second convolution block and the second BN block are Figure 4 The Conv+BN operation shown in the figure is performed, and then the first transformed feature is fused with the second transformed feature and the Conv+BN operation is performed again to obtain the spatial attention map. Finally, the spatial attention map and the channel attention map are fused and the fused attention feature is obtained by the ReLu operation.

[0122] For example, the second semantic information is the thin-layer non-fat-suppressed MR image feature, the third regional feature is the vertebral structure regional feature of the CT image, and the i-th thin-layer non-fat-suppressed MR image feature After the channel attention module, the transformed non-fat-suppressed MR image features are output and channel attention map , After the spatial attention module, the spatial attention map is output , where i is a positive integer, is the non-fat-suppressed MR image feature, is the transformed non-fat-suppressed MR image feature, is the channel attention map, is the spatial attention map. The vertebral structural features of the i-th CT image are represented by For example, after conv and BN, the transformed vertebral structure features of the CT image are obtained , and 、 Multiply point by point to get attention perception features , 、 and Add point by point and pass through ReLu activation function to get the fused attention feature These features refer to the multi-structural features of the spine in MR images, with a focus on the enhanced regional features of the vertebral structure.

[0123] In the above example, the thin layer non-fat suppressed MR image features After convolution, batch normalization, and Relu activation function in the channel attention module, the transformed thin-layer non-fat-suppressed MR image features are obtained. , and finally output the channel attention map , expressed as:

[0124] ;

[0125] ,in, is the maximum pooling, is average pooling, and MLP is multi-layer perceptron.

[0126] In the above example, Input the spatial attention module to generate a spatial attention map, which can be expressed as follows:

[0127] ;

[0128] The attention-aware feature is expressed as: .

[0129] In addition, in one embodiment, referring to Figure 5 The multimodal image rigid-elastic hybrid registration model includes multiple affine matrix estimation modules, affine elastic fusion modules and local rigid constraint modules. Figure 2 The registration and alignment operations in step S30 and step S50 specifically include but are not limited to the following steps:

[0130] S81, inputting multiple images to be registered into a multiple affine matrix estimation module to obtain multi-scale image features and initial deformation fields of each image to be registered, wherein the images to be registered include a preoperative thin-layer fat-suppressed MR image, a preoperative thin-layer non-fat-suppressed MR image, and a preoperative CT image, or the images to be registered include a first registration image and an intraoperative 3D CBCT image;

[0131] S82, based on the same image to be registered, inputting the corresponding multi-scale image features and the initial deformation field into the affine elastic fusion module, performing deformation field splicing based on the initial deformation field to obtain a rigid elastic deformation field, fusing all the rigid elastic deformation fields and outputting the initial registered image;

[0132] S83, input the initial registration image and the registration mask into the local rigid constraint module, and correct the rigid-elastic deformation field of the initial registration image through the local rigid constraint module to obtain the target registration image, wherein, when the image to be registered includes a preoperative spinal image set, the registration mask is a first mask, or when the image to be registered includes an intraoperative three-dimensional CBCT image, the registration mask is a second mask.

[0133] It should be noted that this embodiment uses the images to be registered as the preoperative thin-layer fat-suppressed MR image, the preoperative thin-layer non-fat-suppressed MR image and the preoperative CT image as examples for illustrative explanation. The images to be registered include the first registration image and the intraoperative three-dimensional CBCT image, which can be obtained similarly and will not be repeated later.

[0134] It should be noted that the rigid-elastic hybrid registration model based on multimodal spinal images consists of a deep network framework for rigid-elastic registration of multimodal spinal images (including a multiple affine matrix estimation module and an affine elastic fusion module) and a local rigid constraint module. Its working principle is to first extract the multi-scale image features of the image to be registered through the multiple affine matrix estimation module, and use the rigid transformation estimation parameters to obtain the initial deformation field of each scale. Then, through the affine elastic fusion module, the image features of each scale are spliced ​​with the initial deformation field of each structure obtained by the rigid transformation estimation, the rigid-elastic deformation field is calculated, and the registered image is output. In addition, this embodiment adds a local rigid constraint module to provide strict local rigid constraints on the vertebral structure to reduce the possibility of unreasonable elastic deformation of the vertebral structure during the spinal image registration process.

[0135] In addition, in one embodiment, referring to Figure 5 The multiple affine matrix estimation module includes four consecutively arranged fourth convolution blocks, flat layers, and rigid transformation estimation modules. The rigid transformation estimation module includes multiple dense layers and multiple affine matrices. The dense layers are uniquely corresponding to the affine matrices. Figure 2 Step S81 also specifically includes but is not limited to the following steps:

[0136] S811, sequentially subjecting the image to be registered to convolution processing of each fourth convolution block to obtain multi-scale image features;

[0137] S812, after the multi-scale image features are expanded based on each scale through the flattening layer, the image features at each scale are input into a dense layer of the rigid transformation estimation module, and the rigid deformation field is calculated based on preset rigid transformation estimation parameters;

[0138] S813: Perform weighted averaging on all rigid deformation fields to obtain an initial deformation field.

[0139] It should be noted that the multi-affine matrix estimation module is used to extract multi-scale image features from the images to be registered and simultaneously estimate the affine matrices of each structure and the overall structure. It consists of four fourth-order convolutional blocks, a flattening layer, and N dense layers (N is the number of multi-structures, corresponding to the L4 vertebra, L4 / L5 intervertebral disc, L5 vertebra, L4 nerve root, L5 nerve root, dura mater, ilium, iliac artery and branches, inferior vena cava and branches, psoas major muscle, skin, and lesion target, respectively, N1, N2, N.3..., N9, N10, N11). Each convolutional block consists of a 3D convolutional layer and a LeakyReLU layer. These four convolutional blocks serve as an image feature extraction module, extracting multi-scale image encoding features (Fu represents the smallest high-level image feature, while F1 represents the largest low-level image representation) from the moving image Imoving and the fixed image Ifixed to be registered.

[0140] Exemplarily, the specific implementation of obtaining the initial deformation field in this embodiment is as follows: a multi-resolution network is constructed using a multi-layer convolution + pooling structure; the multimodal image to be registered is input into the network for multi-scale image feature extraction. A flat layer and N pooling layers constitute a rigid transformation estimation module, which is used to calculate the initial deformation field of each scale. The function of the flat layer is to project high-level image features into N affine matrices. The N dense layers are used to predict the affine matrix Tn of each structure, where n is a positive integer. The affine matrix of the overall structure is obtained through global optimization. The N affine dense layers convert the IV affine matrix into N features, and connect them with the u scale to obtain the mixed feature Hu for use in the subsequent affine elastic fusion module. In addition, the registration mask and Smoving of each structure of the moving image to be registered, as well as the registration mask and Sfixed of each structure of the moving image to be registered will be used to weakly supervise the multiple affine matrix estimation module to obtain a more accurate affine matrix for each structure. The specific implementation is as follows: a fully connected network is used to estimate the rigid transformation parameters of each structure, and the cost function adopts the DICE cost of the segmented area corresponding to each rigid area; the rigid deformation field of N features is calculated through the rigid transformation parameters; the rigid deformation field of N features is weightedly averaged by distance correlation to obtain the initial deformation field with the lowest resolution; the initial deformation field of each structure is calculated through upsampling for subsequent rigid / rigid deformation field estimation.

[0141] In addition, in one embodiment, referring to Figure 5 The affine elastic fusion module includes 4 combination units, three sixth convolution blocks and integrated blocks. The combination unit includes an upsampling layer and a fifth convolution block. Figure 2 Step S82 also specifically includes but is not limited to the following steps:

[0142] S821, processing the image features of each scale through the four combination units in sequence to obtain the structural image features of each scale;

[0143] S822, combining the structural image features of the corresponding scale with the initial deformation field to obtain a rigid-elastic deformation field;

[0144] S823: Input each rigid-elastic deformation field into the sixth convolution block for convolution processing, and then input it into the integrated block for fusion to obtain an initial registration image.

[0145] It should be noted that the affine elastic fusion module consists of 4 combination units, three sixth convolution blocks and an integrated block. The specific structure is referenced Figure 5 As shown. The affine elastic fusion module is used to estimate the rigid-elastic deformation field. For each combination unit, the input is the u-scale encoding feature Fu and the mixed feature u, and the output is the u-1-scale combined feature Ou-1. Three convolution blocks and one integrated block are used to fuse multi-resolution rigid-elastic information and estimate the rigid-elastic deformation field. In addition, the registration mask and Smoving of each structure of the moving image to be registered, as well as the registration mask and Sfixed of each structure of the moving image to be registered are used to weakly supervise the affine elastic fusion module, aiming to obtain the optimal rigid-elastic deformation field. Specific implementation: A multi-resolution network is constructed with a multi-layer convolution + upsampling structure; the image features of each structure output by the feature extraction module are spliced ​​with the initial deformation field of each structure output by the rigid transformation estimation module to output the rigid-elastic deformation field; a spatial transformation network is used to calculate the registered image, and the mutual information cost and the DICE cost of the segmented area corresponding to each rigid area are calculated to supervise the network.

[0146] In addition, in one embodiment, referring to Figure 2 , refer to Figure 2 Step S83 also specifically includes but is not limited to the following steps:

[0147] S831, performing least squares regression calculation on the rigid-elastic deformation field based on the registration mask to obtain a reference deformation field;

[0148] S832, constructing a rigid constraint cost function based on the error between the rigid elastic deformation field and the reference deformation field;

[0149] S833: Modify the initial registration image based on the rigid constraint cost function to obtain the target registration image.

[0150] It should be noted that this embodiment is applied to the specific scenario of lumbar intervertebral discs. During the registration process, the vertebral structure is prone to unreasonable elastic deformation, which can affect the accuracy of the registered image. Therefore, this embodiment introduces a local rigid constraint module to ensure appropriate rigid deformation of the bony structure and avoid unreasonable elastic deformation of the vertebral structure during spinal image registration. First, using the registration mask of the segmented image, a least squares regression is performed on the rigid-elastic deformation field of each vertebral structure's rigid region (the output of the rigid-elastic deformation transformation estimation module) to obtain its corresponding rigid deformation field. The error between the original rigid-elastic deformation field and the obtained ideal rigid deformation field is then calculated to form a local rigid constraint cost function. This error is then propagated back to correct the network parameters.

[0151] like Figure 6 As shown, Figure 6 This is a structural diagram of a multimodal image-based percutaneous lumbar disc puncture surgical navigation device according to an embodiment of the present invention. The present invention also provides a multimodal image-based percutaneous lumbar disc puncture surgical navigation device, comprising:

[0152] The processor 601 may be implemented as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0153] The memory 602 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 602 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 602 and are called by the processor 601 to execute the multimodal image-based percutaneous lumbar disc puncture surgical navigation method of the embodiments of this application.

[0154] Input / output interface 603, used to implement information input and output;

[0155] Communication interface 604, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0156] Bus 605 , which transmits information between various components of the device (e.g., processor 601 , memory 602 , input / output interface 603 , and communication interface 604 );

[0157] The processor 601 , the memory 602 , the input / output interface 603 and the communication interface 604 are connected to each other in communication within the device via a bus 605 .

[0158] An embodiment of the present application further provides an electronic device, comprising the multimodal image-based percutaneous lumbar disc puncture surgical navigation device as described above.

[0159] An embodiment of the present application also provides a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned multimodal image-based percutaneous lumbar disc puncture surgical navigation method.

[0160] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory optionally includes a memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of the above-mentioned networks include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof. The device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and are located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.

[0161] Those skilled in the art will appreciate that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, or any suitable combination thereof. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0162] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above implementation. Those skilled in the art can also make various equivalent modifications or substitutions under the shared conditions that do not violate the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.

Claims

1. A multimodal image-based percutaneous lumbar disc puncture surgical navigation method, characterized in that: Applied to a surgical navigation system, the surgical navigation system is pre-configured with a multi-contrast multimodal image segmentation model and a multimodal image rigid-elastic hybrid registration model. The multimodal image-based percutaneous lumbar disc puncture surgical navigation method includes: Acquiring a preoperative spinal image set, wherein the preoperative spinal image set includes a preoperative thin-layer fat-suppressed MR image, a preoperative thin-layer non-fat-suppressed MR image, and a preoperative CT image, wherein the preoperative thin-layer fat-suppressed MR image is acquired based on a preset fat-suppressed sequence, and the preoperative thin-layer non-fat-suppressed MR image is acquired based on a preset non-fat-suppressed sequence; Inputting the preoperative spinal image set into the multi-contrast multimodal image segmentation model for image segmentation, and fusing the image segmentation results to obtain a first mask, wherein the first mask is used to indicate the spinal structure or preoperative lesion target; Inputting the preoperative spinal image set and the first mask into a multimodal image rigid-elastic hybrid registration model, performing a registration and alignment operation, and obtaining a first registered image, wherein the first registered image includes the preoperative thin-layer fat-suppressed MR image, the preoperative thin-layer non-fat-suppressed MR image, and the preoperative CT image that are registered with each other; Acquiring an intraoperative three-dimensional CBCT image, and inputting the intraoperative three-dimensional CBCT image into the multi-contrast multimodal image segmentation model for image segmentation to obtain a second mask, wherein the second mask is used to indicate a vertebral structure; Inputting the second mask, the intraoperative 3D CBCT image, and the first registered image into the multimodal image rigid-elastic hybrid registration model, performing the registration and alignment operation, and obtaining a second registered image, wherein the second registered image includes the first registered image and the intraoperative 3D CBCT image that are registered with each other; A three-dimensional visualization model of the spine is constructed based on the second registered image, and when a target lesion target marked on the three-dimensional visualization model of the spine is acquired, a puncture navigation path is generated based on the target lesion target.

2. The multimodal image-based percutaneous lumbar disc puncture surgical navigation method according to claim 1, characterized in that: The multi-contrast multi-modal image segmentation model includes a first segmentation branch and a second segmentation branch, the first segmentation branch includes a first encoder and a first decoder, and the second segmentation branch includes a second encoder and a second decoder. The preoperative spinal image set is input into the multi-contrast multi-modal image segmentation model for image segmentation, and the image segmentation results are fused to obtain a first mask, including: Inputting the preoperative thin-layer fat-suppressed MR image into the first segmentation branch, extracting first regional features of the preoperative thin-layer fat-suppressed MR image using the first encoder, extracting first semantic information of the first regional features using the first decoder, and determining a first segmentation result based on the first semantic information, wherein the first regional features are used to indicate the spinal structure or preoperative lesion target; Inputting the preoperative thin-layer non-fat-suppressed MR image into the first segmentation branch, and inputting the CT image into the second segmentation branch; extracting a second regional feature of the preoperative thin-layer non-fat-suppressed MR image by the first encoder, extracting second semantic information of the second regional feature by the first decoder, and determining a second segmentation result based on the second semantic information, wherein the second regional feature is used to indicate the spinal structure or the preoperative lesion target; inputting the second semantic information into the second segmentation branch, and extracting a third region feature of the CT image through the second encoder; Performing comparative learning based on the second semantic information and the third region feature to obtain a third segmentation result; The first segmentation result, the second segmentation result, and the third segmentation result are fused to obtain the first mask.

3. The multimodal image-based percutaneous lumbar disc puncture surgical navigation method according to claim 2, characterized in that: The second decoder includes a plurality of attention perception fusion modules, and the third segmentation result is obtained by performing comparative learning based on the second semantic information and the third region feature, including: Inputting the second semantic information into the attention perception fusion module to obtain an attention feature, and upsampling the attention feature and fusing it into the third region feature; extracting third semantic information of the third region feature, and determining a third segmentation result based on the third semantic information and a preset contrast loss function, wherein the third region feature is used to indicate a vertebral structure; The contrast loss function satisfies the following conditions: , is the cross entropy loss function of the first segmentation branch, is the cross entropy loss function of the second segmentation branch, and is the preset magnification factor, is the contrast loss function, is the total loss function of the multi-contrast multimodal image segmentation model.

4. The multimodal image-based percutaneous lumbar disc puncture surgical navigation method according to claim 3, characterized in that: The first encoder and the second encoder perform encoding N times, the first decoder and the second decoder perform decoding N times, the number of the attention perception fusion modules is N, N is a positive integer greater than 2, the second semantic information is input into the attention perception fusion module to obtain an attention feature, and the attention feature is up-sampled and fused into the third region feature, including: Inputting the second semantic information obtained by the first decoder for the first time into the first attention perception fusion module, performing an attention extraction operation, and obtaining the first attention feature; Upsampling the first attention feature, fusing it with the third region feature obtained by the Nth encoding of the second encoder, and then inputting it into the second attention perception fusion module; Inputting the second semantic information obtained by the second decoding of the first decoder into the second attention perception fusion module, performing the attention extraction operation, and obtaining the second attention feature; When the Nth attention feature is obtained, it is fused with the third region feature obtained by the first encoding of the second encoder to obtain the attention feature.

5. The multimodal image-based percutaneous lumbar disc puncture surgical navigation method according to claim 4, characterized in that: The attention perception fusion module includes a channel attention module and a spatial attention module. The channel attention module includes a first convolution block, a first BN block and a first ReLu layer, a first maximum pooling layer, a first average pooling layer, a multi-layer perceptron and a first Sigmoid function. The spatial attention module includes a second convolution block, a second BN block, a third convolution block and a third BN block. The attention extraction operation includes: Determining a segmented image pair, the segmented image pair including the preoperative thin-slice non-fat-suppressed MR image and the preoperative CT image, or including the first registered image and the intraoperative three-dimensional CBCT image; Obtaining a first transformed feature based on the second semantic information by sequentially processing the first convolution block, the first BN block, and the first ReLu layer; Inputting the first transformed features into the first maximum pooling layer and the first average pooling layer respectively, fusing the outputs of the first maximum pooling layer and the first average pooling layer, and then sequentially passing the outputs through the multi-layer perceptron and the first sigmoid function to obtain a channel attention map; Obtain a second transformed feature based on the third region feature passing through the second convolution block and the second BN block in sequence; Inputting the first transformed feature into the spatial attention module, fusing the first transformed feature with the second transformed feature, and then sequentially passing through the third convolution block and the third BN block to obtain a spatial attention map; The attention feature is obtained based on the spatial attention map and the channel attention map.

6. The multimodal image-based percutaneous lumbar disc puncture surgical navigation method according to claim 1, characterized in that: The multimodal image rigid-elastic hybrid registration model includes a multiple affine matrix estimation module, an affine elastic fusion module and a local rigid constraint module. The registration and alignment operation includes: Inputting a plurality of images to be registered into the multiple affine matrix estimation module to obtain multi-scale image features and initial deformation fields of each of the images to be registered, wherein the images to be registered include the preoperative thin-layer fat-suppressed MR image, the preoperative thin-layer non-fat-suppressed MR image, and the preoperative CT image, or the images to be registered include the first registration image and the intraoperative three-dimensional CBCT image; Based on the same image to be registered, the corresponding multi-scale image features and the initial deformation field are input into the affine elastic fusion module, deformation field splicing is performed based on the initial deformation field to obtain a rigid elastic deformation field, and all the rigid elastic deformation fields are fused to output an initial registered image; The initial registration image and the registration mask are input into the local rigid constraint module, and the rigid-elastic deformation field of the initial registration image is corrected by the local rigid constraint module to obtain a target registration image, wherein, when the image to be registered includes the preoperative spinal image set, the registration mask is the first mask, or when the image to be registered includes the intraoperative three-dimensional CBCT image, the registration mask is the second mask.

7. The multimodal image-based percutaneous lumbar disc puncture surgical navigation method according to claim 6, characterized in that: The multiple affine matrix estimation module includes four consecutively arranged fourth convolution blocks, a flat layer, and a rigid transformation estimation module. The rigid transformation estimation module includes multiple dense layers and multiple affine matrices. The dense layers uniquely correspond to the affine matrices. The multi-scale image features and initial deformation fields of each of the images to be registered are obtained, including: Subjecting the image to be registered to convolution processing of each of the fourth convolution blocks in sequence to obtain the multi-scale image features; After the multi-scale image features are expanded based on each scale through the flattening layer, the image features of each scale are input into one of the dense layers of the rigid transformation estimation module, and a rigid deformation field is calculated based on preset rigid transformation estimation parameters; The initial deformation field is obtained by performing weighted averaging on all the rigid deformation fields.

8. The multimodal image-based percutaneous lumbar disc puncture surgical navigation method according to claim 7, characterized in that: The affine elastic fusion module includes four combination units, three sixth convolution blocks and an integrated block. The combination unit includes an upsampling layer and a fifth convolution block. The deformation field is spliced ​​based on the initial deformation field to obtain a rigid elastic deformation field. All the rigid elastic deformation fields are fused to output an initial registration image, including: The image features of each scale are sequentially processed by the four combination units to obtain the structural image features of each scale; Splicing the structural image features of the corresponding scale with the initial deformation field to obtain a rigid elastic deformation field; Each rigid-elastic deformation field is input into the sixth convolution block for convolution processing, and then input into the integrated block for fusion to obtain the initial registration image.

9. The multimodal image-based percutaneous lumbar disc puncture surgical navigation method according to claim 8, characterized in that: The method of correcting the rigid-elastic deformation field of the initial registration image by the local rigid constraint module to obtain a target registration image includes: performing least squares regression calculation on the rigid-elastic deformation field based on the registration mask to obtain a reference deformation field; constructing a rigid constraint cost function based on an error between the rigid elastic deformation field and the reference deformation field; The initial registration image is modified based on the rigid constraint cost function to obtain the target registration image.

10. A multimodal image-based percutaneous lumbar disc puncture surgical navigation device, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to perform the multimodal image-based percutaneous lumbar disc puncture surgical navigation method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Spine registration method, device and equipment and computer storage medium

    CN113538533A

  • Deep learning point cloud lumbar registration method for minimally invasive spine surgery navigation

    CN115049709A