A Cross-Modal Spinal Image Registration Method, System and Device

Through the image transformation method based on deep reinforcement learning, the fitted function curve and spatial transformation network are used to solve the problem of modal differences and noise interference in cross-modal data registration, and an efficient and adaptive image registration effect is achieved.

CN116664634BActive Publication Date: 2025-06-27BEIJING CHAOYANG HOSPITAL CAPITAL MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310762419.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-27
Publication Date
2025-06-27
Estimated Expiration
2043-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively overcome the modal differences in cross-modal data, resulting in unsatisfactory new data registration results, and traditional methods have shortcomings in dealing with noise interference.

Method used

The image transformation method based on deep reinforcement learning is adopted to adaptively realize image registration by fitting the function curve and spatial transformation network, reducing dependence on the data set.

Benefits of technology

It realizes efficient registration of cross-modal images, reduces noise interference, and can adaptively process new data without relying on a large amount of training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116664634B_ABST
    Figure CN116664634B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-modal spinal image registration method, system and device, belonging to the technical field of image registration. By continuously measuring the evaluation indexes of the sampling action and the spatial transformation network, the present invention avoids excessive rejection of features by the sampling action, so as to achieve good removal of irrelevant features in the image registration process without relying on the prior distribution provided by the data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image registration, and more specifically, to a cross-modal spinal image registration method, system and device. Background Art

[0002] Image registration is a method of aligning different images spatially, aiming to map them into a common reference framework for further comparison and analysis. In medical image diagnosis, image registration technology is widely used for aligning different modality images, such as aligning MRI images, CT images and X-ray images.

[0003] Traditional image registration methods usually rely on rigid or non-rigid models. Among them, the rigid model assumes that the geometric shapes between images remain unchanged, while the non-rigid model can handle the situation of shape variation. Traditional non-rigid image registration methods usually rely on a large amount of data sets to train and optimize algorithms. However, this method has some problems, such as it takes a lot of time and labor costs to collect and label data, and it is difficult to solve the registration problem of new data.

[0004] In recent years, image registration methods based on deep learning have received extensive attention. These methods transform the image registration problem into an optimization problem and use neural networks to learn the corresponding relationships between images, so as to achieve accurate registration. However, this method requires a large amount of training data to train the neural network. Even unsupervised learning also requires unlabeled data sets to learn data features. There are also some practical application problems, such as the registration of new data and noise interference.

[0005] Other registration methods rely on the selection of feature points. Whether it is the feature point detection and matching based on Scale-Invariant Feature Transform (SIFT) or other improved algorithms, when facing cross-modal tasks, due to too many differences between different modalities, the matching effect is often not ideal.

[0006] Therefore, how to overcome the modality differences of cross-modal data, reduce the noise interference of new data registration, and provide a new cross-modal spinal image registration method, system and device are urgent problems to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a cross-modal spinal image registration method, system and device.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] The present invention first discloses a cross-modal spinal image registration method, including the following steps:

[0010] Image acquisition step:

[0011] Obtain preoperative CT images and corresponding intraoperative optical images;

[0012] Image preprocessing step:

[0013] Invert the pixel values in the preoperative CT images to obtain inverted CT images;

[0014] Based on the intraoperative optical images, obtain the corresponding intraoperative grayscale images;

[0015] Image transformation step based on deep reinforcement learning idea:

[0016] Input the inverted CT images and intraoperative grayscale images into a network model based on the reinforcement learning idea, use the proportion of the same number of pixels in the total pixels as the evaluation index, and use the fitting function curve and the spatial transformation network as two actions in the action space;

[0017] Take the inverted CT images as the images to be executed, execute the two actions in the action space, and obtain the sampled action CT images and spatial transformation CT images after execution;

[0018] Perform similarity measurement calculations based on the evaluation index for the sampled action CT images and the spatial transformation CT images respectively with the intraoperative grayscale images, and take the sampled action CT image or the spatial transformation CT image with a larger index as the new image to be executed;

[0019] Execute the two actions in the action space again for the new image to be executed, and perform similarity measurement calculations based on the evaluation index for the execution results again with the intraoperative grayscale images until the evaluation index of the execution results meets the set threshold, and take the execution results that meet the set threshold as the final CT images to be registered; the execution results include the sampled action CT images or the spatial transformation CT images after executing the two actions in the action space;

[0020] Image fusion step:

[0021] Fuse the CT images to be registered and the intraoperative grayscale images to obtain the registered fused images.

[0022] Preferably, in the image acquisition step, obtaining the preoperative CT images specifically includes:

[0023] Synthesize the taken preoperative dicom files through mimcs to obtain preoperative CT images.

[0024] Preferably, the image preprocessing step further includes adjusting the sizes of the inverted CT images and the intraoperative grayscale images to the same size.

[0025] Preferably, acquiring the corresponding intraoperative grayscale image based on the intraoperative optical image specifically includes extracting the corresponding intraoperative grayscale image from the intraoperative optical image through a trained Unet neural network.

[0026] Preferably, in the image transformation step based on the deep reinforcement learning concept, obtaining the sampled action CT image specifically includes the following steps:

[0027] Perform a fifth-order polynomial curve fitting on the horizontal and vertical coordinates of the pixel points whose pixel values ​​in the inverted CT image are greater than a set threshold value to obtain a function curve;

[0028] The distances from all pixels whose pixel values ​​are not 0 to the corresponding horizontal coordinates on the function curve are calculated, and the pixels exceeding the set distance threshold are eliminated to obtain a sampled motion CT image.

[0029] Preferably, in the image transformation step based on the deep reinforcement learning concept, obtaining the final CT image to be registered also includes:

[0030] The new image to be executed executes the two actions in the action space again, and the execution result is again compared with the intraoperative grayscale image for similarity measurement based on the evaluation index, until the number of executions meets the set threshold, and the sampled action CT image or spatial transformation CT image with a higher evaluation index in the last execution result is used as the final CT image to be registered.

[0031] Preferably, in the image transformation step based on the deep reinforcement learning concept, obtaining a spatially transformed CT image specifically includes the following steps:

[0032] The inverted CT image is passed through the spatial transformation network for 50 iterations to obtain the spatially transformed CT image.

[0033] Preferably, the loss function of the spatial transformation network includes normalized cross-correlation loss, gradient smoothing loss and information entropy loss.

[0034] The present invention also discloses a cross-modal spinal image registration system, wherein the system includes a computer program, and when the computer program is executed, the cross-modal spinal image registration method described in any one of the above items of the present invention is implemented.

[0035] The present invention also discloses a cross-modal spinal image registration device, which includes a computer system. The computer system is used for the cross-modal spinal image registration method described in any one of the above items of the present invention.

[0036] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a cross-modal spinal image registration method, system and device, which have the following beneficial effects:

[0037] The present invention uses a sampling function based on a fitted curve to remove irrelevant features such as ribs in an image. At the same time, the present invention also utilizes a non-rigid registration method based on a spatial transformation network and improves its loss function, increasing the information entropy loss of the floating image, which can solve the problem of non-linear deformation of the image.

[0038] The present invention fits the overall trend of the spine through a function curve, thus avoiding the deficiency of falling into the differences between modalities in the traditional feature point matching algorithm. By continuously comparing the metrics of the sampling function and the spatial transformation network and iteratively improving the similarity, the registration of two images is adaptively realized without the need to provide prior knowledge or prior distribution through a data set. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0040] Figure 1 It is a schematic flowchart of the method provided by the embodiment of the present invention;

[0041] Figure 2 It is a schematic diagram of an intraoperative optical image provided by the embodiment of the present invention;

[0042] Figure 3 It is a schematic diagram of a preoperative CT image provided by the embodiment of the present invention;

[0043] Figure 4 It is a schematic diagram of a fused image after image registration provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0045] The embodiment of the present invention first provides a non-rigid registration method for intraoperative optical images and preoperative CT two-dimensional screenshots of the spine. By sampling, features of non-spinal parts such as the pelvis and ribs are removed, thus realizing image registration without relying on a data set and complex digital image knowledge. This embodiment is mainly used for the identification of intraoperative spinal vertebrae of scoliosis patients.

[0046] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:

[0047] First, collect the intraoperative optical image of the patient and the preoperative CT screenshot. Extract the spinal gray-scale image from the intraoperative optical image through the trained Unet neural network, and then adjust it to the same size as the preoperative screenshot, which are I1 and I2 respectively.

[0048] Secondly, set the proportion of the number of same pixels after binarization of the two images in the total number of pixels as the evaluation index to measure the similarity of the two images. The detailed explanation of the evaluation index is as follows: The initial score is recorded as 0. For each pixel at the same position on the two binarized images, judge whether the pixel values of the two are equal. If they are equal, the score is increased by 1. Divide the finally statistically obtained score by the image size (total number of pixels) to obtain the evaluation index. The ideal value of the evaluation index is 1.

[0049] Then, set that there are two actions in the action space A, which are a1: fitting the function curve and sampling, and a2: STNs spatial transformation network.

[0050] Next, perform the two actions a1 and a2 on I2 respectively to obtain two new I2s, measure the similarity with I1 through the evaluation index respectively, and take the image with the higher similarity as the new I2.

[0051] Repeat the above step until the current evaluation index is greater than 0.9 or the number of iterations is greater than 10, then terminate and output the fused image of I2 and I1.

[0052] The present invention relies on the idea of deep reinforcement learning. By continuously measuring the evaluation index of the sampling action and the spatial transformation network, it avoids excessive elimination of features by the sampling action, so as to achieve the good removal of irrelevant features such as ribs and pelvis without relying on the prior distribution provided by the data set.

[0053] The sampling function is designed based on the fact that the gray-scale image of the spine approximates a function curve. Specifically, for the points with pixel values greater than 8 after inverting the pixel values of the two-dimensional CT cross-sectional image, the abscissa and ordinate are curve-fitted by a fifth-degree polynomial curve through the functions in the numpy library to obtain a function curve, and then calculate the distance from all points with non-zero pixel values to the points corresponding to the abscissa on the function curve. To avoid excessive elimination of features by the sampling function, only 30% of the points with the farthest distance are regarded as irrelevant features and eliminated (set the pixel values to 0).

[0054] Spatial Transformer Networks (STNs) is a convolutional neural network architecture model proposed by Jaderberg et al. By transforming the input images, it reduces the impact of spatial diversity of data to improve the classification accuracy of the convolutional network model, rather than by changing the network structure. STNs have good robustness and possess spatial invariance such as translation, scaling, rotation, perturbation, and bending. Here, we implement the affine transformation between two images through STNs.

[0055] Neural networks need to work by optimizing the loss function. In this invention, the affine transformation between two images is implemented through the spatial transformation network and then registration is performed. The loss function of the spatial transformation network is set as follows: normalized cross-correlation loss + decay factor * (gradient smoothing loss + entropy loss). Among them, the normalized cross-correlation loss is used to measure the similarity between two images, the gradient smoothing loss is used to smooth the deformation field, and the entropy loss is used to avoid excessive loss of information in the images to be registered.

[0056] In this embodiment, the number of iterations of STNs is set to 50.

[0057] Before taking each step in the action space, each of the two actions is performed once, and the similarities between them and the intraoperative spinal gray-scale image are compared to determine whether the sampling function has excessively removed features. If not, continue sampling. Otherwise, perform the affine transformation through the spatial transformation network.

[0058] Generally speaking, the embodiment of this invention designs a sampling function to remove irrelevant features instead of learning its distribution through the dataset or finding feature points through SIFT, and adaptively determines when to stop sampling by continuous comparison without having to learn from the dataset. This avoids problems such as insufficient or biased datasets, while both supervised and unsupervised learning rely on sufficient and unbiased datasets.

[0059] As Figure 1 shown, the specific implementation steps of this invention are as follows:

[0060] 1. Collect the intraoperative optical images of the patient. As the intraoperative optical images shown, pass the intraoperative optical images through the trained Unet neural network to obtain the intraoperative spinal gray-scale image I1 Figure 2 shown, pass the intraoperative optical images through the trained Unet neural network to obtain the intraoperative spinal gray-scale image I1

[0061] 2. Synthesize the dicom files taken by the patient before surgery through mimcs to obtain the three-dimensional CT image of the patient's spine. Take a screenshot in 3-matic to obtain the final preoperative CT image I2. The obtained preoperative CT image is as Figure 3 shown.

[0062] 3. Resize the two images to (500, 206) and invert the pixel values ​​of I2 to obtain the inverted CT image.

[0063] The CT image after the I2 pixel value is inverted (inverted CT image) is used as the image to be executed, and two actions in the action space (fitting the function curve and sampling and STNs spatial transformation network) are executed to obtain the sampled action CT image and the spatial transformation CT image after execution; it mainly includes steps 4 and 5.

[0064] 4. Execute the fitting function curve and sampling action: The CT image after the I2 pixel value is inverted passes through the sampling function. The sampling function selects all pixel points in the image with pixel values ​​greater than 8 for 5th-order polynomial fitting, and then calculates the distance from all points with non-0 pixel values ​​to the corresponding horizontal coordinate point on the function curve. The 30% points with the farthest distance are regarded as irrelevant features and eliminated (the pixel value is set to 0) to obtain the sampling action CT image.

[0065] 5. Execute the STNs spatial transformation network action: pass the CT image after the I2 pixel value is inverted through the STNs function, splice it into a size of (500,500), enter the STNs neural network, and perform 50 iterations to obtain the spatially transformed CT image.

[0066] 6. Perform similarity measurement calculation based on the evaluation index on the sampling action CT image obtained in step 4 and the spatial transformation CT image obtained in step 5 and the intraoperative spine grayscale image I1, and select the sampling action CT image or the spatial transformation CT image with a larger index as the new image to be executed to execute the two actions in the action space, that is, select the image with a higher similarity to I1 as the new image to be executed.

[0067] 7. Repeat steps 4, 5, and 6 until the evaluation index is greater than 0.9 or the number of repetitions is greater than 10. In a sense, repeating steps 4, 5, and 6 can be seen as an optimization process as follows: the evaluation index is improved when passing through 4 and 5 respectively. In step 6, by measuring which image obtained in step 4 and 5 has a greater improvement in the evaluation index (that is, the index value is higher), and using the image with higher improvement as the image for the next cycle, continue the three steps, thereby further improving the evaluation index. When it is improved to 0.9, it can be considered that the two images have basically achieved registration (that is, the similarity is high enough).

[0068] In another embodiment, when the number of repetitions is greater than 10, it can be considered that the number of iterations is high enough, and the loop should be terminated to prevent the program from crashing.

[0069] 8. The final CT image to be registered (optimized I2) is fused with I1 to obtain the image after image registration. The registered image is as follows:Figure 4 as shown

[0070] In other embodiments, the present invention also discloses a cross-modal spinal image registration system, which includes a computer program. When the computer program is executed, it is used to implement the cross-modal spinal image registration method described in any one of the above of the present invention.

[0071] In additional embodiments, the present invention also discloses a cross-modal spinal image registration device, such as an electronic device, a computer device, etc. The device includes a computer system, and the computer system is used for the cross-modal spinal image registration method described in any one of the above of the present invention.

[0072] In this specification, the various embodiments are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for the relevant parts.

[0073] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A cross-modal spinal image registration method, characterized in that, The following steps are involved: Image acquisition steps: Obtain preoperative CT images and corresponding intraoperative optical images; Image preprocessing steps: Inverting the pixel values ​​in the preoperative CT image to obtain an inverted CT image; Acquire a corresponding intraoperative grayscale image based on the intraoperative optical image; Image transformation steps based on deep reinforcement learning ideas: The inverted CT images and intraoperative grayscale images are input into the network model based on reinforcement learning. The ratio of the number of identical pixels to the total pixels is used as the evaluation index, and the fitting function curve and sampling and spatial transformation networks are used as the two actions in the action space. The inverted CT image is used as the image to be executed, two actions in the action space are executed, and a sampled action CT image and a space-transformed CT image after execution are obtained; The sampling motion CT image and the spatial transformation CT image are respectively subjected to similarity measurement calculation based on the evaluation index with the intraoperative grayscale image, and the sampling motion CT image or the spatial transformation CT image with a larger index is used as a new image to be executed; The two actions in the action space are executed again on the new image to be executed, and the execution result is again calculated with the intraoperative grayscale image based on the similarity measurement of the evaluation index until the evaluation index of the execution result meets the set threshold, and the execution result that meets the set threshold is used as the final CT image to be registered; the execution result includes a sampled action CT image or a spatially transformed CT image after executing the two actions in the action space; Image fusion steps: The CT image to be registered and the intraoperative grayscale image are fused to obtain a registered fused image.

2. The cross-modal spinal image registration method according to claim 1, wherein In the image acquisition step, obtaining the preoperative CT image specifically includes: The preoperative DICOM files taken were synthesized through MIMCS to obtain the preoperative CT images.

3. The cross-modal spinal image registration method according to claim 1, wherein The image preprocessing step also includes adjusting the size of the inverted CT image and the intraoperative grayscale image to a consistent size.

4. The cross-modal spinal image registration method according to claim 1, characterized in that In the image preprocessing step, a corresponding intraoperative grayscale image is obtained based on the intraoperative optical image, specifically including extracting the corresponding intraoperative grayscale image from the intraoperative optical image through a trained Unet neural network.

5. The cross-modal spinal image registration method according to claim 1, wherein In the image transformation step based on the deep reinforcement learning idea, the sampling action CT image is obtained, which specifically includes the following steps: Perform a fifth-order polynomial curve fitting on the horizontal and vertical coordinates of the pixel points whose pixel values ​​in the inverted CT image are greater than a set threshold value to obtain a function curve; The distances from all pixels whose pixel values ​​are not 0 to the corresponding horizontal coordinates on the function curve are calculated, and the pixels exceeding the set distance threshold are eliminated to obtain a sampled motion CT image.

6. The cross-modal spinal image registration method according to claim 1, wherein In the image transformation step based on the idea of ​​deep reinforcement learning, the final CT image to be registered is obtained, which also includes: The new image to be executed executes the two actions in the action space again, and the execution result is again compared with the intraoperative grayscale image for similarity measurement based on the evaluation index, until the number of executions meets the set threshold, and the sampled action CT image or spatial transformation CT image with a higher evaluation index in the last execution result is used as the final CT image to be registered.

7. The cross-modal spine image registration method according to claim 1, wherein In the image transformation step based on the deep reinforcement learning idea, the spatial transformation CT image is obtained, which specifically includes the following steps: The inverted CT image is passed through a spatial transformation network for 50 iterations to obtain a spatially transformed CT image.

8. The cross-modal spine image registration method according to claim 7, wherein The loss function of the spatial transformation network includes a normalized cross-correlation loss, a gradient smoothing loss, and an information entropy loss.

9. A cross-modal spinal image registration system, characterized in that, The system includes a computer program which, when executed, implements the cross-modal spinal image registration method according to any one of claims 1-8.

10. A cross-modal spinal image registration device, characterized in that, The device includes a computer system which is used to implement the cross-modal spinal image registration method according to any one of claims 1-8.

Citation Information

Patent Citations

  • SAR image registration method based on SIFT and normalized mutual information

    CN103839265A

  • Cross-modal medical image registration method and computer readable storage medium

    CN112232362A