Method and Device for Arbitrary Modal Image Registration Based on Neural Network

By random mode mapping and neural network optimization of input data, high accuracy and efficiency of arbitrary modal image registration are achieved, and the problems of data acquisition difficulties and modal dependence in the prior art are solved, reducing the workload.

CN114627167BActive Publication Date: 2025-05-30GUANGZHOU RAYDOSE MEDICAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210177986.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-05-30
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

The existing neural network-based image registration methods require a large amount of paired data for training, especially in the medical field, which leads to high workload and low efficiency due to the difficulty in obtaining data and pair matching. At the same time, there are many modalities in images, and the existing methods require training models for various modal combinations, which is a huge workload.

Method used

A method of arbitrary modal image registration based on neural network is proposed. By randomly modal mapping of input data, the image before and after deformation is generated, and the output deformation field is inputted into the neural network to generate an output deformation field. The neural network is optimized by combining the random deformation field and the output deformation field to obtain a training model. This method can ignore the influence of modality, improve the accuracy of registration, reduce the number of models built, and reduce the workload.

Benefits of technology

Through random mode mapping and neural network optimization, high accuracy and efficiency of image registration are achieved, the dependence on different modes is reduced, the workload is reduced, and the problem of data acquisition and pairwise matching is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627167B_ABST
    Figure CN114627167B_ABST
Patent Text Reader

Abstract

The present invention discloses an arbitrary-modal image registration method and device based on a neural network, which relates to the technical field of image registration. The method of the present invention can simulate the characteristics of substances through the generated random substance information, map the random substance information to obtain a pre-deformation image, and deform and map the random substance information based on a random deformation field to obtain a post-deformation image. The pre-deformation image and the post-deformation image are used as the inputs of the neural network. After the registration of the neural network, an output deformation field is obtained. By comparing the difference between the output deformation field and the random deformation field, the neural network is optimized to obtain a trained training model. The training model obtained by the method of the present invention can ignore the influence of the modality, improve the accuracy of registration, reduce the dependence on training data, and also reduce the number of models to be built and the workload.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image registration, and particularly to an arbitrary-modal image registration method and device based on a neural network. Background Art

[0002] In daily life, in fields such as remote sensing and medical care, it is often necessary to match and superimpose two sets of images obtained at different times, by different imaging devices and methods. This process is called image registration. Traditional registration methods are mainly divided into three categories: feature point registration, rigid registration, and elastic registration. In recent years, machine learning and neural networks have developed rapidly, and registration methods based on neural networks have also emerged. Compared with traditional methods, the registration methods based on neural networks have advantages such as fast prediction speed and good effect.

[0003] However, existing registration methods based on neural networks often require a large amount of paired data for training. In the medical field, the acquisition of such data involves issues such as privacy and ethics, and it is difficult to ensure paired matching of the data. In practical applications, there are numerous modalities of images. Taking medical images as an example, there are images with inconsistent modalities such as CT, MR-T1, MR-T2, and so on. General registration methods based on neural networks need to train models for various modality combinations respectively. Assuming there are N modalities, then N 2 models need to be trained, and corresponding data needs to be prepared for these models respectively, which is very cumbersome and involves a large amount of work. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. For this purpose, the present invention provides an arbitrary-modal image registration method, device, and storage medium based on a neural network.

[0005] According to the arbitrary-modal image registration method based on a neural network in the first aspect embodiment of the present invention, which is applied to an image registration device and includes a training part and a prediction part. The training part includes: performing random modal mapping on input data to obtain a pre-deformation image, where the input data includes random substance information or real image information; deforming the input data according to a random deformation field and performing random modal mapping to obtain a post-deformation image; inputting the pre-deformation image and the post-deformation image into a neural network to generate an output deformation field; combining the random deformation field and the output deformation field to optimize the neural network to obtain a training model. The prediction part includes: inputting a to-be-registered image and a reference image into the training model to obtain a predicted deformation field; applying the predicted deformation field to the to-be-registered image to obtain a registered image.

[0006] According to some embodiments of the present invention, the random modal mapping of the input data to obtain the pre-deformation image includes: introducing a discontinuous mechanism into continuous noise to obtain the random material information; and using a transformation function to map the random material information to generate the pre-deformation image.

[0007] According to some embodiments of the present invention, the introducing a discontinuous mechanism into continuous noise includes: performing a negation operation on the continuous noise or selecting the continuous noise number corresponding to a specific value among multiple continuous noises at corresponding positions.

[0008] According to some embodiments of the present invention, the deforming the input data according to a random deformation field, performing random modal mapping, and obtaining a post-deformation image includes: combining a plurality of continuous noises to generate the random deformation field; applying the random deformation field to the input data to generate deformed data; and using a transformation function to map the deformed data to generate the post-deformation image; wherein the number of the continuous noises is the same as the dimension of the post-deformation image.

[0009] According to some embodiments of the present invention, the transformation function is represented by the following formula:

[0010] p(x) → cos((r 1 + 0.5)πp(x)+r 2 )

[0011] wherein the above formula represents a mapping relationship, p(x) is the value of the random material information or the pixel value of the real image, and is normalized to the range of [-1, 1], π is the circumference ratio, r 1 and r 2 are random numbers uniformly distributed in the range of [0, 1). By selecting different random numbers, the modalities of the mapped images are different; for the random material information obtained based on a plurality of continuous noises, the transformation function for realizing random modal mapping may include a sequence with a plurality of random numbers, that is, using the random number sequence as the mapping relationship in the random modal mapping. Each random number corresponds to an index, and the mapping can be completed by combining the random number with the random material information. The number of random numbers corresponds to the number of continuous noises; when an object includes N types of materials, N continuous noises are correspondingly used for simulation. Therefore, the generated sequence contains N random numbers. For example:

[0012] q(x) = argmax({p j (x), j = 0, …, N - 1})

[0013] p(x) = R(q(x))

[0014] Among them, q(x) represents the random material information formed after selecting the maximum noise value at the corresponding position among multiple consecutive noises, pj(x) represents a certain continuous noise, the argmax function is used to select the maximum value at a certain point, p(x) is the result after mapping, and R represents a sequence with N random numbers.

[0015] According to some embodiments of the present invention, optimizing the neural network by combining the random deformation field and the output deformation field to obtain a training model includes: using the random deformation field and the output deformation field as input values of a loss function, inputting the output value of the loss function into the neural network, and optimizing the neural network to obtain the training model.

[0016] According to some embodiments of the present invention, the neural network is built based on the U-net architecture. Five levels of convolutional layers are set in the neural network. Two consecutive convolutional blocks are set in each convolutional layer except the last level, that is, two consecutive convolutional blocks are set in each of the first four levels of convolutional layers, and there is one consecutive convolutional block in the last level of convolutional layer. Each consecutive convolutional block includes three consecutive convolutional operations; pooling and upsampling are performed between levels. After passing through the convolutional layer, pooling, and upsampling, a differential and integral operation is performed through a differential and integral layer to obtain the output deformation field; an activation function is used in each convolutional layer except the last level; consecutive convolutional block A in the previous level is pooled to consecutive convolutional block A in the next level, consecutive convolutional block A in the fourth level of convolutional layer is pooled to consecutive convolutional block C, consecutive convolutional block C is upsampled to consecutive convolutional block B in the fourth level of convolutional layer, and consecutive convolutional block B in the next level is upsampled to consecutive convolutional block B in the previous level; and the number of output channels of each level of convolutional layer increases in multiples, but the number of output channels of the last two levels of convolutional layers is the same; consecutive convolutional blocks perform three consecutive 3*3*3 convolutional operations and use 2*2*2 pooling and 2*2*2 upsampling.

[0017] According to some embodiments of the present invention, the activation function is the Leaky ReLU function, and its slope in the negative interval is 0.2.

[0018] An image registration device according to an embodiment of the second aspect of the present invention includes a processor and a memory communicatively connected to the processor; a computer program that can run on the processor is stored on the memory, and when the computer program is executed by the processor, the method described in the embodiment of the first aspect is implemented.

[0019] A computer storage medium according to an embodiment of the third aspect of the present invention stores computable executable instructions for executing the method described in the embodiment of the first aspect.

[0020] The method of the present invention can simulate the characteristics of a substance through the generated random substance information, obtain a pre-deformation image by performing a random modal mapping on the random substance information, and moreover, deform the random substance information based on a random deformation field and perform a random modal mapping to obtain a post-deformation image. After the random modal mapping, the modalities of the pre-deformation image and the post-deformation image are different. The pre-deformation image and the post-deformation image are used as the inputs of a neural network, that is, pre-deformation images and post-deformation images of different modalities are input into the neural network. During the process of training the neural network a large number of times, the neural network is further optimized. The obtained training model can ignore the influence of the modality, improve the accuracy of registration, and moreover, can reduce the number of models to be built and reduce the workload.

[0021] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, in which:

[0023] Figure 1 is a schematic diagram of the steps of the training part of the image registration method according to an embodiment of the present invention;

[0024] Figure 2 is a schematic flow diagram of the image registration method according to an embodiment of the present invention;

[0025] Figure 3 is a schematic diagram of the architecture of a neural network according to an embodiment of the present invention;

[0026] Figure 4 is a schematic diagram of the effect of introducing a discontinuous mechanism into continuous noise according to an embodiment of the present invention;

[0027] Figure 5 is another schematic diagram of the effect of introducing a discontinuous mechanism into continuous noise according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0029] In the description of the present invention, "several" means one or more, "multiple" means two or more, "greater than", "less than", "exceeding", etc. are understood not to include the recited number, and "above", "below", "within", etc. are understood to include the recited number. If the terms "first" and "second" are described, they are only used to distinguish technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0030] In the description of the present invention, unless otherwise clearly defined, terms such as "set", "installed", "connected", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in the present invention in combination with the specific content of the technical solution.

[0031] Traditional registration methods such as feature point registration can extract image feature points, match feature point pairs through similarity measurement, and obtain the spatial coordinate transformation parameters of the image according to the matched feature point pairs; while rigid registration regards the object to be registered as a rigid body and only uses means such as translation, rotation, scaling, and shearing to achieve registration; elastic registration regards the object to be registered as a deformable elastic object or fluid and achieves registration by solving elastic mechanics or fluid mechanics equations.

[0032] The speed of traditional method registration is average, and its effect is not as good as the registration method using neural network. However, the registration method based on neural network also has defects in practical applications: due to the inconsistent modalities of images, models need to be trained separately for various modality combinations, resulting in a large amount of work, low efficiency of the neural network registration method, and the need for a large amount of data.

[0033] From the perspective of images, modality can be understood as the mapping from matter to pixel values. Different matters are mapped to different pixel values, that is, different modalities correspond to different mapping methods. For example, CT, MR-T1, and MR-T2 images correspond to different modalities, and the pixel values corresponding to the same matter in different modalities are also different. For example, in CT images, the bone cortex is a high signal (large pixel value), while water and fat are low signals (small pixel value); in MR-T1 images, the bone cortex and water are low signals, while fat is a high signal; in MR-T2 images, the bone cortex is a low signal and water and fat are high signals. Generally, an object contains multiple matters, the object has a size, and the matters therein also have size differences. The matters are continuously distributed, and there are obvious boundaries between the matters. Therefore, from the perspective of images, different matters correspond to different pixel values, and the pixel values are continuously distributed in patches, but the patches are not continuous with each other.

[0034] The method of the present invention can simulate the above characteristics through the generated random material information, perform random modal mapping on the random material information or real image information to obtain the pre-deformation image, and deform the random material information based on the random deformation field and perform random modal mapping to obtain the post-deformation image. After random modal mapping, the modalities of the pre-deformation image and the post-deformation image are different. The pre-deformation image and the post-deformation image are used as the inputs of the neural network, that is, the pre-deformation image and the post-deformation image of different modalities are input into the neural network. During the process of training the neural network a large number of times, the neural network is further optimized, and the obtained training model can ignore the influence of the modality, improve the accuracy of registration, and can also reduce the number of models to be built and the workload.

[0035] The present invention will be further described below with reference to the accompanying drawings.

[0036] Refer to Figure 1 , in some embodiments of the present invention, the image registration method based on a neural network includes a training part and a prediction part. The training part at least includes the following steps:

[0037] Step S100: Perform random modal mapping on the input data to obtain the pre-deformation image.

[0038] It can be understood that the pre-deformation image can be obtained from a simulated image constructed by random material information or from a real image, that is, the input data can be random material information or real image information. Therefore, in some embodiments, to obtain the pre-deformation image, random material information can be generated first, and then random modal mapping is performed on the random material information. After mapping, the pre-deformation image can be obtained. In some embodiments, it can also be to perform random modal mapping on the received real image information to finally obtain the pre-deformation image.

[0039] It should be noted that random modal mapping is performed on both the random material information and the real image information to change the original modality. For example, when real image information is received, these real image information can be CT, MR-T1, MR-T2, etc. The modalities of the real images are different. Performing random modal mapping on the real images changes their modalities, that is, the modality of the pre-deformation image input into the neural network is different from the original modality.

[0040] Step S200: Deform the input data according to the random deformation field and perform random modal mapping to obtain the post-deformation image.

[0041] It can be understood that as another input of the neural network, the post-deformation image requires the input data to be deformed. For example, the random deformation field is applied to the random material information or the real image information to obtain the deformed data, and then random modal mapping is performed on the deformed data to obtain the post-deformation image.

[0042] It should be noted that, taking real image information as an example, the modality of the deformed data obtained by applying a random deformation field to a real image is the same as that of the real image. That is, the deformed data after the real image is deformed is still one of CT, MR-T1, MR-T2, etc. However, through the processing of random modality mapping, the modalities of the pre-deformation image and the post-deformation image change. The modalities of the two are different and also different from their original modalities. That is, the modality of the pre-deformation image is different from the modality of the post-deformation image, the modality of the pre-deformation image is different from the modality of the real image, and the modality of the post-deformation image is different from the modality of the deformed data. Inputting the pre-deformation image and the post-deformation image after random modality mapping into a neural network and continuously training the neural network enables the obtained training model to adapt to different modalities, ignores the influence brought by the modalities, and thus pays more attention to the deformation, making the image registration more accurate. Moreover, the applicability of the neural network is improved, and at the same time, there is no need to construct a large number of models, reducing the workload of registration.

[0043] Step S300: Input the pre-deformation image and the post-deformation image into a neural network to generate an output deformation field.

[0044] Step S400: Optimize the neural network by combining the random deformation field and the output deformation field to obtain a training model.

[0045] It can be understood that the pre-deformation image and the post-deformation image are used as the inputs of the neural network. After being calculated by the neural network, an output deformation field can be obtained. Optimize the neural network by combining the random deformation field and the output deformation field, thereby improving the accuracy of the output of the neural network, so that the output of the training model obtained after continuous training can be closer to the actual deformation.

[0046] The prediction part at least includes: inputting the image to be registered and the reference image into the training model to obtain a predicted deformation field; applying the predicted deformation field to the image to be registered to obtain a registered image.

[0047] The training model can output the deformation relationship between the image to be registered and the reference image, that is, the predicted deformation field. By applying the predicted deformation field to the image to be registered, the image to be registered is deformed to obtain a registered image, completing the registration. It can be understood that the image to be registered and the reference image are different, and it is difficult or even impossible to perform a reference comparison between them. However, by finding the deformation relationship between the two through the training model, the image to be registered is input into the training model as the pre-deformation image, and the reference image is input into the training model as the post-deformation image. The obtained predicted deformation field acts on the image to be registered to make it deformed, generating a registered image, and the registered image can be used for reference comparison with the reference image.

[0048] For example, in the application scenario of medical image registration, in order to find the changes in the lesion site of the lungs, it is necessary to make a reference comparison between two images. However, due to factors such as the different shooting angles of the two images and the different postures of the person being photographed during the two shootings, it is difficult or even impossible to directly compare the two images. The two images are the image to be registered and the reference image respectively. For example, if the image to be registered is a CT image, which is an image of the lungs during inspiration taken before the lesion of the person being photographed, and the reference image is an MR image, which is an image of the lungs during expiration taken after the lesion of the person being photographed. There is a deformation relationship of the lungs in the inspiration and expiration states between the image to be registered and the reference image. When the two are input into the training model, the training model can ignore the influence brought by the different modalities of the CT image and the MR image, and focus on the deformation relationship between the two, so as to output the predicted deformation field. Then, the predicted deformation field is applied to the image to be registered, so that the deformed image of the image to be registered is obtained, that is, the image in the inspiration state becomes the image in the expiration state, and the registration is completed. At this time, the registered image can be compared with the reference image for reference, which is convenient for observing the changes in the lesion site.

[0049] Therefore, the method of the present invention can effectively solve the problems of difficult data acquisition and few paired data, and there is no need to construct a large number of models for different modalities, effectively reducing the workload. The training model can ignore the influence brought by the modality and improve the accuracy of image registration.

[0050] In some embodiments of the present invention, for the random material information, by introducing a discontinuous mechanism into the continuous noise, the generated random material information can simulate the characteristics of the material, that is, it has continuity within the same material and discontinuity between different materials. It should be noted that the continuous noise is continuously changing noise, which is ordered noise and satisfies the continuous characteristic in a neighborhood of any point on it. The fractal noise formed by superimposing multiple continuously changing noises also belongs to a kind of continuous noise. Perlin (Berlin) noise in the continuous noise can be selected, and multiple Perlin noises are superimposed to form fractal noise, and a discontinuous mechanism is introduced into the fractal noise.

[0051] After introducing the discontinuous mechanism into the fractal noise, the random material information is obtained, and the random material information is subjected to random modality mapping through a transformation function. For example, the data of the random material information is mapped to pixel values to obtain the deformed image. By generating random material information that can simulate the characteristics of the material, the input of the neural network is not limited to real images, which can effectively solve the problem of difficult data acquisition. Moreover, random modality mapping is also performed on the random material information, so that the neural network can adapt to different modalities, reduce the number of models built, and reduce the workload.

[0052] In some embodiments, the generation of random material information can be an inversion operation on continuous noise, so as to obtain random material information that satisfies the discontinuous feature between slices of the mapped pixel values. The inversion operation can be understood as taking the inversion with a certain value in the continuous noise as the boundary, as shown in the following formula:

[0053] p(x) → max(p(x), a - p(x))

[0054] Where p(x) represents continuous noise, such as fractal noise, and a represents a certain value in the continuous noise. The maximum value of p(x) and a - p(x) is selected and then normalized to the range of [-1, 1] for easy calculation. It should be noted that multiple inversion operations can be performed to make the generated deformed image more complex. It can be noted that when a is taken as 0, it is an absolute value operation on the continuous noise.

[0055] The distribution of substances is generally continuous, while it is discontinuous between different substances. Therefore, continuous noise can be used to simulate continuous substances, and by taking the absolute value of the continuous noise, the discontinuous feature between different substances can be simulated.

[0056] For example, in a two-dimensional coordinate system, after the continuous noise undergoes an absolute value operation, its first derivative is discontinuous on both sides of the zero crossing point. The generated random material information is also discontinuous in the pre-deformation image generated after mapping. As shown in reference Figure 4 , the pre-deformation image is represented by a 256-level grayscale image. Different pixel values correspond to different grayscale slices, different grayscale slices represent different substances, and there are boundaries between different grayscale slices. It should be noted that the continuous noise can be normalized, such as normalizing the continuous noise to the range of [-1, 1] for easy calculation.

[0057] In some embodiments, the continuous noise numbers corresponding to specific values in multiple continuous noises at corresponding positions are selected to introduce a discontinuous mechanism. It should be noted that the specific value is a selection index for introducing the discontinuous mechanism, and the specific value can be the maximum value or the second largest value.

[0058] For example, the number of the continuous noise with the largest noise value among multiple continuous noises at the corresponding position is selected. It can be understood that, when multiple continuous noises, such as fractal noise, are selected, the number of continuous noises corresponds to the number of substances, and the multiple continuous noises can be numbered. For example, 10 continuous noises are selected and numbered from 0 to 9 in sequence. In the two-dimensional coordinate system, the continuous noise is distributed on the coordinate system. At a corresponding position, such as a coordinate point, each continuous noise corresponds to a noise value at the coordinate point. When the noise value of the continuous noise numbered 1 at the coordinate point is the largest among the 10 noise values ​​at the coordinate, the number of the continuous noise numbered 1 at the coordinate is selected. Due to the characteristics of fractal noise, the noise value of the continuous noise is continuous within a neighborhood of the coordinate point, that is, it can be considered that the noise values ​​of some points in the neighborhood of the coordinate point also meet the specific value conditions. Therefore, number 1 is selected in the neighborhood to form a set of number 1. From the perspective of the image, refer to Figure 5 , the numbers are matched with the pixel values ​​to form an image, and the position of the above coordinate point on the image is a sheet area of ​​a certain pixel value, that is, the sheet area is a certain pixel value (corresponding to the set of number 1). After traversing all the coordinate points, the random material information is obtained.

[0059] In some embodiments of the present invention, for the generation of random deformation field, multiple continuous noises can be combined to form a random deformation field, which provides a deformation field for random material information and real image, and the deformation field (such as random deformation field, output deformation field, predicted deformation field, etc.) is applied to the image. It can be understood that the deformation field provides deformation direction and deformation amount for the deformation of the image. Moreover, the number of continuous noises is the same as the dimension of the deformed image. For example, if the dimension of the deformed image is 3, the number of continuous noises is 3, which respectively provide deformation direction and deformation amount for the random material information and the real image. Therefore, in the three-dimensional coordinate system, the three continuous noises respectively provide deformation amounts along the three directions of X, Y, and Z for the random material information and the real image. The generated random deformation field is applied to the input data to generate deformation data, and the deformed image is obtained by performing random modal mapping on the deformation data. For example, the deformation data is mapped using a transformation function, so that the deformation data is mapped to corresponding pixel value information to generate the deformed image. The deformation data is also mapped so that its mode can be suitable for neural networks, which can reduce the number of models to be built and reduce the workload.

[0060] In some embodiments of the present invention, a transformation function may be used to map the modes and implement random mode mapping. It should be noted that the transformation function represents a mapping relationship. For example, it may be expressed as follows:

[0061] p(x)→cos((r 1(+0.5)πp(x)+r 2 )

[0062] Among them, the above formula represents a mapping relationship. p(x) is the value of random material information or the pixel value of a real image, and is normalized to the range of [-1, 1]. π is the pi, and r 1 and r 2 are random numbers uniformly distributed in [0, 1). By selecting different random numbers, the modalities of the mapped images are also different. The above mapping relationship can fully consider the non-linear and non-monotonic characteristics of modality mapping. After the mapping is completed, the corresponding image can be obtained by selecting the corresponding pixel value for the mapped value.

[0063] In some embodiments, for the random material information obtained based on multiple consecutive noises, the transformation function for implementing random modality mapping may include a sequence with multiple random numbers, that is, a random number sequence is used as the mapping relationship in random modality mapping. Each random number corresponds to an index, and the mapping can be completed by combining the random number with the random material information. Among them, the number of random numbers corresponds to the number of consecutive noises. It can be understood that when an object includes N substances, N consecutive noises are correspondingly used for simulation. Therefore, the generated sequence contains N random numbers. For example:

[0064] q(x) = argmax({p j (x), j = 0,...., N - 1})

[0065] p(x) = R(q(x))

[0066] Among them, q(x) represents the random material information formed after selecting the maximum noise value at the corresponding position among multiple consecutive noises. p j (x) represents a certain consecutive noise. The argmax function is used to select the maximum value at a certain point. p(x) is the result after mapping, and R represents a sequence with N random numbers. It should be noted that in order to increase the authenticity of the result, random noise can also be added.

[0067] Refer to Figure 1 and Figure 2, in some embodiments of the present invention, the neural network is further optimized. The neural network is optimized by combining a random deformation field and an output deformation field. For example, the random deformation field and the output deformation field are used as the input values of a loss function, and then the output of the loss function is input into the neural network to optimize the neural network and obtain a training model. The loss function is used to calculate the difference between the calculation result of each iteration of the neural network and the real value, so as to continuously optimize the neural network, making the output of the training model closer to the actual deformation. Then, by using the output of the training model on the image pair to be registered, the registered image can be obtained, and the image registration is completed. It should be noted that the loss function can be the MAE (Mean Absolute Error) function or the MSE (Mean Squared Error) function.

[0068] Referring to Figure 3 , in some embodiments of the present invention, the neural network is built based on the U-net architecture. Five levels of convolutional layers are set in the neural network. Two consecutive convolutional blocks are set in each convolutional layer except the last one. That is, there are two consecutive convolutional blocks in each of the first four levels of convolutional layers, and there is one consecutive convolutional block in the last convolutional layer. And each consecutive convolutional block includes three consecutive convolutional operations. Moreover, pooling and upsampling are performed between levels. After convolution, pooling and upsampling, differential and integral operations are performed through a differential and integral layer to obtain the output deformation field. It should be noted that an activation function is used in each convolutional layer except the last one. Introducing the activation function can optimize the properties of the output of the neural network and make it more continuous.

[0069] For example, in the first four levels of convolutional layers, there are consecutive convolutional block A and consecutive convolutional block B, and in the last convolutional layer, there is consecutive convolutional block C. Then, consecutive convolutional block A in the upper level is pooled to consecutive convolutional block A in the lower level, consecutive convolutional block A in the fourth-level convolutional layer is pooled to consecutive convolutional block C, consecutive convolutional block C is upsampled to consecutive convolutional block B in the fourth-level convolutional layer, and consecutive convolutional block B in the lower level is upsampled to consecutive convolutional block B in the upper level. And the number of output channels of each convolutional layer increases in multiples, but the number of output channels of the last two convolutional layers is the same. For example, it increases by 2 times. The number of output channels of the first convolutional layer is 8, the number of output channels of the second convolutional layer is 16, the number of output channels of the third convolutional layer is 32, and the number of output channels of the fourth and fifth convolutional layers is 64; or the number of output channels of the first convolutional layer is 16, the number of output channels of the second convolutional layer is 32, the number of output channels of the third convolutional layer is 64, and the number of output channels of the fourth and fifth convolutional layers is 128.

[0070] The continuous convolution block performs three consecutive 3×3×3 convolution operations, and 2×2×2 pooling and 2×2×2 upsampling are adopted. It should be noted that the continuous convolution block can perform three consecutive 5×5×5 convolution operations, three consecutive 7×7×7 convolution operations, etc. Moreover, for three-dimensional images, two consecutive convolution operations can also be adopted. Similarly, for two-dimensional images, two consecutive convolution operations or three consecutive convolution operations can also be adopted. Furthermore, the number of levels of the convolution layer is at least set to 4 levels, and can be set to 4 levels, 5 levels, 6 levels, etc. according to the hardware conditions of the training device.

[0071] In some embodiments of the present invention, the activation function is Leaky ReLU (Leaky Rectified Linear Unit, leaky rectified linear unit), and the activation function Leaky ReLU is expressed by the mathematical formula:

[0072]

[0073] And the slope of the activation function Leaky ReLU in the negative interval is 0.2, that is, a = 0.2. It should be noted that the activation function can also be the ReLU (Rectified Linear Unit, rectified linear unit) function.

[0074] In some embodiments of the present invention, an image registration device is provided. The image registration device includes a processor and a memory. Among them, the memory stores a computer program that can run on the processor, and when the computer program is executed by the processor, it can implement the image registration method described in the above embodiments.

[0075] In the embodiments of the present invention, a computer-readable storage medium is also provided, which stores computer-executable instructions for executing the image registration method described in the above embodiments.

[0076] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof. In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be executed by the cooperation of several physical components. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable storage medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that communication media typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and may include any information delivery medium.

[0077] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.

Claims

1. An arbitrary-modal image registration method based on a neural network, characterized in that, applied to an image registration device, including a training part and a prediction part, the training part includes: Performing a random modal mapping on the input data to obtain a pre-deformation image, where the input data includes random material information or real image information; Deforming the input data according to a random deformation field and performing a random modal mapping to obtain a post-deformation image; Inputting the pre-deformation image and the post-deformation image into a neural network to generate an output deformation field; Combining the random deformation field and the output deformation field to optimize the neural network to obtain a training model; The prediction part includes: Inputting the image to be registered and the reference image into the training model to obtain a predicted deformation field; Applying the predicted deformation field to the image to be registered to obtain a registered image; The performing a random modal mapping on the input data to obtain a pre-deformation image includes: Introducing a discontinuous mechanism into continuous noise to obtain the random material information; Using a transformation function to map the random material information to generate the pre-deformation image; The introducing a discontinuous mechanism into continuous noise includes performing an inversion operation on the continuous noise or selecting the continuous noise number corresponding to a specific value among multiple continuous noises at the corresponding position; The deforming the input data according to a random deformation field and performing a random modal mapping to obtain a post-deformation image includes: Combining multiple continuous noises to generate the random deformation field; Applying the random deformation field to the input data to generate deformed data; Using a transformation function to map the deformed data to generate a post-deformation image; wherein, the number of the continuous noises is the same as the dimension of the post-deformation image; The transformation function is expressed by the following formula: p(x) → cos((r 1 + 0.5)πp(x) + r 2 ) ; Among them, the above formula represents a mapping relationship, p(x) is the value of the random material information or the pixel value of the real image, and is normalized to the range of [-1, 1]. π is the pi, r 1 and r 2 are random numbers uniformly distributed in [0, 1). By selecting different random numbers, the modalities of the mapped images are also different. For the random material information obtained based on multiple consecutive noises, the transformation function for realizing the random modality mapping may include a sequence with multiple random numbers, that is, using the random number sequence as the mapping relationship in the random modality mapping. Each random number corresponds to an index, and the mapping can be completed by combining the random number with the random material information. Among them, the number of random numbers corresponds to the number of consecutive noises. When an object includes N kinds of materials, N consecutive noises are correspondingly used for simulation. Therefore, the generated sequence contains N random numbers. For example: q(x) = argmax({p j (x), j = 0,...., N - 1}); p(x) = R(q(x)); Among them, q(x) represents the random material information formed after selecting the maximum noise value at the corresponding position among multiple consecutive noises, and p j (x) represents a certain continuous noise, the argmax function is used to select the maximum value at a certain point, p(x) is the result after mapping, and R represents a sequence with N random numbers.

2. The arbitrary-modal image registration method based on a neural network according to claim 1, characterized in that, The combining the random deformation field and the output deformation field to optimize the neural network to obtain a training model includes: The random deformation field and the output deformation field are used as input values of a loss function, and the output value of the loss function is input into the neural network to optimize the neural network to obtain the training model.

3. The arbitrary-modal image registration method based on a neural network according to claim 1, characterized in that, The neural network is built based on the U-net architecture. Five levels of convolutional layers are set in the neural network. Two consecutive convolutional blocks are set in each level of convolutional layer except the last one, that is, two consecutive convolutional blocks are in each of the first four levels of convolutional layers, and there is one consecutive convolutional block in the convolutional layer of the last level. And each consecutive convolutional block includes three consecutive convolutional operations; pooling and upsampling are performed between levels. After convolutional layer pooling and upsampling, differential and integral operations are performed through a differential and integral layer to obtain the output deformation field; an activation function is used in each level of convolutional layer except the last one; consecutive convolutional block A and consecutive convolutional block B are included in the convolutional layers of the first four levels, and consecutive convolutional block C is included in the convolutional layer of the last level. Then, consecutive convolutional block A in the upper level pools to consecutive convolutional block A in the lower level, consecutive convolutional block A in the fourth-level convolutional layer pools to consecutive convolutional block C, consecutive convolutional block C upsamples to consecutive convolutional block B in the fourth-level convolutional layer, and consecutive convolutional block B in the lower level upsamples to consecutive convolutional block B in the upper level; and the number of output channels of each level of convolutional layer increases in multiples, but the number of output channels of the last two levels of convolutional layers is the same; the consecutive convolutional block performs three consecutive 3*3*3 convolutional operations and uses 2*2*2 pooling and 2*2*2 upsampling.

4. The arbitrary-modal image registration method based on a neural network according to claim 3, characterized in that the activation function is the Leaky ReLU function, and its slope in the negative interval is 0.

2.

5. An image registration device, characterized in that it includes: a processor and a memory communicatively connected to the processor; a computer program is stored on the memory and can run on the processor. When the computer program is executed by the processor, the method described in any one of claims 1 to 4 is implemented.

6. A computer-readable storage medium stores computable executable instructions, and the computable executable instructions are used to execute the method described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Unsupervised intravascular ultrasound image registration method based on neural network

    CN112150425A