Variational Image Registration Method, Device, Equipment and Medium Based on Cross-Attention Mechanism
By introducing a variational neural network with cross attention mechanism in variable heart registration, the problem of lack of long-distance feature modeling of images is solved, and the accuracy of cardiac image registration is significantly improved.
Patent Information
- Application Number
- CN202310334095.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-03-23
AI Technical Summary
In variable heart registration, the registration accuracy is low due to the lack of long-distance feature modeling of the image and the difference between the log likelihood of the objective function ELBO and the input image pair.
A variational neural network using a cross attention mechanism encodes the image through the T2T layer and the Transformer layer, fuses the multi-scale features of the floating image and the fixed image, and decodes the deformation field of the floating image using a radial basis function.
The accuracy of cardiac image registration is significantly improved, and the gap between the objective function and log likelihood is reduced by effectively modeling the long-distance and global characteristics of the image.
Smart Images

Figure CN116342670B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a variational image registration method, device, equipment and medium with a cross-attention mechanism. Background Art
[0002] Cardiac motion estimation is crucial for evaluating cardiac function, detecting heart diseases, and understanding cardiac biomechanics. Variable image registration is a key technology for cardiac motion estimation.
[0003] Probabilistic generative registration models play a crucial role in clinical applications. However, there are some problems with variational image registration methods. On the one hand, traditional convolutions are limited in representing long-range relationships between image features. On the other hand, there is a gap between the objective function ELBO (Evidence Lower Bound) and the log-likelihood of the input image pair in the variational model, which both lead to a reduction in registration accuracy.
[0004] Therefore, there is an urgent need for an image registration scheme to solve the technical problem of low registration accuracy in variational cardiac registration due to the lack of long-range feature modeling in the image and the gap between the objective function ELBO and the log-likelihood of the input image pair. Summary of the Invention
[0005] The main objective of the present invention is to provide a variational image registration method, device, equipment and medium with a cross-attention mechanism, aiming to solve the technical problem of low registration accuracy in variational cardiac registration due to the lack of long-range feature modeling in the image and the gap between the objective function ELBO and the log-likelihood of the input image pair.
[0006] To achieve the above objective, in the first aspect of the present invention, a variational image registration method with a cross-attention mechanism is provided, including: receiving a cardiac floating image and a fixed image; processing the cardiac floating image and the fixed image by using a pre-trained variational neural network based on a cross-attention mechanism to obtain a latent variable of a variational model; inputting the latent variable into a pre-constructed radial basis function, decoding and outputting a deformation field of the floating image, and determining a registered image according to the deformation field; the variational neural network based on a cross-attention mechanism includes: at least one T2T layer and at least one Transformer layer for encoding a training set of images, the Transformer layer includes an encoder and a decoder, the fixed image features output by the encoder of the Transformer layer are input into the decoder of the Transformer layer to form cross-attention with the floating image features, and the cumulative posterior corresponding to the importance-weighted evidence lower bound is used as the most prior loss function in the variational model to output the latent variable distribution of the variational model, so as to output the latent variable of the variational model.
[0007] Further, if there are at least two T2T layers, all the T2T layers are cascaded; if there are at least two Transformer layers, all the Transformer layers are cascaded, and the output data after the cascading of the T2T layers is used as the input of the cascaded Transformer layers.
[0008] Further, the T2T layer includes at least one Transformer module. The T2T layer is used to divide an image into an overlapping two-dimensional feature sequence with local information modeling and then input it into the Transformer module, and the output feature sequence is restored to three-dimensional image features through three-dimensional transformation.
[0009] Further, there is at least one encoder and at least one decoder in the Transformer layer. If there is more than one encoder and decoder in the Transformer layer, the encoders in the Transformer layer are cascaded, and the decoders in the Transformer layer are cascaded; and the fixed image features output by the encoder of the Transformer layer are input into the decoder of the Transformer layer to form cross-attention with the floating image features.
[0010] Further, the training method of the variational neural network based on the cross-attention mechanism includes: calculating the cumulative posterior corresponding to the importance-weighted evidence lower bound according to the pre-acquired registered cardiac image and the fixed image as the loss function of the most prior in the variational model; adjusting the parameters of the pre-constructed untrained variational neural network based on the cross-attention mechanism according to the loss function to obtain a trained variational neural network based on the cross-attention mechanism.
[0011] A second aspect of the present invention provides an image registration device, including: a receiving module, configured to receive an image training set of a cardiac floating image and a fixed image; a processing module, configured to process the cardiac floating image and the fixed image by using a trained variational neural network based on the cross-attention mechanism, wherein the variational neural network based on the cross-attention mechanism includes: at least one T2T layer and at least one Transformer layer for encoding the image training set, wherein the fixed image features output by the encoder in the Transformer layer are input into the Transformer decoder to form cross-attention with the floating image features to output the latent variable distribution of the variational model; a determining module, a radial basis function for decoding the latent variable to output the deformation field of the floating image, and determining the registered image according to the deformation field.
[0012] A third aspect of the present invention provides an electronic device, including: at least one processor and a memory; wherein, the memory stores computer-executable instructions; the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the variational image registration method of the cross-attention mechanism described in any one of the above.
[0013] A fourth aspect of the present invention provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the variational image registration method of the cross-attention mechanism described in any one of the above is implemented.
[0014] The present invention provides a variational image registration method, device, electronic device and storage medium of a cross-attention mechanism, and the beneficial effects are as follows: in the registration method, a variational neural network structure based on a cross-attention mechanism is used to fuse multi-scale features of a floating image and a fixed image, and different-scale features are fused and decision-making is carried out. Among them, the introduction of the T2T layer is beneficial to the feature modeling of local information of the model, and the global information modeling between the floating image and the fixed image is realized by using the Transformer cross-attention mechanism. At the same time, the cumulative posterior corresponding to the importance-weighted evidence lower bound is used as the most prior loss function in the variational model, so that the distribution of the latent variables estimated by the network is more accurate. The present invention significantly improves the accuracy of cardiac image registration. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.
[0016] Figure 1 It is a schematic flowchart of an image registration method shown according to an exemplary embodiment of the present invention;
[0017] Figure 2 For the present invention Figure 1 It is a schematic structural diagram of a variational neural network based on a cross-attention mechanism of the image registration method shown in the embodiment of the present invention;
[0018] Figure 3 For the present invention Figure 2 It is a schematic structural diagram of a T2T module in a variational neural network based on a cross-attention mechanism shown in the embodiment of the present invention;
[0019] Figure 4 It is a schematic structural diagram of an image registration device shown according to an exemplary embodiment of the present invention;
[0020] Figure 5 This is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present invention. Specific embodiments
[0021] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0022] The present invention aims to provide a variational image registration method, device, equipment, and storage medium with a cross-attention mechanism. By introducing a Transformer cross-attention model, it solves the problem that the global features are not rich enough in variational cardiac registration due to the lack of long-distance feature modeling in images. By optimizing the cumulative posterior corresponding to the importance-weighted evidence lower bound (iwELBO) as the loss function of the most prior in the variational model, the estimation of the latent variable distribution in the variational model is made more accurate, thereby improving the registration accuracy.
[0023] Figure 1 This is a schematic flowchart of a variational image registration method with a cross-attention mechanism shown according to another exemplary embodiment of the present invention. As Figure 1 shown, the registration method provided in this embodiment includes the following steps:
[0024] S201. Receive an image training set.
[0025] More specifically, the image training set includes: a cardiac floating image and a fixed image. The floating image refers to the image of the heart at the end-diastolic (ED) phase of the cardiac cycle, and the fixed image refers to the image of the heart at the end-systolic (ES) phase of the cardiac cycle. Each pair of images can be obtained through annotation in the corresponding expert database.
[0026] In this embodiment, a training sample set is constructed using the short-axis cardiac images of Cine MR and the corresponding image pairs annotated in the expert database.
[0027] After obtaining the training sample set, the image size can be adjusted to be consistent by cropping. The intensity of the images is adjusted using whitening processing to reduce the redundancy of the input images. And through spatial transformation methods such as translation, rotation, and elastic transformation, the training samples in the original training sample set are spatially transformed to generate new training samples, which are the enhanced training sample set.
[0028] S202. Construct a variational neural network based on the cross-attention mechanism.
[0029] More specifically, the variational neural network based on the cross-attention mechanism includes at least one T2T layer and at least one Transformer layer. At least one Transformer layer is cascaded. The T2T layer includes at least one Transformer module. Among them, the T2T layer divides the image into overlapping two-dimensional feature sequences with local information modeling and then inputs them into the Transformer module. The output feature sequence is restored to three-dimensional image features through three-dimensional transformation. The Transformer layer includes at least one Transformer encoder and at least one Transformer decoder. Among them, at least one Transformer encoder is cascaded, and at least one Transformer decoder is cascaded. The fixed image features output by the encoder in the Transformer layer are input into the Transformer decoder to form cross-attention with the floating image features. At least one T2T layer and at least one Transformer layer are used to encode the image training set, and a radial basis function is used to decode the latent variables output by the Transformer to output the deformation field of the floating image.
[0030] Figure 2 For the present invention Figure 1 Schematic diagram of the structure of the variational neural network based on the cross-attention mechanism of the registration method shown in the embodiment. As Figure 2 shown, the variational neural network based on the cross-attention mechanism includes one T2T layer, three Transformer layers, and a radial basis function. Among them, the three Transformer layers are directly cascaded. Each Transformer layer contains a Transformer encoder and a Transformer decoder. One T2T layer and three Transformer layers constitute an encoding layer, and the encoding layer is used to encode the received image training samples to output latent variables. A radial basis function constitutes a decoding layer, and the decoding layer is used to decode the latent variables obtained after encoding by the encoding layer to output the deformation field of the floating image.
[0031] Figure 2 Schematic diagram of the structure of the Transformer layer in the variational neural network based on the cross-attention mechanism shown in the embodiment. As Figure 2As shown, the Transformer layer consists of a Transformer encoder and a Transformer decoder. The Transformer encoder is composed of two layer normalization layers, one multi-head self-attention layer, and one multi-layer perceptron. The Transformer decoder is composed of three layer normalization layers, two multi-head self-attention layers, and one multi-layer perceptron.
[0032] Figure 3 For the present invention Figure 2 The embodiment shows a schematic structural diagram of the T2T layer in the variational neural network based on the cross-attention mechanism. As Figure 3 shown, the T2T layer consists of three Unfold transformations, two Transformer modules, and two Reshape transformations. Among them, the Transformer module is composed of two layer normalization layers, one multi-head self-attention layer, and one multi-layer perceptron. The channel dimension of the Transformer module is 32.
[0033] The image is subjected to the Unfold transformation through the following formula:
[0034]
[0035] where H and W represent the length and width of the input image, P represents the size of the edge padding, K represents the size of the sliding window, S represents the step size of the sliding window, and l o represents the length of the output sequence.
[0036] In the above T2T layer, an image training sample with a size of n×n and 1 channel is subjected to an Unfold transformation with a sliding window size of 7 and a step size of 4 to obtain n / 4×n / 4 image patches of size 49 with overlapping local information modeling, and they are concatenated into a two-dimensional image sequence. Then, a new image sequence with global information modeling is obtained through a Transformer module, and the sequence is subjected to a Reshape transformation to obtain a processed three-dimensional image feature with a size of n / 4×n / 4 and 32 channels. This image feature is subjected to an Unfold transformation with a sliding window size of 3 and a step size of 2, a Transformer module, and a Reshape transformation to obtain a three-dimensional image feature with a size of n / 8×n / 8 and 32 channels. Finally, the three-dimensional image feature is subjected to a third Unfold transformation with a sliding window size of 3 and a step size of 2 to obtain a feature sequence with a length of n / 16×n / 16 and an image patch size of 288, which is used as the input of the Transformer layer.
[0037] In the above Transformer layer, the floating image and the fixed image respectively pass through the T2T layer to output feature sequences. The feature sequences corresponding to the fixed image and the floating image are respectively input into the Transformer encoder and decoder after adding position encoding information. The feature sequence corresponding to the fixed image is sent into the multi-head self-attention module after layer normalization. The multi-head self-attention is calculated according to the following formula:
[0038]
[0039] MSA = [SA 1 ; …; SA h
[0040] where Q, K, V are obtained by linear transformation of the input sequence, and the attention mechanism can calculate weights for each sequence. After passing through the softmax activation function, these weights are dot-multiplied by V to obtain the weights of the input sequence. MSA represents multi-head self-attention, which is obtained by cascading self-attention at different scales. In the encoder, Q, K, V all come from the linear transformation of the input sequence corresponding to the fixed image. In the decoder, Q comes from the linear transformation of the input sequence corresponding to the floating image, while K and V come from the linear transformation of the input sequence corresponding to the fixed image. This cross-attention operation helps to model the relationship between the floating image and the target image. The decoder outputs the distribution of the latent variable through the last layer of multi-layer perceptron.
[0041] The latent variable is input into the radial basis function to obtain the deformation field of the floating image. The floating image obtains the deformed heart image through the deformation field.
[0042] The deformation field of the floating image is determined according to the following formula:
[0043]
[0044] where v is the pixel value, p i is the control point, and ||v - p i || 2 is the Euclidean distance. is a radial basis function with a support set of c, and there are s different support sets in total is the latent variable whose distribution needs to be estimated.
[0045] After determining the structure of the variational neural network based on the cross-attention mechanism, continue to set the parameters of the variational neural network based on the cross-attention mechanism. The parameters of the variational neural network based on the cross-attention mechanism are the parameters of all multi-layer perceptrons. For the parameter w i in the network, use the adjustment amount Δw obtained by backpropagation in S206i , adjusted to w' i = w i + Δw i . During the first iteration, the network parameters are set to random values.
[0046] S203. Use the image training set to train the variational neural network based on the cross-attention mechanism to obtain the image after floating image registration.
[0047] S204. Calculate the cumulative posterior corresponding to the importance-weighted evidence lower bound (iwELBO) based on the registered heart image and the fixed image as the loss function of the most prior in the variational model.
[0048] More specifically, the most prior is derived according to the following formula:
[0049]
[0050] where represents the cumulative posterior corresponding to the importance-weighted evidence lower bound (iwELBO). Assume the most prior is p * (z), when , the above objective function takes the maximum value.
[0051] Calculate the network loss function according to the following formula:
[0052]
[0053]
[0054] where M and F represent the floating image and the fixed image respectively, and p 0 (z) is a given prior distribution, and its form is where B is a diagonal matrix with as elements, q(z∣F,M) is the variational posterior distribution estimated by the network. μ(F,M) and ∑(F,M) represent the mean and variance in the variational posterior distribution q(z∣F,M) respectively. Among them represents the density ratio of the most prior and the given prior.
[0055] represents the deformed heart image and the similarity measure between the fixed image F, and the specific calculation formula is as follows:
[0056]
[0057] where N is the number of image pixels, Ω is the image domain; L v is the local image domain centered on the pixel v, and is to subtract L v from the image intensity after the average intensity.
[0058] where λ is a hyperparameter. When a larger λ is used, the similarity term sim(F, M(f z )) is relatively larger in the loss function, which means that the iwELBO is more inclined to the sample z k , which enables the registered heart image and the fixed image to achieve better similarity and is conducive to the network predicting a more accurate latent variable z.
[0059] S205. Determine whether the number of iterations has reached the set value. If the judgment result is yes, go to S207; otherwise, go to S206.
[0060] S206. Backpropagate to calculate the adjustment amount of the parameters of the radial basis variational neural network based on the cross-attention mechanism.
[0061] More specifically, use the learning rate determined by the adaptive stochastic gradient algorithm, calculate the derivative of the parameters of the radial basis variational neural network based on the cross-attention mechanism using the loss function with shape constraints, and multiply the derivative by the learning rate to obtain the adjustment amount Δw of the network parameters i .
[0062] That is, assume the i-th network parameter is w i , calculate the adjustment amount of the network parameters Then the network parameter w w is adjusted to:
[0063] w′ i = w i + γΔw i
[0064] where γ is the learning rate and is automatically determined according to the adaptive stochastic gradient descent algorithm.
[0065] S207. Receive the floating heart image and the fixed image.
[0066] S208. Use the trained image training set to process the floating heart image and the fixed image to obtain the registered image of the floating image.
[0067] In the registration method provided in this embodiment, by using the encoding-decoding structure of the variational neural network based on the cross-attention mechanism, the local and global information of the image pair is modeled through the T2T module and the Transformer layer respectively, effectively solving the problem of long-distance and short-distance feature modeling of images. The latent variable output by the Transformer layer is decoded by the radial basis function to obtain the deformation field, solving the problem of abnormal deformation field caused by non-parametric transformation. Finally, by introducing the cumulative posterior corresponding to the importance-weighted evidence lower bound (iwELBO) as the loss function with the highest priority in the variational model in the network loss function, the distribution of the latent variable estimated by the network is made more accurate. The present invention realizes cardiac image registration and can significantly improve the accuracy of cardiac image registration.
[0068] Figure 4 FIG. is a schematic structural diagram of an image registration device according to an exemplary embodiment of the present invention. As Figure 4 shown, the registration device provided by the present invention includes:
[0069] A receiving module 301, configured to receive an image training set of a cardiac floating image and a fixed image,
[0070] A processing module 302, configured to process the cardiac floating image and the fixed image by using a trained variational neural network based on the cross-attention mechanism, wherein the variational neural network based on the cross-attention mechanism includes: at least one T2T module and at least one Transformer layer for encoding the image training set;
[0071] A determining module 303, configured to decode the registration deformation field according to the radial basis function, and determine the image after registration of the floating image according to the deformation field.
[0072] Optionally, the device further includes:
[0073] A construction module 304, configured to construct a variational neural network based on the cross-attention mechanism;
[0074] A training module 305, configured to train the variational neural network based on the cross-attention mechanism by using the image training set to output the deformation field of the floating image;
[0075] An adjustment module 306, configured to calculate the cumulative posterior corresponding to the importance-weighted evidence lower bound (iwELBO) as the loss function with the highest priority in the variational model by using the registered cardiac image and the fixed image, and adjust the parameters of the variational neural network based on the cross-attention mechanism according to the loss function to output the trained variational neural network based on the cross-attention mechanism.
[0076] Optionally, at least one T2T module is cascaded, at least one Transformer layer is cascaded, and the output data of the T2T module is used as the input of the next part of the Transformer layer.
[0077] Optionally, the T2T layer includes at least one Transformer module. Among them, the T2T layer divides the image into an overlapping two-dimensional feature sequence with local information modeling and then inputs it into the Transformer module. The output feature sequence is restored to three-dimensional image features through three-dimensional transformation.
[0078] Optionally, the Transformer layer includes at least one Transformer encoder and at least one Transformer decoder. Among them, at least one Transformer encoder is cascaded, and at least one Transformer decoder is cascaded. Among them, the fixed image features output by the encoder in the Transformer layer are input into the Transformer decoder to form cross-attention with the floating image features.
[0079] Optionally, the determination module is specifically configured to:
[0080] Decode the deformation field according to the radial basis function;
[0081] Obtain the registered heart image according to the deformation field.
[0082] Optionally, the adjustment module is specifically configured to:
[0083] Calculate the importance-weighted evidence lower bound model and the loss function with implicit prior according to the registered heart image and the fixed image;
[0084] Adjust the parameters of the variational neural network based on the cross-attention mechanism according to the loss function to output the trained variational neural network based on the cross-attention mechanism.
[0085] Optionally, the determination module is specifically configured to:
[0086] Determine the deformation field of the floating image according to the following formula:
[0087]
[0088] where v is the pixel value, p i is the control point, ||v - p i || 2 is the Euclidean distance. is a radial basis function with a support set of c, and there are s different support sets in total is a latent variable whose distribution needs to be estimated.
[0089] The adjustment module is specifically used for:
[0090] Derive the most prior according to the following formula:
[0091]
[0092] is the cumulative posterior corresponding to the importance-weighted evidence lower bound (iwELBO). Assume the most prior is p * (z), when the above objective function takes the maximum value.
[0093] Calculate the network loss function according to the following formula:
[0094]
[0095] where M and F represent the floating image and the fixed image respectively, and p 0 (z) is a given prior distribution in the form of where B is a diagonal matrix with as elements, q(z∣F,M) is the variational posterior distribution estimated by the network. μ(F,M) and ∑(F,M) represent the mean and variance in the variational posterior distribution q(z∣F,M) respectively. Among them represents the density ratio of the most prior and the given prior.
[0096] represents the deformed heart image and the similarity measure between the fixed image F, and the specific calculation formula is as follows:
[0097]
[0098] where N is the number of image pixels, Ω is the image domain; L v is the local image domain centered on the pixel v, and are the image intensities after subtracting the average intensity of L v .
[0099] where λ is a hyperparameter. When using a larger λ, the similarity term sim(F,M(f z )) is relatively larger in the loss function, which means that iwELBO is more inclined to the sample z k , which enables the registered heart image and the fixed image to achieve better similarity and is conducive to the network predicting more accurate latent variable z.
[0100] Figure 5This is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present invention. As Figure 5 shown, the electronic device 400 of this embodiment includes: a processor 401 and a memory 402, wherein,
[0101] The memory 402 is used to store computer-executable instructions;
[0102] The processor 401 is used to execute the computer-executable instructions stored in the memory to implement each step executed by the receiving device in the above embodiment. For details, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0103] Optionally, the memory 402 can be either independent or integrated with the processor 401.
[0104] When the memory 402 is independently provided, the electronic device 400 further includes a bus 403 for connecting the memory 402 and the processor 401.
[0105] An embodiment of the present invention also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the above splitting method is implemented.
[0106] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or modules can be in electrical, mechanical or other forms.
[0107] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0108] In addition, in each embodiment of the present invention, the functional modules can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0109] When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0110] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily all essential to the present invention.
[0111] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0112] The above is the description of a variational image registration method, device, electronic device, and storage medium provided with a cross-attention mechanism according to the present invention. For those skilled in the art, according to the idea of the embodiments of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A variational image registration method based on cross-attention mechanism, characterized in that, it includes: Receiving a cardiac floating image and a fixed image; Processing the cardiac floating image and the fixed image by using a pre-trained variational neural network based on cross-attention mechanism to obtain the latent variable of the variational model; Inputting the latent variable into a pre-constructed radial basis function, decoding and outputting the deformation field of the floating image, and determining the registered image according to the deformation field; The variational neural network based on cross-attention mechanism includes: at least one T2T layer and at least one Transformer layer, which are used for encoding the image training set. The Transformer layer includes an encoder and a decoder. The fixed image features output by the encoder of the Transformer layer are input into the decoder of the Transformer layer to form cross-attention with the floating image features. The cumulative posterior corresponding to the importance-weighted evidence lower bound is used as the most prior loss function in the variational model to output the latent variable distribution of the variational model and output the latent variable of the variational model; Wherein, if there are at least two T2T layers, all the T2T layers are cascaded. If there are at least two Transformer layers, all the Transformer layers are cascaded. The output data after the cascading of the T2T layers is used as the input of the cascaded Transformer layers; The T2T layer includes at least one Transformer module. The T2T layer is used to divide the image into overlapping two-dimensional feature sequences with local information modeling and then input them into the Transformer module. The output feature sequences are restored to three-dimensional image features through three-dimensional transformation; Both the encoder and the decoder of the Transformer layer have at least one. If there is more than one encoder and decoder of the Transformer layer, the encoders of the Transformer layer are cascaded, and the decoders of the Transformer layer are cascaded; and the fixed image features output by the encoder of the Transformer layer are input into the decoder of the Transformer layer to form cross-attention with the floating image features.
2. The variational image registration method based on cross-attention mechanism according to claim 1, characterized in that, The training method of the variational neural network based on cross-attention mechanism includes: Calculating the cumulative posterior corresponding to the importance-weighted evidence lower bound as the most prior loss function in the variational model according to the pre-obtained registered cardiac image and the fixed image; Adjusting the parameters of the pre-constructed untrained variational neural network based on cross-attention mechanism according to the loss function to obtain a trained variational neural network based on cross-attention mechanism.
3. A variational image registration device based on cross-attention mechanism, characterized in that, it includes: A receiving module, which is used to receive the image training set of the cardiac floating image and the fixed image; A processing module, configured to process the cardiac floating image and the fixed image by using a trained variational neural network based on a cross-attention mechanism, wherein the variational neural network based on the cross-attention mechanism includes: at least one T2T layer and at least one Transformer layer for encoding an image training set, wherein the fixed image features output by the encoder in the Transformer layer are input into the Transformer decoder to form cross-attention with the floating image features, so as to output a variational model latent variable distribution; A determination module, a radial basis function is used to decode the latent variable so as to output a deformation field of the floating image, and the registered image is determined according to the deformation field; Wherein, if there are at least two T2T layers, all the T2T layers are cascaded; if there are at least two Transformer layers, all the Transformer layers are cascaded, and the output data after cascading of the T2T layers is used as the input of the cascaded Transformer layers; The T2T layer includes at least one Transformer module, and the T2T layer is configured to divide an image into overlapping two-dimensional feature sequences with local information modeling and then input them into the Transformer module, and the output feature sequences are restored to three-dimensional image features through three-dimensional transformation; Both the encoder and the decoder of the Transformer layer have at least one, and if there is more than one encoder and decoder of the Transformer layer, the encoders of the Transformer layer are cascaded, and the decoders of the Transformer layer are cascaded; and the fixed image features output by the encoder of the Transformer layer are input into the decoder of the Transformer layer to form cross-attention with the floating image features.
4. An electronic device, characterized in that, it includes: at least one processor and a memory; wherein, the memory stores computer execution instructions; the at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the variational image registration method based on the cross-attention mechanism according to any one of claims 1 to 2.
5. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores computer execution instructions, and when the processor executes the computer execution instructions, the variational image registration method based on the cross-attention mechanism according to any one of claims 1 to 2 is implemented.