A generalizable implicit neural representation unsupervised medical image registration method and system
Patent Information
- Application Number
- CN202610930609.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-06-26
AI Technical Summary
尽管基于无监督深度学习图像配准取得了优异的配准性能,尤其是基于Transformer的图像配准方法取得了SOTA的配准性能,但是基于无监督深度学习图像配准方法配准性能仍可能受到网络结构设计的影响
本发明结合隐式神经表示与Transformer优点,提出了一种创新性的基于Transformer超网络的可泛化隐式神经表示方法用于医学图像配准,旨在兼顾高精度配准性能与良好的计算效率及泛化能力。
Smart Images

Figure CN122453876B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image registration technology, specifically relating to a generalizable implicit neural representation unsupervised medical image registration method and system. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Deformable medical image registration aims to construct spatial mapping relationships between medical images at different time points, imaging devices, or imaging modalities, achieving precise alignment at the pixel or voxel level. Its core objective is to capture the potential correspondences between different images through a non-rigid transformation model, thereby establishing semantically consistent spatial correspondences between images. Traditional registration methods typically rely on predefined geometric models, interpolation strategies, and similarity metrics, seeking the optimal transformation between fixed and moving images through numerical optimization. However, these methods are highly sensitive to initial alignment, parameter selection, and image quality, easily getting trapped in local optima, and have low computational efficiency, especially performing poorly in high-resolution 3D images or large-scale datasets. Furthermore, the deformation complexity of different organs and the differences in intensity distribution among multiple modalities further increase the modeling difficulty of traditional methods.
[0004] Transformer-based image registration methods have shown advantages in accuracy, stability, and generalization ability, making them a hot research area in deep learning image registration. While unsupervised deep learning-based image registration has achieved excellent performance, especially Transformer-based methods which have achieved state-of-the-art (SOTA) performance, the performance of these methods can still be affected by the network architecture design. Specifically, for image pairs with complex deformations or subtle local differences, unsupervised deep learning-based methods often struggle to accurately model their deformation relationships or effectively capture local correspondences. Furthermore, unsupervised deep learning-based image registration methods have difficulty modeling the complex nonlinear correspondences between all image pairs, thus affecting the overall registration accuracy and the model's generalization ability to some extent.
[0005] To address the aforementioned issues, some studies have proposed deformation image registration methods based on implicit neural representation (INR, also known as neural fields). However, these methods typically employ pairwise optimization, leading to low computational efficiency and lengthy training times, making them unsuitable for real-time or large-scale medical image processing. Furthermore, INR-based deformation image registration methods generally lack generalization ability; each registration requires refactoring the neural field parameters. While this "fitting each image pair individually" mechanism can capture complex deformation relationships between image pairs, it incurs enormous computational overhead and lengthy optimization processes, especially when processing high-resolution 3D medical images. Therefore, improving the optimization efficiency and generalization ability of INR-based deformation image registration methods remains a crucial problem that urgently needs to be solved. Summary of the Invention
[0006] To address the aforementioned problems, this invention proposes a generalizable implicit neural representation unsupervised medical image registration method and system. This invention enables efficient deformation registration with strong generalization ability, balancing high-precision registration performance with good computational efficiency and generalization ability.
[0007] According to some embodiments, the present invention adopts the following technical solution: A generalizable unsupervised medical image registration method based on implicit neural representations includes the following steps: Obtain fixed and moving images from historical medical images, form image pairs, and use them as training data to construct a transformer-based supernetwork; A hypernetwork is trained using training data to encode image pairs and extract data labels in order to capture structural and texture information between images. The implicit neural representation weights are viewed as a set of column vectors of the weight matrix of each layer, and a learnable weight label is created for each column vector; The learnable weight labels and the extracted data labels are input into the hypernetwork. The hypernetwork uses a multi-head self-attention mechanism to fuse the information of the image pair with the learnable weight labels and outputs a vector representation corresponding to each learnable weight label. The output is mapped to the implicit neural representation weights based on the initial position of the weight labels, a deformation field is constructed, and end-to-end backpropagation training is performed through the registration loss function to optimize the parameters of the supernetwork and complete the training of the supernetwork. The fixed and moving images of the target to be registered are obtained to form image pairs. The trained supernetwork is used to generate the optimal implicit neural representation weights, thereby achieving registration based on implicit neural representation and obtaining the registration result.
[0008] As an alternative implementation, the process of encoding image pairs includes: encoding images of size . The stationary and moving images are stitched together to form two channels as input. The input stationary and moving volumes are then divided into non-overlapping blocks, totaling [number missing]. There are 1 block, and the size of each block is 1. Block vectors are data tags.
[0009] As a further defined implementation, a linear projection layer is used to project each block vector onto a feature representation of arbitrary dimensions: ; in, and Base mark It represents a block. This represents a linear embedding layer. This is the output.
[0010] As an alternative implementation, the implicit neural representation weights are viewed as a set of column vectors of the weight matrix of each layer, and the process of creating learnable weight labels for each column vector includes: treating the weight matrix of each hidden layer as a set of column vectors. Viewed as a set of column vectors, complete parameters It is jointly represented by a set of column vectors, and for each column vector, a corresponding initialization flag is introduced. This refers to the learnable weight labels. The dimension of the hidden layer.
[0011] As an alternative implementation, the supernetwork is jointly modeled through three types of interaction mechanisms: constructing data feature representations between registered image pairs through the interaction between data tags; transferring data feature information between registered image pairs to the weights of the supernetwork through the interaction between data tags and learnable weight tags; and capturing the potential relationships between different weights in the neural field through the interaction between learnable weight tags, and representing the output vector of the learnable weight tags as weight tags.
[0012] As an alternative implementation, the hypernetwork is a weight matrix. Each A separate fully connected layer was set up to map the corresponding weight labels to... In the neural field weight set, the column vectors are grouped using a weight grouping strategy. Each column vector in the weight matrix is divided into groups, and a separate label is assigned to each group to calculate the weights, thus obtaining the complete neural field weight set. .
[0013] As an alternative implementation, the supernetwork includes 12 alternating Transformer encoders consisting of multi-head self-attention and MLP blocks, with a LayerNorm layer applied before each multi-head self-attention and MLP block and a residual connection applied after the multi-head self-attention and MLP block.
[0014] As an alternative implementation, the trained supernetwork is used to generate the optimal implicit neural representation weights, thereby achieving registration based on the implicit neural representation. In this process, two average pooling layers and two convolutional layers are used to downsample to 1 / 16 of the original size. Each convolutional layer uses a 3×3×3 kernel with a stride of 2 and padding of 1, and each convolutional layer is followed by a LeakyReLU activation function to introduce nonlinearity.
[0015] As an alternative implementation, the trained hypernetwork is used to generate optimal implicit neural representation weights, thereby achieving registration based on implicit neural representations through functional analysis. The features extracted by the convolution operation are mapped to a high-dimensional embedding, where the functional is expressed as: ; in For functional index, Given the number of functionals, the coordinates are encoded using Fourier mapping to obtain: ; in, In order to be in Uniform sampling in a logarithmic manner within the range. It is the control parameter for the highest frequency. The larger the value, the more sensitive the model is to high-frequency signals.
[0016] A generalizable implicit neural representation unsupervised medical image registration system includes: The preprocessing module is configured to acquire fixed and moving images from historical medical images, form image pairs, and use them as training data to construct a transformer-based supernetwork. The data label extraction module is configured to train a supernetwork using training data, encode image pairs, and extract data labels to capture structural and texture information between images. The weight labeling module is configured to treat the implicit neural representation weights as a set of column vectors of the weight matrix of each layer, and to create a learnable weight label for each column vector. The dual-label fusion module is configured to input the learnable weight labels and the extracted data labels into the supernetwork. The supernetwork uses a multi-head self-attention mechanism to fuse the information of the image pair with the learnable weight labels and outputs a vector representation corresponding to each learnable weight label. The hypernetic network training module is configured to map the output results to the implicit neural representation weights based on the initial position of the weight labels, construct the deformation field, and perform end-to-end backpropagation training through the registration loss function, thereby optimizing the parameters of the hypernetic network and completing the training of the hypernetic network. The implicit neural representation registration module is configured to acquire the fixed and moving images of the target to be registered, form image pairs, generate optimal implicit neural representation weights using the trained hypernetwork, and then achieve registration based on implicit neural representation to obtain the registration result.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention combines the advantages of implicit neural representation and Transformer, proposing an innovative generalizable implicit neural representation method based on Transformer supernetworks for medical image registration, aiming to balance high-precision registration performance with good computational efficiency and generalization ability.
[0018] In the training phase, this invention encodes the image pairs to be registered as data tags to capture the structural and texture information between the images. Simultaneously, to generate the corresponding deformation field network weights in the implicit neural representation, the INR weights are treated as a set of column vectors of the weight matrix of each layer, and a learnable weight tag is created for each column vector. These learnable weight tags, along with the data tags extracted from the image pairs, are input into a Transformer-based supernetwork. The supernetwork uses a multi-head self-attention mechanism to fuse the information of the image pairs with the learnable weight tags and outputs a vector representation corresponding to each learnable weight tag. Based on the initial positions of these tags, the output results are mapped to the INR weights, thereby constructing a deformation field generation network. Throughout the process, end-to-end backpropagation training is performed using a designed registration loss function to optimize the parameters of the Transformer supernetwork. In the inference phase, the trained Transformer supernetwork can quickly generate the optimal weights of the INR network based on any new input image pair, achieving efficient and highly generalizable deformation registration.
[0019] This invention constructs a Transformer-based supernetwork and, through joint modeling of the complex structural relationships between image pairs and the INR weight mapping rules during the training phase, realizes for the first time a generalizable implicit neural representation unsupervised medical image registration method, filling a research gap in this direction.
[0020] This invention constructs a Transformer-based supernetwork that integrates image pair information and INR structure priors. Data labels and learnable weight labels are used as inputs to the supernetwork, and a multi-head self-attention mechanism is employed to model the interaction between image features and INR parameters. Therefore, the proposed Transformer-based supernetwork effectively improves the robustness of the weight generation process to input images lacking visible features.
[0021] This invention treats the column vectors of the weight matrices in each layer of the INR as independent learning units and introduces learnable weight labels. By explicitly modeling each column vector, it achieves precise expression and control of the INR weight structure. This modeling approach helps improve registration accuracy and the controllability of INR weights.
[0022] This invention enables the Transformer-based supernetwork, trained to quickly generate INR weights for unseen image pairs during the inference phase without re-optimization. This transforms the implicit neural representation-based deformation registration method from pairwise optimization to a fast inference approach that does not require retraining, significantly improving processing efficiency and generalization ability.
[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0024] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0025] Figure 1 This is a schematic diagram of the overall framework of an image registration method according to one embodiment; Figure 2 This is a flowchart illustrating an image registration method according to one embodiment. Detailed Implementation
[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0027] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0028] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0029] Where there is no conflict, the embodiments and features described in this application may be combined with each other.
[0030] Example 1 First, some explanation is needed. Existing unsupervised deep learning-based registration methods learn mapping relationships from a large number of image pairs using neural networks. They optimize parameters using a loss function and apply the mapping relationships to the images to be registered through a spatial transformation network, achieving fast and efficient registration. Classical unsupervised deep learning-based registration methods typically use convolutional neural networks for image registration. However, due to the locality limitation of convolution operations, this approach struggles to model the dependencies between long-distance features in the image, thus affecting registration accuracy and global consistency.
[0031] In recent years, the multi-head self-attention mechanism of Transformers has overcome the limitation of the receptive field of CNNs due to its ability to model dependencies between arbitrary locations globally. Therefore, unsupervised deep learning-based image registration methods have gradually shifted from traditional CNN frameworks to novel registration networks centered around Transformers.
[0032] While unsupervised deep learning-based registration methods have improved registration performance to some extent by introducing new techniques and advanced network structures, pre-trained neural networks cannot generate complex deformations for every pair of images in unseen image datasets, resulting in limited performance improvements. The proposed method constructs the deformation field generation network based on the weights of the registered image pairs, enabling the model to generate accurate complex deformations even on unseen data, thereby significantly improving registration accuracy.
[0033] Deformation image registration methods based on implicit neural representations aim to achieve high-precision modeling of complex nonlinear deformations by constructing an independent neural network model for each image to be registered. However, existing methods lack or do not fully generalize, resulting in low computational efficiency and slow convergence speed, making them difficult to meet clinical needs.
[0034] like Figure 1As shown, this invention provides a generalizable registration method based on a Transformer-based hypernetwork implicit neural representation image registration model (referred to as such in this paper). The Transformer-based hypernetwork generates INR parameters based on the input fixed and moving images. The grid coordinates of the moving image are used to generate a high-dimensional embedding through a coordinate encoding module as the input to the INR. Finally, the output of the INR is upsampled to obtain the final deformation field.
[0035] A generalizable unsupervised medical image registration method based on implicit neural representations, such as Figure 2 As shown, it includes the following steps: Obtain fixed and moving images from historical medical images, form image pairs, and use them as training data to construct a transformer-based supernetwork; A hypernetwork is trained using training data to encode image pairs and extract data labels in order to capture structural and texture information between images. The implicit neural representation weights are viewed as a set of column vectors of the weight matrix of each layer, and a learnable weight label is created for each column vector; The learnable weight labels and the extracted data labels are input into the hypernetwork. The hypernetwork uses a multi-head self-attention mechanism to fuse the information of the image pair with the learnable weight labels and outputs a vector representation corresponding to each learnable weight label. The output is mapped to the implicit neural representation weights based on the initial position of the weight labels, a deformation field is constructed, and end-to-end backpropagation training is performed through the registration loss function to optimize the parameters of the supernetwork and complete the training of the supernetwork. The fixed and moving images of the target to be registered are obtained to form image pairs. The trained supernetwork is used to generate the optimal implicit neural representation weights, thereby achieving registration based on implicit neural representation and obtaining the registration result.
[0036] The following is a detailed description.
[0037] Unsupervised deformation image registration based on deep learning refers to the process of learning the optimal spatial transformation relationship between two images through a deep neural network without the supervision of registration labels. It can be represented as follows: ; in, and These represent stationary and moving images, respectively. Represents the deformation field. Mapping arrive . This indicates that a distorted image is generated by twisting and moving the image through a deformation field. This represents the optimal deformation field. Loss function. It measures the similarity between a fixed image and a distorted image. Hyperparameters used to constrain the smoothness of the deformation field This plays a moderating role in maintaining the balance between image matching accuracy and deformation field smoothness. In existing deformation image registration methods based on implicit neural representations, the deformation field... From a learnable parameter neural function with weights Also known as implicit neural representation parameterization, one typical example is... A multilayer perceptron is employed. Furthermore, the learnable parameters... It consists of a set of matrices: ; in, This represents the current layer number of the multilayer perceptron. The depth of the multilayer perceptron is given. Furthermore, biases are incorporated into these matrices. Therefore, given both fixed and moving images, the goal of existing deformation image registration methods based on implicit neural representations is to obtain the optimal... ,make Calculate However, Fitting to a given registration image typically requires gradient descent optimization from scratch for this type of method, resulting in inefficiency and a lack of generalization.
[0038] To address this issue, a generalizable implicit neural representation-based unsupervised deformable image registration method based on a Transformer supernetwork is proposed. The goal of the proposed method is to train a Transformer-based supernetwork that can infer from unseen image pairs... weight .
[0039] Specifically, such as Figure 1 As shown, the size is The stationary and moving images are stitched together into two channels as input. Following a strategy similar to Vision Transformer, the fixed and moving volumes of the input are segmented into non-overlapping 3D patches (blocks or small pieces) using an Image Tokenizer, resulting in a total of... There are 1 block, and the size of each block is 1. ,in Typically set to 4, these patch vectors are represented as data tokens. Then, a linear projection layer is used to project each token onto a feature representation of arbitrary dimensions (denoted as C): ; in, and Base mark It represents "patch". This represents a linear embedding layer. This is the output. For decoding. parameters In this embodiment, the weight matrix of each hidden layer is... Viewed as a set of column vectors, complete parameters This can be represented jointly by these column vector sets. For each column vector, this embodiment introduces a corresponding initialized token. (A learnable vector parameter, (The dimension of the hidden layer). Figure 1 The orange squares represent learnable weight tokens.
[0040] In this embodiment, data tokens and learnable weight tokens are jointly input into a Transformer-based supernetwork. This supernetwork jointly models data through three types of interaction mechanisms: first, it constructs data feature representations between registered image pairs through interactions between data tokens; second, it transfers data feature information between registered image pairs to the weights of the supernetwork through interactions between data tokens and learnable weight tokens; and third, it captures the potential relationships between different weights in the neural field through interactions between learnable weight tokens. Finally, the output vector of the learnable weight tokens is represented as the weight tokens.
[0041] Due to different weight matrices They may have different column vector dimensions, therefore, for each A separate fully connected layer was set up to map the corresponding weight tokens to... The column vectors in the weight matrix are used, and to balance computational cost and accuracy, a weight grouping strategy is adopted, dividing the column vectors in each weight matrix into groups and assigning a separate token to each group to calculate the weight. Finally, a complete set of neural field weights is obtained. .
[0042] For the implicit neural representation registration network, it takes a 3D spatial coordinate grid of the same size as the fixed and moving images as input. (Note that in this embodiment, the fixed and moving images are only used for inference.) Similarity metric calculation and warped image generation. Through coordinate encoding, the 3D spatial coordinate grid is mapped to a high-dimensional embedding, then input into the INR and the output is reshaped to obtain a resolution of [resolution missing]. The feature map, where the weights of INR are... Estimation is performed using a Transformer-based hypernetwork. TGINRMorph uses two average pooling modules to downsample the input registered image pairs to generate two feature maps with resolutions of [resolutions to be filled in]. and The output features are then concatenated with the INR output features via skip connections to fully utilize feature information at different resolutions. This embodiment employs PixelShuffle to upsample the output features to overcome feature degradation, ultimately obtaining a Stationary Velocity Field (SVF).
[0043] The velocity field is integrated using scaling and squaring methods to obtain the deformation field. This deformation field is then used to distort the moving image through a spatial transformation network to achieve the registration result. Finally, the parameters of the entire registration network are updated via backpropagation using a calculated loss function.
[0044] In this embodiment, the Transformer-based hypernetwork comprises 12 alternating Transformer encoders consisting of multi-head self-attention and MLP blocks. A LayerNorm layer is applied before each multi-head self-attention and MLP block, and a residual connection is applied after each multi-head self-attention and MLP block. The first layer of the Transformer encoder is... The layer output can be written as: ; ; in Indicates the first in the encoder The image output of the layer is represented by a representation tensor. The self-attention mechanism is calculated as follows: ; in, yes It is a matrix obtained through linear mapping. yes or The feature dimensions. Initial input. and Finally, parallel execution. This self-attention operation is then extended to multi-head self-attention (MSA).
[0045] In summary, firstly, after concatenating the fixed and moving images into a data label sequence, this embodiment concatenates the data label sequence with a learnable weight label sequence as input to the Transformer-based hypernetwork. Then, the Transformer encoder extracts a set of latent tokens, where each latent token corresponds to an input learnable token. It is worth emphasizing that the multi-head self-attention mechanism in the Transformer encoder possesses permutation equivariance, thus eliminating the need for prior settings regarding local features of the input data and the order of latent tokens. During training, each latent token can autonomously learn and model the local semantic information of the input data, while simultaneously capturing global contextual information through the multi-head attention mechanism, thereby achieving a global representation of the input data instance.
[0046] like Figure 1 As shown, the proposed method takes a 3D spatial coordinate grid as input and uses implicit neural representation to model the continuous velocity field. ,in Indicates the INR weights. For spatial points in the grid. The proposed method can generate the corresponding velocity vector. Due to GPU memory limitations, inputting the complete 3D coordinate mesh into the implicit neural representation registration network is not feasible. Therefore, this embodiment uses two average pooling layers and two convolutional layers to downsample to 1 / 16 of the original size. Each convolutional layer uses a 3×3×3 kernel with a stride of 2 and padding of 1, and each convolutional layer is followed by a LeakyReLU activation function to introduce non-linearity.
[0047] This embodiment chooses to use convolution for downsampling instead of trilinear interpolation to downsample the 3D coordinate grid because convolution can introduce nonlinear operations and effectively extract high-frequency features while allowing interaction between voxels in high-dimensional space.
[0048] Furthermore, in order for the INR to perceive spatial differences and high-frequency details, thereby more accurately fitting complex spatial variations, this embodiment employs a series of optional functionals. The features extracted by the convolution operation are mapped to a high-dimensional embedding, where the functional is expressed as: ; in For functional index, Let be the number of functionals. By encoding the coordinates using Fourier mapping, we can obtain: ; in, In order to be in Uniform sampling in a logarithmic manner within the range. It is the control parameter for the highest frequency. The larger the value, the more sensitive the model is to high-frequency signals.
[0049] For the IRN, this embodiment employs a 5-layer multilayer perceptron, where each hidden layer has 256 hidden units. Specifically, the IRN is designed as follows: ; ; ; in, They represent the first Layer weights, biases, and activation functions. In this embodiment, the activation function... Gaussian Error Linear Units (GELU) are used.
[0050] Example 2 A generalizable implicit neural representation unsupervised medical image registration system includes: The preprocessing module is configured to acquire fixed and moving images from historical medical images, form image pairs, and use them as training data to construct a transformer-based supernetwork. The data label extraction module is configured to train a supernetwork using training data, encode image pairs, and extract data labels to capture structural and texture information between images. The weight labeling module is configured to treat the implicit neural representation weights as a set of column vectors of the weight matrix of each layer, and to create a learnable weight label for each column vector. The dual-label fusion module is configured to input the learnable weight labels and the extracted data labels into the supernetwork. The supernetwork uses a multi-head self-attention mechanism to fuse the information of the image pair with the learnable weight labels and outputs a vector representation corresponding to each learnable weight label. The hypernetic network training module is configured to map the output results to the implicit neural representation weights based on the initial position of the weight labels, construct the deformation field, and perform end-to-end backpropagation training through the registration loss function, thereby optimizing the parameters of the hypernetic network and completing the training of the hypernetic network. The implicit neural representation registration module is configured to acquire the fixed and moving images of the target to be registered, form image pairs, generate optimal implicit neural representation weights using the trained hypernetwork, and then achieve registration based on implicit neural representation to obtain the registration result.
[0051] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of one or more computer-usable storage media (including, but not limited to, disk storage, etc.) containing computer-usable program code. CD - ROM It takes the form of a computer program product implemented on (such as optical memory, etc.).
[0052] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0053] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0054] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0055] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made by those skilled in the art without creative effort within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A generalizable unsupervised medical image registration method using implicit neural representations, characterized in that, Includes the following steps: Obtain fixed and moving images from historical medical images, form image pairs, and use them as training data to construct a transformer-based supernetwork; A hypernetwork is trained using training data to encode image pairs and extract data labels in order to capture structural and texture information between images. The implicit neural representation weights are viewed as a set of column vectors of the weight matrix of each layer, and a learnable weight label is created for each column vector; The learnable weight labels and the extracted data labels are input into the hypernetwork. The hypernetwork uses a multi-head self-attention mechanism to fuse the data labels and learnable weight labels, and outputs a vector representation corresponding to each learnable weight label. The output is mapped to the implicit neural representation weights based on the initial position of the weight labels, a deformation field is constructed, and end-to-end backpropagation training is performed through the registration loss function to optimize the parameters of the supernetwork and complete the training of the supernetwork. The fixed and moving images of the target to be registered are obtained to form image pairs. The trained hypernetwork is used to generate the optimal implicit neural representation weights, thereby achieving registration based on implicit neural representation and obtaining the registration result. The supernetwork is jointly modeled through three types of interaction mechanisms: constructing data feature representations between registered image pairs through the interaction between data tags; transferring data feature information between registered image pairs to the weights of the supernetwork through the interaction between data tags and learnable weight tags; and capturing the potential relationships between different weights in the neural field through the interaction between learnable weight tags, and representing the output vector of the learnable weight tags as weight tags.
2. The generalizable implicit neural representation unsupervised medical image registration method as described in claim 1, characterized in that, The process of encoding image pairs includes: encoding images of size 10 ... The stationary and moving images are stitched together to form two channels as input. The input stationary and moving volumes are then divided into non-overlapping blocks, totaling [number missing]. There are 1 block, and the size of each block is 1. Block vectors are data tags.
3. The generalizable implicit neural representation unsupervised medical image registration method as described in claim 2, characterized in that, Each block vector is projected onto a feature representation of arbitrary dimensions using a linear projection layer: in, and Base mark It represents a block. This represents a linear embedding layer. This is the output.
4. The generalizable implicit neural representation unsupervised medical image registration method as described in claim 1, characterized in that, The process of treating the implicit neural representation weights as a set of column vectors of the weight matrix of each layer, and creating learnable weight labels for each column vector, includes: treating the weight matrix of each hidden layer as a set of column vectors of .... Viewed as a set of column vectors, complete parameters It is represented jointly by a set of column vectors, and for each column vector, a corresponding initialization flag is introduced. This refers to the learnable weight labels. The dimension of the hidden layer.
5. The generalizable implicit neural representation unsupervised medical image registration method as described in claim 1, characterized in that, The hypernetwork is a weight matrix. Each A separate fully connected layer was set up to map the corresponding weight labels to... In the neural field weight set, the column vectors are grouped using a weight grouping strategy. Each column vector in the weight matrix is divided into groups, and a separate label is assigned to each group to calculate the weights, thus obtaining the complete neural field weight set. .
6. The generalizable implicit neural representation unsupervised medical image registration method as described in claim 1, characterized in that, The supernetwork comprises 12 alternating Transformer encoders consisting of multi-head self-attention and MLP blocks, with a LayerNorm layer applied before each multi-head self-attention and MLP block and a residual connection applied after each multi-head self-attention and MLP block.
7. The generalizable implicit neural representation unsupervised medical image registration method as described in claim 1, characterized in that, In the process of generating optimal implicit neural representation weights using the trained supernetwork and then achieving registration based on implicit neural representation, two average pooling layers and two convolutional layers are used to downsample to 1 / 16 of the original size. Each convolutional layer uses a 3×3×3 kernel with a stride of 2 and padding of 1, and each convolutional layer is followed by a LeakyReLU activation function to introduce nonlinearity.
8. The generalizable implicit neural representation unsupervised medical image registration method as described in claim 1, characterized in that, The trained hypernetwork is used to generate optimal implicit neural representation weights, thereby enabling registration based on implicit neural representations through functional analysis. The features extracted by the convolution operation are mapped to a high-dimensional embedding, where the functional is expressed as: in For functional index, Given the number of functionals, the coordinates are encoded using Fourier mapping to obtain: in, In order to be in Uniform sampling in a logarithmic manner within the range. It is the control parameter for the highest frequency. The larger the value, the more sensitive the model is to high-frequency signals.
9. A generalizable implicit neural representation unsupervised medical image registration system, characterized in that, include: The preprocessing module is configured to acquire fixed and moving images from historical medical images, form image pairs, and use them as training data to construct a transformer-based supernetwork. The data label extraction module is configured to train a supernetwork using training data, encode image pairs, and extract data labels to capture structural and texture information between images. The weight labeling module is configured to treat the implicit neural representation weights as a set of column vectors of the weight matrix of each layer, and to create a learnable weight label for each column vector. The dual-label fusion module is configured to input the learnable weight labels and the extracted data labels into the supernetwork. The supernetwork uses a multi-head self-attention mechanism to fuse the data labels and learnable weight labels and outputs a vector representation corresponding to each learnable weight label. The hypernetic network training module is configured to map the output results to the implicit neural representation weights based on the initial position of the weight labels, construct the deformation field, and perform end-to-end backpropagation training through the registration loss function, thereby optimizing the parameters of the hypernetic network and completing the training of the hypernetic network. The implicit neural representation registration module is configured to acquire the fixed and moving images of the target to be registered, form image pairs, generate the optimal implicit neural representation weights using the trained hypernetwork, and then achieve registration based on implicit neural representation to obtain the registration result. The supernetwork is jointly modeled through three types of interaction mechanisms: constructing data feature representations between registered image pairs through the interaction between data tags; transferring data feature information between registered image pairs to the weights of the supernetwork through the interaction between data tags and learnable weight tags; and capturing the potential relationships between different weights in the neural field through the interaction between learnable weight tags, and representing the output vector of the learnable weight tags as weight tags.
Citation Information
Patent Citations
Unsupervised deformable three-dimensional medical image registration method and device
CN118314175A
Implicit neural network three-dimensional geological modeling method based on multi-source data fusion
CN121962482A