A CT-CBCT deformable registration method and system based on prompt feedback convolutional neural network
By introducing a method based on prompt-strengthening convolutional neural network in CT-CBCT deformation registration, combined with human evaluation feedback and arrow prompts, the problem of difficulty in obtaining labels and neglecting local large deformation is solved, and the accuracy and usability of deformation registration are improved.
Patent Information
- Application Number
- CN202410776891.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-17
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-06-17
AI Technical Summary
There are problems in CT-CBCT deformation registration that are difficult to obtain labels, local large deformations are easily ignored, and the inability to quickly evaluate feedback and correct and improve registration results.
The method of strengthening convolutional neural network based on prompts is adopted, and the CT-CBCT deformation registration system is improved by introducing human evaluation feedback and arrow prompts, which solves the problem of difficulty in obtaining labels and neglecting local large deformations, and allows users to quickly evaluate and correct the registration results.
The accuracy and availability of CT-CBCT deformation registration are improved, and the deformation field overlap and folding problems caused by simply relying on mean square error of grayscale value is solved, which enhances the registration effect of the core area and reduces the dependence on high-difficulty annotation.
Smart Images

Figure CN118982564B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image registration and fusion in nuclear technology applications, and in particular to a CT-CBCT deformation registration method and system based on a prompt feedback convolutional neural network. Background Art
[0002] Accurate image guidance is one of the important technologies for implementing precision radiotherapy. With the mature application of CBCT technology in image-guided radiotherapy, the registration and fusion of CT and CBCT has become one of the effective means to improve the accuracy of radiotherapy. However, unlike simple rigid registration, which only moves and rotates in three dimensions, different pixels may have different deformation amplitudes and directions under deformation registration.
[0003] Grayscale-based deep learning algorithms are mainly concentrated in two mainstream methods: supervised and unsupervised. To build a deformable registration system through supervised methods, it is often necessary to manually outline the region of interest (ROI) before model training, and the accuracy of the outline directly affects the accuracy of the deformable registration, especially in CBCT images with many artifact scattering, it is particularly difficult to accurately outline the ROI. Although the unsupervised method does not require additional manual annotation, it is often used to deform the whole, so the key deformation of a key part is often not given enough attention.
[0004] Although there have been studies using various convolutional neural networks and their deformation techniques (UNet or CycleGAN), combined with various auxiliary registration methods such as: constructing a transition network between CBCT and CT through a model, changing kilovolt CBCT (KV-CBCT) to megavolt CBCT (MV-CBCT), etc., the accuracy of deformation registration has been further improved from a theoretical level, but the key element of these methods - manual annotation, is often difficult to obtain easily in clinical applications, so it cannot be truly applied to radiotherapy. In addition, related work has not solved the problem of doctors manually correcting or improving the deformation registration results of private images that are not very effective. Summary of the invention
[0005] Purpose of the invention: The purpose of the present invention is to provide a CT-CBCT deformation registration method and system based on a prompt-enhanced convolutional neural network to solve the problems existing in CT-CBCT deformation registration, such as the difficulty in obtaining key annotations, the easy neglect of local large deformations, and the inability to quickly evaluate and feedback the private images during use, and to correct and improve the registration results.
[0006] Technical solution: The present invention provides a CT-CBCT deformation registration method based on a prompt-enhanced convolutional neural network, comprising the following steps:
[0007] The method for CT-CBCT deformable registration includes the following steps:
[0008] (1) Acquire data, divide the acquired data into data sets and perform preliminary processing;
[0009] (1.1) Data collection and data set division: the data set is divided into training set, validation set and test set in a ratio of 7:1:2, and the CT-CBCT image pair of the same patient can only appear in one of the sets;
[0010] (1.2) Preprocess the data set;
[0011] (1.3) Extract the ROI rectangle from the preprocessed data set;
[0012] (2) Construct a prompt-enhanced convolutional neural network;
[0013] (2.1) Build the benchmark image codec module;
[0014] (2.2) Add a prompt encoder and add a prompt encoder module with an attention mechanism to the constructed network body;
[0015] (2.3) Modify the image encoder and add a spatial transformation layer after the output layer of the image decoder;
[0016] (3) Training of the model without prompts: The extracted local CT-CBCT image pairs are used as input, the weighted MSE and the second-order norm of the deformation field are used as the loss function, and the network is trained without prompts;
[0017] (4) Arrow prompt model training: The loss function remains unchanged, the optimal model parameters obtained in (3) are used as the initial parameters, and arrow prompts are added to continue model training;
[0018] (5) Reasoning and self-training: fine-tune the arrow prompt model based on the image pairs input by the user, the arrow prompts, and the feedback evaluation of the deformation registration results.
[0019] Furthermore, the step (1.2) of preprocessing the data set further includes:
[0020] (1.2.1) Perform spacing alignment, resampling, rigid alignment, cropping, 2D slice alignment, and sampling on CT-CBCT image pairs in three dimensions;
[0021] (1.2.2) Take a 2*2 pixel block as a pixel block, and set the IDs of all pixel blocks in the image from left to right and from top to bottom in a row-first manner, with the IDs ranging from 0, 1...N.
[0022] Furthermore, the step (1.3) of extracting ROI rectangles in the same frame of the preprocessed data set further includes: (1.3.1) synchronously selecting the ROIs in the CT-CBCT image pair by rectangular frames, and it is necessary to ensure that the rectangular frames are all captured at the same position and that the respective ROIs are completely contained in the respective rectangular frames;
[0023] (1.3.2) Fill the local image to 256*256 by evenly filling 0 on all sides;
[0024] (1.3.3) The local images after zero padding are encoded from top to bottom and from left to right. Each local image can obtain a 255*255 pixel block matrix, and the pixel blocks are numbered from (0,0) to (254,254).
[0025] Furthermore, the step (2.1) of building a reference image codec module further includes:
[0026] (2.1.1) The benchmark image encoder and image decoder use the UNet architecture to build the model benchmark body;
[0027] (2.1.2) The image encoder is modified by adding an embedding layer after the image input layer, and the embedding output matrix size is set to 256*256;
[0028] (2.1.3) If there is a hint input, a matrix bitwise addition operator is set after the embedding layer of the image encoder to add the output matrix of the embedding layer of the image encoder and the output matrix of the hint encoder bitwise. The result of the addition is input to the convolution module of the image encoder for encoding.
[0029] Furthermore, the step (2.2) of adding a prompt encoder further comprises:
[0030] (2.2.1) Add the first layer of embedding module to the prompt encoder, and the embedding output matrix size is 256*256;
[0031] (2.2.2) Add a second layer of position encoding module to the prompt encoder, and the position encoding output matrix size is also 256*256;
[0032] (2.2.3) Add a third layer of attention module to the prompt encoder, and the attention output matrix size is 256*256;
[0033] Furthermore, the step (2.3) of transforming the image encoder further includes:
[0034] (2.3.1) A spatial transformation layer is added after the output layer of the image decoder. This layer receives two inputs: the deformation field output by the image decoder and the floating image. The linear interpolation method is used to calculate the pixel values of the fixed grid points based on the floating image and the deformation field to complete the output of the registered image. This layer does not involve network parameter updates and does not participate in back-propagation calculations.
[0035] Furthermore, the step (3) of silent model training further includes:
[0036] (3.1) Each time, BS (the size of a training batch, ranging from 2 to 64) local image pairs are randomly selected from the training set to form a training batch. The weighted sum of the average grayscale value MSE of the normalized deformed image and the fixed image and the second-order norm of the deformation field is used as the loss function. The weighting ratio is 1000:10. The Adam optimizer is used, and the epochs is set to M1·N. train / BS, where N train represents the total number of image pairs in the training set, M1 represents a positive integer, usually 2-10, BS is the size of a training batch, and the learning rate for the first 50% of epochs is λ1 (λ1 can be 1*10 -4 or 0.5*10 -4 ) The learning rate of the last 50% of epochs is λ1 / 10, and each epoch is iterated 150 to 300 times, that is, steps / epoch is set to 150-300, and the model is trained on a single A40 card;
[0037] (3.2) The model adopts an optimized update archiving method, and the model parameters are updated and saved only when the model after each round of training shows better results in the validation set;
[0038] (3.3) Gradually adjust the weighted ratio, keep the other hyperparameters unchanged, continue to train the model, and save the optimal model.
[0039] Furthermore, the arrow prompt model training in step (4) further includes:
[0040] (4.1) Obtaining arrow tips: Each pair of local images is superimposed and displayed with a transparency of 30%-70% to ensure that the ROI edges of the foreground and background images can be clearly distinguished; professional doctors mark the major deformation processes in the superimposed images with arrows. The starting point of the arrow should be at the core position of the key deformation area of the ROI of the floating image (generally the deformation core), and the end point of the arrow should be at the core position of the key deformation area of the ROI of the fixed image (generally the deformation core). The length and direction of the arrow represent the deformation degree and direction of this core position. Finally, only the annotations with prompt arrow masks are retained;
[0041] (4.2) The optimal model parameters obtained by training without prompting are used as the initial parameters of the network except the prompt encoder. The Adam optimizer is used, and the learning rate is set to λ2 (λ2 can be set to 0.5*10 -4 or 1*10 -5 ), epochs set to M2 N train / BS, where N train Represents the total number of image pairs in the training set, M2 represents a positive integer, generally 2-10, BS is the size of each training batch, 2-64, steps / epoch is set to 150-300, the input is a local image pair and the arrow prompt mask annotation corresponding to this image pair; each random sampling forms a training batch, and the loss function of the model is still set to a batch-normalized weighted sum of the average MSE of the pixel values of the deformed image and the fixed image and the second-order norm of the average deformation field, with a weighted ratio of 1000:10, and then the model training is carried out;
[0042] (4.3) The model adopts an optimized update archiving method, and the model parameters are updated and saved only when the model after each round of training shows better results in the validation set;
[0043] (4.4) Gradually adjust the weighting ratio, keep the other hyperparameters unchanged, continue to train the model, and save the optimal model.
[0044] Furthermore, the step (5) of fine-tuning the arrow prompt model further includes:
[0045] (5.1) Freeze the image encoder and prompt encoder parameters, open all parameters of the image decoder, and use the weighted loss function of MSE and the second-order norm of the deformation field. The weight coefficient uses the value finally determined in (4.4). The Adam optimizer is used, and the learning rate is λ3 (λ3 is 1*10 -5 or 0.5*10 -5 ), the training step of each image pair is set to 300-500;
[0046] (5.2) The model is fine-tuned based on the local image pairs, arrow prompts, and deformation effect feedback actually input by the user; (5.3) After fine-tuning is completed, the deformation registration results before and after fine-tuning are output. If the user selects the result after fine-tuning, the fine-tuned model is retained. If the user does not select the result after fine-tuning, the model is not retained.
[0047] The present invention also provides a system for CT-CBCT deformation registration based on prompt feedback convolutional neural network, comprising: a data processing module: mainly used for synchronous acquisition of local images on CT-CBCT image pairs, and performing pre-processing operations such as resampling, filling, cropping, rigid alignment, and layer alignment on the global image pairs;
[0048] Registration model training module: mainly used for the pre-built neural network, using the pre-processed data obtained by the data processing module as the training data set and the verification data set, using the arrow prompts provided by the user, to complete the model training with arrow prompts, and obtain the optimal prompt feedback convolutional neural network model, and can accept out-of-set image pairs, perform deformation registration through the model, and output the deformed registered image and deformation field;
[0049] Effect feedback module: It is mainly used to visualize the image and deformation field after deformation registration, and provide human evaluation feedback function based on the above two results, allowing users to tell the system in an interactive way whether they are satisfied with the deformation registration results of this round.
[0050] Reasoning and self-training module: It is mainly used for users to perform deformable registration on their private image pairs after the training is completed and the benchmark model is correctly loaded. According to the user's arrow prompts and result feedback, the model can be automatically fine-tuned on these new out-of-set image pairs to further improve the registration effect of private image pairs.
[0051] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are as follows:
[0052] (1) By introducing a human evaluation feedback method, the model output is fitted in a direction that humans feel is better, solving the deformation field overlap, folding and other distortion problems caused by simply taking the mean square error of grayscale value as the optimization target, and improving the registration effect;
[0053] (2) By selecting the same frame for CT-CBCT, the local image is deformed and registered based on the prompt, which solves the problem of large local deformation being covered in the global image and further improves the registration effect of the core area;
[0054] (3) By introducing arrow prompts, the problem that users cannot intervene in deformation registration is solved; users can quickly privatize the model data to their own hospitals, their own treatment machines, and their own radiotherapy systems by providing personalized arrow prompts.
[0055] (4) The problem of reliance on difficult annotation in the existing technology is solved. The effect of deformable registration can be improved by only a simple and easy-to-implement arrow prompt, which improves the usability of deformable registration technology in practical application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is an overall flow chart of the CT-CBCT deformation registration system of the prompt feedback convolutional neural network of the present invention;
[0057] Figure 2 It is a diagram of the prompt feedback convolutional neural network architecture of the present invention;
[0058] Figure 3 It is a structural schematic diagram of the CT-CBCT deformable registration system based on the prompt feedback convolutional neural network of the present invention. DETAILED DESCRIPTION
[0059] The present invention is described below with specific examples, but is not intended to be limiting of the present invention.
[0060] Example 1
[0061] The following is combined with Figures 1 to 3 The technical solution of the present invention is further described.
[0062] The CT-CBCT deformable registration method based on a prompt feedback convolutional neural network described in an embodiment of the present invention is as follows: Figure 1 As shown, the following steps are included:
[0063] (1) Acquire data, divide the acquired data into data sets and perform preliminary processing, including the following steps:
[0064] (1.1) Data collection and data set division: In order to improve the robustness of the benchmark model, data were collected from multiple hospitals at the provincial, municipal and county levels, with different treatment systems such as Varian and Elekta. The data were mainly collected from the chest and abdomen, which are parts of the body with large changes in the range of organs. The planned CT images and CBCT images before 1-5 treatments were collected at the patient level. The total number of patients was no less than 200, forming the overall data set. In order to ensure the reliability of model training and the non-correlation of data distribution, the data set was divided into training set, validation set and test set in a ratio of 7:1:2, and it was required that the image of the same patient could only appear in one set.
[0065] (1.2) Preprocess the data set:
[0066] (1.2.1) Due to the differences in CT / CBCT imaging principles, the physical information in their images is often different. In order to align the spacing, the CT / CBCT images are first resampled, and the spacing and size of the CBCT are changed with CT as a reference. In order to reduce the error caused by the positioning difference, 3D rigid registration is required to rigidly align CT-CBCT. After alignment, valid 2D layer image pairs are extracted (because the 3D images of CBCT images are spherical, the incomplete images at both ends are considered invalid and discarded).
[0067] (1.2.2) Each pair of CT and CBCT images is encoded by pixel blocks. Considering the continuity of pixels and the interpolation after deformation registration, the pixel block size is set to 2*2, that is, a pixel block includes four adjacent pixels. Assume that the image size is 2N i *2M i(If the length and width of the image are not an even number of pixels, they are padded with 0s.) The subscript i represents the i-th image pair. The pixel block id containing the four pixels with coordinates (0,0), (0,1), (1,0), and (1,1) is set to 0; the pixel block id containing the four pixels with coordinates (1,0), (1,1), (2,0), and (2,1) is set to 1. Similarly, the i-th global image is divided into N pixels of size i *M i The pixel block matrix, the id of a specific pixel block is (n i ,m i ), where n i Indicates that the pixel block is located in the nth row of the pixel block matrix, m i Indicates that the pixel block is located in the mth column of the pixel block matrix, 0<n i <N i , 0<m i <M i ;
[0068] (1.3) Perform ROI rectangle frame extraction on the preprocessed data set:
[0069] (1.3.1) In order to focus the deformable registration on a specific local area or a key area, use the supporting tools to draw a rectangular frame of the same size at the same position of the preprocessed CT / CBCT image pair, and ensure that the corresponding key area in the CT / CBCT falls completely within the same set of rectangular frames;
[0070] (1.3.2) Fill the local image to 256*256 by evenly filling 0 on all sides;
[0071] (1.3.3) The local images after zero padding are encoded from top to bottom and from left to right. Each local image can obtain a 255*255 pixel block matrix, and the pixel blocks are numbered from (0,0) to (254,254).
[0072] (2) Constructing a prompt-enhanced convolutional neural network, including the following steps;
[0073] (2.1) Use standard UNet to build benchmark image encoder and image decoder, including:
[0074] (2.1.1) The encoder is mainly composed of a convolution module and a downsampling module:
[0075] Each convolution module includes two 3*3 convolution kernels. Each convolution operation is followed by a batch normalization layer. At the same time, in order to prevent overfitting, a DropOut layer with a parameter of 0.3 is followed by a ReLU activation layer after the batch normalization layer. Each convolution module is followed by a maximum pooling layer with a step size of 2 to reduce the size of the feature map while retaining important information. From the image input to the output of the encoding convolution module, a total of 5 convolution modules are required. The number of output channels of each convolution module is 64, 128, 256, 512 and 1024 respectively.
[0076] The decoder consists of 4 deconvolution blocks, each of which contains a 2*2 transposed convolutional layer. The number of output channels is 512, 256, 128, and 64, which is the opposite order of the encoder convolutional module. After the transposed convolutional layer, there is a convolutional module (after two 3*3 convolutional layers, a batch normalization layer, and a ReLU activation layer); followed by a convolutional layer with 64 input channels, 2 output channels, a convolution kernel size of 1*1, a stride of 1, and a padding of 1. The first channel of the output represents the x-direction offset of each pixel in the deformation field, and the second channel represents the y-direction offset of each pixel in the deformation field.
[0077] Each layer in the decoder is connected to the corresponding layer in the encoder. The outputs of the first to fourth layers of the encoder are connected to the output layer of the transposed convolution through skip connections, which helps to combine the semantic information of the bottom and high layers.
[0078] (2.1.2) The image encoder is modified by adding an embedding layer after the image input layer. The size of the embedding output matrix is 256*256. By introducing embedding coding, the robustness of image coding in different hospitals, different treatment systems, and different parts is enhanced.
[0079] (2.2) Add prompt encoder module, including:
[0080] (2.2.1) Add an Embedding layer for the prompt image to the prompt encoder. The output matrix size of the Embedding layer is 256*256. After prompt Embedding, the Layer Norm layer normalizes the output matrix to a normal distribution with a mean of 0 and a mean square error of 1.
[0081] (2.2.2) Add a position encoding module for the arrow in the prompt to the prompt encoder. Since the starting and ending points of the arrow do not necessarily coincide with the pixel points, the pixel block coordinates of the starting and ending points of the arrow are used to represent the coordinates of the arrow starting point, arrow end point, and arrow direction vector. The sine and cosine encoding method is used to encode the pixel block coordinates of the starting point of the arrow in the prompt, the pixel block coordinates of the arrow end point, and the arrow direction vector according to the pixel block coordinates according to the following calculation formula. The spatial coordinates of each dimension are converted into a 256-dimensional vector;
[0082]
[0083] Among them, pos(x, y) represents the position in the pixel block matrix, x is the horizontal coordinate of the pixel block, y is the vertical coordinate of the pixel block, 256 represents the dimension of PE, 2i and 2j represent the dimensions calculated by row and by column for even numbers, and 2i+1 and 2j+1 represent the dimensions calculated by row and by column for odd numbers (2i≤256, 2i+1≤256, 2j≤256, 2j+1≤256). After the calculation of formulas (1) to (4), the spatial coordinates of each point can be converted into several vectors of size 256:
[0084] PEX pos(xy) :[PEy0, PEy1, PEy2,...,PEy 255 ] x
[0085] PEY pos(xy) :[PEx0,PEx1,PEx2,...,PEx 155 ] y
[0086] Then PEX pos(x,y) With PEY pos(x,y Add them bit by bit to the x-th row and y-th column of the 256*256 dimensional matrix output by prompt embedding.
[0087] (2.2.3) Add a self-attention module to the prompt encoder. The module first consists of three independent 1*1 convolutional layers (stride is 1). After the linear mapping of the following formula, the three outputs are Q matrix, K matrix, and V matrix, with a size of 256*256*C, where C is the number of output channels:
[0088] Q(x)=W q *x
[0089] K(x)=W k *x
[0090] V(x)=W v *x
[0091] Then flatten Q(x), K(xx), and V(x) to change their size to 65536*C; transpose the K matrix to get K T Matrix, multiplied by Q matrix, outputs a result matrix of size 65536*65536, and then normalizes the result by softmax to get the attention weight β matrix:
[0092]
[0093] β of the above formula j,i It is used to represent the relationship weight of the i-th position to the j-th position, and N represents the number of feature positions, which is 65536 here.
[0094] Then perform matrix multiplication on the attention weight matrix and the V matrix to obtain an output matrix O of size 65536*C, which is reshaped to 256*256*C:
[0095]
[0096] Finally, by setting a parameter gamma, the matrix O is linearly superimposed on the input X of the self-attention layer to obtain the output matrix Y of the self-attention layer with a final size of 256*256*C:
[0097] Y=gamma*0+X
[0098] (2.3) Add a spatial transformation network layer to the output layer of the decoder, including:
[0099] (2.3.1) A spatial transformation layer is added after the image decoder output layer. This layer receives two inputs: the deformation field matrix output by the image decoder and the local floating image after zero padding. The floating image is deformed according to the following formula to obtain the value of each pixel / voxel position of the floating image:
[0100]
[0101] Where p represents the pixel / voxel point at an integer position, p′=p+u(p) represents the point after point p is deformed according to the deformation field, and Z(p′) represents the integer neighbor voxel of p′, which is 4 domain points in 2D processing and 8 domain points in 3D. d q d | measures the distance between point p′ and point q, while 1|p′ d q d| reflects that the closer the distance between point p′ and a neighbor voxel is, the greater the weight of this neighbor voxel in calculating the interpolation value of point p′. Through such a spatial transformation network, the deformation of the floating image is completed and the It is directable.
[0102] (3) The specific steps of training the no-prompt model include:
[0103] (3.1) Each time, BS is randomly selected from the training set (the size of a training batch, ranging from 2 to 64), and the weighted sum of the average grayscale value MSE of the normalized deformed image and the fixed image and the second-order norm of the deformation field is used as the loss function. The weighted ratio is 1000:10. The loss function of the Mth pair of images is calculated by the following code:
[0104]
[0105] Where N represents the total number of pixels. It represents the deformation field, which can add deformation variables of each dimension to the coordinates of the input point p to generate deformation. D represents the dimension of the deformation field. If it is a 3D image deformation, it is 3, if it is a 2D image deformation, it is 2. The subscript i represents each dimension of the deformation field, the subscript , represents pixel-by-pixel calculation, α and β are 1000 and 10, and the total loss is the average loss of each image pair in a batch.
[0106] Adopt Adam optimizer, set epochs to M1·n train / BS, where N train represents the total number of image pairs in the training set, M1 represents a positive integer ranging from 2 to 10, BS is the size of a training batch, and the learning rate for the first 50% of epochs is λ1 (λ1 is 1*10 -4 or 0.5*10 -4 ), the learning rate of the last 50% of epochs is λ1 / 10, each epoch iterates 150-300 times, that is, steps / epoch is set to 150-300, and the model is trained on a single A40 card;
[0107] (3.2) The model adopts an optimized update archiving method, and the model parameters are updated and saved only when the model after each round of training shows better results in the validation set;
[0108] (3.3) Gradually adjust the weighted ratio, keep the other hyperparameters unchanged, continue to train the model, and save the optimal model.
[0109] (4) The specific steps of arrow prompt model training include:
[0110] (4.1) Obtaining arrow prompts: Each pair of local images is superimposed and displayed with a transparency of 30%-70% to ensure that the ROI edges of the foreground and background images can be clearly distinguished; professional doctors mark the major deformation processes in the superimposed images with arrows. The starting point of the arrow should be at the core position of the ROI key deformation area of the floating image (taking the deformation core), and the end point of the arrow should be at the core position of the ROI key deformation area of the fixed image (taking the deformation core). The length and direction of the arrow represent the deformation degree and direction of this core position. Finally, only the annotations with prompt arrow masks are retained.
[0111] (4.2) The optimal model parameters obtained by training without prompt are used as the initial parameters of the network except the prompt encoder. The Adam optimizer is used, and the learning rate is set to λ2 (λ2 is set to 0.5*10 -4 or 1*10 -5 ), epochs set to M2 N train / BS, where N train represents the total number of image pairs in the training set, M2 represents a positive integer, ranging from 2 to 10, BS is the size of each training batch, ranging from 2 to 64, steps / epoch is set to 150-300, and the input is a local image pair and the arrow prompt mask annotation corresponding to this image pair; each random sampling forms a training batch, and the loss function of the model is still set to a batch-normalized weighted sum of the average MSE of the pixel values of the deformed image and the fixed image and the second-order norm of the average deformation field, with a weighted ratio of 1000:10, followed by model training;
[0112] (4.3) The model adopts an optimized update archiving method, and the model parameters are updated and saved only when the model after each round of training shows better results in the validation set;
[0113] (4.4) Gradually adjust the weighting ratio, keep the other hyperparameters unchanged, continue to train the model, and save the optimal model.
[0114] (5) The specific steps of fine-tuning the arrow prompt model include:
[0115] (5.1) Freeze the image encoder and prompt encoder parameters, and open all the parameters of the image decoder. The loss function still uses the weighted MSE and the second-order norm of the deformation field. The weighted coefficient uses the value finally determined in (4.4). The Adam optimizer is used, and the learning rate is λ3 (λ3 is 1*10 -5 or 0.5*10 -5 ), the training step of each image pair is set to 300-500;
[0116] (5.2) When the system receives feedback from the user and is not satisfied with the deformation registration effect, the system will collect the image pair input by the user and the arrow prompt on it, and perform fine-tuning on the new input according to the hyperparameter settings in (5.1); (5.3) After fine-tuning is completed, the deformation registration results before and after fine-tuning are output. If the user selects the result after fine-tuning, the fine-tuning model will be retained. If the user does not select the result after fine-tuning, it will not be retained;
[0117] like Figure 3 As shown, an embodiment of the present invention also provides a CT-CBCT deformable registration system based on a prompt feedback convolutional neural network, including a data processing module, a registration model training module, an effect feedback module, and a reasoning and self-training module.
[0118] Data processing module: mainly used for synchronous acquisition of local images on CT-CBCT image pairs, and preprocessing operations such as resampling, filling, cropping, rigid alignment, and layer alignment on global image pairs;
[0119] Registration model training module: mainly used for the pre-built neural network, using the pre-processed data obtained by the data processing module as the training data set and the verification data set, using the arrow prompts provided by the user, to complete the model training with arrow prompts, and obtain the optimal prompt feedback convolutional neural network model, and can accept out-of-set image pairs, perform deformation registration through the model, and output the deformed registered image and deformation field;
[0120] Effect feedback module: It is mainly used to visualize the deformed image and deformation field, and provide human evaluation feedback based on the above two results, allowing users to tell the system in an interactive way whether they are satisfied with the deformation registration results of this round.
[0121] Reasoning and self-training module: It is mainly used for users to perform deformable registration on their private image pairs after the training is completed and the benchmark model is correctly loaded. According to the user's arrow prompts and result feedback, the model can be automatically fine-tuned on these new out-of-set image pairs to further improve the registration effect of private image pairs.
Claims
1. A CT-CBCT deformable registration method based on a cue-feedback convolutional neural network, characterized in that: The following steps are involved: (1) Acquire data, divide the acquired data into data sets and perform preliminary processing; (1.1) Data collection and data set division: the data set is divided into training set, validation set and test set in a ratio of 7:1:2, and the CT-CBCT image pair of the same patient can only appear in one of the sets; (1.2) Preprocess the data set; (1.3) Extract the ROI rectangle in the same frame of the preprocessed data set; (2) Constructing a prompt-enhanced convolutional neural network; (2.1) Build the benchmark image codec module; (2.2) Add a prompt encoder and add a prompt encoder module with an attention mechanism to the constructed network body; (2.3) Modify the image encoder and add a spatial transformation layer after the output layer of the image decoder; (3) Silent model training: The extracted local CT-CBCT image pairs are used as input, and the weighted MSE and the second-order norm of the deformation field are used as the loss function to train the network in a silent manner; (4) Arrow prompt model training: The loss function remains unchanged, and the optimal model parameters obtained in (3) are used as the initial parameters. Arrow prompts are added to continue training the model to obtain the final weight coefficient value. The starting point of the arrow is at the core position of the key deformation area of the ROI of the floating image, and the end point of the arrow should be at the core position of the key deformation area of the ROI of the fixed image; (5) Reasoning and self-training: fine-tune the arrow prompt model based on the image pairs input by the user, the arrow prompts, and the feedback evaluation of the deformation registration results.
2. The CT-CBCT deformable registration method based on prompt feedback convolutional neural network according to claim 1, characterized in that: The step (1.2) of preprocessing the data set includes: (1.2.1) Perform spacing alignment, resampling, rigid alignment, cropping, 2D slice alignment, and sampling on CT-CBCT image pairs in three dimensions; (1.2.2) Take 2*2 pixels as a pixel block, and set the ID of all pixel blocks in the image from left to right and from top to bottom in a row-first manner, with the ID ranging from 0, 1...N.
3. The CT-CBCT deformable registration method based on prompt feedback convolutional neural network according to claim 1, characterized in that: The step (1.3) extracts the ROI rectangle frame from the preprocessed data set, including: (1.3.1) Select the ROIs in the CT-CBCT image pair by rectangular boxes synchronously. Make sure that the rectangular boxes are captured at the same position and that the corresponding ROIs are completely contained in the rectangular boxes. (1.3.2) Fill the local image to 256*256 by evenly filling 0 on all sides; (1.3.3) The local image after zero-filling is encoded in a row-first manner, from top to bottom and from left to right. Each local image obtains a 255*255 pixel block encoding matrix, and the pixel blocks are numbered from (0,0) to (254,254).
4. The CT-CBCT deformable registration method based on prompt feedback convolutional neural network according to claim 1, characterized in that: The step (2.1) of building a reference image codec module includes: (2.1.1) Build the network entity based on the UNet architecture; (2.1.2) Add an embedding layer after the image pair input layer, and set the embedding output matrix size to 256*256; (2.1.3) If there is a hint input, a matrix bitwise addition operator is set after the embedding layer of the image encoder to add the output matrix of the embedding layer of the image encoder and the output matrix of the hint encoder bitwise, and the result is input to the convolution module of the image encoder for encoding.
5. The CT-CBCT deformable registration method based on prompt feedback convolutional neural network according to claim 1, characterized in that: The step (2.2) of adding a prompt encoder step comprises: (2.2.1) Add an embedding module to the prompt encoder, and the embedding output matrix size is 256*256; (2.2.2) Add a position encoding module to the prompt encoder, and the position encoding output matrix size is also 256*256; (2.2.3) Add an attention module to the hint encoder. The attention output matrix size is 256*256, which is used to capture the relationship between each pixel block and all other pixel blocks.
6. The CT-CBCT deformable registration method based on prompt feedback convolutional neural network according to claim 1, characterized in that: The step (2.3) of transforming the image encoder includes: (2.3.1) A spatial transformation layer is added after the output layer of the image decoder. This layer receives two inputs: the deformation field output by the image decoder and the floating image. The linear interpolation method is used to calculate the pixel values of the fixed grid points after deformation based on the floating image and the deformation field to complete the output of the registered image. This layer does not involve network parameter updates and does not participate in back-propagation calculations.
7. The CT-CBCT deformable registration method based on prompt feedback convolutional neural network according to claim 1, characterized in that: The step (3) of training the silent model includes: (3.1) Each time, BS local image pairs are randomly selected from the training set to form a training batch. BS is the number of training batches. The value range of BS is 2 to 64. The weighted sum of the average gray value MSE of the normalized deformed image and the fixed image and the second-order norm of the deformation field is used as the loss function. The weighting ratio is 1000:
10. M The loss function for the image pair is calculated using the following code: Where N represents the total number of pixels. Represents the deformation field, for the input point The deformation is generated by adding deformation variables of each dimension to the coordinates of Denotes the dimension of the deformation field, subscript i Denotes each dimension of the deformation field, subscript j Indicates pixel-by-pixel calculation, and are weighted coefficients, set to 1000 and 10 respectively. The total loss is the average loss of each image pair in a batch. It represents the deformed image in which each pixel of the moving image is displaced under the deformation field, m is the moving image, and f is the fixed image; The Adam optimizer is used, and the epochs are set to M1·N train / BS, where N train represents the total number of image pairs in the training set, M1 represents a positive integer, ranging from 2 to 10, BS is the size of a training batch, ranging from 2 to 64, and the learning rate for the first 50% of epochs is λ1, where λ1 is 1*10 -4 or 0.5*10 -4 , the learning rate of the last 50% of epochs is λ1 / 10, and each epoch is iterated 150-300 times, that is, steps / epoch is set to 150-300, and the model is trained on a single A40 card; (3.2) The model is archived in an optimized update mode, and the model parameters are updated and saved only when the model after each round of training shows better results in the validation set; (3.3) Gradually adjust the weighted ratio, keep the other hyperparameters unchanged, continue training the model, and save the optimal model.
8. The CT-CBCT deformable registration method based on prompt feedback convolutional neural network according to claim 1, characterized in that: The step (4) of training the arrow prompt model includes: (4.1) Obtaining arrow tips: Overlay each pair of local images with a transparency of 30%-70% to ensure that the ROI edges of the foreground and background images can be clearly distinguished; professional doctors mark the major deformation processes in the overlaid images with arrows, and only the annotations with arrow masks are retained; (4.2) The optimal model parameters obtained by training without prompts are used as the initial parameters of the network except the prompt encoder. The Adam optimizer is used, and the learning rate is set to λ2, where λ2 is set to 0.5*10 -4 or 1*10 -5 , epochs set to M2 N train / BS, where N train represents the total number of image pairs in the training set, M2 represents a positive integer, ranging from 2 to 10, BS is the size of each training batch, ranging from 2 to 64, steps / epoch is set to 150-300, and the input is a local image pair and the arrow prompt mask annotation corresponding to this image pair; each random sampling forms a training batch, and the loss function of the model is still set to a batch-normalized weighted sum of the average MSE of the pixel values of the deformed image and the fixed image and the second-order norm of the average deformation field, with a weighted ratio of 1000:10, and then the model training is carried out; (4.3) The model is archived in an optimized update mode, and the model parameters are updated and saved only when the model after each round of training shows better results in the validation set; (4.4) Gradually adjust the weighted ratio, keep the other hyperparameters unchanged, continue training the model, and save the optimal model.
9. The CT-CBCT deformable registration method based on prompt feedback convolutional neural network according to claim 1, characterized in that: The step (5) of fine-tuning the arrow prompt model includes: (5.1) Freeze the image encoder and prompt encoder parameters, and open all the parameters of the image decoder. The loss function still uses the weighted MSE and the second-order norm of the deformation field. The weight coefficient uses the value finally determined in (4). The Adam optimizer is used, and the learning rate is λ3, where λ3 is 1*10 -5 or 0.5*10 -5 , the training step for each image pair is set to 300-500; (5.2) Fine-tune the model based on the local image pairs, arrow prompts, and deformation effect feedback actually input by the user; (5.3) After fine-tuning is completed, the deformation registration results before and after fine-tuning are output. If the user selects the result after fine-tuning, the fine-tuning model will be retained. If the result after fine-tuning is not selected, it will not be retained.
10. A CT-CBCT deformable registration system based on a prompt feedback convolutional neural network, characterized in that: include: Data processing module: used to synchronously acquire local images on CT-CBCT image pairs, and perform resampling, filling, cropping, rigid alignment, and layer alignment preprocessing operations on global image pairs; Registration model training module: used for the pre-built neural network, using the pre-processed data obtained by the data processing module as the training data set and the verification data set, using the arrow prompts provided by the user, completing the model training with arrow prompts, obtaining the optimal prompt feedback convolutional neural network model, and being able to accept out-of-set image pairs, perform deformation registration through the model, and output the image and deformation field after deformation registration; the acquisition of the arrow prompts, superimposing and displaying each pair of local image pairs, with a transparency of 30%-70%, to ensure that the ROI edges of the foreground and background images can be clearly distinguished; the starting point of the arrow is at the core position of the ROI key deformation area of the floating image, and the end point of the arrow should be at the core position of the ROI key deformation area of the fixed image; Effect feedback module: used to visualize the image and deformation field after deformation registration, and provide human evaluation feedback based on the above two results, allowing users to inform the system in an interactive way whether they are satisfied with the deformation registration results of this round; Inference and self-training module: After the training is completed and the benchmark model is correctly loaded, the user can perform deformable registration on their private image pairs. According to the user's arrow prompts and result feedback, the model can be automatically fine-tuned on these new out-of-set image pairs to further improve the registration effect of private image pairs.
Citation Information
Patent Citations
CT-CBCT image deformation registration method
CN112862873A
Medical image processing system and method for interventional operation
WO2023216947A1