An Image Non-Rigid Registration Method and System Based on Deep Learning
Through the image non-rigid body registration method based on deep learning, the image feature vector sequence is calculated using independent encoder and decoder, which solves the problem of poor image registration effect in the prior art and achieves more accurate image registration.
Patent Information
- Application Number
- CN202411719711.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-11-28
AI Technical Summary
In image registration, the prior art assumes that all parts in the image meet the same deformation model, and it is impossible to adaptively select a deformation model that meets the actual situation based on the deformation differences of each endangered organ, resulting in poor registration effect.
Using a deep learning-based image non-rigid body registration method, a training set is generated by acquiring the first image of the user and the second image obtained by the first image rigid body registration, and a non-rigid body registration model is established. A independent encoder and decoder are used to calculate feature vector sequences of different scales of the image, and the image matching position information is determined based on these feature vector sequences.
It effectively solves the problem of poor registration effect due to inconsistent coverage of the body of the reference image and the moving image, and achieves more accurate image registration, especially when there is a large movement amplitude in the target area.
Smart Images

Figure CN119228860B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image non-rigid registration, and particularly relates to a method and system for image non-rigid registration based on deep learning. Background Art
[0002] Radiotherapy is one of the important means for treating cancer, and it is necessary to accurately locate the tumor positions and surrounding critical organs of each patient. While irradiating the lesion area with high dose, the damage to the critical organs should be reduced, so as to improve the personalized treatment effect. During different fractionated treatments, due to factors such as setup errors, organ movement and deformation, etc., it will lead to omission of the target area or higher dose irradiation of the nearby critical organs, thus affecting the prognosis of radiotherapy patients. Therefore, it is necessary to transform the doses between different fractions to the same reference space through image deformation registration for dose superposition, so as to monitor and evaluate the total actual irradiated dose of the patient during the whole radiotherapy period. Therefore, image deformation registration has important research significance in the field of radiotherapy.
[0003] Image registration is to convert medical images collected at different times and by different devices into a unified spatial coordinate system, so that the image information at the same spatial position corresponds to the same anatomical structure. Due to the differences in different modality imaging methods and principles, there are inherent appearance differences in different modality images. For example, the intensity of CBCT images is inconsistent with the intensity (CT value) of traditional CT images, and the structure representation of MR images is inconsistent with that of CT images, etc. Therefore, finding an effective and accurate image similarity metric is a challenging problem.
[0004] Traditional image registration assumes that all parts within an image conform to the same deformation model, and it is unable to adaptively select a deformation model that conforms to the actual situation based on the deformation differences of each organ at risk. For example, the free-form deformation model assumes that the deformation fields at various positions in space can be controlled by the deformation fields of certain control points, and the deformation fields at all positions are obtained by interpolation using the deformation fields at sparse control points. Elastic deformation regards all organs as elastic entities and satisfies the Navier-Cauchy partial differential equation. The viscous fluid flow model assumes that the deformation field conforms to the characteristics of viscous fluid, and the deformation field is constrained by the Navier-Stokes equation. The optical flow model assumes that the deformation field of each pixel point is the temporal accumulation of its motion speed and conforms to the requirements of the Lagrange transport equation. Currently, the commonly used model is the free-form deformation model, such as the deformation image registration models in commercial software RayStation and MIM Maestro software. This model has the advantage of short processing time, but the disadvantages are that it cannot well preserve the topological structure and it is difficult to add regularization constraints to the deformation field. For example, Marchant's research found that CT images registered based on free-form deformation NiftyReg cannot accurately deform and register in cases where there are large differences in bladder filling in cervical cancer patients, movement of small intestine loops, or the presence of air sacs in the rectum. When there is a large range of motion in the target area, most deformation registration algorithms based on intensity methods also cannot well reproduce CT values in the image. Summary of the Invention
[0005] Based on this, an image non-rigid registration method and system based on deep learning are provided in the embodiments of the present invention, aiming to solve the problem of poor registration effect caused by inconsistent body coverage ranges of the reference image and the moving image.
[0006] In the first aspect of the embodiments of the present invention, an image non-rigid registration method based on deep learning is provided. The method includes:
[0007] Obtain a first image of the user and a second image obtained by rigid registration of the first image to generate a training set;
[0008] Establish a non-rigid registration model, and input the first image and the second image into the non-rigid registration model for training to obtain a target model;
[0009] Input a to-be-registered image of the same type as the second image into the target model and output a registration result;
[0010] The non-rigid registration model includes a decoder, and a first encoder and a second encoder that are independent of each other;
[0011] Both the first encoder and the second encoder are composed of a number of sliding window transformation modules, which are used to calculate the feature vector sequences of different scales of the first image and the second image respectively;
[0012] The decoder is used to determine the position information of the voxels on the second image that match each position on the first image according to the feature vector sequences. Among them, the global attention processing is performed on the feature vector sequence with the smallest size, and the window-internal attention processing is performed on the feature vector sequences of other sizes.
[0013] Further, after the first image and the second image are input into the non-rigid registration model, they first undergo linear embedding processing and rearrangement processing to convert the images into feature vector sequences.
[0014] Further, by combining the sliding window transformation modules, a sliding window transformation module group is obtained. The sliding window transformation module groups in the first encoder and the second encoder match. Among them, one end of the output of the sliding window transformation module group is connected to the next-level sliding window transformation module group through a sub-block merging layer, and the other end is connected to the corresponding interactive attention module. When the sliding window transformation module group is the last-level sliding window transformation module group, the output is only connected to the first cross-attention module, and the first cross-attention module is used to perform global attention processing on the feature vector sequence with the smallest size.
[0015] Further, the second cross-attention module is used to perform window-internal attention processing on the feature vector sequences of other sizes, and the first cross-attention module and the second cross-attention module are connected through an upsampling module.
[0016] Further, the inputs of the first cross-attention module and the second cross-attention module are the feature vector sequences output by the corresponding-level sliding window transformation module group of the first image, the feature vector sequences output by the corresponding-level sliding window transformation module group of the second image, and the position labels corresponding to the feature vector sequences output by the corresponding-level sliding window transformation module group of the second image.
[0017] Further, in the step of establishing the non-rigid registration model and inputting the first image and the second image into the non-rigid registration model for training to obtain the target model, the non-rigid registration model parameters are randomly initialized, and the parameters are optimized under the guidance of the loss function. Among them, the loss function is composed of an image similarity measure, an annotation constraint, and a matching position continuity constraint, and the expression is:
[0018]
[0019] Where Loss represents the total loss, and α, β, and γ are non - negative real numbers used to control the contribution degree of different losses to the overall loss. represents the loss of image similarity measure. represents the annotation constraint loss of the same organs in the first image and the second image. represents the matching position continuity constraint loss.
[0020] The second aspect of the embodiments of the present invention provides an image non - rigid registration system based on deep learning for implementing the image non - rigid registration method based on deep learning provided in the first aspect. The system includes:
[0021] An acquisition module, configured to acquire the first image of the user and the second image obtained by rigid registration of the first image, and generate a training set.
[0022] A training module, configured to establish a non - rigid registration model, and input the first image and the second image into the non - rigid registration model for training to obtain a target model.
[0023] An input module, configured to input a to - be - registered image of the same type as the second image into the target model and output a registration result.
[0024] The non - rigid registration model includes a decoder, and independent first and second encoders.
[0025] Both the first encoder and the second encoder are composed of several sliding window transformation modules, and are used to calculate the feature vector sequences of different scales of the first image and the second image respectively.
[0026] The decoder is used to determine the position information of the voxels on the second image that match each position on the first image according to the feature vector sequences. Among them, the feature vector sequence with the smallest size is subjected to global attention processing, and the feature vector sequences of other sizes are subjected to in - window attention processing.
[0027] The third aspect of the embodiments of the present invention provides a computer - readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the image non - rigid registration method based on deep learning provided in the first aspect.
[0028] The fourth aspect of the embodiments of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the image non - rigid registration method based on deep learning provided in the first aspect.
[0029] A method and system for non-rigid registration of images based on deep learning provided in an embodiment of the present invention. The method generates a training set by obtaining a first image of a user and a second image obtained by rigid registration of the first image; establishes a non-rigid registration model, and inputs the first image and the second image into the non-rigid registration model for training to obtain a target model; inputs a to-be-registered image of the same type as the second image into the target model and outputs a registration result. Specifically, the non-rigid registration model includes a decoder, and independent first and second encoders; both the first encoder and the second encoder are composed of several sliding window transformation modules, which are used to calculate feature vector sequences of different scales of the first image and the second image respectively; the decoder is used to determine the position information of the voxels on the second image that match each position on the first image according to the feature vector sequences. Among them, the feature vector sequence with the smallest size is subjected to global attention processing, and the feature vector sequences of other sizes are subjected to in-window attention processing. It should be noted that due to the use of an independent encoder design, this registration model can be used not only for images of the same modality collected from the same patient at different times, but also for registering different modality images collected from the same patient. In addition, considering the computational efficiency and video memory occupancy in the decoding network, the global attention processing and in-window attention processing methods are used to process the feature vector sequences of different sizes respectively, and finally the problem of poor registration effect caused by inconsistent body coverage ranges of the first image and the second image is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 FIG. 6 is a flowchart of an implementation of a method for non-rigid registration of images based on deep learning provided in Embodiment 1 of the present invention;
[0031] Figure 2 FIG. 10 is a schematic diagram of the network structure of the non-rigid registration model;
[0032] Figure 3 FIG. 14 is a comparison result diagram of the annotations on CBCT and the annotations on the registered CT;
[0033] Figure 4 FIG. 18 is a structural block diagram of a system for non-rigid registration of images based on deep learning provided in Embodiment 3 of the present invention;
[0034] Figure 5 FIG. 22 is a structural block diagram of an electronic device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0036] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are only for the purpose of illustration.
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0038] Embodiment 1
[0039] Please refer to Figure 1 , Figure 1 which shows the implementation flowchart of a deep learning-based image non-rigid registration method provided in Embodiment 1 of the present invention. The deep learning-based image non-rigid registration method specifically includes steps S01 to S03.
[0040] Step S01: Obtain the user's first image and the second image obtained by rigidly registering the first image, and generate a training set.
[0041] Specifically, the first image and the second image are the same-modal images or different-modal images of the same patient collected at different times, and the two images have undergone rigid registration processing and have the same size and resolution, and there is no missing image information in the body parts covered by the first image and the second image, that is, the points in the body area on the first image of interest can all find matching positions in the second image. In this embodiment, the first image refers to the reference image, and the second image refers to the moving image, that is, the reference image and the moving image with the same size and resolution are used as the training set of the subsequent non-rigid registration model.
[0042] Step S02: Establish a non-rigid registration model, and input the first image and the second image into the non-rigid registration model for training to obtain a target model.
[0043] Please refer to Figure 2, which is a schematic diagram of the network structure of the non-rigid registration model. Among them, the deformation field vector is the matching position information of each pixel point of the first image on the second image, which is a three-dimensional vector. Since the first image is a 3D image, the size of the deformation field is a matrix of D×H×W×3. It should be noted that the non-rigid registration model includes a first encoder, a second encoder, and a decoder; both the first encoder and the second encoder are composed of several sliding window transformation modules, which are used to calculate the feature vector sequences of different scales of the first image and the second image respectively. Specifically, by combining the sliding window transformation modules, a sliding window transformation module group is obtained. The sliding window transformation module groups in the first encoder and the second encoder match. Among them, one end of the output of the sliding window transformation module group is connected to the next-level sliding window transformation module group through a sub-block merging layer, and the other end is connected to the corresponding interactive attention module. When the sliding window transformation module group is the last-level sliding window transformation module group, the output is only connected to the first cross-attention module, and the first cross-attention module is used to perform global attention processing on the feature vector sequence with the smallest size. In addition, the second cross-attention module is used to perform in-window attention processing on the feature vector sequences of other sizes. The first cross-attention module and the second cross-attention module are connected through an upsampling module. It should be noted that both the first cross-attention module and the second cross-attention module are cross-transformation modules without multi-layer perceptrons and residual connections. Furthermore, the inputs of the first cross-attention module and the second cross-attention module are the feature vector sequences output by the corresponding-level sliding window transformation module group of the first image, the feature vector sequences output by the corresponding-level sliding window transformation module group of the second image, and the position labels corresponding to the feature vector sequences output by the corresponding-level sliding window transformation module group of the second image.
[0044] It can be understood that the image encoders of the reference image and the moving image are independent of each other, so image registration between different modalities is supported. For example, the reference image is a CBCT image, the moving image is a CT image, or the reference image is an MR image, and the moving image is a CT image.
[0045] The decoder is used to determine the position information of the voxels on the second image that match each position on the first image according to the sequence of feature vectors. Among them, the sequence of feature vectors with the smallest size is subjected to global attention processing, and the sequences of feature vectors with other sizes are subjected to in-window attention processing. It should be noted that if global attention processing is performed on the sequence of feature vectors at each level of the reference image and the moving image, it will consume a large amount of computing time and occupy more storage space. Therefore, only the sequence of feature vectors at the highest level (the smallest size) is subjected to global attention processing, that is, searching for matching positions within the global scope. In other level processing, in-window attention is used to control the computational amount and occupied space of the model. After obtaining the high-level matching position information, it is used as the center of the search window at a larger size as the initial, and in-window attention processing is performed, so as to gradually restore the matching position information at each scale, and finally complete the estimation of the matching position.
[0046] Specifically, after the first image and the second image are input into the non-rigid registration model, they first undergo linear embedding processing and rearrangement processing to convert the images into a sequence of feature vectors. The purpose of the linear embedding processing and the rearrangement processing is to unify the image features of size N channels D×H×W into a sequence composed of D / P×H / P×W / P vectors of length C after dimensionality reduction. Specifically, the purpose of the linear embedding processing is dimensionality reduction, and the purpose of the rearrangement processing is to represent the four-dimensional tensor of channel-depth-row-column as depth×row×column-channel, such that a sequence of vectors composed of depth×row×column vectors of length channel is formed.
[0047] In this embodiment, in the global attention processing of the decoding process, the reference image encoder provides q(1, Dr / 32, Hr / 32, Wr / 32, 128), the moving image encoder provides k(1, Dr / 32, Hr / 32, Wr / 32, 128), and the position label corresponding to the moving image feature vector sequence is used as v(1, Dr / 32, Hr / 32, Wr / 32, 3). The position P4’(1, Dr / 32, Hr / 32, Wr / 32, 3) of the moving image feature vector sequence that matches each reference image feature vector sequence is calculated. Here, Dr represents the number of voxels in the depth direction, Hr represents the number of voxels in the height direction, and Wr represents the number of voxels in the width direction. After the position vector sequence P4’ is rearranged into the feature of (1, 3, Dr / 32, Hr / 32, Wr×32), it is upsampled to the position information P3 of (1, 3, Dr / 16, Hr / 16, Wr / 16). The position information P3 is then vectorized into (1, Dr / 16, Hr / 16, Wr / 16, 3) to provide the window center position information for the window attention processing within the upper-level encoded feature. Then, each reference image feature vector sequence is used as q(1, Dr / 6, Hr / 16, Wr / 16, 1, 64). With the corresponding P3 position information as the center, the moving image feature vectors within the 5×5×5 window of the moving image features form v(1, Dr / 6, Hr / 16, Wr / 16, 125, 64), and the position labels within the 5×5×5 window corresponding to the moving image feature vector sequence are also used as v(1, Dr / 6, Hr / 16, Wr / 16, 125, 3). The window attention processing obtains P3’(1, Dr / 16, Hr / 16, Wr / 16, 3). And so on, continuously obtaining more refined matching position information until the position information of (1, 3, Dr / 4, Hr / 4, Wr×4) is obtained. Finally, it is directly upsampled by 4 times to obtain the moving image matching position information (1, 3, Dr, Hr, Wr). According to the matching position information provided by the model, resampling the moving image can obtain the result after non-rigid registration.
[0048] To further introduce the non-rigid registration model, the processing process after the image is sent into the network is introduced according to the processing steps. It should be noted that although Figure 2 the linear encoding window size of 4×4×4 is given, and the two independent encoders have four levels of feature outputs, which are obtained through 2, 4, 4, and 4 sliding window transformation modules respectively. In practical applications, the window size of linear encoding, the number of encoder levels, the number of stacked sliding window transformation modules within each level, and the window size of the window mutual attention processing in the decoder can be designed according to the GPU video memory size, etc. The processing process specifically includes:
[0049] 1. The input reference image has dimensions (Dr, Hr, Wr). It is evenly divided into 4×4×4 sub - blocks. The image data within each sub - block is arranged in order as a vector of length 64. After being processed by the linear transformation layer, a vector output of length 16 is obtained. Thus, the input reference image can be represented as a tensor R1 of (1, Dr / 4, Hr / 4, Wr / 4, 16), which is then superimposed with the learnable position - encoding tensor P0 of the same size.
[0050] 2. R1 is processed by 2 consecutive sliding window transformation modules with a window size of 5×5×5 to obtain the first - level reference feature vector output R1'.
[0051] 3. R1' is reduced to a vector R2 of (1, Dr / 8, Hr / 8, Wr / 8, 32) through sub - block merging. This module first rearranges R1' into a tensor of (1, 16, Dr / 4, Hr / 4, Wr / 4), divides it into sub - blocks in a 2×2×2 manner, arranges 8 vectors of length 16 within the sub - block into a vector of length 128. This vector is processed by the linear transformation layer to obtain a vector of length 32, and a total of (Dr / 8, Hr / 8, Wr / 8) vectors of length 32 are obtained, thus forming a tensor R2 of size (1, Dr / 8, Hr / 8, Wr / 8, 32).
[0052] 4. R2 is processed by 4 consecutive sliding window transformation modules with a window size of 5×5×5 to obtain the second - level reference feature vector output R2'.
[0053] 5. R2' is reduced to a vector R3 of (1, Dr / 16, Hr / 16, Wr / 16, 64) through sub - block merging. This module first rearranges R2' into a tensor of (1, 32, Dr / 8, Hr / 8, Wr / 8), divides it into sub - blocks in a 2×2×2 manner, arranges 8 vectors of length 32 within the sub - block into a vector of length 256. This vector is processed by the linear transformation layer to obtain a vector of length 64, and a total of (Dr / 16, Hr / 16, Wr / 16) vectors of length 64 are obtained, namely vector R3.
[0054] 6. R3 is processed by 4 consecutive sliding window transformation modules with a window size of 5×5×5 to obtain the third - level reference feature vector output R3'.
[0055] 7. R3' is reduced to a vector R4 of (1, Dr / 32, Hr / 32, Wr / 32, 128) through sub - block merging.
[0056] 8. R4 is processed by 4 consecutive sliding window transformation modules with a window size of 5×5×5 to obtain the fourth - level reference feature vector output R4'.
[0057] 9. Similar to the design of the reference image encoding network, the moving image is processed by the moving image encoding network to obtain the first-level moving feature vector output M1', the second-level moving feature vector output M2', the third-level moving feature vector output M3', and the fourth-level moving feature vector output M4'.
[0058] 10. Then, global interactive attention processing is performed to gradually restore the matching position information. The fourth-level reference feature vector output R4' is used as the q input of the global interactive attention, the fourth-level moving feature vector output M4' is used as the k input of the global interactive attention, and the natural position label P4 of the fourth-level moving feature vector output is used as the v input of the global interactive attention to obtain the position P4' in the moving feature vector output M4' that matches the reference feature vector output R4'. The size of the natural position label sequence P4 is (1, Dr / 32, Hr / 32, Wr / 32, 3), and the label information of (0, 0, 0, 0, :) is (0, 0, 0), the label information of (0, 0, 0, 1, :) is (0, 0, 1), the label information of (0, 0, 1, 0, :) is (0, 1, 0), the label information of (0, 1, 0, 0, :) is (1, 0, 0), and so on. P4' is the weighted sum of the position information in P4, with a size of (1, Dr / 32, Hr / 32, Wr / 32, 3); the weight is obtained by calculating R4' and M4' through the attention mechanism.
[0059] 11. Then, P4' is rearranged into a vector of (1, 3, Dr / 32, Hr / 32, Wr / 32) and then upsampled by 2×2×2 to obtain a tensor of (1, 3, Dr / 16, Hr / 16, Wr / 16). After this tensor is multiplied by 2, it is rearranged into a position tensor P3' of (1, Dr / 16, Hr / 16, Wr / 16, 3).
[0060] 12. According to the position in the third-level moving feature vector output M3' that matches the third-level reference feature vector output R3' at the (0, 0, 0) position provided in P3', the rounded processing result is used as the center of the window, and the features within the 5×5×5 window of M3' are taken and arranged in order into a vector of (125, 64). Each position refers to the above operation; a new tensor MW3 can be obtained, with a size of (1, Dr / 16, Hr / 16, Wr / 16, 125, 64); following the process of rearranging the vectors within the window on M3, the label information within the window can be taken from the third-level natural position label sequence P3 and rearranged to obtain a new tensor PW3.
[0061] 13. Rearrange the third - level reference feature vector output R3’ into (1, Dr / 16, Hr / 16, Wr / 16, 1, 64) as the q input of the cross - attention within the third - level window, use MW3 as the k input of the cross - attention within the third - level window, and use PW3 as the v input of the cross - attention within the third - level window to obtain the third - level matching position output P3’. The size of P3’ is (1, Dr / 16, Hr / 16, Wr / 16, 3).
[0062] 14. Then, rearrange P3’ into a vector of (1, 3, Dr / 16, Hr / 16, Wr / 16), and after 2×2×2 upsampling, obtain a tensor of (1, 3, Dr / 8, Hr / 8, Wr / 8). After multiplying this tensor by 2, rearrange it into a position tensor P2’ of (1, Dr / 8, Hr / 8, Wr / 8, 3).
[0063] 15. According to the position in the second - level moving feature vector output M2’ that matches the second - level reference feature vector output R2’ at the (0, 0, 0) position provided in P2’, take the rounded - off result as the center of the window, take the features within the 5×5×5 window of M2’, and arrange them in order into a (125, 32) vector. Refer to the above operations for each position; a new tensor MW2 can be obtained, and its size is (1, Dr / 8, Hr / 8, Wr / 8, 125, 32); following the process of rearranging the vectors within the window on M2, the window - based label information can be rearranged in the second - level natural position label sequence P2 to obtain a new tensor PW2.
[0064] 16. Rearrange the second - level reference feature vector output R2’ into (1, Dr / 8, Hr / 8, Wr / 8, 1, 32) as the q input of the cross - attention within the second - level window, use MW2 as the k input of the cross - attention within the second - level window, and use PW2 as the v input of the cross - attention within the second - level window to obtain the second - level matching position output P2’. The size of P2’ is (1, Dr / 8, Hr / 8, Wr / 8, 3).
[0065] 17. Then, rearrange P2’ into a vector of (1, 3, Dr / 8, Hr / 8, Wr / 8), and after 2×2×2 upsampling, obtain a tensor of (1, 3, Dr / 4, Hr / 4, Wr / 4). After multiplying this tensor by 2, rearrange it into a position tensor P1’ of (1, Dr / 4, Hr / 4, Wr / 4, 3).
[0066] 18. According to the position in the first-level moving feature vector output M1' that matches the first-level reference feature vector output R1' at the (0, 0, 0) position provided in P1', take its rounded processing result as the center of the window, take the features within the 5×5×5 window of M1', and arrange them in order as a (125, 16) vector. Refer to the above operations for each position; a new tensor MW1 can be obtained, with a size of (1, Dr / 4, Hr / 4, Wr / 4, 125, 16). Imitating the process of rearranging the vectors within the window on M1, the label information within the window can be taken from the first-level natural position label sequence P1 and rearranged to obtain a new tensor PW1;
[0067] 19. Rearrange the first-level reference feature vector output R1' as (1, Dr / 4, Hr / 4, Wr / 4, 1, 16) as the q input for the first-level window-internal interactive attention, take MW1 as the k input for the first-level window-internal interactive attention, and take PW1 as the v input for the first-level window-internal interactive attention to obtain the first-level matching position output P1', with a size of (1, Dr / 4, Hr / 4, Wr / 4, 3);
[0068] 20. Then rearrange P1' into a vector of (1, 3, Dr / 4, Hr / 4, Wr / 4), and after upsampling with 4×4×4, obtain a tensor of (1, 3, Dr, Hr, Wr). After multiplying this tensor by 4, rearrange it into a matching position P result of (1, Dr, Hr, Wr, 3).
[0069] It should be noted that the position label sequences P1, P2, P3, and P4 do not need to be learned and can be directly calculated according to the size of the image features. The parameters in other processing modules all need to be learned, and the optimal parameters that can achieve the non-rigid registration effect are obtained through optimization by the optimizer. After the model training is completed, the model parameters are fixed, and the registration position result can be obtained through the above processing.
[0070] In this embodiment, the non-rigid registration model parameters are randomly initialized, and parameter optimization is carried out under the guidance of the loss function. Among them, the loss function consists of an image similarity measure, a labeling constraint, and a matching position continuity constraint, and the expression is:
[0071]
[0072] Among them, Loss represents the total loss, and α, β, and γ are non-negative real numbers used to control the contribution degree of different losses to the overall loss. represents the loss of the image similarity measure. represents the labeling constraint loss of the same organ in the first image and the second image. represents the matching position continuity constraint loss.
[0073] Specifically, the mutual information I(A, B) measures the correlation degree between two random variables and can also be used to calculate image similarity. Here, A is the reference image, and B is the moving image obtained by deforming using the deformation field. The formula for mutual information is as follows:
[0074]
[0075]
[0076] Among them, represents the entropy of the reference image, represents the entropy of the moving image obtained by deforming using the deformation field. They are respectively calculated from the marginal probability density , of the image. is the cross-entropy, which is calculated from the joint probability density of the two. The marginal probability density is calculated using the statistical histogram, is calculated using the joint statistical histogram.
[0077] The DSC (Dice similarity coefficient), also known as the Dice coefficient, is used to measure the overlap degree between two binary images A and B. Among them, A represents the binary image of the annotation of a certain critical organ on the reference image, where the critical organ part is labeled as 1 and other parts are labeled as 0; B represents the binary image of the same critical organ marked on the moving image after deformation using the deformation field, where the critical organ is labeled as 1 and other parts are labeled as 0. The formula for DSC is as follows:
[0078]
[0079] The higher the proportion of the overlap between the annotation on the moving image after resampling according to the matching position information obtained by registration and the annotation on the reference image, the larger the DSC and the better the registration effect. If it is used as a loss, it needs to be written as
[0080]
[0081] to constrain the rationality of the matching position estimation, and it is required that there should be no folding in the deformed moving image.
[0082] Calculate the Jacobian matrix of the matching position at each pixel position on the reference image using the registration network:
[0083]
[0084] Represented as the x - direction component of the three - dimensional deformation field, Represented as the y - direction component of the three - dimensional deformation field, Represented as the z - direction component of the three - dimensional deformation field. If no folding phenomenon occurs in the deformation field, it is necessary to make the value of the Jacobian determinant greater than 0, that is
[0085]
[0086] To ensure that no folding phenomenon occurs in the deformation field, it is necessary to examine the value of the Jacobian determinant at each pixel position. If this value is greater than zero, the determinant loss is at least zero; if the value of the determinant is less than or equal to zero, the loss is the largest, with a value of 1. The overall loss of the deformation field is the average of the determinant losses at all points.
[0087]
[0088] Step S03: Input the registration - to - be image of the same type as the second image into the target model and output the registration result.
[0089] In summary, for the image non - rigid registration method based on deep learning in the above embodiments of the present invention, this method generates a training set by obtaining the user's first image and the second image obtained by rigid registration of the first image; establishes a non - rigid registration model, inputs the first image and the second image into the non - rigid registration model for training to obtain a target model; inputs the registration - to - be image of the same type as the second image into the target model and outputs the registration result. Specifically, the non - rigid registration model includes a first encoder, a second encoder, and a decoder; both the first encoder and the second encoder are composed of several sliding window transformation modules, which are used to calculate the feature vector sequences of different scales of the first image and the second image respectively; the decoder is used to determine the position information of the voxels on the second image that match each position on the first image according to the feature vector sequences. Among them, global attention processing is performed on the feature vector sequence with the smallest size, and window - within attention processing is performed on the feature vector sequences of other sizes, which can effectively solve the problem of poor registration effect caused by the inconsistent body coverage ranges of the first image and the second image.
[0090] Embodiment 2
[0091] In the second embodiment of the present invention, in order to verify the feasibility of the image non-rigid registration method based on deep learning in the first embodiment of the present invention, the CBCT images and CT images of 186 patients are used as the training set. The CBCT images are used as the reference images. After rigidly registering the CT images, the CBCT images and CT images are uniformly resampled into images with a size of 160×320×320 to form training samples. Since all samples have been unified in size, the batchsize can be set to be greater than 1 during the training process and can be determined according to the video memory size of the graphics card. Then, the following steps are performed:
[0092] 1. The input reference image has a size of (160, 320, 320). It is evenly divided into 4×4×4 sub-blocks. The image data within each sub-block is arranged in order as a vector with a length of 64. After being processed by the linear transformation layer, a vector output with a length of 16 is obtained. In this way, the input reference image can be represented as a tensor R1 of (1, 40, 80, 80, 16), and then it is superimposed with the learnable position encoding P0 tensor of the same size;
[0093] 2. R1 is processed by 2 consecutive sliding window transformation modules with a window size of 5×5×5 to obtain the first-level reference feature vector output R1';
[0094] 3. R1' is reduced to a vector R2 of (1, 20, 40, 40, 32) through sub-block merging processing. This module first rearranges R1' into a tensor of (1, 16, 40, 80, 80), divides it into sub-blocks in the 2×2×2 manner, arranges 8 vectors with a length of 16 within the sub-block into a vector with a length of 128. This vector is processed by the linear transformation layer to obtain a vector with a length of 32. A total of (20, 40, 40) vectors with a length of 32 are obtained, which constitutes a tensor R2 with a size of (1, 20, 40, 40, 32);
[0095] 4. R2 is processed by 4 consecutive sliding window transformation modules with a window size of 5×5×5 to obtain the second-level reference feature vector output R2';
[0096] 5. R2' is reduced to a vector R3 of (1, 10, 20, 20, 64) through sub-block merging processing. This module first rearranges R2' into a tensor of (1, 32, 20, 40, 40), divides it into sub-blocks in the 2×2×2 manner, arranges 8 vectors with a length of 32 within the sub-block into a vector with a length of 256. This vector is processed by the linear transformation layer to obtain a vector with a length of 64. A total of (10, 20, 20) vectors with a length of 64 are obtained, namely the vector R3;
[0097] 6. R3 is processed by 4 consecutive sliding window transformation modules with a window size of 5×5×5 to obtain the third-level reference feature vector output R3';
[0098] 7. R3’ is reduced to a vector R4 of (1, 5, 10, 10, 128) through sub-block merging processing;
[0099] 8. R4 is processed through 4 consecutive sliding window transformation modules with a window size of 5×5×5 to obtain a fourth-level reference feature vector output R4’;
[0100] 9. Similar to the design of the reference image encoding network, the moving image is processed through the moving image encoding network to obtain a first-level moving feature vector output M1’, a second-level moving feature vector output M2’, a third-level moving feature vector output M3’, and a fourth-level moving feature vector output M4’;
[0101] 10. Then, global interactive attention processing is performed to gradually restore the matching position information. The fourth-level reference feature vector output R4’ is used as the q input of the global interactive attention, the fourth-level moving feature vector output M4’ is used as the k input of the global interactive attention, and the natural position label P4 of the fourth-level moving feature vector output is used as the v input of the global interactive attention to obtain the position P4’ in the moving feature vector output M4’ that matches the reference feature vector output R4’; The size of the natural position label sequence P4 is (1, Dr / 32, Hr / 32, Wr / 32, 3), and the label information of its (0, 0, 0, 0, :) is (0, 0, 0), the label information of (0, 0, 0, 1, :) is (0, 0, 1), the label information of (0, 0, 1, 0, :) is (0, 1, 0), the label information of (0, 1, 0, 0, :) is (1, 0, 0), and so on. P4’ is the weighted sum of the position information in P4 and has a size of (1, Dr / 32, Hr / 32, Wr / 32, 3); The weights are obtained by calculating R4’ and M4’ through the attention mechanism.
[0102] 11. Then, P4’ is rearranged into a vector of (1, 3, 5, 10, 10) and then upsampled by 2×2×2 to obtain a tensor of (1, 3, 10, 20, 20). After this tensor is multiplied by 2, it is rearranged into a position tensor P3’ of (1, 10, 20, 20, 3);
[0103] 12. According to the position in the third-level moving feature vector output M3' that matches the third-level reference feature vector output R3' at the (0, 0, 0) position provided in P3', take the rounded result as the center of the window, take the features within the 5×5×5 window of M3', arrange them in order as a (125, 64) vector, and refer to the above operations for each position; a new tensor MW3 can be obtained, with a size of (1, 10, 20, 20, 125, 64); following the process of rearranging the vectors within the window on M3, the label information within the window can be taken from the third-level natural position label sequence P3 and rearranged to obtain a new tensor PW3;
[0104] 13. Rearrange the third-level reference feature vector output R3' into (1, 10, 20, 20, 1, 64) as the q input of the third-level window-internal interactive attention module, use MW3 as the k input of the third-level window-internal interactive attention module, and use PW3 as the v input of the third-level window-internal interactive attention module to obtain the third-level matching position output P3', with a size of (1, 10, 20, 20, 3);
[0105] 14. Then rearrange P3' into a vector of (1, 3, 10, 20, 20), and after upsampling by 2×2×2, obtain a tensor of (1, 3, 20, 40, 40). After multiplying this tensor by 2, rearrange it into a position tensor P2' of (1, 20, 40, 40, 3);
[0106] 15. According to the position in the second-level moving feature vector output M2' that matches the second-level reference feature vector output R2' at the (0, 0, 0) position provided in P2', take the rounded result as the center of the window, take the features within the 5×5×5 window of M2', arrange them in order as a (125, 32) vector, and refer to the above operations for each position; a new tensor MW2 can be obtained, with a size of (1, 20, 40, 40, 125, 32); following the process of rearranging the vectors within the window on M2, the label information within the window can be taken from the second-level natural position label sequence P2 and rearranged to obtain a new tensor PW2;
[0107] 16. Rearrange the second-level reference feature vector output R2' into (1, 20, 40, 40, 1, 32) as the q input of the second-level window-internal interactive attention, use MW2 as the k input of the second-level window-internal interactive attention, and use PW2 as the v input of the second-level window-internal interactive attention to obtain the second-level matching position output P2', with a size of (1, 20, 40, 40, 3);
[0108] 17. Next, P2’ is rearranged into a vector of (1, 3, 20, 40, 40) and then undergoes 2×2×2 upsampling to obtain a tensor of (1, 3, 40, 80, 80). After multiplying this tensor by 2, it is rearranged into a position tensor P1’ of (1, 40, 80, 80, 3).
[0109] 18. According to the position in the first-level moving feature vector output M1’ that matches the first-level reference feature vector output R1’ at the (0, 0, 0) position provided in P1’, its rounded processing result is used as the center of the window. The features within the 5×5×5 window of M1’ are taken and arranged in order as a (125, 16) vector; each position refers to the above operation; a new tensor MW1 can be obtained, with its size being (1, 40, 80, 80, 125, 16); following the process of rearranging the vectors within the window on M1, the label information within the window can be taken and rearranged in the first-level natural position label sequence P1 to obtain a new tensor PW1.
[0110] 19. The first-level reference feature vector output R1’ is rearranged into (1, 40, 80, 80, 1, 16) as the q input for the first-level window-internal interactive attention, MW1 is used as the k input for the first-level window-internal interactive attention, and PW1 is used as the v input for the first-level window-internal interactive attention to obtain the first-level matching position output P1’, with the size of P1’ being (1, 40, 80, 80, 3).
[0111] 20. Next, P1’ is rearranged into a vector of (1, 3, 40, 80, 80) and then undergoes 4×4×4 upsampling to obtain a tensor of (1, 3, 160, 320, 320). After multiplying this tensor by 4, it is rearranged into a matching position P result of (1, 160, 320, 320, 3).
[0112] 21. During the training process, Batchsize = 2, the initial learning rate is set to 1e-4, and StepLR is used to adjust the learning rate every 5 epochs, with the adjustment ratio being 0.9. Here, StepLR is a learning rate adjustment class in PyTorch, which can adjust the learning rate according to the given step after each epoch. The loss function uses a combination of mutual information similarity loss and continuity loss for training. The weight of the similarity loss is α = 10, and the weight of the continuity loss is γ = 2 for training, and the number of epochs is set to 50.
[0113] 22. After the training is completed, the model parameters are determined, the test data and its annotation information are processed, various metrics of the reference CBCT image and the registered CT image are analyzed, as well as the coincidence degree of the annotation information on the reference CBCT and the annotation information of the registered CT, and the registration result evaluation is completed.
[0114] Please refer to Figure 3 , which is a comparison result diagram of the labels on CBCT and the labels on CT after registration. Specifically, the lavender areas on the left and right sides are the pelvis, the lavender area at the bottom is the rectum, the green area in the middle is the sigmoid colon, and the pink and green areas above the middle are labeled as the small intestine. Among them, the pink is the label on CBCT, and the green is the label information on CT. According to the matching position information calculated in the embodiments of the present invention, the resampled CT image and the label information on CT are compared with the label information on CBCT, and it is found that: the coincidence degree of the label information of the pelvis is very good, the coincidence degree of the label information of the rectum is very good, but the coincidence degree of the label information of the sigmoid colon and the small intestine is relatively low (the coincidence degree of the pink and green labels is low).
[0115] Embodiment 3
[0116] Please refer to Figure 4 , Figure 4 which is a structural block diagram of an image non-rigid registration system based on deep learning provided in Embodiment 3 of the present invention. The image non-rigid registration system 200 based on deep learning includes: an acquisition module 21, a training module 22, and an input module 23, where:
[0117] The acquisition module 21 is used to acquire the first image of the user and the second image obtained by rigid registration of the first image, and generate a training set;
[0118] The training module 22 is used to establish a non-rigid registration model, and input the first image and the second image into the non-rigid registration model for training to obtain a target model. After the first image and the second image are input into the non-rigid registration model, they first undergo linear embedding processing and rearrangement processing to convert the images into a sequence of feature vectors. In addition, the parameters of the non-rigid registration model are randomly initialized, and parameter optimization is performed under the guidance of a loss function. Among them, the loss function is composed of an image similarity measure, a label constraint, and a matching position continuity constraint, and the expression is:
[0119]
[0120] where Loss represents the total loss, and α, β, and γ are non-negative real numbers used to control the contribution degree of different losses to the overall loss, represents the loss of the image similarity measure, represents the label constraint loss of the same organ in the first image and the second image, represents the matching position continuity constraint loss;
[0121] The input module 23 is used to input the to-be-registered image of the same type as the second image into the target model and output the registration result;
[0122] The non-rigid registration model includes a first encoder, a second encoder, and a decoder;
[0123] Both the first encoder and the second encoder are composed of a number of sliding window transformation modules, which are used to calculate the feature vector sequences of different scales of the first image and the second image respectively. Among them, by combining the sliding window transformation modules, a sliding window transformation module group is obtained. The sliding window transformation module groups in the first encoder and the second encoder are matched. Among them, one end of the output of the sliding window transformation module group is connected to the next-level sliding window transformation module group through a sub-block merging layer, and the other end is connected to the corresponding interactive attention module. When the sliding window transformation module group is the last-level sliding window transformation module group, the output is only connected to the first cross-attention module. The first cross-attention module is used to perform global attention processing on the feature vector sequence with the smallest size, and the second cross-attention module is used to perform in-window attention processing on the feature vector sequences of other sizes. The first cross-attention module and the second cross-attention module are connected through an upsampling module. In addition, the inputs of the first cross-attention module and the second cross-attention module are the feature vector sequences output by the sliding window transformation module group of the corresponding level of the first image, the feature vector sequences output by the sliding window transformation module group of the corresponding level of the second image, and the position labels corresponding to the feature vector sequences output by the sliding window transformation module group of the corresponding level of the second image;
[0124] The decoder is used to determine the position information of the voxels on the second image that match each position on the first image according to the feature vector sequences, where the feature vector sequence with the smallest size is subjected to global attention processing, and the feature vector sequences of other sizes are subjected to in-window attention processing.
[0125] Embodiment 4
[0126] On the other hand, the present invention also proposes an electronic device. Please refer to Figure 5 , which shows the electronic device in Embodiment 4 of the present invention, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor. When the processor 10 executes the computer program 30, it implements the above-mentioned image non-rigid registration method based on deep learning.
[0127] Among them, in some embodiments, the processor 10 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips, which are used to run the program code stored in the memory 20 or process data, such as executing an access restriction program, etc.
[0128] Among them, the memory 20 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 20 can be an internal storage unit of the electronic device in some embodiments, such as the hard disk of the electronic device. The memory 20 can also be an external storage device of the electronic device in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc. equipped on the electronic device. Further, the memory 20 can also include both an internal storage unit and an external storage device of the electronic device. The memory 20 can be used not only to store application software and various types of data of the electronic device, but also to temporarily store data that has been output or will be output.
[0129] It should be noted that Figure 5 The structure shown does not constitute a limitation on the electronic device. In other embodiments, the electronic device may include fewer or more components than shown in the figure, or combine certain components, or have a different component arrangement.
[0130] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the image non-rigid registration method based on deep learning as described above.
[0131] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0132] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0133] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0134] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0135] The above embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several variations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A non-rigid image registration method based on deep learning, characterized in that: The method comprises: Acquire a first image of the user and a second image obtained by rigid body registration of the first image to generate a training set; Establishing a non-rigid registration model, and inputting the first image and the second image into the non-rigid registration model for training to obtain a target model; Inputting an image to be registered of the same type as the second image into the target model, and outputting a registration result; The non-rigid registration model includes a decoder, and a first encoder and a second encoder that are independent of each other; The first encoder and the second encoder are both composed of a plurality of sliding window transformation modules, and are used to respectively calculate feature vector sequences of different scales of the first image and the second image; The decoder is used to determine the position information of the voxels on the second image that match each position on the first image according to the feature vector sequence, wherein the feature vector sequence with the smallest size is subjected to global attention processing, and the feature vector sequences of other sizes are subjected to in-window attention processing; After the first image and the second image are input into the non-rigid registration model, they are first subjected to linear embedding processing and rearrangement processing to convert the images into feature vector sequences; The sliding window transformation module group is obtained by combining the sliding window transformation module groups, and the sliding window transformation module groups in the first encoder and the second encoder are matched, wherein one end of the output of the sliding window transformation module group is connected to the sliding window transformation module group of the next level through the sub-block merging layer, and the other end is connected to the corresponding interactive attention module, and when the sliding window transformation module group is the last level sliding window transformation module group, the output is only connected to the first cross attention module, and the first cross attention module is used to perform global attention processing on the feature vector sequence with the smallest size; The second cross attention module is used to perform in-window attention processing on feature vector sequences of other sizes, and the first cross attention module and the second cross attention module are connected through an upsampling module; The inputs of the first cross-attention module and the second cross-attention module are the feature vector sequence output by the sliding window transformation module group of the corresponding level of the first image, the feature vector sequence output by the sliding window transformation module group of the corresponding level of the second image, and the position labels corresponding to the feature vector sequence output by the sliding window transformation module group of the corresponding level of the second image.
2. The deep learning-based non-rigid image registration method according to claim 1, characterized in that: In the step of establishing a non-rigid registration model and inputting the first image and the second image into the non-rigid registration model for training to obtain a target model, the parameters of the non-rigid registration model are randomly initialized and optimized under the guidance of a loss function, wherein the loss function is composed of an image similarity measure, a labeling constraint, and a matching position continuity constraint, and the expression is: Among them, Loss is expressed as the total loss, α, β, and γ are non-negative real numbers used to control the contribution of different losses to the overall loss. Expressed as image similarity measure loss, Denoted as the label constraint loss of the same organ in the first image and the second image, It is expressed as the matching position continuity constraint loss.
3. A deep learning-based image non-rigid registration system, characterized in that: For implementing the deep learning-based non-rigid image registration method according to any one of claims 1 to 2, the system comprises: An acquisition module, used to acquire a first image of a user and a second image obtained by rigid body registration of the first image, and generate a training set; A training module, used for establishing a non-rigid registration model, and inputting the first image and the second image into the non-rigid registration model for training to obtain a target model; An input module, used for inputting an image to be registered of the same type as the second image into the target model, and outputting a registration result; The non-rigid registration model includes a decoder, and a first encoder and a second encoder that are independent of each other; The first encoder and the second encoder are both composed of a plurality of sliding window transformation modules, and are used to respectively calculate feature vector sequences of different scales of the first image and the second image; The decoder is used to determine the position information of voxels on the second image that match each position on the first image based on the feature vector sequence, wherein the feature vector sequence with the smallest size is subjected to global attention processing, and the feature vector sequences of other sizes are subjected to in-window attention processing.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the deep learning-based non-rigid image registration method as described in any one of claims 1-2 is implemented.
5. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for non-rigid image registration based on deep learning as described in any one of claims 1 to 2 is implemented.
Citation Information
Patent Citations
Multi-size window Transform network cloth image registration method and system based on cross attention
CN116934820A