Image registration method and related equipment

By combining style transfer and rigid registration techniques in image registration in radiation therapy, the problem of poor registration effect in the existing technology when image offset is large or style difference is large, higher accuracy and stability are achieved, and suitable for a variety of scenarios.

CN120070164APending Publication Date: 2025-05-30MEVION MEDICAL EQUIPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510129789.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, it is difficult to achieve the best results in the registration range and accuracy at the same time when image registration is registered in radiation therapy, especially when image offset is large or the style of real X-ray images and DRR images are large.

Method used

Using a method combining style transfer and rigid registration, by adding an image style transfer network on the basis of a single rigid registration neural network, using the style transfer network to migrate the real X-ray image style to a style similar to the DRR image, generate DRR-like images, and combine the structural loss and pose loss of the image when training the rigid registration neural network.

Benefits of technology

It improves the accuracy and stability of image registration, can process X-ray images under different patients, different lesion types and different equipment conditions, enhances the applicability of registration, and maintains a faster registration speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070164A_ABST
    Figure CN120070164A_ABST
Patent Text Reader

Abstract

The invention discloses an image registration method and related equipment, and relates to the technical field of radiotherapy, and the method comprises the steps: obtaining a current position X-Ray image group of a patient, and the X-Ray image group comprises X-Ray images in two different directions; inputting the X-Ray image group into a style migration neural network to obtain a DRR-like image group; inputting the DRR-like image group into a rigid registration neural network to obtain a predicted posture of the patient; an image style migration network is added on the basis of a single rigid registration neural network, so that style migration and rigid registration are combined, meanwhile, when the rigid registration neural network is trained, the structure loss and the attitude loss of an image are combined, the registration precision and stability are improved, and meanwhile the high registration speed is kept; by using the differentiable DRR image, the network can better carry out back propagation during training, so that the network is more effectively trained and the registration performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of radiotherapy, and particularly to an image registration method and related equipment. Background Art

[0002] In the field of medical image processing, image registration is a crucial technology, which aims to superimpose two or more images taken at different times, in different modalities, or from different perspectives, so that the corresponding points in these images can correspond one by one; this technology is particularly important in radiotherapy (RT), because it can ensure that the radiation accurately irradiates the tumor tissue while minimizing the impact on surrounding healthy tissues.

[0003] Traditional non-AI image registration algorithms usually rely on optimizers to register images; although these methods have the characteristics of stability and high precision in some cases (the precision can reach the sub-millimeter level), they are often limited by the performance of the optimizer and it is difficult to achieve the best results in both the registration range and registration precision at the same time; in addition, for cases where the image offset is large, traditional algorithms may not be able to achieve good registration because the optimizer is prone to falling into local optima.

[0004] In recent years, with the rapid development of deep learning technology, deep learning-based registration algorithms have gradually emerged; compared with traditional registration algorithms, deep learning-based methods can still register images with large offsets. However, due to the "black box" nature of deep learning itself, the registration effect may be unstable, especially for cases not included in the training data, the registration effect may be poor. In addition, although deep learning-based registration algorithms have advantages in registration speed and range, their registration precision often cannot reach the sub-millimeter level, which is a key requirement in radiotherapy.

[0005] Refer to the attached Figure 1 , in the current neural network-based 2D / 3D rigid registration method, digital reconstructed radiography (DRR) is mainly used for training; however, due to the large style difference between actual X-ray images and DRR images, the network trained using DRR images does not perform well when processing real X-ray images; this problem limits the effect of deep learning-based registration algorithms in practical applications. Summary of the Invention

[0006] Based on the above problems, the object of the present invention is to provide an image registration method and related devices. An image style transfer network is added on the basis of a single rigid registration neural network, aiming to combine style transfer and rigid registration. At the same time, when training the rigid registration neural network, the structural loss and pose loss of the image are combined to improve the registration accuracy and stability, while maintaining a relatively fast registration speed; by using differentiable DRR images, the network can better perform backpropagation during training, thereby training the network more effectively and improving the registration performance.

[0007] The object of the present invention is achieved by the following technical solutions:

[0008] In the first aspect, the present application provides an image registration method for predicting the pose of a patient, and the method includes:

[0009] Obtain a group of X-Ray images of the patient's current position, and the group of X-Ray images includes X-Ray images in two different directions;

[0010] Input the group of X-Ray images into a style transfer neural network to obtain a group of DRR-like images;

[0011] Input the group of DRR-like images into a rigid registration neural network to obtain the predicted pose of the patient.

[0012] Preferably, the style transfer neural network is trained with multiple X-ray images and corresponding DRR images.

[0013] Preferably, based on the architecture of CycleGAN, train the style transfer neural network; the architecture of CycleGAN includes a generator and a discriminator; optimize the parameters of the generator and the discriminator by minimizing the first total loss function;

[0014] Among them, the first total loss function is determined by the CycleGAN loss and the SCC loss; the SCC loss is the structural similarity loss.

[0015] Preferably, through the patient's planned CT image, obtain multiple groups of DRR image groups and their corresponding poses; train the rigid registration neural network with multiple groups of the DRR image groups and their corresponding poses.

[0016] Preferably, perform forward geometric projection on the patient's planned CT image in the space of the setup coordinate system, and generate differentiable DRR images through the pytorch framework.

[0017] Preferably, the step of obtaining multiple groups of DRR image groups and their corresponding poses through the patient's planned CT image includes:

[0018] Based on the planned CT image of the patient, obtain the reference position of the patient and the pose at the reference position;

[0019] Move the planned CT image of the patient in different directions in the three-dimensional space based on the reference position, obtain the moved positions, and record the movement parameters;

[0020] Simulate DR sampling in two directions at each position to obtain a group of DRR images at that position;

[0021] Based on the movement parameters and the pose at the reference position, obtain the poses corresponding to the groups of DRR images at each moved position.

[0022] Preferably, training the rigid registration neural network with multiple groups of the DRR image groups and the poses corresponding to them includes:

[0023] Perform model training by minimizing the second total loss function to optimize the parameters of the rigid registration neural network; the second total loss function includes pose error and structure error, where the structure error includes a rotation part error and a translation part error.

[0024] Preferably, the second total loss function is obtained by the following formula:

[0025]

[0026]

[0027] Among them, Loss2 is the loss function of the rigid registration neural network, that is, the second total loss function; is the pose error, used to measure the difference between the predicted pose and the true pose T; is the structure error; T is the true transformation matrix; is the predicted transformation matrix; f is the image feature; drr_graditude is the DRR gradient image; drr_feature is the DRR feature image; |||| is the norm; ZNCC is the normalized cross-correlation coefficient; R is the true rotation matrix; is the predicted rotation matrix; T pre_pose is the predicted pose part, T gd_pose is the true pose part; is the inverse matrix of the predicted pose part; t is the true translation vector; is the predicted translation vector; is the rotation part error; is the translation part error; is the logarithmic mapping difference between the transpose of the predicted rotation matrix and the true rotation matrix.

[0028] In a second aspect, the present application provides an image registration device for predicting the posture of a patient, and the device includes:

[0029] an X-Ray image acquisition module for acquiring a group of X-Ray images of the patient's current position, where the group of X-Ray images includes X-Ray images in two different directions;

[0030] a DRR-like image acquisition module for inputting the group of X-Ray images into a style transfer neural network to obtain a group of DRR-like images;

[0031] a posture acquisition module for inputting the group of DRR-like images into a rigid registration neural network to obtain the predicted posture of the patient.

[0032] In a third aspect, the present application provides a positioning system, and the system includes:

[0033] the image registration device of the present application for determining the posture of the patient during the positioning process.

[0034] In a fourth aspect, the present application provides a treatment system, and the system includes:

[0035] a radiotherapy device for performing radiotherapy on the patient through a treatment plan;

[0036] the image registration device of the present application for determining the posture of the patient during the positioning process.

[0037] In a fifth aspect, the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor implements the computer program, it implements the functions of any system of the present invention or executes the steps of any method of the present invention.

[0038] In a sixth aspect, the present invention provides a computer-readable storage medium, characterized in that the storage medium stores computer instructions, and when a computer reads the computer instructions, the computer implements the functions of any system of the present invention or executes the steps of any method of the present invention.

[0039] Compared with the prior art, the beneficial effects of the present invention at least include: by combining a style transfer neural network and a rigid registration neural network, it is possible to first transfer the style of a real X-ray image to a style similar to that of a DRR image to generate a pseudo-DRR image, effectively reducing the style difference between the X-ray image and the DRR image, enabling the subsequent rigid registration neural network to process the input image more accurately, thereby improving the registration accuracy; compared with traditional non-AI registration algorithms, the present application avoids the problem of the optimizer falling into local optima and can achieve better results in both the registration range and the registration accuracy. The introduction of the style transfer neural network enables this method to process X-ray images under different patients, different lesion types, and different device conditions, enhancing the applicability of registration; when training the rigid registration neural network, combining the structural loss and the pose loss of the image helps to improve the performance of the network; the structural loss can ensure that the registered image is consistent with the original image in structure, while the pose loss can ensure that the registered pose is close to the real pose; by performing forward geometric projection in the space of the positioning coordinate system and using the PyTorch framework to generate a differentiable DRR image, it is possible to enable the registration neural network to better perform backpropagation during training, thereby training the network more effectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a schematic diagram of conventional image registration according to an embodiment of the present invention;

[0041] Figure 2 is a schematic diagram of the image registration method according to an embodiment of the present invention;

[0042] Figure 3 is a schematic diagram of the style transfer neural network according to an embodiment of the present invention;

[0043] Figure 4 is a schematic diagram of the DRR projection according to an embodiment of the present invention;

[0044] Figure 5 is a schematic diagram of the rigid registration neural network according to an embodiment of the present invention;

[0045] Figure 6 is a schematic diagram of the pose prediction result according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this invention will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art; like reference numerals in the figures denote the same or similar structures, and thus their repetitive description will be omitted.

[0047] The words describing the expression position and direction in the present invention are all illustrated by taking the attached drawings as examples, but can also be changed according to needs, and all the changes made are included in the protection scope of the present invention.

[0048] Referring to the attached Figure 2 , an embodiment of the present application provides an image registration method for predicting the posture of a patient, and the method includes:

[0049] Obtain a group of X-Ray images of the patient's current position, where the group of X-Ray images includes X-Ray images in two different directions;

[0050] Input the group of X-Ray images into a style transfer neural network to obtain a group of DRR-like images;

[0051] Input the group of DRR-like images into a rigid registration neural network to obtain the predicted posture of the patient.

[0052] The working principle and effect of the above technical solution are as follows: First, a group of X-Ray images of the patient's current position in two different directions (such as orthogonal directions, front-back and left-right) are obtained through a medical device (digital X-ray), and the group of X-Ray images includes X-Ray images in two different directions; the images reflect the projection of the patient's bone structure in the current posture; the obtained group of X-Ray images is input into a style transfer neural network; this neural network is trained to be able to convert the style of real X-Ray images into a style similar to that of digital reconstructed radiographs (DRR), thereby generating a group of DRR-like images; the purpose of style transfer is to reduce the style difference between real X-Ray images and DRR images, so that the subsequent rigid registration neural network can process these images more effectively; subsequently, the generated group of DRR-like images is input into a rigid registration neural network; the network is a neural network specially trained for image registration; the rigid registration neural network extracts features from the input group of DRR-like images, including structural features, texture features, etc. in the images; then, through the internal calculation and matching mechanism of the network, these features are compared and matched with the patterns and knowledge learned by the network during the training process to find the posture information that best matches the input image; finally, the network outputs the predicted posture of the patient, and this predicted posture includes translations in the X, Y, and Z directions and rotations around the X, Y, and Z directions, that is, a registration result with six degrees of freedom. This result can accurately reflect the patient's current posture information and provide an important basis for precise positioning and treatment in medical processes such as radiotherapy.

[0053] In summary, the image registration method of the embodiments of the present application effectively improves the accuracy and stability of patient pose prediction by combining style transfer and rigid registration techniques. The style transfer neural network reduces the image style differences, enabling the rigid registration neural network to process images more accurately and output reliable pose prediction results.

[0054] In some embodiments, the style transfer neural network is trained with multiple X-ray images and corresponding DRR images;

[0055] In some embodiments, the training of the style transfer neural network with multiple X-ray images and corresponding DRR images includes:

[0056] Obtain multiple original samples under different patients, different lesion types, and different device conditions; the original samples include original X-ray images and corresponding DRR images;

[0057] Perform data augmentation on the original samples to generate diverse training samples;

[0058] Train the style transfer neural network with the original samples and the generated samples.

[0059] Among them, the performing data augmentation on the original samples to generate diverse training samples includes:

[0060] Perform one or more preprocessing operations on the original X-ray images; the preprocessing operations include geometric transformations (such as scaling, translation), morphological operations (such as erosion, dilation), and inserting known lesion patterns (such as different types of tumors);

[0061] For each preprocessed X-ray image, perform the same operations on its corresponding DRR image using the same parameters, and ensure that the lesion patterns are fused in the same position and manner;

[0062] After completing the preprocessing, perform registration checks to ensure that the key anatomical structures are still accurately aligned in the X-ray images and DRR images;

[0063] Determine the proportion of training samples (the degree of data augmentation) according to the occurrence frequency of lesions; for example, common lesions can account for a larger proportion in the training set, while rare lesions are appropriately increased but not excessively.

[0064] The working principle and effect of the above technical solution are: Obtain the original X-ray images and corresponding DRR images from multiple patients, different lesion types, and different device conditions as training samples. These samples cover a wide range of lesion situations and device differences, which helps to train a more robust neural network;

[0065] Preprocess the original X-ray image to obtain the processed X-ray image. The preprocessing includes performing transformations at different scales (such as scaling and translation) on the original X-ray image and combining morphological operations (such as erosion and dilation) to simulate X-ray images taken at different distances; establish a knowledge base containing various lesion patterns and their common locations. Through the knowledge base, insert known lesion patterns, that is, incorporate known lesions or abnormal patterns (such as different tumors) into normal images to create synthetic X-ray images with specific pathological features; by inserting known lesion patterns, such as different types of tumors, to simulate real clinical situations, so as to better perform style transfer when facing actual X-ray images with lesions.

[0066] While performing geometric transformation and lesion insertion on the X-ray image, perform exactly the same geometric transformation on the corresponding DRR image using the same transformation parameters, and fuse the lesion pattern at the same position and in the same manner to ensure the correspondence and consistency between the X-ray image and the DRR image, enabling the style transfer neural network to learn the style mapping relationship between the two under various changing situations.

[0067] Perform registration checks after transformation and fusion to ensure that the key anatomical structures are still accurately aligned.

[0068] Control the geometric transformation parameters, including:

[0069] Limit the scaling ratio to ensure that there is no excessive magnification or reduction that would change the essential features of the lesion or anatomical structure.

[0070] Set the maximum translation distance to prevent the lesion pattern from being moved to an impossible position.

[0071] Control the rotation angle, especially for directional structures (such as bones), to maintain their natural pose.

[0072] Erosion and dilation are only used to simulate minor imaging artifacts and not to significantly change the image content. The possibilities under actual imaging conditions should be considered when using them;

[0073] By performing multi-dimensional preprocessing on the original X-ray images, such as different scale transformations, morphological operations, and inserting known lesion patterns; the scaling and translation at different scales simulate the imaging effects at different shooting distances, the morphological operations reproduce imaging artifacts, and inserting lesion patterns complements various pathological features, greatly enriching the data diversity and providing a large amount of diverse materials for network learning, enabling it to accurately capture the style differences and mapping relationships between X-ray and DRR images; when performing geometric transformations and inserting lesions on the X-ray images, the corresponding DRR images are processed synchronously using the same parameters, and it is ensured that the lesion fusion methods and positions are consistent. In this way, the correspondence and consistency between the two images are maintained, and the style transfer neural network can focus on learning style mapping rather than being trapped by interference factors such as image misalignment and mismatched lesion positions, improving both the learning efficiency and effect; performing the key step of registration inspection can confirm that the key anatomical structures are still accurately aligned after the transformation and fusion operations, preventing the misalignment of key information caused by errors in data preprocessing, ensuring the high quality of the training data, laying a solid foundation for subsequent network training, and avoiding the "pollution" of the network learning process by incorrect data; the fine adjustment of the geometric transformation parameters avoids excessive scaling, unreasonable translation, and out-of-control rotation, and at the same time standardizes the degree of use of morphological operations; it not only preserves the essential features of the lesions and anatomical structures but also conforms to the actual imaging conditions, making the generated data close to real clinical images, enabling the style transfer network to learn realistic style mapping rather than unrealistic false relationships.

[0074] Refer to the appendix Figure 3 In some embodiments, based on the architecture of CycleGAN, the style transfer neural network is trained; the architecture of CycleGAN includes a generator and a discriminator; the parameters of the generator and the discriminator are optimized by minimizing the first total loss function;

[0075] Among them, the first total loss function is determined by the CycleGAN loss and the SCC loss; the SCC loss is the structural similarity loss;

[0076] In some embodiments, the first total loss function is obtained by the following formula:

[0077]

[0078] Loss1 = L CycleGAN + λ SCC ·L SCC

[0079] Among them, Loss1 is the total loss function of the style transfer neural network, that is, the first total loss function; λ SCC is a hyperparameter used to balance the influence of the two losses, L SCC is the structural similarity loss; L CycleGANis the CycleGAN loss, including the cycle consistency loss and the adversarial loss; N X-Ray is the total number of pixels in the X-Ray image; N DRR is the total number of pixels in the DRR image; G DRR (I X-Ray ) represents the process of converting the X-Ray image I X-Ray into a Digitally Reconstructed Radiograph (DRR) image through a generator network GRR; the purpose of the generator network GRR is to learn the mapping from the X-Ray image to the DRR image; SCC(I X-Ray ) represents calculating the structural similarity index between the original X-Ray image and the reference image; the structural similarity index is a metric for measuring the visual similarity between two images, taking into account the brightness, contrast, and structural information of the images; SCC(G DRR (I X-Ray )) represents calculating the structural similarity index between the DRR image converted by the generator network and the reference image; G X-Ray (I DRR ) represents the process of converting the Digitally Reconstructed Radiograph (DRR) image I DRR into an X-Ray image through a generator network G x-Ray ; the purpose of the generator network G x-Ray is to learn the mapping from the DRR image to the X-Ray image; SCC(I DRR ) represents calculating the structural similarity index between the original DRR image I DR- and the reference image; SCC(G X-Ray (I DRR )) - SCC(I DRR ) represents calculating the structural similarity index between the X-Ray image converted by the generator network and the reference image.

[0080] The working principle of the above technical solution is as follows: The above solution constructs a pixel adaptive network based on CycleGAN to achieve style transfer from X-ray images to DRR images. Its core goal is to train the generator and discriminator so that the generator can learn the mapping relationship from X-ray images to DRR images, while ensuring the consistency and high quality of the generated images in terms of content, structure, etc. The original X-ray image is converted into a DRR image through the generator, and this DRR image can be converted back into an X-ray-like image through the generator; similarly, the original DRR image can also be converted and reverse-converted through a similar process.

[0081] The total loss function of the style transfer neural network, i.e., the first total loss function, includes the CycleGAN loss and the SCC loss. is the coefficient hyperparameter of the SCC loss function, which is used to balance the influence of the two losses.

[0082] The parameters of the generator and discriminator are optimized by minimizing the first total loss function;

[0083] Identity Loss: Ensure that when the input image already belongs to the target domain, the generator does not change the image content;

[0084] Adversarial Loss: The discriminator is used to distinguish between real images and generated images to improve the quality of the generated images;

[0085] Cycle Loss: Ensure that the converted image can be reversed back to the original image through reverse conversion to maintain the consistency of the image content;

[0086] SCC Loss(L ScC ): Structural similarity loss, which is used to measure the structural similarity between the generated image and the target image to improve the quality of image conversion; The SCC Loss is calculated by comparing the structural similarity between the generated image and the target image;

[0087] During the training process, by continuously inputting a large number of X-ray images and DRR images, the difference between the network output and the target is calculated according to the above loss function, and the parameters of the generator and discriminator are updated through the backpropagation algorithm to minimize the total loss function; As the training progresses, the generator gradually learns an effective mapping from X-ray images to DRR images, and can generate high-quality, style-similar and structurally consistent DRR images, providing a good data basis for subsequent tasks such as image registration.

[0088] Through the Adversarial Loss, the discriminator continuously prompts the generator to improve the quality of the generated images, making the generated DRR images closer to real DRR images in style, thus achieving high-quality style transfer from X-ray images to DRR images; This provides a more reliable data basis for subsequent medical analysis and processing based on DRR images. For example, in radiotherapy planning, more accurate DRR images help to more accurately determine the location and shape of tumors, thereby improving the accuracy of radiotherapy.

[0089] The introduction of SCC Loss (Structural Similarity Loss) ensures a high degree of structural similarity between the generated image and the target image, meaning that during the style transfer process, important structural information of the image is retained, and key anatomical structures and other information will not be lost or distorted due to the change in style; for example, when a doctor views the image after style transfer, they can clearly identify the structural relationship between the lesion site and its surrounding tissues without being misled by the change in image style.

[0090] Identity Loss ensures that when the input image already belongs to the target domain, the generator will not make unnecessary modifications to it, which helps to stabilize the training process of the network; in the initial stage of training, the network parameters may fluctuate greatly. Identity Loss can limit this fluctuation to a certain extent, prevent the generator from overfitting or producing abnormal outputs, thereby making the training process smoother and more controllable, and improving the efficiency and success rate of training.

[0091] By minimizing the first total loss function to optimize the parameters of the generator and discriminator, various factors are comprehensively considered, including the fidelity of the style (reflected by the CycleGAN loss), the structural similarity (reflected by the SCC loss), etc. This comprehensive design of the loss function enables the network to be optimized in multiple dimensions, thereby learning a more comprehensive and accurate image mapping relationship. The optimization of the parameters is more effective, can better adapt to different types and styles of input images, and improves the generalization ability of the network.

[0092] In some embodiments, multiple groups of DRR image groups and their corresponding poses are obtained from the planned CT images of the patient; a rigid registration neural network is trained with the multiple groups of DRR image groups and their corresponding poses; the pose is a 6-degree-of-freedom pose; that is, the translational degrees of freedom and rotational degrees of freedom in the X, Y, and Z directions.

[0093] Refer to the appendix Figure 4 In some embodiments, the planned CT image of the patient is subjected to a forward geometric projection in the space of the setup coordinate system, and a differentiable DRR image is generated through the pytorch framework.

[0094] The working principle and effects of the above technical solution are as follows: The planned CT image of the patient contains detailed anatomical structure information inside the patient's body, which is presented in the form of tomographic images. The planned CT image of the patient is subjected to a forward geometric projection in the space of the positioning coordinate system. This process simulates the physical process of X-rays passing through the patient's body and forming a projection on the imaging plane. In this way, a group of DRR (Digitally Reconstructed Radiograph) images in two different directions (preferably orthogonal directions) at the same position of the CT image can be generated. The DRR images are similar to actual X-ray images, but they are calculated from CT data and can reflect the projection of the patient's body from a specific perspective.

[0095] Use the PyTorch framework to generate differentiable DRR images. Differentiability enables the calculation of gradients through the backpropagation algorithm during the training process, thereby optimizing network parameters. Specifically: First, read the planned CT image data of the patient, which is usually organized in a three-dimensional tensor format to adapt to the calculation requirements of PyTorch. For example, use the torchvision library or numpy combined with the torch.from_numpy method to convert the CT image data into a torch.Tensor while retaining its spatial dimension information.

[0096] The pixel value range of CT images is relatively large. To accelerate model convergence and training stability, it is necessary to normalize the image data. A common approach is to map the pixel values to a specific interval, such as [0,1] or [-1,1]. A linear transformation formula can be used to complete the normalization operation.

[0097] Based on the physical imaging principle, construct a mathematical model for forward geometric projection, which involves simulating the process of X-rays penetrating the three-dimensional volume data corresponding to the CT image data from different angles. At the code level, coordinate transformation needs to be considered. For example, convert the voxel coordinates of the CT image according to parameters such as the projection direction and the source-to-detector distance to determine the projection position of each voxel on the detector plane. These calculations are usually implemented in the form of matrix operations.

[0098] Adopt classical algorithms such as Ray Casting to generate DRR images. In PyTorch, use loops and tensor operations to gradually simulate the process of light passing through the CT volume data, accumulate the attenuation values of each light passing through the voxels, and finally form a two-dimensional projection image. Since the entire process is completed with tensor operations, PyTorch will automatically trace the computational graph to prepare for subsequent derivative calculation.

[0099] To train the projection model or use the generated DRR images for other tasks later, it is necessary to define a loss function. For example, if we want the generated DRR images to be as similar as possible to the real DRR images, we can use the mean squared error (MSE) loss function torch.nn.MSE_loss, or more complex loss metrics such as structural similarity loss (SSIM).

[0100] After the model generates DRR images during forward propagation, calculate the value of the loss function, and then call loss.backward(). PyTorch's automatic differentiation engine will, based on the computational graph, traverse the previous calculation steps backward from the loss function, calculating the gradient of each tensor involved in the operation with respect to the loss; using these gradients, update the model parameters through an optimizer (such as torch.optim.Adam), gradually iteratively optimizing the projection model to continuously improve the quality of the generated DRR images while maintaining the differentiability of the entire process, enabling the model to have a clear gradient direction guidance during improvement.

[0101] After completing the above model construction and parameter optimization training, given new CT image data input, through the trained forward geometric projection model, high-quality and differentiable DRR images can be output; these DRR images are not only an effective presentation form of medical image data, but also due to their differentiability, when combined with other neural network modules later, they can be smoothly jointly trained and parameter fine-tuned. For example, as the input of a rigid registration neural network, they can continue to participate in more complex medical image analysis tasks.

[0102] During the process of generating DRR images, record the pose information corresponding to each group of DRR images at the same time. Here, the pose is a 6-degree-of-freedom pose, that is, it includes the translational degrees of freedom in the X, Y, and Z directions and the rotational degrees of freedom around the X, Y, and Z directions;

[0103] Input multiple generated groups of DRR images and their corresponding 6-degree-of-freedom poses into the rigid registration neural network; the neural network extracts features from the input DRR images, and these features include information such as edges, textures, and gray-scale distributions in the images. There is a certain potential mapping relationship between these features and the pose information; during the training process, the network continuously adjusts its internal parameters to learn the mapping rule from DRR image features to 6-degree-of-freedom poses.

[0104] Obtaining multiple sets of DRR images with 6 - degree - of - freedom poses from the patient's planned CT images means covering all possible position and angle changes of the object in three - dimensional space; this comprehensive data generation method simulates various potential positioning states of the patient's body in the real world, providing a data basis highly fitting the actual application scenario for the subsequent neural network training of this patient, enabling the trained rigid registration neural network to handle various pose situations; a large number of samples with different poses are input for training, which can effectively avoid the model overfitting to a single data pattern, enhance the generalization ability of the model, and enable it to still complete the registration task stably and accurately when facing the complex and diverse patient poses in real clinical practice.

[0105] Generating differentiable DRR images with the help of the PyTorch framework opens the door wide for the backpropagation algorithm in neural network training; backpropagation depends on calculating the gradient of the loss function with respect to the network parameters, and differentiable DRR images allow for the accurate calculation of these gradients, enabling the network to quickly and accurately adjust the parameters during training, accelerating the convergence speed, reducing the training time and computational resource consumption, and improving the overall training efficiency. Differentiability gives developers greater flexibility to optimize and adjust the model architecture. When it is found that the model prediction effect is not ideal, or when new optimization strategies need to be implemented, due to the differentiability of DRR images, developers can more smoothly modify the network structure and adjust the loss function, continuously iteratively optimizing the rigid registration neural network without affecting the coherence of the entire training process, and improving the model performance continuously.

[0106] In some embodiments, obtaining multiple sets of DRR image groups and their corresponding poses through the planned CT images of the patient includes:

[0107] Obtaining the reference position of the patient and the pose at the reference position according to the planned CT image of the patient;

[0108] Moving the patient's planned CT image in different directions in three - dimensional space based on the reference position to obtain the moved positions and record the movement parameters;

[0109] Simulating DR sampling in two directions at each position to obtain the DRR image group at that position;

[0110] Obtaining the poses corresponding to the DRR image groups at each moved position through the movement parameters and the pose at the reference position;

[0111] Training the rigid registration neural network for this patient through the multiple sets of DRR images and corresponding poses of this patient obtained by simulation;

[0112] Inputting the class - DRR image group into the rigid registration neural network to obtain the predicted pose of the patient.

[0113] The working principle of the above technical solution is as follows: The planned CT image of the patient is a three-dimensional tomographic scan data set that presents the anatomical structure of the patient's body in all directions. First, a simulated random sampling of the patient's posture is performed, that is, a reference position (usually the tumor center position) is determined from the patient's planned CT image. This reference position can be regarded as the starting reference point for subsequent operations. At the same time, the corresponding posture at this reference position is obtained. This posture contains 6 degrees of freedom information, that is, the translation values in the X, Y, and Z directions and the rotation angles around the X, Y, and Z directions, providing an initial posture data basis for subsequent position changes; taking the reference position as the origin, the planned CT image is moved in different directions in the three-dimensional space. This movement simulates the possible position offsets of the patient's body in the actual scenario. The movement directions cover the X, Y, and Z coordinate axes and their combined directions. Each movement will bring the CT image to a new position. At this time, the system accurately records the movement parameters corresponding to each movement, such as how many millimeters it has moved in the X-axis direction and how many degrees it has rotated around the Y-axis, etc. These movement parameters are the key intermediate data for subsequent derivation of the DRR image posture; the movement distance does not exceed 3 cm; random movements are made within 3 cm in each direction from the reference position to obtain multiple new positions after movement.

[0114] For the CT image at the reference position and each new position after movement, simulated DR (Digital Reconstruction Radiography) sampling is performed from two specific directions (preferably orthogonal directions). The process of simulated DR sampling is essentially a physical process of simulating X-rays penetrating the body tissues represented by the CT image and then forming a projection on a specific plane, and finally obtaining a set of DRR images at this position. Since sampling is performed at different positions, each set of DRR images generated carries unique information corresponding to the position, reflecting the projection differences caused by position changes.

[0115] Using the previously recorded movement parameters and the initial posture of the reference position, the postures corresponding to each set of DRR images at the moved positions can be calculated. For example, if the rotation angle around the Z-axis in the reference posture is 0°, and there is a movement parameter record of a 10° rotation around the Z-axis in a subsequent movement, then for the DRR image generated corresponding to this movement, the rotation angle of its posture around the Z-axis is 10°. Combining the translation and rotation parameter adjustments in other directions, the 6-degree-of-freedom posture of this DRR image can be completely determined.

[0116] Train a rigid registration neural network for this patient with multiple sets (such as 10,000 sets) of DRR image sets and corresponding postures obtained through simulation;

[0117] After the rigid registration neural network is trained, the DRR image group generated by the style transfer network is used as input and fed into the trained rigid registration neural network. Based on the previously learned feature-pose mapping relationship, the network quickly analyzes the features of the DRR image group and then outputs the corresponding predicted patient pose, which is also six degrees of freedom and accurately reflects the patient's current body pose, providing a key reference for subsequent medical operations. The pose prediction result refers to Appendix Figure 6 .

[0118] The effects of the above technical solution are as follows: Different patients vary greatly in terms of body structure, lesion location and scope, and even physiological functions. Starting from the planned CT images of the patient to be located, the exclusive reference position and pose information of the patient are obtained, and the corresponding DRR images and pose data are generated, fully considering the unique anatomical characteristics of the patient. The rigid registration neural network trained with such personalized data can output a registration result that accurately fits the actual situation of the patient, laying a solid foundation for subsequent precise treatment and meeting the personalized medical needs to the greatest extent. The rigid registration neural network obtained through training can quickly and accurately register the images during the treatment process, reducing the time and labor costs required for manual registration. In addition, accurate registration helps to reduce errors during the treatment process and improve the safety and reliability of the treatment: Training based on the patient's own CT data can significantly reduce the registration errors caused by individual differences and ensure that the results output by the model are more reliable. By moving the CT images in different directions in the three-dimensional space and simulating DR sampling at each position, a large number of DRR images with different perspectives and poses can be obtained. This not only increases the diversity of the training data but also improves the generalization ability of the model to various imaging conditions. Combining the movement parameters with the pose at the reference position ensures that each generated DRR image has accurate pose information. Using the patient's simulated data for training can quickly generate a large number of training samples without relying on actual X-ray images, greatly shortening the data preparation time and reducing the cost.

[0119] Refer to Appendix Figure 5 , in some embodiments, the training of the rigid registration neural network with multiple groups of the DRR image groups and their corresponding poses includes:

[0120] Extract image features through Efficient Network, which is the encoder part of the network;

[0121] Use the Decoder to reconstruct the features to restore the image information;

[0122] During the feature extraction process, use the multi-head self-attention mechanism to capture the relationships between different regions in the image;

[0123] Further process the predicted pose using a multi-layer perceptron;

[0124] Train the model by minimizing the second total loss function to optimize the rigid registration neural network parameters; the second total loss function includes pose error and structure error, where the structure error includes a rotation part error and a translation part error;

[0125] Use the trained rigid registration neural network model to predict a new pose.

[0126] In some embodiments, the second total loss function is obtained by the following formula:

[0127]

[0128] When ZNCC approaches 1, it indicates that the pixel gray value change trends of the two images are exactly the same, that is, the similarity is very high. To unify the physical meaning of the loss function, ZNCC here represents the result of 1 - ZNCC;

[0129]

[0130] and Based on the SE3 theory under Lie algebra, by mapping rotation and translation to a unified linear Lie algebra space, the numerical scale difference between rotation and translation is eliminated.

[0131] where Loss2 is the loss function of the rigid registration neural network; where Loss2, the loss function of the rigid registration neural network, is the second total loss function; is the pose error, used to measure the difference between the predicted pose and the true pose T; is the structure error; is the true transformation matrix; is the predicted transformation matrix; f is the image feature; drr_graditude is the DRR gradient image; drr_feature is the DRR feature image; |||| is the norm; ZNCC is the normalized cross-correlation coefficient; R is the true rotation matrix; is the predicted rotation matrix; T pre_pose is the predicted pose part, T gd_pose is the true pose part; is the inverse matrix of the predicted pose part; t is the true translation vector; is the predicted translation vector; is the rotation part error; is the translation part error; is the logarithmic mapping difference between the transpose of the predicted rotation matrix and the true rotation matrix.

[0132] The working principle and effect of the above technical solution are as follows: using Efficient Network as an encoder to extract features from multiple input DRR images; Efficient Network has an efficient convolution operation architecture, which can quickly capture key feature information in DRR images, such as bone contours, tissue density differences, etc., to provide a data basis for subsequent processing. The Decoder part takes over the features extracted by the encoder and tries to reversely reconstruct these features back to image information; this process is similar to "restoring a puzzle", by learning the association and combination between features, the abstract features are restored to a complete image, the purpose is to assist in subsequent posture prediction, so that the model can better understand the intrinsic relationship between image content and posture. In the feature extraction link, the multi-head self-attention mechanism begins to play a role; it can focus on different areas of the DRR image from the perspective of multiple "attention heads" at the same time, and capture the complex association relationship between these areas; for example, it may be found that the rotation angle of a certain bone structure in the image will have a specific relationship with the displacement of the surrounding soft tissue. This cross-regional relationship capture helps the model to establish a more accurate posture understanding model and improve its sensitivity to posture changes. The predicted posture information obtained in the previous steps is further processed by a multi-layer perceptron; the multi-layer perceptron has powerful nonlinear fitting capabilities, which can fine-tune and optimize the preliminary predicted posture, explore the high-order posture features hidden in the data, and make the final output predicted posture more accurate.

[0133] The rigid registration neural network is trained by minimizing the second total loss function Loss2; Loss2 consists of multiple parts, among which Measure the difference between the predicted posture and the actual posture, and intuitively reflect the accuracy of the predicted posture; Taking into account the structural error, the rotation and translation errors quantify the degree to which the predicted posture deviates from the true posture in the rotation and translation dimensions respectively; 10×Zncc(drr_graditude,drr_feature) constrains the model prediction from the perspective of image feature correlation. Here, ZNCC takes 1-ZNCC. The closer the pixel grayscale change trends of the two images are, the smaller the value of this item is, which prompts the model to generate more similar images. and Based on the SE3 theory under Lie algebra, the rotation and translation are mapped to a unified linear Lie algebra space, which cleverly solves the problem of different numerical scales of rotation and translation, allowing the network to consider the two posture change factors more balanced during optimization and avoid training bias caused by scale differences. During the training process, the calculated Loss2 value is used through the back propagation algorithm to transfer the error from the output layer back to each layer of the network, prompting the network to adjust the parameters of components such as Efficient Network, Decoder, and Multilayer Perceptron, so that the model gradually reduces Loss2 in each iteration and improves the prediction accuracy.

[0134] The embodiment of the present application provides an image registration device for predicting a patient's posture, the device comprising:

[0135] An X-Ray image acquisition module, used to acquire an X-Ray image group of the patient's current position, wherein the X-Ray image group includes X-Ray images in two different directions;

[0136] A DRR-like image acquisition module, used for inputting the X-Ray image group into a style transfer neural network to obtain a DRR-like image group;

[0137] The posture acquisition module is used to input the DRR-like image group into the rigid registration neural network to obtain the patient's predicted posture.

[0138] In some embodiments, a style transfer neural network is trained using a plurality of X-ray images and corresponding DRR images;

[0139] In some embodiments, the training of a style transfer neural network using a plurality of X-ray images and corresponding DRR images includes:

[0140] Acquire multiple original samples under different patients, different lesion types and different equipment conditions; the original samples include original X-ray images and corresponding DRR images;

[0141] Perform data augmentation on the original samples to generate diversified training samples;

[0142] The style transfer neural network is trained using the original samples and the generated samples.

[0143] The data enhancement of the original samples to generate diversified training samples includes:

[0144] Perform one or more preprocessing operations on the original X-ray image; the preprocessing operations include geometric transformation (such as scaling, translation), morphological operations (such as corrosion, expansion), and insertion of known lesion patterns (such as different types of tumors);

[0145] For each pre - processed X - ray image, perform the same operations on its corresponding DRR image using the same parameters, and ensure that the lesion patterns are fused at the same location and in the same way;

[0146] After completing the pre - processing, perform a registration check to ensure that the key anatomical structures are still accurately aligned in the X - ray image and the DRR image;

[0147] Determine the proportion of training samples (the degree of data augmentation) according to the occurrence frequency of lesions; for example, common lesions can occupy a larger proportion in the training set, while rare lesions are appropriately increased but not excessively.

[0148] In some embodiments, based on the architecture of CycleGAN, train the style transfer neural network; the architecture of CycleGAN includes a generator and a discriminator; optimize the parameters of the generator and the discriminator by minimizing the first total loss function;

[0149] Among them, determine the first total loss function through the CycleGAN loss and the SCC loss; the SCC loss is the structural similarity loss;

[0150] In some embodiments, the first total loss function is obtained by the following formula:

[0151]

[0152] Loss1 = L CycleGAN + λ SCC ·L SCC

[0153] Among them, Loss1 is the total loss function of the style transfer neural network, that is, the first total loss function; λ SCC is a hyperparameter used to balance the influence of the two losses, L SCC is the structural similarity loss; L CycleGAN is the CycleGAN loss, including the cycle consistency loss and the adversarial loss; N X-Ray is the total number of pixels of the X - Ray image; N DRR is the total number of pixels of the DRR image; g DRR (I X-Ray ) represents the process of converting the X - Ray image I X-Ray into a Digitally Reconstructed Radiograph (DRR) image through a generator network GRR; the purpose of the generator network GRR is to learn the mapping from the X - Ray image to the DRR image; SCC(I X-Ray) represents calculating the structural similarity index between the original X-Ray image and the reference image; the structural similarity index is a metric for measuring the visual similarity between two images, taking into account the brightness, contrast, and structural information of the images; SCC(G DRR (I X-Ray )) represents calculating the structural similarity index between the DRR image converted by the generator network and the reference image; G X-Ray (I DRR ) represents the process of converting the digitally reconstructed radiograph (DRR) image I DRR into an X-Ray image through a generator network G X-Ray ; the purpose of the generator network G X-Ray is to learn the mapping from the DRR image to the X-Ray image; SCC(I DRR ) represents calculating the structural similarity index between the original DRR image I DRR and the reference image; SCC(G X-Ray (I DRR )) - SCC(I DRR ) represents calculating the structural similarity index between the X-Ray image converted by the generator network and the reference image.

[0154] In some embodiments, multiple groups of DRR image groups and their corresponding poses are obtained through the patient's planned CT image; a rigid registration neural network is trained through multiple groups of the DRR image groups and their corresponding poses; the pose is a 6-degree-of-freedom pose; that is, the translational degrees of freedom and rotational degrees of freedom in the X, Y, and Z directions.

[0155] In some embodiments, the patient's planned CT image is geometrically projected forward in the space of the positioning coordinate system, and a differentiable DRR image is generated through the pytorch framework.

[0156] In some embodiments, the obtaining of multiple groups of DRR image groups and their corresponding poses through the patient's planned CT image includes:

[0157] Obtaining the reference position of the patient and the pose at the reference position according to the patient's planned CT image;

[0158] Moving the patient's planned CT image in different directions in the three-dimensional space based on the reference position to obtain the moved positions, and recording the movement parameters;

[0159] Simulating DR sampling in two directions at each position to obtain the DRR image group at that position;

[0160] Obtaining the poses corresponding to the DRR image groups at each moved position through the movement parameters and the pose at the reference position;

[0161] Train a rigid registration neural network for this patient by using multiple groups of DRR images of this patient obtained through simulation and the corresponding poses.

[0162] Training the rigid registration neural network by using multiple groups of the DRR image groups and the poses corresponding thereto includes:

[0163] Extract image features through Efficient Network, which is the encoder part of the network;

[0164] Use a Decoder to reconstruct the features to restore the image information;

[0165] During the feature extraction process, use a multi-head self-attention mechanism to capture the relationships between different regions in the image;

[0166] Use a multi-layer perceptron to further process the predicted pose;

[0167] Perform model training by minimizing the second total loss function to optimize the parameters of the rigid registration neural network; the second total loss function includes a pose error and a structure error, wherein the structure error includes a rotation part error and a translation part error;

[0168] Use the trained rigid registration neural network model to predict a new pose.

[0169] In some embodiments, the second total loss function is obtained by the following formula:

[0170]

[0171] When ZNCC approaches 1, it indicates that the pixel gray-scale change trends of the two images are completely consistent, that is, the similarity is very high. To unify the physical meaning of the loss function, ZNCC here represents the result of 1 - ZNCC;

[0172]

[0173] and Based on the SE3 theory under Lie algebra, by mapping rotation and translation to a unified linear Lie algebra space, the numerical scale differences between rotation and translation are eliminated.

[0174] wherein, Loss2 is the loss function of the rigid registration neural network; wherein, Loss2 is the loss function of the rigid registration neural network, that is, the second total loss function; is the pose error, which is used to measure the predicted pose and the difference between the true pose T; is the structure error; is the true transformation matrix; is the predicted transformation matrix; f is the image feature; drr_graditude is the DRR gradient image; drr_feature is the DRR feature image; |||| is the norm; ZNCC is the normalized cross-correlation coefficient; R is the true rotation matrix; is the predicted rotation matrix; T pre_pose is the predicted pose part, T gd_pose is the true pose part; is the inverse matrix of the predicted pose part; t is the true translation vector; is the predicted translation vector; is the rotation part error; is the translation part error; is the logarithmic mapping difference between the transpose of the predicted rotation matrix and the true rotation matrix.

[0175] The working principle and effect of the above technical solution are the same as those in the method embodiment of this application, and will not be elaborated here.

[0176] An embodiment of this application provides a positioning system, and the system includes:

[0177] The image registration device described in the embodiment of this application; used to determine the patient's pose during positioning.

[0178] An embodiment of this application further provides a treatment system, and the system includes:

[0179] A radiotherapy device, used to perform radiotherapy on a patient through a treatment plan;

[0180] The image registration device described in the embodiment of this application; used to determine the patient's pose during positioning.

[0181] An embodiment of the present invention further provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the steps of any method described in the embodiment of the present invention or the functions of the system described in the embodiment of the present invention.

[0182] An embodiment of the present invention further provides a computer-readable storage medium, which is used to store a computer program, and when the computer program is executed, it implements the steps of the method in the embodiment of the present invention. Its specific implementation manner is the same as the implementation manner and the achieved technical effect described in the above method embodiment, and some contents will not be elaborated.

[0183] In the present invention, a readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0184] The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above. The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the C language or similar programming languages. The program code can be executed entirely on the user computing device, partially on an associated device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0185] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention, and all such changes should fall within the protection scope of the claims of the present invention.

Claims

1. An image registration method for predicting a patient's posture, characterized in that The method comprises: Acquire an X-Ray image group of the patient's current position, wherein the X-Ray image group includes X-Ray images in two different directions; Inputting the X-Ray image group into a style transfer neural network to obtain a DRR-like image group; The DRR-like image group is input into a rigid registration neural network to obtain the predicted posture of the patient.

2. The image registration method according to claim 1, characterized in that: The style transfer neural network is trained by multiple X-ray images and corresponding DRR images.

3. The image registration method according to claim 2, characterized in that: Based on the architecture of CycleGAN, the style transfer neural network is trained; the architecture of CycleGAN includes a generator and a discriminator; the parameters of the generator and the discriminator are optimized by minimizing the first total loss function; Among them, the first total loss function is determined by CycleGAN loss and SCC loss; the SCC loss is a structural similarity loss.

4. The image registration method according to claim 1, characterized in that: A plurality of DRR image groups and their corresponding postures are obtained through the planned CT images of the patient; and a rigid registration neural network is trained through the plurality of DRR image groups and their corresponding postures.

5. The image registration method according to claim 4, characterized in that: The planned CT image of the patient is forward geometrically projected in the space of the positioning coordinate system, and a differentiable DRR image is generated through the pytorch framework.

6. The image registration method according to claim 4, characterized in that: The method of obtaining a plurality of DRR image groups and the postures corresponding thereto through the planned CT image of the patient comprises: Obtaining a reference position of the patient and a posture at the reference position according to the planned CT image of the patient; The planned CT image of the patient is moved in different directions in the three-dimensional space based on the reference position, the position after the movement is obtained, and the movement parameters are recorded; Simulate DR sampling in two directions at each position to obtain a DRR image group at that position; The posture of the DRR image group at each position after the movement is obtained by moving the parameters and the posture at the reference position.

7. The image registration method according to claim 4, characterized in that: The method of training a rigid registration neural network using a plurality of DRR image groups and their corresponding postures includes: The model is trained by minimizing a second total loss function to optimize the rigid registration neural network parameters; the second total loss function includes a posture error and a structural error, wherein the structural error includes a rotation error and a translation error.

8. The image registration method according to claim 7, characterized in that: The second total loss function is obtained by the following formula: Among them, Loss2 is the loss function of the rigid registration neural network, that is, the second total loss function; is the attitude error, used to measure the predicted attitude The difference between the actual posture T; is the structural error; T is the real transformation matrix; is the predicted transformation matrix; f is the image feature; drr_graditude is the DRR gradient image; drr_feature is the DRR feature image; ∥∥ is the norm; ZNCC is the normalized mutual correlation coefficient; R is the real rotation matrix; is the predicted rotation matrix; T pre_pose is the predicted posture part, T gd_pose For the real posture part; is the inverse matrix of the predicted posture part; t is the real translation vector; is the predicted translation vector; is the rotation error; is the translation error; is the logarithmic mapping difference between the transpose of the predicted rotation matrix and the true rotation matrix.

9. An image registration device for predicting a patient's posture, characterized in that The device comprises: An X-Ray image acquisition module, used to acquire an X-Ray image group of the patient's current position, wherein the X-Ray image group includes X-Ray images in two different directions; A DRR-like image acquisition module, used for inputting the X-Ray image group into a style transfer neural network to obtain a DRR-like image group; The posture acquisition module is used to input the DRR-like image group into the rigid registration neural network to obtain the patient's predicted posture.

10. A positioning system, characterized in that: The system comprises: The image registration device of claim 9 is used to determine the patient's posture during positioning.

11. A treatment system, characterized in that: The system comprises: Radiotherapy devices, used to deliver radiation therapy to patients according to treatment plans; The image registration device of claim 9 is used to determine the patient's posture during positioning.

12. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor executes the steps of any method described in claims 1-8 or implements the functions of any system described in claims 10-11 when implementing the computer program.

13. A computer-readable storage medium, characterized in that: The storage medium stores computer instructions. When a computer reads the computer instructions, the computer executes the steps of any method described in claims 1-8 or implements the functions of any system described in claims 10-11.