Radiotherapy position verification data registration method and device, equipment and storage medium
By using deep learning technology to automatically process CBCT and CT image data and generate spatial transformation parameters according to the DICOM-RT standard, the problem of low efficiency and accuracy in radiotherapy positioning verification and registration is solved. This achieves high-precision, automated positioning error correction, thereby improving radiotherapy efficacy and safety.
Patent Information
- Application Number
- CN202511292321.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-11
AI Technical Summary
In existing radiotherapy techniques, the efficiency and accuracy of position verification and registration are not high, which leads to an increase in the workload of medical staff and a decrease in the accuracy of data analysis, and is prone to recording errors or omissions.
Employing deep learning and image processing techniques such as CycleGAN modality transformation, adaptive histogram matching, 3D-ResNet-34-FPN architecture, TPS-Net, and RANSAC-SVD singular value decomposition, this system automatically processes CBCT and CT image data to generate spatial transformation parameter sequences conforming to the DICOM-RT standard.
It significantly improves registration accuracy and automation, reduces the workload of medical staff and subjective errors, provides real-time and accurate positioning error correction parameters, and ensures that the radiation dose is accurately delivered to the target area to protect the surrounding normal tissues.
Smart Images

Figure CN120823250B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of radiotherapy, and particularly relates to a radiotherapy position verification data registration method and device, equipment and a storage medium. BACKGROUND
[0002] In radiotherapy, position verification is a core link to ensure treatment accuracy and safety. It obtains real-time image data of a patient through CBCT scanning and performs registration with a positioning CT image to correct positioning errors. In this process, registration results such as positioning error data and images generated by registration have important guiding value for clinical treatment. In clinical practice, statistical analysis of positioning errors mainly relies on manual recording mode. Specifically, after each CBCT registration is completed, medical staff need to manually extract key parameters from the registration result interface, including registration mode (such as rigid / non-rigid), error value (translation and rotation components), and patient identification information, and enter them into the statistical system one by one. Subsequently, the manually recorded data also needs to be reorganized, summarized and analyzed and evaluated.
[0003] For radiotherapy centers with a large number of daily admissions, the above-mentioned manual recording method and the reorganization of a large amount of radiotherapy-related data are time-consuming and labor-intensive, increasing the workload of medical staff. For long-term data processing of medical staff, it is easy to increase the error rate, and human operation is easy to introduce recording errors or omissions, thereby reducing the accuracy of subsequent data analysis. SUMMARY
[0004] The main purpose of the present application is to provide a radiotherapy position verification data registration method, device, equipment and storage medium to solve the problem of low efficiency and low accuracy of traditional position verification and registration in the prior art.
[0005] In order to achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0006] A radiotherapy position verification data registration method, the radiotherapy position verification data is derived from CBCT verification image data of a patient to be treated on a radiotherapy device, the patient to be treated has CT image data, and the registration method comprises:
[0007] Step S1: converting the CBCT verification image data and the CT image data into body data of a preset matrix size through CycleGAN modal conversion to obtain CBCT body data and CT body data;
[0008] Step S2, locally histogram matching the CBCT volume data and the CT volume data respectively by an adaptive histogram matching algorithm, and organ-specific gray scale normalization processing the CBCT volume data and the CT volume data respectively based on a pre-segmentation mask, to obtain standardized CBCT volume data and standardized CT volume data in a unified format;
[0009] Step S3, inputting the standardized CBCT volume data and the standardized CT volume data respectively into a pre-trained 3D-ResNet-34-FPN architecture, and obtaining a CBCT feature pyramid and a CT feature pyramid respectively through learning and training of the 3D-ResNet-34-FPN architecture;
[0010] Step S4, taking the sampling feature points in the CT feature pyramid as a control point coordinate set, taking the sampling feature points in the CBCT feature pyramid as a target point coordinate set, and generating a deformation field from the control point coordinate set to the target point coordinate set through a TPS-Net architecture;
[0011] Step S5, converting the deformation field into a zero-centered displacement field through a 3D-CNN network, and extracting all affine transformation parameters of the displacement field through RANSAC-SVD singular value decomposition;
[0012] Step S6, encapsulating all the affine transformation parameters into a spatial transformation parameter sequence conforming to the DICOM-RT standard, and sending to the radiotherapy device.
[0013] Beneficial effects:
[0014] The application first reduces the modal difference between CBCT and CT images through CycleGAN modal conversion and organ-specific gray scale normalization, effectively laying a unified data foundation for subsequent registration; then uses a 3D-ResNet-34-FPN architecture to extract multi-level features from the standardized data, and combines a TPS-Net to generate a deformation field that accurately reflects local deformation, significantly improving registration accuracy, especially effectively capturing non-linear changes such as soft tissue deformation and organ displacement; finally, through 3D-CNN displacement field normalization and RANSAC-SVD decomposition, robustly extracting global affine transformation parameters, and encapsulating them into a DICOM-RT standard protocol for direct transmission to the radiotherapy device, not only automating the traditional manual operation process, reducing the workload and subjective errors of medical staff, but also improving the reliability and repeatability of the registration result, so that the application can provide real-time and accurate positioning error correction parameters for online image-guided radiotherapy (IGRT), ensuring that the radiation dose can be more accurately projected to the target area, thereby improving the treatment effect while protecting the surrounding normal tissue, so that the application has high calculation accuracy and high automation degree.
[0015] As a further improvement of the present application, step S1, the CBCT verification image data and the CT image data are unified into volume data of a preset matrix size through CycleGAN modal conversion, to obtain CBCT volume data and CT volume data, comprising:
[0016] Step S11, converting the CBCT verification image data and the CT image data into DICOM standard format;
[0017] Step S12, aligning the spatial coordinate system of the CBCT verification image data with the spatial coordinate system of the CT image data through a non-rigid registration method;
[0018] Step S13, defining two generators of the CycleGAN modal conversion based on a 3D U-Net architecture, and defining two discriminators of the CycleGAN modal conversion based on a 3D PatchGAN;
[0019] Step S14, converting the CBCT verification image data into CT style data through one of the generators, and converting the CT image data into CBCT style data through the other generator;
[0020] Step S15, calculating the adversarial loss based on the Wasserstein distance between the CBCT style data and the CBCT verification image data through one of the discriminators, and calculating the adversarial loss based on the Wasserstein distance between the CT style data and the CT image data through the other discriminator;
[0021] Step S16, defining the L1 loss of the CycleGAN modal conversion, and simultaneously iterating the two adversarial losses through back propagation until the L1 loss reaches a minimum value;
[0022] Step S17, when the L1 loss reaches the minimum value, obtaining the CT style data of the last iteration through one of the generators, and the CT style data and the CT image data are volume data unified into a preset matrix size.
[0023] Beneficial effects:
[0024] Steps S11 to S17 effectively solve the core problems of large modal difference, inconsistent gray distribution and low soft tissue contrast between CBCT and CT images due to different imaging principles through CycleGAN modal conversion technology. Specifically, by non-rigid registration to align the spatial coordinate system and the generative adversarial network, the CBCT data is converted into CT style data, which significantly reduces the noise, artifacts and resolution difference between the two, improves the data consistency, and uses the adversarial loss of Wasserstein distance and L1 cycle consistency loss to ensure that the generated image not only retains the anatomical structure authenticity, but also has the high contrast characteristics of CT images, reducing the registration error caused by the unsatisfactory CBCT image quality. The whole process does not need manual intervention or strictly paired training data, and automatically outputs the body data of uniform matrix size through the end-to-end deep learning framework, providing standardized input for subsequent registration and reducing the subjective error and work burden of traditional manual operation.
[0025] As a further improvement of the present application, step S2, the CBCT body data and the CT body data are respectively subjected to local histogram matching by an adaptive histogram matching algorithm, and organ-specific gray standardization processing is performed on the CBCT body data and the CT body data based on a pre-segmentation mask, to obtain standardized CBCT body data and standardized CT body data in a unified format, comprising:
[0026] Step S21, defining a sliding window of a preset voxel size, and traversing the standardized CBCT body data through the sliding window at a preset sliding speed;
[0027] Step S22, calculating the gray histogram of the current block once every time the sliding window slides to obtain the pixel distribution probability of the current block;
[0028] Step S23, calculating the cumulative distribution function of each block by formula (1) respectively:
[0029] (1);
[0030] Wherein, is the cumulative distribution function, is the upper limit of the gray value, is the gray value, , is the probability distribution of the gray value;
[0031] Step S24, taking the CT body data as the traversed object, repeating steps S21 to S23 to obtain the cumulative distribution function of the CT body data;
[0032] Step S25: Based on the two cumulative distribution functions of the same block, construct a grayscale mapping table according to equation (2):
[0033] (2);
[0034] in, The first CBCT volume data A mapping function that maps each grayscale value to all grayscale values in the CT body data. For the CBCT volume data of the first The cumulative distribution function of grayscale values, For the CT body data number The cumulative distribution function of grayscale values, The first CT body data One grayscale value;
[0035] Step S26: Extract at least one organ ROI region from the CT volume data and the CBCT volume data respectively using a 3D U-Net architecture pre-trained with a preset organ dataset;
[0036] Step S27: The mean and standard deviation of grayscale values for each organ ROI region are calculated, and organ-specific grayscale standardization is performed using equation (3) to obtain standardized CBCT volume data and standardized CT volume data in a unified format.
[0037] (3);
[0038] in, The normalized grayscale value. These are voxel coordinate values. coordinates The grayscale value of the voxel. The standard deviation of voxel gray values within the organ's ROI region. The mean voxel gray value within the organ's ROI region. To preset the standard deviation of grayscale of the target organ, The preset grayscale mean of the target organ.
[0039] Beneficial effects: Steps S21 to S27 effectively solve the problems of inconsistent gray scale distribution, low organ contrast, and modal statistical property deviation between CBCT and CT images caused by differences in imaging mechanisms through adaptive histogram matching and organ-specific gray scale standardization processing. Specifically, by calculating the cumulative distribution function block by block through a sliding window and constructing a gray scale mapping table, local histogram matching of CBCT and CT data is achieved, significantly reducing global gray scale deviation caused by differences in scanning parameters and equipment, and improving the comparability of cross-modal data; based on the pre-trained 3D U-Net, the organ ROI region is extracted, and dynamic standardization is performed on the gray mean and standard deviation of different organs, eliminating the gray scale distribution differences of the same organ in different modalities, especially improving the low contrast problem of soft tissue in CBCT, and providing a higher structural consistency data basis for subsequent feature extraction; through the above processing, CBCT and CT data have uniform gray scale statistical properties at the organ level, reducing the mismatch caused by inconsistent gray scale in subsequent deformation field calculation, especially improving the capture accuracy of nonlinear deformation (such as organ displacement and soft tissue deformation). The whole process does not require manual intervention, and the local block parameters are adjusted adaptively by the algorithm, avoiding the over-smoothing of heterogeneous regions by traditional global standardization, enhancing the adaptability of the algorithm to different anatomical sites (such as bone structure and soft tissue), and providing a reliable and repeatable solution for large-scale clinical data processing.
[0040] As a further improvement of the present application, step S3, respectively input the standardized CBCT volume data and the standardized CT volume data into the pre-trained 3D-ResNet-34-FPN architecture, and through the learning and training of the 3D-ResNet-34-FPN architecture, CBCT feature pyramids and CT feature pyramids are obtained, including:
[0041] Step S31, dividing the standardized CBCT volume data into overlapping sub-blocks with a preset overlap rate, and respectively converting each overlapping sub-block into a four-dimensional tensor;
[0042] Step S32, defining the network structure of the 3D-ResNet-34-FPN architecture based on a preset strategy, and sequentially inputting the four-dimensional tensor into the network structure to obtain a plurality of layer tensor features through multi-stage feature extraction;
[0043] Step S33, sampling the tensor features through high-level feature upsampling principle to obtain the CBCT feature pyramid;
[0044] Step S34, taking the standardized CT volume data as the executed subject, repeating steps S31 to S33 to obtain the CT feature pyramid.
[0045] Beneficial effects:
[0046] Steps S31-S34 significantly improve the consistency of CBCT and CT images in space and semantic level through multi-scale feature extraction and fusion of 3D-ResNet-34-FPN architecture. Specifically, the deep convolutional stages C3-C5 of the residual network ResNet-34 are used to extract feature maps of different resolutions, which can capture organ boundaries, large organ structures and global anatomical context, solving the problem that single-scale features cannot balance local details and global context. The top-down path of the FPN feature pyramid network is used to fuse high-level semantic features through upsampling and low-level high-resolution features, enhancing the sensitivity to small anatomical structures such as nerves and blood vessels while preserving spatial accuracy and reducing local mismatch in registration. Then, through overlapping sub-block division and four-dimensional tensor conversion, the consistency of CBCT and CT data in input structure is ensured, and combined with multi-stage feature extraction and upsampling, the generated feature pyramid strictly corresponds in spatial dimension, providing a high-precision coordinate mapping basis for subsequent deformation field calculation. Finally, the pre-trained 3D-ResNet-34-FPN architecture avoids the computational overhead of training from scratch, and its hierarchical feature design adapts to deformation modeling of different size organs, enhancing the algorithm's generalization ability for diverse anatomical structures.
[0047] As a further improvement of the present application, step S4, the sampling feature points in the CT feature pyramid are taken as the control point coordinate set, the sampling feature points in the CBCT feature pyramid are taken as the target point coordinate set, and the deformation field from the control point coordinate set to the target point coordinate set is generated through the TPS-Net architecture, including:
[0048] Step S41, uniformly sampling feature point coordinates from each layer of the CT feature pyramid to generate a control point coordinate set;
[0049] Step S42, sampling feature point coordinates of the same position from the corresponding layers of the CBCT feature pyramid and the CT feature pyramid to obtain a feature point coordinate set;
[0050] Step S43, generating a deformation field from the control point coordinate set to the target point coordinate set by minimizing the bending energy of TPS.
[0051] Advantages:
[0052] Steps S41 to S43 realize high-precision, multi-scale nonlinear registration by a thin plate spline (TPS) deformation field generation method based on a feature pyramid. Specifically, feature points are uniformly sampled from corresponding levels of the CT / CBCT feature pyramid to ensure that the control point set and the target point set cover full-scale spatial information from organ boundaries (high-resolution layers) to global anatomical context (low-resolution layers), solving the problem of insufficient capture of small structures or large-scale deformation by single-resolution features. Then, a TPS is used to minimize bending energy to generate a deformation field, and the radial basis function of the TPS effectively fits local organ displacement, soft tissue deformation and other complex deformations, avoiding the limitations of rigid or affine registration in elastic tissue matching and significantly reducing registration errors. Meanwhile, the corresponding layer sampling strategy based on the feature pyramid forces the spatial semantic consistency of the CT and CBCT feature points, ensuring that the deformation field meets the anatomical constraints and avoiding abnormal deformation or topological structure errors. Moreover, the analytical solution of the TPS solves the deformation parameters by a linear equation system, combined with the precomputed feature point coordinate set, avoiding the computational overhead caused by iterative optimization and supporting real-time registration requirements in clinical practice.
[0053] As a further improvement of the present application, step S5 converts the deformation field into a zero-centered displacement field by a 3D-CNN network, and then extracts all affine transformation parameters of the displacement field by RANSAC-SVD singular value decomposition, including:
[0054] Step S51 extracts multi-scale features of the deformation field by the convolution layer of the pre-trained 3D-CNN network;
[0055] Step S52 converts the multi-scale features into an initial displacement field with the same resolution as the deformation field by deconvolution operation;
[0056] Step S53 performs maximum value normalization on the initial displacement field to obtain a normalized displacement field;
[0057] Step S54 flattens the three-dimensional vector of the normalized displacement field into a matrix form and performs RANSAC-SVD singular value decomposition to obtain a left singular matrix and a right singular matrix;
[0058] Step S55 extracts all affine transformation parameters from the left singular matrix and the right singular matrix.
[0059] Beneficial effects:
[0060] Steps S51 to S55 realize accurate conversion from a nonlinear deformation field to high-precision affine parameters through 3D-CNN displacement field refinement and RANSAC-SVD parameter robust extraction. Specifically, the multi-scale spatial features of the deformation field are extracted through the convolution layer of the pre-trained 3D-CNN, the deconvolution operation generates a displacement field matching the resolution of the original deformation field, and then the maximum value normalization obtains a zero-centered displacement field, which effectively suppresses the abnormal displacement caused by local deformation being too large or feature matching error, and improves the smoothness and physical rationality of the displacement field; then the RANSAC algorithm is used to filter the outliers in the displacement field through random sampling consistency, and the left / right singular matrices are accurately solved by combining SVD singular value decomposition, and the 12-degree-of-freedom affine transformation parameters such as rotation, translation and scaling are extracted, which is not sensitive to noise and local deformation, and significantly improves the robustness and accuracy of parameter estimation; the generated affine parameters directly correspond to the mechanical correction instructions required by the radiotherapy equipment, and through the outlier elimination mechanism of RANSAC, the interference of soft tissue elastic deformation on the calculation of global positioning parameters is avoided, and the output result meets the physical constraints of clinical positioning error correction. The whole process does not require manual intervention, automatically converts the displacement field into a DICOM-RT standard compatible affine parameter sequence, provides a plug-and-play spatial transformation parameter for the radiotherapy equipment, reduces the subjective error and operation burden of traditional manual analysis of the deformation field, and improves the efficiency and reliability of radiotherapy positioning correction.
[0061] As a further improvement of the present application, step S6, all affine transformation parameters are packaged into a spatial transformation parameter sequence conforming to the DICOM-RT standard and sent to the radiotherapy equipment, including:
[0062] Step S61, according to the AffineTransformSequence label of the DICOM-RT standard, the translation parameters of all affine transformation parameters are mapped to the TranslationVector field, and the rotation / scaling parameters are mapped to the RotationMatrix field;
[0063] Step S62, assign a DICOM label, value representation and value length to each TranslationVector field and each RotationMatrix field, and the packaged spatial transformation parameter sequence is obtained;
[0064] Step S63, send the spatial transformation parameter sequence to the radiotherapy equipment.
[0065] Beneficial effects:
[0066] Steps S61 to S63 realize seamless conversion from algorithm output to clinical device instructions through the DICOM-RT standardized parameter packaging and transmission mechanism. According to the AffineTransformSequence label of the DICOM-RT protocol, the affine transformation parameters are accurately mapped to the standardized fields, such as the translation parameters to the TranslationVector and the rotation / scaling parameters to the RotationMatrix, to ensure that the output sequence is fully compatible with all radiotherapy devices supporting DICOM-RT (such as Varian system and Elekta system), avoiding interface conflicts caused by private protocols; and by assigning strict DICOM labels, value representations (VR) and value lengths (VL) to each field, the metadata integrity of the spatial transformation parameter sequence is guaranteed, eliminating the risk of device parsing failure caused by data format errors or transmission packet loss, while the packaging process does not require human intervention, reducing the errors introduced by subjective operations; the packaged parameter sequence is directly sent to the radiotherapy device through the network interface, and the device can be parsed and applied to mechanical correction (such as treatment bed translation and gantry rotation) in real time, realizing closed-loop automation from calculation to execution, reducing the delay of traditional manual recording and input from minutes to milliseconds, and significantly improving the real-time performance of online image-guided radiotherapy. The built-in verification mechanism and log recording function of DICOM-RT standard ensure the reliability and traceability of the parameter transmission process, providing a data basis for clinical quality control.
[0067] To achieve the above object, the application further provides the following technical solutions.
[0068] A registration device for radiotherapy position verification data, the registration device is applied to the registration method as described above, and the registration device comprises:
[0069] A dual-image body data acquisition module is configured to convert CBCT verification image data and CT image data into body data of a preset matrix size through CycleGAN modality conversion, to obtain CBCT body data and CT body data.
[0070] A dual-image standardized body data acquisition module is configured to perform local histogram matching on the CBCT body data and the CT body data respectively through an adaptive histogram matching algorithm, and perform organ-specific gray scale standardization processing on the CBCT body data and the CT body data respectively based on a pre-segmentation mask, to obtain standardized CBCT body data and standardized CT body data in a unified format.
[0071] a dual-image feature pyramid acquisition module, configured to input the standardized CBCT volume data and the standardized CT volume data into a pre-trained 3D-ResNet-34-FPN architecture respectively, and obtain a CBCT feature pyramid and a CT feature pyramid respectively through learning and training of the 3D-ResNet-34-FPN architecture;
[0072] a dual-image deformation field mapping module, configured to take the sampling feature points in the CT feature pyramid as a control point coordinate set, take the sampling feature points in the CBCT feature pyramid as a target point coordinate set, and generate a deformation field from the control point coordinate set to the target point coordinate set through a TPS-Net architecture;
[0073] an affine transformation parameter extraction module, configured to convert the deformation field into a zero-centered displacement field through a 3D-CNN network, and extract all affine transformation parameters of the displacement field through RANSAC-SVD singular value decomposition;
[0074] a spatial transformation parameter sequence packaging and sending module, configured to package all the affine transformation parameters into a spatial transformation parameter sequence conforming to a DICOM-RT standard, and send the spatial transformation parameter sequence to the radiotherapy device.
[0075] Please combine the technical features to explain the technical effects and how to solve the technical problems
[0076] To achieve the above purpose, the present application also provides the following technical solutions:
[0077] An electronic device includes a processor, and a memory coupled to the processor, the memory storing program instructions executable by the processor; the processor implements the registration method as described above when executing the program instructions stored in the memory.
[0078] To achieve the above purpose, the present application also provides the following technical solutions:
[0079] A computer-readable storage medium, the computer-readable storage medium stores program instructions, the program instructions are executed by the processor to achieve the registration method as described above.
[0080] Effective effect:
[0081] The application unifies CBCT verification image data and CT image data into preset matrix size volume data through CycleGAN modal conversion to obtain CBCT volume data and CT volume data; the CBCT volume data and the CT volume data are respectively subjected to local histogram matching through an adaptive histogram matching algorithm, and are respectively subjected to organ-specific gray standardization processing based on a pre-segmentation mask to obtain standardized CBCT volume data and standardized CT volume data in a unified format; the standardized CBCT volume data and the standardized CT volume data are respectively input into a pre-trained 3D-ResNet-34-FPN architecture, and CBCT feature pyramids and CT feature pyramids are respectively obtained through learning and training of the 3D-ResNet-34-FPN architecture; sampling feature points in the CT feature pyramids are taken as a control point coordinate set, sampling feature points in the CBCT feature pyramids are taken as a target point coordinate set, and a deformation field from the control point coordinate set to the target point coordinate set is generated through a TPS-Net architecture; the deformation field is converted into a zero-centered displacement field through a 3D-CNN network, and all affine transformation parameters of the displacement field are extracted through RANSAC-SVD singular value decomposition; all the affine transformation parameters are packaged into a spatial transformation parameter sequence conforming to a DICOM-RT standard, and are sent to a radiotherapy device. The application first reduces the modal difference between CBCT and CT images through CycleGAN modal conversion and organ-specific gray standardization, effectively laying a unified data foundation for subsequent registration; then multi-level features are extracted from the standardized data through a 3D-ResNet-34-FPN architecture, and a deformation field accurately reflecting local deformation is generated through a TPS-Net, which significantly improves the registration accuracy, especially effectively capturing nonlinear changes such as soft tissue deformation and organ displacement; finally, through 3D-CNN displacement field normalization and RANSAC-SVD decomposition, global affine transformation parameters are robustly extracted and packaged into a DICOM-RT standard protocol for direct transmission to a radiotherapy device, which not only automates the traditional manual operation process, reduces the workload and subjective errors of medical staff, but also improves the reliability and repeatability of the registration result, so that the application can provide real-time and accurate positioning error correction parameters for online image-guided radiotherapy (IGRT), ensuring that the radiation dose can be more accurately projected to the target area, thereby improving the treatment effect while protecting the surrounding normal tissue, so that the application has high calculation accuracy and high automation degree. BRIEF DESCRIPTION OF DRAWINGS
[0082] Figure 1 A step flowchart schematic diagram of one embodiment of a radiotherapy body position verification data registration method of the application;
[0083] Figure 2 A functional module schematic diagram of one embodiment of a radiotherapy body position verification data registration device of the application;
[0084] Figure 3 Structure diagram of an embodiment of an electronic device of the present application;
[0085] Figure 4 Structure diagram of an embodiment of a storage medium of the present application. DETAILED DESCRIPTION
[0086] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0087] The terms “first”, “second”, “third” in the present application are only for descriptive purpose, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with “first”, “second”, “third” can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of “plurality” is at least two, such as two, three, etc., unless otherwise explicitly and specifically limited. All directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present application are only used to explain the relative position relationship, movement condition, etc. between the components in a certain posture (as shown in the drawings), and if the certain posture changes, the directional indications also change accordingly. In addition, the terms “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.
[0088] In this document, reference to “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase “in an embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. Those skilled in the art will appreciate that embodiments described herein can be combined with other embodiments.
[0089] As Figure 1 shown, the present embodiment provides an embodiment of a registration method of radiotherapy position verification data, in which the radiotherapy position verification data is derived from CBCT verification image data of a patient to be treated on a radiotherapy device, and the patient to be treated has CT image data.
[0090] Preferably, in order to ensure the calculation accuracy and accuracy of the embodiment, the CBCT verification image data and the CT image data of the patient need to meet the following conditions at the same time: spatial resolution ≤0.5mm isotropic, image matrix at least 512*512 specification, and density resolution reaching 0.2% gray difference.
[0091] Specifically, the registration method comprises the following steps:
[0092] Step S1: unify the CBCT verification image data and the CT image data into 3D volume data with a preset matrix size through CycleGAN modal conversion, to obtain CBCT volume data and CT volume data.
[0093] Preferably, the core idea of step S1 is to realize cross-modal image conversion under non-paired data by using a cycle-consistent generative adversarial network (CycleGAN) to eliminate the style difference between CBCT and CT images, thereby laying a foundation for subsequent accurate registration.
[0094] Specifically, for the CBCT verification image data and the CT image data of the same patient, they can be regarded as two independent data sets (Domain X and Domain Y), and do not need to be strictly corresponding layer by layer, but should come from the same anatomical part (such as head and neck, chest and abdomen). Then, the original DICOM format CBCT and CT data are converted into NumPy array or tensor format convenient for processing in deep learning framework such as PyTorch and TensorFlow, so as to ensure that all image data have the same 512*512 matrix size.
[0095] Step S2: locally histogram matching the CBCT volume data and the CT volume data respectively by using an adaptive histogram matching algorithm, and performing organ-specific gray standardization processing on the CBCT volume data and the CT volume data respectively based on a pre-segmentation mask, to obtain standardized CBCT volume data and standardized CT volume data in a unified format.
[0096] Preferably, the embodiment step S2 adopts limited adaptive histogram equalization (CLAHE) as a basic algorithm, and combines LAB color space conversion for optimization, so as to achieve a better optimization effect.
[0097] The parameter settings of CLAHE are as follows:
[0098] The contrast limiting factor is 3.0, to prevent over-enhancing noise.
[0099] The grid partition size is 8*8*8, to be more fine for 3D volume data.
[0100] Histogram statistics: 0.5% and 99.5% quantiles are used to truncate outliers.
[0101] where the LAB color space conversion is performed by converting the input data from grayscale space to LAB space, and only the L channel is equalized.
[0102] Preferably, in the data preprocessing link, the input volume data needs to be subjected to 3D Gaussian filtering (σ = 1.0) to remove noise.
[0103] Step S3, respectively input the standardized CBCT volume data and the standardized CT volume data into the pre-trained 3D-ResNet-34-FPN architecture, and through the learning training of the 3D-ResNet-34-FPN architecture, respectively obtain the CBCT feature pyramid and the CT feature pyramid.
[0104] Preferably, the standardized CBCT / CT volume data (size example: 512x512x120) needs to be converted into a 4D tensor (Batchx1xDxHxW). If a multi-channel input (such as a multi-organ mask) is used, the number of channels can be expanded to N.
[0105] Preferably, the residual block of the 3D-ResNet-34-FPN architecture adopts a 3D convolution kernel, and the 3D-ResNet-34-FPN architecture is as shown in the following pseudo code:
[0106] lass BasicBlock(nn.Module):
[0107] def __init__(self, in_channels, out_channels, stride=1):
[0108] super().__init__()
[0109] self.conv1 = nn.Conv3d(in_channels, out_channels, kernel_size=3, stride=stride, padding=1)
[0110] self.bn1 = nn.BatchNorm3d(out_channels)
[0111] self.relu = nn.ReLU()
[0112] self.conv2 = nn.Conv3d(out_channels, out_channels, kernel_size=3, padding=1)
[0113] self.bn2 = nn.BatchNorm3d(out_channels)
[0114] # Downsample to adapt dimensions
[0115] self.downsample = nn.Sequential(
[0116] nn.Conv3d(in_channels, out_channels, kernel_size=1, stride=stride),
[0117] nn.BatchNorm3d(out_channels)
[0118] ) if in_channels!= out_channels else None
[0119] def forward(self, x):
[0120] identity = x
[0121] out = self.conv1(x)
[0122] out = self.bn1(out)
[0123] out = self.relu(out)
[0124] out = self.conv2(out)
[0125] out = self.bn2(out)
[0126] if self.downsample:
[0127] identity = self.downsample(x)
[0128] out += identity
[0129] out = self.relu(out)
[0130] return out
[0131] Step S4, taking the sampled feature points in the CT feature pyramid as the control point coordinate set, the sampled feature points in the CBCT feature pyramid as the target point coordinate set, generating the deformation field from the control point coordinate set to the target point coordinate set through the TPS-Net architecture.
[0132] Preferably, the spatial alignment of the two sets of coordinate sets can be ensured through the ImagePositionPatient field in the DICOM label.
[0133] Preferably, the TPS-Net architecture can be implemented through the following pseudo code:
[0134] class TPSNet(nn.Module):
[0135] def __init__(self, num_points):
[0136] super().__init__()
[0137] # Control point encoder (extract local deformation features)
[0138] self.encoder = nn.Sequential(
[0139] nn.Linear(3, 64), # input is coordinate (x, y, z)
[0140] nn.ReLU(),
[0141] nn.Linear(64, 128) )
[0143] # Affine transformation prediction branch
[0144] self.affine_branch = nn.Linear(128, 12) # output 12-dimensional affine parameters
[0145] # Non-rigid deformation branch
[0146] self.nonrigid_branch = nn.Linear(128, num_points * 3) # output N 3D displacements
[0147] def forward(self, control_pts, target_pts):
[0148] # concatenate control points and target points as input [batch, N, 6]
[0149] inputs = torch.cat([control_pts, target_pts], dim=-1)
[0150] # Feature encoding
[0151] features = self.encoder(inputs) # [batch, N, 128]
[0152] # Predict affine parameters
[0153] affine_params = self.affine_branch(features.mean(dim=1)) #[batch, 12]
[0154] # Predict non-rigid displacements
[0155] displacements = self.nonrigid_branch(features) # [batch, N,3]
[0156] return affine_params, displacements
[0157] Step S5, the deformation field is converted into a zero-centered displacement field by a 3D-CNN network, and then all affine transformation parameters of the displacement field are extracted by RANSAC-SVD singular value decomposition.
[0158] Preferably, since the deformation field generated by the TPS-Net is four-dimensional (DxHxWx3) and stores the displacement vector (dx, dy, dz) of each voxel, it is necessary to convert the deformation field into a 5D tensor [Batch, 3, Depth, Height, Width] to adapt to the input requirements of the 3D-CNN.
[0159] Preferably, the Python code of the 3D-CNN architecture is as follows:
[0160] import torch
[0161] import torch.nn as nn
[0162] import torch.nn.functional as F
[0163] class DeformationToDisplacement(nn.Module):
[0164] def __init__(self, in_channels=3, out_channels=3):
[0165] super().__init__()
[0166] # Encoder (Feature Extraction)
[0167] self.encoder = nn.Sequential(
[0168] nn.Conv3d(in_channels, 64, kernel_size=3, padding=1),
[0169] nn.ReLU(),
[0170] nn.MaxPool3d(2),
[0171] nn.Conv3d(64, 128, kernel_size=3, padding=1),
[0172] nn.ReLU(),
[0173] nn.MaxPool3d(2),
[0174] nn.Conv3d(128, 256, kernel_size=3, padding=1),
[0175] nn.ReLU() )
[0177] # Decoder (Displacement Field Generation)
[0178] self.decoder = nn.Sequential(
[0179] nn.ConvTranspose3d(256, 128, kernel_size=3, stride=2,padding=1, output_padding=1),
[0180] nn.ReLU(),
[0181] nn.ConvTranspose3d(128, 64, kernel_size=3, stride=2,padding=1, output_padding=1),
[0182] nn.ReLU(),
[0183] nn.Conv3d(64, out_channels, kernel_size=3, padding=1), # Outputs a 3-channel displacement field
[0184] nn.Tanh() # Constrain to [-1, 1] )
[0186] def forward(self, x):
[0187] x = self.encoder(x)
[0188] x = self.decoder(x)
[0189] return x
[0190] Step S6: Encapsulate all affine transformation parameters into a spatial transformation parameter sequence conforming to the DICOM-RT standard and send it to the radiotherapy equipment.
[0191] Further, step S1, unifying the CBCT verification image data and CT image data into volume data of a preset matrix size through CycleGAN modal transformation, to obtain CBCT volume data and CT volume data, specifically includes the following steps:
[0192] Step S11: Convert the CBCT verification image data and CT image data into the DICOM standard format.
[0193] Preferably, the two sets of image data need to be normalized first. This is done by normalizing the pixel values to a specific range, such as [-1, 1], to accelerate training convergence, as follows:
[0194] I normalized =(II min ) / (I max -I min )*2-1, where I is the original image data, I min and I max These are the minimum and maximum pixel values of the image data, calculated separately for CBCT verification image data and CT image data.
[0195] Step S12: Align the spatial coordinate system of the CBCT verification image data with the spatial coordinate system of the CT image data using a non-rigid registration method.
[0196] Step S13, define two generators of CycleGAN modal conversion based on 3D U-Net architecture, and define two discriminators of CycleGAN modal conversion based on 3D PatchGAN.
[0197] Preferably, CycleGAN modal conversion requires two generators (G and F) and two discriminators (D X and D Y ).
[0198] Wherein, the generator G and the generator F generally adopt U-Net structure, the generator G is responsible for converting the CBCT image (Domain X) into CT style image (Domain Y), and the generator F performs reverse conversion CT image→CBCT style image. The intention of adopting the encoding-decoding structure of U-Net in this embodiment is to effectively preserve the spatial information of the two sets of image data.
[0199] Wherein, the discriminator D X and the discriminator D Y generally use PatchGAN (or similar structure of CNN). The discriminator D Y is used to distinguish between real CT images and false CT images generated by the generator G; and the discriminator D X is used to distinguish between real CBCT images and false CBCT images generated by the generator F.
[0200] Step S14, convert the CBCT verification image data into CT style data through one of the generators, and convert the CT image data into CBCT style data through the other generator.
[0201] Step S15, calculate the adversarial loss based on the Wasserstein distance between the CBCT style data and the CBCT verification image data through one of the discriminators, and calculate the adversarial loss based on the Wasserstein distance between the CT style data and the CT image data through the other discriminator.
[0202] Step S16, define the L1 loss of the CycleGAN modal conversion, and simultaneously iterate the two adversarial losses through back propagation until the L1 loss reaches the minimum value.
[0203] Preferably, the training target of the L1 loss is to optimize the generator and the discriminator at the same time, which can be realized by training the adversarial loss and the cycle consistency loss. Wherein, the adversarial loss is distinguished between real CT and false CT images generated by the generator G through the discriminator D X , and the discriminator D YThe true CBCT image and the false CBCT image generated by the generator F are distinguished; the cycle consistency loss is L1 (recov CBCT, original CBCT) + L1 (recov CT, original CT); the back propagation is realized by jointly optimizing the parameters of the generator and the discriminator, the learning rate is initially set to 0.0002, the Adam optimizer is used, and it is expected to reach convergence through 100-200 epochs.
[0204] In step S17, when the L1 loss reaches the minimum value, the CT style data of the last iteration is obtained through one of the generators, and the CT style data and the CT image data are the volume data with a unified preset matrix size.
[0205] Preferably, the input CBCT and CT volume data are converted into.nii or.hdr format of a unified modality through the generators F and G respectively, and it is ensured that the matrix specifications of the output volume data are consistent with the preset specifications, and the voxel values need to be uniformly mapped to the range of [0, 1] or [-1, 1].
[0206] The steps S11 to S17 of the embodiment effectively solve the core problems of large modality difference, inconsistent gray distribution and low soft tissue contrast between CBCT and CT images caused by different imaging principles through the CycleGAN modality conversion technology. Specifically, the CBCT data is converted into CT style data through non-rigid registration of the spatial coordinate system and the generative adversarial network, which significantly reduces the noise, artifacts and resolution difference between the two, improves the data consistency, and uses the adversarial loss of the Wasserstein distance and the L1 cycle consistency loss to ensure that the generated image not only retains the anatomical structure authenticity, but also has the high contrast characteristic of the CT image, reducing the registration error caused by the unsatisfactory CBCT image quality. The whole process does not need manual intervention or strictly paired training data, and the end-to-end deep learning framework automatically outputs the volume data with a unified matrix size, provides a standardized input for subsequent registration, reduces the subjective error and work burden of traditional manual operation.
[0207] Further, in step S2, the CBCT volume data and the CT volume data are respectively subjected to local histogram matching through an adaptive histogram matching algorithm, and the CBCT volume data and the CT volume data are respectively subjected to organ-specific gray standardization processing based on a pre-segmentation mask, to obtain standardized CBCT volume data and standardized CT volume data in a unified format, which specifically includes the following steps:
[0208] In step S21, a sliding window with a preset voxel size is defined, and the standardized CBCT volume data is traversed through the sliding window at a preset sliding speed.
[0209] Preferably, the division strategy of step S21 is to divide the image into overlapping local blocks (e.g. 16x16x16 voxels) by sliding window, and traverse the whole volume data by sliding window.
[0210] Step S22, calculate the gray level histogram of the current block once per sliding window, and obtain the pixel distribution probability of the current block.
[0211] Preferably, the gray level histogram needs to be calculated independently for each block, and the gray level range is [0, 255].
[0212] Step S23, calculate the cumulative distribution function of each block by formula (1) respectively:
[0213] (1).
[0214] Wherein, is the cumulative distribution function, is the upper limit of the gray level, is the th gray level, , is the probability distribution of the th gray level.
[0215] Preferably, is equal to the number of pixels of the th gray level divided by the total number of pixels of the current block.
[0216] Step S24, taking the CT volume data as the traversed object, repeat steps S21 to S23 to obtain the cumulative distribution function of the CT volume data.
[0217] Step S25, based on the two cumulative distribution functions of the same block, construct the gray level mapping table according to formula (2):
[0218] (2);
[0219] Wherein, is the mapping function of the th gray level of the CBCT volume data to all gray levels of the CT volume data, is the cumulative distribution function of the th gray level of the CBCT volume data, is the cumulative distribution function of the th gray level of the CT volume data, is the th gray level of the CT volume data.
[0220] Preferably, the mapping table The bilinear interpolation smoothing is adopted for the overlapping area to avoid the gray level jump between blocks. The particle swarm optimization (PSO) algorithm can be introduced to adaptively adjust the division granularity of the local block. The particle position represents the block size, and the fitness function is the KL divergence of the global histogram. The particle position is updated iteratively to find the optimal block division scheme, such as optimizing from 16x16 to 12x12, 8x8, etc.
[0221] In step S26, the 3D U-Net architecture pre-trained by the preset organ data set is used to extract at least one organ ROI region of the CT volume data and the CBCT volume data, respectively.
[0222] Preferably, the organ ROI region is an organ ROI region of interest.
[0223] For example, for CBCT data, xLSTM-U-Net is used to segment ribs and soft tissue; for CT data, residual U-Net is used to segment the lung.
[0224] In step S27, the gray mean and standard deviation of each organ ROI region are calculated, and the organ-specific gray normalization processing is performed by formula (3) to obtain the normalized CBCT volume data and the normalized CT volume data in a unified format, respectively:
[0225] (3);
[0226] wherein, is the normalized gray value, is the voxel coordinate value, is the coordinate of the voxel, is the voxel gray standard deviation in the organ ROI region, is the voxel gray mean in the organ ROI region, is the preset target organ gray standard deviation, is the preset target organ gray mean.
[0227] Preferably, and are both preset statistical values of the target organ, which can be obtained by statistics.
[0228] Preferably, the gray scale characteristics of different organs can be normalized based on the pre-segmentation mask to eliminate the internal gray scale inconsistency of the organs. The pre-segmentation mask can be set by using a supervised segmentation model. Specifically, the generator structure of CycleGAN in step S1 is combined, the pre-trained weights of the MS-COCO dataset are loaded, the organ bounding box is detected by the YOLOv8 (such as predict (source = CT_data, conf = 0.25)), and then the bounding box is input into the SAM model to generate a high-precision mask (masks = predictor.predict (box = bbox)).
[0229] The steps S21 to S27 of the embodiment effectively solve the problems of inconsistent gray scale distribution, low organ contrast and modal statistical property deviation between CBCT and CT images caused by differences in imaging mechanisms. Specifically, the cumulative distribution function is calculated and the gray scale mapping table is constructed by using a sliding window to calculate the cumulative distribution function block by block, realizing the local histogram matching of CBCT and CT data, significantly reducing the global gray scale deviation caused by differences in scanning parameters and equipment, and improving the comparability of cross-modal data. Based on the pre-trained 3D U-Net, the organ ROI region is extracted, and the gray mean and standard deviation of different organs are dynamically standardized, eliminating the gray scale distribution difference of the same organ in different modalities, especially improving the low contrast problem of soft tissue in CBCT, and providing a data basis with higher structural consistency for subsequent feature extraction. Through the above processing, the CBCT and CT data have unified gray scale statistical properties at the organ level, reducing the mismatch caused by inconsistent gray scale in subsequent deformation field calculation, especially improving the capture accuracy of nonlinear deformation (such as organ displacement and soft tissue deformation). The whole process does not require manual intervention, and the local block parameters are adjusted adaptively by the algorithm, avoiding the over-smoothing of heterogeneous regions by traditional global standardization, enhancing the adaptability of the algorithm to different anatomical parts (such as bone structure and soft tissue), and providing a reliable and repeatable solution for large-scale clinical data processing.
[0230] Further, in step S3, the standardized CBCT volume data and the standardized CT volume data are respectively input into a pre-trained 3D-ResNet-34-FPN architecture, and through the learning and training of the 3D-ResNet-34-FPN architecture, CBCT feature pyramids and CT feature pyramids are respectively obtained, including the following steps:
[0231] In step S31, the standardized CBCT volume data is divided into overlapping sub-blocks with a preset overlap rate, and each overlapping sub-block is converted into a four-dimensional tensor.
[0232] Preferably, the four-dimensional tensor is the 4D tensor in the above (Batch x 1 x D x H x W).
[0233] Step S32, define the network structure of the 3D-ResNet-34-FPN architecture based on the preset strategy, and input the four-dimensional tensor into the network structure in sequence, and obtain a plurality of layer tensor features through multi-stage feature extraction.
[0234] Preferably, the four convolution stages of ResNet-34 output feature maps C1-C5, and the actual FPN usually selects C3, C4 and C5, corresponding to 1 / 8, 1 / 16 and 1 / 32 downsampling, and the specific style is shown in Table 1:
[0235] Table 1: Tensor feature C3-C5 feature map.
[0236]
[0237] It is worth noting that C1-C2 are usually discarded by FPN to reduce the amount of calculation because of too high resolution and weak semantics.
[0238] Step S33, sample the tensor feature through the high-level feature upsampling principle to obtain the CBCT feature pyramid.
[0239] Preferably, the top-down path can be selected from the highest layer C5, and gradually upsampled to the adjacent low layer resolution through 3D deconvolution (such as bilinear interpolation), and fused with the layer feature, and the pseudo code is as follows:
[0240] Pseudo code: FPN top-down fusion
[0241] P5 = nn.Conv3d(C5_channels, feature_size, kernel_size=1) # 1x1x1 convolution to adjust channel
[0242] P4 = nn.Conv3d(C4_channels, feature_size, kernel_size=1) + F.upsample(P5, scale_factor=2)
[0243] P3 = nn.Conv3d(C3_channels, feature_size, kernel_size=1) + F.upsample(P4, scale_factor=2)
[0244] Preferably, in order to cover a larger scale range, P6 can be additionally generated by performing step 2 convolution on C5:
[0245] P6 = nn.Conv3d(C5_channels, feature_size, kernel_size=3, stride=2,padding=1) # downsample to 1 / 64
[0246] Preferably, the final output pyramid levels: P3 (1 / 8), P4 (1 / 16), P5 (1 / 32), P6 (1 / 64) are used to adapt to the deformation modeling of different size organs, for example, P3 for small organs (optic nerve), P6 for large range chest.
[0247] In step S34, the CT feature pyramid is obtained by repeating steps S31 to S33 with the standardized CT volume data as the execution subject.
[0248] The steps S31 to S34 of the embodiment significantly improve the consistency of CBCT and CT images in space and semantic level through multi-scale feature extraction and fusion of the 3D-ResNet-34-FPN architecture. Specifically, the deep convolutional stages C3-C5 of the residual network ResNet-34 are used to extract feature maps of different resolutions, which capture organ boundaries, large organ structures and global anatomical context, respectively, solving the problem that single-scale features cannot balance local details and global context. The top-down path of the FPN feature pyramid network is used to fuse high-level semantic features through upsampling with low-level high-resolution features, enhancing the sensitivity to small anatomical structures such as nerves and blood vessels, while preserving spatial precision and reducing local mismatch in registration. Then, through overlapping sub-block division and four-dimensional tensor conversion, the consistency of CBCT and CT data in input structure is ensured, and combined with multi-stage feature extraction and upsampling, the generated feature pyramid strictly corresponds in spatial dimension, providing a high-precision coordinate mapping basis for subsequent deformation field calculation. Finally, the pre-trained 3D-ResNet-34-FPN architecture avoids the computational overhead of training from scratch, and its hierarchical feature design adapts to the deformation modeling of different size organs, enhancing the generalization ability of the algorithm to diverse anatomical structures.
[0249] Further, in step S4, the sampling feature points in the CT feature pyramid are taken as the control point coordinate set, the sampling feature points in the CBCT feature pyramid are taken as the target point coordinate set, and the deformation field from the control point coordinate set to the target point coordinate set is generated through the TPS-Net architecture, which includes the following steps:
[0250] In step S41, the feature point coordinates are uniformly sampled from each layer of the CT feature pyramid to generate the control point coordinate set.
[0251] In step S42, the feature point coordinates of the same position are sampled from the corresponding layers of the CBCT feature pyramid and the CT feature pyramid to obtain the feature point coordinate set.
[0252] Step S43: Minimize the bending energy to generate the deformation field from the control point coordinate set to the target point coordinate set using TPS.
[0253] Preferably, TPS achieves deformation field smoothing by minimizing bending energy, and the deformation function is:
[0254] .
[0255] in, For deformation function, Let be the spatial coordinate vector of the point to be interpolated. It is a global translation vector. This is the affine transformation matrix, typically with 12 degrees of freedom. For the first The non-rigid deformation weights of each control point are given, and the sum of all non-rigid deformation weights is 0. It is a radial basis function, and .
[0256] Preferably, since the solution process is complex and there are already mature existing technologies for solving this function, this embodiment provides the Python code involved in the solution, and no longer provides a detailed solution process:
[0257] import numpy as np
[0258] def tps_coefficients(source_pts, target_pts):
[0259] n = len(source_pts)
[0260] K = np.zeros((n, n))
[0261] for i in range(n):
[0262] for j in range(n):
[0263] r = np.linalg.norm(source_pts[i] - source_pts[j])
[0264] if r > 0:
[0265] K[i, j] = r**2 * np.log(r**2) # U(r) = r² ln(r²)
[0266] P = np.hstack([np.ones((n, 1)), source_pts]) # [1, x, y]
[0267] L_top = np.hstack([K, P])
[0268] L_bottom = np.hstack([PT, np.zeros((3, 3))])
[0269] L = np.vstack([L_top, L_bottom])
[0270] V_x = np.hstack([target_pts[:, 0], np.zeros(3)]) # [x1',...,x n ',0,0,0]
[0271] V_y = np.hstack([target_pts[:, 1], np.zeros(3)]) # [y1',...,y n ',0,0,0]
[0272] # Solve the system of linear equations
[0273] W_a_x = np.linalg.solve(L, V_x) # [w1...w n , a1, a x [ , a_y] for X
[0274] W_a_y = np.linalg.solve(L, V_y) # Same as above for Y
[0275] return W_a_x, W_a_y
[0276] source_pts = np.array([[0,0], [1,0], [0,1]]) # Control points (sources)
[0277] target_pts = np.array([[0,0], [1.1,0], [0,1.1]]) # Control points (targets)
[0278] coef_x, coef_y = tps_coefficients(source_pts, target_pts)
[0279] The steps S41 to S43 of the embodiment are implemented by a thin plate spline (TPS) deformation field generation method based on a feature pyramid, to achieve high-precision and multi-scale nonlinear registration. Specifically, feature points are uniformly sampled from corresponding levels of the CT / CBCT feature pyramid, to ensure that the control point set and the target point set cover full-scale spatial information from the organ boundary (high-resolution layer) to the global anatomical context (low-resolution layer), to solve the problem of insufficient capture of small structures or large-scale deformation by single-resolution features; then, a TPS is used to minimize the bending energy to generate a deformation field, and the radial basis function of the TPS effectively fits complex deformations such as local organ displacement and soft tissue deformation, to avoid the limitations of rigid or affine registration in elastic tissue matching and significantly reduce registration errors; meanwhile, the corresponding layer sampling strategy based on the feature pyramid forces the spatial semantic consistency of the CT and CBCT feature points, to ensure that the deformation field meets the anatomical constraints and avoids abnormal deformation or topological structure errors; and the analytical solution of the TPS solves the deformation parameters by a linear equation system, in combination with the precomputed feature point coordinate set, to avoid the computational overhead caused by iterative optimization and support the real-time registration requirements in clinical practice.
[0280] Further, in step S5, the deformation field is converted into a zero-centered displacement field by a 3D-CNN network, and then all affine transformation parameters of the displacement field are extracted by RANSAC-SVD singular value decomposition, including the following steps:
[0281] In step S51, multi-scale features of the deformation field are extracted by a convolution layer of a pre-trained 3D-CNN network.
[0282] Preferably, the multi-scale features in step S51 of the embodiment are the 5D tensor described above.
[0283] Preferably, sudden displacement caused by feature point matching errors can be filtered out by a maximum organ displacement > 20 mm.
[0284] In step S52, an initial displacement field with the same resolution as the deformation field is converted from the multi-scale features by a deconvolution operation.
[0285] Preferably, the Python code of the deconvolution operation is as follows:
[0286] import torch
[0287] import torch.nn as nn
[0288] import torch.nn.functional as F
[0289] class MedicalDeconv3D(nn.Module):
[0290] def __init__(self, in_channels=3, out_channels=3):
[0291] super().__init__()
[0292] # Encoder (down-sampling)
[0293] self.encoder = nn.Sequential(
[0294] nn.Conv3d(in_channels, 64, kernel_size=3, padding=1),
[0295] nn.InstanceNorm3d(64),
[0296] nn.LeakyReLU(0.2),
[0297] nn.MaxPool3d(2),
[0298] nn.Conv3d(64, 128, kernel_size=3, padding=1),
[0299] nn.InstanceNorm3d(128),
[0300] nn.LeakyReLU(0.2),
[0301] nn.MaxPool3d(2) )
[0303] # Decoder (deconvolution up-sampling)
[0304] self.decoder = nn.Sequential(
[0305] nn.ConvTranspose3d(128, 64, kernel_size=4, stride=2,padding=1),
[0306] nn.InstanceNorm3d(64),
[0307] nn.LeakyReLU(0.2),
[0308] nn.ConvTranspose3d(64, 32, kernel_size=4, stride=2,padding=1),
[0309] nn.InstanceNorm3d(32),
[0310] nn.LeakyReLU(0.2),
[0311] nn.Conv3d(32, out_channels, kernel_size=3, padding=1),
[0312] nn.Tanh() # Outputs the normalized displacement field )
[0314] def forward(self, x):
[0315] x = self.encoder(x)
[0316] x = self.decoder(x)
[0317] return x
[0318] Preferably, the 3D transposed convolutional layer uses nn.ConvTranspose3d to achieve 3D upsampling, and the output resolution is guaranteed by the combination of stride=2 and kernel_size=4.
[0319] Step S53: Normalize the initial displacement field to obtain the normalized displacement field.
[0320] Preferably, the Python code for maximum value normalization is as follows:
[0321] def normalize_displacement(field, organ_mask, max_displacement=10.0):
[0322] # Calculate the maximum displacement within the mask
[0323] max_val = torch.max(field * organ_mask, dim=[1,2,3])[0]
[0324] # Dynamic Normalization
[0325] normalized = field / (max_val + 1e-6) * 2 - 1 # Map to [-1,1]
[0326] return normalized
[0327] Step S54, flatten the three-dimensional vector of the normalized displacement field into a matrix form and perform RANSAC-SVD singular value decomposition to obtain the left and right singular matrices.
[0328] Preferably, the Python code for matrixing the displacement field is as follows:
[0329] def prepare_for_svd(displacement_field):
[0330] # Flatten the 3D displacement field into a 2D matrix
[0331] x, y, z = displacement_field.shape
[0332] return displacement_field.view(x * y * z, 3)
[0333] Preferably, 3 non-collinear points are randomly selected from the flattened matrix, and the centroids of the two groups of points are calculated. The covariance matrix is constructed according to the centroids, and SVD decomposition is performed on the covariance matrix. The specific Python code is as follows:
[0334] def ransac_svd(displacement_matrix, max_iters=1000, threshold=0.1):
[0335] """
[0336] RANSAC-SVD algorithm implementation
[0337] Input: displacement_matrix (N×3 matrix)
[0338] Output: optimal rotation matrix R, translation vector t, inlier index, left and right singular matrices U and Vt
[0339] """
[0340] best_R = None
[0341] best_t = None
[0342] best_inliers = []
[0343] best_U = None
[0344] best_Vt = None
[0345] for _ in range(max_iters):
[0346] # 1. Randomly sample 3 points
[0347] sample_indices = np.random.choice(len(displacement_matrix), 3, replace=False)
[0348] sample = displacement_matrix[sample_indices]
[0349] # 2. Compute centroid and centerize
[0350] mu_p = np.mean(sample, axis=0)
[0351] centered_p = sample - mu_p
[0352] # 3. SVD decomposition
[0353] H = centered_p.T @ centered_p # Covariance matrix
[0354] U, S, Vt = np.linalg.svd(H)
[0355] # 4. Compute rotation matrix (handle reflection case)
[0356] R = Vt.T @ U.T
[0357] if np.linalg.det(R) < 0:
[0358] Vt[-1, :] *= -1
[0359] R = Vt.T @ U.T
[0360] # 5. Inlier detection
[0361] errors = np.linalg.norm(displacement_matrix - (displacement_matrix @ R.T), axis=1)
[0362] inliers = np.where(errors < threshold)[0]
[0363] # 6. Update best model
[0364] if len(inliers) > len(best_inliers):
[0365] best_inliers = inliers
[0366] best_R = R
[0367] best_U = U
[0368] best_Vt = Vt
[0369] # Recompute SVD using all inliers
[0370] if len(best_inliers) > 0:
[0371] inlier_points = displacement_matrix[best_inliers]
[0372] mu_p = np.mean(inlier_points, axis=0)
[0373] centered_p = inlier_points - mu_p
[0374] H = centered_p.T @ centered_p
[0375] best_U, _, best_Vt = np.linalg.svd(H)
[0376] return best_R, best_U, best_Vt, best_inliers
[0377] From the above, it can be seen that best_U is the left singular matrix and best_Vt is the right singular matrix.
[0378] Step S55 extracts all affine transformation parameters from the left singular matrix and the right singular matrix.
[0379] Preferably, the Python code for extracting all affine transformation parameters is as follows:
[0380] def extract_affine_params(displacement_field):
[0381] # Generate voxel grid coordinates (x,y,z)
[0382] d, h, w = displacement_field.shape[2:]
[0383] coords = np.mgrid[:d, :h, :w].transpose(1,2,3,0).reshape(-1, 3)
[0384] # Target coordinates = Source coordinates + displacement
[0385] src_coords = coords
[0386] dst_coords = src_coords + displacement_field.cpu().numpy().reshape(-1, 3)
[0387] # Centering
[0388] src_centroid = np.mean(src_coords, axis=0)
[0389] dst_centroid = np.mean(dst_coords, axis=0)
[0390] src_centered = src_coords - src_centroid
[0391] dst_centered = dst_coords - dst_centroid
[0392] # SVD decompose H matrix
[0393] H = src_centered.T @ dst_centered
[0394] U, S, VT = np.linalg.svd(H)
[0395] # Compute affine matrix A and translation b
[0396] A = VT.T @ U.T
[0397] if np.linalg.det(A) < 0: # Ensure rotation is non-reflection
[0398] VT[-1, :] *= -1
[0399] A = VT.T @ U.T
[0400] b = dst_centroid - A @ src_centroid
[0401] return A, b # Returns the affine matrix and translation vector
[0402] In this embodiment, steps S51 to S55 achieve precise conversion from nonlinear deformation fields to high-precision affine parameters through 3D-CNN displacement field refinement and RANSAC-SVD parameter robust extraction. Specifically, multi-scale spatial features of the deformation field are extracted through the convolutional layers of a pre-trained 3D-CNN. Deconvolution operations generate a displacement field with a resolution matching the original deformation field, and then the maximum value is normalized to obtain a zero-centered displacement field. This effectively suppresses abnormal displacements caused by excessive local deformation or feature matching errors, improving the smoothness and physical rationality of the displacement field. Then, the RANSAC algorithm is used to randomly sample and filter outout points in the displacement field, and the left / right singular matrices are accurately solved by combining SVD singular value decomposition. From these matrices, 12-DOF affine transformation parameters such as rotation, translation, and scaling are extracted. This process is insensitive to noise and local deformation, significantly improving the robustness and accuracy of parameter estimation. The generated affine parameters directly correspond to the mechanical correction instructions required by the radiotherapy equipment, and the outlier removal mechanism of RANSAC avoids the interference of soft tissue elastic deformation on the calculation of global positioning parameters, ensuring that the output results meet the physical constraints of clinical positioning error correction. The entire process requires no manual intervention and automatically converts the displacement field into a DICOM-RT standard-compatible affine parameter sequence, providing plug-and-play spatial transformation parameters for radiotherapy equipment. This reduces the subjective errors and operational burden of traditional manual analysis of deformation fields, and improves the efficiency and reliability of radiotherapy positioning correction.
[0403] Further, step S6 encapsulates all affine transformation parameters into a spatial transformation parameter sequence conforming to the DICOM-RT standard and sends it to the radiotherapy equipment, specifically including the following steps:
[0404] Step S61: Based on the AffineTransformSequence tag of the DICOM-RT standard, map the translation parameters of all affine transformation parameters to the TranslationVector field and the rotation / scaling parameters to the RotationMatrix field.
[0405] Preferably, the AffineTransformSequence tag (e.g., (3006,00A0)) encapsulates the parameters, the translation parameters are mapped to the TranslationVector field (e.g., (3006,00A2), and the rotation / scaling parameters are mapped to the RotationMatrix field (e.g., (3006,00A4)).
[0406] Step S62, assign DICOM tag, value representation, value length to each TranslationVector field and each RotationMatrix field, and encapsulate to get the spatial transformation parameter sequence.
[0407] Preferably, the spatial transformation parameter sequence comprises DICOM tag (Tag), value representation (VR) and value length (VL).
[0408] Step S63, send the spatial transformation parameter sequence to the radiotherapy device.
[0409] The steps S61-S63 of the embodiment realize seamless conversion from algorithm output to clinical device instruction through the parameter encapsulation and transmission mechanism standardized by DICOM-RT. Specifically, according to the AffineTransformSequence tag of the DICOM-RT protocol, the affine transformation parameters are accurately mapped to the standardized fields, such as the translation parameters to TranslationVector and the rotation / scaling parameters to RotationMatrix, to ensure that the output sequence is fully compatible with all radiotherapy devices supporting DICOM-RT (such as Varian system and Elekta system), avoiding interface conflicts caused by private protocols; and by assigning strict DICOM tag, value representation (VR) and value length (VL) to each field, the metadata integrity of the spatial transformation parameter sequence is guaranteed, eliminating the risk of device parsing failure caused by data format error or transmission packet loss, while the encapsulation process does not require manual intervention, reducing the errors introduced by subjective operation; the encapsulated parameter sequence is directly sent to the radiotherapy device through the network interface, and the device can be parsed and applied to mechanical correction (such as treatment bed translation and gantry rotation) in real time, realizing closed-loop automation from calculation to execution, reducing the delay of traditional manual recording and input from minutes to milliseconds, and significantly improving the real-time performance of online image-guided radiotherapy. The built-in verification mechanism and log recording function of DICOM-RT standard ensure the reliability and traceability of the parameter transmission process, providing a data basis for clinical quality control.
[0410] The embodiment unifies CBCT verification image data and CT image data into preset matrix size volume data through CycleGAN modal conversion, obtains CBCT volume data and CT volume data; performs local histogram matching on the CBCT volume data and the CT volume data respectively through an adaptive histogram matching algorithm, and performs organ-specific gray scale normalization processing on the CBCT volume data and the CT volume data respectively based on a pre-segmentation mask, to obtain standardized CBCT volume data and standardized CT volume data in a unified format; the standardized CBCT volume data and the standardized CT volume data are respectively input into a pre-trained 3D-ResNet-34-FPN architecture, and through learning and training of the 3D-ResNet-34-FPN architecture, CBCT feature pyramids and CT feature pyramids are respectively obtained; sampling feature points in the CT feature pyramids are taken as a control point coordinate set, sampling feature points in the CBCT feature pyramids are taken as a target point coordinate set, and a deformation field from the control point coordinate set to the target point coordinate set is generated through a TPS-Net architecture; the deformation field is converted into a zero-centered displacement field through a 3D-CNN network, and all affine transformation parameters of the displacement field are extracted through RANSAC-SVD singular value decomposition; all the affine transformation parameters are packaged into a spatial transformation parameter sequence conforming to the DICOM-RT standard, and are sent to a radiotherapy device. The embodiment first reduces the modal difference between CBCT and CT images through CycleGAN modal conversion and organ-specific gray scale normalization, effectively laying a unified data foundation for subsequent registration; then multi-level features are extracted from the standardized data using the 3D-ResNet-34-FPN architecture, and a deformation field accurately reflecting local deformation is generated in combination with the TPS-Net, which significantly improves the registration accuracy, especially effectively capturing nonlinear changes such as soft tissue deformation and organ displacement; finally, through 3D-CNN displacement field normalization and RANSAC-SVD decomposition, global affine transformation parameters are robustly extracted and packaged into a DICOM-RT standard protocol for direct transmission to a radiotherapy device, which not only automates the traditional manual operation process, reduces the workload and subjective errors of medical staff, but also improves the reliability and repeatability of the registration result, so that the embodiment can provide real-time and accurate positioning error correction parameters for online image-guided radiotherapy (IGRT), ensuring that the radiation dose can be more accurately projected to the target area, thereby improving the treatment effect while protecting the surrounding normal tissue, so that the embodiment has high calculation accuracy and high automation degree.
[0411] As shown in Figure 2 , the present example provides an embodiment of a registration device for radiotherapy position verification data, which is applied to the registration method of the above-mentioned embodiments in the present embodiment.
[0412] Specifically, the registration device comprises, in sequence and electrically or in signal connection, a dual-image volume data acquisition module 1, a dual-image standardized volume data acquisition module 2, a dual-image feature pyramid acquisition module 3, a dual-image deformation field mapping module 4, an affine transformation parameter extraction module 5, and a spatial transformation parameter sequence packaging and sending module 6.
[0413] The dual-image volume data acquisition module 1 is configured to unify the CBCT verification image data and the CT image data into volume data of a preset matrix size through CycleGAN modal conversion, to obtain CBCT volume data and CT volume data; the dual-image standardized volume data acquisition module 2 is configured to perform local histogram matching on the CBCT volume data and the CT volume data respectively through an adaptive histogram matching algorithm, and to perform organ-specific gray scale standardization processing on the CBCT volume data and the CT volume data respectively based on a pre-segmentation mask, to obtain standardized CBCT volume data and standardized CT volume data in a unified format; the dual-image feature pyramid acquisition module 3 is configured to input the standardized CBCT volume data and the standardized CT volume data respectively into a pre-trained 3D-ResNet-34-FPN architecture, to obtain a CBCT feature pyramid and a CT feature pyramid respectively through learning and training of the 3D-ResNet-34-FPN architecture; the dual-image deformation field mapping module 4 is configured to take the sampling feature points in the CT feature pyramid as a control point coordinate set, take the sampling feature points in the CBCT feature pyramid as a target point coordinate set, and generate a deformation field from the control point coordinate set to the target point coordinate set through a TPS-Net architecture; the affine transformation parameter extraction module 5 is configured to convert the deformation field into a zero-centered displacement field through a 3D-CNN network, and to extract all affine transformation parameters of the displacement field through RANSAC-SVD singular value decomposition; and the spatial transformation parameter sequence packaging and sending module 6 is configured to package all the affine transformation parameters into a spatial transformation parameter sequence conforming to the DICOM-RT standard, and to send the spatial transformation parameter sequence to a radiotherapy device.
[0414] Further, the dual-image volume data acquisition module 1 specifically comprises, in sequence and electrically or in signal connection, a first dual-image volume data acquisition unit, a second dual-image volume data acquisition unit, a third dual-image volume data acquisition unit, a fourth dual-image volume data acquisition unit, a fifth dual-image volume data acquisition unit, a sixth dual-image volume data acquisition unit, and a seventh dual-image volume data acquisition unit; the seventh dual-image volume data acquisition unit is electrically or in signal connection with the dual-image standardized volume data acquisition module 2.
[0415] The first dual-image body data acquisition unit is configured to convert the CBCT verification image data and the CT image data into a DICOM standard format; the second dual-image body data acquisition unit is configured to align a spatial coordinate system of the CBCT verification image data with a spatial coordinate system of the CT image data by using a non-rigid registration method; the third dual-image body data acquisition unit is configured to define two generators of the CycleGAN modal conversion based on a 3D U-Net architecture, and to define two discriminators of the CycleGAN modal conversion based on a 3D PatchGAN; the fourth dual-image body data acquisition unit is configured to convert the CBCT verification image data into CT style data by using one of the generators, and to convert the CT image data into CBCT style data by using the other generator; the fifth dual-image body data acquisition unit is configured to calculate an adversarial loss based on a Wasserstein distance between the CBCT style data and the CBCT verification image data by using one of the discriminators, and to calculate an adversarial loss based on a Wasserstein distance between the CT style data and the CT image data by using the other discriminator; the sixth dual-image body data acquisition unit is configured to define an L1 loss of the CycleGAN modal conversion, and to simultaneously iterate the two adversarial losses by using back propagation until the L1 loss reaches a minimum value; and the seventh dual-image body data acquisition unit is configured to obtain the CT style data at the last iteration by using one of the generators when the L1 loss reaches the minimum value, and the CT style data and the CT image data are body data with a unified preset matrix size.
[0416] Further, the dual-image standardized body data acquisition module 2 specifically comprises, which are sequentially electrically connected or signal connected, a first dual-image standardized body data acquisition unit, a second dual-image standardized body data acquisition unit, a third dual-image standardized body data acquisition unit, a fourth dual-image standardized body data acquisition unit, a fifth dual-image standardized body data acquisition unit, a sixth dual-image standardized body data acquisition unit, and a seventh dual-image standardized body data acquisition unit; the first dual-image standardized body data acquisition unit is electrically connected or signal connected with the seventh dual-image body data acquisition unit, and the seventh dual-image standardized body data acquisition unit is electrically connected or signal connected with the dual-image feature pyramid acquisition module 3.
[0417] The first dual-image standardized body data acquisition unit is configured to define a sliding window with a preset voxel size, and to traverse the standardized CBCT body data by using the sliding window at a preset sliding speed.
[0418] The second dual-image standardized body data acquisition unit is configured to calculate a gray histogram of a current block once every time the sliding window slides, to obtain a pixel distribution probability of the current block.
[0419] The third dual-image standardized body data acquisition unit is configured to calculate a cumulative distribution function of each block by using formula (1):
[0420] (1).
[0421] in, The cumulative distribution function is... This represents the upper limit of the grayscale value. For the first Each grayscale value , For the first The probability distribution of grayscale values.
[0422] The fourth dual-image normalized volume data acquisition unit is used to repeatedly execute steps S21 to S23 with CT volume data as the traversed object to obtain the cumulative distribution function of CT volume data.
[0423] The fifth dual-image normalized volume data acquisition unit is used to construct a grayscale mapping table based on two cumulative distribution functions of the same block according to equation (2):
[0424] (2).
[0425] in, The first CBCT volume data A mapping function that maps individual grayscale values to all grayscale values in CT body data. For CBCT volume data The cumulative distribution function of grayscale values, For CT body data The cumulative distribution function of grayscale values, The first CT body data Each grayscale value.
[0426] The sixth dual-image normalized volume data acquisition unit is used to extract at least one organ ROI region from CT volume data and CBCT volume data respectively through a 3D U-Net architecture pre-trained with a preset organ dataset.
[0427] The seventh dual-image standardized volume data acquisition unit is used to calculate the mean and standard deviation of grayscale values for each organ's ROI region, and then perform organ-specific grayscale standardization processing using equation (3) to obtain standardized CBCT volume data and standardized CT volume data in a unified format.
[0428] (3).
[0429] in, The normalized grayscale value. These are voxel coordinate values. coordinates The grayscale value of the voxel. a standard deviation of grayscale of voxels in the organ ROI region, a mean value of grayscale of voxels in the organ ROI region, a preset target organ grayscale standard deviation, a preset target organ grayscale mean value.
[0430] Further, the double-image feature pyramid acquisition module 3 specifically comprises a first double-image feature pyramid acquisition unit, a second double-image feature pyramid acquisition unit, a third double-image feature pyramid acquisition unit, and a fourth double-image feature pyramid acquisition unit which are sequentially electrically or signal connected; the first double-image feature pyramid acquisition unit is electrically or signal connected with the seventh double-image standardized volume data acquisition unit, and the fourth double-image feature pyramid acquisition unit is electrically or signal connected with the double-image deformation field mapping module 4.
[0431] The first double-image feature pyramid acquisition unit is configured to divide the standardized CBCT volume data into overlapping sub-blocks with a preset overlap rate, and convert each overlapping sub-block into a four-dimensional tensor; the second double-image feature pyramid acquisition unit is configured to define a network structure of the 3D-ResNet-34-FPN architecture based on a preset strategy, and sequentially input the four-dimensional tensors into the network structure to obtain a plurality of layer tensor features through multi-stage feature extraction; the third double-image feature pyramid acquisition unit is configured to sample the tensor features through a high-layer feature upsampling principle to obtain a CBCT feature pyramid; and the fourth double-image feature pyramid acquisition unit is configured to take the standardized CT volume data as an execution subject, repeatedly execute steps S31 to S33 to obtain a CT feature pyramid.
[0432] Further, the double-image deformation field mapping module 4 specifically comprises a first double-image deformation field mapping unit, a second double-image deformation field mapping unit, and a third double-image deformation field mapping unit which are sequentially electrically or signal connected; the first double-image deformation field mapping unit is electrically or signal connected with the seventh double-image standardized volume data acquisition unit, and the third double-image deformation field mapping unit is electrically or signal connected with the affine transformation parameter extraction module 5.
[0433] The first double-image deformation field mapping unit is configured to uniformly sample feature point coordinates from each layer of the CT feature pyramid to generate a control point coordinate set; the second double-image deformation field mapping unit is configured to sample feature point coordinates of the same position from corresponding layers of the CBCT feature pyramid and the CT feature pyramid to obtain a feature point coordinate set; and the third double-image deformation field mapping unit is configured to generate a deformation field of the control point coordinate set to the target point coordinate set through TPS minimization of bending energy.
[0434] Further, the affine transformation parameter extraction module 5 specifically comprises a first affine transformation parameter extraction unit, a second affine transformation parameter extraction unit, a third affine transformation parameter extraction unit, a fourth affine transformation parameter extraction unit, and a fifth affine transformation parameter extraction unit connected in sequence in electrical or signal connection. The first affine transformation parameter extraction unit is electrically or signal connected with the third dual-image deformation field mapping unit, and the fifth affine transformation parameter extraction unit is electrically or signal connected with the spatial transformation parameter sequence packaging and sending module 6.
[0435] The first affine transformation parameter extraction unit is configured to extract multi-scale features of the deformation field through a convolution layer of a pre-trained 3D-CNN network; the second affine transformation parameter extraction unit is configured to convert the multi-scale features into an initial displacement field with the same resolution as the deformation field through a deconvolution operation; the third affine transformation parameter extraction unit is configured to perform maximum value normalization on the initial displacement field to obtain a normalized displacement field; the fourth affine transformation parameter extraction unit is configured to flatten the three-dimensional vectors of the normalized displacement field into a matrix form and perform RANSAC-SVD singular value decomposition to obtain a left singular matrix and a right singular matrix; and the fifth affine transformation parameter extraction unit is configured to extract all affine transformation parameters from the left singular matrix and the right singular matrix.
[0436] Further, the spatial transformation parameter sequence packaging and sending module 6 specifically comprises a first spatial transformation parameter sequence packaging and sending unit, a second spatial transformation parameter sequence packaging and sending unit, and a third spatial transformation parameter sequence packaging and sending unit connected in sequence in electrical or signal connection. The first spatial transformation parameter sequence packaging and sending unit is electrically or signal connected with the fifth affine transformation parameter extraction unit.
[0437] The first spatial transformation parameter sequence packaging and sending unit is configured to map the translation parameters of all affine transformation parameters to the TranslationVector field and the rotation / scaling parameters to the RotationMatrix field according to the AffineTransformSequence tag of the DICOM-RT standard; the second spatial transformation parameter sequence packaging and sending unit is configured to assign a DICOM tag, a value representation, and a value length to each TranslationVector field and each RotationMatrix field, and package the spatial transformation parameter sequence obtained after packaging; and the third spatial transformation parameter sequence packaging and sending unit is configured to send the spatial transformation parameter sequence to a radiotherapy device.
[0438] It should be noted that the present embodiment is a functional module embodiment based on the above-mentioned method embodiment. The preferred, expanded, limited, exemplified, and principle explained parts of the present embodiment can be seen in the above-mentioned embodiment, and the present embodiment will not be described again.
[0439] The embodiment unifies CBCT verification image data and CT image data into preset matrix size volume data through CycleGAN modal conversion, obtains CBCT volume data and CT volume data; performs local histogram matching on the CBCT volume data and the CT volume data respectively through an adaptive histogram matching algorithm, and performs organ-specific gray scale normalization processing on the CBCT volume data and the CT volume data respectively based on a pre-segmentation mask, to obtain standardized CBCT volume data and standardized CT volume data in a unified format; the standardized CBCT volume data and the standardized CT volume data are respectively input into a pre-trained 3D-ResNet-34-FPN architecture, and through learning and training of the 3D-ResNet-34-FPN architecture, CBCT feature pyramids and CT feature pyramids are respectively obtained; sampling feature points in the CT feature pyramids are taken as a control point coordinate set, sampling feature points in the CBCT feature pyramids are taken as a target point coordinate set, and a deformation field from the control point coordinate set to the target point coordinate set is generated through a TPS-Net architecture; the deformation field is converted into a zero-centered displacement field through a 3D-CNN network, and all affine transformation parameters of the displacement field are extracted through RANSAC-SVD singular value decomposition; all the affine transformation parameters are packaged into a spatial transformation parameter sequence conforming to the DICOM-RT standard, and are sent to a radiotherapy device. The embodiment first reduces the modal difference between CBCT and CT images through CycleGAN modal conversion and organ-specific gray scale normalization, effectively laying a unified data foundation for subsequent registration; then multi-level features are extracted from the standardized data through the 3D-ResNet-34-FPN architecture, and a deformation field accurately reflecting local deformation is generated through the TPS-Net, which significantly improves the registration accuracy, especially effectively capturing nonlinear changes such as soft tissue deformation and organ displacement; finally, through 3D-CNN displacement field normalization and RANSAC-SVD decomposition, global affine transformation parameters are robustly extracted and packaged into a DICOM-RT standard protocol for direct transmission to a radiotherapy device, not only automating the traditional manual operation process and reducing the workload and subjective errors of medical staff, but also improving the reliability and repeatability of the registration result, so that the embodiment can provide real-time and accurate positioning error correction parameters for online image-guided radiotherapy (IGRT), ensuring that the radiation dose can be more accurately projected to the target area, thereby improving the treatment effect while protecting the surrounding normal tissue, so that the embodiment has high calculation accuracy and high automation degree.
[0440] Figure 3 FIG. 1 is a structural schematic diagram of an electronic device according to an embodiment of the present application. Figure 3 As shown in FIG. 1, the electronic device 7 includes a processor 71 and a memory 72 coupled to the processor 71.
[0441] The memory 72 stores program instructions for implementing the federated learning-based government data group collaborative energy-saving method of any of the above embodiments.
[0442] The processor 71 is configured to execute the program instructions stored in the memory 72 to perform the federated learning-based government data group collaborative energy-saving.
[0443] The processor 71 can also be referred to as a CPU (Central Processing Unit). The processor 71 can be an integrated circuit chip having a processing capability of signals. The processor 71 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application-Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0444] Further, Figure 4 FIG. 8 is a structural schematic diagram of a storage medium according to an embodiment of the present application. Figure 4 The storage medium 8 according to the embodiment of the present application stores program instructions 81 capable of implementing all the above methods. The program instructions 81 can be stored in the above storage medium in the form of a software product, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The above storage medium includes a U disk, a mobile hard disk, a ROM (Read-Only Memory), a RAM (Random Access Memory), a magnetic disk or an optical disk, and various media capable of storing program codes, or a terminal device such as a computer, a server, a mobile phone, a tablet, etc.
[0445] In the several embodiments provided in the present application, it should be understood that the disclosed system, system and method can be implemented in other ways. For example, the above-described system embodiments are merely schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be omitted or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, systems or units, and can be electrical, mechanical, signal or other forms.
[0446] In addition, the various functional units in the embodiments of the present application can be integrated in one processing unit, or each can exist physically as a separate unit, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a software functional unit. The above is only an implementation of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method of registering radiotherapy position verification data, said radiotherapy position verification data derived from CBCT verification image data of a patient to be treated on a radiotherapy device, said patient to be treated having CT image data, characterized in that, The registration method comprises: Step S1, unify the CBCT verification image data and the CT image data into volume data of a preset matrix size through CycleGAN modal conversion, to obtain CBCT volume data and CT volume data; Step S2, respectively perform local histogram matching on the CBCT volume data and the CT volume data through an adaptive histogram matching algorithm, and respectively perform organ-specific gray scale normalization processing on the CBCT volume data and the CT volume data based on a pre-segmentation mask, to obtain standardized CBCT volume data and standardized CT volume data in a unified format; Step S3, respectively input the standardized CBCT volume data and the standardized CT volume data into a pre-trained 3D-ResNet-34-FPN architecture, and respectively obtain a CBCT feature pyramid and a CT feature pyramid through learning and training of the 3D-ResNet-34-FPN architecture; Step S4, take sampling feature points in the CT feature pyramid as a control point coordinate set, take sampling feature points in the CBCT feature pyramid as a target point coordinate set, and generate a deformation field from the control point coordinate set to the target point coordinate set through a TPS-Net architecture; Step S5, convert the deformation field into a zero-centered displacement field through a 3D-CNN network, and then extract all affine transformation parameters of the displacement field through RANSAC-SVD singular value decomposition; Step S6, encapsulate all the affine transformation parameters into a spatial transformation parameter sequence conforming to a DICOM-RT standard, and send to the radiotherapy device; Step S1, unify the CBCT verification image data and the CT image data into volume data of a preset matrix size through CycleGAN modal conversion, to obtain CBCT volume data and CT volume data, comprising: Step S11, convert the CBCT verification image data and the CT image data into a DICOM standard format; Step S12, align a spatial coordinate system of the CBCT verification image data with a spatial coordinate system of the CT image data through a non-rigid registration method; Step S13, define two generators of the CycleGAN modal conversion based on a 3D U-Net architecture, and define two discriminators of the CycleGAN modal conversion based on a 3D PatchGAN; Step S14, convert the CBCT verification image data into CT style data through one of the generators, and convert the CT image data into CBCT style data through the other generator; Step S15, calculate an adversarial loss based on a Wasserstein distance between the CBCT style data and the CBCT verification image data through one of the discriminators, and calculate an adversarial loss based on a Wasserstein distance between the CT style data and the CT image data through the other discriminator; Step S16, define an L1 loss of the CycleGAN modal conversion, and simultaneously iterate the two adversarial losses through back propagation until the L1 loss reaches a minimum value. Step S17, when the L1 loss reaches the minimum value, obtaining the CT style data of the last iteration by one of the generators, and the CT style data and the CT image data are the volume data of a unified preset matrix size.
2. The registration method of claim 1, wherein, Step S2, performing local histogram matching on the CBCT volume data and the CT volume data respectively by an adaptive histogram matching algorithm, and performing organ-specific gray scale normalization processing on the CBCT volume data and the CT volume data respectively based on a pre-segmentation mask, to obtain standardized CBCT volume data and standardized CT volume data in a unified format, including: Step S21, defining a sliding window of a preset voxel size, and traversing the standardized CBCT volume data at a preset sliding speed through the sliding window; Step S22, calculating the gray scale histogram of the current block once every time the sliding window slides to obtain the pixel distribution probability of the current block; Step S23, calculating the cumulative distribution function of each block respectively by formula (1): (1); wherein, is the cumulative distribution function, is the upper limit of the gray value, is the gray value, , is the probability distribution of the gray value, . Step S24, taking the CT volume data as the traversed object, repeating steps S21 to S23 to obtain the cumulative distribution function of the CT volume data; Step S25, constructing a gray scale mapping table according to formula (2) based on the cumulative distribution functions of the same block: (2); wherein is a mapping function mapping the i-th gray value of the CBCT volume data to all gray values of the CT volume data, is a mapping function mapping the i-th gray value of the CBCT volume data to all gray values of the CT volume data, is a cumulative distribution function of the i-th gray value of the CBCT volume data, is a cumulative distribution function of the i-th gray value of the CBCT volume data, is a cumulative distribution function of the i-th gray value of the CT volume data, is a cumulative distribution function of the i-th gray value of the CT volume data, is a cumulative distribution function of the i-th gray value of the CT volume data, is a cumulative distribution function of the i-th gray value of the CT volume data, Step S26, extracting at least one organ ROI region of the CT volume data and the CBCT volume data respectively through a 3D U-Net architecture pre-trained by a preset organ data set; Step S27, performing organ-specific gray scale normalization processing on the gray mean and standard deviation of each organ ROI region respectively, and obtaining standardized CBCT volume data and standardized CT volume data in a unified format respectively by formula (3): (3); wherein, is a normalized gray value, is a voxel coordinate value, is a coordinate of a voxel, is a standard deviation of a gray value of a voxel in an organ ROI region, is a mean value of a gray value of a voxel in an organ ROI region, is a preset target organ gray standard deviation, is a preset target organ gray mean value.
3. The registration method of claim 1, wherein, Step S3, inputting the standardized CBCT volume data and the standardized CT volume data into a pre-trained 3D-ResNet-34-FPN architecture respectively, and obtaining a CBCT feature pyramid and a CT feature pyramid respectively through learning and training of the 3D-ResNet-34-FPN architecture, including: Step S31, dividing the standardized CBCT volume data into overlapping subblocks with a preset overlap rate, and converting each overlapping subblock into a four-dimensional tensor respectively; Step S32, defining the network structure of the 3D-ResNet-34-FPN architecture based on a preset strategy, and inputting the four-dimensional tensors into the network structure in sequence to obtain a plurality of layer tensor features through multi-stage feature extraction; Step S33, sampling the tensor features through a high-level feature upsampling principle to obtain the CBCT feature pyramid; Step S34, taking the standardized CT volume data as the execution subject, repeating steps S31 to S33 to obtain the CT feature pyramid.
4. The registration method of claim 1, wherein, Step S4, taking the sampling feature points in the CT feature pyramid as a control point coordinate set and the sampling feature points in the CBCT feature pyramid as a target point coordinate set, generating a deformation field from the control point coordinate set to the target point coordinate set through a TPS-Net architecture, including: Step S41, uniformly sampling feature point coordinates from each layer of the CT feature pyramid to generate a control point coordinate set; Step S42, sampling feature point coordinates of the same position from the corresponding layers of the CBCT feature pyramid and the CT feature pyramid to obtain a feature point coordinate set; Step S43, generating a deformation field from the control point coordinate set to the target point coordinate set by minimizing the bending energy through TPS.
5. The registration method of claim 1, wherein, Step S5, converting the deformation field into a zero-centered displacement field through a 3D-CNN network, and then extracting all affine transformation parameters of the displacement field through RANSAC-SVD singular value decomposition, including: Step S51, extracting multi-scale features of the deformation field through the convolution layer of the pre-trained 3D-CNN network; Step S52, converting the multi-scale features into an initial displacement field with the same resolution as the deformation field through the deconvolution operation; Step S53, performing maximum normalization on the initial displacement field to obtain a normalized displacement field; Step S54, flattening the three-dimensional vector of the normalized displacement field into a matrix form and performing RANSAC-SVD singular value decomposition to obtain a left singular matrix and a right singular matrix; Step S55, extracting all affine transformation parameters from the left singular matrix and the right singular matrix.
6. The registration method of claim 1, wherein, Step S6, packaging all affine transformation parameters into a spatial transformation parameter sequence conforming to the DICOM-RT standard and sending it to the radiotherapy device, including: Step S61, mapping the translation parameters of all affine transformation parameters to the TranslationVector field and the rotation / scaling parameters to the RotationMatrix field according to the AffineTransformSequence tag of the DICOM-RT standard; Step S62, assigning DICOM tags, value representations, and value lengths to each TranslationVector field and each RotationMatrix field to obtain the spatial transformation parameter sequence after packaging; Step S63, sending the spatial transformation parameter sequence to the radiotherapy device.
7. A registration apparatus of radiotherapy position verification data, the registration apparatus being applied to the registration method according to any one of claims 1 to 6, characterized in that, The registration device comprises: A dual-image body data acquisition module configured to convert CBCT verification image data and CT image data into body data with a preset matrix size through CycleGAN modal conversion to obtain CBCT body data and CT body data; A dual-image standardized body data acquisition module configured to perform local histogram matching on the CBCT body data and the CT body data through an adaptive histogram matching algorithm, and perform organ-specific gray scale standardization processing on the CBCT body data and the CT body data based on a pre-segmentation mask to obtain standardized CBCT body data and standardized CT body data in a unified format; a dual-image feature pyramid acquisition module, configured to input the standardized CBCT volume data and the standardized CT volume data into a pre-trained 3D-ResNet-34-FPN architecture respectively, and obtain a CBCT feature pyramid and a CT feature pyramid respectively through learning and training of the 3D-ResNet-34-FPN architecture; a dual-image deformation field mapping module, configured to take the sampling feature points in the CT feature pyramid as a control point coordinate set, take the sampling feature points in the CBCT feature pyramid as a target point coordinate set, and generate a deformation field from the control point coordinate set to the target point coordinate set through a TPS-Net architecture; an affine transformation parameter extraction module, configured to convert the deformation field into a zero-centered displacement field through a 3D-CNN network, and extract all affine transformation parameters of the displacement field through RANSAC-SVD singular value decomposition; a spatial transformation parameter sequence packaging and sending module, configured to package all the affine transformation parameters into a spatial transformation parameter sequence conforming to a DICOM-RT standard, and send the spatial transformation parameter sequence to the radiotherapy device.
8. An electronic device, comprising: The computer readable storage medium stores program instructions, and the program instructions are executed by the processor to implement the registration method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores program instructions, and the program instructions are executed by the processor to implement the registration method according to any one of claims 1 to 6.
Citation Information
Patent Citations
CT and CBCT registration method based on pseudo CT
CN114820730A