3D image reconstruction model training method, 3D image reconstruction model training device, 3D image reconstruction method and 3D image reconstruction device
By dicing 2D Xray images and establishing correspondence, using neural network model training, the problem of too large memory limitation and solution space when generating 3D segmented images in the prior art is solved, and efficient 3D image reconstruction and accurate reconstruction of spinal bone model are achieved.
Patent Information
- Application Number
- CN202311810137.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has too much memory limitation and solution space when generating 3D segmented images on 2D Xray images, resulting in extremely low convergence speed of the model, making it difficult to effectively solve the 3D reconstruction problem of the spinal bone model.
By dicing the 2D Xray image into multiple positive side 2D sub-maps, and establishing the correspondence between the 2D sub-maps and the target 3D segmented image, using neural network model training, the model learns the correspondence between the 2D anchor point and the 3D anchor point until the model converges, and obtains the 3D image reconstruction model.
This reduces the video memory requirement during neural network model training, improves the convergence speed of the model, realizes higher-precision 3D image reconstruction, and solves the 3D reconstruction problem of spinal bone model.
Smart Images

Figure CN120218154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a 3D image reconstruction model training method and device, and a 3D image reconstruction method and device. Background Art
[0002] X-ray is a commonly used, painless and fast medical imaging method. The principle of X-ray is that during the X-ray examination, different tissues in the patient's body absorb X-rays at different rates. During the X-ray examination, electromagnetic waves pass through the patient and are absorbed at different rates, projecting the radiation density of the patient's volume onto the exposed photographic film, thereby generating a 2D image.
[0003] The basic principle of CT imaging is to use an X-ray beam to scan a certain thickness of the human body. Simply put, CT scans the human body to produce multiple slices, and then reconstructs the slice information to obtain detailed three-dimensional information of the human body. Although CT can provide detailed three-dimensional information, it will bring additional money and safety costs. CT imaging is more expensive than taking a single X-ray, and it will expose patients to an order of magnitude higher level of radiation. Therefore, in clinical medicine, doctors generally use X-rays more commonly than CT.
[0004] When a doctor gets a two-dimensional X-ray, he can understand the three-dimensional structure in the two-dimensional plane X-ray based on prior knowledge. Therefore, in the human field, with sufficient prior knowledge, the three-dimensional human body structure can be understood from the two-dimensional X-ray. Therefore, how to reconstruct a three-dimensional spinal bone model from a two-dimensional X-ray image, obtain more data in space, and better classify the degree of scoliosis of the spine has become a technical problem that needs to be solved in the field of spinal surgery planning.
[0005] Generating 3D segmented images from 2D X-ray means that the input information includes the frontal and lateral X-ray images and X-ray parameter information (resolution, IOS center, tube distance information), and then using the trained neural network model to output the 3D segmentation results. However, the existing technology has a very low convergence speed due to the large solution space and the limitation of the video memory performance of the hardware device for training the neural network model.
[0006] In order to facilitate understanding of the technical solution of this application, several existing technical terms are explained as follows:
[0007] 1. DRR (Digitally Reconstructed Radiograph), the full name is Digital Reconstruction Radiograph. Simply put, DRR is to generate 2D images (equivalent to DR images, that is, coronal images or sagittal images) from 3D CT volume data (multiple cross-sectional image data) through mathematical simulation algorithms.
[0008] Obtaining a DRR image is an important pre - step in medical image registration. Its main purpose is: through CT three - dimensional images, to obtain simulated X - ray images, and this process is also called digital image reconstruction.
[0009] (1) First, when performing image registration, the two entities to be registered should have the same dimension. That is, either both are 2D images for planar registration; or both are 3D models for spatial registration. Therefore, when performing 2D - 3D registration, it is necessary to reduce the dimension of the 3D model to two - dimensional and then perform 2D - 2D registration to complete.
[0010] (2) Second, whether taking X - rays or CT scans, there is radiation, especially in treatment methods that require real - time observation. It is impossible to continuously perform DR or CT scans throughout the diagnosis and treatment process. Then, the DRR algorithm can be used. After having CT image data, through this algorithm, 3D - 2D conversion into a DRR image is achieved, obtaining an effect similar to an X - ray film (equivalent to flattening the CT), saving the patient the step of taking another X - ray. Finally, the generated DRR image is used for registration (the registration should be between the generated DRR image and the existing X - ray image).
[0011] (3) General scenarios for registration:
[0012] The patient had a CT scan before surgery and a DR scan during surgery. Therefore, it is necessary to register the DR image during surgery with the CT image before surgery to ensure the accuracy of the surgical position, that is, to be consistent with the pre - operative diagnosis and treatment. Due to what is described in (1), the registration must be in the same dimension. Therefore, the later - taken CT image needs to be processed into 2D using the DRR algorithm and then registered with the previously taken DR (i.e., the existing X - ray picture).
[0013] In addition, during the treatment process, the patient generally takes a DR first to preliminarily check the lesion location. After diagnosis, if surgical treatment is required, the patient needs to undergo further examinations, such as taking a CT scan. However, the results of the two scans taken before and after generally have an offset due to the different positions of the patient during the two detections. Therefore, it is necessary to register the images taken before and after, that is, the images taken before and after formulating the treatment plan.
[0014] Second, an affine transformation, also known as an affine mapping, refers to in geometry, a linear transformation of a vector space followed by a translation, transforming it into another vector space.
[0015] It is a linear transformation from one coordinate system to another, which preserves the "flatness" of two-dimensional graphics (a straight line remains a straight line after transformation) and "parallelism" (the relative positional relationship between graphics remains unchanged, parallel lines remain parallel, and the positional order of points on a straight line remains unchanged).
[0016] Any affine transformation can be expressed in the form of "multiplying by a matrix (linear transformation) and then adding a vector (translation)".
[0017] III. Dice is the most frequently used metric in medical image competitions. It is a metric for measuring the similarity of sets, usually used to calculate the similarity between two samples, with a value range of [0, 1]. It is often used for image segmentation in medical images. The best result of segmentation is 1, and the worst result is 0. Summary of the Invention
[0018] The purpose of the present invention is to overcome the above technical deficiencies and provide a method and device for training a 3D image reconstruction model, as well as a 3D image reconstruction method and device, to solve the problem of difficult convergence caused by video memory limitations and too large a solution space when generating 3D segmentation images from 2D X-ray images in related technologies.
[0019] To achieve the above technical objectives, the present invention adopts the following technical solutions:
[0020] According to the first aspect of the present invention, there is provided a method for training a 3D image reconstruction model, including:
[0021] Cut the 2D X-ray image into multiple 2D sub-images of the front and side views;
[0022] Establish the corresponding relationship between the 2D anchor points on the 2D X-ray image and the 3D anchor points on the target 3D segmentation image, and perform position normalization on the corresponding 3D anchor points according to the coordinates of the 2D anchor points, so that the 2D anchor points on the 2D X-ray image and the 3D anchor points on the 3D segmentation image are in one-to-one correspondence;
[0023] Establish the corresponding relationship between the 2D sub-images and the target 3D segmentation image according to the one-to-one correspondence between the 2D anchor points and the 3D anchor points;
[0024] Use the 2D sub-images and the position encoding of the 2D anchor points as inputs, and the target 3D segmentation image and the 3D anchor points after position normalization as the ground truth to train a pre-constructed neural network model, so that the neural network model learns the corresponding relationship between the 2D anchor points and the 3D anchor points, and the corresponding relationship between the 2D sub-images and the 3D segmentation image, until the model converges to obtain a 3D image reconstruction model;
[0025] The output of the 3D image reconstruction model is the predicted 3D segmentation image and its 3D anchor points, and the loss function of the 3D image reconstruction model is related to the ground truth and the output result of the model.
[0026] Preferably, the correspondence between the 2D anchor points on the 2D X-ray image and the 3D anchor points on the target 3D segmentation image is established through the following steps, including:
[0027] Select 3D anchor points on the 3D CT image;
[0028] Generate a 3D sphere according to the 3D anchor points;
[0029] According to the DRR hyperparameters and the 3D sphere, use the digital reconstructed radiograph (DRR) algorithm to obtain the 2D X-ray images of the anteroposterior and lateral views of the part to be detected and the 2D heatmaps of the anteroposterior and lateral views;
[0030] Determine the local maximum values on the 2D heatmap as the 2D anchor points corresponding one-to-one to the 3D anchor points.
[0031] Preferably, the correspondence between the 2D anchor points on the 2D X-ray image and the 3D anchor points on the target 3D segmentation image is established through the following steps, including:
[0032] Select 3D anchor points on the 3D CT image;
[0033] According to the 3D anchor points, DRR hyperparameters and the following formula (1), use the affine matrix to obtain the 2D anchor point coordinates corresponding to the 3D anchor points, specifically:
[0034]
[0035] Where, P achor3d represents the 3D anchor point coordinates, which is a known quantity; represents the affine matrix, which is a known quantity; P anchor2d represents the 2D anchor point coordinates;
[0036] Where, P achor3d is a 1*4 vector, and the first three values in the vector respectively represent the X-axis coordinate, Y-axis coordinate and Z-axis coordinate of the 3D anchor point, and the fourth value is a constant 1; is a 4*6 affine matrix, P anchor2d is a 1*6 vector, and the six values in the vector respectively represent the anteroposterior 2D anchor point coordinates and the lateral 2D anchor point coordinates, expressed as: [anteroposterior (x, y, 1), lateral (x, y, 1)].
[0037] Preferably, the position normalization process is performed on the corresponding 3D anchor points according to the coordinates of the 2D anchor points, specifically:
[0038] According to the coordinates of the 2D anchor points, the coordinates of the corresponding 3D anchor points are determined to be on the fixed points of the corresponding 3D segmentation image.
[0039] According to a second aspect of the present invention, there is provided a 3D image reconstruction method, including:
[0040] Obtain 2D X-ray images of the front and side views of the part to be detected;
[0041] Predict 3D anchor points on the 3D segmentation image based on the approximate projection of the 2D anchor points on the 2D X-ray image, and then obtain the corrected 2D anchor points based on the 3D anchor points, so that the 3D anchor points correspond one-to-one with the corrected 2D anchor points;
[0042] Input the 2D X-ray image and the position encoding of the corrected 2D anchor points into the 3D image reconstruction model trained according to the above method for model prediction to obtain the 3D segmentation image of each bone joint;
[0043] According to the 3D anchor points on the 3D segmentation image, stitch the 3D segmentation image output by the model to obtain the image reconstruction result of the whole part to be detected.
[0044] Preferably, the predicting 3D anchor points on the 3D segmentation image based on the approximate projection of the 2D anchor points on the 2D X-ray image, and then obtaining the corrected 2D anchor points based on the 3D anchor points, includes:
[0045] Predict 2D candidate key points on the front and side 2D X-ray images based on the key point detection algorithm Is a 1*6 vector, and the six values in the vector respectively represent the front view 2D anchor point coordinates and the side view 2D anchor point coordinates, expressed as: [front view (x, y, 1), side view (x, y, 1)];
[0046] In order to make the front view 2D anchor point coordinates and the side view 2D anchor point coordinates correspond to a unique 3D anchor point, the following steps are performed, including:
[0047] First select As the center, a set PP with a radius less than R
[0048]
[0049] Use the pseudo-inverse of the affine matrix to Project into 3D space to obtain
[0050] Among them, Pinv is the matrix pseudo-inverse;
[0051] Because the projection is singular, re-project the 3D image to obtain
[0052]
[0053] The points obtained by reprojection relative to the original with the smallest distance are determined as the modified 2D anchor points:
[0054] According to the third aspect of the present invention, there is provided a 3D image reconstruction model training device, including:
[0055] A cutting module that cuts the 2D X-ray image into multiple 2D sub-images of the anteroposterior and lateral views;
[0056] An establishing module for establishing the correspondence between the 2D anchor points on the 2D X-ray image and the 3D anchor points on the target 3D segmentation image, and performing position normalization processing on the corresponding 3D anchor points according to the coordinates of the 2D anchor points, so that the 2D anchor points on the 2D X-ray image and the 3D anchor points on the 3D segmentation image are in one-to-one correspondence;
[0057] It is also used to establish the correspondence between the 2D X-ray image and the target 3D segmentation image according to the one-to-one correspondence between the 2D anchor points and the 3D anchor points;
[0058] A training module for using the 2D sub-images and the position encodings of the 2D anchor points as inputs, and using the target 3D segmentation image and the 3D anchor points after position normalization as the ground truth to train a pre-constructed neural network model, so that the neural network model learns the correspondence between the 2D anchor points and the 3D anchor points, and the correspondence between the 2D sub-images and the 3D segmentation image, until the model converges to obtain a 3D image reconstruction model;
[0059] The output of the 3D image reconstruction model is the predicted 3D segmentation image and its 3D anchor points, and the loss function of the 3D image reconstruction model is related to the ground truth and the output result of the model.
[0060] According to the fourth aspect of the present invention, there is provided a 3D image reconstruction device, including:
[0061] An acquisition module for acquiring 2D X-ray images of the anteroposterior and lateral views of the part to be detected;
[0062] A correction module for predicting the 3D anchor points on the 3D segmentation image based on the approximate projection of the 2D anchor points on the 2D X-ray image, and then obtaining the corrected 2D anchor points based on the 3D anchor points, so that the 3D anchor points and the corrected 2D anchor points are in one-to-one correspondence;
[0063] A reconstruction module, configured to use the 2D X-ray image and the position encoding of the corrected 2D anchor points as inputs, and input them into a 3D image reconstruction model trained according to the above method for model prediction, so as to obtain 3D segmentation images of each bone joint;
[0064] It is also configured to stitch the 3D segmentation images output by the model according to the 3D anchor points on the 3D segmentation images, so as to obtain an image reconstruction result of the overall part to be detected.
[0065] According to a fifth aspect of the present invention, there is provided an electronic device, including:
[0066] A processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0067] The memory is used to store a computer program;
[0068] The processor is configured to implement the above method when executing the program stored on the memory.
[0069] According to a sixth aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the above method.
[0070] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0071] Before model training, by establishing a one-to-one correspondence between the 2D anchor points on the 2D X-ray image and the 3D anchor points on the target 3D segmentation image, the bone structures on the complex and large-structured 2D X-ray image are split into simple segmentation structures, so that both the 2D X-ray image and the target 3D segmentation image can be split out for training. At the same time, different bone joints are decoupled, so that the video memory capacity required for training the neural network model is reduced, and the convergence speed is faster.
[0072] In addition, because the 2D anchor points and the 3D anchor points uniquely determine the relationship between the two, it is easier to perform data augmentation, and finally the algorithm can train a 3D image reconstruction model with higher accuracy.
[0073] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 is a flowchart of a method for training a 3D image reconstruction model shown according to an exemplary embodiment;
[0075] Figure 2It is a schematic diagram showing the simulation generation of anteroposterior and lateral 2D X-ray images from a CT 3D image using the DRR algorithm according to an exemplary embodiment;
[0076] Figure 3 It is a flowchart showing the establishment of the correspondence between 2D and 3D anchors according to an exemplary embodiment;
[0077] Figure 4 It is a flowchart showing the establishment of the correspondence between 2D and 3D anchors according to another exemplary embodiment;
[0078] Figure 5 It is a flowchart showing the establishment of the correspondence between 2D and 3D anchors according to another exemplary embodiment;
[0079] Figure 6 It is a schematic diagram showing the splitting result based on 3D anchors according to an exemplary embodiment;
[0080] Figure 7 It is a schematic diagram showing the splitting result based on 2D anchors according to an exemplary embodiment;
[0081] Figure 8 It is a schematic diagram showing the training of a neural network model according to an exemplary embodiment;
[0082] Figure 9 It is a schematic flowchart showing a 3D image reconstruction method according to an exemplary embodiment;
[0083] Figure 10 It is a schematic flowchart showing a 3D image reconstruction method according to an exemplary embodiment;
[0084] Figure 11 It is a schematic block diagram showing a 3D image reconstruction model training device according to an exemplary embodiment;
[0085] Figure 12 It is a schematic block diagram showing a 3D image reconstruction device according to an exemplary embodiment;
[0086] Figure 13 A schematic block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0087] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0088] As described in the background art above, in the related art, due to the limited detection range of X-ray images, the generalization ability of the model is poor when facing various lesion data, and the recognition accuracy of the position and bending posture of the spine is poor.
[0089] In order to effectively solve the problems in the related art, the present invention provides a training method and device for a feature point extraction model of a part to be detected, and a method and device for automatically generating surgical planning parameters, which will be specifically described below.
[0090] Embodiment 1
[0091] Figure 1 is a flowchart of a 3D image reconstruction model training method shown according to an exemplary embodiment. As Figure 1 shown, the method includes:
[0092] Step S11: Cut the 2D X-ray image into multiple 2D subgraphs of the anteroposterior and lateral views;
[0093] Step S12: Establish the correspondence between the 2D anchor points on the 2D X-ray image and the 3D anchor points on the target 3D segmentation image, and perform position normalization on the corresponding 3D anchor points according to the coordinates of the 2D anchor points, so that the 2D anchor points on the 2D X-ray image and the 3D anchor points on the 3D segmentation image are in one-to-one correspondence;
[0094] Step S13: Establish the correspondence between the 2D subgraph and the target 3D segmentation image according to the one-to-one correspondence between the 2D anchor points and the 3D anchor points;
[0095] Step S14: Use the 2D subgraph and the position encoding of the 2D anchor points as inputs, and use the target 3D segmentation image and the 3D anchor points after position normalization as the ground truth to train a pre-constructed neural network model, so that the neural network model learns the correspondence between the 2D anchor points and the 3D anchor points, and the correspondence between the 2D subgraph and the 3D segmentation image, until the model converges to obtain a 3D image reconstruction model;
[0096] The output of the 3D image reconstruction model is the predicted 3D segmentation image and its 3D anchor points, and the loss function of the 3D image reconstruction model is related to the ground truth and the output result of the model.
[0097] It should be noted that the technical solution provided in this embodiment runs in the controller of the medical device in specific practice, or is loaded and run in an electronic device connected to the controller. The controller of the medical device executes the corresponding method by calling the program stored in the electronic device.
[0098] It can be understood that for the technical solution provided in this embodiment, before model training, by establishing a one-to-one correspondence between the 2D anchor points on the 2D Xray image and the 3D anchor points on the target 3D segmentation image, the bone structures on the complex and large-structured 2D Xray image are split into simple segmentation structures, so that both the 2D Xray image and the target 3D segmentation image can be split for training (as Figure 6 shown, based on the 3D anchor points, the 3D CT image can be split into 3D segmentation images of multiple single-segment bone structures; as Figure 7 shown, based on the 2D anchor points, the coronal 2D Xray image can be split into multiple pairs of anteroposterior 2D sub-images). At the same time, different bone joints are decoupled, reducing the video memory capacity required for neural network model training and achieving a faster convergence speed.
[0099] In addition, since the 2D anchor points and the 3D anchor points uniquely determine the relationship between the two, data augmentation becomes easier, and finally the algorithm can train a 3D image reconstruction model with higher accuracy.
[0100] In specific practice, in step S11, cutting the 2D Xray image into multiple anteroposterior 2D sub-images includes:
[0101] Referring to Figure 2 , according to the 3D CT image and DRR hyperparameters, the DRR algorithm simulates and generates anteroposterior 2D Xray images, where Figure 2 the left side of the bottom row in
[0102] is the anterior 2D Xray image, and the right side is the lateral 2D Xray image. Figure 3 In a possible implementation, referring to
[0103] select 3D anchor points on the 3D CT image;
[0104] generate a 3D sphere according to the 3D anchor points;
[0105] according to the DRR hyperparameters and the 3D sphere, use the digital reconstruction radiography (DRR) algorithm to obtain the anteroposterior 2D Xray image and anteroposterior 2D heat map of the part to be detected;
[0106] determine the local maximum values on the 2D heat map as the 2D anchor points corresponding one-to-one to the 3D anchor points.
[0107] In another possible implementation, referring to Figure 4, the correspondence between the 2D anchor points on the 2D X-ray image and the 3D anchor points on the target 3D segmentation image is established through the following steps, including:
[0108] Select 3D anchor points on the 3D CT image;
[0109] According to the 3D anchor points, DRR hyperparameters and the following formula (1), the 2D anchor point coordinates corresponding to the 3D anchor points are obtained using the affine matrix, specifically:
[0110]
[0111] where P achor3d represents the 3D anchor point coordinates, which are known quantities; represents the affine matrix, which is a known quantity; P anchor2d represents the 2D anchor point coordinates;
[0112] where P achor3d is a 1*4 vector, and the first three values in the vector respectively represent the X-axis coordinate, Y-axis coordinate, and Z-axis coordinate of the 3D anchor point, and the fourth value is the constant 1; is a 4*6 affine matrix, P anchor2d is a 1*6 vector, and the six values in the vector respectively represent the anterior-posterior 2D anchor point coordinates and the lateral 2D anchor point coordinates, expressed as: [anterior-posterior (x, y, 1), lateral (x, y, 1)].
[0113] In specific practice, in step S12, position normalization processing is performed on the corresponding 3D anchor points to make the 2D anchor points on the 2D X-ray image correspond one-to-one with the 3D anchor points on the 3D segmentation image, including:
[0114] According to the coordinates of the 2D anchor points, the coordinates of the corresponding 3D anchor points are determined at fixed points on the corresponding 3D segmentation image (it should be noted that the fixed point here can be the center point of the 3D segmentation image or a non-center point, such as fixing the 3D anchor point to the voxel position with z, y, x coordinates of (18, 18, 18)).
[0115] It is understandable that when the position of the 2D anchor changes, the position of the mapped 3D anchor also changes. Taking the spine as the part to be detected, in an ideal state: one single vertebral segment of the spine corresponds to one 2D anchor, and this 2D anchor is mapped onto the target 3D segmentation image of the single vertebral segment. It is hoped that this 3D anchor is located at the fixed point position of the single vertebral segment. However, due to possible errors in the selection of the 2D anchor, for example, it may be 3 mm to the left of the expected position, or 3 mm to the right of the expected position, or 3 mm above the expected position, or 3 mm below the expected position. This will cause a position deviation of the mapped 3D anchor, and the coordinate position will not be at the fixed point position of the single vertebral segment. In order not to affect the stitching of the final target 3D segmentation image, the technical solution provided in this embodiment is to fix the coordinates of the mapped 3D anchor to the fixed point of the target 3D segmentation image. That is, assuming that due to the position perturbation of the 2D anchor, the coordinates of the mapped 3D anchor are not at the fixed point of the target 3D segmentation image, by moving the target 3D segmentation image, the position of the 3D anchor is aligned to the fixed point of the 3D segmentation image before movement.
[0116] In this way, the 3D anchors on all the 3D segmentation images of the single vertebral segments are at the fixed point positions of their respective images, ensuring that the spine obtained after stitching is a smooth curve without abnormal points.
[0117] In specific practice, in step S14, the position encoding of the 2D subgraph and the 2D anchor is used as the input, and the target 3D segmentation image and the 3D anchor after position normalization are used as the ground truth to train the pre-constructed neural network model, so that the neural network model learns the correspondence between the 2D anchor and the 3D anchor, and the correspondence between the 2D subgraph and the 3D segmentation image until the model converges to obtain a 3D image reconstruction model, as shown in Figure 8 shown.
[0118] In specific practice, assuming that the 2D anchor is known, the following two methods can be used to let the neural network model to be trained know the coordinate information of the 2D anchor:
[0119] First, in the CNN network model, map the 3D sphere onto 2D, regard this information as a 2D heat map, and superimpose it on the 2D subgraph to let the CNN network model learn this information.
[0120] Second, in the transformer network model, directly use position embedding to perform position encoding on the 2D anchor.
[0121] See Figure 8, the pre - constructed neural network model includes: CNN network model, transformer network model, encoder, decoder, etc.
[0122] Solution 1: For multiple input cut front - side views, the corresponding heatmap can be expressed by the following formula:
[0123] I ij =(P cor -(i,j)) 2 , where (0 < i < hight, 0 < j < widht)
[0124] P cor represents the pixel coordinates of the 2D anchor point, and I ij represents the pixel value of the heatmap coordinate (i,j); hight represents the height of the image, and widht represents the width of the image; the heatmap characterizes the positional relationship of the 2D anchor point on the front - side bitmap.
[0125] Input: 1) Multiple cut front - side views (2D sub - images); 2) The heatmap into the CNN network model, and the output is obtained: [256, 256, 3]. These three channels respectively represent the original image, histogram equalization, and anchor point information I. Input this multi - dimensional feature vector into the encoder for convolution and downsampling to compress the information, and finally compress it into [128, 8, 8, 8], where 128 represents the feature information and [8, 8, 8] represents the spatial information. Then use the decoder for 4 times of upsampling to finally obtain the 3D segmentation information of [8, 128, 128, 128]. Calculate the loss function by comparing the segmented 3D bone structure output by the model with the ground truth of the 3D bone structure.
[0126] Solution 2: For multiple input cut front - side views, the corresponding heatmap can be expressed by the following formula:
[0127] I ij =(P cor -(i,j)) 2 , where (0 < i < hight, 0 < j < widht)
[0128]
[0129]
[0130] P cor represents the pixel coordinates of the 2D anchor point, and I ij represents the pixel value of the heatmap coordinate (i,j); hight represents the height of the image, and widht represents the width of the image; PE is the position encoding function, which encodes each position value into 2k feature values. d model is a hyperparameter that controls the frequency, and in this paper, it is 1000.
[0131] Input the following: 1) multiple cut front and side views (2D sub - images); 2) position - encoding vectors into the transformer network model to obtain the output: image information of [256, 256, 2]. Use the image - chunking strategy of vit (Vision Transformer) to chunk the input image X into [16, 32, 32, 32], where 16 represents the feature dimension and [32, 32, 32] represents the spatial information.
[0132] Use position embedding to perform position encoding on the 2D anchor points: X = X + Pe;
[0133] Take the image after position embedding as the input and input it into the transfomer network model, and then use the down - sampling and up - sampling strategies of unetr++. First, compress the information to [128, 8, 8, 8] through the encoder, and then perform up - sampling to finally obtain 3D segmentation information of [8, 128, 128, 128]. Calculate the loss function by comparing the segmented 3D bone structure output by the model with the ground truth of the 3D bone structure.
[0134] It should be noted that for the technical solution provided in this embodiment, how to construct a neural network model for training is not the innovation point of the present invention. Figure 8 The given neural network model is just an example. Neural network models with other structures, as long as they can implement the technical solution of this embodiment, can be applied to the above - mentioned steps S11 and S13, and finally obtain a high - precision 3D image reconstruction model.
[0135] Embodiment 2
[0136] Figure 9 and Figure 10 is a flowchart of a 3D image reconstruction model prediction method shown according to an exemplary embodiment, as Figure 9 and Figure 10 shown. The method includes:
[0137] Step S21: Obtain 2D X - ray images of the front and side views of the part to be detected;
[0138] Step S22: Predict 3D anchor points on the 3D segmentation image based on the approximate projection of 2D anchor points on the 2D X - ray image, and then obtain the corrected 2D anchor points based on the 3D anchor points, so that the 3D anchor points correspond one - to - one with the corrected 2D anchor points;
[0139] Step S23: Use the 2D X-ray image and the position encoding of the corrected 2D anchor points as inputs, and input them into the 3D image reconstruction model trained according to the above method for model prediction to obtain the 3D segmentation images of each bone joint;
[0140] Step S24: According to the 3D anchor points on the 3D segmentation images, stitch the 3D segmentation images output by the model to obtain the image reconstruction result of the overall part to be detected.
[0141] It should be noted that, in specific practice, the technical solution provided in this embodiment runs in the controller of the medical device, or is loaded and runs in the electronic device connected to the controller. The controller of the medical device executes the corresponding method by calling the program stored in the electronic device.
[0142] It can be understood that, for the technical solution provided in this embodiment, the 3D image reconstruction model trained by the above method can restore the 2D X-ray images of the anterior-posterior and lateral views of the part to be detected to the 3D image of the bone joint area to be treated. Since, before the model training of the 3D image reconstruction model, by establishing a one-to-one correspondence between the 2D anchor points on the 2D X-ray and the 3D anchor points on the 3D segmentation images, the bone structures on the complex and large-structured 2D X-ray images are split into simple 3D segmentation structures, and finally the accurate restoration of the 3D image is realized.
[0143] In a possible implementation, see Figure 5 , in step S22, predicting the 3D anchor points on the 3D segmentation images based on the approximate projection of the 2D anchor points on the 2D X-ray image, and then obtaining the corrected 2D anchor points based on the 3D anchor points, including:
[0144] Predicting 2D candidate key points based on the key point detection algorithm on the anterior-posterior and lateral 2D X-ray images Is a 1*6 vector, and the six values in the vector respectively represent the anterior-posterior 2D anchor point coordinates and the lateral 2D anchor point coordinates, expressed as: [anterior-posterior (x, y, 1), lateral (x, y, 1)];
[0145] In order to make the anterior-posterior 2D anchor point coordinates and the lateral 2D anchor point coordinates correspond to a unique 3D anchor point, the following steps are performed, including:
[0146] First select As the center, a set PP with a radius less than R
[0147]
[0148] Using the pseudo-inverse of the affine matrix to project Into 3D space to obtain
[0149] Among them, Pinv is the matrix pseudo-inverse;
[0150] Because the projection is singular, re-project the 3D image to obtain
[0151]
[0152] The point obtained by re-projection relative to the original with the minimum distance is determined as the modified 2D anchor point:
[0153] It can be understood that the 2D X-ray image includes a frontal view and a lateral view. Ideally, the 2D anchor points on the frontal view and the 2D anchor points on the lateral view should be mapped to the same 3D anchor point on the 3D segmentation image. However, in specific practice, due to errors, the 2D anchor points on the frontal view and the 2D anchor points on the lateral view often map to a 3D anchor point respectively, which will result in a non-unique correspondence between the 2D anchor points and the 3D anchor points. To avoid this situation, the technical solution provided in this embodiment predicts the 3D anchor points on the 3D segmentation image based on the approximate projection of the 2D anchor points on the 2D X-ray image, and then obtains the modified 2D anchor points based on the 3D anchor points, so as to ensure a one-to-one correspondence between the 3D anchor points and the modified 2D anchor points, and improve the accuracy of the model image reconstruction result.
[0154] Embodiment 3
[0155] Figure 11 is a schematic block diagram of a 3D image reconstruction model training device 100 shown according to an exemplary embodiment, as Figure 11 shown. The device 100 includes:
[0156] A cutting module 101 that cuts the 2D X-ray image into multiple 2D sub-images of the front and side views;
[0157] A building module 102 for establishing the correspondence between the 2D anchor points on the 2D X-ray image and the 3D anchor points on the target 3D segmentation image, and performing position normalization processing on the corresponding 3D anchor points according to the coordinates of the 2D anchor points, so that the 2D anchor points on the 2D X-ray image and the 3D anchor points on the 3D segmentation image are in one-to-one correspondence;
[0158] It is also used to establish the correspondence between the 2D X-ray image and the target 3D segmentation image according to the one-to-one correspondence between the 2D anchor points and the 3D anchor points;
[0159] A training module 103 is configured to use the position encodings of the 2D sub - graphs and 2D anchors as inputs, and the target 3D segmentation image and the 3D anchors after position normalization as the ground truth to train a pre - constructed neural network model, so that the neural network model learns the correspondence between the 2D anchors and 3D anchors, and the correspondence between the 2D sub - graphs and the 3D segmentation image until the model converges, obtaining a 3D image reconstruction model;
[0160] The output of the 3D image reconstruction model is the predicted 3D segmentation image and its 3D anchors, and the loss function of the 3D image reconstruction model is related to the ground truth and the output result of the model.
[0161] It can be understood that in the technical solution provided in this embodiment, before model training, by establishing a one - to - one correspondence between the 2D anchors on the 2D X - ray image and the 3D anchors on the target 3D segmentation image, the bone structures on the complex and large - scale 2D X - ray image are split into simple segmentation structures, so that both the 2D X - ray image and the target 3D segmentation image can be split for training (as Figure 6 shown, based on the 3D anchors, the 3D CT image can be split into 3D segmentation images of multiple single - segment bone structures; as Figure 7 shown, based on the 2D anchors, the coronal 2D X - ray image can be split into multiple pairs of 2D sub - graphs of the anteroposterior view). At the same time, different bone joints are decoupled, reducing the video memory capacity required for neural network model training and making the convergence speed faster.
[0162] In addition, since the 2D anchors and 3D anchors uniquely determine the relationship between them, data augmentation becomes easier, and finally the algorithm can train a 3D image reconstruction model with higher accuracy.
[0163] Embodiment Four
[0164] Figure 12 is a schematic block diagram of a 3D image reconstruction device 200 shown according to an exemplary embodiment. As Figure 12 shown, the device 200 includes:
[0165] An acquisition module 201 is configured to acquire 2D X - ray images of the anteroposterior view of the part to be detected;
[0166] A correction module 202 is configured to predict the 3D anchors on the 3D segmentation image based on the approximate projection of the 2D anchors on the 2D X - ray image, and then obtain the corrected 2D anchors based on the 3D anchors, so that the 3D anchors and the corrected 2D anchors are in one - to - one correspondence;
[0167] A reconstruction module 203, configured to use the 2D X-ray image and the encoded positions of the corrected 2D anchor points as inputs, and input them into a 3D image reconstruction model trained according to the above method for model prediction, to obtain 3D segmentation images of each bone joint;
[0168] It is also configured to stitch the 3D segmentation images output by the model according to the 3D anchor points on the 3D segmentation images, to obtain an image reconstruction result of the overall part to be detected.
[0169] It can be understood that, for the technical solution provided in this embodiment, the 3D image reconstruction model trained by the above method can restore the 2D X-ray images of the anteroposterior and lateral views of the part to be detected to 3D images of the bone joint area to be treated. Since, before model training, the 3D image reconstruction model established a one-to-one correspondence between the 2D anchor points on the 2D X-ray and the 3D anchor points on the 3D segmentation images, splitting the bone structures on the complex and large-structured 2D X-ray images into simple 3D segmentation structures, ultimately achieving accurate restoration of the 3D images.
[0170] Embodiment Five
[0171] Refer to Figure 13 , an electronic device shown according to an exemplary embodiment includes:
[0172] A processor 701, a communication interface 702, a memory 703, and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704;
[0173] The memory 703 is configured to store a computer program;
[0174] The processor 701 is configured to implement the above method when executing the program stored in the memory.
[0175] It can be understood that, for the technical solution provided in this embodiment, the 3D image reconstruction model trained by the above method can restore the 2D X-ray images of the anteroposterior and lateral views of the part to be detected to 3D images of the bone joint area to be treated. Since, before model training, the 3D image reconstruction model established a one-to-one correspondence between the 2D anchor points on the 2D X-ray and the 3D anchor points on the 3D segmentation images, splitting the bone structures on the complex and large-structured 2D X-ray images into simple 3D segmentation structures, ultimately achieving accurate restoration of the 3D images.
[0176] Embodiment Six
[0177] A non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the above method.
[0178] It can be understood that for the technical solution provided in this embodiment, the 3D image reconstruction model obtained by training through the above method can restore the 2D X-ray images of the anteroposterior and lateral views of the part to be detected into 3D images of the bone joint area to be treated. Since before the model training of the 3D image reconstruction model, by establishing a one-to-one correspondence between the 2D anchor points on the 2D X-ray and the 3D anchor points on the 3D segmentation image, the bone structures on the 2D X-ray images with complex and large structures are split into simple 3D segmentation structures, and finally the accurate restoration of the 3D images is realized.
[0179] Certainly, those of ordinary skill in the art can understand that all or part of the processes in implementing the above embodiment methods can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The described program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a memory, a magnetic disk, an optical disc, etc.
[0180] The specific implementation manners of the present invention described above do not constitute a limitation to the protection scope of the present invention. Any other corresponding changes and deformations made according to the technical concept of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for training a 3D image reconstruction model, characterized in that, Including: Cutting the 2D X-ray image into multiple anteroposterior and lateral 2D sub-images; Establishing the correspondence between the 2D anchor points on the 2D X-ray image and the 3D anchor points on the target 3D segmentation image, and performing position normalization on the corresponding 3D anchor points according to the coordinates of the 2D anchor points, so that the 2D anchor points on the 2D X-ray image and the 3D anchor points on the 3D segmentation image are in one-to-one correspondence; Establishing the correspondence between the 2D sub-images and the target 3D segmentation image according to the one-to-one correspondence between the 2D anchor points and the 3D anchor points; Taking the 2D sub-images and the position encoding of the 2D anchor points as inputs, and taking the target 3D segmentation image and the 3D anchor points after position normalization as the ground truth, training a pre-constructed neural network model, so that the neural network model learns the correspondence between the 2D anchor points and the 3D anchor points, and the correspondence between the 2D sub-images and the 3D segmentation image, until the model converges, to obtain a 3D image reconstruction model; The output of the 3D image reconstruction model is the predicted 3D segmentation image and its 3D anchor points, and the loss function of the 3D image reconstruction model is related to the ground truth and the output result of the model.
2. The method according to claim 1, wherein Establishing the correspondence between the 2D anchor points on the 2D X-ray image and the 3D anchor points on the target 3D segmentation image through the following steps, including: Selecting 3D anchor points on the 3D CT image; Generating 3D spheres according to the 3D anchor points; Using the digital reconstructed radiograph (DRR) algorithm to obtain the anteroposterior and lateral 2D X-ray images and anteroposterior and lateral 2D heat maps of the part to be detected according to the DRR hyperparameters and the 3D spheres; Determining the local maxima on the 2D heat map as the 2D anchor points corresponding to the 3D anchor points one by one.
3. The method according to claim 1, wherein Establishing the correspondence between the 2D anchor points on the 2D X-ray image and the 3D anchor points on the target 3D segmentation image through the following steps, including: Selecting 3D anchor points on the 3D CT image; According to the 3D anchor points, DRR hyperparameters and the following formula (1), using the affine matrix to obtain the coordinates of the 2D anchor points corresponding to the 3D anchor points, specifically: Among them, P achor3d represents the 3D anchor coordinates, which are known quantities; represents the affine matrix, which is a known quantity; P anchor2d represents the 2D anchor coordinates; Among them, P achor3d is a 1×4 vector. The first three values in the vector respectively represent the X-axis coordinate, Y-axis coordinate, and Z-axis coordinate of the 3D anchor point, and the fourth value is the constant 1; is a 4×6 affine matrix, P anchor2d is a 1×6 vector. The six values in the vector respectively represent the frontal 2D anchor point coordinates and the lateral 2D anchor point coordinates, expressed as: [frontal (x, y, 1), lateral (x, y, 1)].
4. The method according to claim 1, wherein Performing position normalization on the corresponding 3D anchor points according to the coordinates of the 2D anchor points, specifically: According to the coordinates of the 2D anchor points, determining the coordinates of the corresponding 3D anchor points on the fixed points of the corresponding 3D segmentation image.
5. A 3D image reconstruction method, characterized in that, Including: Obtaining the anteroposterior and lateral 2D X-ray images of the part to be detected; Predicting the 3D anchor points on the 3D segmentation image based on the approximate projection of the 2D anchor points on the 2D X-ray image, and then obtaining the corrected 2D anchor points based on the 3D anchor points, so that the 3D anchor points and the corrected 2D anchor points are in one-to-one correspondence; Taking the 2D X-ray image and the position encoding of the corrected 2D anchor points as inputs, and inputting them into the 3D image reconstruction model trained by the method according to any one of claims 1 to 4 for model prediction, to obtain the 3D segmentation images of each bone joint; Suturing the 3D segmentation images output by the model according to the 3D anchor points on the 3D segmentation image to obtain the overall image reconstruction result of the part to be detected.
6. The method according to claim 5, characterized in that Predicting 3D anchor points on a 3D segmentation image based on the approximate projection of 2D anchor points on a 2D X-ray image, and then obtaining corrected 2D anchor points based on the 3D anchor points, including: Predict 2D candidate key points on the 2D X-ray images in the anteroposterior and lateral views based on the key point detection algorithm It is a 1×6 vector, and the six values in the vector respectively represent the anteroposterior 2D anchor coordinates and the lateral 2D anchor coordinates, expressed as: [anteroposterior (x, y, 1), lateral (x, y, 1)]; To make the 3D anchor points corresponding to the anterior-posterior 2D anchor point coordinates and the lateral 2D anchor point coordinates unique, the following steps are performed, including: First, select The set PP centered at and with a radius less than R Project using the pseudo-inverse of the affine matrix to obtain in 3D space Among them, Pinv is the matrix pseudo-inverse; Because the projection is singular, the 3D image is reprojected to obtain The point with the minimum distance obtained by reprojection relative to the original is determined as the modified 2D anchor point:
7. A 3D image reconstruction model training device, characterized in that, Including: A cutting module that cuts the 2D X-ray image into multiple anterior-posterior and lateral 2D sub-images; A building module for establishing the correspondence between 2D anchor points on the 2D X-ray image and 3D anchor points on the target 3D segmentation image, and performing position normalization on the corresponding 3D anchor points according to the coordinates of the 2D anchor points, so that the 2D anchor points on the 2D X-ray image and the 3D anchor points on the 3D segmentation image are in one-to-one correspondence; It is also used to establish the correspondence between the 2D X-ray image and the target 3D segmentation image according to the one-to-one correspondence between the 2D anchor points and the 3D anchor points; A training module for using the 2D sub-images and the position encoding of the 2D anchor points as inputs, and the target 3D segmentation image and the 3D anchor points after position normalization as the ground truth to train a pre-constructed neural network model, so that the neural network model learns the correspondence between the 2D anchor points and the 3D anchor points, and the correspondence between the 2D sub-images and the 3D segmentation image, until the model converges to obtain a 3D image reconstruction model; The output of the 3D image reconstruction model is the predicted 3D segmentation image and its 3D anchor points, and the loss function of the 3D image reconstruction model is related to the ground truth and the output result of the model.
8. A 3D image reconstruction device, characterized in that, Including: An acquisition module for acquiring anterior-posterior and lateral 2D X-ray images of the part to be detected; A correction module for predicting 3D anchor points on a 3D segmentation image based on the approximate projection of 2D anchor points on a 2D X-ray image, and then obtaining corrected 2D anchor points based on the 3D anchor points, so that the 3D anchor points and the corrected 2D anchor points are in one-to-one correspondence; A reconstruction module for using the 2D X-ray image and the position encoding of the corrected 2D anchor points as inputs, and inputting them into the 3D image reconstruction model trained according to the method described in any one of claims 1 to 4 for model prediction to obtain the 3D segmentation image of each bone joint; It is also used to stitch the 3D segmentation images output by the model according to the 3D anchor points on the 3D segmentation image to obtain the image reconstruction result of the whole part to be detected.
9. An electronic device, characterized in that, Including: A processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor is used to implement the method described in any one of claims 1 to 6 when executing the programs stored on the memory.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method described in any one of claims 1-6.