Three-dimensional model reconstruction method and device for medical image, equipment and medium
By constructing a medical image dataset with shooting angle labels and iteratively removing images with large error indices, the motion artifact problem caused by respiratory motion in cone-beam CT scanning was solved, and high-quality 3D model reconstruction was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-13
AI Technical Summary
During cone-beam CT scanning, anatomical displacement caused by respiratory movements results in motion artifacts in the 3D model reconstruction, affecting the accuracy of clinical procedures.
By constructing a medical image dataset with shooting angle labels, iteratively performing 3D model reconstruction and projection processing, removing images with error indices greater than a preset threshold, and optimizing the dataset to eliminate motion artifacts introduced by inconsistencies in anatomical state.
Accurate and reliable 3D reconstruction results were obtained, motion artifacts were reduced, and the accuracy of clinical diagnosis and surgical guidance was improved.
Smart Images

Figure CN121661260A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical imaging technology, and in particular to a method, apparatus, device and medium for three-dimensional model reconstruction of medical images. Background Technology
[0002] Cone-beam computed tomography (CBCT) plays a vital role in clinical diagnosis and surgical guidance, and its application is typically based on the reconstruction of high-quality 3D models. These 3D models are obtained by reconstructing 2D medical images acquired from different angles, and can be further used to generate digital projection images for visualization.
[0003] However, cone-beam CT scans typically take several seconds to tens of seconds, inevitably covering multiple respiratory cycles. Respiratory motion causes displacement of anatomical structures, resulting in images acquired at different times corresponding to different anatomical states. Reconstructing using these images with inconsistent anatomical states leads to motion artifacts in the 3D model. The projected images generated from this model will also exhibit motion artifacts, just like the actual anatomical structures, thus affecting the accuracy of subsequent clinical procedures. Summary of the Invention
[0004] To overcome the aforementioned problems in the prior art, this disclosure provides a method, apparatus, device, and medium for three-dimensional model reconstruction of medical images. Specifically, this disclosure is achieved through the following technical solution: According to a first aspect of the embodiments of this specification, a method for reconstructing a three-dimensional model of a medical image is provided, comprising: A dataset of medical images is constructed, wherein the medical images are images acquired by a cone-beam CT device within a specified time period, the specified time period being the end-expiratory or end-inspiratory time period obtained based on the respiratory signal analysis of the acquired object; the medical images are labeled, the labels being used to indicate the shooting angle of the medical images; Iteratively execute the following steps: Construct a 3D model based on the dataset; For each frame of medical image in the dataset, the 3D model is projected according to the shooting angle of the medical image to generate a projected image corresponding to the medical image; Determine the error index between each frame of the medical image and the corresponding projection image; Medical images with error indices exceeding a preset threshold are removed from the dataset. If there is no error index greater than the preset threshold, the iteration ends and the 3D model is output.
[0005] According to a second aspect of the embodiments of this specification, a three-dimensional model reconstruction apparatus for medical images is provided, comprising: A dataset construction module is used to construct a dataset of medical images; the medical images are images acquired by a cone-beam CT device within a specified time period, the specified time period being the end-expiratory or end-inspiratory time period obtained based on the respiratory signal analysis of the acquired object; the medical images are labeled, the labels being used to indicate the shooting angle of the medical images; The model building module is used to build a 3D model based on the dataset; The projection processing module is used to project the 3D model onto each frame of medical image in the dataset according to the shooting angle of the medical image, and generate a projection image corresponding to the medical image. The error analysis module is used to determine the error index between each frame of the medical image and the corresponding projection image. The dataset management module is used to remove medical images corresponding to error indicators that exceed a preset threshold from the dataset. The scheduling module is used to control the sequential iteration of the model building module, projection processing module, error analysis module, and dataset management module, and to end the iteration and output the 3D model when there is no error index greater than the preset threshold.
[0006] According to a third aspect of the embodiments of this specification, an electronic device is provided, including a processor and a machine-readable storage medium storing machine-executable instructions executable by the processor, the processor being prompted by the machine-executable instructions to perform the method as described in the first aspect.
[0007] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, wherein the storage medium stores a computer program that, when executed by a processor, implements the method described in the first aspect.
[0008] This embodiment utilizes respiratory gating technology to construct a medical image dataset with shooting angle labels acquired within a specified time period. Then, iterative processes are performed to construct a 3D model and optimize the dataset. In each iteration, the 3D model is reconstructed based on the current dataset, and corresponding projected images are generated according to the shooting angle of each image. By calculating the error index between each frame of the original medical image and its corresponding projected image, medical images with error indices exceeding a preset threshold are identified and removed; these images are considered inconsistent with the overall anatomical state represented by the current model. Through iterative execution of reconstruction, projection, comparison, and removal, the consistency of the anatomical state represented by the images within the dataset is continuously optimized. Finally, when the error index corresponding to all medical images in the dataset does not exceed the preset threshold, the iteration ends and the 3D model is output. This embodiment can eliminate motion artifacts introduced by the inconsistency in the anatomical state represented by the acquired medical images in the dataset, thereby obtaining accurate and reliable 3D reconstruction results. Attached Figure Description
[0009] Figure 1 This is a schematic flowchart illustrating an exemplary embodiment of a three-dimensional model reconstruction method for medical images. Figure 2 This is a schematic diagram of the structure of the C-arm of a cone-beam CT device, as exemplarily shown in the embodiments of this specification. Figure 3 This is a schematic diagram illustrating the process of frame interpolation for a dataset, as exemplarily shown in the embodiments of this specification. Figure 4 This is a schematic diagram of the structure of the second neural network model exemplarily shown in the embodiments of this specification; Figure 5 This is a schematic diagram of a three-dimensional model reconstruction device for medical images, exemplarily shown in the embodiments of this specification. Figure 6 This is a schematic diagram illustrating the structure of an electronic device as exemplified in an embodiment of this specification. Detailed Implementation
[0010] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0011] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0012] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0013] CBCT (Cone Beam Computed Tomography) has significant applications in modern medical practice. CBCT, hereinafter referred to as cone-beam CT, can not only be used for preoperative diagnostic planning but also integrated into treatment or interventional procedures for real-time or near-real-time image guidance, providing technical support for dynamically monitoring target location and ensuring precise treatment implementation.
[0014] In cone-beam computed tomography (CBCT) applications, generating a 3D model is fundamental for subsequent image-guided imaging. Specifically, this typically involves first acquiring several medical images of the patient's target area at different shooting angles. Then, a 3D model of the target area is reconstructed based on these images. After reconstruction, the 3D model is digitally projected onto the desired angle to obtain a digitally reconstructed 2D projected image, which is then displayed. The quality of this 2D projected image—its consistency with the actual acquired medical images in terms of anatomical structure—directly reflects the geometric accuracy of the 3D model.
[0015] However, in actual clinical practice, cone-beam CT scans take several seconds to tens of seconds to complete, inevitably covering multiple respiratory cycles of the patient. Respiratory motion causes continuous and periodic displacement of internal organs and other anatomical structures during the scan. This displacement means that medical images acquired at different time points and shooting angles actually correspond to different states of three-dimensional anatomical structures. When using these inconsistent medical images for 3D reconstruction, the ideal premise upon which reconstruction algorithms are usually based—that "all projections depict the same stationary object"—no longer holds true. The reconstructed 3D model will contain motion artifacts, specifically manifested as blurring, ghosting, or stripes. Consequently, when digital simulation projection is performed based on this 3D model, the generated projected image will have motion artifacts compared to the actual state of the patient's target area, making it impossible to guarantee the accuracy of subsequent operations based on this projected image.
[0016] To address these issues, a respiratory gating method can be used to reconstruct a more accurate 3D model by selecting medical images from a specified respiratory phase. However, this approach requires ensuring that the projection data selected based on the respiratory signals is of "high quality." If inaccurate respiratory signal extraction or patient micro-movements result in the initial reconstruction dataset containing medical images inconsistent with the specified respiratory phase, the resulting 3D model will also exhibit motion artifacts, manifesting as blurring, ghosting, and striping distortions. Furthermore, involuntary movements such as the patient's heartbeat and bowel movements can also contribute to the same problems. These motion artifacts cannot be automatically identified in previous solutions, compromising the quality of the final 3D model used for clinical evaluation, which may contain hidden errors.
[0017] It's worth noting that while traditional spiral CT can also employ respiratory gating technology, the difference in imaging principles between it and cone-beam CT leads to different sensitivities to respiratory motion. Traditional spiral CT has an extremely fast single-turn scan speed; for example, a single rotation can be completed in 0.2 seconds. This allows for the rapid acquisition of all the projection data required for reconstruction, ensuring that all projection data correspond to essentially the same anatomical state and effectively suppressing motion artifacts. In contrast, cone-beam CT has a longer single-turn scan time, making it impossible to guarantee that projections at different angles are acquired from the same precise moment within the same respiratory cycle. Therefore, even with gating, the projection data may still correspond to slightly different anatomical states, making motion artifacts more likely during reconstruction. This fundamental difference presents a greater challenge for cone-beam CT in achieving high-quality 3D reconstruction.
[0018] Therefore, a quality control mechanism is urgently needed in clinical practice to ensure the consistency of anatomical structure in medical images used for 3D model reconstruction. This specification's embodiments utilize a preliminarily reconstructed 3D model to generate simulated projection images, which are then compared with each frame of medical images acquired for reconstructing the model. This comparison allows for an objective assessment of the current 3D model's ability to interpret each frame of the original medical images. Those original medical images with excessively large errors in the simulated projection indicate that their anatomical state does not match the "average" or "expected" state represented by the model. Their presence is the cause of impaired 3D model quality and motion artifacts in its projection images. These excessively erroneous original medical images are removed from the modeling dataset, and an accurate and reliable 3D model significantly reduced by respiratory motion artifacts is constructed accordingly.
[0019] Accordingly, embodiments of this specification provide a method for reconstructing a three-dimensional model of a medical image. The C-arm is used to obtain medical images based on X-rays penetrating a target human body region.
[0020] like Figure 1 As shown, Figure 1 This is a schematic flowchart illustrating a method for reconstructing a three-dimensional model of a medical image, as shown in an embodiment of this specification. The three-dimensional model reconstruction method includes: S100: Construct a dataset of medical images. The medical images are images acquired by a cone-beam CT scanner within a specified time period. The specified time period is the end-expiratory or end-inspiratory time period obtained based on the respiratory signal analysis of the acquired subject.
[0021] The medical images in this dataset were acquired using cone-beam CT equipment within a specified time period. Cone-beam CT equipment refers to an imaging device capable of using cone-beam X-ray scanning and reconstructing three-dimensional images using cone-beam CT technology. Its specific forms include, but are not limited to: DSA (digital subtraction angiography) equipment with an integrated C-arm, mobile C-arm equipment, and cone-beam CT equipment with a ring gantry structure. The subject being scanned refers to a human or animal undergoing cone-beam CT scanning. The acquired medical images target specific areas of the subject, such as the head, chest, abdomen, or limbs.
[0022] by Figure 2 Taking the C-arm shown as an example, Figure 2This is a schematic diagram of the C-arm of a cone-beam CT scanner, exemplarily illustrated in an embodiment of this specification. The C-arm includes an X-ray tube 202 and a detector 201, which together constitute the imaging system. When the focal point within the X-ray tube 202 is energized, it emits X-rays that expand in a cone shape. This X-ray beam penetrates the target area of the object being examined. As the X-rays penetrate the target area, due to the different attenuation characteristics of different internal tissues and structures, the penetrated X-ray beam carries spatial distribution information of the internal structures of the human body. The detector 201 receives this penetrated X-ray and converts it into electrical signals. These electrical signals undergo a series of subsequent processing steps to ultimately generate a two-dimensional medical image of the target area.
[0023] This specified time period is determined based on the analysis of the respiratory signals of the subject, specifically the end-expiratory or end-inspiratory phase of the respiratory cycle. This time period is chosen because, in imaging of areas such as the lungs, images corresponding to the end-inspiratory phase allow the lungs to be fully expanded, thus more clearly reflecting pathological changes within the lungs, such as tumors, inflammation, or emphysema. Therefore, such images have higher application value in clinical diagnosis. Furthermore, the range of motion of the internal anatomical structures is relatively minimal at this time, helping to reduce differences in anatomical displacement caused by respiratory movements between images, thereby providing a more consistent initial data foundation for subsequent 3D reconstruction.
[0024] It should be noted that due to the irregularity of individual breathing patterns, the duration of each respiratory cycle may constantly change. Therefore, the specific time period identified through respiratory signal analysis may not correspond precisely to the theoretical "end-expiratory period" or "end-inspiratory period." Thus, this specific time period is essentially one or more time intervals deemed suitable for image acquisition, calibrated after analyzing the respiratory signals.
[0025] In the dataset, medical images are labeled with their capture angles. Each frame of the medical image in the dataset is accompanied by corresponding label information. The label records the capture angle of the imaging system at the moment the image was acquired. Simultaneously with the acquisition of the medical image, the spatial orientation parameters of the imaging system (i.e., the capture angle) are recorded and uniquely associated with the image as a label. When constructing the dataset, the capture angle corresponding to each frame of the medical image can be automatically identified without manual matching or correction. The capture angle label precisely indicates the capture angle to which the X-ray source and detector components rotated around the object being acquired when the image was acquired. These images with capture angle labels are organized according to the acquisition order or angle order, forming the initial dataset for subsequent iterative reconstruction and quality control processes.
[0026] Iteratively execute the following steps: S102: Construct a three-dimensional model based on the dataset.
[0027] Each frame of medical image in the dataset reflects the attenuation of X-rays by various tissues within the target area along each X-ray path. Because each frame carries a precise shooting angle label, the spatial geometric relationships corresponding to that frame's projection data are clearly defined. The reconstruction algorithm utilizes these two-dimensional projection data sets from different angles with well-defined geometric locations, and through specific mathematical transformations and backprojection operations, solves for the attenuation level of each voxel in three-dimensional space, ultimately synthesizing a complete three-dimensional volumetric image.
[0028] For example, the reconstruction algorithm can employ analytical algorithms, such as the FDK algorithm. The FDK algorithm works by converting two-dimensional projection data into a three-dimensional image space. First, each frame of the two-dimensional projection image acquired by cone-beam CT scanning is preprocessed, including necessary corrections and filtering. The filtering step aims to eliminate image blurring that might result from simple backprojection, thereby recovering sharper edges and details. Subsequently, the algorithm backprojects the filtered projection data into a three-dimensional matrix in the three-dimensional image space along the corresponding X-ray path, according to the shooting angle at which it was acquired. This backprojection process is repeated at all projection angles, ultimately accumulating to form a complete three-dimensional volumetric image.
[0029] Based on the foregoing explanation, the dataset used for reconstruction at the start of the iteration is a collection of medical images corresponding to a specified time period, filtered through respiratory signal analysis. Therefore, the reconstructed 3D model should ideally reflect the average anatomical state of the target site during a relatively stable respiratory phase. However, due to potential errors in respiratory signal extraction, or the possibility of minor patient movements within the specified time period, the 3D model constructed in the first iteration may still contain motion artifacts introduced by inconsistencies in anatomical states between medical images. The 3D model constructed in the first iteration will serve as the starting point for subsequent quality verification and iterative optimization.
[0030] S104: For each frame of medical image in the dataset, the 3D model is projected according to the shooting angle of the medical image to generate a projected image corresponding to the medical image.
[0031] After constructing the 3D model for this iteration, each frame of the original medical image in the dataset is analyzed. The shooting angle label defines the spatial relationship between the cone-beam CT scanner's motion mechanism and the object being acquired during the original acquisition; specifically, using the aforementioned C-arm as an example, this could be the spatial relationship between the X-ray tube, the object being acquired, and the detector. Subsequently, using this shooting angle as a virtual imaging condition, digital X-ray projection calculations are performed on the reconstructed 3D model. This calculation process simulates the process of X-rays emanating from a virtual focal point, penetrating the 3D model, and ultimately forming a projected image on the virtual detector plane. The generated projected image is commonly referred to as a digitally reconstructed radiographic image.
[0032] S106: Determine the error index between each frame of the medical image and the corresponding projection image.
[0033] Ideally, the original acquired medical image and the simulated projection image should depict the same anatomical structure projected at the same angle. However, in practical applications, due to potential errors in respiratory signal extraction or the possibility of minute movements by the patient within a specified time period, this specification uses an error index between the medical image and the corresponding projection image to indicate their spatial alignment.
[0034] Error metrics are used to measure the similarity between two images. For example, error metrics may include calculating the cosine similarity and normalized Euclidean distance between the feature matrices of the two images.
[0035] After S106, the system determines whether there is an error index greater than a preset threshold (S108). If so, proceed to step S110: remove medical images with error indices exceeding a preset threshold from the dataset. Then, return to step S102.
[0036] The magnitude of the error index reflects the compatibility between the original medical image and the current 3D model: a smaller error index value indicates that the original medical image is closer to the model's simulated projection image, meaning the data is more consistent with other data used to build the model; conversely, a larger error index value indicates a significant deviation between the anatomical information contained in the original medical image and the "average" or "expected" state represented by the model. This deviation suggests that the original medical image may have been acquired at different respiratory phases or during unexpected movement, representing anomalous data that disrupts model consistency.
[0037] Therefore, each frame of the original medical image can be filtered using a preset threshold. The error index calculated in step S106 for each frame of the original medical image is compared with this preset threshold. If the error index of any medical image frame exceeds the preset threshold, that image frame is determined to be low-quality data or an outlier. This determination means that the anatomical structure information contained in that image frame differs significantly from the overall state represented by the current 3D model, and cannot be corrected by simple registration. Therefore, the image frame with the excessively large error index is removed from the current dataset, preventing it from participating in subsequent iterative reconstruction calculations, thus purifying the data source for modeling.
[0038] If not, execute S112: end the iteration and output the 3D model.
[0039] If there is no error index greater than the preset threshold, it means that the difference between all existing images in the dataset of the current iteration and the current 3D model is within an acceptable range. That is, the 3D model reconstructed based on the current dataset can well explain each frame of the original medical image, and the two have achieved a high degree of consistency.
[0040] Therefore, if no error index exceeding the preset threshold is found, the iteration ends and the 3D model is output. The output 3D model can be used for subsequent clinical diagnostic analysis, surgical planning, or treatment guidance. This output 3D model is the result of the above reconstruction-verification-screening iteration mechanism, thus more accurately reflecting the anatomical structure of the target site of the collected object at a specified respiratory phase, and significantly suppressing motion artifacts.
[0041] It's worth noting that the reason for performing multiple iterations is that a single reconstruction and screening process usually cannot obtain a high-quality dataset with completely consistent anatomical states in one go. The iterative mechanism allows for incremental optimization of the dataset and model. Specifically, in the first iteration, although the initial dataset used to reconstruct the 3D model has been screened for respiratory phases, it may still contain unqualified images due to respiratory signal extraction errors or slight patient movements, resulting in motion artifacts in the initially reconstructed 3D model. While the error index calculated based on this model can identify some obvious abnormalities, it may not fully expose all inconsistencies, or the error judgment may be biased due to model distortion itself. By removing images with errors exceeding a threshold, the dataset is optimized for the first time. Using the optimized dataset for the second round of reconstruction, the generated 3D model is closer to the true anatomical structure of the target respiratory phase, and its "average state" is more representative. Based on this improved model, another projection comparison can identify unqualified images in the remaining images that do not conform to this improved state and may have been previously masked, thus enabling a new round of removal. This process repeats itself, with each iteration reconstructing a more accurate 3D model based on a more consistent dataset. This more accurate model then filters out more hidden anomalies, forming a progressively converging positive feedback optimization mechanism.
[0042] This embodiment utilizes respiratory gating technology to construct a medical image dataset with shooting angle labels acquired within a specified time period. Then, iterative processes are performed to construct a 3D model and optimize the dataset. In each iteration, the 3D model is reconstructed based on the current dataset, and corresponding projected images are generated according to the shooting angle of each image. By calculating the error index between each frame of the original medical image and its corresponding projected image, medical images with error indices exceeding a preset threshold are identified and removed; these images are considered inconsistent with the overall anatomical state represented by the current model. Through iterative execution of reconstruction, projection, comparison, and removal, the consistency of the anatomical state represented by the images within the dataset is continuously optimized. Finally, when the error index corresponding to all medical images in the dataset does not exceed the preset threshold, the iteration ends and the 3D model is output. This embodiment can eliminate motion artifacts introduced by the inconsistency in the anatomical state represented by the acquired medical images in the dataset, thereby obtaining accurate and reliable 3D reconstruction results.
[0043] Based on the foregoing description, medical images are images acquired by a cone-beam CT scanner within a specified time period. This specification further provides embodiments for determining the specified time period and acquiring medical images accordingly through prospective gating. As one or more embodiments of this specification, the medical images are images acquired by the cone-beam CT scanner rotating around the object being acquired within the specified time period; the end-expiratory time period or the end-inspiratory time period is predicted based on the respiratory signal before acquiring the medical images.
[0044] In the embodiments described in this specification, the identification of a specified time period is a preliminary process for image acquisition. Specifically, before the cone-beam CT device initiates the formal scanning acquisition, it continuously monitors multiple complete respiratory cycles of the subject. This can be achieved by acquiring continuous respiratory signals through sensors in a wearable device or by analyzing pre-scan images. By analyzing the respiratory signals of these multiple respiratory cycles, the periodicity, phase, and amplitude characteristics of the subject's current breathing pattern can be obtained. Based on this, the upcoming respiratory cycle can be predicted, thereby pre-calculating the theoretical time windows for multiple future end-expiratory or end-inspiratory phases. The set of these predicted time windows is defined as the specified time period.
[0045] Subsequently, the cone-beam CT scanner controls image acquisition during the scan based on this specified time period. The cone-beam CT scanner only triggers the acquisition of medical images within the predicted specified time period to obtain the medical images used for reconstruction, while acquisition is paused outside the specified time period. Through this prediction-based, prospective gating acquisition, the cone-beam CT scanner can actively and selectively capture medical images corresponding to the phases of expected minimal respiratory motion.
[0046] In practice, the cone-beam CT scanner first predicts the end-expiratory (or end-inspiratory) time periods over multiple future cycles based on respiratory signals. Taking the end-expiratory time period as an example, at the start of the scan, the motion mechanism of the cone-beam CT scanner, such as the C-arm, enters a rotatable ready state. When an end-expiratory time period arrives, the motion mechanism rotates to the first preset imaging angle position and triggers the acquisition of medical images at that position. After acquiring images at this angle, the gantry continues to rotate to the next equally spaced imaging angle and acquires medical images. This acquisition process during rotation is repeated until the end of the end-expiratory time period. Outside of the end-expiratory time period, the motion mechanism stops moving and waits for the next predicted end-expiratory time period to arrive before repeating the above process.
[0047] Based on the foregoing, the embodiments in this specification attempt to actively constrain image acquisition activities within a predicted relatively stable respiratory phase. It should be noted that, since the duration of each respiratory cycle may continuously change, the specified time period predicted by analyzing respiratory signals and used to trigger acquisition may not correspond precisely to the theoretical "end-expiratory or end-inspiratory time period," and the predicted stable phase itself may also contain subtle movements. Therefore, the embodiments in this specification aim to initially improve the consistency of projection data in the respiratory phase dimension from the perspective of acquisition strategy, while also considering the sampling distribution in angular space, thereby providing a relatively optimized data foundation for subsequent reconstruction processes.
[0048] This specification further provides embodiments for determining a specified time period and acquiring medical images accordingly through retrospective gating. As one or more embodiments of this specification, the end-expiratory time period or the end-inspiratory time period is obtained by analyzing the respiratory signals within the imaging time period.
[0049] The medical images are images selected from a plurality of candidate medical images within the specified time period; the plurality of candidate medical images are images acquired by the cone-beam CT device within the specified shooting time period.
[0050] The definition of "specified time period" in the embodiments of this specification differs from the prospective approach. Here, the specified time period refers to multiple end-expiratory or end-inspiratory time periods actually experienced by the human body, identified after the recording of respiratory signals of the subject within a preset shooting time period. The "specified time period" is a temporal interval that has already occurred, determined after retrospective analysis of respiratory signals during the entire completed scanning period (i.e., the "shooting time period") containing multiple respiratory cycles.
[0051] Correspondingly, the method of acquiring the "medical images" used to construct the dataset also changes. In this mode, the cone-beam CT scanner first performs a complete, typically continuous rotational scan. During this entire "image acquisition time," a series of projected images within a specific angular range are acquired at preset angular intervals. These images constitute "a number of candidate medical images." Simultaneously, external sensors from a synchronous monitoring device, such as a wearable detector, ensure that the acquisition of respiratory signals is synchronized with this image acquisition process in time. After the scan, the complete respiratory signals are analyzed and all actual end-expiratory or end-inspiratory time points are located, thereby determining the corresponding "specified time period." Finally, based on the timestamps, a subset of images whose acquisition times fall within the aforementioned "specified time period" is selected from all the "candidate medical images." This selected subset of images serves as the dataset used for subsequent 3D model reconstruction.
[0052] Therefore, retrospective gating can be summarized as "acquiring all data first, then analyzing and filtering," avoiding the mechanical control complexities of waiting for specific respiratory phases during scanning. However, it relies on accurate analysis of respiratory signals after acquisition and precise time synchronization between images and signals. This method aims to filter images with relatively consistent respiratory states from all projection data, providing a foundation for subsequent reconstruction.
[0053] Based on the foregoing embodiments, prospective gating places higher demands on device motion control. To achieve multi-angle acquisition within a stable timeframe, the C-arm typically needs to move within a predicted specified time period. This requires control logic involving "rotation acquisition-waiting within a specified time period," as well as accelerating / decelerating rotation within the specified time period to achieve equally spaced sampling and stopping movement, ensuring that medical images with uniform angular distribution are acquired within the short specified timeframe. Retrospective gating, on the other hand, has relatively simpler requirements for device motion control. The device can typically perform continuous rotation acquisition at preset uniform speed and equal angular intervals, without needing to start / stop or change speed due to respiratory phases, resulting in more conventional and stable motion control. This equal angular interval is typically set to 1° to 2°, and this specification does not impose this limitation.
[0054] Regarding scan duration, prospective gating, since it only performs acquisition within a specified time period, theoretically only acquires medical images within that time period. Therefore, it typically only requires rotating around the subject once (i.e., 360°) to acquire a sufficient number of medical images with uniform angular intervals. Retrospective gating, on the other hand, requires acquiring medical images over multiple complete respiratory cycles to ensure sufficient data for subsequent screening. To obtain uniform angular sampling from the screened data, it usually requires multiple consecutive rotations to accumulate the chance of each angle being acquired at a stable phase across multiple cycles. During this process, the subject experiences a relatively higher radiation dose. In both methods, the number of rotations around the subject or the total rotation angle can be set according to the actual situation.
[0055] As one or more embodiments of this specification, the medical images are arranged in the dataset according to the order of the shooting angles. To meet the basic requirements of reconstruction algorithms for the continuity and uniformity of projection data spatial sampling, the medical images in the dataset are arranged according to the order of their shooting angles. This order typically corresponds to the increasing or decreasing angles during the rotational scanning process of a cone-beam CT device, thereby forming an image set with a clear sequence in angular space.
[0056] In the process of acquiring images using respiratory gating, prospective gating only triggers acquisition during specific respiratory phases, which may introduce errors due to device movement. Conversely, retrospective gating can result in irregular patient breathing leading to uneven distribution of effective acquisition moments, both of which could result in excessively large shooting angle intervals between adjacent frames in the final dataset. To address this issue, the steps prior to constructing the 3D model also include: For any two adjacent medical images, the difference in the shooting angles of the two adjacent medical images is determined. If the difference is greater than a preset angle difference threshold, the dataset is subjected to frame interpolation to obtain an optimized dataset.
[0057] Excessive angular intervals mean that there is a lack of sufficient projection data within a certain angular range. This can lead to insufficient information in the reconstruction algorithm for that region, potentially introducing undersampling artifacts and reducing the geometric accuracy and image clarity of the 3D model.
[0058] This specification's embodiments introduce a method to enhance data completeness in the reconstruction process. Specifically, the dataset, already arranged by shooting angle, can be sequentially traversed. For any two adjacent medical images, the difference between their shooting angles is calculated. Then, this difference is compared with a preset angle difference threshold. This angle difference threshold is set based on the maximum angular interval allowed by the reconstruction algorithm theory, aiming to ensure the required angular sampling density for reconstruction.
[0059] When the difference in shooting angle between two adjacent medical images is detected to be greater than a preset angle difference threshold, it is determined that there is a sampling gap in this angle range. At this time, frame interpolation processing of the dataset is triggered. The purpose of frame interpolation processing is to re-acquire or generate one or more new interpolated frame images corresponding to the intermediate angles between the angles corresponding to the two original images, and insert them into the original order of the dataset to fill the gaps in angle sampling. This makes the distribution of projected angles of the entire dataset more uniform and continuous, thereby providing a data foundation that meets the requirements of the algorithm's sampling theorem for subsequent reconstruction steps, suppressing reconstruction artifacts caused by insufficient angle sampling, and improving the quality and reliability of the final 3D model.
[0060] Based on the foregoing description, medical images can be re-acquired or generated as interpolated images in a dataset. However, re-acquiring images requires operator intervention, increases the total examination time, and the subjects being examined need to accumulate a higher radiation dose. Therefore, the interpolation method for generating medical images has advantages such as high efficiency, no additional radiation, and a high degree of automation. As one or more embodiments of this specification, Figure 3 This is a schematic diagram illustrating the frame interpolation process of a dataset, as exemplarily shown in the embodiments of this specification. The frame interpolation process of the dataset includes: S301: Determine at least one target shooting angle that needs to be interpolated based on the difference; For each target shooting angle, the following steps are performed iteratively: S302: Construct a reference 3D model based on the dataset; S303: Project the reference 3D model according to the target shooting angle to generate a reference projection image; S304: Optimize the visual features of the reference projection image using a neural network to obtain an interpolated image; S305: Add the interpolated image to the dataset; S306: Determine the error index between the interpolated image and the reference projection image; then perform a judgment on whether the error index is greater than the preset interpolation error threshold (S307). If so, return to step S302; If not, execute S308: End iteration.
[0061] When the projection angle sampling interval is identified as too large, the embodiments in this specification do not supplement the data by controlling the CT equipment to perform physical acquisition again. Instead, they use image generation technology based on existing data to synthesize the required interpolated images and supplement them into the dataset.
[0062] Specifically, the frame interpolation process begins by determining the angle positions that need to be filled. First, the shooting angle sequence of all images in the dataset is analyzed to locate the intervals where the angle interval between adjacent images exceeds the angle difference threshold. Then, one or more specific angle values that meet the uniform sampling requirements and are missing within these intervals are calculated. These angle values are defined as the target shooting angles that need to be interpolated.
[0063] For each target shooting angle, an iterative optimization process is executed to generate high-quality interpolated images. In each iteration, a 3D model is first reconstructed based on all images in the current dataset as a reference 3D model for this iteration. Subsequently, the reference 3D model is projected according to the current target shooting angle using a digital ray projection method to generate a simulated image at that angle, i.e., the reference projected image. This image reflects the theoretical projection that the current 3D model "should" have at that angle. The above process is similar to the implementation of steps S102 and S104, and will not be explained in detail here.
[0064] Next, the reference projection image is optimized to improve its consistency with the real captured image in terms of details such as texture, noise, and grayscale, thereby obtaining an interpolated image that more closely approximates the actual physical capture effect. The optimization process can employ, for example, a deep learning model. Using the reference projection image as input, and based on the pre-trained mapping relationship between the input and output images, an optimized, more realistic interpolated image is output.
[0065] After obtaining the optimized interpolated image, it is temporarily added to the dataset, and its quality is immediately evaluated. The evaluation method compares this newly generated interpolated image with its source (i.e., the reference projection image in this iteration) and calculates the error metric between the two. This error metric can include mean squared error (MSE), structural similarity (SSIM), etc., to measure whether the optimization process improves the realism of the image while maintaining consistency with the anatomical structure represented by the 3D model constructed in this iteration.
[0066] The quality assessment can be implemented by setting a preset interpolation error threshold as the quality standard. If the error index is not greater than the preset threshold, the iteration ends. If the calculated error index is not greater than a second preset threshold, it indicates that the generated interpolated image is of acceptable quality and meets the final requirements for inclusion in the dataset, and the iteration process for that target interpolation angle ends. If the error index exceeds the threshold, it may be necessary to adjust the optimization parameters or start a new round of iteration based on the updated dataset until an interpolated image that meets the quality requirements is generated. This iterative closed loop ensures the reliability and effectiveness of the synthesized image, thereby effectively improving the angular sampling integrity of the dataset without increasing the additional scanning burden.
[0067] Based on the aforementioned frame interpolation process, an embodiment further provides an enhanced implementation of the frame interpolation process. In the iterative process consisting of steps S102 to S110 in the aforementioned embodiment, the dataset mentioned is a collection of medical images of a single respiratory phase (end of expiration or end of inspiration); when generating the interpolated image, a reference projection image can be generated using information from a single respiratory phase, or a reference projection image can be generated by fusing information from multiple respiratory phases.
[0068] This enhanced implementation can be specifically applied between steps S302 and S303 in the iterative process as a particular method for constructing a better reference projection image. Specifically, in the iteration, step S302 can be based on multiple datasets collected retrospectively using a gating method, where the medical images in each dataset were acquired by a cone-beam CT device at different specified time periods, corresponding to different respiratory phases. For example, the first dataset contains images acquired during the end-inspiratory phase, and the second dataset contains images acquired during the end-expiratory phase. Based on these different datasets, multiple reference 3D models can be independently reconstructed, each model representing a specific respiratory phase state.
[0069] Based on this, multiple reference 3D models representing different respiratory phases can be combined. Specifically, taking the construction of the master dataset in steps S102 to S110, which relies on multiple "end-expiratory time periods," as an example, these time periods are discretely distributed in the scanning sequence and have time intervals between them. Although they all belong to the end-expiratory phase, due to the incomplete regularity of respiration and the possible subtle state differences between different cycles, relying solely on these discrete end-expiratory projections to reconstruct the reference 3D model may result in uncertainty or ambiguity in the anatomical state represented by the reference 3D model in regions with sparse angle sampling, especially in such areas.
[0070] At this point, referencing another independent dataset is highly valuable, such as an image collection containing multiple "end-inspiratory phases." End-inspiratory and end-expiratory phases represent two distinct but physiologically related extreme states in the respiratory cycle. First, a spatiotemporal registration algorithm based on metrics such as mutual information is used to spatially align the reference 3D model reconstructed from the end-expiratory dataset with the reference 3D model reconstructed from the end-inspiratory dataset. This step aims to eliminate the overall anatomical spatial offset between the two temporal models caused by respiratory motion, placing them in the same geometric coordinate system. It is understood that this embodiment can also use datasets from other temporal phases to perform the above process. After spatial alignment is completed, a joint digital preprojection operation is performed. For example, in the three-dimensional spatial domain, deformation interpolation is performed between the registered end-expiratory 3D model and the end-inspiratory 3D model based on an estimate of the current target interpolation angle's position within the respiratory cycle (possibly based on the relative position of its acquisition timestamp within the end-expiratory interval), generating a more accurate reference 3D model. Subsequently, a digital preprojection is performed on this reference 3D model, which incorporates dual-temporal information, for the current target acquisition angle.
[0071] The resulting reference projection image is based not only on discrete end-expiratory information but also incorporates the corresponding end-inspiratory state as a reference point on the continuous motion trajectory. This is equivalent to supplementing the end-expiratory model with constraint information from another temporal phase, thereby generating a more stable and anatomically accurate reference projection image. This image will serve as input to the subsequent optimization step (S304), enabling the final interpolated image to be more smoothly and reliably integrated into the dataset sequence, improving the robustness and output quality of the interpolation process under discrete temporal sampling conditions.
[0072] Furthermore, regarding step S304, the process of optimizing the reference projection image using a neural network includes: The target shooting angle and the reference projection image are input into the trained first neural network model to obtain the interpolated image. The trained first neural network model is trained through several first samples, which include input sample images and output sample images. The input sample image is a projection image obtained by projecting a pre-constructed sample 3D model according to a specified shooting angle. The sample 3D model is pre-constructed using images of preset sample objects obtained from several different shooting angles. The output sample image is an image of the preset sample object acquired by a cone-beam CT device at the specified shooting angle.
[0073] The embodiments in this specification optimize the reference projection image by using a pre-trained first neural network model to generate an idealized reference projection image based on the digital projection of the 3D model. This idealized reference projection image has visual characteristics and noise structure that are closer to the real physical acquisition image, thereby improving the realism and usability of the final generated interpolated image.
[0074] To train this first neural network model, embodiments of this specification construct a deep learning model as the execution of image optimization. This deep learning model may employ an architecture capable of feature extraction and detail reconstruction, such as the U-Net model with an attention mechanism, ResNet with a residual learning structure, or generative adversarial networks and their variants capable of driving highly realistic image generation. The learning objective of this deep learning model is to achieve a mapping from the input sample image to the output sample image.
[0075] To achieve this goal, deep learning models need to be trained under supervision on large-scale, high-quality sample sets. For each training sample in the sample set, a 3D model of the sample needs to be constructed first. This 3D model can be constructed by acquiring a large number of medical images of a predefined sample object at the same respiratory phase from various shooting angles, and then removing the medical images from the specified angles to obtain the 3D model constructed when the specified angle is missing. Then, a 3D model is constructed using the modeling dataset as a sample, and the projected image of this 3D model from the specified angle is obtained as the input sample image, while the previously removed medical images from the specified angle are used as the output sample images.
[0076] By training on paired images, deep learning models can gradually learn to automatically correct systematic biases in projection values in input sample images and fill in anatomical details and textures missing due to model simplification or insufficient information, ultimately outputting an interpolated image that is closer to real-world capture in both visual features and physical meaning.
[0077] As one or more embodiments of this specification, determining the error index between each frame of the medical image and the corresponding projected image includes: The feature matrices of the medical image and the projected image are obtained through the intermediate layer of the trained second neural network model. The second neural network model is trained under supervision using several medical images with segmentation mask labels as samples. The goal of the supervised training is to output a segmentation result consistent with the segmentation mask labels based on the input medical image. The intermediate layer in the second neural network model is used to output the feature matrix of the medical image. The error index is obtained based on the feature matrices of the medical image and the projected image.
[0078] This specification describes an embodiment that implements feature extraction using a pre-trained second neural network model. This second neural network model is a pre-trained medical image segmentation model, such as U-Net, ResNet, or other convolutional neural networks with encoder-decoder architectures or deep residual connections. When constructing this model, supervised training can be performed using a large number of medical image samples with anatomical structure annotations. These samples can be real images acquired by cone-beam CT equipment or simulated projection images obtained through digital projection reconstruction of a 3D model. The anatomical structure annotations are provided in the form of segmentation masks, labeled as specific organs or tissues in the image. For example, when the input image is a spinal image, the annotation includes segmentation masks containing tissue structures such as vertebral bodies and spinous processes. During training, the model learns to extract feature matrices related to anatomical structures from the input image and outputs corresponding segmentation results based on these feature matrices. In this process, the model learns the mapping relationship from the original input image to the feature matrices.
[0079] Taking U-Net as an example, Figure 4 This is a schematic diagram of the structure of a second neural network model exemplified in an embodiment of this specification. The original function of this second neural network model is to receive an input image, obtain the feature matrix of the input image through its internal intermediate layers, and finally output the corresponding segmentation result based on this feature matrix.
[0080] In this method, the input layer to the intermediate layer of the second neural network model is used as a feature extractor. In application, the medical image and the projected image are input into the model respectively, and the output generated by its intermediate layer during processing is obtained. Specifically, the intermediate layer refers to the last convolutional layer in its encoder path, and its output is the required feature matrix. These feature matrices carry key anatomical semantic information from the image.
[0081] Accordingly, based on the two feature matrices obtained from the medical image and the projection image respectively, a similarity measure between them is calculated. For example, the feature matrix is specifically a three-dimensional or four-dimensional feature matrix, and the similarity measure can be cosine similarity and / or normalized Euclidean distance. The calculated result is defined as an error index, used to quantitatively evaluate the degree of consistency between the original acquired image and the digitally reconstructed projection in terms of anatomical structure.
[0082] The embodiments in this specification also provide Figure 5 The diagram illustrates an exemplary three-dimensional model reconstruction apparatus for medical images. The three-dimensional model reconstruction apparatus includes: The dataset construction module 501 is used to construct a dataset of medical images; the medical images are images acquired by a cone-beam CT device within a specified time period, the specified time period being the end-expiratory time period or the end-inspiratory time period obtained based on the respiratory signal analysis of the acquired object; the medical images are labeled, and the labels are used to indicate the shooting angle of the medical images; Model building module 502 is used to build a 3D model based on the dataset; The projection processing module 503 is used to perform projection processing on the three-dimensional model according to the shooting angle of each medical image in the dataset, and generate a projection image corresponding to the medical image. Error analysis module 504 is used to determine the error index between each frame of the medical image and the corresponding projection image. The dataset management module 505 is used to remove medical images corresponding to error indicators that exceed a preset threshold from the dataset. The scheduling module 506 is used to control the model building module, projection processing module, error analysis module and dataset management module to execute iteratively in sequence, and to end the iteration and output the three-dimensional model when there is no error index greater than the preset threshold.
[0083] The embodiments in this specification also provide Figure 6 The diagram illustrates the structure of an exemplary electronic device. The electronic device includes a processor 601 and a machine-readable storage medium 602; the machine-readable storage medium 602 stores machine-executable instructions that can be executed by the processor 601, which in turn cause the processor 601 to perform the method shown in any of the above embodiments.
[0084] like Figure 6 As shown, at the hardware level, the electronic device includes a processor 601, a system bus 603, a network interface, memory, and a machine-readable storage medium 602, and may also include other hardware required for business operations. The processor 601 reads the corresponding computer program from non-volatile memory into memory and then runs it to implement the method shown in any of the above embodiments. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0085] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method shown in any of the above embodiments.
[0086] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0087] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0088] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0089] The embodiments of the subject matter and functional operation described in this specification can be implemented in the following ways: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, the program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof.
[0090] The processing and logic described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output.
[0091] Suitable computers for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0092] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.
[0093] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.
[0094] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0095] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0096] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for reconstructing a three-dimensional model of a medical image, characterized in that, include: A dataset of medical images is constructed, wherein the medical images are images acquired by a cone-beam CT device within a specified time period, the specified time period being the end-expiratory or end-inspiratory time period obtained based on the respiratory signal analysis of the acquired object; the medical images are labeled, the labels being used to indicate the shooting angle of the medical images; Iteratively execute the following steps: Construct a 3D model based on the dataset; For each frame of medical image in the dataset, the 3D model is projected according to the shooting angle of the medical image to generate a projected image corresponding to the medical image; Determine the error index between each frame of the medical image and the corresponding projection image; Medical images with error indices exceeding a preset threshold are removed from the dataset. If there is no error index greater than the preset threshold, the iteration ends and the 3D model is output.
2. The three-dimensional model reconstruction method according to claim 1, characterized in that, The medical image is an image acquired by the cone-beam CT device rotating around the subject within the specified time period; the end-expiratory time period or the end-inspiratory time period is predicted based on the respiratory signal before the acquisition of the medical image.
3. The three-dimensional model reconstruction method according to claim 1, characterized in that, The end-expiratory time period or the end-inspiratory time period is obtained by analyzing the respiratory signals within the shooting time period; The medical images are images selected from a plurality of candidate medical images within the specified time period; the plurality of candidate medical images are images acquired by the cone-beam CT device within the specified shooting time period.
4. The three-dimensional model reconstruction method according to claim 1, characterized in that, The medical images are arranged in the dataset in order of the shooting angle; The steps prior to constructing the 3D model also include: For any two adjacent medical images, the difference in the shooting angles of the two adjacent medical images is determined. If the difference is greater than a preset angle difference threshold, the dataset is subjected to frame interpolation to obtain an optimized dataset.
5. The three-dimensional model reconstruction method according to claim 4, characterized in that, The frame interpolation process on the dataset includes: Based on the difference, at least one target shooting angle that requires frame interpolation is determined; for each target shooting angle, the following steps are iteratively performed: Construct a reference 3D model based on the dataset; The reference 3D model is projected according to the target shooting angle to generate a reference projection image; The visual features of the reference projection image are optimized using a neural network to obtain the interpolated image; Add the interpolated image to the dataset; Determine the error index between the interpolated image and the reference projection image. If the error index is not greater than a preset interpolation error threshold, the iteration ends.
6. The three-dimensional model reconstruction method according to claim 5, characterized in that, The process of optimizing the reference projection image using a neural network includes: The target shooting angle and the reference projection image are input into the trained first neural network model to obtain interpolated images. The trained first neural network model is obtained by training with several first samples. The first samples include input sample images and output sample images. The input sample images are projection images obtained by projecting the sample 3D model according to a specified shooting angle. The sample 3D model is pre-constructed using images of preset sample objects obtained at several different shooting angles. The output sample images are images of the preset sample objects acquired by the cone-beam CT device at the specified shooting angle.
7. The three-dimensional model reconstruction method according to claim 1, characterized in that, The determination of the error index between each frame of the medical image and the corresponding projected image includes: The feature matrices of the medical image and the projected image are obtained through the intermediate layer of the trained second neural network model. The second neural network model is trained under supervision using several medical images with segmentation mask labels as samples. The goal of the supervised training is to output a segmentation result consistent with the segmentation mask labels based on the input medical image. The intermediate layer in the second neural network model is used to output the feature matrix of the medical image. The error index is obtained based on the feature matrices of the medical image and the projected image.
8. A three-dimensional model reconstruction device for medical images, characterized in that, include: A dataset construction module is used to construct a dataset of medical images; the medical images are images acquired by a cone-beam CT device within a specified time period, the specified time period being the end-expiratory or end-inspiratory time period obtained based on the respiratory signal analysis of the acquired object; the medical images are labeled, the labels being used to indicate the shooting angle of the medical images; The model building module is used to build a 3D model based on the dataset; The projection processing module is used to project the 3D model onto each frame of medical image in the dataset according to the shooting angle of the medical image, and generate a projection image corresponding to the medical image. The error analysis module is used to determine the error index between each frame of the medical image and the corresponding projection image. The dataset management module is used to remove medical images corresponding to error indicators that exceed a preset threshold from the dataset. The scheduling module is used to control the sequential iteration of the model building module, projection processing module, error analysis module, and dataset management module, and to end the iteration and output the 3D model when there is no error index greater than the preset threshold.
9. An electronic device, characterized in that, The method includes a processor and a machine-readable storage medium storing machine-executable instructions that can be executed by the processor, the processor being prompted by the machine-executable instructions to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Complex fabric surface three-dimensional reconstruction system and method under non-single visual angle
CN110415332A
Method and device for generating three-dimensional CT (Computed Tomography) image
CN117710573A
Techniques for Suppression of Motion Artifacts in Medical Imaging
US20170156690A1
Apparatus and method for removing breathing motion artifacts in CT scans
WO2019183562A1
Method and device for testing rendering results
WO2023068817A1