Reference marking system, three-dimensional reconstruction method and electronic device
By using markers with different non-coplanar features in X-ray and natural light imaging for pose estimation, the problem of insufficient accuracy and robustness in 3D reconstruction in existing technologies is solved, and high-precision 3D spatial information extraction and reconstruction is achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LI YAO
- Filing Date
- 2025-09-17
- Publication Date
- 2026-04-23
AI Technical Summary
Existing deep learning-based 3D reconstruction methods for X-ray images lack accuracy and robustness in the absence of clear reference points, making it particularly difficult to achieve accurate 3D reconstruction on ordinary X-ray machines.
A reference marking system is adopted, which includes at least four non-coplanar markers with different target features and full-view projection specificity. Pose estimation is performed through deep learning and geometric algorithms to provide accurate spatial position and orientation information.
It enables high-precision extraction and reconstruction of three-dimensional spatial information from a single or sparse angle in X-ray and natural light imaging, improving spatial positioning accuracy and reconstruction efficiency.
Smart Images

Figure CN2025121961_23042026_PF_FP_ABST
Abstract
Description
Reference marker system, 3D reconstruction method and electronic equipment
[0001] Relevant publicly available cross-references
[0002] This disclosure claims priority to Chinese Patent Application No. 2024114543990, filed on October 17, 2024, entitled "Reference Marking System, Three-Dimensional Reconstruction Method and Electronic Device", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of three-dimensional reconstruction technology, and in particular to a reference marker system, a three-dimensional reconstruction method, and an electronic device. Background Technology
[0004] In recent years, significant progress has been made in deep learning-based 3D reconstruction technology of X-ray images. Existing deep learning-based 3D reconstruction methods for X-ray images include methods that use convolutional neural networks to reconstruct 3D models from single X-ray images and deep learning algorithms that use X-ray images of mammalian skulls to reconstruct 3D CT images. These methods mainly rely on prior knowledge of human anatomy, and the accuracy and robustness of the reconstruction still face challenges in the absence of clear reference points.
[0005] In the field of machine vision, Fiducial Marker systems such as ArUco (Augmented Reality Uniform Coordinate Object) markers are widely used in camera pose estimation and augmented reality. Fiducial Marker systems are primarily used for localization and calibration; they are typically markers placed in fixed locations to help sensors or cameras determine their positional relationship relative to these markers. ArUco markers are a class of Fiducial Marker systems composed of a series of black and white squares. However, these systems are mainly designed for natural light imaging and are not suitable for penetrating imaging modalities such as X-rays. Recently, Cai et al. proposed a method called X-Gaussian, which significantly improves the accuracy and speed of X-ray 3D reconstruction by introducing 3D Gaussian point clouds and radiation intensity modeling mechanisms. However, like previous methods, the X-Gaussian method still assumes that the camera pose information is known and accurate. This assumption limits these methods to applications only on devices that can provide accurate pose information, such as CT scanners. For ordinary X-ray machines, due to the lack of accurate pose information, these methods are difficult to apply directly. Therefore, achieving accurate 3D reconstruction on a wider range of X-ray imaging devices remains a challenge.
[0006] Public content
[0007] The purpose of this disclosure is to provide a reference marker system, a 3D reconstruction method, and an electronic device to provide accurate spatial position and attitude information, thereby improving spatial positioning accuracy and reconstruction efficiency.
[0008] In a first aspect, this disclosure provides a reference marking system, which includes marking components;
[0009] The marking component includes at least four non-coplanar markers with different target features. The target features of the markers have projection stability. The marking component has projection specificity from all angles, and the two-dimensional projection image of the marking component corresponds one-to-one with its three-dimensional spatial pose.
[0010] Optionally, the target features include one or more of relative size, color, grayscale, and surface texture.
[0011] Optionally, the marker component has a rotation center for providing the projection determination point, the rotation center being used as the origin of the normalized projected coordinate system, and the actual size of the marker being used to provide normalization parameters for the normalized projected coordinate system.
[0012] Optionally, the center of rotation may include a sphere or intersecting line segments.
[0013] Optionally, the reference marker system is used for X-ray imaging, and the marker assembly is a non-coplanar rigid structure consisting of at least four spheres of different sizes, the spheres being made of a material that is X-ray radiable but has a certain degree of permeability.
[0014] Optionally, the reference marking system is applied to natural light imaging. The marking component is a non-coplanar rigid structure composed of at least four spheres of different colors. Each sphere is fixed by a preset material. The transparency of the preset material reaches a preset transparency threshold and the refractive distortion rate is less than a preset distortion rate threshold.
[0015] Optionally, there may be multiple marker components.
[0016] Secondly, this disclosure also provides a three-dimensional reconstruction method based on the reference marker system of the first aspect, comprising:
[0017] Acquire a target projection image captured by a target imaging device on a test scene in which a reference marker system is placed;
[0018] The pose of the target imaging device is estimated based on the target projection image to obtain the pose information of the target imaging device.
[0019] Based on the pose information of the target imaging device, the three-dimensional reconstruction of the object to be imaged in the scene under test is performed on the target projection image to obtain the three-dimensional reconstruction result.
[0020] Optionally, the marking component has a rotation center for providing the projection determination point; the pose of the target imaging device is estimated based on the target projection image to obtain the pose information of the target imaging device, including:
[0021] Feature extraction is performed on the target projection image to obtain the original projection position information of each marker and the rotation center, as well as the projection size information of the reference marker; wherein, the reference marker is one of the markers.
[0022] Based on the original projection position information of the rotation center and the projection size information and actual size information of the reference markers, the original projection position information of each marker is normalized in terms of coordinate system and size to obtain the target projection position information of each marker in the normalized projection coordinate system; where the origin of the normalized projection coordinate system is the rotation center.
[0023] Based on the target projection position information of each marker, the pose information of the target imaging device is calculated by a preset pose estimation algorithm; wherein, the pose estimation algorithm includes deep learning algorithm and / or geometric algorithm.
[0024] Thirdly, this disclosure also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the three-dimensional reconstruction method of the second aspect.
[0025] Fourthly, this disclosure also provides a computer-readable storage medium on which a computer program is stored, the computer program being executed by a processor to perform the three-dimensional reconstruction method of the second aspect.
[0026] The reference marker system, 3D reconstruction method, and electronic device provided in this disclosure include a marker component. The marker component comprises at least four non-coplanar markers with distinct target features, the target features of which exhibit projection stability. The marker component possesses projection specificity across all viewing angles, and the two-dimensional projected image of the marker component corresponds one-to-one with its three-dimensional spatial pose. This reference marker system is applicable to X-ray imaging and natural light imaging. Because the target features of the markers in the marker component of this reference marker system exhibit projection stability, accurate identification of the projection positions of each marker can be achieved, thereby providing accurate spatial position and pose information for 3D reconstruction. Furthermore, due to the projection specificity of the marker component across all viewing angles, and the one-to-one correspondence between the two-dimensional projected image of the marker component and its three-dimensional spatial pose, accurate spatial pose information can be provided, improving spatial positioning accuracy and reconstruction efficiency. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0028] Figure 1 is a structural design diagram of a reference marker system suitable for natural light imaging provided in an embodiment of this disclosure;
[0029] Figure 2 is a projection diagram of some attitudes of the reference marking system shown in Figure 1;
[0030] Figure 3 shows the deep learning results of the spatial pose corresponding to the reference labeling system shown in Figure 1;
[0031] Figure 4 shows the prediction accuracy of the reference marking system shown in Figure 1.
[0032] Figure 5 is a structural design diagram of a reference marker system suitable for X-ray imaging provided in an embodiment of this disclosure;
[0033] Figure 6 is a projection diagram of some attitudes of the reference marking system shown in Figure 5;
[0034] Figure 7 is a structural design diagram of another reference marker system suitable for X-ray imaging provided in an embodiment of this disclosure;
[0035] Figure 8 is a projection diagram of some attitudes of the reference marking system shown in Figure 7;
[0036] Figure 9 shows the deep learning results of the spatial pose of a reference marker system suitable for X-ray imaging provided in an embodiment of this disclosure;
[0037] Figure 10 is a prediction accuracy diagram corresponding to a reference marker system suitable for X-ray imaging provided in an embodiment of this disclosure;
[0038] Figure 11 is a flowchart illustrating a three-dimensional reconstruction method provided in an embodiment of this disclosure;
[0039] Figure 12 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0040] The technical solutions of this disclosure will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of this disclosure, not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0041] Currently, there is a lack of a unified spatial reference marker system applicable to all viewing angles in X-ray imaging and natural light imaging. Existing polyhedral structure markers based on surface QR codes have limitations in all-view applications and X-ray imaging. In particular, in X-ray imaging, there is a lack of standardized reference objects that can provide accurate spatial position and attitude information, which limits the ability to perform accurate 3D reconstruction from sparse angle X-ray images. Based on this, the reference marker system, 3D reconstruction method, and electronic device provided in this disclosure are applicable to multimodal imaging of X-ray imaging and natural light imaging. When the reference marker system based on the combination of non-coplanar feature points with all-view projection specificity is applied to 3D reconstruction, it can achieve high-precision extraction and reconstruction of 3D spatial information from a single or sparse angle.
[0042] To facilitate understanding of this embodiment, a reference marking system disclosed in this disclosure will first be described in detail.
[0043] This disclosure provides a reference marking system, which includes a marking component. The marking component includes at least four non-coplanar markers with different target features. The target features of the markers have projection stability. The marking component has projection specificity from all angles, and the two-dimensional projection image of the marking component corresponds one-to-one with its three-dimensional spatial pose.
[0044] The aforementioned reference marker system can be applied to X-ray imaging, natural light imaging, and multimodal applications combining natural light and X-rays, as long as the target features of each marker are different and the projection stability is ensured under the corresponding imaging conditions.
[0045] Because the target features of the markers in the marker components of the reference marker system possess projection stability, accurate identification of the projected positions of each marker can be achieved, thus providing accurate spatial position and pose information for 3D reconstruction. Furthermore, due to the projection specificity of the marker components across all viewing angles, the 2D projected image of each marker component corresponds one-to-one with its 3D spatial pose, ensuring the uniqueness of PnP (especially EPnP) solutions. This allows for high-precision 3D reconstruction from a single or sparse perspective, improving spatial positioning accuracy and reconstruction efficiency. Here, PnP stands for Perspective-n-Point, a method for solving the correspondence between 3D and 2D points; EPnP stands for Efficient Perspective-n-Point, an efficient algorithm for solving the PnP problem.
[0046] Optionally, the aforementioned marking components can be one or more. When the reference marking system includes multiple marking components, the placement angles of each marking component can be different, which can reduce unpredictable situations caused by occlusion between marking objects within the marking components; multiple marking components with different placement angles can refer to and correct each other, improving the robustness and accuracy of the reference marking system.
[0047] Optionally, the target features mentioned above may include one or more of relative size, color, grayscale, and surface texture. Color, grayscale, and surface texture can all be used in reference marking systems for natural light imaging, while relative size and grayscale (i.e., X-ray reproducibility) can both be used in reference marking systems for X-ray imaging. Therefore, relative size can be applied to both X-ray imaging and natural light imaging. Specifically, target features can be spheres of different sizes, markers of different colors, areas with different grayscale values, or special textures / patterns, or any combination of these features.
[0048] It should be noted that the embodiments disclosed herein do not limit the geometry of the target feature and the marker (i.e., the marker is not limited to a sphere). In other embodiments, the target feature may be other features, and the marker may be other geometric shapes.
[0049] When there are multiple marker components, these components can be identical, distinct, or partially identical. Within different marker components, the relative positions and / or target features of the markers differ. Differences in target features include different types and / or different feature values. For example, two marker components might choose relative size and color as target features, respectively; or two marker components might both choose relative size as target features, but the sizes of their markers are different.
[0050] Optionally, the aforementioned marking component has a rotation center for providing the projection determination point, the rotation center serving as the origin of the normalized projected coordinate system, and the actual size of the marker providing normalization parameters for the normalized projected coordinate system. By normalizing the two-dimensional projected image of the marking component to the normalized projected coordinate system, scaling resistance and translation resistance can be achieved, thus giving the reference marking system scaling resistance and translation resistance.
[0051] Optionally, the aforementioned center of rotation may include a sphere or intersecting line segments. The sphere may be a marker in the marking component, or it may not be a marker. The intersecting line segments may consist of at least two perpendicularly intersecting line segments, or they may consist of at least two non-perpendicularly intersecting line segments. It should be noted that this disclosure does not limit the type of object used for the center of rotation; that is, the center of rotation is not limited to a sphere or intersecting line segments, as long as it is easily identifiable from its two-dimensional projection image.
[0052] The rotation center of the marker component can be a structure that provides a projection determination point, such as a sphere or intersecting line segments. This structure can be used as the origin of the projection coordinates. By combining the projection size information of the marker (such as the diameter of the projection circle corresponding to the spherical marker) and the actual size information, the projection coordinates can be normalized, making the projection coordinates resistant to scaling and translation.
[0053] When there are multiple marker components, each marker component normalizes its own projection and performs its own pose estimation.
[0054] Optionally, when the above-described reference marking system is applied to X-ray imaging, the marking assembly can be a non-coplanar rigid structure composed of at least four spheres of different sizes. The spheres are made of a material that can be visualized by X-rays but has a certain degree of transmittance. Certain transmittance means that the transmittance is within a preset transmittance range; the transmittance range can be set according to actual needs and is not limited here, for example, a transmittance range of 20% to 80%.
[0055] In one possible implementation, the reference marker system consists of at least five spheres of different sizes, one of which serves as a rotation center and provides a reference for normalizing the projected coordinates. The spheres are connected by connecting rods made of X-ray low-transparency material, while the spheres are made of X-ray transmissible but permeable material to maintain distinguishability in cases of overlapping and obscuring.
[0056] In another possible implementation, the reference marker system consists of at least four spheres of different sizes, each fixed by a link; the intersection of the links serves as the rotation center of the reference marker system and is designated as the origin of the coordinate system, which, together with the diameter of the specified sphere, provides a reference for the normalization of the projected coordinates; the link is made of an X-ray morphing material, and the spheres are made of an X-ray morphing material with a certain degree of permeability to facilitate computer vision recognition.
[0057] Optionally, when the aforementioned reference marking system is applied to natural light imaging, the marking component can be a non-coplanar rigid structure composed of at least four spheres of different colors. Each sphere is fixed by a preset material, the transparency of which reaches a preset transparency threshold and the refractive distortion rate is less than a preset distortion rate threshold. Both the transparency threshold and the distortion rate threshold can be set according to actual needs and are not limited here; for example, the transparency threshold could be 95% and the distortion rate threshold could be 20%.
[0058] Optionally, the preset material can be polycarbonate, which has the characteristics of high transparency and low refractive distortion, and can achieve a rigid structure while maintaining visibility from all angles.
[0059] In one possible implementation, the reference marker system consists of four non-coplanar spheres of different colors forming a rigid structure. These spheres are arranged non-coplanarly in three-dimensional space to ensure projection specificity across all viewing angles. A sphere of a different color is placed at the center of rotation as an additional marker, and its two-dimensional projection center serves as the origin of the coordinate system. Together with the projection diameter, it provides a reference for normalizing the projection coordinates.
[0060] When using the above-mentioned reference marker system for 3D reconstruction, the following steps may be included: a) placing the reference marker system near the object to be imaged; b) acquiring one or more 2D projection images containing the reference marker system and the object to be imaged; c) identifying and extracting features of the reference marker system from the 2D projection images; d) determining the 3D pose of the reference marker system based on the extracted features using a pre-trained deep learning algorithm or a geometric algorithm such as PnP; e) using the determined 3D pose information of the reference marker system as a spatial reference to derive the pose information of imaging devices such as cameras, and calculating translation information in combination with the pose information. The translation information and pose information together provide the extrinsic parameter matrix of the camera in 3D reconstruction to reconstruct the 3D structure of the object to be imaged.
[0061] A key feature of this disclosure is the use of the rotational attitude of the reference marker system about its own rotational center to predict the rotational attitude of the reference marker system about its own rotational center. Specifically, we first perform pre-training and geometric calculations on the rotation of the reference marker system. Then, using a transformation matrix, we can equivalently convert the rotational attitude of the reference marker system into the rotational attitude of the imaging device. This method not only simplifies the training process but also improves the flexibility and accuracy of the system in practical applications.
[0062] To facilitate understanding, the relationship between the rotation of the reference marker system and the rotation of the imaging device around the reference marker system is described below.
[0063] The two rotation scenarios are as follows: a. The rotation of the reference marker system around its own rotation center (used for pre-training and geometric calculation); b. The rotation of the imaging device around the rotation center of the reference marker system (real-world application scenario).
[0064] Definitions: Rm is the rotation matrix of the reference marker system relative to its initial position; Rc is the rotation matrix of the imaging device relative to its initial position; T is the transformation matrix from the coordinate system of the reference marker system to the coordinate system of the imaging device.
[0065] Therefore, the conversion from reference marker system rotation to imaging device rotation can be expressed by the following formula:
[0066] Rc=T×Rm×T -1 ;
[0067] Among them, T -1 It is the inverse matrix of T.
[0068] For continuous rotation, at two times t1 and t2, the rotation change of the marker is ΔRm. Then, the corresponding camera rotation change ΔRc can be expressed as: ΔRc=Rc(t2)×Rc(t1) -1 =(T×Rm(t2)×T^ -1 )×(T×Rm(t1)×T -1 ) -1 .
[0069] The equivalence is explained as follows:
[0070] The equivalence between the rotation of the reference marker system and the rotation of the imaging device is based on the following key points:
[0071] a) Relative motion: In a visual system, the visual effects produced by the rotation of the reference marker system and the rotation of the imaging device are relative.
[0072] b) Coordinate transformation: Through appropriate coordinate transformations (such as T and T in the above formula), -1 This allows for transformations between the reference marker system coordinate system and the imaging device coordinate system. This enables us to convert a rotational representation in one system to an equivalent representation in another.
[0073] c) Invariance: The fundamental properties of rotation (such as angles and axes) remain unchanged during this transformation. This means that we can use the rotation of the reference marking system to accurately infer the rotation of the imaging device.
[0074] d) Computational simplification: Generally, it is easier to simulate and compute a rotating object observed by a fixed imaging device than a fixed object observed by a moving imaging device. Therefore, we can use data from the rotation of the reference marker system to train the model and then apply it to actual imaging device motion scenarios.
[0075] e) Generalization ability: This equivalence allows the system to generalize better. While training uses data from a reference-labeled system rotation, the system can be equally applied to situations involving imaging device motion, increasing its flexibility and applicability.
[0076] By leveraging this equivalence, the system can be trained and developed in a simplified environment while still being effectively applied to real-world imaging device motion scenarios, which is a significant advantage of this approach.
[0077] This rotational attitude conversion method has the following advantages: ① It simplifies the pre-training process; ② It improves the system's adaptability and versatility; ③ It reduces computational complexity in practical applications. Potential applications: Augmented reality and robot vision, etc.
[0078] This disclosure relates to a novel reference marker system for multimodal imaging and its application method. The system comprises at least four non-coplanar, feature-distinct region points, forming a spatial structure with full-view projection specificity. Through its two-dimensional projection in imaging, combined with computer vision and deep learning algorithms, the position, size, and pose information of the marker in three-dimensional space can be accurately inferred. This disclosure provides a high-precision spatial reference for X-ray films and natural light images, reducing the computational load of pose estimation and improving its accuracy. The pose information estimated based on this method can be applied to multiple fields such as medical imaging, industrial non-destructive testing, and augmented reality, helping to improve spatial positioning accuracy and reconstruction efficiency; it can also be used in scenarios such as 3D-2D re-registration, surgical navigation, and equipment calibration.
[0079] To facilitate understanding, the above reference marking system and its applications will be described in detail below.
[0080] 1. Natural Light Imaging Reference Marker System
[0081] Design Description: a) The marker component consists of four non-coplanar spheres of different colors, forming a rigid structure. A sphere of a different color is placed at the center of rotation as an additional marker, its 2D projection center serving as the origin, together with the projection diameter, to provide a normalized reference for the 2D projection coordinates (or other designs such as combinations of vertical line segments can be used to provide the rotation center location). b) These spheres are non-coplanar in 3D space to ensure projection specificity across all viewing angles. c) To achieve a rigid structure while maintaining full-angle visibility, polycarbonate can be used to fix the spheres. Polycarbonate's high transparency and low refractive distortion make it ideal for this application. d) To improve the system's robustness, two or more marker components placed at a preset angle can be used to avoid simultaneous occlusion. e) In the projection diagram, the center sphere's center serves as the origin, and the sphere's diameter acts as a scale to provide a normalized projection coordinate system. This design is resistant to scaling and translation, and can be used for deep learning or geometric algorithms to calculate rotational attitude.
[0082] Figure 1 shows a structural design diagram of a reference marker system suitable for natural light imaging. Markers A, B, C, and D indicate four non-coplanar spheres, and sphere O is located at the rotation center. Each of the five spheres (A, B, C, D, and O) has a different color. Figure 2 shows projection diagrams of the reference marker system shown in Figure 1 under various poses. It is worth noting that the spheres in each projection diagram in Figure 2 also have different colors in the actual scene; however, due to mapping requirements, their actual colors are not shown here. Figure 3 shows the deep learning results for the spatial poses corresponding to the reference marker system shown in Figure 1, and Figure 4 shows the prediction accuracy diagram for the reference marker system shown in Figure 1.
[0083] Figure 3 shows the model loss curves corresponding to the training loss and validation loss under the natural light imaging reference label system. As shown in Figure 3, both the training loss and validation loss decrease rapidly and stabilize after about 20 training epochs. This indicates that the model learns well and there is no obvious overfitting or underfitting.
[0084] Figure 4 shows a comparison between predicted and actual values under the natural light imaging reference marker system. As shown in Figure 4, the actual prediction results of the model highly overlap with the ideal prediction results (i.e., perfectly accurate predictions), with most points concentrated near the diagonal. This indicates that the actual predictions of the model are very close to the ideal predictions, and the prediction accuracy is very high.
[0085] In summary, Figures 3 and 4 illustrate the training process and prediction results of a deep learning model for projection and pose prediction. The loss curves show that the model converges well, while the comparison between predicted and true values visually demonstrates the high accuracy of the model's predictions. The high degree of overlap between the actual and ideal prediction results indicates that the model performs excellently on projection and pose prediction tasks, providing accurate and reliable support for the geometric calibration of natural light imaging systems.
[0086] One possible implementation steps of the aforementioned natural light imaging reference marker system are as follows: a) Fabricating the marker components of the reference marker system: Using five spheres of different colors, fix them into a rigid structure using polycarbonate material. b) Placing the reference marker system: Placing one or more marker components in the scene to be tested (containing the object to be imaged). c) Image acquisition: Capturing scene images containing the reference marker system using a camera or other natural light imaging device. d) Feature extraction: Identifying and extracting the color and position information of the spheres from the image, where color is used to distinguish each sphere, and position information can be the planar coordinate information of the sphere's center projection. e) Pose estimation: Using a pre-trained deep learning model or geometric algorithms such as EPnP, calculating the camera's pose information (such as a rotation matrix) relative to the markers based on the extracted features. Translation information, such as translation vectors, can also be further calculated based on the pose information. f) Application: Using the calculated pose information, translation information, and other pose information for applications such as augmented reality and robot navigation.
[0087] 2. X-ray imaging reference marker system
[0088] Considering the unique characteristics of X-ray imaging, especially the fact that C-arm X-ray machines mainly rotate around the head / tail (Cranial / Caudal) axis in clinical applications, two X-ray imaging reference marker systems were designed.
[0089] Design Scheme 1: a) The reference marker system consists of at least five spheres of different sizes, with one sphere serving as the center of rotation to provide a reference for coordinate system normalization. b) The connecting rods use a low-X-ray radiopaque material, while the spheres use a material that is X-ray radiopaque but has some translucency to maintain distinguishability under overlapping and obscuring conditions. Figure 5 shows the structural design of the reference marker system under this scheme, and Figure 6 is a projection diagram of some orientations of the reference marker system shown in Figure 5.
[0090] Design Scheme 2: a) The main structure of the reference marker system is constructed using X-ray-developable connecting rods. b) The reference marker system includes at least four spheres of different sizes to provide dimensional scales and additional spatial reference points. c) The intersection points of the connecting rods in the projection are identified using computer vision algorithms and designated as the origin of the coordinate system. These intersections, along with the diameters of the specified spheres, provide a reference for normalizing the projected coordinates. d) The connecting rods are made of X-ray-developable material. Figure 7 shows the structural design of the reference marker system under this scheme, and Figure 8 shows the projection diagram of the reference marker system shown in Figure 7 in some poses.
[0091] Figure 9 shows the model loss curves corresponding to the training loss and validation loss under the X-ray imaging reference marker system, and Figure 10 shows the comparison between the predicted and actual values under the X-ray imaging reference marker system. Figures 9 and 10 illustrate the training process and prediction results of a deep learning model for X-ray projection and pose. The loss curves show that the model converges well, while the comparison between the predicted and actual values intuitively demonstrates the high accuracy of the model's predictions.
[0092] As shown in Figure 10, the actual prediction results of the model highly overlap with the ideal prediction results, with most points concentrated near the diagonal. This indicates that the actual predictions of the model are very close to the ideal predictions, and the prediction accuracy is very high. The model performs excellently in projection and attitude prediction tasks and can provide accurate and reliable support for the geometric calibration of X-ray imaging systems.
[0093] One possible implementation steps of the above-mentioned X-ray imaging reference marker system are as follows: a) Fabricating the marker system:
[0094] For Option 1: A connecting rod is made of carbon fiber reinforced polymer (CFRP) to precisely fix the sphere to the carbon fiber structure (i.e., the connecting rod), forming the designed non-coplanar rigid structure. CFRP has the characteristics of good X-ray transmittance, high strength, and lightweight, and is almost invisible on X-ray images, avoiding interference with sphere identification. Spheres of different sizes are made using titanium alloys (such as Ti-6Al-4V) or tungsten alloys. These materials have moderate radioactivity under X-rays, do not completely block X-rays, and also have good biocompatibility and corrosion resistance.
[0095] For Option 2: The connecting rod uses X-ray reproducible material for easy computer vision recognition. Spheres of different sizes are also made using titanium alloy (such as Ti-6Al-4V) or tungsten alloy.
[0096] b) Placement of a reference marker system: Taking orthopedic surgery (such as long bone fracture reduction and internal or external fixation surgery) as an example, the reference marker system is fixed in an appropriate position near the object to be imaged before or during surgery, or fixed to the bone cortex with bone screws, etc., to ensure that the long axis is consistent with the head-to-tail axis and the relative position with the anatomical structure is fixed. The lightweight design of the reference marker system does not impose an additional burden on the object to be imaged. c) X-ray imaging: X-ray images containing the reference marker system are acquired using, for example, a C-arm X-ray machine. d) Feature extraction: The positional information of spheres of different sizes is identified and extracted from the X-ray images (i.e., the geometric center in the projection is extracted, which can be achieved by using efficient geometric calculation methods such as Epnp, or by calculating each pixel in the entire projection image or by deep learning). Even in the case of partial overlap, the spheres remain identifiable due to the properties of the material. The rotation center point combined with the scale information provided by the sphere diameter is used as a normalization reference, making the coordinate information obtained from the projection image resistant to scaling and translation. e) Pose estimation: Using pre-trained deep learning algorithms and / or geometric algorithms such as Epnp, the pose information (i.e., attitude and translation information) of the X-ray camera is calculated based on the extracted features. f) Applications: The calculated pose information is used in medical applications such as surgical navigation and radiotherapy positioning to improve surgical accuracy and patient safety. It can also be applied to industrial non-destructive testing, CT geometric calibration, etc.
[0097] The above design fully considers the characteristics of X-ray imaging, providing a high-precision and robust space reference system for medical and other industrial, as well as scientific research applications. Through careful selection of materials and structural design, the embodiments of this disclosure significantly improve the identification accuracy and reliability under various X-ray imaging conditions, providing an important tool for precise medical diagnosis and industrial applications.
[0098] This disclosure also provides a three-dimensional reconstruction method based on the above-described reference marker system, which can be executed by an electronic device with data processing capabilities. Referring to Figure 11, a flowchart of a three-dimensional reconstruction method is shown, which mainly includes the following steps S1110 to S1130:
[0099] Step S1110: Obtain the target projection image captured by the target imaging device for the test scene where the reference marker system is placed.
[0100] The aforementioned test scene includes an object to be imaged, which can be a human body or other objects requiring three-dimensional imaging. The target projection image can be a two-dimensional projection image under natural light or a two-dimensional projection image under X-rays.
[0101] Step S1120: Estimate the pose of the target imaging device based on the target projection image to obtain the pose information of the target imaging device.
[0102] Step S1130: Based on the pose information of the target imaging device, perform three-dimensional reconstruction of the scene to be tested on the target projection image to obtain the three-dimensional reconstruction result.
[0103] Optionally, the marking component described above has a rotation center for providing the projection determination point; based on this, step S1120 may include the following sub-steps 1121 to 1123:
[0104] Sub-step 1121: Extract features from the target projection image to obtain the original projection position information of each marker and the rotation center, as well as the projection size information of the reference marker; wherein, the reference marker is one of the markers.
[0105] Sub-step 1122: Based on the original projection position information of the rotation center and the projection size information and actual size information of the reference markers, the original projection position information of each marker is normalized in coordinate system and size to obtain the target projection position information of each marker in the normalized projection coordinate system; wherein, the origin of the normalized projection coordinate system is the rotation center.
[0106] Sub-step 1123: Based on the target projection position information of each marker, calculate the pose information of the target imaging device using a preset pose estimation algorithm; wherein, the pose estimation algorithm includes a deep learning algorithm and / or a geometric algorithm.
[0107] In 3D reconstruction, pose estimation of the imaging device includes the calculation of position changes (i.e., translation vectors) and rotational attitude (rotational attitude can be obtained and labeled for each monocular vision image). Optionally, the rotational attitude is first calculated using deep learning algorithms and / or geometric algorithms (requiring the use of coordinate information after coordinate system normalization and size normalization), and then the translation vector is calculated based on the obtained rotational attitude (using only the coordinate information after size normalization).
[0108] The basic principle of pose estimation is to use the position information of a marker in a two-dimensional projection image to infer its position and pose in three-dimensional space. This disclosure provides various methods to achieve this goal:
[0109] 1. Deep learning methods: a) Extract features from 2D projected images using convolutional neural networks (CNNs) to accurately locate the 2D coordinates of markers. b) Utilize a pre-trained multilayer perceptron (MLP) network to map these 2D coordinates to position and pose parameters in 3D space. c) Optimize the network by minimizing the error between the predicted 3D parameters and the true parameters.
[0110] 2. Geometric computation methods (such as the EPnP algorithm): a) Extract the 2D coordinates of the markers from the projected image. b) Establish the correspondence between the 2D projected image coordinates and the 3D world coordinates. c) Use the EPnP algorithm to solve the perspective-n-point (PnP) problem to calculate the position and orientation of the imaging device. d) Improve the estimation accuracy through iterative optimization.
[0111] 3. A hybrid approach combining deep learning and geometric computation: a) Using deep learning networks to extract and refine the 2D coordinates of markers. b) Inputting these high-precision 2D coordinates into a geometric algorithm (such as EPnP) for pose estimation. c) Fine-tuning and optimizing the pose estimation results of the geometric algorithm using a deep learning network. d) Through end-to-end training, enabling the deep learning network to learn geometric constraints and improve overall accuracy. This hybrid approach combines the adaptability of deep learning with the precision of geometric algorithms, maintaining high accuracy under various imaging conditions while providing interpretability and robustness.
[0112] For calculating the rotational attitude under the deep learning model: First, the target features of the markers are used to identify the markers. After identification, they are labeled as, for example, A, B, C, D, and the two-dimensional projected coordinates of the center position of the markers are recorded. These two-dimensional projected coordinates are input into the deep learning model after coordinate system normalization and size normalization. The output of the deep learning model is the rotational attitude of the reference marker system around its own rotation center (represented by rotation matrix, Euler angles, or quaternions). Then, the rotational attitude of the imaging device is converted into the equivalent rotation.
[0113] The above process mainly includes the following steps:
[0114] A) Size Normalization: First, the position of the marker is normalized using its projected size in the 2D projection image. This step aims to eliminate differences in projected size caused by distance variations, making subsequent calculations more accurate. Taking a sphere as an example, the specific implementation is as follows: 1. The actual size of the sphere (i.e., its true diameter) is known; 2. Measure the projected diameter of the specified sphere in the 2D projection image; 3. Using the projected diameter as a scale, and taking the projection of the center of the sphere at the rotation center in the 2D projection image as the origin of the coordinate system, the coordinates of each sphere on the 2D image plane are normalized.
[0115] B) Target Feature Processing: Target features (such as color information) of markers are used to distinguish and identify different markers. This is crucial for correctly matching 3D world coordinates and 2D image coordinates. Taking the color of a sphere as an example, the specific implementation is as follows: Perform color analysis on the 2D projection image to identify the color of each sphere; match the identified colors with predefined color identifiers.
[0116] For the EPNP algorithm, taking a sphere as the marker as an example, the pose estimation process can be as follows:
[0117] 1. Input: a) 3D point set: The position of the sphere in the world coordinate system (a preset reference coordinate system for the reference marking system) with the rotation center as the origin (i.e., the three-dimensional coordinate information of each sphere in the world coordinate system under the reference attitude, which is known in advance). b) 2D point set: The normalized position of the sphere on the two-dimensional image plane. c) Intrinsic parameter matrix of the imaging device: Describes the internal characteristics of the imaging device (such as focal length, principal point, etc.).
[0118] 2. EPNP Algorithm Steps: a) Select 4 control points to represent all 3D points. b) Represent the 3D points as the centroid coordinates of these 4 control points. c) Establish a system of linear equations (related to the intrinsic parameter matrix of the imaging device) to describe the relationship between the 3D points and their 2D projections. d) Solve this system of linear equations to obtain the positions of the control points in the imaging device coordinate system. e) Recover the rotation and translation of the imaging device from the positions of the control points.
[0119] 3. Output: The rotation matrix R and translation vector t of the imaging device relative to the world coordinate system. The rotation matrix R describes the orientation of the imaging device; the translation vector t describes its position. The output of the EPnP algorithm is the precise orientation of the imaging device.
[0120] Alternatively, taking a sphere as an example, the translation vector between two two-dimensional projected images can be calculated as follows:
[0121] 1. Normalize the size of the two two-dimensional projection images: Use the diameter of the reference sphere as a scale to unify the scale of the two two-dimensional projection images.
[0122] 2. Calculate the 2D coordinate changes of the rotation center in the normalized image: Let the normalized coordinates of the rotation center in the first image be (x1, y1), and the normalized coordinates of the rotation center in the second image be (x2, y2). Then the coordinate changes are: Δx = x2 - x1, Δy = y2 - y2. 1 .
[0123] 3. Calculate the translation vector by combining the attitude change of the imaging device: Let the rotation matrix between the two shots be R, and the position of the rotation center in the world coordinate system be P(X,Y,Z). Then the translation vector T can be solved by the following equation: R×P+T=P+(Δx,Δy,Δz); where Δz can be estimated by other information (such as focal length change), or assumed to be 0 in some application scenarios.
[0124] 4. Solve for the translation vector: T = (Δx, Δy, Δz) - (RI) × P, where I is a 3 × 3 identity matrix.
[0125] This disclosure not only significantly improves the accuracy of reconstructing 3D structures from single or sparse angle images, but also provides X-ray imaging with functionality similar to ArUco markers in natural light images. This is of great value for medical procedures requiring precise spatial positioning, CT geometric calibration, industrial inspection, and augmented reality applications. Furthermore, because this system requires only a small amount of projection for precise positioning and reconstruction in X-ray imaging, it enables the development of low-cost, portable 3D imaging devices based on existing X-ray systems, allowing for rapid adoption in clinical practice, especially in areas with limited medical resources.
[0126] As shown in Figure 12, an electronic device 1200 provided in this embodiment includes: a processor 1201, a memory 1202 and a bus. The memory 1202 stores a computer program that can run on the processor 1201. When the electronic device 1200 is running, the processor 1201 and the memory 1202 communicate through the bus. The processor 1201 executes the computer program to realize the above-mentioned three-dimensional reconstruction method.
[0127] Optionally, the memory 1202 and processor 1201 described above can be general-purpose memory and processor, without specific limitations.
[0128] In one possible implementation, the aforementioned electronic device 1200 includes an input terminal, a host computer, and a display device; wherein, the input terminal is used to input projected images captured by the imaging device, such as natural light images obtained by natural light imaging or X-ray images obtained by X-ray imaging; the host computer is used to execute the aforementioned computer program to realize the three-dimensional reconstruction of the object to be imaged, and output the three-dimensional reconstructed image to the display device; the display device may be a display screen, an AR (augmented reality) device, or a VR (virtual reality) device, etc., for displaying the three-dimensional reconstructed image.
[0129] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the three-dimensional reconstruction method described in the preceding method embodiments. The computer-readable storage medium includes various media capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), RAM, magnetic disk, or optical disk.
[0130] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0131] In all examples shown and described herein, any specific values should be interpreted as merely exemplary and not as limitations; therefore, other examples of exemplary embodiments may have different values.
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this disclosure. Industrial applicability
[0134] The above scheme enables precise identification of the projected positions of each marker, thereby improving the accuracy of pose prediction. Based on the predicted pose information, high-precision 3D reconstruction can be performed from a single or sparse angle, improving spatial positioning accuracy and reconstruction efficiency. Furthermore, the predicted pose information can also be used in 3D-2D re-registration, surgical navigation, equipment calibration, and other scenarios, showing broad research prospects.
Claims
1. A reference mark system, characterized in that The reference marking system includes marking components; The marking component includes at least four non-coplanar markers with different target features, and the target features of the markers have projection stability; The marking component has projection specificity across the entire field of view, and the two-dimensional projection image of the marking component corresponds one-to-one with its three-dimensional spatial pose.
2. The reference mark system of claim 1, wherein, The target features include one or more of the following: relative size, color, grayscale, and surface texture.
3. A reference mark system according to claim 1 or 2, characterized in that The marking component has a rotation center for providing a projection determination point, the rotation center serving as the origin of a normalized projected coordinate system, and the actual size of the marker is used to provide normalization parameters for the normalized projected coordinate system.
4. The reference mark system of claim 3, wherein, The center of rotation may be a sphere or an intersecting line segment.
5. A reference mark system according to any one of claims 1-4, characterized in that The reference marker system is used for X-ray imaging. The marker assembly is a non-coplanar rigid structure composed of at least four spheres of different sizes. The spheres are made of a material that can be visualized by X-rays but has a certain degree of transmittance. The certain degree of transmittance refers to the transmittance being within a preset transmittance range.
6. The reference mark system of claim 5, wherein, The reference marker system consists of at least five spheres of different sizes, one of which serves as the center of rotation and provides a reference for the normalization of projected coordinates. The spheres are connected by connecting rods made of X-ray low-transparency material, while the spheres are made of X-ray transmissible material with a certain degree of permeability.
7. A reference mark system according to claim 5 or 6, characterized in that The reference marking system consists of at least four spheres of different sizes, each sphere being fixed by a connecting rod; the intersection of the connecting rods serves as the rotation center of the reference marking system. The rotation center is set as the origin of the coordinate system, and the origin of the coordinate system and the diameter of the specified sphere together provide a reference for the normalization of the projected coordinates; the connecting rod is made of X-ray imaging material, and the sphere is made of X-ray imaging material with a certain degree of permeability.
8. The reference mark system according to any one of claims 1 to 7, wherein The reference marking system is applied to natural light imaging. The marking component is a non-coplanar rigid structure composed of at least four spheres of different colors. Each sphere is fixed by a preset material. The transparency of the preset material reaches a preset transparency threshold and the refractive distortion rate is less than a preset distortion rate threshold.
9. The reference marking system according to claim 8, characterized in that, The reference marking system consists of four non-coplanar spheres of different colors, forming a rigid structure, with each sphere arranged non-coplanarly in three-dimensional space. A sphere of a different color is placed at the center of rotation as an additional marker. The center of rotation's two-dimensional projection circle serves as the origin of the coordinate system, and together with the projection diameter, it provides a reference for the normalization of the projected coordinates.
10. The reference marking system according to claim 8, characterized in that, The preset material is polycarbonate.
11. The reference marking system according to any one of claims 1-10, characterized in that, The tagging components are multiple.
12. The reference marking system according to claim 11, characterized in that, The placement angles of each marker component are different.
13. The reference marking system according to claim 11 or 12, characterized in that, The marker components may be the same or different, or partially the same; In different marker components, the relative positions and / or target features of each marker are different; among them, the different target features include different types of target features and / or different feature values.
14. The reference marking system according to any one of claims 11-13, characterized in that, Each marker component normalizes its own projection and performs its own pose estimation.
15. A three-dimensional reconstruction method based on the reference marker system according to any one of claims 1-14, characterized in that, include: Acquire a target projection image captured by the target imaging device on the test scene where the reference marker system is placed; Based on the target projection image, the pose of the target imaging device is estimated to obtain the pose information of the target imaging device; Based on the pose information of the target imaging device, the target projection image is used to perform three-dimensional reconstruction of the object to be imaged in the scene under test, and the three-dimensional reconstruction result is obtained.
16. The three-dimensional reconstruction method according to claim 15, characterized in that, The marking component has a rotation center for providing a projection determination point; the step of estimating the pose of the target imaging device based on the target projection image to obtain the pose information of the target imaging device includes: Feature extraction is performed on the target projection image to obtain the original projection position information corresponding to each of the markers and the rotation center, as well as the projection size information of the reference marker; wherein, the reference marker is one of the markers: Based on the original projection position information of the rotation center and the projection size information and actual size information of the reference marker, the original projection position information of each marker is normalized in coordinate system and size to obtain the target projection position information of each marker in the normalized projection coordinate system; wherein, the origin of the normalized projection coordinate system is the rotation center. Based on the target projection position information of each of the markers, the pose information of the target imaging device is calculated by a preset pose estimation algorithm; wherein, the pose estimation algorithm includes a deep learning algorithm and / or a geometric algorithm.
17. The three-dimensional reconstruction method according to claim 16, characterized in that, The step of calculating the pose information of the target imaging device using a preset pose estimation algorithm includes: The rotational orientation is determined using deep learning algorithms and / or geometric algorithms. The translation vector is determined based on the rotational orientation.
18. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the three-dimensional reconstruction method according to any one of claims 15-17.
Citation Information
Patent Citations
Method of determining the position of an object using projections of markers or struts
CN105051786A
Marker positioning method, geometric spacing measuring method and device
CN113888664A
Reference marking system, three-dimensional reconstruction method and electronic equipment
CN119399274A
Projection system, image processing apparatus, and calibration method
US20160295184A1