A tooth three-dimensional reconstruction method and device based on multi-view geometric prior and mirror Gaussian representation

By employing a multi-view geometric estimation and mirror Gaussian representation method for 3D tooth reconstruction, the problems of unstable reconstruction and insufficient detail in tooth scenes are solved, achieving high-quality reconstruction of teeth and adjacent soft tissues.

CN122134946APending Publication Date: 2026-06-02NANJING STOMATOLOGICAL HOSPITAL

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING STOMATOLOGICAL HOSPITAL
Filing Date
2026-04-29
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing 3D tooth reconstruction techniques are unstable under conditions of high reflectivity, repetitive textures, and sparse perspectives, and they fail to adequately express the surface continuity of soft tissue areas such as the gums.

Method used

Initial point cloud and camera parameters are obtained by multi-view geometric estimation, and optimized reconstruction is performed by combining mirror Gaussian representation. Global alignment and depth loss constraints are used to suppress mirror reflection interference and improve the reconstruction effect of teeth and adjacent soft tissues.

Benefits of technology

It improves the geometric accuracy and visualization quality of teeth and adjacent soft tissues, enhances the stability and detail of reconstruction results, and is suitable for fine reconstruction of enamel and gingival regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134946A_ABST
    Figure CN122134946A_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for 3D tooth reconstruction based on multi-view geometric priors and specular Gaussian representation. The method first acquires multi-view image data containing teeth and gingiva. A multi-view geometric estimation module predicts and globally aligns dense 3D point maps to obtain an initial point cloud, camera intrinsic and extrinsic parameters, and a depth map. Based on this, a Gaussian primitive set is initialized and input into a specular Gaussian reconstruction module, where relevant parameters are combined to optimize the reconstruction. This module introduces depth and normal losses based on the color target, supervises and reduces the weight of highlight regions within the region of interest, and finally outputs the 3D reconstruction result. This invention can improve the stability and geometric accuracy of tooth reconstruction under high reflectivity and weak texture, enhance specular reflection and highlight expression capabilities, and balance the detailed reconstruction of tooth hard tissue with the geometric and appearance effects of gingival soft tissue, thereby improving the integrity of oral scene reconstruction and its clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of dental 3D reconstruction and computer vision technology, and particularly relates to a method and apparatus for dental 3D reconstruction based on multi-view geometric priors and mirror Gaussian representation. Background Technology

[0002] Three-dimensional dental models have significant applications in dental restoration, orthodontic analysis, implant planning, digital case management, and medical teaching. Current three-dimensional dental reconstruction methods typically rely on intraoral scanning equipment, structured light equipment, or multi-view geometry-based reconstruction methods. While the former can acquire relatively detailed geometric shapes, it is costly, requires skilled operation, and in some scenarios, demands high patient cooperation. The latter, although offering more flexible data acquisition, often suffers from problems such as unstable camera pose estimation, insufficient recovery of geometric details, and fragmented reconstruction of local areas under conditions of high tooth surface reflectivity, repetitive local textures, frequent occlusion, and weak textures and significant surface continuity changes in soft tissue areas such as the gums.

[0003] In recent years, neural network-based multi-view 3D reconstruction methods can directly predict consistent 3D geometric information across multiple images and recover initial point clouds and camera parameters, providing geometric priors for subsequent 3D representation optimization. The core principle of this multi-view geometric prior is to directly predict the 3D point map corresponding to each pixel in each input viewpoint as a reference coordinate system, and then use the geometric correspondence between different viewpoints to perform global alignment of multiple 3D point maps in a unified coordinate system, thereby simultaneously recovering the scene geometry and camera pose. This approach differs from the traditional multi-view geometry processing chain that relies on feature point extraction, feature matching, fundamental matrix or essential matrix estimation, triangulation, and incremental pose recovery. It can provide a more stable and dense geometric prior under conditions of high reflectivity, weak texture, and repetitive texture in tooth scenes. Meanwhile, Gaussian splashing-based 3D representation methods have attracted widespread attention due to their high rendering efficiency and strong detail representation capabilities. However, traditional Gaussian representation typically uses low-order spherical harmonic functions to describe view-dependent appearance, which is insufficient for expressing the high-frequency changes of specular reflection and anisotropic specular highlights on the tooth surface. Specular Gaussian representation, on the other hand, separates diffuse reflection components from specular reflection components and uses specular Gaussian functions to describe the high-frequency specular response that varies with the viewing direction, making it more suitable for appearance modeling of dental scenes.

[0004] Therefore, there is an urgent need to provide a 3D reconstruction scheme that can make full use of multi-view geometric priors and improve the surface continuity representation of soft tissue regions such as gingiva and dental scenes, in order to improve the geometric accuracy, surface stability and visualization quality of the reconstruction results of teeth and adjacent soft tissues. Summary of the Invention

[0005] Purpose of the invention: The technical problem to be solved by the present invention is to address the shortcomings of the prior art by providing a method and device for three-dimensional reconstruction of teeth based on multi-view geometric prior and mirror Gaussian representation. This solves the problems of unstable reconstruction, geometric inaccuracy and insufficient detail expression caused by high reflectivity, texture repetition and sparse view in the three-dimensional reconstruction of tooth scenes in the prior art, and improves the overall reconstruction effect on soft tissue areas such as gingiva.

[0006] The method includes the following steps:

[0007] Step 1: Obtain multi-view image data of the oral cavity region to be reconstructed. The oral cavity region to be reconstructed includes hard dental tissue and soft tissue adjacent to the gums. The multi-view image data includes dental images and keyframe images extracted from dental video.

[0008] Step 2: Input the multi-view image data into the multi-view geometric estimation module to predict the dense 3D point map with each input view as the reference coordinate system, and perform global alignment of the dense 3D point map based on the cross-view geometric correspondence to obtain the initial point cloud, camera intrinsic parameters, camera extrinsic parameters and depth map.

[0009] Step 3: Initialize the Gaussian set based on the initial point cloud, depth map, and normal map;

[0010] Step 4: Input the Gaussian set of Gaussian elements into the mirror Gaussian reconstruction module, and perform optimized reconstruction by combining camera intrinsic parameters, camera extrinsic parameters and depth map. The mirror Gaussian reconstruction module introduces depth loss and normal loss on the basis of color reconstruction target, and performs supervision within the oral cavity region of interest mask and performs weight reduction processing on the region corresponding to the specular mask.

[0011] Step 5: Output the 3D reconstruction results of the oral cavity region to be reconstructed.

[0012] In step 1, the multi-view image data is preprocessed, including: keyframe filtering, image cropping, color normalization, oral region of interest segmentation, tooth region segmentation, background removal, viewpoint order adjustment, highlight region detection, and highlight mask generation; wherein, the oral region of interest includes the visible area of ​​the anterior teeth and the adjacent visible gingival area.

[0013] In step 2, the multi-view geometry estimation module predicts dense 3D point maps corresponding to the viewpoints of the input image. Each pixel in the dense 3D point map corresponds to a 3D coordinate. The dense 3D point maps under different viewpoints are transformed to a unified coordinate system through cross-viewpoint geometric correspondence and global alignment to recover the scene geometry and camera pose. This method differs from traditional multi-view geometry methods that rely on manual feature extraction, feature matching, triangulation, and incremental pose solving. It can directly obtain consistent dense geometric priors across viewpoints, thereby reducing mismatches caused by high reflectivity, weak texture, and repetitive textures in teeth. Bundle adjustment refinement is performed on the globally aligned camera parameters and 3D points to improve the consistency of camera pose and sparse geometry. The depth map is obtained by projecting 3D points in a unified coordinate system.

[0014] In step 3, the initialization method of the Gaussian tuple includes: using the three-dimensional point coordinates in the initial point cloud as the center position of the Gaussian tuple; determining the scale and rotation parameters of the Gaussian tuple using the point neighborhood distribution, depth confidence and normal direction; determining the color parameters of the Gaussian tuple using image color information; and determining the opacity parameters of the Gaussian tuple using viewpoint visibility and reconstruction confidence.

[0015] In step 4, the specular Gaussian reconstruction module employs a Gaussian representation with view-dependent appearance modeling capabilities to characterize the specular reflection and anisotropic highlights of the tooth surface. Each Gaussian unit simultaneously possesses a diffuse color component and a specular reflection component that varies with the viewing direction. The changes in brightness and color response of the tooth surface caused by changes in the viewing direction are described by a specular Gaussian function. This specular Gaussian representation differs from the traditional Gaussian splashing method, which only uses low-order spherical harmonic functions to describe view-dependent appearance, and can better express the high-frequency highlights and anisotropic reflections of the tooth surface.

[0016] In step 4, the target loss function L of the mirror Gaussian reconstruction module includes pixel-level color loss. Structural similarity loss and depth loss , represented as:

[0017] ,

[0018] ,

[0019] ,

[0020] ,

[0021] in, This represents the actual color value of the i-th pixel. This represents the rendered color value of the i-th pixel. For structural similarity evaluation functions, This represents the prior depth value corresponding to the i-th pixel. Let represent the rendering depth value corresponding to the i-th pixel, and N represent the number of pixels participating in the supervision. and This is the loss weighting coefficient.

[0022] In step 4, the depth loss is used to constrain the consistency between the rendering depth and the depth map obtained in step 2, so as to enhance the geometric stability and surface restoration effect of highly reflective areas, weak texture areas and soft tissue areas such as gums; and the color reconstruction term is calculated within the oral cavity region of interest mask, and the region corresponding to the specular mask is subject to reduced weight supervision to reduce the interference of strong specular reflection on geometric optimization.

[0023] In step 4, the optimized reconstruction is performed using differentiable rendering based on the camera intrinsic and extrinsic parameters obtained in step 2, and the search space of Gaussian units is constrained by the initial point cloud, thereby reducing the initial optimization bias and improving the convergence stability under sparse view conditions. During the training process, background randomization and phased resolution optimization strategies are adopted to suppress edge artifacts and balance reconstruction speed and detail recovery.

[0024] In step 5, the three-dimensional reconstruction results include: tooth point cloud model, tooth Gaussian representation model, tooth mesh model, textured tooth surface model, tooth and gingiva joint reconstruction model, real-time interactive browsing model for anterior tooth appearance display, and visualization model for subsequent occlusal analysis, restoration design, and digital diagnosis and treatment display.

[0025] The present invention also provides a three-dimensional tooth reconstruction device based on the method described above, comprising:

[0026] The system comprises the following modules: a data acquisition module for acquiring multi-view image data of the tooth object to be reconstructed; a preprocessing module for generating a region of interest mask and a specular mask; a geometry initialization module for predicting a dense 3D point map with each input viewpoint as a reference coordinate system based on the multi-view image data, and obtaining an initial point cloud, camera intrinsic parameters, camera extrinsic parameters, and a depth map through cross-view geometric correspondence and global alignment; a Gaussian parameter initialization module for initializing a Gaussian primitive set based on the initial point cloud, camera intrinsic parameters, camera extrinsic parameters, and depth map; and a mirror Gaussian reconstruction module for optimizing the Gaussian primitive set based on color reconstruction loss and depth loss to output the 3D reconstruction result of the tooth.

[0027] Compared with the prior art, the present invention has at least the following beneficial effects:

[0028] (1) The present invention obtains the initial point cloud and camera parameters through multi-view geometric estimation, providing a stable geometric prior for subsequent Gaussian optimization; the multi-view geometric estimation is based on dense 3D point map prediction and global alignment, which is different from the traditional feature matching and triangulation process, and can reduce the ambiguity of pose solving under the conditions of high tooth reflectivity, repetitive texture and weak texture.

[0029] (2) The present invention introduces depth loss on the basis of color reconstruction target, and combines supervision in region of interest and highlight region weighting strategy, which can enhance the geometric consistency of reconstruction results, reduce the interference of specular reflection on geometric optimization, and improve the stable expression of tooth detail area and weak texture area.

[0030] (3) The present invention uses a mirror Gaussian to represent the specular reflection and anisotropic specular highlights on the tooth surface, which is different from the traditional Gaussian splashing method that uses a low-order spherical harmonic function for view-related modeling. It can better express the high-frequency specular details on the tooth surface. At the same time, through multi-view geometric prior initialization, bundle adjustment refinement and depth consistency constraints, the optimization convergence ability and reconstruction quality under the condition of limited input view are improved.

[0031] (4) This invention is not only applicable to the fine reconstruction of hard tissue areas such as tooth enamel, but also has a good geometric restoration effect and appearance continuity for soft tissue areas such as gums, which can improve the integrity of the overall oral scene reconstruction. Attached Figure Description

[0032] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention above and in other aspects will become clearer.

[0033] Figure 1 This is a schematic diagram of the overall process of the method of the present invention.

[0034] Figure 2 This is a schematic diagram comparing the reconstruction results of the tooth region using the method of the present invention and existing methods. Detailed Implementation

[0035] This invention provides a method for three-dimensional tooth reconstruction based on multi-view geometric priors and mirror Gaussian representation, such as... Figure 1 As shown, it includes the following steps:

[0036] Step 1, Data Acquisition and Preprocessing. Multi-view images of the tooth to be reconstructed and the adjacent visible gingival region are acquired from different angles using a camera, mobile phone, or oral imaging device. When the input is video, keyframes are extracted from the video to form an image set. Subsequently, a region of interest mask and a specular mask are generated for each view image. The region of interest mask is used to define the effective supervision range of the tooth and adjacent gingival region, and the specular mask is used to mark areas with strong specular reflection on the tooth surface.

[0037] In this step, the results obtained include a set of multi-view images { The preprocessing includes oral cavity region of interest masks and specular masks. Compared to direct supervision of the entire image, the above preprocessing can suppress the interference of the external oral cavity background, lip boundaries, and saturated specular highlights on geometric estimation and appearance optimization, providing stable input for subsequent multi-view geometric prior estimation and specular Gaussian optimization.

[0038] Step 2, Multi-view geometric prior estimation. Take any two images... and A multi-view geometry estimation module is input. In one embodiment, the multi-view geometry estimation module employs a shared-weight visual encoder (first branch decoder) and an information interaction decoder with cross-attention (second branch decoder) to perform joint inference on the input image pairs. take over and Finally, the overall output is represented as = ( , , , This means that the outputs are a point plot and a confidence plot in the same reference coordinate system; the outputs of the two branch regression heads are represented as follows:

[0039] ,

[0040] in, and Two input images and The corresponding dense 3D point map, and Two input images and The corresponding confidence maps are used, and both point maps are expressed in a first-person view reference coordinate system. Each pixel corresponds one-to-one with a 3D point, so image texture, pixel correspondence, and local scene geometry are uniformly encoded into the point map representation. and These represent the regression heads of the first branch and the second branch, respectively. , ..., Indicates the first branch decoder from layer 0 to layer 1. The feature sequence output by the layer; , ..., This indicates the second branch decoder from layer 0 to layer 1. The feature sequence output by the layer.

[0041] In this embodiment, the transformation relationship of the point graph between different reference coordinate systems is expressed as follows:

[0042] ,

[0043] in, This represents the point plot corresponding to the nth image in the reference coordinate system of the mth image. This represents the point plot of the nth image in its own reference coordinate system. and Let represent the camera extrinsic transformations of the m-th image and the n-th image, respectively. This represents the homogeneous coordinate mapping function.

[0044] In this embodiment, the multi-view geometric estimation module uses a confidence-weighted three-dimensional regression objective during training, which imparts different weights to easily predictable and difficult-to-predict regions, as expressed as:

[0045] ,

[0046] in, This represents the confidence-weighted loss in a three-dimensional regression. Indicates the view index. Indicates pixel index, Indicates the first The set of effective pixels in each viewpoint Indicates the first From the perspective of the first The 3D regression error corresponding to each effective pixel Indicates the first Prediction confidence of each valid pixel, This represents the weighting coefficient of the confidence regularization term.

[0047] Global alignment is further performed on the point maps obtained from different image pairs, so that multiple local point maps are represented in a unified coordinate system. The optimized form is expressed as:

[0048]

[0049] in, This represents the set of optimal point graphs after global alignment optimization. This represents the global point plot variable to be optimized. This represents the set of rigid body transformations to be optimized. This represents the set of scaling coefficients to be optimized. Representing image pairs The corresponding scaling factor, Indicates the image pair index. This represents the confidence weight of the corresponding pixel. Indicates the first The first perspective in a unified coordinate system The global dot map element corresponding to a pixel Indicating an image pair The corresponding rigid body transformation Indicating an image pair In the The local dot map element corresponding to the pixel at the Indicating the Euclidean norm

[0050] By further constraining the globally aligned dot map to a pinhole camera model, the camera intrinsic parameters , camera extrinsic parameters and depth map of each view can be recovered, where the depth map can be directly obtained from the coordinates of the three-dimensional points corresponding to the dot map at each pixel:

[0051] ,

[0052] where represents the camera intrinsic parameters of the nth view, represents the camera extrinsic parameters of the nth view, represents the depth map of the nth view, represents the depth value at the pixel coordinates ([[]] , j) in the nth view, represents the result of extracting the direction coordinates of the three-dimensional point corresponding to the dot map of the nth view at the pixel coordinates ([[]] , j).

[0053] In this step, the obtained results include: the dense dot maps and confidence maps corresponding to pairwise views, the globally consistent dot map set, the initial point cloud, the camera intrinsic parameters, the camera extrinsic parameters, and the depth maps of each view. Compared with the traditional process of first performing local feature extraction, feature matching, geometric verification, triangulation, and incremental pose recovery, the above dot map regression and global alignment method can directly establish cross-view consistent dense geometric relationships, reduce false matches under conditions of high tooth specularity, weak texture, and repetitive texture, and improve the pose stability and initialization geometric quality under the condition of few views.

[0054] Step 3, Gaussian parameter initialization. Construct an initial Gaussian basis element set based on the globally consistent point cloud and camera parameters obtained in Step 2 , where each Gaussian basis element is denoted as [[ID=6!]]and includes the center position , rotation parameter , scale , opacity and local features used for modeling the appearance of mirrors The Gaussian covariance matrix is ​​expressed as:

[0055] ,

[0056] in, Indicates the first A high-ranking dollar, Indicates the first A three-dimensional covariance matrix of Gaussian elements. Indicated by rotation parameters A defined rotation matrix, Indicated by scale parameter The resulting scale matrix, where T represents the transpose.

[0057] In this embodiment, point cloud coordinates are used as the initial Gaussian center, the spatial distribution of the point neighborhood and multi-view visibility are used to initialize scale, rotation, and opacity, and image color is used to initialize diffuse color. Gaussian units are primarily distributed in the visible anterior teeth region and adjacent visible gingival regions. This results in an initial Gaussian representation suitable for differentiable rendering optimization. Compared to schemes that rely solely on random initialization or overly sparse point cloud initialization, this approach reduces floating points in the early training stages and improves subsequent convergence stability.

[0058] Step 4, Mirror Gaussian Optimization. The Gaussian primitive set obtained in Step 3 is input into the mirror Gaussian reconstruction module for differentiable rendering and parameter updates under camera intrinsic and extrinsic constraints. For each Gaussian primitive, it is first projected onto the 2D image plane based on the 3D covariance, which is expressed as:

[0059] ,

[0060] in, Indicates the first The two-dimensional covariance matrix after Gaussian elements are projected onto the image plane The Jacobian matrix represents the projection transformation. This represents the view transformation matrix from the world coordinate system to the camera coordinate system.

[0061] The color of pixel p is obtained through Gaussian rasterization and alpha blending, and the rendering representation is as follows:

[0062] ,

[0063] in, This represents the rendered color at pixel position p, where p represents the pixel coordinates on the image plane. This represents the set of Gaussian elements that contribute to pixel p. Indicates the first Transmittance corresponding to each Gaussian element Indicates the first The contribution of each Gaussian element to the opacity of pixel p. Indicates the first The color of each high-score element.

[0064] The mirror Gaussian module uses an anisotropic spherical Gaussian field to model the view-related appearance, and the basic mirror function is expressed as:

[0065] ,

[0066] in, Let ν represent the anisotropic spherical Gaussian function, and let ν represent the input unit direction vector. , , ] represents mutually orthogonal local coordinate axes, where For the tangential axis, As the secondary tangential axis, The main lobe axis, λ and μ represent the lobe axis along the main lobe axis, respectively. Sharpness parameters along the axial direction and along Sharpness parameter in the axial direction. Indicates the amplitude parameter. Indicates the smoothing term, " represents the dot product of vectors. Indicates input direction On the tangential axis Projection in direction, Indicates input direction On the secondary tangential axis Projection in direction, This represents the natural exponential function.

[0067] In this embodiment, instead of directly using the anisotropic spherical Gaussian field to output the final color, it is first used as a mirror latent feature, and then the mirror color is obtained through a decoding function. The overall appearance is decomposed into:

[0068] ,

[0069] ,

[0070] in, Indicates the first The total color of each Gaussian element, Indicates the first The diffuse reflection color of a Gaussian element, Indicates the first The mirror color of the Gaoski element Indicates the direction of view. This represents the latent mirror feature obtained from an anisotropic spherical Gaussian field. This represents the mirror feature decoding function. The location code indicates the viewing direction. The above-mentioned specular appearance representation can more accurately express the high-frequency specular highlights and anisotropic reflections produced by the tooth surface as the viewing direction changes, while maintaining a continuous appearance representation of soft tissue areas such as the gums.

[0071] To suppress floating points caused by initial sparsity and early overfitting in real oral cavity scenes, a phased optimization strategy from low resolution to high resolution is adopted during training. The resolution scheduling is represented as follows:

[0072] ,

[0073] in, For the first The training resolution corresponding to the next iteration. This is the initial training resolution. For the final training resolution, The threshold iteration steps are increased to improve resolution. In this embodiment, the low-resolution image is first rapidly optimized, and then switched to a high-resolution stage to restore the details of the tooth edges, incisal edges, and adjacent gingival regions. At the same time, background randomization can be combined to further suppress edge artifacts caused by a fixed background.

[0074] In this embodiment, the target loss function of the mirror Gaussian reconstruction module includes pixel-level color loss. Structural similarity loss and depth loss Joint loss function Represented as:

[0075] ,

[0076] ,

[0077] ,

[0078] ,

[0079] in, This represents the actual color value of the i-th pixel. This represents the rendered color value of the i-th pixel. For structural similarity evaluation functions, This represents the prior depth value corresponding to the i-th pixel. Let represent the rendering depth value corresponding to the i-th pixel, and N represent the number of pixels participating in the supervision. and This is the loss weighting coefficient.

[0080] In this embodiment, color and depth supervision are calculated only within the region of interest in the oral cavity, and a weighted sampling or weight scaling strategy is adopted for the region corresponding to the specular mask. The above processing does not change the basic expression of the loss function, but it can reduce the disturbance of strong specular highlights on the optimization direction, thereby improving the surface stability of highly reflective areas and the geometric continuity of soft tissue areas such as the gums.

[0081] Step 5, Reconstruction Output. After completing the iterative optimization in Step 4, the 3D reconstruction results of the teeth are output. These results include: a mirror Gaussian model for real-time viewing from new perspectives, a target image rendered by perspective, depth visualization results, and point cloud models, mesh models, or textured surface models obtained by further transformation from Gaussian representation when needed. The final results enable detailed appearance reconstruction of the anterior tooth hard tissue surface and achieve good geometric restoration and appearance expression effects in the adjacent visible gingival area. These results can be used for digital restorative design, orthodontic analysis, case presentation, and clinical decision support.

[0082] This embodiment also provides a tooth 3D reconstruction device based on multi-view geometric priors and mirror Gaussian representation, including a data acquisition module, a preprocessing module, a multi-view geometric prior module, a Gaussian parameter initialization module, and a mirror Gaussian reconstruction module. The data acquisition module acquires multi-view image data of the tooth object to be reconstructed; the preprocessing module generates a region of interest mask and a specular mask; the multi-view geometric prior module outputs point maps, confidence maps, globally consistent point clouds, camera intrinsic parameters, camera extrinsic parameters, and depth maps, and can refine camera parameters using bundle adjustment; the Gaussian parameter initialization module initializes the Gaussian unit set based on the geometric priors; the mirror Gaussian reconstruction module optimizes the Gaussian units based on view-related appearance modeling, region of interest supervision, specular weighting, and depth constraints to output the tooth 3D reconstruction result. In addition to being suitable for tooth hard tissue reconstruction, the device is also suitable for joint reconstruction of adjacent gingival soft tissue regions.

[0083] In one specific embodiment, 35 multi-view images of the anterior teeth in a closed state were acquired, with an image resolution of 1920×1080. The reconstruction targets were the visible area of ​​the anterior teeth and the adjacent visible gingival area. First, oral region of interest (ROI) masks and specular masks were generated for all images, retaining only the effective supervision information corresponding to the visible area of ​​the anterior teeth and the adjacent gingival area. Then, multi-view geometric estimation was performed on pairwise combinations of images to obtain corresponding point maps and confidence maps. Global alignment was then performed to obtain initial point clouds, camera intrinsic parameters, camera extrinsic parameters, and depth maps in a unified coordinate system. If necessary, bundle adjustment refinement was performed on the globally aligned camera parameters and 3D points to further improve cross-view consistency. Next, mirror Gaussian primitives were initialized according to the geometric priors, and a two-stage training method was used for optimization: the first stage performed fast convergence at a lower resolution, and the second stage restored details at a higher resolution. During training, color and depth supervision were calculated only within the ROI of the oral cavity, and a weighting was applied to the specular mask area. Background randomization was also used to suppress edge artifacts. After training, a real-time viewable 3D reconstructed tooth model was obtained. Figure 2 As shown, this embodiment can better restore the geometric and appearance details of the incisal edge of the tooth and the adjacent gingival region in the comparison of tooth region reconstruction results. This embodiment solves the problems of pose instability, geometric fragmentation and specular distortion caused by high reflectivity, weak texture, repetitive texture and insufficient continuous texture in soft tissue areas in tooth scenes. Compared with the traditional scheme based on local feature matching and low-order spherical harmonic appearance modeling, it can more stably restore the geometric and appearance details of the incisal edge of the tooth, the labial specular region and the adjacent gingival region.

[0084] This invention provides a method and apparatus for three-dimensional tooth reconstruction based on multi-view geometric priors and mirror Gaussian representation. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A method for three-dimensional tooth reconstruction based on multi-view geometric priors and mirror Gaussian representation, characterized in that, Includes the following steps: Step 1: Obtain multi-view image data of the oral cavity region to be reconstructed. The oral cavity region to be reconstructed includes hard dental tissue and soft tissue adjacent to the gums. The multi-view image data includes dental images and keyframe images extracted from dental video. Step 2: Input the multi-view image data into the multi-view geometric estimation module to predict the dense 3D point map with each input view as the reference coordinate system, and perform global alignment of the dense 3D point map based on the cross-view geometric correspondence to obtain the initial point cloud, camera intrinsic parameters, camera extrinsic parameters and depth map. Step 3: Initialize the Gaussian set based on the initial point cloud, depth map, and normal map; Step 4: Input the Gaussian set of Gaussian elements into the mirror Gaussian reconstruction module, and perform optimized reconstruction by combining camera intrinsic parameters, camera extrinsic parameters and depth map. The mirror Gaussian reconstruction module introduces depth loss and normal loss on the basis of color reconstruction target, and performs supervision within the oral cavity region of interest mask and performs weight reduction processing on the region corresponding to the specular mask. Step 5: Output the 3D reconstruction results of the oral cavity region to be reconstructed.

2. The method according to claim 1, characterized in that, In step 1, the multi-view image data is preprocessed, including: keyframe filtering, image cropping, color normalization, oral region of interest segmentation, tooth region segmentation, background removal, viewpoint order adjustment, highlight region detection, and highlight mask generation; wherein, the oral region of interest includes the visible area of ​​the anterior teeth and the adjacent visible gingival area.

3. The method according to claim 2, characterized in that, In step 2, the multi-view geometric estimation module predicts the dense 3D point map corresponding to the viewpoint of the input image. Each pixel in the dense 3D point map corresponds to a 3D coordinate. The dense 3D point map under different viewpoints is transformed into a unified coordinate system through cross-view geometric correspondence and global alignment to restore the scene geometry and camera pose. The depth map is obtained by projecting three-dimensional points in a unified coordinate system.

4. The method according to claim 3, characterized in that, In step 3, the initialization method of the Gaussian tuple includes: using the three-dimensional point coordinates in the initial point cloud as the center position of the Gaussian tuple; determining the scale and rotation parameters of the Gaussian tuple using the point neighborhood distribution, depth confidence and normal direction; determining the color parameters of the Gaussian tuple using image color information; and determining the opacity parameters of the Gaussian tuple using viewpoint visibility and reconstruction confidence.

5. The method according to claim 4, characterized in that, In step 4, the specular Gaussian reconstruction module adopts a Gaussian representation with view-dependent appearance modeling capabilities to characterize the specular reflection and anisotropic specular highlights of the tooth surface; wherein, each Gaussian element has both a diffuse color component and a specular reflection component that varies with the viewing direction, and the changes in brightness and color response of the tooth surface caused by the change in viewing direction are described by a specular Gaussian function.

6. The method according to claim 5, characterized in that, In step 4, the target loss function L of the mirror Gaussian reconstruction module includes pixel-level color loss. Structural similarity loss and depth loss , is represented as: , , , , in, This represents the actual color value of the i-th pixel. This represents the rendered color value of the i-th pixel. For structural similarity evaluation functions, This represents the prior depth value corresponding to the i-th pixel. Let represent the rendering depth value corresponding to the i-th pixel, and N represent the number of pixels participating in the supervision. and This is the loss weighting coefficient.

7. The method according to claim 6, characterized in that, In step 4, the depth loss is used to constrain the consistency between the rendering depth and the depth map obtained in step 2, so as to enhance the geometric stability and surface restoration effect of highly reflective areas, weak texture areas and soft tissue areas such as gums; and the color reconstruction term is calculated within the oral cavity region of interest mask, and the region corresponding to the specular mask is subject to reduced weight supervision to reduce the interference of strong specular reflection on geometric optimization.

8. The method according to claim 7, characterized in that, In step 4, the optimized reconstruction is performed using differentiable rendering based on the camera intrinsic and extrinsic parameters obtained in step 2, and the search space of Gaussian units is constrained using the initial point cloud. Background randomization and phased resolution optimization strategies are employed during training.

9. The method according to claim 8, characterized in that, In step 5, the three-dimensional reconstruction results include: tooth point cloud model, tooth Gaussian representation model, tooth mesh model, textured tooth surface model, tooth and gingiva joint reconstruction model, real-time interactive browsing model for anterior tooth appearance display, and visualization model for subsequent occlusal analysis, restoration design, and digital diagnosis and treatment display.

10. A three-dimensional tooth reconstruction device based on the method described in any one of claims 1 to 9, characterized in that, include: The data acquisition module is used to acquire multi-view image data of the tooth object to be reconstructed; The preprocessing module is used to generate oral cavity region of interest masks and specular masks; The geometry initialization module is used to predict a dense 3D point map with each input viewpoint as the reference coordinate system based on multi-view image data, and obtain the initial point cloud, camera intrinsic parameters, camera extrinsic parameters, and depth map through cross-view geometric correspondence and global alignment; the Gaussian parameter initialization module is used to initialize the Gaussian unit set based on the initial point cloud, camera intrinsic parameters, camera extrinsic parameters, and depth map; the mirror Gaussian reconstruction module is used to optimize the Gaussian unit set based on color reconstruction loss and depth loss to output the 3D reconstruction result of the tooth.