A real-time positioning and guiding system and method for orthodontic bracket based on augmented reality and intraoral scanning

CN122805389APending Publication Date: 2026-09-25LANZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611024691.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

上述途径在一定程度上推动了托槽定位精度的提升,但各自存在固有局限:徒手操作受人类视觉分辨力限制,间接粘接技术流程冗长且转移过程易引入误差,而现有增强现实方案在口腔动态环境下的跟踪稳定性、配准精度和抗遮挡能力仍显不足,尚未形成从数据采集到术中引导的完整闭环系统

Benefits of technology

本发明通过构建从口内扫描、深度学习单牙分割、自动托槽位姿规划,到SLAM与点到面ICP融合配准、深度图遮挡渲染及实体托槽视觉反馈的完整闭环系统,有效解决了上述问题:将托槽定位精度提升至亚毫米级(静态配准<0.15mm),消除对实体转移托盘的依赖,减少加工成本与流程环节;通过SLAM/IMU/EKF融合及粗精两级配准策略,在动态口腔环境和频繁遮挡下保持稳定跟踪(延迟<20ms,帧率60FPS);并引入实体托槽中心与姿态的实时偏差量化反馈,当偏差达标时自动提示,减少医师目测判断依赖。由此,本发明显著提高了托槽粘接的精度、效率和一致性,降低了操作门槛和椅旁时间(平均节省约30%),实现了正畸托槽定位的全流程数字化可追溯引导。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122805389A_ABST
    Figure CN122805389A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on real-time positioning guiding system and method of dental orthodontic bracket of augmented reality and intraoral scanning, belong to the field of dental orthodontics, comprising: data acquisition unit, data processing unit and augmented reality display unit.The method obtains three-dimensional dentition model by intraoral scanning, and the tooth mark is output by single tooth segmentation of deep learning, and ideal bracket attachment point is calculated and three-dimensional guide template is generated;SLAM tracking view point pose is combined with coarse registration and point-to-plane ICP fine registration in operation, template and real-time oral point cloud are aligned, superimposed and displayed on perspective display screen, and entity bracket visual feedback is introduced to realize bonding verification.The application realizes the transformation of bracket positioning from traditional experience dependence to sub-millimeter level precision, real-time, anti-shielding digital closed-loop guidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of orthodontics, and particularly relates to a real-time positioning and guidance system and method for orthodontic brackets based on augmented reality and intraoral scanning. Background Technology

[0002] Currently, orthodontic bracket bonding and positioning mainly rely on three approaches: first, traditional manual operation, where dentists visually determine bracket positions using clinical experience and simple mechanical tools (such as positioning forceps and Boone gauges); second, digital indirect bonding technology, which uses preoperative tooth arrangement design and 3D-printed transfer trays to achieve external positioning of the brackets before overall transfer; and third, the recently emerging augmented reality-assisted navigation, which attempts to overlay preoperatively designed virtual bracket images onto real dental surfaces to provide visual guidance. While these approaches have improved bracket positioning accuracy to some extent, each has inherent limitations: manual operation is limited by human visual resolution; indirect bonding technology is lengthy and prone to introducing errors during transfer; and existing augmented reality solutions still lack sufficient tracking stability, registration accuracy, and anti-occlusion capabilities in dynamic oral environments, failing to form a complete closed-loop system from data acquisition to intraoperative guidance.

[0003] However, the aforementioned existing technologies still face the following prominent problems: First, the linear error of manual positioning is generally above 0.5mm and the angular deviation reaches 2-5 degrees, which is difficult to meet the sub-millimeter level orthodontic mechanical requirements; Second, indirect bonding technology involves multiple processes such as model segmentation, tray printing and clinical placement, which are time-consuming and costly, and tray deformation or incomplete placement can easily lead to final positioning deviation; Third, frequent patient micro-movements, breathing and physician operations in the oral cavity can cause visual tracking of existing AR systems to be easily lost when there is large-area occlusion, and the dynamic registration accuracy and refresh rate are insufficient, and virtual images are prone to drift or delay. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a real-time positioning and guidance system for orthodontic brackets based on augmented reality and intraoral scanning, comprising: The data acquisition unit is used to acquire a three-dimensional digital model of the patient's dentition. The data processing unit is used to receive the three-dimensional digital model, perform single-tooth segmentation on the dental arch model and output the tooth identification of each tooth, calculate the ideal bracket attachment point and bracket posture of each tooth based on the tooth identification, tooth surface geometric features and preset orthodontic rules, and generate a three-dimensional digital guide template containing the ideal bracket attachment point and virtual bracket model. The augmented reality display unit includes a processor, a perspective display screen, a depth camera, an RGB camera, and an inertial measurement unit. The depth camera and the RGB camera are used to acquire three-dimensional depth information and two-dimensional texture images of the patient's oral cavity in real time. The processor has built-in synchronous positioning and mapping front-end, real-time registration module, state fusion module, and rendering module. It is used to spatially register the three-dimensional digital guide template with the real-time acquired two-dimensional texture image and the local point cloud constructed based on the three-dimensional depth information, and to overlay the registered three-dimensional digital guide template on the perspective display screen.

[0005] To address the aforementioned technical problems, this invention also provides a method for real-time positioning and guidance of orthodontic brackets based on augmented reality and intraoral scanning, comprising: Obtain a three-dimensional digital model of the patient's dentition; The three-dimensional digital model is preprocessed and segmented into individual teeth to obtain an independent three-dimensional model and tooth identifier for each tooth; Based on the tooth identification and the calculation rules for the ideal bracket attachment point of the corresponding tooth position, the candidate bonding area and long axis direction of the buccal side of each tooth are extracted. A local best fitting plane is fitted within the candidate bonding area of ​​the buccal side. The ideal bracket attachment point and bracket posture are calculated based on the normal vector of the local best fitting plane and the long axis direction, and a three-dimensional digital guide template is generated. Import the three-dimensional digital guide template into the augmented reality display device; The augmented reality display device uses a camera to capture two-dimensional texture images and three-dimensional depth information of the patient's oral cavity in real time, and constructs a local three-dimensional point cloud based on the three-dimensional depth information. By synchronously positioning and mapping front-end tracking the viewpoint pose of the augmented reality display device, the three-dimensional digital guide template is spatially registered with the local three-dimensional point cloud and the transformation matrix is ​​calculated. Based on the transformation matrix and display calibration parameters, the three-dimensional digital guide template is rendered and superimposed on the perspective display screen of the augmented reality display device.

[0006] Optionally, when preprocessing and segmenting the three-dimensional digital model, the imported three-dimensional digital model is subjected to denoising, smoothing and mesh repair processing; according to the input data form, a mesh segmentation network or a point cloud segmentation network is used to segment the whole dentition model into independent single tooth models and output the tooth identifier of each tooth.

[0007] Optionally, when calculating the ideal bracket attachment point, the calculation rule for the ideal bracket attachment point of the corresponding tooth position is called according to the tooth identification; the buccal surface of the tooth is extracted and a candidate bonding area located in the center of the buccal surface and whose curvature change meets the preset conditions is determined; the local best fitting plane is fitted in the candidate bonding area, and the normal vector of the local best fitting plane is calculated as the initial normal vector of the bracket base plate; according to the preset orthodontic treatment technical requirements, the projection position of the bracket center point on the tooth surface is calculated in combination with the long axis direction of the tooth.

[0008] Optionally, when spatially registering the three-dimensional digital guide template with the local three-dimensional point cloud, the three-dimensional digital guide template and the local three-dimensional point cloud are coarsely registered using point cloud features and an initial transformation matrix is ​​output; the initial transformation matrix is ​​used as input, and a point-to-surface iterative nearest point algorithm is used for fine registration.

[0009] Optionally, a robust kernel function is introduced to suppress the impact of outliers on registration accuracy; the viewpoint pose output by the synchronous positioning and mapping front end, the predicted pose output by the inertial measurement unit, and the registration pose output by the fine registration are fused by extended Kalman filtering or nonlinear optimization methods to output a stable display transformation matrix.

[0010] Optionally, a visual feedback mechanism is also included: the center point and attitude of the physical tray base plate are identified and estimated by the two-dimensional texture image captured by the RGB camera and the depth map captured by the depth camera; the three-dimensional Euclidean distance between the center point of the physical tray base plate and the center point of the virtual tray base plate in the three-dimensional digital guide template is calculated as the distance deviation; the attitude deviation of the physical tray relative to the virtual tray in the Tip, Torque, and Rotation directions is calculated; when the distance deviation is less than a preset distance threshold and the angle deviations in the three directions are all less than a preset angle threshold, the three-dimensional digital guide template changes the display state or emits a prompt sound.

[0011] Optionally, before acquiring two-dimensional texture images and three-dimensional depth information of the patient's oral cavity in real time, camera extrinsic calibration and display perspective calibration are performed to establish the mapping relationship between the RGB camera coordinate system, depth camera coordinate system, display screen coordinate system and human eye coordinate system; the perspective display projection matrix is ​​solved by a single-point active alignment algorithm to ensure that the imaging position of the virtual image on the retina is consistent with the position of the real object in space.

[0012] On the other hand, the present invention also provides an electronic device including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.

[0013] On the other hand, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method.

[0014] Compared with the prior art, the present invention has the following advantages and technical effects: This invention effectively solves the aforementioned problems by constructing a complete closed-loop system, from intraoral scanning, deep learning single-tooth segmentation, and automatic bracket pose planning, to SLAM and point-to-surface ICP fusion registration, depth map occlusion rendering, and visual feedback of the physical bracket. It improves bracket positioning accuracy to sub-millimeter level (static registration <0.15mm), eliminates dependence on physical transfer trays, and reduces processing costs and steps. Through SLAM / IMU / EKF fusion and a coarse-to-fine registration strategy, it maintains stable tracking (latency <20ms, frame rate 60FPS) under dynamic oral environments and frequent occlusion. Furthermore, it introduces real-time quantitative feedback on the deviation between the physical bracket center and posture, automatically prompting when the deviation reaches a certain level, reducing reliance on visual judgment by the physician. Therefore, this invention significantly improves the accuracy, efficiency, and consistency of bracket bonding, reduces the operational threshold and chairside time (saving an average of approximately 30%), and achieves fully digital, traceable guidance for orthodontic bracket positioning. Attached Figure Description

[0015] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 2 This is a flowchart of the algorithm for tooth segmentation and FA point calculation in the data processing unit of this invention. Figure 3 This is a flowchart illustrating the logic of the real-time registration and tracking module in an embodiment of the present invention. Detailed Implementation

[0016] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0017] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0018] Example 1 like Figure 1As shown, this embodiment provides a real-time positioning and guidance system for orthodontic brackets based on augmented reality and intraoral scanning, including: As a specific embodiment of the present invention, the real-time positioning and guidance system for orthodontic brackets based on augmented reality and intraoral scanning includes a data acquisition unit, a data processing unit, and an augmented reality display unit, the specific configuration of each unit being as follows.

[0019] In this embodiment, the data acquisition unit mainly consists of a high-precision intraoral scanner, used to acquire three-dimensional surface geometry data of the patient's entire dentition non-contactly before surgery. The data format includes, but is not limited to, STL, PLY, or OBJ. This data records the anatomical morphology of the tooth crown, the curvature of the buccal surface, and the gingival margin, serving as the absolute spatial coordinate reference for subsequent digital tooth arrangement and the generation of virtual guide templates.

[0020] The data processing unit is typically a workstation or cloud computing cluster configured with a high-performance graphics computing core, used to process point cloud / mesh data and run deep learning models. This unit specifically includes a data preprocessing module, an intelligent single-tooth segmentation module, and a pose planning and template generation module. The data preprocessing module denoises the imported STL mesh, using the Laplacian smoothing algorithm to remove scanning noise and employing Poisson reconstruction technology to repair local holes and improve mesh continuity. The intelligent single-tooth segmentation module calls a preset segmentation network based on the imported data's form: when the imported data retains mesh facets and facet features such as normal vectors, curvature, and shape diameter function (SDF), a MeshSegNet-like mesh segmentation network is preferred; when the imported data is sampled as a point cloud and uses point coordinates and normal vectors as input, a PointNet++-like point cloud segmentation network can be used. This segmentation network outputs the tooth ID and gingival label corresponding to each facet or point, thereby obtaining an independent 3D model and independent coordinate system for each tooth; network training uses a combination of weighted cross-entropy loss and Dice coefficient loss to improve the segmentation robustness of tooth boundary regions. The pose planning and template generation module is based on orthodontic clinical theories such as Andrews' six elements, Roth or MBT straight wire techniques, and combines the tooth ID of a segmented single tooth, the buccal candidate bonding area, and the clinical long axis to automatically calculate the ideal bracket attachment point (FA point) and spatial pose (including Tip, Torque, Rotation parameters). The tooth ID is used to call a preset tooth position rule table to distinguish the clinical long axis definition, buccal candidate bonding area, and FA point correction rules for different tooth positions such as incisors, canines / premolars, and molars. A local plane is fitted within the candidate bonding area, and the normal vector of this local plane is used as the initial normal vector of the bracket base to generate a three-dimensional digital guide template containing a semi-transparent virtual bracket, dynamic crosshair positioning lines, and auxiliary alignment boxes.

[0021] The augmented reality display unit is a see-through head-mounted display (HMD) with spatial awareness capabilities, such as the HoloLens 2. This unit includes a multimodal sensor array, a real-time processing engine, an optical display module, a virtual / real occlusion processing module, and an interaction module. The multimodal sensor array includes a ToF or depth camera for depth perception, an RGB camera for texture feature extraction, and an inertial measurement unit (IMU) for motion prediction and compensation. The real-time processing engine runs a SLAM front-end, a real-time registration module, a state fusion module, and a 3D rendering engine. The SLAM front-end continuously estimates the viewpoint pose of the augmented reality display device based on RGB images, depth information, and IMU data. The real-time registration module uses the SLAM pose and coarse registration results as initial values ​​to spatially align the 3D digital guidance template with the real-time captured local dentition point cloud, and outputs a 4×4 transformation matrix from the preoperative 3D dentition model coordinate system to the current camera coordinate system or the local point cloud coordinate system. The state fusion module fuses the SLAM pose, IMU predicted pose, and ICP registration pose using methods such as Extended Kalman Filter (EKF) to reduce jitter and provide a stable display pose to the rendering engine. The optical display module uses holographic waveguides or freeform surface technology to project the digital guidance scheme onto the surgeon's retina, enabling optical fusion of virtual information and the real oral cavity scene. The virtual-real occlusion processing module generates an occlusion mask based on the real-time depth map acquired by the depth camera, and performs depth testing and region culling on the virtual guide template pixels in the rendering pipeline. For optical see-through head-mounted displays, real-scene light directly enters the human eye; therefore, this module does not use the traditional Z-buffer within the same virtual scene as the sole occlusion criterion. Instead, it uses the real-time depth map to determine the front-to-back relationship of real teeth, oral soft tissue, doctor's hands, and instruments relative to the virtual guide template, achieving occlusion rendering that conforms to the real operation scenario. The interaction module primarily uses pre-set display parameters and voice control, allowing doctors to adjust the transparency of the virtual guide template, display mode, or switch target teeth via voice commands. Gesture recognition is only used for auxiliary operations in non-critical bonding stages to prevent doctors from interrupting continuous positioning and clamping actions during bracket bonding positioning due to gesture operations.

[0022] Example 2 This embodiment provides a real-time positioning and guidance method for orthodontic brackets based on augmented reality and intraoral scanning, including: First, an intraoral scanner is used to acquire a three-dimensional digital model of the patient's dentition.

[0023] Then, a mapping relationship is established between the 3D dental arch model coordinate system, image acquisition coordinate system, display coordinate system, and human eye observation coordinate system. The 3D digital model is preprocessed and individual tooth segmentation is performed to obtain an independent 3D model and tooth ID for each tooth. The aforementioned mapping relationship mainly includes the calibration relationship between the RGB camera coordinate system, depth camera coordinate system, display screen coordinate system, and human eye observation coordinate system, as well as the perspective projection relationship from the display screen coordinate system to the human eye observation coordinate system. The spatial transformation relationship between the preoperative 3D dental arch model coordinate system and the intraoperative real-time camera coordinate system or real-time local dental arch point cloud coordinate system is not preset as a fixed extrinsic parameter, but is solved in real-time during the subsequent real-time spatial registration step through PnP initial pose estimation, FPFH coarse registration, and point-to-surface ICP fine registration. For individual tooth segmentation, a preset network is selected based on the input data format: MeshSegNet-like networks are used for mesh data that retains surface features, and PointNet++-like networks are used for point cloud data. A weighted combination of cross-entropy loss function and Dice coefficient loss function is used for optimization.

[0024] Next, based on the tooth ID, geometric center, and clinical long axis of each tooth, the ideal bracket attachment point and bracket posture are determined, generating a 3D digital guide template. The method for determining the ideal bracket attachment point includes: calling the FA point calculation rule for the corresponding tooth position based on the tooth ID; extracting the buccal surface of the tooth and identifying a candidate bonding area located in the center of the buccal surface with minimal curvature change and suitable for bracket base fitting; fitting a local best-fit plane within the candidate bonding area and calculating the normal vector of this local plane as the initial normal vector for the bracket base; calculating the projection position of the bracket center point on the tooth surface based on Andrews' six-element rule or Roth's straight wire orthodontic technique, combined with the tooth's long axis direction. After generating the 3D digital guide template containing the ideal bracket attachment point and a virtual bracket model, this 3D digital guide template is imported into an augmented reality display device.

[0025] Subsequently, the patient's oral cavity images and depth data are acquired in real time using the camera of the augmented reality display device. Before constructing the local 3D point cloud, the depth data undergoes quality control. Quality control includes evaluating the proportion of effective depth points, hole regions, and outlier distribution. For depth maps that meet quality requirements, temporal smoothing, outlier removal, hole filling, and RGB image synchronization are performed to construct the local 3D point cloud for the current viewpoint. If the depth data quality is insufficient, a prompt is made to re-acquire or use the effective pose from the previous frame; low-quality point clouds are not directly input for ICP registration. Prior to this step, camera extrinsic calibration and display perspective calibration are performed, establishing the mapping relationship between the camera coordinate system, display coordinate system, and human eye coordinate system to ensure that the imaging position of the virtual image on the retina is consistent with the spatial position of the real object.

[0026] Subsequently, the 3D digital guide template and the local 3D point cloud are spatially registered in real time, and the transformation matrix is ​​calculated. The SLAM front-end is responsible for continuous viewpoint pose tracking of the augmented reality display device and provides the initial camera pose value to the registration module. The registration module adopts a two-stage strategy from coarse to fine. In the coarse registration stage, point cloud features such as Fast Point Feature Histogram (FPFH) are used to obtain a 4×4 initial transformation matrix from the 3D digital guide template to the real-time oral cavity point cloud. If there are insufficient feature matching points, relocalization is triggered or the previous valid pose is retained. In the fine registration stage, the initial transformation matrix output by the coarse registration is used as input. On this basis, the Iterative Closest Point (ICP) algorithm for point-to-plane is adopted, and robust kernel functions such as Huber Loss are introduced to suppress the influence of outliers. Finally, the SLAM pose, IMU prediction, and ICP output are fused by state fusion methods such as EKF to obtain a stable real-time display transformation.

[0027] Then, based on the aforementioned transformation matrix and the aforementioned display projection / perspective calibration parameters, the three-dimensional digital guide template is rendered and superimposed on the perspective display screen of the augmented reality display device, ensuring visual alignment with the patient's actual teeth. This step also includes a visual feedback mechanism: the system identifies the outline of the physical bracket base through RGB images and estimates the center point and orientation of the physical bracket base by combining the depth map; the distance deviation is defined as the three-dimensional Euclidean distance between the center point of the physical bracket base and the center point of the virtual bracket base, and the angle deviation is defined as the orientation deviation of the physical bracket relative to the virtual bracket in the Tip, Torque, and Rotation directions. When the distance deviation is less than a preset threshold (e.g., 0.5 mm, preferably 0.2 mm) and the angle deviations in all three directions are less than a preset threshold (e.g., 2 degrees), the virtual guide template changes color or emits a prompt sound to confirm accurate positioning.

[0028] Finally, guided by augmented reality overlay and visual feedback indicating compliance, the dentist uses tweezers or bracket clamping tools to hold the physical bracket in place and then cements it to the corresponding position on the patient's tooth surface. This method emphasizes that the system provides positioning, deviation assessment, and compliance feedback, reducing reliance on the operator's experience-based visual judgment.

[0029] As a specific embodiment of the present invention, the clinical usage process is as follows.

[0030] During the digital modeling and design phase, the physician uses an intraoral scanner to obtain the patient's dental model; the data processing unit segments the dental model and calculates the ideal bracket position for each tooth based on orthodontic principles; the system generates a composite 3D scene containing virtual teeth, virtual brackets, and positioning markers.

[0031] In the environmental perception and initial registration stage, the doctor wears AR glasses and looks at the patient's mouth. The AR glasses camera captures real-time images. The system uses feature point detection algorithms such as ORB and SIFT to find feature points in the real-time images that match the 3D digital model or pre-rendered template image. The initial pose of the camera is obtained by solving the PnP problem. Then, geometric coarse registration is performed using point cloud features such as FPFH of the real-time depth point cloud, and the initial transformation matrix is ​​output for fine registration.

[0032] During the real-time tracking and fine registration stage, based on the coarse registration, the system starts the SLAM engine to perform continuous viewpoint pose tracking. At the same time, it uses a depth camera to acquire the local 3D point cloud of the current viewpoint and performs point-to-surface ICP registration with the pre-stored complete dental model point cloud. To cope with the patient's head movement and slight shaking of the display device, the system combines IMU data to perform motion prediction and fuses the SLAM pose, ICP pose and IMU prediction results through EKF.

[0033] During the augmented reality guidance and feedback phase, the system precisely overlays the rendered virtual brackets and positioning crosshairs onto the patient's real teeth. The dentist can see the relative position of the virtual brackets to the tooth surface through AR glasses. When the physical brackets enter the field of view, the system identifies the bracket base outline using RGB images and estimates the center point and orientation of the physical bracket using a depth map. It then calculates in real-time the center distance deviation between the physical and virtual brackets, as well as the angular deviations in the Tip, Torque, and Rotation directions. If the deviation is within acceptable limits, the virtual marker turns green or emits a prompt sound; otherwise, the warning color remains and the adjustment direction is displayed.

[0034] During the bonding process, the physician, guided by AR, adjusts the physical bracket to align with the virtual image, and maintains the position to complete the light curing fixation after the system provides feedback that the target has been met.

[0035] Example 3 This embodiment provides a real-time positioning and guidance method for orthodontic brackets based on augmented reality and intraoral scanning, including: As a specific embodiment of the present invention, the hardware configuration of the real-time positioning and guidance system for orthodontic brackets based on augmented reality and intraoral scanning is as follows.

[0036] In this embodiment, the system hardware mainly consists of three parts: a 3D intraoral scanner using 3Shape TRIOS 4, which has color scanning capabilities and a scanning accuracy of up to 10 micrometers, used to acquire a three-dimensional model of the patient's entire dentition before surgery; a high-performance computing workstation equipped with an Intel Core i9 processor, 64GB of memory, and an NVIDIA RTX 4090 graphics card, used to run deep learning segmentation networks and high-load 3D rendering and registration algorithms; and an AR head-mounted display using a Microsoft HoloLens 2, which displays a binocular 2K resolution with a 52-degree field of view. The sensors include four visible light tracking cameras, one 1MP depth camera (ToF), one 8MP RGB camera, an accelerometer, a gyroscope, and a magnetometer. The processor is a Qualcomm Snapdragon 850 and a holographic processing unit (HPU) 2.0.

[0037] As a specific embodiment of the present invention, the data processing and scheme generation algorithm is as follows: Figure 2 As shown, the specific process is as follows: In the data import step, STL format oral scan data is read. In the mesh repair step, the Laplacian smoothing algorithm is used to remove scanning noise, and the Poisson reconstruction algorithm is used to repair mesh holes. In the tooth segmentation step, an improved MeshSegNet network is used, inputting the geometric features of the tooth mesh (normal vector, curvature, shape diameter function SDF), and outputting the semantic label (tooth ID or gingiva) for each facet; when the input is processed into point cloud data, the PointNet++-like point cloud segmentation network is used to output the tooth ID label for the corresponding point. The loss function uses a combination of Weighted Cross-Entropy Loss and Generalized Dice Loss to address the imbalance problem between tooth and gingiva samples. The calculation formula is as follows: ; In the FA point calculation step, for a segmented single tooth, the tooth ID is first read and the corresponding FA point calculation rule is called; then, its minimum bounding box (OBB) is calculated to determine the principal axis direction, the buccal mesh region is filtered according to the normal vector direction, and a candidate bonding region with small curvature change is selected in the center of the buccal side; a local best-fit plane is fitted within the candidate bonding region, and the normal vector of this local plane is used as the initial normal vector of the bracket base; finally, the FA point height, axial tilt, and rotation direction are adjusted according to Andrews' planar guide rule. In the guided model generation step, a semi-transparent virtual bracket model (STL) is generated at the calculated FA points, and crosshairs are added. The major axis of the crosshairs coincides with the major axis of the tooth, and the minor axis is parallel to the incisal edge of the tooth.

[0038] As a specific embodiment of the present invention, the real-time registration and tracking method is as follows: Figure 3 As shown. During initialization, when the physician looks at the patient's mouth, the RGB camera captures an image. The system extracts ORB (Oriented Fast and Rotated BRIEF) feature points from the image and matches these feature points with a multi-angle template image pre-rendered from the 3D model. The initial pose of the camera is obtained by solving the PnP problem. Subsequently, a 4×4 coarse registration transformation matrix is ​​obtained by matching the FPFH features between the real-time depth point cloud and the digital guide template point cloud.

[0039] During real-time tracking and fine registration, the system utilizes the Lucas-Kanade optical flow method to track feature points and rapidly update the camera pose. The local point cloud acquired by the HoloLens depth camera from the current viewpoint undergoes depth quality assessment, outlier removal, and hole filling before being compared with a pre-stored complete dental model point cloud for point-to-surface ICP registration. ICP uses the aforementioned coarse registration transformation matrix as its initial value, and the objective function is to minimize the point-to-surface distance error. ; in For real-time points, For the corresponding points of the model, for The normal vector at that location.

[0040] To eliminate jitter caused by ICP calculation, an extended Kalman filter (EKF) is introduced. The state vector includes position, velocity, attitude quaternions, and angular velocity. The prediction phase utilizes IMU data, while the update phase uses the pose calculated by ICP and the viewpoint pose output by the SLAM front-end for correction. Static registration accuracy is supported by the initial coarse registration value, the convergence result of point-to-area ICP, and the suppression of abnormal depth points by the Huber robust kernel. Dynamic latency is jointly affected by SLAM tracking, KD-Tree nearest neighbor search, and EKF state fusion.

[0041] As a specific embodiment of the present invention, the clinical operation process includes the following steps. In preoperative preparation, the physician performs an intraoral scan of the patient and uploads the data. The cloud server automatically completes the segmentation and bracket layout design, generating an AR guide file in .glb format. During device wearing, the physician wears HoloLens 2, opens the dedicated APP, and scans the patient's face or QR code to load the corresponding case data. During calibration, the physician performs OST display calibration and camera extrinsic parameter checks to ensure that the holographic image is aligned with the line of sight and the spatial position of the actual dentition. During recognition and locking, after the patient opens their mouth, the physician observes the patient's dentition, the system recognizes the tooth features, and the virtual bracket quickly adheres to the tooth surface. During operation, the physician applies an etching agent and adhesive to the tooth surface, uses tweezers to pick up the bracket with the base adhesive applied, and moves the bracket through the AR glasses to make it overlap with the green virtual bracket; the system identifies the center and posture of the physical bracket base plate in real time, calculates the three-dimensional distance deviation between it and the center of the virtual bracket base plate, as well as the angular deviations in the Tip, Torque, and Rotation directions. When the distance deviation is less than a preset threshold (e.g., 0.5mm, preferably 0.2mm) and the angle deviation meets the preset requirements, the virtual crosshair thickens and brightens, while a "beep" sound is emitted. During curing and inspection, the position remains unchanged, and an assistant performs light curing. After all brackets are bonded, the system can perform a full-mouth scan to generate a heat map showing the overall bonding error.

[0042] As a specific embodiment of the present invention, the anomaly handling and security mechanism includes tracking loss handling and occlusion rendering handling. In tracking loss handling, when the patient turns their head significantly or is completely obscured by their hand, causing tracking loss, the system displays a "Tracking Lost" message and retains the valid position of the last frame. Once the dental arch feature points reappear, the system invokes the relocalization process, retrieving ORB feature points matching the current RGB image from pre-stored multi-angle template images or keyframes, restoring the camera pose through PnP solving, and restarting coarse registration and ICP fine registration with the restored pose. Under the hardware conditions of this embodiment, the relocalization recovery time can be controlled within 50ms. In occlusion rendering handling, to prevent virtual brackets from floating above the physician's hands or instruments, the system uses a real-time depth map from a depth camera to generate an occlusion mask. If the depth value of the doctor's hand or instrument is less than the depth value of the corresponding virtual bracket pixel, the virtual pixels in that area will be removed during rendering to achieve a realistic occlusion effect. If there are obvious holes or outliers in the depth map, the holes will be filled and the outliers will be removed before being used for occlusion judgment.

[0043] As a specific embodiment of the present invention, under the hardware and clinical testing conditions of this embodiment, the key performance indicators of the system include: static registration accuracy less than 0.15mm, dynamic tracking latency less than 20ms, frame rate of 60FPS, and single-tooth segmentation accuracy (Dice) greater than 95%, saving approximately 30% of chairside time per patient on average compared to traditional bonding. Specifically, the single-tooth segmentation accuracy is supported by a MeshSegNet / PointNet++ segmentation network and a combination of cross-entropy loss and Dice loss training; the static registration accuracy is supported by FPFH coarse registration, point-to-surface ICP, Huber robust kernel function, and EKF fusion; and the dynamic latency is supported by SLAM tracking, KD-Tree nearest neighbor search, and IMU prediction compensation.

[0044] As a specific embodiment of the present invention, the camera extrinsic parameter calibration and display calibration method is as follows. To ensure that the virtual image is accurately superimposed on the real object, this embodiment establishes a precise coordinate system transformation relationship. For OST (Optical See-Through) display calibration, due to differences in interpupillary distance (IPD) and wearing position among different users, the projection position of the virtual image on the retina may deviate. The system adopts the SPAAM (Single Point Active Alignment Method): the system displays a series of crosshair cursors on the screen, and the user moves their head to make the cursors overlap and confirm with known physical markers in the real world; the system collects at least 6 sets of matching point pairs and uses direct linear transformation (DLT) to solve for the projection matrix P. This projection matrix, together with the transformation matrix obtained from registration, is used in the subsequent rendering steps for perspective rendering of the virtual guide template.

[0045] Camera extrinsic calibration is used to determine the transformation relationship between the RGB camera coordinate system and the depth camera coordinate system, as well as the transformation relationship between the camera coordinate system and the display coordinate system. This calibration determines the camera-display extrinsic relationship and does not use the physician's handheld tweezers or physical tray as the mechanical end being calibrated. Typically, the aforementioned camera and display extrinsic parameters are completed at the factory; however, to improve accuracy, physicians can also use a dedicated calibration board (Checkerboard) for recalibration.

[0046] As a specific embodiment of the present invention, the data structure and communication protocol are as follows. In the point cloud data structure, a KD-Tree (K-Dimensional Tree) is used to organize the point cloud data, reducing the time complexity of Nearest Neighbor Search from O(N) to O(log N), significantly improving ICP registration speed. In network communication, the data processing unit and AR glasses can use Wi-Fi 5 or Wi-Fi 6 as connection methods; the transmission protocol uses WebRTC to transmit real-time video streams and MQTT protocol to transmit control commands (such as voice commands or gesture operations in non-adhesive key stages); the 3D model transmission uses the Draco compression algorithm, with a compression ratio of up to 1:20, to reduce loading time.

[0047] As a specific embodiment of the present invention, the system also introduces a cloud-based federated learning optimization mechanism. To continuously improve the accuracy of tooth segmentation and feature recognition, the parameters of the single-tooth segmentation module are globally and collaboratively updated through the federated learning mechanism, achieving model optimization without uploading original case images. During local training, each computing workstation deployed in the hospital uses local anonymized case data to fine-tune the segmentation network (MeshSegNet / PointNet++). During parameter aggregation, each hospital only uploads the trained model parameters or gradients to the cloud central server, without uploading original case images, to protect patient privacy. During global updates, the cloud server performs a weighted average (FedAvg algorithm) on the parameters from various locations to update the global model before distributing it to each hospital. Through this method, the system can continuously optimize using multi-center data, improving its ability to identify rare dental malformations.

[0048] As a specific embodiment of the present invention, the system is designed with an open data interface, enabling compatibility with mainstream orthodontic software and hardware. Regarding intraoral scanner compatibility, it supports STL, PLY, and OBJ format data exported from mainstream intraoral scanners such as 3Shape, iTero, and Carestream. Regarding tooth alignment software compatibility, it supports importing tooth alignment plans (Setup Models) designed by software such as 3Shape Ortho Analyzer and Maestro 3D; the system can automatically parse XML format tooth alignment data and extract the target bracket positions. Regarding the bracket database, it has a built-in 3D model library of brackets from mainstream brands such as Damon, 3M, and AO; after the dentist selects the bracket model, the system automatically loads the corresponding virtual model.

[0049] This invention connects intraoral scan data processing, deep learning single-tooth segmentation, automatic calculation of bracket pose based on tooth ID and buccal candidate bonding area, SLAM and point-to-surface ICP fusion registration, depth map occlusion masking, and physical bracket positioning feedback into a complete closed loop, which is different from separate tooth segmentation, separate virtual bracket planning, or separate augmented reality oral display solutions.

[0050] On the other hand, this embodiment also provides an electronic device, including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.

[0051] On the other hand, this embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method.

[0052] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A real-time positioning and guidance system for orthodontic brackets based on augmented reality and intraoral scanning, characterized in that, include: The data acquisition unit is used to acquire a three-dimensional digital model of the patient's dentition. The data processing unit is used to receive the three-dimensional digital model, perform single-tooth segmentation on the dental arch model and output the tooth identification of each tooth, calculate the ideal bracket attachment point and bracket posture of each tooth based on the tooth identification, tooth surface geometric features and preset orthodontic rules, and generate a three-dimensional digital guide template containing the ideal bracket attachment point and virtual bracket model. The augmented reality display unit includes a processor, a perspective display screen, a depth camera, an RGB camera, and an inertial measurement unit. The depth camera and the RGB camera are used to acquire three-dimensional depth information and two-dimensional texture images of the patient's oral cavity in real time. The processor has built-in synchronous positioning and mapping front-end, real-time registration module, state fusion module, and rendering module. It is used to spatially register the three-dimensional digital guide template with the real-time acquired two-dimensional texture image and the local point cloud constructed based on the three-dimensional depth information, and to overlay the registered three-dimensional digital guide template on the perspective display screen.

2. A method for real-time positioning and guidance of orthodontic brackets based on augmented reality and intraoral scanning, characterized in that, include: Obtain a three-dimensional digital model of the patient's dentition; The three-dimensional digital model is preprocessed and segmented into individual teeth to obtain an independent three-dimensional model and tooth identifier for each tooth; Based on the tooth identification and the calculation rules for the ideal bracket attachment point of the corresponding tooth position, the candidate bonding area and long axis direction of the buccal side of each tooth are extracted. A local best fitting plane is fitted within the candidate bonding area of ​​the buccal side. The ideal bracket attachment point and bracket posture are calculated based on the normal vector of the local best fitting plane and the long axis direction, and a three-dimensional digital guide template is generated. Import the three-dimensional digital guide template into the augmented reality display device; The augmented reality display device uses a camera to capture two-dimensional texture images and three-dimensional depth information of the patient's oral cavity in real time, and constructs a local three-dimensional point cloud based on the three-dimensional depth information. By synchronously positioning and mapping front-end tracking the viewpoint pose of the augmented reality display device, the three-dimensional digital guide template is spatially registered with the local three-dimensional point cloud and the transformation matrix is ​​calculated. Based on the transformation matrix and display calibration parameters, the three-dimensional digital guide template is rendered and superimposed on the perspective display screen of the augmented reality display device.

3. The method according to claim 2, characterized in that, When preprocessing and segmenting individual teeth in the three-dimensional digital model, the imported three-dimensional digital model is subjected to denoising, smoothing and mesh repair processing; according to the input data form, a mesh segmentation network or a point cloud segmentation network is used to segment the whole dentition model into independent single tooth models and output the tooth identifier of each tooth.

4. The method according to claim 2, characterized in that, When calculating the ideal bracket attachment point, the calculation rule for the ideal bracket attachment point of the corresponding tooth position is called according to the tooth identification; the buccal surface of the tooth is extracted and a candidate bonding area located in the center of the buccal surface and whose curvature change meets the preset conditions is determined; the local best fitting plane is fitted in the candidate bonding area, and the normal vector of the local best fitting plane is calculated as the initial normal vector of the bracket base plate. Based on the pre-set orthodontic treatment technical requirements, the projection position of the bracket center point on the tooth surface is calculated in combination with the long axis direction of the tooth.

5. The method according to claim 2, characterized in that, When spatially registering the 3D digital guide template with the local 3D point cloud, the point cloud features are used to perform coarse registration between the 3D digital guide template and the local 3D point cloud and output an initial transformation matrix; with the initial transformation matrix as input, the point-to-surface iterative nearest point algorithm is used for fine registration.

6. The method according to claim 5, characterized in that, A robust kernel function is introduced to suppress the impact of outliers on registration accuracy; the viewpoint pose output by the synchronous positioning and mapping front end, the predicted pose output by the inertial measurement unit, and the registration pose output by the fine registration are fused by extended Kalman filtering or nonlinear optimization method to output a stable display transformation matrix.

7. The method according to claim 2, characterized in that, It also includes a visual feedback mechanism: identifying and estimating the center point and orientation of the physical tray base plate through the two-dimensional texture image captured by the RGB camera and the depth map captured by the depth camera; The three-dimensional Euclidean distance between the center point of the physical bracket base plate and the center point of the virtual bracket base plate in the three-dimensional digital guide template is calculated as the distance deviation. The attitude deviation of the physical bracket relative to the virtual bracket in the three directions of Tip, Torque, and Rotation is calculated. When the distance deviation is less than a preset distance threshold and the angle deviations in the three directions are all less than a preset angle threshold, the three-dimensional digital guide template changes its display state or emits a prompt sound.

8. The method according to claim 2, characterized in that, Before acquiring two-dimensional texture images and three-dimensional depth information of the patient's oral cavity in real time, camera extrinsic parameter calibration and monitor perspective calibration are performed to establish the mapping relationship between the RGB camera coordinate system, depth camera coordinate system, display screen coordinate system and human eye coordinate system. The perspective display projection matrix is ​​solved by a single-point active alignment algorithm, so that the imaging position of the virtual image on the retina is consistent with the spatial position of the real object.

9. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, characterized in that, When the processor executes the computing program, it implements the method of any one of claims 2-8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 2-8.