A multimodal vision and augmented reality based dynamic orthodontic bracket navigation method, system, and device

By combining multimodal vision and augmented reality technologies with preoperative 3D patient data and real-time intraoral image processing, personalized and high-precision bracket positioning is achieved, solving the problems of large bracket positioning errors and complex operations in existing technologies, and improving the convenience and accuracy of the bonding process.

CN122156550APending Publication Date: 2026-06-05SHANGHAI NINTH PEOPLES HOSPITAL SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI NINTH PEOPLES HOSPITAL SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
Filing Date
2026-03-25
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing technologies have problems such as large positioning errors, complex operation, high cost, and difficulty in achieving personalized positioning and real-time adjustment during bracket bonding, especially in complex oral environments where it is difficult to achieve high-precision and convenient bracket positioning.

Method used

By acquiring the patient's preoperative oral 3D data, a 3D digital target model is generated. Multimodal vision and augmented reality technologies are used to collect intraoral image data in real time, calculate the spatial transformation matrix, render the virtual bracket bonding mark on the real tooth surface in real time, and combine inertial measurement unit data to perform high-frequency pose prediction to achieve dynamic 3D registration.

Benefits of technology

It achieves personalized and accurate bracket positioning, reduces operational difficulty and time costs, improves recognition accuracy and stability in complex oral environments, and provides sub-millimeter registration accuracy and real-time visual guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122156550A_ABST
    Figure CN122156550A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of oral orthodontic navigation, and particularly relates to a dynamic orthodontic bracket navigation method, system and device based on multi-modal vision and augmented reality, which comprises the following steps: S1: generating a three-dimensional digital target model and extracting a virtual anatomical feature point set; S2: collecting RGB image data and depth image data in the patient's mouth, using a pre-trained multi-modal large model for real-time inference, segmenting the target teeth and extracting a real-time anatomical feature point set; S3: calculating a spatial transformation matrix based on the virtual anatomical feature point set and the real-time anatomical feature point set, and converting the virtual bracket bonding mark to the real-time camera coordinate system according to the spatial transformation matrix; S4: real-time rendering and locking the converted virtual bracket bonding mark on the real tooth surface. The present application realizes high-precision and low-delay bracket positioning navigation in a complex intraoral environment, overcomes interference such as strong reflection, saliva coverage and instrument obstruction, improves the accuracy and operation convenience of orthodontic bracket bonding, and simplifies the clinical process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of orthodontic navigation technology, and in particular to a dynamic orthodontic bracket navigation method, system and device based on multimodal vision and augmented reality. Background Technology

[0002] With the rapid development of digital dentistry and artificial intelligence technologies, orthodontic treatment is showing a significant trend towards digitalization and intelligentization. Against this backdrop, to improve the accuracy and reliability of bracket bonding, existing technologies have mainly formed two mainstream approaches: direct bonding and indirect bonding.

[0003] Direct bonding relies on the operator's experience in judging the tooth shape, clinical crown center, and long axis direction within the patient's mouth. While this method is simple and requires no additional auxiliary devices, the operator's reliance on visual positioning is prone to errors under complex conditions such as the confined space of the mouth, strong light, and saliva reflection. Such positioning deviations can lead to inaccurate positioning of key parameters such as bracket height and axial tilt, affecting the mechanical performance of the subsequent archwire, increasing the probability of restarting adjustments during treatment, and prolonging the overall treatment cycle.

[0004] Indirect bonding techniques determine the ideal bracket position using a preoperative model for tooth alignment, then transfer the entire bracket to the intraoral cavity via a transfer tray. This method shifts some operations to the laboratory, improving positioning accuracy to some extent. However, this approach adds a cumbersome process to the laboratory, involving model making, tooth alignment, and transfer tray fabrication, resulting in higher material and labor costs per treatment. Furthermore, the transfer tray may deform during fabrication, storage, and intraoral placement, introducing new positioning errors. In addition, this approach lacks a dynamic adjustment mechanism based on intraoperative conditions and is difficult to implement dynamically.

[0005] With the widespread use of intraoral scanners, cone-beam computed tomography (CBCT), and digital orthodontic design software, clinicians can obtain relatively complete three-dimensional data of the dentition and ideal bracket placement plans preoperatively. However, most existing digital solutions only provide static three-dimensional model references before surgery. During the actual bonding operation, the operator still needs to rely on memory or constantly compare the three-dimensional plan on the external screen with the patient's oral cavity, lacking real-time in-situ visual guidance during the dynamic operation. Some navigation solutions require the use of additional markers, positioning clamps, or complex calibration procedures, increasing the operation steps and time costs, making it difficult to balance accuracy, real-time performance, and clinical ease of operation.

[0006] In actual clinical practice, the required positioning accuracy varies significantly between different teeth and different locations. The anterior aesthetic zone demands extremely high bracket positioning accuracy, while the posterior zone requires relatively lower accuracy. However, once existing bracket positioning schemes are determined, it is difficult to adopt differentiated guidance strategies for different teeth; all teeth are treated according to a uniform standard. This forces operators to maintain overall control according to the highest accuracy requirements, resulting in a significant increase in operational difficulty and time consumption, and significant shortcomings in terms of economy and practicality. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality, comprising: S1: Obtain the patient's preoperative oral cavity three-dimensional data, generate a three-dimensional digital target model based on the oral cavity three-dimensional data, and extract a set of virtual anatomical feature points from the three-dimensional digital target model; S2: Real-time acquisition of RGB image data and depth image data in the patient's mouth using augmented reality devices; real-time inference of the RGB image data and depth image data using a pre-trained multimodal large model; segmentation of the target tooth and extraction of its real-time anatomical feature point set in the real-time camera coordinate system. S3: Based on the virtual anatomical feature point set and the real-time anatomical feature point set, calculate the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system. According to the spatial transformation matrix, transform the virtual bracket bonding mark in the three-dimensional digital target model to the real-time camera coordinate system. The transformation relationship satisfies... ,in The coordinates of the virtual bracket bonding mark in the world coordinate system. For rotation matrix, It is a translation vector. The coordinates are in the transformed real-time camera coordinate system; S4: Using the augmented reality device, the converted virtual bracket bonding mark is rendered and locked in real time on the corresponding real tooth surface in the doctor's field of vision in a three-dimensional overlay manner.

[0008] Preferably, in step S1, generating a three-dimensional digital target model based on the oral cavity three-dimensional data includes: Import the patient’s cone-beam computed tomography (CBCT) data and intraoral scan data, and perform rigid registration and fusion using the iterative nearest point algorithm to generate a three-dimensional digital model containing the complete dental anatomy. Virtual tooth arrangement is performed on the three-dimensional digital model, the coordinates of the clinical crown center point and the long axis vector of the tooth body of each target tooth are calculated, and the ideal three-dimensional bonding position of the bracket is determined based on the coordinates of the clinical crown center point and the long axis vector of the tooth body, thereby generating the three-dimensional digital target model containing the preset bracket pose information of each tooth. A world coordinate system is established, and the virtual anatomical feature point set and virtual bracket bonding identifier are extracted from the three-dimensional digital target model. The virtual anatomical feature point set includes at least the enamel-cementum boundary point cloud, incisal edge point cloud, or cusp point cloud; the virtual bracket bonding identifier includes at least the bracket center point coordinates or the local coordinate point set of the bracket boundary contour line.

[0009] Preferably, in step S2, real-time inference is performed on the RGB image data and the depth image data using a pre-trained multimodal large model, including: The visual large model pre-trained based on a self-supervised learning framework is invoked to extract features from the RGB image data and depth map data. The self-supervised learning framework includes a mask autoencoder, whose pre-training objective is to minimize the reconstruction loss of the mask image patch. During the encoding stage of the large visual model, the semantic features of the RGB image and the spatial geometric features of the depth image are fused through a cross-attention mechanism to output a high-dimensional feature map. The high-dimensional feature map is restored to a pixel-level semantic segmentation mask by the decoder of the large visual model. The semantic segmentation mask, combined with depth information, maps two-dimensional pixels back to three-dimensional space and extracts the real-time anatomical feature point set of the real teeth in the current field of view. The formula for calculating the reconstruction loss is as follows: Where M is the set of image patches to be masked. For the original image patch, To reconstruct image patches.

[0010] Preferably, the multimodal large model is optimized using a joint loss function during the downstream task training phase to ensure the segmentation accuracy of the target tooth boundary. The calculation formula for the joint loss function is as follows: ,in, Dice loss is used to improve the overlap of tooth region segmentation. Cross-entropy loss is used to optimize pixel-level classification accuracy.

[0011] Preferably, in step S3, the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system is calculated, including: Using singular value decomposition or nonlinear optimization algorithm, point cloud matching is performed on the virtual anatomical feature point set and the real-time anatomical feature point set with anatomical correspondence, and the initial spatial transformation matrix that minimizes the mean square error between the virtual anatomical feature point set and the real-time anatomical feature point set is obtained. By integrating data from the inertial measurement unit built into the augmented reality device, and using extended Kalman filtering or visual inertial odometry, high-frequency pose prediction is performed between consecutive video frames, continuously updating the rotation matrix. With translation vector .

[0012] Preferably, calculating the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system further includes: Construct a loss term that includes semantic segmentation and 3D registration loss term The total loss function is calculated using the following formula: , and These are adaptive weighting coefficients; The system continuously monitors the confidence level of visual features within the intraoral field of view. This confidence level is evaluated based on the confidence map of a semantic segmentation mask or the inlier rate of feature point matching. When the target tooth is detected to be obscured by instruments, covered in saliva, or subject to strong glare, causing the visual feature confidence level to fall below a preset threshold, the system automatically reduces the confidence level. The weight and increase The weights are adjusted so that the system relies more on inertial measurement unit data for pose deduction when visual features are insufficient, thus maintaining the stability of the spatial transformation matrix.

[0013] Preferably, in step S4, the real-time rendering and locking onto the corresponding real tooth surface in the doctor's field of vision includes: Based on the binocular camera model of the augmented reality device, the world coordinates of the virtual bracket bonding mark are transformed to the left-eye camera coordinate system and the right-eye camera coordinate system, respectively. Perspective projection is performed on the transformed coordinates to generate left and right eye display images with binocular parallax. The virtual bracket bonding mark is occluded using the depth image data. The rendering transparency or occlusion relationship of the virtual mark is determined by comparing the projection depth of the virtual mark point with the depth value of the real object inside the mouth. Temporal smoothing filtering is applied to the projection coordinates of the virtual bracket bonding mark between consecutive frames to eliminate rendering graphics jitter caused by pose jitter. Furthermore, display latency is reduced by combining time warp technology with inertial measurement unit data, thus ensuring that the virtual mark is stably locked onto the real tooth surface.

[0014] Preferably, in step S4, the virtual bracket bonding mark includes: The virtual bracket bonding markings include at least one of the bracket outline, center point mark, crosshair, or tooth long axis; the augmented reality device is optical see-through AR glasses, and the real-time rendering is achieved through the optical display module of the AR glasses.

[0015] Based on the same concept, the present invention also provides a dynamic orthodontic bracket navigation system based on multimodal vision and augmented reality, comprising: A three-dimensional target model construction module is used to acquire the patient's preoperative oral three-dimensional data, generate a three-dimensional digital target model based on the oral three-dimensional data, and extract a set of virtual anatomical feature points from the three-dimensional digital target model; The real-time multimodal feature extraction module is used to collect RGB image data and depth image data in the patient's mouth in real time through augmented reality devices, and use a pre-trained multimodal large model to perform real-time inference on the RGB image data and the depth image data to segment the target tooth and extract its real-time anatomical feature point set in the real-time camera coordinate system. The dynamic registration and coordinate transformation module is used to calculate the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system based on the virtual anatomical feature point set and the real-time anatomical feature point set; and to transform the virtual bracket bonding mark in the three-dimensional digital target model to the real-time camera coordinate system according to the spatial transformation matrix. The augmented reality rendering guidance module, through the augmented reality device, renders and locks the converted virtual bracket bonding mark in real time onto the corresponding real tooth surface in the doctor's field of vision in a three-dimensional overlay manner.

[0016] Based on the same concept, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores computer-readable instructions, which, when executed by the processor, cause the processor to perform the steps of the dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality as described in any one of the embodiments.

[0017] Compared with the prior art, the beneficial effects of the present invention are: (1) This invention acquires the patient’s preoperative oral cavity three-dimensional data and generates a three-dimensional digital target model containing a set of virtual anatomical feature points, thereby establishing a high-precision preoperative virtual benchmark, providing accurate spatial reference and visual guidance data source for subsequent real-time navigation, thus laying a digital foundation for the entire navigation process and ensuring the personalization and accuracy of bracket positioning.

[0018] (2) This invention uses a pre-trained multimodal large model to infer the real-time acquired intraoral RGB-D image data, segment the target teeth and extract the real-time anatomical feature point set, so as to achieve robust feature extraction in complex intraoral environments (strong reflection, saliva coverage, instrument occlusion). The multimodal large model integrates RGB texture features and depth geometric features for anti-interference processing, thereby ensuring that key anatomical landmarks of teeth can still be stably identified under poor visual conditions, and improving the reliability of the system in clinical practice.

[0019] (3) This invention calculates the spatial transformation matrix based on the virtual anatomical feature point set and the real-time anatomical feature point set, and transforms the virtual bracket bonding mark to the real-time camera coordinate system to achieve dynamic three-dimensional registration between the virtual model and the real teeth. Combined with the inertial measurement unit data, high-frequency pose prediction is performed to ensure that when the patient moves slightly or the doctor's perspective changes, the virtual guide mark can follow and stably lock onto the real tooth surface in milliseconds, achieving sub-millimeter registration accuracy.

[0020] (4) The present invention renders and locks the converted virtual bracket bonding mark in real time on the real tooth surface in the doctor's field of vision in a three-dimensional superposition method, realizing the in-situ visual guidance of what you see is what you get. The doctor does not need to frequently switch between the mouth and the external screen. The doctor can directly align the physical bracket with the virtual outline in the field of vision to complete the bonding, thereby effectively reducing the difficulty of operation, shortening the chairside operation time, and relieving the doctor's visual fatigue and cervical spine pressure. It has the precision of indirect bonding and the convenience of direct bonding. Attached Figure Description

[0021] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention.

[0022] Figure 1 This is a flowchart of a dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality according to the present invention; Figure 2 This is another flowchart of a dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality according to the present invention; Figure 3 This is a block diagram of a dynamic orthodontic bracket navigation system based on multimodal vision and augmented reality according to the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. Obviously, the described embodiments are only some, not all, of the embodiments described in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without creative effort are within the scope of protection of this application.

[0024] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a” and “an” used herein, and “the”, may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0025] First Embodiment Please see Figure 1 and Figure 2 As shown, this embodiment provides a dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality, including the following steps: S1: Obtain the patient's preoperative oral cavity 3D data, generate a 3D digital target model based on the oral cavity 3D data, and extract a set of virtual anatomical feature points from the 3D digital target model.

[0026] Preferably, in step S1, generating a three-dimensional digital target model based on oral cavity three-dimensional data includes: The patient's cone-beam computed tomography (CBCT) scan data and intraoral scan data are imported, and rigid registration and fusion are performed using an iterative nearest-point algorithm to generate a three-dimensional digital model containing the complete anatomical structure of the tooth. Specifically, in this embodiment, the CBCT scan data includes voxel images of deep anatomical structures such as tooth roots, alveolar bone, and jawbone, which are used to establish a complete three-dimensional morphology of the tooth and periodontal tissues. The intraoral scan data provides a high-precision three-dimensional mesh model of the crown surface, reflecting the fine geometric features and color texture of the tooth surface. During the registration process, the crown surface reconstructed based on CBCT and the crown surface scanned intraorally are initially aligned by marking feature points (such as cusps and pits) or recognizing curvature features, so that the two roughly overlap. Virtual tooth alignment is performed on a 3D digital model. The coordinates of the clinical crown center point and the long axis vector of the tooth body of each target tooth are calculated. Based on the coordinates of the clinical crown center point and the long axis vector of the tooth body, the ideal 3D bonding position of the bracket is determined, and a 3D digital target model containing the preset bracket pose information of each tooth is generated. Specifically, in this embodiment, based on the segmented tooth model, the malocclusion characteristics such as crowding, rotation angle, and tilt of the current dentition are analyzed. According to the straight wire orthodontic concept and biomechanical principles, the teeth are simulated to move to the ideal position in the software, and the spatial pose parameters (clinical crown center point and long axis vector of the tooth body) of each tooth in the target position are calculated. The clinical crown center point is the geometric center of the most prominent point on the labial surface of the crown, and the long axis vector of the tooth body represents the axial direction of the tooth in three-dimensional space, which is used to determine the axial tilt and torque of the bracket. A world coordinate system is established, and virtual anatomical feature point sets and virtual bracket bonding markers are extracted from the 3D digital target model. The virtual anatomical feature point set includes at least the cementoenamel boundary point cloud, incisal edge point cloud, or cusp point cloud; the virtual bracket bonding markers include at least the coordinates of the bracket center point or the local coordinate point set of the bracket boundary contour line. Specifically, in this embodiment, the center of the dentition or a specific anatomical landmark is used as the origin of the world coordinate system, and the coordinate axis direction is aligned with the cranial positioning plane (such as the orbitoauricular plane); the cementoenamel boundary point cloud is formed by uniformly sampling along the cervical line of each tooth to form a dense point cloud. This region has a stable morphology and is ideal. Registration anchor points; incisal edge / cusp point cloud, i.e., sampling of the line connecting the highest points of the incisal edge of the anterior teeth or the cusp vertex region of the posterior teeth, used to help determine the vertical position of the teeth; extracting visual guidance data (bracket center point coordinates, bracket boundary contour point set) for AR rendering from the 3D digital target model, the bracket center point coordinates are the precise 3D coordinates of the center point of the bracket base on each tooth; the bracket boundary contour point set is the local coordinate point set formed by uniform sampling along the edge of the bracket base to form a closed contour line, used to render the contour line; in addition, auxiliary guidance information such as the long axis of the tooth and the direction line of the bracket groove can also be extracted as auxiliary identifiers.

[0027] S2: Real-time acquisition of RGB and depth image data from the patient's mouth using augmented reality devices, real-time inference of the RGB and depth image data using a pre-trained multimodal large model, segmentation of the target tooth, and extraction of its real-time anatomical feature point set in the real-time camera coordinate system.

[0028] Preferably, in step S2, real-time inference is performed on the RGB image data and depth image data using a pre-trained multimodal large model, including: A large visual model pre-trained based on a self-supervised learning framework is invoked to extract features from RGB image data and depth map data. The self-supervised learning framework includes a mask autoencoder, whose pre-training objective is to minimize the reconstruction loss of mask image patches. Specifically, in this embodiment, the RGB image is processed by an encoder of a visual Transformer or a convolutional neural network to extract high-dimensional semantic features layer by layer. The shallow layers of the network focus on low-level features such as edges and textures, while the deep layers focus on high-level semantic information such as tooth shape and category, outputting an RGB feature map. The depth map is processed by an independent encoder branch to extract three-dimensional geometric features. The depth map reflects the surface undulations of the object. The encoder learns geometric features such as curvature, gradient, and concavity / convexity from distance information and outputs a depth feature map.

[0029] The formula for calculating reconstruction loss is as follows: Where M is the set of image patches to be masked. For the original image patch, To reconstruct image patches.

[0030] In the encoding stage of the large visual model, the semantic features of the RGB image and the spatial geometric features of the depth image are fused through a cross-attention mechanism to output a high-dimensional feature map. Specifically, in this embodiment, RGB features are used as queries, and depth features are used as keys and values. The attention mechanism calculates which geometric information should be focused on from the depth features at each position in the RGB features. This process enables texture features to obtain shape constraints from geometric information, enhancing robustness to reflective areas. Using depth features as queries and RGB features as keys and values ​​allows geometric features to obtain texture boundary information from RGB features, filling in inferences about depth hole regions. A multi-head attention mechanism is adopted to learn multiple complementary fusion relationships from different subspaces, improving the richness of feature representation. The high-dimensional feature map is restored to a pixel-level semantic segmentation mask by the decoder of the large visual model. The semantic segmentation mask, combined with depth information, maps the two-dimensional pixels back to three-dimensional space, extracting the real-time anatomical feature point set of the real teeth in the current field of view. Specifically, in this embodiment, the decoder gradually restores the spatial resolution of the feature map through deconvolution or interpolation upsampling operations, making it close to the size of the original input image. The last layer of the decoder outputs a probability map of each pixel belonging to different categories (gingiva, oral mucosa, instruments, etc., different tooth positions, such as left upper central incisor, right lower first molar, etc.). For each detected two-dimensional anatomical feature point, the depth value of its corresponding pixel position in the synchronous depth map is obtained. Since pixel-level alignment has been completed in step S1, the depth map can be directly indexed according to the two-dimensional coordinates (u,v), and the distance d of the point can be read. Using the intrinsic parameter matrix of the augmented reality device camera, the two-dimensional pixel coordinates are back-projected to three-dimensional space through the pinhole camera model to obtain the three-dimensional coordinates in the real-time camera coordinate system. , , ,in,( , () is the focal length, ( , ) represents the coordinates of the optical center.

[0031] Preferably, the multimodal large model is optimized using a joint loss function during the downstream task training phase to ensure the segmentation accuracy of the target tooth boundary. The calculation formula for the joint loss function is as follows: ,in, Dice loss is used to improve the overlap of tooth region segmentation. Cross-entropy loss is used to optimize pixel-level classification accuracy.

[0032] S3: Based on virtual anatomical feature point sets With real-time anatomical feature point set (Located in the real-time camera coordinate system) Calculate the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system. Based on the spatial transformation matrix, transform the virtual bracket bonding mark in the 3D digital target model to the real-time camera coordinate system. The transformation relationship satisfies... ,in The coordinates of the virtual bracket bonding mark in the world coordinate system. For rotation matrix, It is a translation vector. These are the coordinates in the transformed real-time camera coordinate system.

[0033] Preferably, in step S3, the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system is calculated, including: Singular value decomposition (SVD) or nonlinear optimization algorithms are used to perform point cloud matching between virtual and real-time anatomical feature point sets with anatomical correspondences. The initial spatial transformation matrix that minimizes the mean square error between the virtual and real-time anatomical feature point sets is then calculated. Specifically, in this embodiment, when point pairs have noise or many outliers, nonlinear least squares algorithms such as Levenberg-Marquardt can be used to minimize the reprojection error or point-to-point distance, resulting in the initial spatial transformation matrix. This serves as the initial pose estimate for the current frame.

[0034] More preferably, a singular value decomposition algorithm is used to perform point cloud matching between the virtual anatomical feature point set with anatomical correspondence and the real-time anatomical feature point set, including: Calculate the centroids of the two point sets: , ; Centroid coordinates: , ; Construct the covariance matrix: ; Perform SVD decomposition on H: ; Rotation matrix: (Make sure the determinant is +1; if it is -1, invert it.) Translation vector: .

[0035] By integrating data from the inertial measurement unit (IMU) built into the augmented reality device (including a three-axis accelerometer (measuring linear acceleration) and a three-axis gyroscope (measuring angular velocity)), high-frequency pose prediction is performed between consecutive video frames using extended Kalman filtering or visual inertial odometry, continuously updating the rotation matrix. With translation vector The inertial measurement unit data includes linear acceleration measured by a triaxial accelerometer and angular velocity measured by a triaxial gyroscope.

[0036] More preferably, after transforming the virtual tray bonding mark to the real-time camera coordinate system, the intrinsic parameter matrix K of the AR device's camera is used to... Projecting onto a two-dimensional image plane yields the pixel coordinates of each marker point in the current camera view: .

[0037] Preferably, calculating the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system also includes a dynamic optimization process: Construct a loss term that includes semantic segmentation and 3D registration loss term The total loss function is calculated using the following formula: , and These are adaptive weighting coefficients; The system monitors the confidence level of visual features in the intraoral field of view in real time. The confidence level is evaluated based on the confidence map of the semantic segmentation mask or the inlier rate of feature point matching. When the target tooth is detected to be obscured by instruments, covered by saliva, or subjected to strong glare, causing the visual feature confidence level to fall below a preset threshold, the system automatically reduces the confidence level. The weight and increase The weights are adjusted so that the system relies more on inertial measurement unit data for pose deduction when visual features are insufficient, thus maintaining the stability of the spatial transformation matrix.

[0038] S4: Using augmented reality equipment, the converted virtual bracket bonding mark is rendered in real time and locked onto the corresponding real tooth surface in the doctor's field of vision in a three-dimensional overlay.

[0039] Preferably, in step S4, real-time rendering and locking onto the corresponding real tooth surface in the doctor's field of vision includes: Based on the binocular camera model of the augmented reality device, the world coordinates of the virtual bracket bonding mark are transformed to the left-eye camera coordinate system and the right-eye camera coordinate system, respectively. Perspective projection is then performed on the transformed coordinates to generate left and right eye display images with binocular parallax.

[0040] For each 3D point First, perform perspective division to obtain the normalized device coordinates: Convert normalized coordinates to screen pixel coordinates: Obtain the 2D pixel position (u,v) of each marker point in the current camera view, and the corresponding depth value. .

[0041] The world coordinates of the virtual bracket bonding mark Transform to the left-eye and right-eye camera coordinate systems respectively to obtain the pixel coordinates on the left and right eye display images. , ) and( , The transformation relationship is as follows: , .

[0042] The virtual bracket bonding mark is occluded using depth image data. By comparing the projection depth of the virtual mark point with the depth value of the real object in the mouth, the rendering transparency or occlusion relationship of the virtual mark is determined.

[0043] More preferably, the rendering transparency or occlusion relationship of the virtual identifier is determined, including: like ( The tolerance threshold (e.g., 0.5mm) indicates that the virtual point is located behind the real tooth and should be occluded. Set the rendering opacity to 0 (no rendering) or very low opacity. like This indicates that the virtual point is located in front of the real tooth and should be rendered normally. Set the opacity to 1 (completely opaque).

[0044] Temporal smoothing filtering is applied to the projection coordinates of the virtual bracket bonding mark between consecutive frames to eliminate rendering graphics jitter caused by pose jitter. Furthermore, display latency is reduced by combining time warp technology with inertial measurement unit data, thus ensuring that the virtual mark is stably locked onto the real tooth surface.

[0045] Preferably, in step S4, the virtual bracket bonding marking includes: The virtual bracket bonding markings include at least one of the bracket outline, center point mark, crosshair, or tooth long axis; the augmented reality device is optical see-through AR glasses, and real-time rendering is achieved through the optical display module of the AR glasses.

[0046] Second Embodiment Please see Figure 3 As shown, based on the same concept, the present invention also provides a dynamic orthodontic bracket navigation system based on multimodal vision and augmented reality, comprising: The three-dimensional target model construction module (i.e., the preoperative data module) is used to acquire the patient's preoperative oral three-dimensional data, generate a three-dimensional digital target model based on the oral three-dimensional data, and extract a set of virtual anatomical feature points from the three-dimensional digital target model; The real-time multimodal feature extraction module (i.e., AR perception module) is used to collect RGB image data and depth image data in the patient's mouth in real time through augmented reality devices, and use a pre-trained multimodal large model to perform real-time inference on the RGB image data and the depth image data to segment the target tooth and extract its real-time anatomical feature point set in the real-time camera coordinate system. The dynamic registration and coordinate transformation module (i.e., the core calculation module) is used to calculate the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system based on the virtual anatomical feature point set and the real-time anatomical feature point set; and to transform the virtual bracket bonding mark in the three-dimensional digital target model to the real-time camera coordinate system according to the spatial transformation matrix. The augmented reality rendering guidance module (i.e., the display module) renders and locks the converted virtual bracket bonding mark onto the corresponding real tooth surface in the doctor's field of vision in real time through the augmented reality device in a three-dimensional overlay manner.

[0047] Third Embodiment Based on the same concept, this embodiment also provides a computer device, including a memory and a processor, wherein the memory stores computer-readable instructions, which, when executed by the processor, cause the processor to perform the steps of a dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality as described in the embodiment.

[0048] Based on the same concept, the present invention also provides a storage medium storing computer-readable instructions, characterized in that, when the computer-readable instructions are executed by one or more processors, the one or more processors cause the one or more processors to perform the steps of a dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality as described in any one of the embodiments.

[0049] It is understood that, regarding the aforementioned dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality, if all of these methods are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer server or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.

[0050] Computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0051] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality, characterized in that, Includes the following steps: S1: Obtain the patient's preoperative oral cavity three-dimensional data, generate a three-dimensional digital target model based on the oral cavity three-dimensional data, and extract a set of virtual anatomical feature points from the three-dimensional digital target model; S2: Real-time acquisition of RGB image data and depth image data in the patient's mouth using augmented reality devices; real-time inference of the RGB image data and depth image data using a pre-trained multimodal large model; segmentation of the target tooth and extraction of its real-time anatomical feature point set in the real-time camera coordinate system. S3: Based on the virtual anatomical feature point set and the real-time anatomical feature point set, calculate the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system. According to the spatial transformation matrix, transform the virtual bracket bonding mark in the three-dimensional digital target model to the real-time camera coordinate system. The transformation relationship satisfies... ,in The coordinates of the virtual bracket bonding mark in the world coordinate system. Let be a rotation matrix. It is a translation vector. The coordinates are in the transformed real-time camera coordinate system; S4: Using the augmented reality device, the converted virtual bracket bonding mark is rendered and locked onto the corresponding real tooth surface in the doctor's field of vision in a three-dimensional overlay manner in real time.

2. The dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality according to claim 1, characterized in that, In step S1, a three-dimensional digital target model is generated based on the oral cavity three-dimensional data, including: Import the patient’s cone-beam computed tomography (CBCT) data and intraoral scan data, and perform rigid registration and fusion using the iterative nearest point algorithm to generate a three-dimensional digital model containing the complete dental anatomy. Virtual tooth arrangement is performed on the three-dimensional digital model, the coordinates of the clinical crown center point and the long axis vector of the tooth body of each target tooth are calculated, and the ideal three-dimensional bonding position of the bracket is determined based on the coordinates of the clinical crown center point and the long axis vector of the tooth body, thereby generating the three-dimensional digital target model containing the preset bracket pose information of each tooth. A world coordinate system is established, and the virtual anatomical feature point set and virtual bracket bonding identifier are extracted from the three-dimensional digital target model. The virtual anatomical feature point set includes at least the enamel-cementum boundary point cloud, incisal edge point cloud, or cusp point cloud; the virtual bracket bonding identifier includes at least the bracket center point coordinates or the local coordinate point set of the bracket boundary contour line.

3. The dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality according to claim 1, characterized in that, In step S2, a pre-trained multimodal large model is used to perform real-time inference on the RGB image data and the depth image data, including: The visual large model pre-trained based on a self-supervised learning framework is invoked to extract features from the RGB image data and depth map data. The self-supervised learning framework includes a mask autoencoder, whose pre-training objective is to minimize the reconstruction loss of the mask image patch. During the encoding stage of the large visual model, the semantic features of the RGB image and the spatial geometric features of the depth image are fused through a cross-attention mechanism to output a high-dimensional feature map. The high-dimensional feature map is restored to a pixel-level semantic segmentation mask by the decoder of the large visual model. The semantic segmentation mask, combined with depth information, maps two-dimensional pixels back to three-dimensional space and extracts the real-time anatomical feature point set of the real teeth in the current field of view. The formula for calculating the reconstruction loss is as follows: Where M is the set of image patches to be masked. For the original image patch, To reconstruct image patches.

4. The dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality according to claim 3, characterized in that, The multimodal large model is optimized using a joint loss function during the downstream task training phase to ensure the segmentation accuracy of the target tooth boundary. The calculation formula for the joint loss function is as follows: ,in, Dice loss is used to improve the overlap of tooth region segmentation. Cross-entropy loss is used to optimize pixel-level classification accuracy.

5. The dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality according to claim 1, characterized in that, In step S3, the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system is calculated, including: Using singular value decomposition or nonlinear optimization algorithm, point cloud matching is performed on the virtual anatomical feature point set and the real-time anatomical feature point set with anatomical correspondence, and the initial spatial transformation matrix that minimizes the mean square error between the virtual anatomical feature point set and the real-time anatomical feature point set is obtained. By integrating data from the inertial measurement unit built into the augmented reality device, and using extended Kalman filtering or visual inertial odometry, high-frequency pose prediction is performed between consecutive video frames, continuously updating the rotation matrix. With translation vector .

6. The dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality according to claim 5, characterized in that, Calculating the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system also includes: Construct a loss term that includes semantic segmentation and 3D registration loss term The total loss function is calculated using the following formula: , and These are adaptive weighting coefficients; The system continuously monitors the confidence level of visual features within the intraoral field of view. This confidence level is evaluated based on the confidence map of a semantic segmentation mask or the inlier rate of feature point matching. When the target tooth is detected to be obscured by instruments, covered in saliva, or subject to strong glare, causing the visual feature confidence level to fall below a preset threshold, the system automatically reduces the confidence level. Weight and increase The weights are adjusted so that the system relies more on inertial measurement unit data for pose deduction when visual features are insufficient, thus maintaining the stability of the spatial transformation matrix.

7. The dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality according to claim 1, characterized in that, In step S4, the real-time rendering and locking onto the corresponding real tooth surface in the doctor's field of vision includes: Based on the binocular camera model of the augmented reality device, the world coordinates of the virtual bracket bonding mark are transformed to the left-eye camera coordinate system and the right-eye camera coordinate system, respectively. Perspective projection is performed on the transformed coordinates to generate left and right eye display images with binocular parallax. The virtual bracket bonding mark is occluded using the depth image data. The rendering transparency or occlusion relationship of the virtual mark is determined by comparing the projection depth of the virtual mark point with the depth value of the real object inside the mouth. Temporal smoothing filtering is applied to the projection coordinates of the virtual bracket bonding mark between consecutive frames to eliminate rendering graphics jitter caused by pose jitter. Furthermore, display latency is reduced by combining time warp technology with inertial measurement unit data, thus ensuring that the virtual mark is stably locked onto the real tooth surface.

8. The dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality according to claim 1, characterized in that, In step S4, the virtual bracket bonding mark includes: The virtual bracket bonding markings include at least one of the bracket outline, center point mark, crosshair, or tooth long axis; the augmented reality device is optical see-through AR glasses, and the real-time rendering is achieved through the optical display module of the AR glasses.

9. A dynamic orthodontic bracket navigation system based on multimodal vision and augmented reality, characterized in that, include: A three-dimensional target model construction module is used to acquire the patient's preoperative oral three-dimensional data, generate a three-dimensional digital target model based on the oral three-dimensional data, and extract a set of virtual anatomical feature points from the three-dimensional digital target model; The real-time multimodal feature extraction module is used to collect RGB image data and depth image data in the patient's mouth in real time through augmented reality devices, and use a pre-trained multimodal large model to perform real-time inference on the RGB image data and the depth image data to segment the target tooth and extract its real-time anatomical feature point set in the real-time camera coordinate system. The dynamic registration and coordinate transformation module is used to calculate the spatial transformation matrix from the world coordinate system to the real-time camera coordinate system based on the virtual anatomical feature point set and the real-time anatomical feature point set; and to transform the virtual bracket bonding mark in the three-dimensional digital target model to the real-time camera coordinate system according to the spatial transformation matrix. The augmented reality rendering guidance module, through the augmented reality device, renders and locks the converted virtual bracket bonding mark in real time onto the corresponding real tooth surface in the doctor's field of vision in a three-dimensional overlay manner.

10. A computer device, characterized in that, The system includes a memory and a processor, wherein the memory stores computer-readable instructions that, when executed by the processor, cause the processor to perform the steps of the dynamic orthodontic bracket navigation method based on multimodal vision and augmented reality as described in any one of claims 1 to 8.