Digital mandibular movement trajectory fusion reconstruction method and system
By using a regular smartphone and a simple mirror to achieve simultaneous dual-view data acquisition, combined with AI technology, the problem of existing equipment being complex, costly, and unsuitable for home use has been solved. This enables markerless and non-invasive fusion reconstruction of mandibular movement trajectories, supporting home self-testing and remote diagnosis and treatment, and is applicable to a variety of oral medical applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-21
AI Technical Summary
Existing mandibular motion acquisition devices suffer from problems such as complex equipment, high cost, inability to be used at home, difficulty in remote diagnosis and treatment, and insufficient monocular visual reconstruction accuracy. They cannot meet the needs of clinical teaching and primary healthcare for home self-testing, remote diagnosis and treatment, low cost, easy operation, and multi-view synchronization.
Using a single camera from a regular smartphone and a simple reflector, it achieves simultaneous acquisition of frontal and side views. Combined with AI-powered automatic reconstruction, it completes the markerless and non-invasive reconstruction of the mandibular movement trajectory through tooth semantic segmentation and key point matching, supporting remote diagnosis and treatment via the cloud.
It enables patients to complete mandibular movement video acquisition independently at home, without the need for professional personnel. It is simple to operate, low in cost, and has high reconstruction accuracy. It supports remote diagnosis and treatment and three-dimensional trajectory reconstruction, and is suitable for oral medicine students' practical training, temporomandibular joint disease screening, and orthodontic follow-up.
Smart Images

Figure CN122435153A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of digital oral diagnosis and treatment technology, and in particular relates to a digital mandibular motion trajectory fusion and reconstruction method and system. Background Technology
[0002] Existing mandibular motion acquisition devices are mainly divided into two categories: contact and non-contact. Contact devices, such as mechanical bows and ultrasonic mandibular motion analyzers, require wearing intraoral devices, which have drawbacks such as being invasive, having poor comfort, having accuracy affected by wearing them, and being unsuitable for home use. Non-contact devices mostly use multi-camera arrays, binocular vision, or structured light devices, which generally suffer from problems such as being expensive, bulky, and having complicated calibration. Furthermore, multi-camera perspective recording has synchronization errors, requires professional operation, cannot be self-service, and does not support remote diagnosis and treatment.
[0003] Existing technologies include a markerless mandibular motion dynamic acquisition method that proposes using a multi-view high-speed camera system combined with crown feature points and a 3D model to achieve markerless mandibular motion acquisition. While this approach can achieve high-precision motion acquisition, it still has significant shortcomings: 1. A multi-camera array must be used, which is complex, bulky, and costly. Furthermore, multi-view cameras have synchronization errors, making calibration complex. 2. Multi-camera calibration is cumbersome, requires intraoral scanning, and needs to be operated by professionals, making it unsuitable for patients to collect data themselves; 3. It cannot enable remote diagnosis and treatment or home self-testing, and can only be used in fixed clinics; the hardware threshold is high, and it lacks universality and portability, making it difficult to popularize in primary healthcare institutions and ordinary families. Summary of the Invention
[0004] To overcome the shortcomings of the existing technologies, this invention provides a digital mandibular motion trajectory fusion and reconstruction method and system. It can achieve simultaneous acquisition of frontal and lateral dual-view images using only a single camera of an ordinary smartphone and a simple reflector. This enables label-free, non-invasive, AI-automated reconstruction, and cloud-based remote diagnosis and treatment of digital mandibular motion trajectory fusion and reconstruction. The single camera simultaneously acquires frontal and lateral dual-view images, which makes up for the insufficient accuracy of monocular vision three-dimensional reconstruction. Patients can complete the mandibular motion video acquisition themselves in a home environment without the need for professional personnel to be present.
[0005] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of this invention provides a method for digital mandibular motion trajectory fusion and reconstruction.
[0006] The digital mandibular motion trajectory fusion and reconstruction method includes the following steps: Simultaneous video captures both a frontal intraoral image and a side-view mirror image of the patient using a single camera; Identify and separate the patient's frontal intraoral region and its mirrored lateral region in a synchronized video frame; Semantic segmentation is performed on the teeth in the intraoral region of the frontal view and the mirrored region of the side view, and the first dentition region in the intraoral region of the frontal view and the second dentition region in the mirrored region of the side view are extracted. Key point information matching and fusion are performed on the first and second dentition regions to obtain a fused image; Based on the fused images, the three-dimensional pose of the patient's mandible is solved, and the mandibular motion trajectory is reconstructed.
[0007] A second aspect of the present invention provides a digital mandibular motion trajectory fusion and reconstruction system.
[0008] A digital mandibular motion trajectory fusion and reconstruction system includes: The data acquisition module is configured to simultaneously acquire a synchronous video containing both a frontal intraoral image and a side-view mirror image of the patient using a single camera. The view segmentation module is configured to: identify and separate the patient's frontal intraoral region and the side mirror region in the synchronized video frame; The semantic segmentation module is configured to perform semantic segmentation on the teeth in the frontal intraoral region and the side mirror region, and extract the first dentition region in the frontal intraoral region and the second dentition region in the side mirror region. The matching and fusion module is configured to: perform key point information matching and fusion on the first and second dental arch regions to obtain a fused image; The solution and reconstruction module is configured to: solve the three-dimensional pose of the patient's mandible based on the fused image, and complete the reconstruction of the mandibular motion trajectory.
[0009] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of the digital mandibular motion trajectory fusion and reconstruction method as described in the first aspect of the present invention.
[0010] A fourth aspect of the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the digital mandibular motion trajectory fusion reconstruction method as described in the first aspect of the present invention.
[0011] The above one or more technical solutions have the following beneficial effects: This invention provides a digital mandibular motion trajectory fusion and reconstruction method and system, enabling a single camera to simultaneously capture dual perspectives of the patient's oral cavity, namely, a frontal view and a side view. By acquiring synchronized video containing both the patient's frontal intraoral image and a mirrored side image, the mandibular motion trajectory is reconstructed using tooth semantic segmentation, key point matching and fusion, and 3D pose solving techniques. Patients can independently complete mandibular motion video acquisition at home, without the need for professional personnel, specialized equipment, complex calibration, intraoral scanning, or the application of markers; the operation is simple. Furthermore, the simultaneous acquisition of frontal and side views by a single camera compensates for the insufficient accuracy of monocular vision 3D reconstruction.
[0012] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0013] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0014] Figure 1 This is a flowchart illustrating the method steps of Example 1.
[0015] Figure 2 This is a flowchart of the method in Example 1.
[0016] Figure 3 This is a diagram showing the relative positions of a simplified rearview mirror and a human face in Example 1.
[0017] Figure 4 This is a schematic diagram illustrating the dual-view effect of a single camera simultaneously capturing the frontal and side views of a patient's oral cavity, as shown in Example 1.
[0018] Figure 5 This is a schematic diagram of the tooth image segmentation result in Example 1.
[0019] Figure 6 This is a schematic diagram of the tooth key point localization results in Example 1.
[0020] Figure 7 This is a schematic diagram of the trajectory output in Example 1.
[0021] The attached diagram lists the components represented by each number as follows: 1. Simple reflector; 2. Human face; 3. Same shooting scene. Detailed Implementation
[0022] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0023] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0024] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0025] Example 1 As mentioned earlier, existing methods for acquiring mandibular motion data cannot meet the needs of clinical teaching and primary healthcare for home-based self-testing, remote diagnosis and treatment, low cost, ease of operation, and simultaneous multi-view synchronization. To address the shortcomings of existing mandibular motion acquisition devices, such as complexity, high cost, inability to be used at home, difficulty in remote diagnosis and treatment, insufficient accuracy of monocular visual reconstruction, and poor synchronization of multiple cameras, this invention provides a digital mandibular motion trajectory fusion reconstruction method that uses only a single camera from a regular smartphone with a simple reflector to achieve simultaneous acquisition of frontal and lateral views of the patient's oral cavity, is markerless and non-invasive, features AI-automated reconstruction, and enables remote diagnosis and treatment via the cloud. This achieves the following objectives: 1. Patients can complete the mandibular movement video recording independently at home without the need for professional personnel to be present; 2. A single camera simultaneously acquires frontal and side views, compensating for the insufficient accuracy of monocular vision in 3D reconstruction; 3. No specialized equipment, complex calibration, intraoral scanning, or marker points are required; operation is simple. 4. By combining deep learning to achieve the fusion of tooth segmentation and key point detection, 3D trajectory reconstruction can be automatically completed; 5. Construct a complete remote diagnosis and treatment process, including video capture, encrypted data upload, cloud-based AI analysis, online doctor diagnosis, and report delivery; 6. Supports the fusion of CBCT data to match motion trajectories with anatomical structures, providing precise data for subsequent treatment; 7. Applicable to practical training for dental students, screening for temporomandibular joint disorders, and follow-up after orthodontic and prosthodontic surgery.
[0026] like Figure 1 As shown, the digital mandibular motion trajectory fusion and reconstruction method includes the following steps: Simultaneous video captures both a frontal intraoral image and a side-view mirror image of the patient using a single camera; Identify and separate the patient's frontal intraoral region and its mirrored lateral region in a synchronized video frame; Semantic segmentation is performed on the teeth in the intraoral region of the frontal view and the mirrored region of the side view, and the first dentition region in the intraoral region of the frontal view and the second dentition region in the mirrored region of the side view are extracted. Key point information matching and fusion are performed on the first and second dentition regions to obtain a fused image; Based on the fused images, the three-dimensional pose of the patient's mandible is solved, and the mandibular motion trajectory is reconstructed.
[0027] The implementation process of the above technical solution in this embodiment, in more detail, includes: Further: such as Figure 3 and Figure 4 As shown, a single camera simultaneously acquires a synchronized video containing both a frontal intraoral image and a mirrored side image of the patient. Specifically, this includes: The simple reflector 1 component is positioned at an appropriate angle on one side of the patient's face to present the lateral movement of the patient's jaw when performing standard movements; Position the single camera directly at the patient's face; Ensure that in the single camera's shooting frame, the same shooting frame 3 simultaneously includes the patient's frontal intraoral image and side mirror image, providing both frontal and side views; Specifically, the frontal intraoral image is defined as an image of the patient's oral cavity, dental arch, and mandibular frontal movement area directly captured by the camera; the lateral mirror image is defined as an image of the patient's oral cavity, mandibular lateral movement, and condylar movement projection area presented after reflection by a mirror. The standard movements include opening and closing the mouth at a constant speed 3 times, extending forward once, and moving to the left and right sides once each.
[0028] Furthermore: the simplified reflector assembly includes a plane reflector, the angle between the plane reflector and one side of the patient's face being 30°–45°.
[0029] Furthermore, the single camera simultaneously provides dual perspectives, capturing both the frontal and side views of the patient's oral cavity.
[0030] In this embodiment, the mobile phone acquisition terminal is a regular smartphone with a single camera that supports high-definition video recording. It is used to capture synchronized video including a frontal intraoral image and a side mirror image. The simple reflector assembly is a plane reflector, positioned at a fixed angle of 30°–45° on one side of the face. It reflects the lateral movement of the jaw to the mobile phone camera, allowing the single camera to simultaneously obtain frontal and side views in the same frame.
[0031] Figure 3 The diagram shown illustrates the relative positions of the simplified reflector 1 and the face 2 in this embodiment. Figure 4 The diagram shown is a schematic of the dual-view effect of a single camera simultaneously capturing the frontal and side views of the patient's oral cavity in this embodiment. Figure 3In the process, a simple reflector 1 is set on one side of the face 2. The simple reflector 1 and the face 2 on the side closest to the simple reflector 1 are at a fixed angle of 30°–45°, so that the movement of the lower jaw side is reflected in the simple reflector 1, and thus the shooting effect of presenting two perspectives simultaneously using the same shooting scene 3 of a single camera is achieved.
[0032] Figure 4 In the middle, the single camera uses the camera that comes with a regular smartphone, and the shooting effect can be clearly seen. In the same shooting scene 3, both the frontal and side views of the oral cavity are presented at the same time.
[0033] Furthermore, it also includes image preprocessing: The side mirror area is mirrored and flipped to restore the true view. Perform image distortion correction, scale normalization, tooth region enhancement, and noise reduction; Perform inter-frame alignment and timing synchronization to ensure that the timestamps of the frontal intraorific area and the side mirror area are consistent.
[0034] This step involves automatic dual-view region segmentation, identifying and separating the frontal intraoral region and the mirrored lateral region in the image. Mirror flip correction is performed by horizontally flipping the lateral view to restore the true perspective. Image distortion correction, scale normalization, and tooth region enhancement and noise reduction are also performed. Inter-frame alignment and temporal synchronization ensure consistent timestamps between the two views.
[0035] In specific dual-view automatic region segmentation, the U-Net-AsppAtt model (which integrates dilated spatial convolutional pooling and attention mechanisms) is used for region segmentation.
[0036] 1. Feature extraction and multi-scale perception: The U-Net encoder structure is used to extract image depth features. An ASPP (Atrous Spatial Pyramid Pooling) module is introduced into the high-level feature map. Parallel processing of dilated convolutions with different sampling rates is used to increase the receptive field to adapt to the differences in tooth size caused by distance changes in the video, ensuring complete capture of the large frontal field of view and the local side mirror.
[0037] The Atrous Spatial Pyramid Pooling (ASPP) module is widely used in deep learning, particularly in computer vision tasks, especially image segmentation. The ASPP module was originally proposed in the DeepLab series of papers to improve the accuracy and performance of image segmentation.
[0038] How the ASPP module works: The ASPP module captures multi-scale contextual information by using atrous convolutions with different sampling rates in parallel. Atrous convolution (also known as dilated convolution) is a variant of standard convolution that expands the receptive field without increasing the number of parameters by inserting "holes" (i.e., zero-padding) into the convolution kernel.
[0039] Main components: Parallel dilated convolutions: The ASPP module contains multiple dilated convolutional layers that execute in parallel, each using a different sampling rate (i.e., dilation rate). For example, common configurations include dilated convolutions with sampling rates of 1, 6, 12, and 18.
[0040] Image-level features: There is also a global average pooling layer, which averages all the pixel values of the feature map to obtain a single value, and then transforms this value into a vector with the same number of channels as the original feature map through a 1x1 convolution.
[0041] Merging features: The outputs of all these layers (including the output of the global average pooling layer) are concatenated to form a larger feature map, which is then typically reduced in number of channels by a 1x1 convolution to facilitate integration with subsequent network layers.
[0042] Implementation steps: Global average pooling: Performs global average pooling on the entire feature map.
[0043] Multiple parallel dilated convolutions: Dilated convolutions with different sampling rates are applied to the feature map.
[0044] Concatenation with 1x1 Convolution: Concatenate all outputs and reduce the number of channels by a 1x1 convolution.
[0045] Upsampling: An optional step that upsamples the feature map to the size of the original input so as to align it with the image segmentation mask at the original resolution.
[0046] 2. Enhanced attention mechanism (Attention Gate): An attention mechanism is introduced at the skip connection of the decoder to automatically suppress the feature weights of non-tooth areas such as facial skin, mirror edges, and background, and focus on the high-response dental pixel areas.
[0047] 3. Region determination logic: After the model outputs the mask, it performs separation based on geometric priors, identifying the central region of the image and the larger connected regions as the frontal in-mouth region; and identifying the smaller connected regions on one side of the image that conform to mirror geometry as the side mirror region.
[0048] The final segmentation result is as follows Figure 5 As shown.
[0049] Further: Key point information matching and fusion are performed on the first and second dentition regions to obtain a fused image, specifically including: Key point detection is performed on the first dentition region to obtain the first key point sequence of the first dentition, which covers all key locations in the first dentition region. Key point detection is performed on the second dentition region to obtain the second key point sequence of the second dentition, which covers all key locations in the second dentition region; Match the corresponding key points of the first key point sequence and the second key point sequence: Based on the consistency of the same timestamp and semantic index of the two views, automatically match the i-th key point in the frontal oral region with the i-th key point in the side mirror region (i∈[1,35]). The fusion of the matched first keypoint sequence and the second keypoint sequence: Based on the 30°–45° arrangement angle of the reflector and the mirror flip correction parameters, the side mirror 2D coordinates are projected onto the virtual side camera space to obtain the projection coordinates P'_side of the side mirror 2D coordinates in the virtual side camera space. A confidence-weighted fusion algorithm is used, with the formula P_fused = w1. P_front+w2 P'_side, where P_fused represents the fused keypoint coordinates, P_front represents the keypoint coordinates corresponding to the frontal intraoral region, and w1 and w2 represent weights, which are determined by the keypoint confidence scores; After fusion, the data is smoothed by Kalman filtering to eliminate inter-frame pixel jitter and ensure continuous and smooth trajectory.
[0050] In the above process: Based on the consistency of the same timestamp and semantic index from both perspectives, the i-th key point in the frontal intraoral region is automatically matched one-to-one with the i-th key point in the side mirror region. The specific process includes: When the timestamps and semantic indexes are consistent, the automatic matching of the i-th key point in the frontal oral region and the i-th key point in the side mirror region is completed. When timestamps or semantic indexes are inconsistent, the i-th key point in the frontal oral region does not match the i-th key point in the side mirror region.
[0051] Based on the 30°–45° arrangement angle of the reflector and the mirror flip correction parameters, the 2D coordinates of the side mirror are projected onto the virtual side camera space to obtain the projection coordinates P'_side of the 2D coordinates of the side mirror in the virtual side camera space. In this process, based on the geometric imaging principle of the reflector and the camera projection model, the mirror view is converted into the coordinate projection under the virtual side camera coordinate system. This is a mature monocular dual-view coordinate transformation method in the field of computer vision and is a conventional technical means in this field.
[0052] Keypoint confidence can be determined using existing mature methods such as deep learning heatmap regression maximum value, keypoint response score, and Gaussian distribution probability value.
[0053] It can be understood that the first and second key points mentioned above are a descriptive method to facilitate the distinction between whether a key point originates from the first or second dentition region. In essence, both the first and second key points mentioned above are predefined from 35 key points, such as... Figure 6 As shown, the 35 key points specifically include: 20 points in the maxilla: midpoint of the gingival margin on the labial / buccal surface of the maxillary central incisor, lateral incisor, canine, first premolar, edge point of the contact area of the dental arch, and cusp of the canine.
[0054] Seventeen points on the mandible: the lowest point of the contact area of the mandibular dentition, the midpoint of the gingival margin on the labial / buccal surface of the central incisor, lateral incisor, canine, and first premolar.
[0055] The 35 key points are defined using existing mature methods for defining key points of the dental arch. They are predefined in combination with oral anatomical features and computer vision tracking requirements, taking into account both motion tracking stability and 3D reconstruction accuracy.
[0056] Using ResNet-50 as the backbone network and connecting a heatmap regression head at the end of the network, high-precision positioning of 35 key points can be achieved, which can effectively handle tooth overlap, occlusion and mirror distortion interference during opening and closing.
[0057] In this step, a deep learning model is trained based on a dedicated oral dentition dataset to achieve tooth semantic segmentation, accurately extract the dentition region, and eliminate interference from lips, cheeks, tongue, and background; key point detection outputs 35 key points of the dentition, covering key positions of the entire dentition; temporal tracking is performed to stably track key points in continuous video frames to reduce jitter and loss.
[0058] Furthermore: Based on the fused images, the patient's three-dimensional mandibular pose is solved to reconstruct the mandibular movement trajectory, specifically including: The PnP algorithm was used to solve the three-dimensional pose of the mandible, and the triangulated reconstructed three-dimensional point cloud and motion trajectory were obtained. By using temporal smoothing filtering, denoising, and interpolation, a continuous and smooth three-dimensional mandibular motion trajectory is obtained.
[0059] Since the PnP algorithm is a mature method in existing technology, its application to solving the three-dimensional pose of the mandible in fused images will not be elaborated here.
[0060] In simple terms, in this embodiment, the EPnP algorithm is used to solve the three-dimensional pose of the mandible. The three-dimensional anatomical coordinates of the key points of the teeth are used as 3D reference points, and the fused 2D key points of the current frame are used as projection points. The rotation vector R and translation vector T are solved by combining the camera intrinsic parameter matrix. Through continuous frame pose change calculation, the triangulated reconstructed three-dimensional point cloud and motion trajectory are obtained, and the dynamic motion trajectory reconstruction of the mandible relative to the maxilla is completed.
[0061] In the trajectory reconstruction process, the key point information matching and fusion of dual perspectives are first realized. Then, the three-dimensional pose of the mandible is solved based on the PnP algorithm to obtain the triangulated reconstructed three-dimensional point cloud and motion trajectory. Finally, temporal smoothing filtering, denoising and interpolation are performed to output a continuous and smooth three-dimensional trajectory.
[0062] It also includes: automatically calculating quantitative diagnostic indicators: mouth opening, skewness, lateral displacement distance, motion symmetry, velocity, acceleration, and trajectory smoothness.
[0063] like Figure 7 As shown, the system displays the mandibular motion trajectory curve and three-dimensional reconstruction results from the front and side views. The system can select mandibular motion videos and CT scans for import. It supports the calculation of quantitative diagnostic indicators such as mouth opening, deviation, lateral displacement distance, motion symmetry, motion speed, acceleration, and trajectory smoothness, and can export auxiliary diagnostic reports.
[0064] This also includes fusing the obtained three-dimensional mandibular motion trajectory with multi-source data: Import patient CBCT data or intraoral scan data to construct a personalized three-dimensional dental / condylar model; The ICP registration algorithm is used to accurately match the reconstructed mandibular motion trajectory with the CBCT model, enabling visualization of the motion trajectory on the real anatomical structure, estimation of the condylar motion center position, and automatic identification of occlusal interference points and joint risk points.
[0065] During CBCT model fusion: It supports importing patient CBCT data or intraoral scan data to construct personalized 3D dentofacial / condylar models. Through the ICP registration algorithm, the reconstructed mandibular motion trajectory is accurately matched with the CBCT model, achieving: visualization of the motion trajectory on the real anatomical structure; calculation of the condylar motion center position; automatic identification of occlusal interference points and joint risk points; and providing data support for orthodontic, restorative, and occlusal plate design.
[0066] When this method is applied to remote diagnosis and treatment, the patient's end has functions such as video capture guidance, encrypted data upload, report viewing, and historical record management. The doctor's end has functions such as 3D trajectory visualization, online annotation, historical data comparison, and diagnosis editing. The cloud is responsible for AI inference, data storage, access control, and automatic generation of diagnostic reports. The system supports temporomandibular joint disorder screening, post-orthodontic / restorative surgery follow-up, and practical training assessment for dental students.
[0067] This invention discloses an intelligent diagnostic method for non-invasive acquisition and three-dimensional trajectory reconstruction of mandibular movements that can be performed at home using only a regular smartphone without the need for dedicated hardware. It combines digital oral diagnosis and treatment, computer vision, deep learning, telemedicine, monocular three-dimensional reconstruction, and label-free motion capture technology.
[0068] Finally, a specific application scenario of this embodiment is given: like Figure 2 As shown, this embodiment, in its specific implementation, includes a hardware layer, an image preprocessing layer, a dataset construction layer, a tooth image segmentation layer, a key point detection layer, a trajectory reconstruction layer, an automatic calculation of quantitative diagnostic indicators, a CBCT model fusion layer, and a remote diagnosis and treatment layer. Among them: Hardware layer: Single-camera video recording on mobile phones, with a simple reflector positioned at a fixed angle of 30°-45° on one side of the face, allowing the single camera to simultaneously obtain both front and side views within the same frame.
[0069] Image preprocessing layer: Automatic dual-view region segmentation; implements mirror flip correction, image distortion correction, scale normalization, region enhancement, and noise reduction. Inter-frame alignment and temporal synchronization ensure consistent timestamps between the two views.
[0070] Dataset Construction: Obtain frontal and lateral mandibular movement video data, construct a semantic segmentation and key point detection dataset for deep learning model training.
[0071] U-Net tooth image segmentation model: accurately extracts the dentition region and eliminates background interference from the lips, cheeks, tongue, and palate.
[0072] ResNet-50 key point detection model: outputs multiple key points of the dental arch and locates their precise coordinates.
[0073] Trajectory Reconstruction Layer: Matching and fusing key point information from both perspectives. Temporal smoothing filtering, denoising, and interpolation output a continuous and smooth 3D trajectory. Automatically calculates quantitative diagnostic indicators: mouth opening, skewness, lateral displacement distance, motion symmetry, velocity, acceleration, and trajectory smoothness.
[0074] CBCT Model Fusion Layer (Optional Enhancement Module): Supports importing patient CBCT data or intraoral scan data to build personalized models and visualize motion trajectories on real anatomical structures; provides data support for orthodontic, prosthodontic, and occlusal plate design.
[0075] Remote diagnosis and treatment layer: The patient's end features video capture guidance, encrypted data upload, report viewing, and historical record management. The doctor's end features 3D trajectory visualization, online annotation, historical data comparison, and diagnosis editing. The cloud layer is responsible for AI inference, data storage, access control, and automatic generation of diagnostic reports.
[0076] More detailed: 1. Hardware setup: Use a 50mm×80mm flat mirror with an angle of 30°–45°, placed on one side of the face; 2. The phone camera is pointed directly at the face to ensure that the same image simultaneously includes both the front view of the mouth and a side mirror image; 3. Video parameters: 1080P resolution, 30fps frame rate, recording duration 8 seconds; 4. Standard movements: Open and close the mouth 3 times at a constant speed, extend forward once, and move to the left and right sides once each; 5. Upload: The video is encrypted and uploaded to the cloud server; 6. Dual-view processing: Automatically identifies and segments the front and side regions, and performs mirror flipping, scale normalization, and geometric correction; 7. AI Model: Lightweight segmentation network and key point detection network are jointly used to achieve temporally stable tracking with inference speed ≤300ms / frame; 8. Trajectory Reconstruction: Dual-view information fusion → PnP pose solving → 3D trajectory reconstruction → quantification index calculation; CBCT fusion is also supported (optional): Import DICOM format CBCT data to complete dental model segmentation and ICP registration; 9. Output results: Quantitative indicators such as three-dimensional motion trajectory curve, mouth opening, skewness, lateral displacement, symmetry, and motion speed.
[0077] 10. Remote diagnosis: Doctors view 3D trajectory and quantitative data → online annotation and diagnosis → generate electronic reports; the reports are pushed to the patient's end, supporting historical comparison and long-term follow-up.
[0078] The technical solution of this embodiment has the following technical advantages: 1. The hardware is extremely simple and the cost is very low, requiring only a mobile phone and a rearview mirror. No professional equipment is needed, making it suitable for widespread use. 2. Home-based, non-invasive self-help, no markings, no contact, no doctor's presence required, suitable for all ages including children and the elderly; 3. Monocular dual-view synchronization, breaking through the technical bottleneck of low reconstruction accuracy of single camera; 4. No calibration or scanning required; ordinary users can complete data collection within 1 minute, making it extremely easy to operate. 5. AI fully automated processing, from video to 3D trajectory, is fully automated and requires no human intervention; 6. Comprehensive quantitative indicators directly support temporomandibular joint assessment, orthodontic follow-up, and teaching and training. 7. A complete remote diagnosis and treatment closed loop, truly realizing remote diagnosis, primary care screening, and long-term follow-up; 8. Optional CBCT fusion enables precise matching of motion trajectory and anatomical structure, enhancing clinical treatment value; 9. It can be scaled up and is compatible with multiple terminals such as mini-programs, apps, and web, making it suitable for use in hospitals, primary care institutions, and families.
[0079] Example 2 This embodiment discloses a digital mandibular motion trajectory fusion and reconstruction system.
[0080] A digital mandibular motion trajectory fusion and reconstruction system includes: The data acquisition module is configured to simultaneously acquire a synchronous video containing both a frontal intraoral image and a side-view mirror image of the patient using a single camera. The view segmentation module is configured to: identify and separate the patient's frontal intraoral region and the side mirror region in the synchronized video frame; The semantic segmentation module is configured to perform semantic segmentation on the teeth in the frontal intraoral region and the side mirror region, and extract the first dentition region in the frontal intraoral region and the second dentition region in the side mirror region. The matching and fusion module is configured to: perform key point information matching and fusion on the first and second dental arch regions to obtain a fused image; The solution and reconstruction module is configured to: solve the three-dimensional pose of the patient's mandible based on the fused image, and complete the reconstruction of the mandibular motion trajectory.
[0081] Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.
[0082] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the digital mandibular motion trajectory fusion reconstruction method as described in Embodiment 1 of this disclosure.
[0083] Example 4 The purpose of this embodiment is to provide an electronic device.
[0084] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the digital mandibular motion trajectory fusion reconstruction method as described in Embodiment 1 of this disclosure.
[0085] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0086] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0087] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A digital mandibular motion trajectory fusion and reconstruction method, characterized in that, Includes the following steps: Simultaneous video captures both a frontal intraoral image and a side-view mirror image of the patient using a single camera; Identify and separate the patient's frontal intraoral region and its mirrored lateral region in a synchronized video frame; Semantic segmentation is performed on the teeth in the intraoral region of the frontal view and the mirrored region of the side view, and the first dentition region in the intraoral region of the frontal view and the second dentition region in the mirrored region of the side view are extracted. Key point information matching and fusion are performed on the first and second dentition regions to obtain a fused image; Based on the fused images, the three-dimensional pose of the patient's mandible is solved, and the mandibular motion trajectory is reconstructed.
2. The digital mandibular motion trajectory fusion and reconstruction method as described in claim 1, characterized in that, Simultaneous video capture using a single camera, including both a frontal intraoral image and a mirrored lateral image of the patient, specifically includes: A simple reflector assembly is placed at an appropriate angle on one side of the patient's face to present the lateral movement of the patient's jaw when performing standard movements; Position the single camera directly at the patient's face; Ensure that in a single camera's capture frame, the same frame simultaneously includes a frontal intraoral image of the patient and a side mirror image, providing both frontal and side views; Specifically, the frontal intraoral image is defined as an image of the patient's oral cavity front, dental arch, and mandibular front movement area directly captured by the camera, while the lateral mirror image is defined as an image of the patient's oral cavity side, mandibular lateral movement, and condylar movement projection area presented after reflection by a reflector. The standard movements include opening and closing the mouth at a constant speed 3 times, extending forward once, and moving to the left and right sides once each.
3. The digital mandibular motion trajectory fusion and reconstruction method as described in claim 2, characterized in that, The simplified reflector assembly includes a plane reflector with an angle of 30°–45° between the plane reflector and one side of the patient's face.
4. The digital mandibular motion trajectory fusion and reconstruction method as described in claim 1, characterized in that, The single camera has dual perspectives, capturing both the frontal and side views of the patient's mouth.
5. The digital mandibular motion trajectory fusion and reconstruction method as described in claim 1, characterized in that, This also includes image preprocessing: The side mirror area is mirrored and flipped to restore the true view. Perform image distortion correction, scale normalization, tooth region enhancement, and noise reduction; Perform inter-frame alignment and timing synchronization to ensure that the timestamps of the frontal intraorific area and the side mirror area are consistent.
6. The digital mandibular motion trajectory fusion and reconstruction method as described in claim 1, characterized in that, Key point information matching and fusion are performed on the first and second dentition regions to obtain a fused image, specifically including: Key point detection is performed on the first dentition region to obtain the first key point sequence of the first dentition, which covers all key locations in the first dentition region. Key point detection is performed on the second dentition region to obtain the second key point sequence of the second dentition, which covers all key locations in the second dentition region; Match the corresponding key points of the first key point sequence and the second key point sequence: Based on the consistency of the same timestamp and semantic index of the two views, the i-th key point in the frontal intraoral region and the i-th key point in the side mirror region are automatically matched one-to-one (i∈[1,35]). The fusion of the matched first keypoint sequence and the second keypoint sequence: Based on the 30°–45° arrangement angle of the reflector and the mirror flip correction parameters, the side mirror 2D coordinates are projected onto the virtual side camera space to obtain the projection coordinates P'_side of the side mirror 2D coordinates in the virtual side camera space. A confidence-weighted fusion algorithm is used, with the formula P_fused = w1. P_front+w2 P'_side, where P_fused represents the fused keypoint coordinates, P_front represents the keypoint coordinates corresponding to the frontal intraoral region, and w1 and w2 represent weights, which are determined by the keypoint confidence scores; After fusion, the data is smoothed by Kalman filtering to eliminate inter-frame pixel jitter and ensure continuous and smooth trajectory.
7. The digital mandibular motion trajectory fusion and reconstruction method as described in claim 1, characterized in that, Based on the fused images, the three-dimensional pose of the patient's mandible is solved, and the mandibular motion trajectory is reconstructed, specifically including: The PnP algorithm was used to solve the three-dimensional pose of the mandible, and the triangulated reconstructed three-dimensional point cloud and motion trajectory were obtained. By using temporal smoothing filtering, denoising, and interpolation, a continuous and smooth three-dimensional mandibular motion trajectory is obtained. This also includes fusing the obtained three-dimensional mandibular motion trajectory with multi-source data: Import patient CBCT data or intraoral scan data to construct a personalized three-dimensional dental / condylar model; The ICP registration algorithm is used to accurately match the reconstructed mandibular motion trajectory with the CBCT model, enabling visualization of the motion trajectory on the real anatomical structure, estimation of the condylar motion center position, and automatic identification of occlusal interference points and joint risk points.
8. A digital mandibular motion trajectory fusion and reconstruction system, characterized in that, include: The data acquisition module is configured to simultaneously acquire a synchronous video containing both a frontal intraoral image and a side-view mirror image of the patient using a single camera. The view segmentation module is configured to: identify and separate the patient's frontal intraoral region and the side mirror region in the synchronized video frame; The semantic segmentation module is configured to perform semantic segmentation on the teeth in the frontal intraoral region and the side mirror region, and extract the first dentition region in the frontal intraoral region and the second dentition region in the side mirror region. The matching and fusion module is configured to: perform key point information matching and fusion on the first and second dental arch regions to obtain a fused image; The solution and reconstruction module is configured to: solve the three-dimensional pose of the patient's mandible based on the fused image, and complete the reconstruction of the mandibular motion trajectory.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the digital mandibular motion trajectory fusion and reconstruction method as described in any one of claims 1-7.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the digital mandibular motion trajectory fusion and reconstruction method as described in any one of claims 1-7.