Label-free dynamic acquisition method of mandibular movement

By using a pose calculation method based on feature extraction from unlabeled video images and 3D tooth models, the problems of cumbersome operation and data distortion in mandibular motion acquisition using traditional mechanical devices and modern optical acquisition systems are solved, achieving high-precision mandibular motion reconstruction and occlusion assessment.

CN120852474BActive Publication Date: 2025-12-05NANJING XIAOLING TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511348990.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-12-05
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Traditional mechanical devices and modern optical acquisition systems are cumbersome to operate and have a poor wearing experience when collecting mandibular movements. They also suffer from data distortion caused by loose or obstructed markers, making it impossible to effectively record the dynamic information of the mandible.

Method used

A label-free method was adopted to extract features and calculate poses using video images and 3D tooth models. A 3D gingival arch model was constructed using intraoral 3D scanner and CBCT image data. Combined with the YOLOv5 model and PnP algorithm, label-free 3D reconstruction of mandibular movements was achieved.

Benefits of technology

Data collection is performed under natural conditions, eliminating human interference, improving the physiological accuracy of measurements and user experience, maintaining high-precision reconstruction of mandibular movement trajectory, and providing high-quality clinical occlusal assessment data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852474B_ABST
    Figure CN120852474B_ABST
Patent Text Reader

Abstract

The application provides a label-free mandibular movement dynamic acquisition method, and relates to the field of dynamic reconstruction of jaw position relationship.The application is completely based on video images and tooth three-dimensional models for feature extraction and posture calculation, does not need label points or mechanical device intervention, and can be completed by a subject in a natural state, eliminates human interference sources, and significantly improves physiological measurement accuracy and user experience.A multi-view high-speed camera system is combined with a dental crown point topological structure modeling, multi-view space reconstruction and time sequence compensation are used, stable tracking and high-precision reconstruction are still maintained under the conditions of occlusion and motion blur, and the acquisition failure risk is effectively avoided.The method is based on tooth array overall modeling and posture calculation, can output complete mandibular movement trajectory curves, can accurately express complex movements such as mastication, deviation, occlusion, mandibular lateral rotation, and the like, and provides high-quality original data for clinical occlusion evaluation generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic reconstruction of jaw position relationship, and in particular to a markerless method for dynamic acquisition of mandibular movement. Background Technology

[0002] The recording of mandibular movements has important clinical value in the fields of oral prosthodontics, orthodontics, and maxillofacial surgery.

[0003] Traditional mechanical devices such as facebows and articulators can obtain certain jaw position relationships, but the operation process is cumbersome, the wearing experience is not good, and they can only record static information, which limits their accuracy.

[0004] Although modern optical acquisition systems have introduced marker points and multi-camera tracking technology to improve dynamic capture capabilities, they still inevitably face problems such as wear interference, loose or obstructed marker points, and the resulting data distortion. Summary of the Invention

[0005] Purpose of the invention: To propose a markerless dynamic acquisition method for mandibular movements, which extracts features and calculates posture based on video images and three-dimensional tooth models. It does not require the intervention of markers or mechanical devices, and the subject can complete the acquisition in a natural state, eliminating human interference sources and effectively solving the problems of data distortion caused by loose or occluded markers in existing technologies.

[0006] The present invention proposes a markerless dynamic acquisition method for mandibular movements, comprising the following steps:

[0007] S1. Oral scan data of the subject is obtained by an intraoral 3D scanner, and supplemented by CBCT image data to form a 3D gingival dentition model;

[0008] S2. The three-dimensional gingival dental arch model is segmented, the boundaries of each tooth and gingival tissue are extracted, and the segmented CBCT image data and oral scan data are rigidly registered to the same coordinate system.

[0009] S3. In the segmented and registered three-dimensional gingival dentition model, mark the three-dimensional reference points of each tooth to form a candidate set of anchor points;

[0010] S4. Record image sequences of the subject's jaw movements at multiple shooting points;

[0011] S5. Using the YOLOv5 model combined with the feature point detection algorithm, the image sequence is processed frame by frame to detect and locate the two-dimensional pixel position of the tooth anchor point. Based on the confidence of the detection results, several key points with the highest confidence in the maxillary and mandibular teeth are selected as the motion anchor points of the frame.

[0012] S6. Combine the two-dimensional coordinates of the motion anchor points of each frame image with their corresponding points in the three-dimensional model, and use the PnP algorithm to estimate the three-dimensional spatial pose of the mandible in the current frame.

[0013] S7. Integrate the three-dimensional spatial poses of all mandibles from different shooting points in chronological order to construct a complete mandibular movement trajectory, thereby achieving markerless three-dimensional reconstruction of the dynamic movement process of the subject's mandible.

[0014] In a further embodiment, step S1 specifically includes:

[0015] S1-1. Instruct the subject to maintain a stable, naturally open-mouth posture, and use an intraoral 3D scanner to scan the maxillary dentition, mandibular dentition, and gingival tissue in sequence to obtain complete intraoral scan data.

[0016] S1-2. Use cone-beam CT equipment to scan the subject's head or maxillofacial region to obtain CBCT image data including teeth, tooth roots, alveolar bone, maxilla and mandible and surrounding bony structures.

[0017] S1-3. Standardize the CBCT image data into a unified voxel format, and simultaneously align the intraoral scan data to the coordinates.

[0018] In a further embodiment, step S3 specifically includes:

[0019] S3-1, Denote the point cloud set of each tooth as... , where p is the three-dimensional coordinate of a point on the i-th tooth;

[0020] S3-2, Calculate the normal vector of point p. It is then compared with the unit normal vector k of the occlusal plane to select the set of occlusal surface points that meet the conditions. :

[0021]

[0022] in, For point The normal vector, The unit normal vector of the occlusal plane. For the preset threshold, For the first A collection of point clouds of individual teeth;

[0023] S3-3, In the set of occlusal surface points Internal calculation of local curvature and the point set According to curvature value Sort by largest to smallest:

[0024]

[0025] The n cusps with the highest curvature are selected as the set of cusp reference points for the tooth. :

[0026]

[0027] S3-4. Obtain the boundary points of each tooth by detecting mutations in the local neighborhood normal vector, thus obtaining the set of boundary points. :

[0028]

[0029] S3-5. Combine the curvature cusp with the boundary points to obtain the complete three-dimensional reference set of the i-th tooth:

[0030]

[0031] S3-6, The set of three-dimensional reference points for all teeth is represented as follows:

[0032]

[0033] in, For the i-th tooth, This is the corresponding set of three-dimensional reference points.

[0034] In a further embodiment, in step S4, multiple cameras are installed evenly distributed around the shooting area to simultaneously capture image sequences of the subject opening and closing his mouth, biting, extending forward, and sliding laterally from different perspectives.

[0035] In a further embodiment, in step S5, each candidate ROI output by the YOLOv5 model is processed... Location loss, The classification loss and confidence loss are optimized, with bounding box regression being one of the key components. Loss is defined as:

[0036]

[0037] Overall losses for:

[0038]

[0039] In the formula, The coordinates of the center point of the prediction box. The coordinates of the center point of the true bounding box. For prediction boxes With real frame The intersection and union ratio, The squared Euclidean distance between the center of the predicted bounding box and the center of the ground truth bounding box in the pixel coordinate system. It is the square of the diagonal length of the smallest bounding rectangle that can simultaneously enclose both the ground truth bounding box and the predicted bounding box. This is a parameter for maintaining consistency in aspect ratio between the predicted bounding box and the ground truth bounding box. This is a penalty coefficient used to adjust... The weights; These are the weighting coefficients for the bounding box regression loss. The weighting coefficients for the target confidence loss. For confidence loss, These are the weighting coefficients for the classification loss. For classification loss.

[0040] In a further embodiment, in step S5, after all candidate boxes are detected, non-maximum suppression is first performed according to the category to remove redundant predicted boxes; for each category, all candidate boxes are sorted from high to low according to the prediction confidence (conf), and the box with the highest confidence is selected. As the current retained box, it is then compared with each of the remaining candidate boxes. Crossover ratio:

[0041]

[0042] If the intersection-union ratio is greater than the preset threshold, the two boxes are considered to have overlapping heights and the low-confidence boxes are deleted; repeat this process until all candidate boxes have been processed.

[0043] After nonmaximum suppression is completed, the set of maxillary tooth key points U and the set of mandibular tooth key points L corresponding to the three-dimensional reference points are denoted as follows:

[0044]

[0045]

[0046] in, Indicates the first Key maxillary points and their corresponding detection boxes Indicates the first Each key point of the mandible is assigned a corresponding bounding box, and each key point is assigned a confidence level. ;

[0047] Sort sets U and L separately according to their confidence scores from highest to lowest, and select at least three key points with the highest confidence scores from each set to form two subsets:

[0048]

[0049]

[0050] in, and The specific key points can originate from the same tooth or be distributed among different teeth.

[0051] Will and The set of motion anchor points is obtained by merging. .

[0052] In a further embodiment, step S5 further includes:

[0053] For sets For each maxillary and mandibular motion anchor point, Harris corner detection was performed within the maxillary and mandibular regions of interest (ROIs) respectively. Local geometric feature points highly correlated with the cusp and boundary positions within the ROIs were extracted, and gradients were calculated for the images within the ROIs. Construct a Gaussian weighted second-order matrix:

[0054]

[0055] In the formula, The images are respectively in Gradient in direction; The standard deviation is Gaussian kernel;

[0056] Calculate the Harris response function:

[0057]

[0058] In the formula, Let be the determinant of the matrix, representing the product property of the gradient in two directions; The trace of the matrix, These are empirical constants;

[0059] Non-maximum suppression and thresholding within the ROI The set of candidate corner points of the maxilla is obtained. and the set of candidate angle points for the mandible ; Three-dimensional reference points for each tooth Given that the camera calibration matrix K, rotation matrix R, and translation vector t are projected onto the current frame image plane, the pixel coordinates of the template keypoints are obtained:

[0060]

[0061]

[0062] in, This represents the normalization transformation from homogeneous coordinates to pixel coordinates; The pixel coordinates of key points on the maxillary template; These are the pixel coordinates of key points on the mandibular template.

[0063] In a further embodiment, step S6 specifically includes:

[0064] S6-1. Set the two-dimensional coordinates of the maxillary key points detected in each frame of the image sequence. Two-dimensional coordinate set of key points of the mandible The registered 3D gingival arch model is mapped to a known set of 3D reference points.

[0065]

[0066]

[0067] in, This represents the position vector of the i-th maxillary key point in the 3D gingival dentition model. This represents the position vector of the j-th mandibular key point in the three-dimensional gingival dentition model;

[0068] S6-2, Intrinsic parameter matrix obtained from camera calibration The PnP algorithm is used to solve the maxillary pose matrix of the current frame. and jaw frame pose matrix :

[0069]

[0070]

[0071] In the formula, Let be the rotation matrix of the maxilla, representing the maxillary reference frame relative to the . The rotational relationship of the camera coordinate system; Let be the translation vector of the maxilla, representing the origin of the maxillary reference frame at the _i_th ... The position of the camera in the coordinate system; Let be the rotation matrix of the mandible, representing the mandibular reference frame relative to the _____. The rotational relationship of the camera coordinate system; Let be the translation vector of the mandible, representing the origin of the mandibular reference frame at the _i_th ... The position of the camera in the coordinate system;

[0072] S6-3, Minimize reprojection error:

[0073]

[0074]

[0075] in, Let be the two-dimensional pixel coordinates of the i-th maxillary key point in the image captured by the c-th camera; Let J be the two-dimensional pixel coordinates of the j-th mandibular key point in the image captured by the c-th camera; , These are the coordinates of the maxillary and mandibular reference points in the corresponding 3D model.

[0076] In a further embodiment, in the process of minimizing reprojection error, all three-dimensional points are first processed through EPnP. Represented as four virtual control points Linear combination:

[0077]

[0078] In the formula, These are known coefficients, determined by the geometric relationship between the world coordinate system and the control point system; Depend on or Substitute;

[0079] Substitute into the projection equation:

[0080]

[0081] in, As a scale factor, Let R be the homogeneous coordinates of the pixel, R be the rotation matrix, and t be the translation vector;

[0082] All matching points form a system of linear equations, which can be solved using Singular Value Decomposition (SVD). And derive the rotation matrix. With translation vector ; the result of EPNP As initial values, the LM algorithm is used to solve the problem:

[0083]

[0084] in, Depend on or Substitute;

[0085] After optimization using the LM algorithm, a global pose that meets the predetermined accuracy requirements is obtained. and ;

[0086] Based on the global pose, calculate the movement of the mandible relative to the maxilla:

[0087]

[0088] in, and Let represent the global pose matrices of the maxilla and mandible in the camera's c-coordinate system, respectively. This indicates the rotational relationship between the lower jaw and the upper jaw. This represents the translation vector of the mandible relative to the maxilla.

[0089] In a further embodiment, step S7 specifically includes:

[0090] S7-1. Perform fusion processing on the relative poses obtained from different cameras, and fuse the rotation matrices of each camera. Convert to quaternion And a uniform rotation result is calculated using spherical mean squares:

[0091]

[0092] in, ( () represents the spherical distance between quaternions. These are the weighting coefficients;

[0093] S7-2, The translation part is based on the reprojection error of each camera. The fused pose is obtained by weighting the inverse weights and summing the values. ;

[0094]

[0095] S7-3. Perform time-series smoothing on the fused attitude, using exponential smoothing filtering for the translation component:

[0096]

[0097] S7-4. Rotational components achieve continuous transitions using quaternion spherical linear interpolation:

[0098]

[0099] in, and Control the smoothness of translation and rotation separately;

[0100] S7-5. Obtain the set of three-dimensional motion trajectories of the mandible relative to the maxilla in the video sequence in chronological order. .

[0101] Compared with the prior art, the present invention has the following beneficial effects:

[0102] (1) Feature extraction and pose calculation are based entirely on video images and three-dimensional tooth models. No markers or mechanical devices are required. Subjects can complete the data collection in a natural state, eliminating human interference sources and significantly improving the accuracy of physiological measurements and user experience.

[0103] (2) A three-view high-speed camera system was adopted in combination with crown point topology modeling. Multi-view spatial reconstruction and temporal compensation were used to maintain stable tracking and high-precision reconstruction under occlusion and motion blur conditions, effectively avoiding the risk of acquisition failure.

[0104] (3) Based on the overall modeling and posture calculation of the dental arch, the complete mandibular motion trajectory curve can be output, which can accurately express complex movements such as chewing, deviation, occlusion, and mandibular lateral rotation, providing high-quality raw data for clinical occlusion assessment.

[0105] (4) The system design achieves integrated data link, automatically aligning the crown feature points with the CBCT model, and combining spatial trajectory and timestamp information to achieve fully automatic spatiotemporal registration of "tooth model, CT skeleton, and motion trajectory". A high-fidelity, continuous virtual mandibular motion system can be generated without manual intervention, greatly improving system stability and doctor's work efficiency. Attached Figure Description

[0106] Figure 1 This is a flowchart of the markerless mandibular motion dynamic acquisition method of the present invention.

[0107] Figure 2 This is a schematic diagram of the Unet-ASPP structure of the present invention.

[0108] Figure 3 This is a schematic diagram of the YOLOv5 structure of the present invention.

[0109] Figure 4 This is a segmentation result diagram of the three-dimensional model of the present invention. Detailed Implementation

[0110] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention can be practiced without one or more of these details. In other instances, certain technical features well-known in the art have not been described in order to avoid obscuring the invention.

[0111] This embodiment provides a markerless dynamic acquisition method for mandibular movements, the flowchart of which is shown below. Figure 1 As shown, the specific execution steps are as follows:

[0112] Step 1: A high-precision 3D model of the subject's complete dentition and gingiva is obtained using an intraoral 3D scanner, supplemented by cone-beam computed tomography (CBCT) image data.

[0113] Step 2: Segment the CBCT images and intraoral scan data, extract the boundaries of each tooth and gingival tissue, and rigidly register the segmented CBCT data and intraoral scan data to unify them into the same coordinate system.

[0114] Step 3: In the segmented and registered model, label the three-dimensional reference points of each tooth to form a candidate set of anchor points.

[0115] Step four: Use multi-view video to record the subject's actual mandibular movements, such as opening and closing the mouth, biting, and lateral movements, and collect image sequences.

[0116] Step 5: Using the YOLOv5 model combined with the feature point detection algorithm, the image sequence is processed frame by frame to detect and locate the two-dimensional pixel position of the tooth anchor point. Based on the confidence of the detection results, at least three key points with the highest confidence in the maxillary and mandibular teeth are selected as the motion anchor points of the frame.

[0117] Step 6: Combine the two-dimensional coordinates of the motion anchor points of each frame with their corresponding points in the three-dimensional model, and use the PnP algorithm to estimate the three-dimensional spatial pose of the mandible in the current frame.

[0118] Step 7: Integrate the three-dimensional mandibular poses of all video frames in chronological order to construct a complete mandibular motion trajectory, thereby achieving label-free three-dimensional reconstruction of the subject's dynamic mandibular motion process.

[0119] As a preferred approach, step one involves acquiring a high-precision 3D model of the subject's complete dentition and gingiva using an intraoral 3D scanner, supplemented by cone-beam computed tomography (CBCT) image data. The specific process is as follows:

[0120] First, the subjects were instructed to maintain a stable, naturally open-mouth posture. An intraoral 3D scanner was then used to scan the maxillary dentition, mandibular dentition, and gingival tissue in sequence to obtain data on a complete 3D surface model.

[0121] Next, cone-beam CT scanners were used to scan the subject's head or maxillofacial region to obtain three-dimensional volumetric image data including teeth, tooth roots, alveolar bone, maxilla and mandible and surrounding bony structures.

[0122] After completing the oral scan and CBCT acquisition, the obtained data undergoes quality checks and format conversion. The CBCT images are standardized into a unified voxel format, and the oral scan data is aligned to the coordinates.

[0123] As a preferred approach, step two involves segmenting the CBCT images and intraoral scan data, extracting the boundaries of each tooth and gingival tissue, and rigidly registering the segmented CBCT and intraoral scan data to a single coordinate system. The specific process is as follows:

[0124] First, the CBCT images undergo standardization processing, including voxel value normalization, size interpolation resampling, and data enhancement. The normalization operation linearly maps the original voxel gray values ​​to the [0,1] interval.

[0125]

[0126] in, This represents the original CBCT voxel values. The value after normalization. and These are the maximum and minimum gray levels of the image.

[0127] like Figure 2 As shown, the image segmentation model employs a UNet-based symmetric encoder-decoder structure, and introduces a dilated spatial pyramid pooling (ASPP) module in each skip connection to enhance multi-scale contextual feature representation capabilities. The encoder part consists of multiple layers. The system employs convolutional layers, CBAM attention mechanisms, and residual connections to progressively extract deep features. The decoder fuses these features with corresponding skip connections through upsampling operations, and also incorporates... Convolution, CBAM, and residual connections gradually restore spatial resolution; the ASPP module performs convolution operations in parallel at different dilation rates, fusing multi-scale information to effectively expand the receptive field and improve segmentation accuracy. The ASPP module is as follows:

[0128]

[0129] in, Indicates the input feature map; Indicates use Regular convolution performed by kernels; Indicates the value with expansion rate r Aperture convolution; , , This represents multiple different dilation rates, used to extract contextual information under different receptive fields; This means concatenating multiple convolutional outputs along the channel dimension.

[0130] Following this, CBCT data was fed into the model for training. During the training phase, manually annotated voxel-level label maps were used as supervision signals, with labels including different tissue categories such as maxilla, mandible, teeth, and gingiva. To balance class imbalance and optimize segmentation accuracy, a hybrid loss function was constructed by combining Dice loss, Focal loss, and boundary loss.

[0131]

[0132] in:

[0133] , , These are the weighting coefficients for each loss term, and + + =1;

[0134] The Dice loss, used to measure the overlap between the predicted region and the ground truth label, is defined as:

[0135]

[0136] for The loss, used to reduce the contribution of easily classified samples to the overall loss, focuses on difficult-to-classify samples and is defined as:

[0137]

[0138] Boundary loss, used to improve the accuracy of edge recognition, is defined as:

[0139]

[0140] The total number of voxels involved in the calculation. To predict probabilities, ϵ represents the true label and the numerically stable term. It is the modulation factor. It is the class weight coefficient (used when positive and negative samples are uneven). For real labels The generated signed distance function measures the error distance between the predicted value and the true boundary.

[0141] After training, the entire CBCT image is used as input, and the network outputs a three-dimensional probability map. ,in The number of channels represents the structural category. The segmentation label map is obtained by extracting the channel with the highest probability for each element.

[0142]

[0143] Subsequently, the predicted label map was subjected to 3D connected component extraction and morphological processing to remove small artifact regions and smooth tissue boundaries, ultimately obtaining structured segmentation results including each tooth and gingiva.

[0144] Based on geometric curvature analysis and morphological watershed algorithms, high-precision segmentation of scanning data is achieved. For example... Figure 4 The image shown is a segmentation result of one of the 3D models.

[0145] First, curvature estimation is performed on each vertex of the 3D scanning model to measure local geometric abrupt changes. The mean curvature estimation is used, and the formula is as follows:

[0146]

[0147] in, For the first One vertex, Let it be the set of its first-order neighborhood vertices. , The angle between opposite sides of an adjacent triangular face. Is with The area of ​​the adjacent Voronoi region.

[0148] curvature This represents the degree of geometric abrupt change at each vertex. The interdental spaces typically have high curvature. To more efficiently apply image segmentation algorithms, a three-dimensional curvature map is used. The projection is a two-dimensional pseudo-grayscale image. The definition is as follows:

[0149]

[0150] in This indicates the 3D vertices The depth mapping function projected onto the imaging plane.

[0151] Subsequently, the curvature diagram Image enhancement operations are performed to smooth noise, and local extremum detection is used to extract seed points in the curvature peak region:

[0152]

[0153] in, This represents the local neighborhood centered on the seed candidate point. For the selected first Seed points.

[0154] Obtaining the seed point set Then, the two-dimensional watershed algorithm is used to transform the gaps between teeth into explicit boundaries. First, the curvature map is inverted:

[0155]

[0156] Calculate the Euclidean distance graph:

[0157]

[0158] Distance map Applying the watershed transform, minimize the boundary gradient:

[0159]

[0160] This process completes the image region expansion, obtaining two-dimensional label images for each tooth region. These two-dimensional label images are then mapped back into the original three-dimensional network to achieve precise data segmentation. Based on the position of each triangular facet projected onto the two-dimensional image, it is labeled as the corresponding tooth label.

[0161]

[0162] in, For the third in the 3D mesh A triangular facet, ; This is the inverse mapping from 2D image coordinates to a 3D mesh.

[0163] Rigid registration is performed between the intraoral scan model and the CBCT image to place them in a unified three-dimensional coordinate system. This registration is performed on each three-dimensional point in the intraoral scan model. Apply a rigid transformation to make it spatially correspond to the points in the CBCT image Keep it as consistent as possible.

[0164] A rigid transformation can be expressed as:

[0165]

[0166] in, These are the coordinates of a point in the original oral scan model; For CBCT images and The corresponding target point; These are the points used for registration after transformation; Let be a rotation matrix, satisfying , T is the translation vector.

[0167] The goal of registration is to find the optimal R and T values ​​that minimize the sum of squared distances between the two sets of points, i.e.:

[0168]

[0169] The optimal R and T are obtained using the Singular Value Decomposition (SVD) method, as follows:

[0170] (1) Calculate the centroids of the two sets of points and construct a centralized point set:

[0171]

[0172]

[0173] (2) Construct the covariance matrix and perform singular value decomposition:

[0174]

[0175] (3) Obtain the optimal R and T:

[0176]

[0177]

[0178] Finally, rotation R and translation T are applied to all points in the oral scan model to achieve spatial alignment with the CBCT image.

[0179] As a preferred approach, in the segmented and registered model, the three-dimensional reference points of each tooth are labeled to form a candidate set of anchor points. The specific process is as follows:

[0180] The point cloud set of each tooth is denoted as:

[0181]

[0182] Where p is the three-dimensional coordinate of a point on the i-th tooth.

[0183] Calculate the normal vector of point p It is then compared with the unit normal vector k of the occlusal plane to select the set of occlusal surface points that meet the conditions. :

[0184]

[0185] in, For point The normal vector, The unit normal vector of the occlusal plane. For the preset threshold, For the first A collection of point clouds representing individual teeth.

[0186] In the set of occlusal facets Internal calculation of local curvature and the point set According to curvature value Sort by largest to smallest:

[0187]

[0188] Then, the n points with the highest curvature are selected as the set of reference points for the tooth cusp:

[0189] ={ }

[0190] Furthermore, the boundary points of each tooth are obtained through local neighborhood normal vector mutation detection, resulting in a set of boundary points:

[0191]

[0192] Finally, by combining the curvature cusps and boundary points, a complete three-dimensional reference set for the i-th tooth is obtained:

[0193]

[0194] The set of three-dimensional reference points for all teeth is represented as:

[0195]

[0196] in, For the i-th tooth, This is the corresponding set of three-dimensional reference points.

[0197] As a preferred option, the upper and lower jaws can also be paired up in pairs, and the specific process is as follows:

[0198] In the segmented and registered 3D dental arch model, the maxillary and mandibular teeth are paired one by one based on the FDI tooth position coding system. The FDI coding system uses a two-digit number, where the first digit indicates the quadrant number, with 1 and 2 for the left and right sides of the maxilla, and 3 and 4 for the right and left sides of the mandible. The second digit indicates the sequence number of the tooth in that quadrant, with 1 for the central incisor and so on up to 8 for the third molar.

[0199] Let the set of maxillary teeth be:

[0200]

[0201] The mandibular teeth assembly is as follows:

[0202]

[0203] The FDI code for each tooth is denoted as: or The pairing relationship between the upper and lower jaws can be expressed as:

[0204]

[0205] in, The difference represents the difference in the numbering of the corresponding quadrants of the upper and lower jaws in the FDI coding, for example... , .

[0206] After establishing the pairing relationships, stable 3D reference points are extracted for each tooth. Reference points are selected from the cusps of the teeth on the occlusal surface region that possess stable geometric features. First, the set of occlusal surface points is filtered based on the angle between the normal vector and the occlusal plane. :

[0207]

[0208] in, For point The normal vector, The unit normal vector of the occlusal plane. This is a preset threshold.

[0209] Subsequently Internal calculation of local curvature And select the point of maximum curvature as the three-dimensional reference point of the tooth:

[0210]

[0211] Finally, all paired teeth and their corresponding three-dimensional reference point coordinates are stored as a candidate set of anchor points, represented as:

[0212]

[0213] in, and These are the three-dimensional coordinates of the reference points for the maxilla and mandible, respectively.

[0214] As a preferred approach, step four involves using multi-view video recording to capture the subject's actual mandibular movements, including opening and closing the mouth, biting, and lateral movements, and then acquiring image sequences. The specific process is as follows:

[0215] Multiple industrial-grade high-speed cameras were evenly distributed around the shooting area at approximately 120° angles to simultaneously capture image sequences of the subject's mandibular movements, including opening and closing, biting, protrusion, and lateral sliding, from different perspectives. Before formal data acquisition, the multi-camera system required precise calibration, including camera intrinsic and extrinsic parameters, to ensure spatial consistency and 3D reconstruction accuracy between subsequent images. Intrinsic parameter calibration used the Zhang calibration method, acquiring the internal imaging parameters of each camera, including focal length, by capturing images of a planar calibration board in different poses. , Main point location Radial and tangential distortion coefficients Its model is:

[0216]

[0217] Extrinsic parameter calibration estimates the relative pose relationships between cameras, i.e., the rotation matrix of each camera relative to the reference coordinate system. With the translation vector T, from:

[0218]

[0219] in, It is a three-dimensional point in the camera coordinate system. It is a point in the world coordinate system.

[0220] Following this, to ensure data consistency and the accuracy of subsequent attitude estimation, the camera needs to be fixed on a calibrated rigid support, and uniform exposure and frame rate settings should be used. Exposure time approximately During the data acquisition, the subjects kept their heads relatively still, only making natural jaw movements, while exposing the intraoral dentition area to ensure that key tooth anchor points were always visible. The acquisition system used frame numbers as an index to save multi-view image frames for each moment, forming a complete time-series data.

[0221] As a preferred approach, step five utilizes the YOLOv5 architecture combined with a feature point detection algorithm to process the video images frame by frame, detecting and locating the two-dimensional pixel positions of the tooth anchor points in each frame. Based on the confidence level, at least three key points with the highest confidence are automatically selected as motion anchor points. The specific process is as follows:

[0222] like Figure 3 As shown, the YOLOv5 framework in this embodiment consists of an input layer, a backbone network, a feature fusion network (Neck), and a head. The input layer first performs Mosaic data augmentation, combined with adaptive anchor box calculation and adaptive image filling to improve adaptability to targets at different scales. The backbone network includes depthwise separable convolutions (DWConv2d), a multi-layer BottleneckCSP structure, and a dilated spatial pyramid pooling (ASPP) module to progressively extract and enhance multi-scale features. The Neck network employs a path aggregation structure combining multi-layer BottleneckCSP and feature concatenation (Concat) to achieve the fusion and transfer of features at different scales. The head contains three detection branches at different scales, each consisting of a convolutional layer (Conv2d) and a decoding module, predicting candidate boxes at that scale.

[0223] For each captured image frame, the YOLOv5 model is first used to detect all candidate regions of tooth anchor points for which maxillary and mandibular pairings have been established. Each pair of maxillary and mandibular teeth is output as an independent bounding box, and its pairing group number is recorded in the category label. This is used to maintain the correspondence between maxillary and mandibular anchor points during training and inference. Each candidate bounding box output by the model... pass Location loss, The classification loss and confidence loss are optimized, with bounding box regression being one of the key components. Loss is defined as:

[0224]

[0225] The total loss is:

[0226] +

[0227] in, The coordinates of the center point of the prediction box. The coordinates of the center point of the true bounding box. For prediction boxes With real frame The intersection and union ratio, The squared Euclidean distance between the center of the predicted bounding box and the center of the ground truth bounding box in the pixel coordinate system. It is the square of the diagonal length of the smallest bounding rectangle that can simultaneously enclose both the ground truth bounding box and the predicted bounding box. This is a parameter for maintaining consistency in aspect ratio between the predicted bounding box and the ground truth bounding box. This is a penalty coefficient used to adjust... The weight. These are the weighting coefficients for the bounding box regression loss. The weighting coefficients for the target confidence loss. For confidence loss, These are the weighting coefficients for the classification loss. For classification loss.

[0228] After detecting all candidate boxes, non-maximum suppression (NMS) is first performed on each category to remove redundant predicted boxes. For each category, all candidate boxes are sorted from highest to lowest prediction confidence (conf), and the box with the highest confidence is selected. As the current retained box, it is then compared with each of the remaining candidate boxes. Crossover ratio:

[0229]

[0230] like If the two boxes are considered to overlap in height, the low-confidence boxes are deleted; this process is repeated until all candidate boxes have been processed. After NMS is completed, the sets of maxillary tooth key points and mandibular tooth key points corresponding to the detected 3D reference points (composed of curvature cusps and boundary points) are denoted as follows:

[0231]

[0232]

[0233] in, Indicates the first Key maxillary points and their corresponding detection boxes Indicates the first Each key point of the mandible is assigned a corresponding bounding box, and each key point is assigned a confidence level. .

[0234] Subsequently, sets U and L were sorted from highest to lowest confidence level, and at least three key points with the highest confidence level were selected from each set to form two subsets:

[0235]

[0236]

[0237] in, and The specific key points can originate from the same tooth or be distributed among different teeth.

[0238] Ultimately, and The resulting set of motion anchor points is obtained by merging:

[0239]

[0240] For sets For each maxillary and mandibular motion anchor point, Harris corner detection was performed within the maxillary and mandibular regions of interest (ROIs) respectively. Local geometric feature points highly correlated with the cusp and boundary positions within the ROIs were extracted, and gradients were calculated for the images within the ROIs. Construct a Gaussian weighted second-order matrix:

[0241]

[0242] In the formula, The images are respectively in Gradient in direction; The standard deviation is Gaussian kernel;

[0243] Calculate the Harris response function:

[0244]

[0245] In the formula, Let be the determinant of the matrix, representing the product property of the gradient in two directions; The trace of the matrix, These are empirical constants;

[0246] Non-maximum suppression and thresholding within the ROI The set of candidate corner points of the maxilla is obtained. and the set of candidate angle points for the mandible ; Three-dimensional reference points for each tooth Given that the camera calibration matrix K, rotation matrix R, and translation vector t are projected onto the current frame image plane, the pixel coordinates of the template keypoints are obtained:

[0247]

[0248]

[0249] in, This represents the normalization transformation from homogeneous coordinates to pixel coordinates; The pixel coordinates of key points on the maxillary template; These are the pixel coordinates of key points on the mandibular template.

[0250] To ensure accurate correspondence, corner points detected within the ROI are... , Calculate the corresponding SIFT descriptors respectively and and points in the corresponding template key point domain and To perform the matching, a ratio test was used:

[0251]

[0252]

[0253] in, and Represent the i-th and i-th corner points detected within the maxillary and mandibular ROIs in the current frame, respectively. Two-dimensional pixel coordinates of each corner point and This is the SIFT descriptor vector corresponding to that corner point. and These are the coordinates of the k-th key point of the maxilla and mandible templates in the template library. and For its corresponding SIFT descriptor, and These are the indices of the template keypoints that are closest and second closest to the Euclidean distance of the current frame descriptor, respectively. Represents the Euclidean distance between vectors. This is the threshold for the ratio test.

[0254] Retain the matching points that pass the ratio test to obtain the set of matching feature points:

[0255]

[0256]

[0257] The coarse localization points detected by YOLOv5 are weighted and fused with the points obtained through Harris and SIFT filtering to obtain the final two-dimensional pixel coordinates of each point. Fine localization points for the maxilla and mandible are calculated separately.

[0258]

[0259]

[0260] Define the weighting coefficients:

[0261]

[0262]

[0263] The final fusion formula is:

[0264]

[0265]

[0266] in, and They represent the upper jaw and lower jaw, respectively. , This is the set of matching angle points between the maxillary and mandibular cusps after screening using the SIFT ratio test. , This represents the number of corner points in the corresponding set. , The first maxillary digit in the set The two-dimensional pixel coordinates of the j-th corner point and the j-th corner point of the mandible , These are the ROI areas of the maxilla and mandible detected by YOLOv5, respectively. , These represent the number of feature points obtained by SIFT filtering within the corresponding ROI. , The pixel coordinates of the maxillary and mandibular keypoints for coarse localization in YOLOv5. , To determine the coordinates of key points for precise positioning based on the centroid calculation of corner points, , These are the final two-dimensional pixel coordinates of the fused maxilla and mandible.

[0267] As a preferred approach, step six combines the two-dimensional coordinates of key tooth points in each frame with their corresponding points in the three-dimensional model, and uses the PnP algorithm to estimate the three-dimensional spatial pose of the mandible in the current frame. The specific process is as follows:

[0268] By using the registered 3D intraoral scan model, the 2D coordinates of the key teeth points in each frame of the video sequence can be mapped one-to-one to known 3D reference points. .in, This represents the position vector of the i-th key point of the maxillary tooth in the 3D model. This represents the position vector of the j-th key point of the mandibular tooth in the 3D model.

[0269] Combined with the intrinsic parameter matrix obtained from camera calibration The PnP algorithm is used to solve for the maxillary pose matrix in the current frame:

[0270]

[0271]

[0272] Minimize reprojection error:

[0273]

[0274] in, Let be the rotation matrix of the maxilla, representing the maxillary reference frame relative to the . The rotational relationship of the camera coordinate system. Let be the translation vector of the maxilla, representing the origin of the maxillary reference frame at the _i_th ... Position of the camera in the coordinate system Let be the two-dimensional pixel coordinates of the i-th maxillary key point in the image captured by the c-th camera. For the first The intrinsic parameter matrix of the camera, These are the coordinates of the maxillary reference point in the corresponding 3D model.

[0275] In the process of minimizing reprojection error, all three-dimensional points are first processed through EPnP. Represented as four virtual control points Linear combination:

[0276]

[0277] in, These are known coefficients, determined by the geometric relationship between the world coordinate system and the control point system.

[0278] Substitute this relationship into the projection equation:

[0279]

[0280] in, As a scale factor, Let R be the homogeneous coordinates of the pixel, R be the rotation matrix, and t be the translation vector.

[0281] In this way, all the matching points form a system of linear equations, which can be solved at once using singular value decomposition (SVD). And derive the rotation matrix. With translation vector The result obtained from EPNP As initial values, LM refinement is used to further reduce reprojection errors. The LM algorithm is used to solve the problem.

[0282]

[0283] LM update rules:

[0284]

[0285] in, Let them be rotation vectors and translation vectors. For Jacobian matrices, For reprojection residuals, is the damping factor.

[0286] After LM optimization, high accuracy and stability can be obtained. .

[0287] The pose matrix for the current frame is also solved using the PnP algorithm for the jaw region:

[0288]

[0289] Minimize reprojection error:

[0290]

[0291] Consistent with the process of solving the global pose of the maxilla, after sequentially passing through EPNP and LM, a high-precision and stable solution is obtained. .

[0292] After obtaining the global posture of the maxilla and mandible, calculate the movement of the mandible relative to the maxilla:

[0293]

[0294] in, and Let represent the global pose matrices of the maxilla and mandible in the camera's c-coordinate system, respectively. This indicates the rotational relationship between the lower jaw and the upper jaw. This represents the translation vector of the mandible relative to the maxilla.

[0295] As a preferred approach, step seven integrates the three-dimensional mandibular poses of all video frames in chronological order to construct a complete mandibular motion trajectory, thereby achieving label-free three-dimensional reconstruction of the subject's dynamic mandibular movement process. The specific process is as follows:

[0296] To fuse the relative poses obtained from different cameras, the rotation matrices of each camera are first... Convert to quaternion And a uniform rotation result is calculated using spherical mean squares:

[0297]

[0298] in, ( () represents the spherical distance between quaternions. These are the weighting coefficients.

[0299] The translation part is based on the reprojection error of each camera. Weighted average summation using inverse weights:

[0300]

[0301] In order to achieve a more robust integration posture .

[0302] The fused attitude sequence is subjected to temporal smoothing, and the translation component is smoothed using exponential smoothing filtering:

[0303]

[0304] The rotational component utilizes quaternion spherical linear interpolation to achieve a continuous transition:

[0305]

[0306] in, and Control the smoothness of translation and rotation respectively.

[0307] To suppress jitter caused by noise and detection errors during video acquisition. After the above processing, the set of three-dimensional motion trajectories of the mandible relative to the maxilla in the video sequence is obtained in chronological order:

[0308]

[0309] The logical ideas behind the methods disclosed in the above embodiments can be implemented, in whole or in part, through software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions or computer programs.

[0310] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0311] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0312] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0313] As described above, although the invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the invention itself. Various changes in form and detail may be made without departing from the spirit and scope of the invention as defined in the appended claims.

Claims

1. A markerless mandibular motion dynamic acquisition method, characterized in that, The method comprises the following steps: S1, acquiring the mouth scan data of the subject by an intraoral three-dimensional scanner, supplemented by CBCT image data, to form a three-dimensional gingival dentition model; S2, segmenting the three-dimensional gingival dentition model, extracting the boundaries of each tooth and gingival tissue, and rigidly registering the segmented CBCT image data and the mouth scan data to the same coordinate system; S3, in the segmented and registered three-dimensional gingival dentition model, labeling the three-dimensional reference points of each tooth to form an anchor point candidate set, specifically comprising: S3-1, let the point cloud set of each tooth be denoted as where p is the three-dimensional coordinate of a point on the ith tooth. S3-2, calculating the normal vector of the point p and compared with the unit normal vector k of the occlusal plane, and the occlusal surface point set satisfying the condition is screened out : wherein, is a normal vector of the point , is a unit normal vector of the occlusal plane, is a preset threshold value, is a set of point clouds of the first tooth. S3-3, in the occlusal point set Compute local curvature and sort the point set According to curvature value From large to small Selecting the curvature cusps in the top n ranks of curvatures as the set of cusp reference points of the tooth : S3-4, obtain the boundary points of each tooth through local neighborhood normal vector mutation detection, and obtain a boundary point set : S3-5, combining the curvature sharp point with the boundary point to obtain the complete three-dimensional reference set of the i-th tooth: S3-6, the three-dimensional reference point set of all teeth is represented as: wherein, is the ith tooth, is the corresponding set of three-dimensional reference points; S4, recording the image sequence of the subject's mandibular movement at multiple shooting points; S5, using the YOLOv5 model combined with a feature point detection algorithm to process the image sequence frame by frame, detect and locate the two-dimensional pixel position of the tooth anchor point, and according to the confidence of the detection result, select the highest confidence of several key points in the upper and lower jaws as the motion anchor point of the frame; S6, combining the two-dimensional coordinates of each frame of image motion anchor point with its corresponding point in the three-dimensional model, and using the PnP algorithm to estimate the three-dimensional spatial pose of the current frame of mandible; S7, integrating the three-dimensional spatial poses of all mandibles under different shooting points in time sequence to construct a complete mandibular movement trajectory, realizing the markerless three-dimensional reconstruction of the dynamic movement process of the subject's mandible.

2. The markerless mandibular motion dynamic acquisition method according to claim 1, characterized in that, Step S1 specifically comprises: S1-1, guiding the subject to maintain stability in a natural open-mouth posture, and using an intraoral three-dimensional scanner to sequentially scan the upper and lower dentition and gingival tissue to obtain complete mouth scan data; S1-2, using a cone beam CT device to scan the subject's head or maxillofacial region to obtain CBCT image data containing teeth, tooth roots, alveolar bone, upper and lower jaws, and surrounding bony structures; S1-3, standardizing the CBCT image data to a uniform voxel format, and aligning the coordinates of the mouth scan data.

3. The markerless mandibular motion dynamic acquisition method according to claim 1, characterized in that, In step S4, multiple cameras are evenly distributed around the shooting area and installed to simultaneously shoot image sequences of the subject's opening and closing mouth, occlusion, protrusion, and lateral sliding from different angles.

4. The markerless mandibular motion dynamic acquisition method of claim 1, wherein, In step S5, each candidate box ROI output by the YOLOv5 model is optimized by a position loss, a classification loss, and a confidence loss, wherein the position loss is defined as: Comprehensive loss Is: In the formula, is the center point coordinate of the predicted frame, is the center point coordinate of the real frame, is the predicted frame and the real frame intersection ratio, is the Euclidean distance square of the predicted frame center and the real frame center in the pixel coordinate system, is the square of the diagonal length of the minimum circumscribed rectangle that can simultaneously enclose the real frame and the predicted frame, is the consistency parameter of the width-height ratio of the predicted frame and the real frame, is the penalty coefficient for adjusting the weight of ; is the weight coefficient of the bounding box regression loss, is the weight coefficient of the target confidence loss, is the confidence loss, is the weight coefficient of the classification loss, is the classification loss.

5. The markerless mandibular motion dynamic acquisition method according to claim 4, characterized in that, In step S5, after all candidate boxes are detected, non-maximum suppression is first performed according to the category to remove redundant prediction boxes; For each category, sort all the candidate boxes by predicted confidence conf from high to low, and select the box with the highest confidence As the current reserved box, and in turn calculate its intersection over union with each remaining candidate box : If the intersection over union is greater than a preset threshold, it is considered that the two boxes are highly overlapped and the low confidence box is deleted; repeat the process until all candidate boxes are processed; After non-maximum suppression, the detected upper jaw key point set U corresponding to the three-dimensional reference point and the lower jaw key point set L are respectively recorded as: wherein, represents the i-th upper jaw key point and its corresponding detection box, represents the i-th lower jaw key point and its corresponding detection box, and each key point is provided with a confidence score.​​ Sort the sets U and L according to the confidence value from high to low, and select at least three key points with the highest confidence value from each set to form two subsets: wherein, With The specific key points of the tooth can be derived from the same tooth or distributed among different teeth. Will and The set of motion anchor points is obtained by merging. .

6. The markerless mandibular motion dynamic acquisition method according to claim 5, characterized in that, Step S5 further comprises: For sets For each maxillary and mandibular motion anchor point, Harris corner detection was performed within the maxillary and mandibular regions of interest (ROIs) respectively. Local geometric feature points related to the height of the cusp position were extracted within the ROIs, and gradients were calculated for the images within the ROIs. Construct a Gaussian weighted second-order matrix: wherein respectively the gradient of the image in directions; is a Gaussian kernel with standard deviation of 1. Calculate the Harris response function: wherein is the determinant of the matrix, representing the product property of the gradient in two directions; is the trace of the matrix, is an empirical constant; Non-maximum suppression and thresholding within the ROI , to obtain a set of upper jaw candidate corner points and a set of lower jaw candidate corner points ; three-dimensional reference points of each tooth It is known that the template cusp point pixel coordinates are projected to the current frame image plane by the camera calibration matrix K, the rotation matrix R and the translation vector t: wherein, represents a normalized conversion from homogeneous coordinates to pixel coordinates; is the upper arch template tooth cusp pixel coordinate; is the lower arch template tooth cusp pixel coordinate.

7. The markerless mandibular motion dynamic acquisition method of claim 1, wherein, Step S6 specifically comprises: S6-1, the upper jaw key point two-dimensional coordinate set detected in each frame of the image sequence and the lower jaw key point two-dimensional coordinate set correspond to the known three-dimensional reference point set respectively through the registered three-dimensional gingival dentition model wherein, represents a position vector of the ith maxillary key point in the three-dimensional gingival dentition model, represents a position vector of the jth mandibular key point in the three-dimensional gingival dentition model; S6-2, combine the camera calibration to obtain the intrinsic matrix , using PnP algorithm to solve the current frame of the upper jaw pose matrix and the lower jaw frame pose matrix : wherein is a rotation matrix of the upper jaw, representing the rotational relationship of the upper jaw reference frame with respect to the first station camera coordinate system; is a translation vector of the upper jaw, representing the position of the origin of the upper jaw reference frame in the first station camera coordinate system; is a rotation matrix of the lower jaw, representing the rotational relationship of the lower jaw reference frame with respect to the first station camera coordinate system; is a translation vector of the lower jaw, representing the position of the origin of the lower jaw reference frame in the first station camera coordinate system; S6-3, minimizing the re-projection error: wherein, is the two-dimensional pixel coordinate of the ith maxillary landmark in the image taken by the cth camera; is the two-dimensional pixel coordinate of the jth mandibular landmark in the image taken by the cth camera; , is the coordinate of the corresponding maxillary, mandibular reference point in the three-dimensional model.

8. The markerless mandibular motion dynamic acquisition method according to claim 7, characterized in that, In the process of minimizing the reprojection error, first all three-dimensional points are represented as a linear combination of four virtual control points : wherein are known coefficients determined by the geometric relationship from the world coordinate system to the control point system; are substituted by or are substituted by Substitute into the projection equation: wherein, is a scale factor, is a homogeneous coordinate of the pixel point, R is a rotation matrix, and t is a translation vector; All matching points form a system of linear equations, which can be solved using Singular Value Decomposition (SVD). And derive the rotation matrix. With translation vector ; the result of EPNP As initial values, the LM algorithm is used to solve the problem: wherein by or substitution; After the LM algorithm optimization, the global attitude meeting the predetermined accuracy condition is obtained and ; Based on the global pose, calculate the movement of the lower jaw relative to the upper jaw: wherein, with Rupper and Rlower represent the global pose matrices of the upper and lower jaws in the camera c coordinate system, respectively, Rlowerupper represents the rotational relationship of the lower jaw with respect to the upper jaw, Tlowerupper represents the translation vector of the lower jaw with respect to the upper jaw.

9. The markerless mandibular motion dynamic acquisition method of claim 1, wherein, Step S7 specifically comprises: S7-1, fuse the relative poses obtained by different cameras, convert the rotation matrix of each camera to a quaternion and calculate a unified rotation result using spherical mean ​ wherein is the spherical distance between quaternions, is a weight coefficient;​ S7-2, the translation part is according to the re-projection error of each camera inverse proportional weights are used to average and sum to obtain the fused pose ; S7-3, time sequence smoothing processing is performed on the fused pose, and the translation component adopts exponential smoothing filtering: S7-4, the rotation component is realized by continuous transition of quaternion spherical linear interpolation: wherein, with respectively control the smoothness of translation and rotation; S7-5, acquiring a set of three-dimensional motion trajectories of the mandible relative to the maxilla in the video sequence in chronological order .

Citation Information

Patent Citations

  • Head posture estimation method combined with YOLO-MobilenetV3 face detection

    CN113705521A

  • Mandibular movement modeling method, electronic equipment and storable medium

    CN119228983A

  • Three-dimensional attitude high-precision acquisition and reconstruction system for exercise training

    CN120431178A