Method and system for accurately acquiring intraoral data containing muscle movement information and medium
By combining U-Net and PointNet++ models with flexible registration technology, intraoral muscle movement information is dynamically recorded and processed, solving the problem of insufficient accuracy in traditional impressions. This enables precise data acquisition for tooth loss restoration and improves the accuracy and stability of the restoration margin.
Patent Information
- Application Number
- CN202511004571.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2026-02-27
AI Technical Summary
In existing technologies, the precision and accuracy of traditional movable restoration impressions for missing or damaged teeth are insufficient, especially when the mucosal morphology changes, which makes it impossible to accurately obtain dynamic data, resulting in large errors in the positioning of the edge lines and affecting the restoration effect.
An intraoral data acquisition method incorporating muscle movement information was adopted. By collecting intraoral photographs, the attached gingival region was automatically segmented and edge lines were generated using U-Net and PointNet++ models. Combined with elastic registration technology, the three-dimensional morphological information under muscle movement was dynamically recorded and registered with static data to generate a morphologically complete intraoral data sequence.
It enables precise acquisition of intraoral data under both dynamic and static conditions, simplifies clinical procedures, improves the accuracy and stability of prosthesis margins, reduces technical sensitivity, and promotes the digital and intelligent development of the field of dental implant restoration.
Smart Images

Figure CN121570282A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of dental data processing, in particular to an intraoral data accurate acquisition method containing muscle movement information, an intraoral processing device and a medium. BACKGROUND
[0002] Traditional dentition defect or dentition loss active repair impression needs to put a component similar to a tray into the user's mouth, and the edge shaping material is adsorbed on the edge of the tray. When the user moves the lips, marks will be left on the material, and then the material solidifies to form an impression. However, such an impression has the following problems: large shape difference, the mucosa has different shapes in the natural posture (closed mouth state) and under the action of the pulling force (open mouth state), resulting in insufficient impression accuracy and accuracy. The production of the impression depends greatly on the doctor's skill and experience, especially when accurately positioning the edge line, the difference between experienced doctors and beginners can reach several millimeters.
[0003] Currently, the mouth scanning technology obtains the intraoral data of the user through a scanning head, hoping to replace the traditional impression. However, in the open mouth state, the mucosa shape is different from that in the natural posture, resulting in a deviation between the scanning data and the actual use scene. There is a difference in the edge shape between the traditional impression and the mouth scanning data, and the traditional impression presents the shape in the natural closed mouth state, while the mouth scanning data presents the shape in the open mouth state. In addition, using mouth scanning to replace the traditional silicone rubber impression method, especially in some special scenarios such as the soft tissue shape in the edentulous jaw field, there are problems of great difficulty in implementation. The edentulous jaw relies on soft tissue for positioning, and the most needed for mouth scanning is accurate positioning to obtain more accurate data, therefore, it is extremely necessary to replace the impression taking method with mouth scanning in the prior art. SUMMARY
[0004] The purpose of the present application is to provide an intraoral data accurate acquisition method containing muscle movement information to solve the technical difficulty that the active repair impression cannot be directly obtained by mouth scanning in the prior art.
[0005] An intraoral data accurate acquisition method containing muscle movement information, comprising:
[0006] S1: Collect single-frame texture maps including intraoral photos or dental scanning images, use dental professionals to mark the attached gingival boundary, create a ground truth segmentation map, and create and train an attached gingival area automatic segmentation model after modifying the U-Net;
[0007] S2: First, perform plaster model scanning and manual edge line drawing, then use pointNet++ to generate edge lines, and then create and train an edge line automatic generation model;
[0008] S3: dynamically record the three-dimensional shape information of the user's intraoral muscles in the moving state;
[0009] S4: perform elastic registration on the obtained dynamic data and the conventional static oral scanning data using the landmark guidance to obtain a complete shape intraoral data sequence; and in the elastic registration, the dynamic data of each frame or designated frame is first segmented into attached gingival regions and edge lines using the attached gingival automatic segmentation model and the edge line automatic generation model.
[0010] The method further comprises: S5: obtaining the best edge line of the denture using the elastic registration and the complete shape intraoral data sequence.
[0011] Preferably, S4 further comprises:
[0012] Mark common anatomical landmarks: manually identify and mark the same pre-set N anatomical landmarks on the dynamic scanning data and the static full-mouth scanning model, which have clear anatomical definitions and are relatively stable or predictable in the moving and static states;
[0013] Apply the common anatomical landmarks to the elastic registration algorithm to find a spatial transformation function T using the energy equation E(T) so that the landmarks on the deformed static model fit the corresponding points on the dynamic scanning data;
[0014]
[0015] where, ∑i=1 N : represents the summation of all N paired landmarks (i from 1 to N).
[0016] ||dynamici-T(statici)|| 2 : calculate the square of the Euclidean distance between the i-th dynamic point and the corresponding i-th static point T after transformation;
[0017] λ, called the regularization coefficient or smoothing weight;
[0018] This is a measure of the spatial gradient of the transformation T, denotes the derivative of T, which essentially quantifies the degree of change or irregularity of the transformation T in space,
[0019] Finally, the dynamic data processing and reconstruction step is performed to fuse the continuous point cloud frames obtained after registration to construct a dynamic surface sequence of the moving area within a specific time period.
[0020] Preferably, the elastic registration of the dynamic data of each frame or designated frame to the automatic attached gingiva segmentation model to automatically segment the attached gingiva and the automatic edge line generation model to automatically generate the edge line further comprises:
[0021] The dynamic data of each frame or designated frame is automatically segmented by the automatic attached gingiva segmentation model to automatically segment the attached gingiva, and the attached gingiva is positioned to locate the mucogingival junction line between the movable mucosa of the attached gingiva. The highest point of the mucogingival junction line marked in the static state matches the buccal side of the alveolar ridge, and the error of the marked point is verified before matching. In the energy equation, β is added to control the deformation inhibition intensity of the attached gingiva area, and the area is forced to be close to non-deformation;
[0022] The edge line is automatically generated by the automatic edge line generation model, and the automatic edge line generation model can complete the missing gingival contour in the dynamic scan to improve the completeness of the registration usable data.
[0023] Preferably, the best edge line of the denture is obtained by using the elastic registration and the morphologically complete intraoral data sequence, which further comprises:
[0024] The dynamic-static registration generates a continuous time sequence of deformation fields {T t}(t=1, 2,..., N frames),
[0025] The model generates an initial edge line L_static in the static scan through S2, and the inverse transformation of the deformation field is applied to each frame of dynamic data to obtain a time-varying edge line sequence {L_dynamic,t},
[0026] The position of the edge point p at time t is defined as pt, and its motion range can be calculated by using displacement, time and deformation,
[0027] The gingival surface normal vector and the pre-set buffer value are used to obtain the final safety corridor boundary, and thus the best denture edge position is obtained.
[0028] Preferably, step S2 further comprises:
[0029] When the data is acquired, the edge line of the edentulous jaw is mostly at the mucosal turning point, and the posterior maxillary ridge area is easy to miss during scanning;
[0030] In the labeling stage, the pre-labeling of the maxillary alveolar notch and the molar pad is added; the semi-automatic labeling operation based on curvature mutation is used, if the Gaussian curvature change rate > threshold value θ and is located on the buccal and lingual sides of the alveolar ridge, the point is marked as an edge point, otherwise, if the point is located in the triangular fossa of the molar pad, the point is forced to be marked as an edge point;
[0031] In the model design, the sampling level of PointNet++ is reduced to retain details, the mucosa is deformed under pressure during clinical modeling, and data enhancement with elastic deformation is added during training;
[0032] Obtaining the key areas of the edentulous jaw, including the alveolar ridge, the labial frenulum, the molar pad area, and the upper posterior ridge area,
[0033] Obtaining the edge line of the border of the denture closed area.
[0034] Preferably, the dynamic recording of the three-dimensional morphological information of the user's intraoral muscles in the moving state further comprises:
[0035] The intraoral scanner based on the principle of structured light or confocal imaging is used to dynamically record the three-dimensional morphological information of the user's intraoral muscles in the moving state, including recording the morphological data of the intraoral focus movement area in the moving state,
[0036] The intraoral data in each frame is recorded in the form of motion vectors similar to video encoding;
[0037] Each frame of the dynamic scanning is marked with accurate time coordinates.
[0038] Preferably, the creation of the ground truth segmentation map in S1 further comprises the creation of an automatic attached gingiva segmentation model after the variant of the U-Net.
[0039] The information including the polygon contour line is converted into a binary segmentation mask to realize the generation of the ground truth, and in the mask: 0 and 1 are designed, including the information of the background / non-attached gingiva area as 0: representing teeth, free gingiva, alveolar mucosa, background, and the foreground / attached gingiva area as 1: representing the internal area of the annotated polygon, and ensuring that the mask size is consistent with the size after preprocessing of the original image.
[0040] In the model architecture design, the challenges of attached gingiva segmentation, such as fine boundary, low contrast with surrounding tissues, limited samples, etc., are improved
[0041] ResUNet is used for deeper network and residual connection, the residual connection alleviates gradient disappearance, especially accelerates training to improve performance when the network is deepened, the ordinary convolutional blocks in the encoder and decoder are replaced by residual blocks containing shortcut connections to guide the model to pay attention to the attached gingiva boundary area, suppress the interference of irrelevant areas including teeth and free gingiva, and introduce attention gates on the skip connection path, the gating signal weights the encoder features with high spatial resolution in the shallow layer, and highlights the features related to the target area.
[0042] The present application can also provide a system, comprising:
[0043] An intraoral scanner based on the principle of structured light or confocal imaging is used to dynamically record the three-dimensional morphological information of the user's intraoral muscles in a moving state, including recording the morphological data of the intraoral focus movement area in a moving state,
[0044] The processor is configured to collect single-frame texture maps including intraoral photos or dental scan images, mark the attachment gingival boundary with a dental professional, create a ground truth segmentation map, create and train an automatic attachment gingival area segmentation model after a variant of U-Net, first perform plaster model scanning and manual edge line drawing, then generate edge lines using pointNet++, and then create and train an edge line automatic generation model; in the first user operation, first dynamically record the three-dimensional morphological information of the user's intraoral muscles in a moving state; then, the dynamic data obtained and the conventional static oral scan data are elastically registered using a landmark guide to obtain a complete morphological intraoral data sequence; and in the elastic registration, the dynamic data of each frame or specified frame is first segmented into an attachment gingival area by the automatic attachment gingival area segmentation model and the edge lines are first generated by the edge line automatic generation model.
[0045] A storage medium storing a method as described above.
[0046] A single oral scan can simultaneously obtain "static anatomical morphology + dynamic functional morphology + edge line generation", and truly realize the transition from "experience-driven" to "data-driven" in some scenarios, such as edentulous jaw removable denture, removable partial denture, and implant overdenture. Muscle function shaping is a key impression technique for edentulous or partially edentulous removable restoration, which accurately defines the physiological extension range of the restoration edge through functional movements such as opening and closing the mouth, ensures the appropriate edge length, and thus facilitates the formation of edge closure, ensuring the retention and stability of the restoration. Digital intraoral scanning technology obtains dynamic data of muscle function movement, which can simplify the clinical operation procedure, reduce technical sensitivity, and improve the development of digital and intelligent development in the field of oral implant restoration, and bring precise and efficient restoration treatment to edentulous patients. BRIEF DESCRIPTION OF DRAWINGS
[0047] Figure 1 A principle flowchart of the method for accurately obtaining intraoral data containing muscle movement information of the present application; DETAILED DESCRIPTION
[0048] The present application will be described in detail below with reference to the accompanying drawings.
[0049] A method for accurately obtaining intraoral data containing muscle movement information, comprising:
[0050] S1: Collect single-frame texture maps including intraoral photos or dental scan images, mark the attached gingival boundary by dental professionals, create ground truth segmentation maps, and create and train an automatic attached gingival segmentation model after modifying the U-Net.
[0051] S2: First, perform plaster model scanning and manual edge line drawing, then use pointNet++ to generate edge lines, and then create and train an edge line automatic generation model.
[0052] S3: Dynamically record the three-dimensional morphological information of the user's intraoral muscles in the moving state.
[0053] S4: Perform elastic registration of the obtained dynamic data and conventional static oral scan data using landmark points to obtain a complete morphological intraoral data sequence. In the elastic registration, the dynamic data obtained in each frame or designated frame is first automatically segmented into attached gingival areas by the attached gingival automatic segmentation model and the edge lines are first generated by the edge line automatic generation model. The model in S2 automatically generates accurate edge lines from the intraoral scan data obtained in S4 to guide the production of the active prosthesis.
[0054] Please refer to Figure 1 , which is a flowchart of a method for accurately obtaining intraoral data containing muscle movement information. The specific process is described in detail as follows.
[0055] S110: Creation and training of the attached gingival automatic segmentation model
[0056] Collect single-frame texture maps including intraoral photos or dental scan images, mark the attached gingival boundary by dental professionals, create ground truth segmentation maps, and create and train an automatic attached gingival segmentation model after modifying the U-Net.
[0057] One implementation of step S110 is to first collect intraoral image data, then use professionals to label, and then use the model constructed by training the U-Net variant model to realize automatic segmentation of the attached gingiva.
[0058] The technical difficulties encountered in the implementation of step S110 mainly lie in three aspects: first, the complexity of the oral image, with many interference factors such as the tongue and saliva reflection; second, the boundary between attached gingiva and free gingiva is visually subtle; and third, the interference of the crown region needs special processing. The applicant also tries to consider the sample learning problem. Regarding the model design, the input layer needs to increase the HSV color space conversion because the color information of the gingival tissue is more distinguishable than RGB. The attention gate mechanism is introduced at the jump connection to improve the segmentation accuracy of the small boundary as much as possible; the output layer should generate binary mask and boundary distance map simultaneously considering the clinical practicability. In the layout of the data enhancement strategy, in addition to the conventional rotation and flipping, the applicant also increases the brightness distortion of the simulated saliva reflection and the random mask of the simulated tongue shielding according to the characteristics of the oral image. The consistency between the annotators is evaluated by Cohen's kappa coefficient in the professional annotation link. In addition, the model is as light as possible to be deployed on portable devices in dental clinics. For this purpose, the applicant can use depth separable convolution in the decoder part, which can compress the model volume by 40% with little precision loss. The clinical applicability index is easily overlooked in the final model evaluation. In addition to the Dice coefficient, the applicant also adds the gingival margin point positioning error (in millimeters) which is most concerned by dentists. The gingival margin point positioning error can be converted according to the image resolution in the operation process, so it is necessary to experiment with various image resolutions to experiment with the corresponding gingival margin point positioning error.
[0059] The detailed process of step S110 is as follows: S1111: data collection and preprocessing; S1112: professional annotation and GroundTruth creation; S1113: model architecture design (U-Net variant); S1114: model training and optimization; S1115: evaluation and deployment. The following is a detailed description of each step.
[0060] S1111: Data collection and preprocessing
[0061] First, the data source is obtained, such as oral photos taken by a standard dental intraoral camera. It is necessary to ensure that: it contains sufficient attached gingiva area (usually needs to show teeth, free gingiva, attached gingiva and alveolar mucosa), and the viewing angle is relatively standardized (buccal side, lingual / palatal side). It can also be derived from single-frame color texture images of 3D oral scanning devices (such as iTero, Trios, CEREC) and other intraoral scanning (IOS) texture images. In addition, samples of different users, different gingival biotypes, different health conditions, and different tooth positions are as diverse as possible.
[0062] Next, the preprocessing operation including size normalization and color normalization.
[0063] Size Normalization Step: All images are scaled to a uniform resolution (e.g. 512x512, 768x768). Keeping the original aspect ratio can result in padding or cropping.
[0064] Color Normalization Step: Apply techniques such as z-score standardization (subtract mean per channel and divide by standard deviation) or min-max scaling (scale to [0,1]) to reduce the impact of lighting and device differences. Histogram Matching or standardization based on deep learning style transfer can also be considered.
[0065] In the pre-processing process, an enhancement step can be included or not. Subsequent consideration of coping with sample learning mechanism, the enhancement step is added in the implementation process. Including but not limited to geometric transformation, photometric transformation, elastic deformation and occlusion processing steps. Geometric transformation can use rotation operation (such as ±10-30°), small amplitude translation, scaling, horizontal / vertical flip, etc. Photometric transformation operation includes brightness / contrast adjustment, adding Gaussian noise, simulating motion blur (slight), random Gamma correction, etc. Elastic deformation operation includes simulating slight deformation of gingival tissue, and occlusion operation includes simulating partial occlusion of saliva, tongue or instrument for processing.
[0066] If the intraoral photo is obtained by shooting, the photo needs to be extracted for ROI. Especially the image contains a large amount of irrelevant background such as lips, cheeks, etc. The region of interest (ROI) containing gingiva and teeth can be extracted by using simple threshold, edge detection or pre-trained general segmentation model for detecting tooth area, and then subsequent processing and training are carried out to improve the model focus.
[0067] S1112: Professional annotation and Ground Truth creation step
[0068] The attachment gingiva boundary is annotated by experienced periodontal doctors or strictly trained dental professionals using professional image annotation software (such as ITK-SNAP, 3D Slicer, Labelbox, CVAT, VGG Image Annotator (VIA)).
[0069] The annotator is provided with clear anatomical definition of attached gingiva, including junctional epithelium, end of attached gingiva, beginning of alveolar mucosa. Generally, alveolar mucosa is more red in color, smooth in surface, and movable, while attached gingiva is more pale in color, with orange peel-like mottling, tough in texture, immovable, and corresponding legends. Usually, a continuous, closed contour line similar to a polygonal attached gingiva line is drawn along the attached gingiva-alveolar mucosa junction, and the area outlined by the annotator is the attached gingiva area.
[0070] In general, a detailed annotation protocol document is developed, including boundary definition, special case handling (e.g. inflammation leading to fuzzy boundaries), example images. The same annotator should be highly consistent in their annotations of the same image at different points in time. A random selection of images (e.g. 10-20%) are independently annotated by multiple annotators, and consistency between annotators is assessed as much as possible. This assessment can be performed using the intra-class correlation coefficient or Dice coefficient to assess consistency.
[0071] The annotated polygon contour lines are converted to binary segmentation masks to generate the Ground Truth. In the mask: 0 and 1 can be used to design, such as background / non-attached gingival area is 0: represents teeth, free gingiva, alveolar mucosa, background, etc. The foreground / attached gingival area is 1, representing the internal area of the annotated polygon (attached gingiva), and ensuring that the mask size is consistent with the size of the original image after preprocessing.
[0072] S1113: Model architecture design (U-Net variant):
[0073] This step is based on the classic U-Net, and improvements are made for the challenges of attached gingival segmentation, such as boundary refinement, possibly low contrast with surrounding tissues, limited samples, etc.
[0074] Classic U-Net: In this example: 1) Encoder (downsampling): Extract multi-scale features. 2) Decoder (upsampling): Restore spatial resolution and precise positioning. 3) Bottle neck layer. 4) Final 1x1 Conv + Sigmoid output single-channel probability map. The techniques that can be used to achieve this block are not described in detail.
[0075] The most core of this step is variant selection: such as: using deeper network & residual connection (ResUNet). Residual Blocks alleviate the vanishing gradient, especially when the network is deepened to accelerate training to improve performance. And replace the ordinary convolutional blocks in the encoder and decoder with residual blocks (including shortcut connection). In this example, the guided model focuses on the attached gingival boundary area and suppresses the interference of irrelevant areas (such as teeth and free gingiva). It is particularly helpful to solve the problem of low contrast and fuzzy boundary. Introduce attention gate (Attention Gate) on the skip connection (Skip Connection) path. The gating signal (from the deep, semantically strong decoder features) weights the shallow, spatially high-resolution encoder features, highlighting features related to the target area. In this example, DenseUNet is used: to strengthen feature reuse, reduce gradient disappearance, and improve feature propagation efficiency, which may be more effective with limited data. In each block of the encoder / decoder, the input of each layer is the registration (concatenation) of the output of all previous layers. Also, auxiliary outputs and loss functions are added to the middle layers of the decoder, providing more direct gradient signals to the shallow decoder, which helps improve boundary prediction and accelerate convergence. Additional 1x1 Conv + Sigmoid output layers are added at different scales in the decoder (usually after several upsampling). Calculate auxiliary loss. In this example, improved upsampling is used: for example, transposed convolution (Transposed Convolution) may produce checkerboard artifacts (checkerboard artifacts). Therefore, use bilinear interpolation (Bilinear Upsampling) + convolution (Convolution) instead of transposed convolution for upsampling. Also, use a powerful feature extractor pre-trained on a large-scale dataset (such as ImageNet) as the encoder, and transfer learning can significantly improve performance on small datasets. One solution is to replace the encoder of U-Net with a pre-trained ResNet, VGG, EfficientNet, DenseNet, etc. Note that the first layer of convolution kernel may need to be adjusted to adapt to the number of input channels (RGB).
[0076] S1114. Model training and optimization
[0077] This step includes the design of the loss function and the design of the optimizer. When dealing with a serious imbalance between foreground (attached gingiva) and background areas (attached gingiva usually only accounts for a small part of the image). In this example, the Dice coefficient can be directly optimized.
[0078] The formula is DL = 1 - (2 * |X∩Y| + ε) / (|X| + |Y| + ε), where X is prediction, Y is GT, and ε is a smoothing term.
[0079] If the evaluation finds that the boundary accuracy is the main bottleneck, the applicant can use the Boundary-focused loss (such as BoundaryLoss) to explicitly increase the loss weight of the boundary area.
[0080] The optimizer is designed to prefer Adam or AdamW (Adam with weight decay), such as the initial learning rate: 1e-4 or 1e-3. Learning rate scheduling strategies are designed using ReduceLROnPlateau (reduce LR when validation loss no longer decreases) and CosineAnnealingLR. One training strategy for this example:
[0081] Data division: strictly divide the training set (70-80%), the validation set (10-15%), and the test set (10-15%). Ensure that different images from the same user only exist in one subset (prevent data leakage).
[0082] Batch size: select according to GPU memory (such as 8, 16, 32).
[0083] Training rounds (Epochs): use Early Stopping, stop training when the validation loss no longer decreases for consecutive epochs (such as 10-20), prevent overfitting. Monitor the Dice coefficient or IoU on the validation set.
[0084] Regularization: weight decay (L2 regularization): set in the optimizer. Dropout: can be added between the convolutional layers of the encoder / decoder or the bottleneck layer. Data augmentation: the most effective and necessary regularization method.
[0085] S1115: Evaluation and deployment
[0086] In the evaluation index in this example, the applicant designs the average symmetric surface distance (ASSD) / Hausdorff distance (HD): the key boundary index, calculates the average distance between the predicted boundary and the real boundary point set. The smaller the value, the better. In the visualization process, overlay the predicted mask (outline or semi-transparent cover) and the Ground Truth mask on the test image, intuitively check the segmentation result, especially the error at the boundary.
[0087] S120: First, perform a plaster model scan and manually draw the edge line, use pointNet++ to generate the edge line, and then automatically generate the model edge line.
[0088] By combining 3D scanning, traditional image processing and deep learning (PointNet++), the automatic generation of the edge line of the plaster model is realized, replacing the traditional manual drawing method and improving the efficiency and accuracy.
[0089] High-precision point cloud data of the plaster model is obtained using a structured light scanner or a laser scanner. The feature edges are marked on the point cloud or mesh model, and the labeled data is stored as a subset of the point cloud (each point is labeled with edge = 0 / 1). A training set (e.g., 70%), a validation set (e.g., 20%), and a test set (e.g., 10%) are generated. Then, the data is input into the created model. For example, the entire point cloud is divided into 512x512 overlapping blocks for point cloud block reasoning. The model predicts the edge probability (0-1) at the point level for each block, and the maximum probability is taken in the overlapping area. The global threshold T = 0.7 is used to determine the edge points for block fusion and thresholding. The edge line generation model is generated through the above method. Further optimization and clinical verification can be performed subsequently.
[0090] In this example, the automatic generation of the edge line of the edentulous jaw plaster model is taken as an example. The edentulous jaw has its particularity - there is no tooth interference, but the gingival morphology is more complex. The edge line usually refers to the termination line of the denture bearing area (such as the wobble line, the molar pad area, etc.), which has very high precision requirements (such as ±0.3mm is the clinical bottom line). First, at the data level, the edentulous jaw edge line is usually at the mucosa turning point (such as the labial frenulum), and special attention should be paid to these easily missed areas such as the maxillary posterior ridge during scanning. Second, during the labeling phase, strict adherence to anatomical landmarks (such as the zygomaticomaxillary notch and the molar pad) is required. Finally, in terms of model design, the edge of the edentulous jaw is smoother and more continuous, so the sampling level of PointNet++ can be reduced to preserve details. During clinical modeling, the mucosa is deformed under pressure, and the edge position may drift after scanning. The solution is to add data augmentation with elastic deformation during training or introduce mechanical simulation to generate synthetic data.
[0091] The key areas of the edentulous jaw, including the alveolar ridge, labial frenulum, molar pad area, and maxillary posterior ridge, are obtained, and the edge line of the denture closure area boundary is obtained. In addition to manual labeling, semi-automatic labeling based on curvature mutation can also be used. For example, if the Gaussian curvature change rate > threshold θ (threshold θ is a pre-set value) and is located on the buccal and lingual sides of the alveolar ridge, it is marked as an edge point, otherwise the point is located in the molar pad triangular fossa, it is forcibly marked as an edge point. Through the above method, semi-automatic labeling based on curvature mutation is performed.
[0092] The PointNet++ model is optimized. The input layer uses (x, y, z, normal vector, curvature) as 5-dimensional feature information. The loss function is optimized. The applicant uses a custom loss function class `EdgeFocalLoss`, which inherits from `nn.Module` (the base class of PyTorch neural network modules). The goal of this loss function is to give higher weights to key anatomical regions (such as the posterior area of the molar) during training to address the importance of these regions in the generation of the edge line of the edentulous jaw model. The create_anatomical_weight() function is called to generate a weight map of the same size as the input. Special processing of the posterior area of the molar: this area is given a weight of 3.0 times, and other areas: the base weight is 1.0. Based on the oral anatomy atlas, locate the key region coordinate range, and use morphological dilation to expand the key region boundary.
[0093] Use standard Focal Loss to solve the class imbalance problem.
[0094] FL(pt)=−αt(1−pt)γlog(pt)FL(pt)=−αt(1−pt)γlog(pt) Where: $\alpha_t$ : balance positive and negative sample weights (usually alpha = 0.25) $\gamma$ : focus parameter (usually gamma = 2) $p_t$ : model prediction probability Multiply the base loss value by the anatomical weight map point by point.
[0100] By the above method, the PointNet++ model can be optimized.
[0101] Step S110 and step S120 have no order, and can be performed simultaneously, or S120 can be performed first and S110 can be performed later.
[0102] S130: dynamically record the three-dimensional morphological information of the user's intraoral muscles in the motion state.
[0103] Wherein "dynamic record" is the biggest bottleneck, because the current mainstream oral scanning equipment is static. In high-speed three-dimensional scanning hardware, an intraoral scanner based on structured light or confocal imaging principle is adopted. The extremely high scanning frame rate of several tens of frames per second or even hundreds of frames, the scanning head keeps relative position or moves in a guided manner in the user's mouth, and continuously captures the point cloud or triangular mesh data stream of the local area (the aforementioned "focusing motion area"). The scanning head itself may have slight shaking due to the operator holding it, and the user's head may also have micro-motions. At the same time, the target soft tissue (lips, cheeks, tongue, floor of the mouth, etc.) is in rapid motion, and motion tracking and time stamp synchronization are as synchronized as possible. For example, the scanner integrates an inertial measurement unit (IMU-accelerometer, gyroscope) or optical tracking markers to record the pose changes (position and direction) of the scanning head in real time. An external optical tracking system (such as an infrared camera tracking a reflective ball on the scanning head) is used to provide higher-precision pose information, and each frame of scanning data is associated with an accurate time stamp in software, and the time stamp is as synchronized as possible with the tracking data.
[0104] In summary, an intraoral scanner based on structured light or confocal imaging principle is used to dynamically record the three-dimensional morphological information of the user's oral muscles in the motion state, mainly to record the morphological data of the focusing motion area in the motion state, and to record the data of each frame of intraoral data in the form of motion vectors in video coding; Finally, the time stamp synchronization problem, each frame of dynamic scanning needs to be marked with accurate time coordinates.
[0105] S140: Register the obtained data with conventional static oral scanning data to obtain a complete intraoral data sequence.
[0106] Dynamic oral scanning technology (e.g., recording the morphology of oral soft and hard tissues in the motion state of the user speaking, chewing, smiling, etc.) can only capture three-dimensional data of a limited area (focusing on the motion area). This range is smaller than the complete static full-mouth scan (including all teeth, gums, palate, etc.). In order to obtain complete and meaningful motion information, this local, motion-state scanning data needs to be accurately "pasted" onto the complete, static full-mouth model. This process is alignment or registration.
[0107] The implementation process of the landmark point guided elastic registration is as follows:
[0108] First, mark common anatomic landmarks: On both dynamic scan data and static full-arch scan model, manually identify and mark the same number of anatomic landmarks, in this example, 8. These points should have clear anatomic definition and be relatively stable or predictable in both motion and static state. Examples of landmark points that can be used are: Midpoint of Upper Labial Frenulum, Midpoint of Lower Labial Frenulum, Bilateral Buccal Shelf Area (need to mark one point on each side, total 2 points), Pterygomaxillary Notch (usually one on each side), Retromolar Pad (one on each side). The statici and dynamici are the corresponding points of the same anatomic location on different scans (static vs dynamic). i goes from 1 to N (N = 8 points).
[0109] Next, apply the elastic registration algorithm to these common anatomic landmarks, the goal is to find a spatial transformation function T that can "warp" or "deform" the static model such that the landmark points on the deformed static model T(statici) coincide with the corresponding points on the dynamic scan dynamici as much as possible. At the same time, this deformation itself should be smooth and physically reasonable, cannot produce drastic, discontinuous warping.
[0110] In this invention, an energy equation E(T) is used to achieve this:
[0111]
[0112] This equation defines a criterion ("energy") to measure how good a transformation T is. The goal is to find a transformation T that minimizes E(T). It consists of two terms:
[0113] Data Fidelity Term:∑i=1 n ||dynamici-T(statici)||2
[0114] ∑i=1 N : means summing over all N pairs of landmark points (i goes from 1 to N).
[0115] ||dynamici-T(statici)||2: calculates the square of the Euclidean distance between the ith dynamic point dynamici and the corresponding ith static point on the transformed model T(statici).
[0116] This measures how well the transformation T does in matching the corresponding points (statici and dynamici) known to the applicant. The smaller the value, the more accurate the matching of the corresponding points. This is the primary goal of the registration.
[0117] Regularization Term / Smoothness Term: λ
[0118] Lambda (λ): This is a positive real number, called the regularization coefficient or smoothness weight. It is a key control parameter of the algorithm. This is a measure of the spatial gradient of the transformation T. denotes the derivative of T (gradient in 3D space). Essentially quantifies how drastic or irregular the change in space is for the transformation T. Imagine T as a piece of rubber film, measures how drastic the stretching or bending of this film is. Large, abrupt deformations will result in large This term penalizes (suppresses) those transformations that are too complex, drastic, or discontinuous. The smaller the value, the smoother, more continuous, and more physically reasonable the transformation T is.
[0119] If λ = 0, the algorithm only focuses on point matching (∑||...|| 2 ), and will strive to make all T(statici) exactly equal to dynamici. This will often result in extreme distortion, folding, or discontinuity of the static model near the landmark points (overfitting), especially in the regions between or outside the points, because the algorithm can "arbitrarily" deform the model to satisfy the exact matching of the points. As λ increases, the smoothness constraint becomes more and more important. The algorithm will tend to choose a deformation that is overall smoother and more natural, even if this means that the landmark points are not matched exactly 100% (allowing for small matching errors). λ controls the trade-off between "exactly matching the landmark points" and "keeping the overall deformation smooth and natural". Choosing the right λ is key to obtaining reliable registration results, and needs to be adjusted (possibly through experiments) according to the specific data and application scenario. "Avoiding local distortion" is achieved precisely by adjusting λ.
[0120] The applicant finds λ through several experiments, and then determines the corresponding transformation T such that it 1) aligns the landmark points on the static model to the corresponding points on the dynamic model as well as possible (minimizes the first term), 2) is as smooth and regular as possible itself (minimizes the second term), with the relative importance between the two controlled by λ.
[0121] Again, the dynamic data processing and reconstruction steps are performed.
[0122] The continuously captured high-speed but local point cloud frames are registered in real-time (or post-processing) to the common reference coordinate system implemented above using the pose tracking data of the scanning head. This usually involves iterative closest point (ICP) or feature-based registration algorithm between adjacent frames, combined with IMU / tracking data for constraint and correction, compensating for the scanning head motion.
[0123] The registered continuous point cloud frames are fused to construct a dynamic surface sequence of the moving region in a specific time period. This may need to deal with the problems of blurring, occlusion, etc. caused by motion.
[0124] The resulting dynamic moving region surface sequence is "pasted" onto the complete, high-precision static full-mouth model to obtain dynamic information in the complete anatomical background: as mentioned above, using the pre-labeled common anatomical landmark points (statici, dynamici) on both data, applying an elastic registration algorithm (minimizing the energy function E(T) containing data items and smoothing items), calculating an optimal spatial deformation field T. Apply this calculated deformation field T to the entire static full-mouth model. In this way, the static model "comes to life", and its morphological changes (especially in the moving region covered by dynamic scanning) reflect the actual recorded motion state. The regions of the static model that are not directly captured in dynamic scanning (such as the labial side of the anterior teeth, the middle of the palate, etc.) have their deformation interpolated and calculated reasonably according to the motion of the adjacent regions.
[0125] The applicant can also add that the final output is a deformed static model sequence or a complete oral surface model sequence that changes over time, visually displaying muscle movement (such as cheek and lip movement) and soft tissue deformation (such as displacement of the buccal shelf area and the posterior pad of the molars). This allows for subsequent quantitative analysis of the intraoral state, such as displacement trajectory of specific points, regional strain, motion range measurement, etc.
[0126] In addition, after the model creation and training steps of S110 (automatic gingival region segmentation) and S120 (automatic generation of the marginal line) are completed, the dynamic scanning data of each frame or specified frame can be automatically segmented into the gingival region and the marginal line by the automatic gingival region segmentation model and the marginal line generation model.
[0127] The "membranosal junction" between the attached gingiva and the alveolar mucosa is an important soft tissue anatomical boundary in the oral cavity. Landmarks (such as the "bilateral buccal shelf area") are usually located near the junction of the attached gingiva and mucosa. Automatic segmentation results can directly assist in locating these areas, reducing errors from manual marking. In the static model, the "buccal shelf area" landmark can be defined as: "the highest point of the membranosal junction on the buccal side of the alveolar ridge." The segmentation model can automatically extract this boundary line to assist in manual / automatic marking. Automatic segmentation of the attached gingiva is performed first, ensuring the attached gingiva adheres closely to the bone surface with minimal deformation; while the alveolar mucosa will displace significantly under dynamic conditions. The segmentation results can serve as a physical constraint on the deformation field. New terms can be added to the energy equation.
[0128] E_{\text{new}}(T)=E(T)+\beta\sum_{j\in\text{attached gingival point}}\|T(\text{static}_j)-\text{static}_j\|^2
[0129] β controls the deformation inhibition intensity of the attached gingival region, forcing this region to remain nearly undeformed.
[0130] Energy equation-driven registration:
[0131]
[0132] Parameter function:
[0133] λ: Controls the smoothness of the deformation field (avoiding local distortion)
[0134] β: The forced attachment gingival region (G) is nearly rigid and does not deform.
[0135] The automatic edge line generation model first generates edge lines, which define the geometric boundaries of restorations / natural teeth. In other words, the edge lines (prepared edge / gingival margin) are key geometric feature lines in oral scans. The gingival margin may be obscured during motion. The automatic edge line generation model can complete missing gingival contours in dynamic scans, improving the completeness of usable registration data.
[0136] That is, taking the output of S110 / S120 as the input channel of the registration algorithm (for example: marking the attached gingival area / edge line in the static model as "non-deformable area"), check in the dynamic sequence: whether the deformation of the attached gingival area is close to 0 - this is the gold standard for verifying the physical reasonableness of registration. In summary, the dynamic data preprocessing process of each frame or key frame is added in the method, 1) attached gingival area segmentation: input dynamic scanning point cloud → automatically mark attached gingival area by pre-trained segmentation model (S110) → output binary mask. The attached gingival area is a rigid area with minimal deformation, serving as a registration anchor point. 2) Edge line generation: input dynamic point cloud → extract the abutment boundary by edge line model (S140) → output parameterized curve. The core of this point is to provide continuous geometric constraints to avoid relying solely on discrete points.
[0137] In addition, how to determine the best denture edge position in the dynamic-static registration process is also very critical. The dynamic registration result is used to correct the static edge line, and the soft tissue displacement range is quantified through the dynamic sequence to avoid gingival displacement leading to edge exposure / pressure as much as possible. Before this, the dynamic-static registration generates a continuous time sequence of deformation fields {T t}(t=1, 2,..., N frames), and the initial edge line L_static is generated in the static scan by the S120 model. The inverse transformation of the deformation field is applied to each frame of dynamic data to obtain a time-varying edge line sequence {L_dynamic,t}. The position of the edge point p at time t is defined as pt. The motion range can be calculated using displacement, time, and deformation. The final safe corridor boundary is obtained by considering the buffer margin using the gingival surface normal vector. Thus, the best denture edge position is obtained. In the design of the removable denture, the edge of the molar pad area can be extended to the upper limit of the dynamic safety corridor (using the maximum hyperplasia position to enhance retention), while the anterior area is strictly controlled within the central axis
[0138] In summary, in order to realize the functional remodeling of soft tissue muscles and determine the best position and shape of the denture edge, the oral scanning equipment needs to have the function of recording muscle movement, i.e. the intraoral scanner needs to have the function of dynamically recording the three-dimensional shape of the user's intraoral muscle ligament under movement. The best position and shape of the denture edge are then registered with the conventional static oral scanning data to obtain complete intraoral data with muscle movement function.
[0139] S110 generates the attached gingival segmentation mask, S120 extracts the motion amplitude heat map, and the motion amplitude is quantified: A(v)=maxt∥vt−vstatic∥A(v)=maxt∥vt−vstatic∥ (the displacement extreme value of vertex v at time t. The upper limit of the safety boundary is calculated as: Upper limit (retraction limit): Lupper={v∣A(v)+0.3mm}Lupper={v∣A(v)+0.3mm} Lower bound (proliferation limit): Llower={v∣A(v)−0.2mm}Llower={v∣A(v)−0.2mm} Obtain the best edge line.
[0143] The present application can obtain a precise functional impression acquisition method for edentulous jaws by combining dynamic muscle motion capture and static oral scanning data, and is suitable for digital design and manufacturing of complete dentures and implant dentures. Of course, edentulous jaws are only one application scenario, and can also be applied to missing teeth or other teeth.
[0144] A system comprising:
[0145] An intraoral scanner based on structured light or confocal imaging principles is used to dynamically record the three-dimensional morphological information of the user's intraoral muscles in a moving state, including recording the morphological data of the intraoral focus movement area in a moving state,
[0146] The processor is configured to: collect single-frame texture maps including intraoral photos or dental scan images, mark the attached gingival boundary with a dental professional, create a ground truth segmentation map, create and train an attached gingival area automatic segmentation model after a variant of U-Net; first perform plaster model scanning and manual edge line drawing, then use pointNet++ to generate edge lines, and then create and train an edge line automatic generation model; in the first user operation, first dynamically record the three-dimensional morphological information of the user's intraoral muscles in a moving state; then use the landmark points to guide the elastic registration of the obtained dynamic data and the conventional static oral scanning data to obtain a complete morphological intraoral data sequence; and in the elastic registration, the attached gingival area automatic segmentation model is used to automatically segment the attached gingival area, and the edge line automatic generation model is used to generate the edge line.
[0147] A storage medium stores the steps as above.
[0148] The embodiments of the present application are described in detail above in combination with the drawings, but the present application is not limited to the above-described embodiments. Even if various changes are made to the present application, if the changes fall within the scope of the claims of the present application and equivalent technologies, they still fall within the protection scope of the present application.
Claims
1. A method for accurately acquiring intraoral data containing muscle movement information, characterized in that, include: S1: Collect single-frame texture maps including intraoral photographs or dental scan images, use dental professionals to mark the attached gingival boundaries, create ground truth segmentation maps, and create and train an automatic segmentation model for the attached gingival region after modifying U-Net. S2: First, perform plaster model scanning and manual edge line drawing, then use pointNet++ to generate edge lines, and finally create and train an automatic edge line generation model; S3: Dynamically records the three-dimensional morphological information of the user's intraoral muscles during movement; S4: The obtained dynamic data and conventional static intraoral scan data are flexibly registered using marker-guided methods to obtain a complete intraoral data sequence. In this flexible registration, the dynamic data of each frame or a specified frame are first processed by the attached gingival region automatic segmentation model to automatically segment the attached gingival region and the edge line automatic generation model to generate the edge line. The model in S2 automatically generates precise edge lines based on the intraoral scan data obtained in S4 to guide the fabrication of active restorations.
2. The method of claim 1, further comprising: S5: The optimal edge line of the denture is obtained by using this flexible registration and morphologically complete intraoral data sequence.
3. The method as described in claim 1 or 2, characterized in that, S4 further includes: Marking common anatomical landmarks: On dynamic scan data and static full jaw scan models, manually identify and mark N pre-set anatomical landmarks that have clear anatomical definitions and are relatively stable or predictable in both motion and rest states; The elastic registration algorithm is applied to these common anatomical landmarks, and a spatial transformation function T is found using the energy equation E(T) so that the landmarks on the deformed static model are matched with the corresponding points on the dynamic scan. Where, ∑ i =1 N : This indicates summing over all N paired markers (i from 1 to N). ||dynamici-T(statici)|| 2 : Calculate the square of the Euclidean distance between the i-th dynamic point and the i-th static point T after transformation; λ is called the regularization coefficient or smoothing weight; This is a measure of the spatial gradient of the transformation T. The derivative of T is represented by... Essentially, it quantifies the degree of drasticness or irregularity of the spatial change of transformation T. Finally, dynamic data processing and reconstruction steps are performed to fuse the continuous point cloud frames obtained after registration, and to construct a dynamic surface sequence of the moving region within a specific time period.
4. The method as described in claim 3, characterized in that, In this flexible registration, dynamic data for each frame or a specified frame will be obtained first. The attached gingival region will be automatically segmented by the automatic segmentation model, and the edge line will be automatically generated by the automatic edge line generation model. Further steps include: For each frame or a specified frame of dynamic data, the attached gingival region is automatically segmented using the automatic segmentation model. The membrane-gingival junction line between the movable mucosa of the attached gingiva is located and matched with the highest point of the pre-marked membrane-gingival junction line on the buccal side of the alveolar ridge in the static state. Before matching, the marking error is verified in advance. β is added to the energy equation to control the deformation inhibition intensity of the attached gingival region. The following energy equation is used to drive the registration to force the region to be close to undeformed. Energy equation-driven registration: Parameter function: λ: Controls the smoothness of the deformation field (avoiding local distortion) β: The forced attachment gingival region (G) is nearly rigid and does not deform. The edge line automatic generation model first generates edge lines, which can complete the missing gingival contours in dynamic scanning to improve the integrity of the registration data.
5. The method as described in claim 2, characterized in that, The optimal prosthesis margin obtained using this flexible registration and morphologically complete intraoral data sequence further includes: Dynamic-static registration generates deformation fields {T} for continuous time series. t (t = 1, 2, ..., N frames), The S2 model generates initial edge lines L_static through static scanning, and applies inverse deformation field transformation to each frame of dynamic data to obtain a time-varying sequence of edge lines {L_dynamic,t}. Define the position of edge point p at time t as pt, and calculate its range of motion using displacement, time, and deformation. The final safe corridor boundary is obtained by using the gingival surface normal vector and a pre-set buffer value, thereby obtaining the optimal denture edge position.
6. The method as described in claim 1 or 2, characterized in that, Step S2 further includes: When acquiring data, the edge line of edentulous jaws is mostly at the mucosal transition point, and the posterior maxillary ridge area, which is easily missed, should be collected during scanning. During the annotation phase, annotations are added for the pterygomaxillary notch and the retromolar pad. Using a semi-automatic annotation operation based on curvature mutation, if the Gaussian curvature change rate is greater than the threshold θ and the point is located on the buccal or lingual side of the alveolar ridge, it is marked as an edge point; otherwise, if the point is located in the triangular fossa of the retromolar pad, it is forcibly marked as an edge point. In terms of model design, the sampling levels of PointNet++ are reduced to preserve details. During clinical modeling, the mucosa is deformed by pressure, and data augmentation based on elastic deformation is added during training. This allows for the acquisition of key edentulous areas, including the alveolar ridge crest, labial and buccal frenulum, retromolar pad area, and maxillary posterior ridge area. Obtain the edge line of the denture sealing area.
7. The method as described in claim 1 or 2, characterized in that, The dynamic recording of the three-dimensional morphological information of the user's intraoral muscles during movement further includes: An intraoral scanner based on structured light or confocal imaging principles is used to dynamically record the three-dimensional morphological information of the user's intraoral muscles during movement, including recording the morphological data of the intraoral focused motion area during movement. The intraoral data is recorded in each frame using motion vectors similar to those used in video coding. Each frame of a dynamic scan must be marked with precise time coordinates.
8. The method as described in claim 1, characterized in that, The creation of the ground truth segmentation map in S1, and the creation of the automatic segmentation model for the attached gingival region after modifying U-Net, further includes: The information, including the polygon outline, is converted into a binary segmentation mask to generate Ground Truth. The mask is designed using 0 and 1, with 0 representing the background / non-attached gingival region, which represents the information including teeth, free gingiva, alveolar mucosa, and background. 1 represents the foreground / attached gingival region, which represents the internal region of the labeled polygon. The mask size is ensured to be consistent with the size of the preprocessed original image. When designing the model architecture, improvements were made to address the challenges of attached gingival segmentation, including fine boundaries, potentially low contrast with surrounding tissues, and a limited sample size. ResUNet is used for deeper networks and residual connections. Residual connections alleviate gradient vanishing and accelerate training, especially when the network is deepened, to improve performance. Ordinary convolutional blocks in the encoder and decoder are replaced with residual blocks containing shortcut connections to guide the model to focus on the attached gingival boundary region and suppress interference from irrelevant regions, including teeth and free gingiva. Attention gates are introduced on the skip connection paths. The gated signals are used to weight shallow, spatially high-resolution encoder features to highlight features related to the target region.
9. A system, characterized in that, include: Intraoral scanners based on structured light or confocal imaging principles are used to dynamically record the three-dimensional morphological information of a user's intraoral muscles during movement, including recording morphological data of the intraoral focused motion area during movement. Processor: Configured to: collect single-frame texture maps including intraoral photographs or dental scan images; mark the attached gingival boundary using dental professionals; create ground truth segmentation maps; create and train an automatic attached gingival region segmentation model after modifying U-Net; first, perform plaster model scanning and artificial edge line drawing, then use PointNet++ to generate edge lines, and finally create and train an automatic edge line generation model; during the first user operation, dynamically record the three-dimensional morphological information of the user's intraoral muscles in motion; then, use marker points to guide the elastic registration of these dynamic data with conventional static intraoral scan data to obtain a complete intraoral data sequence; and in this elastic registration, the dynamic data of each frame or specified frame is first used by the automatic attached gingival region segmentation model to automatically segment the attached gingival region and the automatic edge line generation model is first used to generate edge lines.
10. A storage medium storing the method as described in any one of claims 1 to 8.