An Automated Tooth Alignment Method and System Based on Multi-Source Feature Fusion and Dental Arch Trajectory Prediction

By using multi-source feature fusion and dental arch trajectory prediction, a local geometric feature map structure and a global spatial dependency model for teeth are constructed, which solves the problems of inaccurate dental arch trajectory prediction and tooth collision in existing technologies, and realizes efficient automatic tooth alignment.

CN122492822APending Publication Date: 2026-07-31ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV OF TECH
Filing Date
2026-05-15
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing automatic tooth alignment methods do not fully utilize the tooth grid topology, resulting in inaccurate prediction of dental arch trajectory, affecting the overall coordination and alignment of teeth, and making them prone to tooth collisions and abnormal tooth spacing.

Method used

By constructing a local geometric feature map structure for teeth based on multi-source feature fusion, performing graph convolution feature extraction, and modeling global spatial dependencies, the center feature is encoded by combining tooth centroid coordinate data to construct dental arch prediction features. The network is then optimized through collision avoidance loss and dental arch alignment loss function to achieve automated tooth alignment.

Benefits of technology

It improves the accuracy of tooth position prediction, achieves overall coordination of the dentition and alignment with the natural dental arch trajectory, suppresses tooth collision, and enhances the overall coordination and clinical applicability of the tooth arrangement results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492822A_ABST
    Figure CN122492822A_ABST
Patent Text Reader

Abstract

This invention discloses an automatic tooth alignment method and system based on multi-source feature fusion and dental arch trajectory prediction, relating to the field of image processing technology. The method includes acquiring three-dimensional mesh data of the target dentition and extracting tooth point coordinate data, tooth point normal vector data, and tooth centroid coordinate data; constructing a single-tooth topology graph structure based on the tooth point normal vector data and performing graph convolution feature extraction processing to obtain local geometric feature data of the teeth; performing position enhancement processing based on the tooth point coordinate data and performing feature fusion processing with the local geometric feature data of the teeth; performing global spatial dependency modeling processing based on the fused structural feature data to obtain tooth pose prediction data; performing dental arch trajectory prediction processing based on the tooth centroid coordinate data to obtain dental arch curve parameter data; and performing tooth alignment constraint analysis based on the tooth pose prediction data and dental arch curve parameter data to output automatic tooth alignment result data. This invention can improve the accuracy of tooth alignment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, specifically to an automatic tooth alignment method and system based on multi-source feature fusion and dental arch trajectory prediction. Background Technology

[0002] In the process of digital orthodontic treatment, doctors usually need to adjust each tooth individually based on the three-dimensional model of the patient's dental arch in order to achieve dental arch coordination and dental alignment. For example, existing automated tooth alignment methods mostly use tooth point cloud data as input and predict tooth pose through deep learning to achieve automatic tooth alignment.

[0003] Existing technologies have achieved a certain degree of automatic tooth alignment by predicting tooth pose through deep learning, but most of them do not make full use of the tooth mesh topology to predict the dental arch trajectory, which makes the tooth alignment results easily deviate from the natural dental arch trajectory and affect the overall coordination of the teeth. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides an automatic tooth alignment method and system based on multi-source feature fusion and dental arch trajectory prediction.

[0005] To achieve the above objectives, the technical solution of the present invention is as follows:

[0006] In a first aspect, this invention discloses an automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction, comprising the following steps: acquiring three-dimensional mesh data of the target dentition teeth, and extracting tooth point coordinate data, tooth point normal vector data, and tooth centroid coordinate data; analyzing tooth point normal vector data to obtain local geometric feature data of the teeth; performing position enhancement processing based on tooth point coordinate data to obtain tooth position feature data, and performing feature fusion processing based on the local geometric feature data and tooth position feature data to obtain fused structural feature data; performing global spatial dependency modeling processing based on the fused structural feature data to obtain tooth pose prediction data corresponding to each tooth, wherein the tooth pose prediction data includes rotation parameter data and translation parameter data; performing center feature encoding processing and position enhancement processing based on tooth centroid coordinate data to obtain dental arch prediction feature data, thereby performing dental arch trajectory prediction processing to obtain dental arch curve parameter data; performing tooth alignment constraint analysis based on tooth pose prediction data and dental arch curve parameter data to obtain collision avoidance loss data and dental arch alignment loss data, and performing network iterative optimization processing to output automatic tooth alignment result data corresponding to the target dentition teeth.

[0007] Secondly, this invention discloses an automatic tooth alignment system based on multi-source feature fusion and dental arch trajectory prediction, comprising the following modules: a data acquisition module, used to acquire three-dimensional mesh data of the target dentition teeth and extract tooth point coordinate data, tooth point normal vector data, and tooth centroid coordinate data; a local feature extraction module, used to analyze tooth point normal vector data to obtain local geometric feature data of the teeth; a feature fusion analysis module, used to perform position enhancement processing based on tooth point coordinate data to obtain tooth position feature data, and perform feature fusion processing based on tooth local geometric feature data and tooth position feature data to obtain fused structural feature data; and a pose prediction module, used to perform feature fusion based on the fusion... The structural feature data undergoes global spatial dependency modeling to obtain tooth pose prediction data for each tooth, including rotation and translation parameters. The arch trajectory prediction module performs center feature encoding and position enhancement based on tooth centroid coordinate data to obtain arch prediction feature data, which is then used to perform arch trajectory prediction to obtain arch curve parameter data. The iterative optimization module performs tooth alignment constraint analysis based on the tooth pose prediction data and arch curve parameter data to obtain collision avoidance loss data and arch alignment loss data, and performs network iterative optimization to output the automatic tooth alignment result data corresponding to the target dentition.

[0008] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0009] 1. This invention constructs a single-tooth topology graph structure based on tooth point normal vector data and performs graph convolution feature extraction processing to obtain local geometric feature data of the tooth. At the same time, it performs global spatial dependency modeling processing based on fused structural feature data, so that the network can simultaneously perceive the local geometric shape of the tooth and the spatial relationship of the entire dental arch. This improves the accuracy of tooth pose prediction and the overall coordination of the dental arch, effectively solving the problem of poor overall coordination of tooth arrangement results caused by isolated tooth pose prediction in the prior art.

[0010] 2. This invention obtains dental arch prediction feature data by performing center feature encoding and position enhancement processing based on tooth centroid coordinate data, and obtains dental arch curve parameter data by performing dental arch trajectory prediction processing. This enables the network to learn the overall physiological dental arch morphology of the dentition, thereby achieving automatic alignment between the tooth arrangement result and the natural dental arch trajectory. This solves the problems of tooth arrangement deviating from the real dental arch morphology and insufficient overall continuity of the dentition in the prior art.

[0011] 3. By performing rigid body transformation processing on the three-dimensional mesh data of teeth based on rotation parameter data and translation parameter data, and performing collision probability analysis based on the two-dimensional projection data of teeth and the preset Gaussian kernel function, the overlap and interlocking between teeth are constrained, thereby effectively suppressing the collision phenomenon during tooth arrangement and solving the problems of tooth collision and abnormal tooth spacing in the automatic tooth arrangement results of the existing technology.

[0012] 4. This invention constructs a joint loss function based on collision avoidance loss data and dental arch alignment loss data, and performs parameter iterative update processing on the automatic tooth arrangement network based on the joint loss function. This enables the automatic tooth arrangement network to learn the collision-free tooth arrangement and the natural alignment law of the dental arch at the same time, thereby realizing fully automated tooth arrangement processing from inputting the three-dimensional mesh data of the target dentition to outputting the automatic tooth arrangement result data, solving the problem of low efficiency of manual tooth arrangement in the prior art. Attached Figure Description

[0013] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts. Wherein:

[0014] Figure 1 This is a flowchart of the automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction of the present invention;

[0015] Figure 2 This is a diagram illustrating the overall architecture of the network of the present invention;

[0016] Figure 3 This is a schematic diagram of the diagram construction of the present invention;

[0017] Figure 4 This is a schematic diagram of the dental arch alignment loss constraint of the present invention;

[0018] Figure 5 This is a schematic diagram comparing the accuracy curves of the present invention;

[0019] Figure 6 This is a schematic diagram illustrating the qualitative comparison of the tooth arrangement results of the present invention;

[0020] Figure 7 This is a qualitative comparison diagram of the collision avoidance module of the present invention;

[0021] Figure 8 This is a qualitative comparison diagram of the dental archline prediction network of the present invention;

[0022] Figure 9 This is a system architecture diagram of the present invention. Detailed Implementation

[0023] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.

[0024] In the process of digital orthodontic treatment, doctors usually need to adjust each tooth individually based on the three-dimensional model of the patient's dental arch in order to achieve dental arch coordination and dental alignment. However, existing automatic tooth alignment methods have problems such as insufficient extraction of local geometric features of teeth, lack of explicit constraints on dental arch trajectory, and easy collision and overlap during tooth alignment, resulting in poor overall coordination and clinical applicability of the tooth alignment results.

[0025] Therefore, this proposal suggests an automatic tooth alignment method and system based on multi-source feature fusion and dental arch trajectory prediction, including the following steps: acquiring pre-orthodontic 3D mesh data of teeth, extracting tooth point coordinates, point normal vectors, and centroid coordinates; then constructing a graph structure based on the mesh topology of a single tooth, extracting local geometric features of the tooth through a graph convolutional network, and fusing the point coordinate features with the geometric features; using a visual transformation network to perform global spatial dependency modeling on the fused features, outputting the six-degree-of-freedom pose parameters of the teeth; simultaneously encoding the position of the tooth centroid coordinates and inputting them into a dental arch prediction network to obtain dental arch curve parameters to guide the teeth to align along the natural dental arch trajectory; constructing collision avoidance loss and dental arch alignment loss to constrain the overlapping areas of teeth and the degree of dental arch deviation; finally, iteratively optimizing the network parameters through joint loss to output collision-free automatic tooth alignment results that conform to the natural dental arch shape. This approach solves the problem of poor overall coordination of tooth alignment results caused by existing technologies that only perform single-tooth analysis during automatic tooth alignment by performing coordination analysis between teeth and dental arch trajectory prediction.

[0026] like Figure 1 As shown, this is an embodiment provided in this application. Figure 1The flowchart of the automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction of the present invention includes the following steps: acquiring three-dimensional mesh data of the target dentition teeth and extracting tooth point coordinate data, tooth point normal vector data, and tooth centroid coordinate data; analyzing tooth point normal vector data to obtain local geometric feature data of the teeth; performing position enhancement processing based on tooth point coordinate data to obtain tooth position feature data; performing feature fusion processing based on tooth local geometric feature data and tooth position feature data to obtain fused structural feature data; performing global spatial dependency modeling processing based on fused structural feature data to obtain tooth pose prediction data corresponding to each tooth, wherein the tooth pose prediction data includes rotation parameter data and translation parameter data; performing center feature encoding processing and position enhancement processing based on tooth centroid coordinate data to obtain dental arch prediction feature data, thereby performing dental arch trajectory prediction processing to obtain dental arch curve parameter data; performing tooth alignment constraint analysis based on tooth pose prediction data and dental arch curve parameter data to obtain collision avoidance loss data and dental arch alignment loss data, and performing network iterative optimization processing to output automatic tooth alignment result data corresponding to the target dentition teeth.

[0027] In this embodiment, as Figure 2 As shown, Figure 2The diagram shows the overall architecture of the network in this invention. The input data includes the centroid coordinates, point coordinates, and normal vectors of the teeth. The feature encoding module employs a three-branch approach. The first branch takes the tooth's point normal vector as input and encodes it using a local encoder. In this paper, a multi-layer perceptron (MLP) is used for feature encoding across all encoder types. This MLP is then input to the tooth geometry feature extraction network, which extracts the local geometric features of the teeth. The second branch takes the tooth's point coordinates as input. After encoding by the local encoder, these coordinates are fused element-wise with the local geometric features output from the first branch. This fusion is then input to the tooth structure feature extraction network, which captures the global spatial dependencies of the teeth and ultimately outputs the six degrees of freedom (6DoF) pose transformation parameters, achieving accurate estimation of the tooth pose. The third branch takes the tooth's centroid coordinates as input. After encoding by the tooth center encoder, the classic sinusoidal positional encoding (SPE) is introduced to construct the tooth position encoding. This method was first proposed by Vaswani et al. This positional encoding constructs a discriminative and hierarchical embedding representation by applying a set of sine and cosine functions of different frequencies to each position. Because these are non-learnable parameters, they exhibit good translation invariance in sequence modeling tasks, providing necessary sequence positional information for the Transformer structure under unsupervised conditions. After encoding construction, the central encoding features are element-wise added to the tooth position encoding to supplement the prior knowledge of tooth row position for global feature extraction. Simultaneously, the features are input into the arch prediction network to fit the standard arch curve, and the tooth 6DoF pose transformation parameters are constrained by a loss function to ensure the anatomical rationality of the final tooth arrangement result. The first and second branches perform tooth pose prediction well, while the third branch independently performs arch trend prediction. The prediction results from both branches are used as input for subsequent loss calculations. A joint loss function is constructed by weighted summing of collision avoidance loss and arch loss, serving as the overall training objective for the automatic tooth arrangement network. After training, a forward propagation is performed on the target tooth row using the target network parameters, outputting the tooth rotation quaternions and translation corrections. A rigid body transformation is then performed on the original mesh data to obtain the final automatic tooth arrangement result data.

[0028] Due to the irregular topological structure of 3D tooth mesh data, traditional Euclidean space-based convolution methods struggle to effectively model its neighborhood relationships. This approach introduces Geometric Networking (GCN) to learn local geometric features of the tooth mesh and models the local geometric structure based on node adjacency relationships. It's worth noting that this paper does not employ the common k-nearest neighbors (KNN) method in constructing the graph structure. (The following sentence appears to be a separate, unrelated point: "Establishing adjacency relationships based on KNN...") Figure 3 As shown in the left sub-figure, this can easily lead to mesh vertices of different teeth being misclassified as adjacent nodes, thus confusing the independent geometric semantic features of each tooth. The point normal vector feature introduced in this paper is a core geometric attribute characterizing the local morphology of a single tooth surface; its spatial distribution does not require establishing connections between different teeth. Therefore, as... Figure 3 As shown in the right sub-figure, this paper constructs a graph structure based on mesh topology and establishes edge connections only between adjacent mesh vertices within a single tooth. This ensures that graph convolution operations only aggregate local features belonging to the same tooth, effectively avoiding cross-tooth information interference, thereby improving the modeling accuracy of the local geometry of the tooth.

[0029] Based on the above graph structure construction strategy, this paper designs a network for extracting tooth geometric features. For example... Figure 2 As shown, Figure 2 The diagram shows the overall architecture of the network in this invention. This network takes the point normal vector features of tooth mesh data as input and relies on a topological graph to represent the adjacency relationships of tooth vertices, completing refined feature aggregation within an irregular mesh space. The specific structure is as follows: The input features first undergo a point-by-point one-dimensional convolution (Conv1d) in the first layer, performing a linear transformation along the channel dimension to achieve feature upsizing. Subsequently, batch normalization (BN) and a non-linear activation function are applied sequentially to normalize and non-linearly transform the features. Based on this, the features are input into the graph convolutional layer, where they are weighted and aggregated with the features of their neighboring nodes at each vertex to extract local topological information of the teeth, enhancing the geometric expressive power of the features. At the end of the module, the features are further processed by Conv1d, BN, and a non-linear activation function to fuse and non-linearly enhance the contextual features extracted by the graph convolution along the channel dimension, ultimately outputting enhanced geometric features that provide reliable pre-feature support for subsequent feature fusion and tooth pose regression.

[0030] Furthermore, the local geometric feature data of the teeth is obtained. Specifically, the following methods are used: obtain the mesh vertex data corresponding to each tooth in the target dentition and perform adjacency analysis to obtain the set of adjacent vertices inside a single tooth, thereby constructing a single tooth topology graph structure; perform graph convolution aggregation processing on the tooth point normal vector data based on the single tooth topology graph structure to obtain the local topology feature data corresponding to each mesh vertex; and perform channel fusion processing based on each local topology feature data to obtain the local geometric feature data of the teeth.

[0031] In this embodiment, it should be noted that after the tooth itself is reconstructed by oral scanning, a three-dimensional mesh model is formed. The three-dimensional mesh model itself contains vertices, edges, and faces. Therefore, the mesh vertex data is essentially the original vertex coordinate data in the three-dimensional mesh model of the tooth.

[0032] A single-tooth topological graph structure refers to a graph data structure formed by establishing edge connections only between the vertices of the mesh within the same tooth, without establishing edge connections across different teeth. It is obtained as follows: First, the patient's pre-orthodontic 3D dental mesh data is acquired using an oral scanning device (such as an intraoral scanner). Then, a tooth instance segmentation algorithm (such as a deep learning-based 3D mesh segmentation model) is used to segment the dental mesh into sub-mesh for each individual tooth. All vertices of each tooth sub-mesh constitute the set of mesh vertices for that tooth. Adjacency analysis is performed on the vertex set of each tooth sub-mesh, that is, the adjacent vertices of each vertex are determined based on the vertex sharing relationship of the mesh face elements. Vertex pairs within the same tooth that share edges or face relationships are determined as adjacent vertex pairs. This constructs a graph structure containing only edge connections within a single tooth, i.e., a single-tooth topological graph structure. The specific method is as follows: Obtain the segmented sub-mesh data of each tooth in the target dentition. For each tooth's sub-mesh, traverse its triangular facets one by one, extracting the three edges of each facet. Record the two endpoints of each edge as adjacent vertices, forming the adjacency matrix of that tooth, i.e., the set of adjacent vertices within a single tooth. This constructs the topological graph structure of a single tooth. This process strictly limits edge connections to being established only within the same tooth, without introducing any cross-tooth adjacency relationships.

[0033] The set of adjacent vertices within a single tooth refers to the set of all vertices in a tooth submesh that are directly adjacent to a given vertex (sharing an edge). It is obtained by traversing all triangular elements of the tooth submesh, recording the two endpoints of each element's three edges as adjacent vertices, and summing these to obtain the set of adjacent vertices for each vertex, i.e., the set of adjacent vertices within a single tooth.

[0034] Graph convolutional aggregation refers to the operation of aggregating and transforming neighborhood information of graph node features in a Graph Convolutional Network (GCN). Specifically, the method involves using tooth point normal vectors as initial node features, inputting them into a tooth geometric feature extraction network composed of a stacked sequence of one-dimensional convolutional layers, batch normalization (BN) layers, activation function layers (such as ReLU), and graph convolutional layers. Based on the adjacency relationships in the single-tooth topological graph structure, graph convolutional aggregation is performed on each vertex to obtain the local topological feature data corresponding to each grid vertex. It should be noted that the direct output of graph topological aggregation is the local topological feature vector corresponding to each vertex, the number of which is the same as the number of vertices in a single vertex, resulting in a variable-length sequence. To compress the variable-length vertex sequence into a fixed-dimensional representation of a single vertex, this scheme performs channel-by-channel max pooling on all the required vertex feature matrices of a single vertex, taking all vertex feature values ​​in each channel, and outputting a fixed-dimensional feature sparsity, i.e., the fused local geometric feature data corresponding to that tooth. This operation ensures that the output of tooth-level features is consistent with the number of vertices contained in any given tooth, guaranteeing the consistency between subsequent features and the input variance of the global model.

[0035] Local topological feature data refers to the high-dimensional feature vector corresponding to each grid vertex after graph convolution aggregation processing, reflecting the local geometric relationship between that vertex and its neighboring vertices in a single-tooth topological structure. It is obtained by the output of the graph convolution aggregation processing described above.

[0036] Channel fusion processing refers to performing a max pooling aggregation operation along the channel dimension on the local topological feature data of each mesh vertex, aggregating the local feature vectors of all vertices of a single tooth into a feature vector representing the overall local geometric information of the tooth, i.e., the local geometric feature data of the tooth. Specifically, for all vertex feature matrices of a tooth, a channel-wise maximum value operation (Max Pooling) is performed along the vertex dimension, outputting a feature vector of fixed dimensions.

[0037] By performing adjacency analysis on the vertices of the tooth mesh, edge connections between vertices are established only within a single tooth to construct a single-tooth topology graph structure, instead of using the cross-tooth KNN adjacency method. This fundamentally avoids the mutual interference of feature information between vertices of different teeth, ensuring the purity and independence of the local geometric features of each tooth.

[0038] Furthermore, the fused structural feature data is obtained by performing local encoding on the tooth point coordinate data to obtain point coordinate encoded feature data; performing corresponding dimension alignment processing on the point coordinate encoded feature data and the tooth local geometric feature data to obtain feature alignment result data; and performing element-by-element fusion processing on the feature alignment result data to obtain fused structural feature data.

[0039] In this embodiment, local encoding processing is performed on the tooth point coordinate data to obtain point coordinate encoded feature data. Specifically, the three-dimensional point coordinate matrix of each tooth (containing the three-dimensional coordinate values ​​of all vertices of the tooth) is input into the local encoder. The local encoder is composed of a linear layer, a batch normalization layer and an activation function layer stacked in sequence. The low-dimensional coordinates are upgraded to high-dimensional features through forward propagation, and then max pooling is performed along the vertex dimension to output the fixed-dimensional point coordinate encoded feature data corresponding to each tooth.

[0040] The feature alignment result data is obtained through the following method: The point coordinate encoded feature data is encoded by a local encoder (consisting of a linear layer, a batch normalization layer, and an activation function layer) to encode the three-dimensional point coordinates of the teeth. The result is then max-pooled and output. The feature dimension is determined by the number of output channels of the final linear layer of the local encoder, denoted as [missing information]. The local geometric feature data of the teeth is encoded using a point-by-point method by the Geometric Feature Extraction Network (GCN module) and output after max pooling. Its feature dimension is determined by the number of output channels of the GCN, denoted as . ,when and When there is inconsistency, a linear transformation layer (fully connected layer) is used to perform a dimensionality-upgrading mapping on the feature path with lower dimensionality, so that the two feature paths are unified to the same target dimension. This yields the feature matching results, ensuring dimension matching for subsequent element-by-element addition operations.

[0041] The fused structural feature data is obtained by adding the two feature vectors in the feature alignment result data element by element to obtain the fused structural feature data. This fused feature simultaneously encodes the three-dimensional spatial position information of the tooth (from the point coordinate encoding feature) and the local surface geometric information (from the local geometric feature), which serves as the input for subsequent global modeling.

[0042] Furthermore, the tooth pose prediction data corresponding to each tooth is obtained. The specific method is as follows: input the fused structural feature data into the preset visual transformation network model; perform tooth spatial association analysis based on the multi-head self-attention mechanism in the visual transformation network model to obtain tooth global association feature data; perform pose regression processing based on the tooth global association feature data to obtain rotation parameter data and translation parameter data corresponding to each tooth.

[0043] In this embodiment, the Vision Transformer (ViT) network model refers to a neural network model that applies the Transformer architecture to visual feature modeling. In this scheme, the Vision Transformer network model is composed of multiple stacked Vision Transformer modules. Each module includes a multi-head self-attention (MHSA) mechanism and a feed-forward network (FFN) to model the global spatial dependencies between teeth on the input sequence of fused tooth structural features.

[0044] A spatial correlation analysis of teeth is performed to obtain global tooth correlation feature data. The specific method is as follows: the feature sequence is processed sequentially through multiple layers of visual transformation modules. Within each layer, a multi-head self-attention mechanism is used to calculate the pairwise attention weights between all tooth features in the sequence. Based on these attention weights, all tooth features are weighted and aggregated, ensuring that each tooth's features can perceive information from all other teeth in the dentition. After multi-layer stacking, the global tooth correlation feature data is output. The calculation process of the multi-head self-attention mechanism uses a scaled dot product attention formula, as detailed below:

[0045] ;

[0046] In the formula, Represents global correlation feature data of teeth. Represents the query matrix. Represents the key matrix. Represents a value matrix, The dimension of the key vector. Indicates matrix transpose. This represents the normalized exponential function.

[0047] Pose regression processing is performed based on global tooth association feature data to obtain rotation parameter data and translation parameter data corresponding to each tooth. Specifically, the global tooth association feature data is used as input, and the fully connected layer (linear regression layer) outputs the rotation quaternion (rotation parameter data) and three-dimensional translation vector (translation parameter data) corresponding to each tooth, which together constitute the six-degree-of-freedom pose prediction data of each tooth.

[0048] By introducing a Vision Transformer network model to model the global spatial dependencies of the fused structural feature sequences of each tooth, and utilizing a multi-head self-attention mechanism, the features of each tooth can perceive the position and geometric information of all other teeth in the dentition, overcoming the shortcomings of existing methods such as weak global dependency modeling, isolated tooth pose prediction, and lack of overall dentition coordination. The scaling operation (divided by the square root of the key vector dimension) in the multi-head self-attention mechanism ensures the numerical stability of the attention weight calculation process, avoids gradient vanishing or exploding problems, and improves the training convergence quality. Representing rotation parameters with rotation quaternions avoids redundant parameters in the rotation matrix and gimbal lock problems, improving the numerical stability and regression accuracy of pose parameters. This ensures that the pose prediction of each tooth fully considers its relative positional relationship in the entire dentition spatial structure, improving the overall coordination and accuracy of the tooth arrangement results.

[0049] The Vision Transformer (ViT) network model is based on the Transformer encoding of the ViT encoder. It consists of three groups of visual transformation modules that move forward sequentially. Each group of modules includes a multi-head self-attention mechanism (MHSA) and a feedforward neural network (FFN). Through a multi-layer structure, it gradually realizes local feature blocks and global structures.

[0050] The network input includes two sets of features: the first is fused structural feature data (obtained from step three, encoding the three-dimensional spatial position and local surface geometry of the teeth); the second is position enhancement features (obtained by adding the tooth centroid coordinates, encoded by the tooth center encoder, to the tooth position encoding element-wise, encoding the overall sequence position prior of the dentition). The visual transformation module adds the output elements of the two modules and inputs them into another visual transformation module, using a third-head self-attention mechanism to model the global spatial dependencies between all teeth in the dentition. Finally, the fully connected layer regresses and outputs the rotation quaternion data (rotation parameter data) and three-dimensional translation parameters (translation parameter data) for each tooth, collectively forming the six-DOF pose prediction data for the teeth.

[0051] Furthermore, the dental arch curve parameter data is obtained by performing center feature encoding on the tooth centroid coordinate data to obtain centroid encoded feature data; performing position enhancement processing on the centroid encoded feature data based on the preset sine and cosine position encoding rules to obtain dental arch predicted feature data; and performing polynomial curve fitting processing on the dental arch predicted feature data to obtain dental arch curve parameter data.

[0052] In this embodiment, the centroid encoded feature data is obtained by calculating the arithmetic mean of the coordinates of all vertices in the sub-grid of each tooth to obtain the three-dimensional centroid coordinates of that tooth. The three-dimensional centroid coordinates of all teeth are then combined into a matrix, input into the tooth center encoder, and the high-dimensional centroid encoded feature data corresponding to each tooth is output.

[0053] Based on a preset sine and cosine positional encoding rule, positional enhancement processing is performed on the centroid-encoded feature data to obtain dental arch prediction feature data. Specifically, according to the preset sine and cosine positional encoding rule, a corresponding fixed positional encoding vector is generated for the sequential position of each tooth in the dental arch (e.g., the first tooth, the second tooth, etc.). The positional encoding vector is then added element-wise to the centroid-encoded feature data of the corresponding tooth to output the dental arch prediction feature data. The sine and cosine positional encoding uses the Transformer positional encoding formula, applying a sine function for even dimensions and a cosine function for odd dimensions. The positional index is divided by the frequency of the corresponding dimension, and then a trigonometric function is input to generate the positional encoding vector.

[0054] The dental arch prediction feature data is input into the dental arch prediction network, which consists of a linear layer, an activation function layer, a batch normalization layer, and a one-dimensional convolutional layer. After forward propagation, the dental arch curve coefficient data is output. Then, after normalization amplification and inverse normalization recovery processing, the final dental arch curve parameter data is obtained.

[0055] By performing central feature encoding on the centroid coordinates of teeth, the three-dimensional centroid coordinates are upgraded to high-dimensional feature vectors, fully expressing the spatial semantic information of the teeth. On this basis, sine and cosine position encoding is superimposed to inject the relative position information of each tooth within the dentition sequence, enabling the network to distinguish teeth with the same spatial location but different sequence positions, thus compensating for the lack of sequence perception capability in pure position encoding. The arch prediction feature data after the fusion of these two information streams is used as input to the arch curve prediction network, enabling the network to predict fourth-order polynomial arch curve parameters end-to-end based on complete position semantics and sequence position information. The arch curve parameters in this scheme are adaptively predicted by the neural network based on the patient's actual dentition data, exhibiting stronger individual adaptability and making the predicted arch curve more closely match the patient's true physiological arch morphology.

[0056] Furthermore, polynomial curve fitting is performed based on the predicted dental arch feature data. Specifically, the following steps are taken: generating corresponding dental arch curve coefficient data based on the predicted dental arch feature data; performing normalization and amplification processing on the dental arch curve coefficient data to obtain training dental arch parameter data; performing dental arch curve fitting processing on the training dental arch parameter data to obtain dental arch fitting result data; and performing inverse normalization recovery processing on the dental arch fitting result data to obtain dental arch curve parameter data.

[0057] In this embodiment, the corresponding dental arch curve coefficient data is generated by inputting the dental arch prediction feature data into a dental arch prediction network consisting of a linear layer, an activation function layer, a batch normalization layer, and a one-dimensional convolutional layer. The original predicted values ​​of the coefficients of each order of the fourth-order polynomial are directly output through forward propagation, which are the dental arch curve coefficient data.

[0058] For each order coefficient in the dental arch curve coefficient data, a corresponding preset scaling factor is multiplied to obtain the training dental arch parameter data. The scaling factor is predetermined through statistical analysis (mean and standard deviation statistics) of all real dental arch curve coefficients in the training dataset and stored as a fixed hyperparameter, remaining consistent during the training and inference phases. Using each order coefficient in the training dental arch parameter data as polynomial coefficients, several points are uniformly sampled within a preset horizontal axis sampling range, and substituted into a fourth-order polynomial function to calculate the corresponding vertical axis coordinates, obtaining the dental arch fitting result data (i.e., the discrete point sequence of the dental arch curve). During the inference phase, the training dental arch parameter data (or its corresponding coefficients) is divided by the scaling factor used during training to restore each order coefficient to a polynomial coefficient value with consistent physical dimensions in the real coordinate system, obtaining the dental arch curve parameter data, which serves as the basis for subsequent calculation of dental arch alignment loss and generation of tooth arrangement results.

[0059] Furthermore, collision avoidance loss data is obtained through the following methods: rigid body transformation is performed on the 3D mesh data of the teeth based on rotation and translation parameter data to obtain tooth pose transformation result data; a preset projection plane mapping process is performed on the tooth pose transformation result data to obtain 2D projection data of the teeth; point distance analysis is performed on the 2D projection data of adjacent teeth to obtain tooth distance distribution data; and collision probability analysis is performed on the tooth distance distribution data and a preset Gaussian kernel function to obtain collision avoidance loss data.

[0060] In this embodiment, the automatic tooth alignment network model outputs predicted parameters after forward propagation of the target dentition in the current training batch. These are the predicted rotational quaternions and translational support values ​​for each tooth output by the Tooth Structure Feature Extraction Network (ViT) in the current batch, rather than the actual labeled values. The collision avoidance loss is calculated entirely based on the network's predicted pose, which transforms the tooth mesh. This allows the loss gradient to propagate back to the network parameters through rigid body transformation, driving the network to autonomously learn and output collision-free alignment poses.

[0061] Performing rigid body transformation on 3D tooth mesh data to obtain tooth pose transformation results refers to the process of using rotation parameter data (rotation quaternions) and translation parameter data (3D translation vectors) to perform rotation and translation transformations on the coordinates of each vertex in the 3D tooth mesh data, transforming the teeth from their initial position to their predicted arrangement position. Specifically, the rotation quaternions are converted into rotation matrices; each vertex coordinate vector is left-multiplied by the rotation matrix and then added to the translation vector to obtain the transformed vertex coordinates. All transformed vertices constitute the tooth pose transformation results.

[0062] Tooth pose transformation result data refers to the set of coordinates of each tooth's 3D mesh point under the predicted alignment pose, obtained after rigid body transformation of the tooth's 3D mesh data. Performing a preset projection plane mapping process yields 2D tooth projection data, which refers to the process of projecting the coordinates of the tooth's 3D vertices onto a preset 2D plane (i.e., the XZ plane, corresponding to the horizontal occlusal plane of the dental arch). Specifically, for each vertex in the tooth pose transformation result data, its X-axis and Z-axis coordinate components are taken, while the Y-axis coordinate component is discarded, resulting in the 2D projection coordinates corresponding to each vertex. The 2D projection coordinates of all vertices constitute the tooth's 2D projection data. The tooth's 2D projection data refers to the set of 2D point coordinates obtained after projecting each vertex of the tooth's 3D mesh onto the XZ plane, reflecting the contour distribution of the teeth within the horizontal occlusal plane.

[0063] The point-to-point distance analysis based on the two-dimensional projection data of adjacent teeth refers to calculating the Euclidean distance between each projection point of one tooth and each projection point of another tooth in the two-dimensional projection data of adjacent teeth, and obtaining the set of distance values ​​between all projection point pairs of the adjacent teeth, that is, the tooth distance distribution data.

[0064] The loss function in this scheme is a point-level collision avoidance loss function based on a two-dimensional projected Gaussian kernel. This loss is based on the following geometric property: when there is excessive proximity or overlap between teeth, the distance between their projected points on the XZ plane (i.e., the occlusal plane) will be significantly reduced. In each batch, for any pair of adjacent teeth... and First, the vertices of the 3D mesh of the teeth are rotated using predicted rotation quaternions. With translation vector Perform a rigid body transformation and project it onto the XZ plane. Then, use the Gaussian kernel function as a measure of the degree of overlap. The collision loss is calculated as follows:

[0065] ;

[0066] In the formula, This represents the numerical value of collision loss. This represents the batch size; in this article, the value is 8. Indicates the batch number. A represents a list of adjacent pairs of teeth, used to describe pairs of teeth that have spatial constraints in the dental arch topology. Its element form is: ; It refers to the number of teeth. Teeth The first on the projection plane A vector of point coordinates, Indicates the index of the point involved in the loss calculation. P represents the number of dots on each tooth. Teeth The first on the projection plane A vector of point coordinates, Indicates the index of the point involved in the loss calculation. , The distance between the two projection points is represented by σ, which represents the bandwidth of the Gaussian kernel function and controls the overlap sensitivity; its value is a positive real number. In this experiment, Set as Where tooth i and tooth Adjacent, together forming a list of adjacent pairs of teeth.

[0067] This loss function encourages maintaining appropriate distances between adjacent tooth point sets, reducing tooth collisions or overlaps in the arrangement results, thereby enhancing the physical plausibility and clinical applicability of the model's generated results. This paper considers a complete dentition, so T is set to 28.

[0068] Collision avoidance loss data is obtained by calculating the square of the Euclidean distance between all samples, all adjacent tooth pairs, and all projected point pairs in a batch, then inputting the result into a Gaussian kernel function, and averaging the Gaussian kernel values ​​of all point pairs to obtain the batch-level average collision loss.

[0069] By projecting a 3D tooth mesh onto the XZ plane (horizontal occlusal plane) in 2D, the 3D tooth collision detection problem is transformed into point-to-point distance analysis on a 2D plane, significantly reducing the computational complexity of collision detection. Simultaneously, the XZ plane projection, representing the most observed occlusal plane perspective in clinical orthodontics, can strongly penalize any degree of tooth overlap at the damage function level, while avoiding excessive constraints on well-separated teeth. The damage function design considers both the equivalence of collisions and tolerance for the alignment of normal teeth. Through continuous monitoring of collision damage during network training, the network learns to autonomously output overlapping tooth postures, fundamentally eliminating problems such as tooth overlap and impaction that violate clinical requirements in existing automatic tooth alignment methods, thus improving the clinical feasibility and safety of the alignment results.

[0070] Furthermore, the specific method for obtaining the dental arch alignment loss data is as follows: generating predicted dental arch trajectory data based on dental arch curve parameter data; extracting the centroid coordinate data of each tooth; performing a preset direction deviation analysis based on the centroid coordinate data of each tooth and the predicted dental arch trajectory data to obtain dental arch deviation data; and performing mean loss calculation processing based on the dental arch deviation data to obtain dental arch alignment loss data.

[0071] In this embodiment, the centroid of the tooth is used as the basis for tooth position analysis.

[0072] The method for generating predicted dental arch trajectory data based on dental arch curve parameter data is as follows: The coefficients of each order of the fourth-order polynomial in the dental arch curve parameter data are used as parameters of the dental arch curve function to construct the dental arch curve function.

[0073] ;

[0074] In the formula, This represents the lateral coordinate of the dental arch along the direction of the dental arch. This represents the corresponding longitudinal coordinate of the dental arch. These represent the coefficients of the fourth to constant terms in the dental arch curve parameter data, i.e., the dental arch curve parameters.

[0075] Uniform sampling is performed within a preset range of dental arch lateral coordinates to obtain multiple sampling lateral coordinate points:

[0076] ;

[0077] The corresponding longitudinal coordinates of the dental arch are calculated as follows:

[0078] ;

[0079] Obtain the dental arch trajectory points:

[0080] ;

[0081] Among them, the actual dental arch curve parameters The method for obtaining it is as follows:

[0082] ;

[0083] Thus, all tooth centroids are obtained, and then a dental arch fitting point set is established to obtain the centroid of each tooth: This forms a sequence of dental arch fitting points.

[0084] Perform a fourth-order polynomial least squares fit:

[0085] ;

[0086] We can obtain the result through matrix solving. .

[0087] The dental arch deviation data is obtained by extracting the X-axis coordinate value from the centroid coordinate data of each tooth, substituting it into the dental arch curve function to obtain the corresponding dental arch Z-axis predicted value, calculating the difference between the actual Z-axis coordinate of each tooth centroid and the dental arch Z-axis predicted value, and obtaining the dental arch deviation value in the Z-axis direction for each tooth.

[0088] The dental arch alignment loss data is obtained by performing a mean squared deviation calculation on the Z-axis deviation values ​​of all teeth to obtain the dental arch correction loss data. This loss is calculated using the centroid of the tooth as a reference, calculating its deviation from the predicted dental arch curve in the Z-axis direction. This constrains the tooth arrangement to approach and align with the fitted dental arch trajectory. The dental arch curve is modeled using a fourth-order polynomial. The final dental arch alignment loss is defined as the mean error between the predicted curve and the actual tooth centroid in the Z-axis direction, using the following formula:

[0089] ;

[0090] in, The value represents the dental arch loss, B represents the batch size (taken as 8 in this paper), and T represents the number of centroids of the teeth (taken as 28). Let z be the z-coordinate vector of the centroid of the i-th tooth predicted by the model in the b-th batch. This is the z-coordinate vector of the corresponding centroid on the actual dental arch.

[0091] By explicitly defining the common loss of the dental arch as the deviation of the centroid of each tooth in the longitudinal direction of the Z-axis (i.e., the longitudinal direction perpendicular to the tangent of the dental arch) from the predicted dental arch curve, a constraint on the longitudinal deviation of the dentition is provided. This allows the network to provide precise guidance on the rationality of the arrangement of each tooth along the tangent of the dental arch during training, addressing the issue of the transverse distribution of tooth centroids in the longitudinal direction of the dental arch. This creates a greater steepness, driving the network to prioritize the correction of teeth with the most severe non-overlap, thus improving the adaptability of the dental arch alignment during the process. By implementing dual-dimensional constraints of transverse non-collision and longitudinal arch fit, the final tooth arrangement meets both the clinical safety requirement of non-overlapping teeth and the anatomical requirement of overall dentition fit to the natural physiological morphology of the dental arch, improving the accuracy of the automatic tooth arrangement results.

[0092] like Figure 4 As shown, Figure 4 This is a schematic diagram of the arch alignment loss constraint of the present invention. The red line represents the arch curve predicted by the arch line prediction network, and the red arrow represents the constraint of the arch alignment loss on the tooth arrangement. This loss can effectively enhance the consistency and regularity of the teeth arrangement along the arch, reduce misalignment and abrupt arrangement, and thus improve the overall structural rationality of the tooth arrangement model.

[0093] Furthermore, the automatic tooth alignment result data is obtained as follows: a joint loss function is constructed based on collision avoidance loss data and arch alignment loss data; the automatic tooth alignment network is subjected to parameter iterative update processing based on the joint loss function to obtain target network parameter data; tooth pose prediction processing is performed on the target dentition based on the target network parameter data to obtain target tooth pose data; and tooth arrangement transformation processing is performed on the 3D mesh data of the teeth based on the target tooth pose data to obtain automatic tooth alignment result data.

[0094] In this embodiment, a joint loss function is constructed, specifically by: combining collision avoidance loss data... Data on loss of dental arch correction Weighted summation is used to construct a joint loss function, which serves as the overall optimization objective of the automatic tooth alignment network. The joint loss function is constructed as follows:

[0095] ;

[0096] in, Denotes the joint loss function. This indicates a collision to avoid damage. Indicates dental arch alignment loss. Indicates the collision loss weight. This indicates the weight of dental arch loss.

[0097] The collision loss weight and dental arch loss weight can be obtained through a validation set grid search method. Specifically, both the collision loss weight and dental arch loss weight are set to 1. Then, on the validation set, the following are calculated: ADD (Average Distance of Model Points), collision (number of tooth collisions), and ADE (Arch Deviation Error). If the collision is too high, it is increased. The collision loss weight is increased if ADE is high. Dental arch loss weights are then applied, followed by a grid search method:

[0098] ;

[0099] The combined experiment selects the parameter that minimizes the overall error of the validation set, thus obtaining... Collision loss weights Dental arch loss weight.

[0100] The joint loss function calculates the gradient of all learnable parameters in the automatic tooth alignment network (implemented through the backpropagation algorithm). The network parameters are iteratively updated using a gradient descent optimization algorithm (such as the Adam optimizer). The above process is repeated until the joint loss function converges, and the target network parameter data is obtained.

[0101] During the inference phase, the automatic tooth alignment network is loaded with the target network parameter data. The 3D mesh data of the target dentition is input into the network for forward propagation, and the target tooth pose data (rotation quaternions and translation control) corresponding to each tooth is output. Using the rotation quaternions and translation corrections in the target tooth position data, a rigid body transformation is performed on each vertex in the original 3D mesh data of each tooth (first converting the quaternion into a rotation matrix, left-multiplying the vertex coordinates, and then adding translation support), transforming each tooth to the predicted arrangement position. The results are then summarized to obtain the complete 3D dentition mesh data after alignment, i.e., the automatic tooth alignment result data.

[0102] Target tooth pose data refers to the six-degree-of-freedom pose parameters (rotation quaternions and translation movements) of each tooth output after forward propagation of the target dentition using the target network parameter data during the inference phase, which is the final predicted tooth arrangement pose.

[0103] By constructing collision avoidance and dental arch loss as a joint loss function, multi-objective effective supervision of the automatic tooth alignment network is achieved. This allows the network to simultaneously learn and output collision-free tooth alignment poses that conform to the natural dental arch shape during a single training iteration. This enables independent manual intervention in collision handling and dental arch correction, improving the automation and efficiency of the tooth alignment process. Through an iterative parameter update mechanism involving backpropagation and progressive recovery optimization, the network can autonomously learn tooth alignment patterns from a large amount of clinical orthodontic data, continuously improving its generalization ability as the training data scale increases.

[0104] It should be noted that the 3D dental data used in this experiment came from the 3D dental scan database (3DDSD). The database contains 300 pairs of 3D dental arch data before and after orthodontic treatment. Each sample is represented by a 3D mesh of a single tooth, manually annotated by a professional orthodontist, including the segmented 3D mesh and tooth number for each tooth. This paper divides the dataset into training, validation, and test sets in a 7.0:1.5:1.5 ratio, containing 350, 75, and 75 sets of dental arch data, respectively. The training set is used for model parameter learning, the validation set for hyperparameter tuning and model selection, and the test set for final quantitative performance evaluation.

[0105] This scheme improves convergence efficiency and reduces sensitivity to the initial learning rate by dynamically adjusting the momentum estimate of the gradient and adaptively assigning a learning rate to each parameter. The initial learning rate is set to... The weight decays to The batch size was set to 8. During training, a cosine annealing learning rate strategy combined with stochastic gradient descent with warm restarts (SGDR) was employed. Each cycle was 100 epochs long, for a total of 500 training epochs. Within each cycle, the learning rate decreased from an initial value of 1×10⁻⁴ to a minimum of 1×10⁻⁶ using a cosine function, and was reset to the initial value at the end of the cycle to encourage the model to escape local optima. The choice of cycle length was based on observations of changes in training loss and key metrics. Experiments showed that 100 epochs achieved a good balance between local convergence and global exploration, improving model performance and training efficiency.

[0106] This scheme was evaluated against PSTN, TANet, and TaligNet on a unified test set using various metrics, including average point distance, rotation angle error, translation error, and the dental arch deviation error proposed in this paper.

[0107] The average distance of model points (ADD) measures the average distance between the predicted teeth and the actual teeth, and its calculation process is shown below:

[0108] ;

[0109] in, The first set of pre-orthodontic tooth data Let R and t be the coordinates of a point, respectively, and R and t be the predicted rotation quaternion and translation vector, respectively. and The value is the true value, and N is the number of tooth sampling points.

[0110] Rotation error (Rot) measures the angular deviation between the predicted rotation matrix and the true rotation matrix. Its calculation process is shown below:

[0111] ;

[0112] Where B represents the batch size, i.e., the number of samples input in each training iteration, and M represents the number of teeth in a single sample. In this paper, a complete dentition is considered, so the value is 28. Let represent the true rotation quaternion of the j-th tooth in the b-th sample. This represents the rotational quaternion of the corresponding tooth predicted by the model.

[0113] Translation error (Trans) measures the deviation between the predicted translation vector and the true translation vector. Its calculation process is shown below:

[0114] ;

[0115] Arch deviation error (ADE) measures the distance between the centroid of the tooth and the actual arch curve. Its calculation process is shown below:

[0116] ;

[0117] Where N is the number of teeth. This represents the shortest distance from the tooth centroid to the dental arch curve. Cosine similarity of axis (CSA) measures the directional consistency between the predicted and actual rotations. The corresponding rotation axis vectors are extracted from both the predicted and actual rotation quaternions, and then the cosine similarity between them is calculated as the CSA value.

[0118] Table 1 Comparison of Tooth Arrangement Results

[0119] Model ADD / mm CSA Rot / (º) Trans / mm ADE / mm TaligNet 1.612 0.793 6.241 0.290 1.692 PSTN 1.391 0.818 8.033 0.156 1.795 TANet 1.183 0.842 6.501 0.093 1.307 This article's method 1.089 0.861 6.277 0.074 1.113

[0120] As shown in Table 1, the proposed method outperforms the comparative methods in most metrics. The improvement in ADD and ADE metrics is particularly significant, indicating a clear advantage in ensuring the overall structural harmony and alignment accuracy of teeth.

[0121] like Figure 5 As shown, Figure 5 This diagram illustrates the accuracy curves of the present invention, showing a comparison of the accuracy of different methods under different ADD thresholds. The solid red line represents the method described in this paper.

[0122] like Figure 6 As shown, Figure 6 This diagram illustrates the qualitative comparison of tooth arrangement results. The first to third rows represent the tooth arrangement effects of the entire dentition, mandibular dentition, and maxillary dentition, respectively. From left to right, the diagram shows the input, TAligNet, PSTN, TANet, the proposed method, and the actual results. Orange, red, and blue boxes mark typical comparison areas for the entire dentition, mandibular dentition, and maxillary dentition, respectively. The comparison results show that existing comparison methods still suffer from problems such as misaligned teeth or uneven interdental spacing during tooth arrangement. In contrast, the proposed method outputs a more regular and physiologically consistent tooth arrangement structure, with uniform interdental spacing, a smooth and continuous overall dental arch morphology, and tooth arrangement results that better conform to clinical orthodontic tooth arrangement principles.

[0123] To analyze the modeling effect of different dentition regions, the sampling points were divided into anterior tooth region (incisors and canines) and posterior tooth region (premolars and molars) according to the anatomical position of the teeth. The error index of the corresponding region was calculated, and the statistical results are shown in Table 2.

[0124] Table 2. Quantitative Table of Predictive Errors for Dental Arch Line

[0125] area MAEarch / mm RMSEarch / mm Full dentition 0.42 0.58 Anterior teeth region 0.39 0.54 Posterior teeth region 0.45 0.61

[0126] This approach also includes validation of the effectiveness of collision avoidance loss. To verify the role of collision avoidance loss in tooth alignment optimization, an ablation experiment was conducted. The collision avoidance loss constraint was removed, leaving only the base loss, and the model was retrained under completely identical training configurations. As shown in Table 3, the number of colliding teeth is defined as the average number of teeth colliding in each group of dentition. After removing the collision avoidance loss constraint, the phenomena of tooth overlap and impaction increased significantly, resulting in a significant decrease in both the structural rationality and clinical usability of the tooth alignment results.

[0127] Table 3. Quantitative Comparison of Tooth Arrangement Quality by Collision Module

[0128] Model ADD / mm CSA Rot / (º) Trans / mm collision There is a collision module 1.089 0.861 6.277 0.074 0.115 Collision-free module 1.179 0.842 6.523 0.133 3.235

[0129] like Figure 7 As shown, Figure 7 This is a qualitative comparative diagram of the collision avoidance module of the present invention. The left side shows the predicted tooth arrangement results without introducing collision avoidance loss, and the right side shows the optimized results after introducing this loss. Typical areas are marked with red boxes. It is easy to see that without constraints, there is a tendency for obvious overlap or insufficient spacing between adjacent teeth. After introducing collision avoidance loss constraints, tooth collision and impaction are significantly suppressed, and the interdental space arrangement is more reasonable.

[0130] To verify the role of the archline prediction network in tooth alignment optimization, this study designed a corresponding ablation experiment. The archline prediction network was removed, and the model was retrained under completely identical training configurations.

[0131] like Figure 8 As shown, Figure 8 This is a qualitative comparative schematic diagram of the dental arch line prediction network of the present invention. The diagram shows a visual comparison of tooth arrangement with and without the dental arch line prediction network (projected onto the XZ plane). The red dots represent the centroid of each tooth, and the blue curve represents the actual dental arch curve. The comparison shows that without the dental arch line prediction network, the centroids of some teeth deviate significantly from the standard dental arch trajectory, resulting in more pronounced localized arrangement deviations. After introducing the dental arch line prediction network, the overall distribution of the centroids of each tooth more closely matches the actual dental arch curve. This confirms that the network can effectively guide the teeth to arrange themselves along the physiological dental arch morphology, thereby improving the overall consistency of the dental arch structure.

[0132] like Figure 9 As shown, Figure 9 The system architecture diagram of this invention is shown below. The automatic tooth alignment system based on multi-source feature fusion and dental arch trajectory prediction provided in this application includes the following modules: a data acquisition module, used to acquire three-dimensional mesh data of the target dentition teeth and extract tooth point coordinate data, tooth point normal vector data, and tooth centroid coordinate data; a local feature extraction module, used to analyze tooth point normal vector data to obtain local geometric feature data of the teeth; a feature fusion analysis module, used to perform position enhancement processing based on tooth point coordinate data to obtain tooth position feature data, and perform feature fusion processing based on the local geometric feature data and tooth position feature data to obtain fused structural feature data; and a pose prediction module, used for... Global spatial dependency modeling is performed based on the fused structural feature data to obtain tooth pose prediction data for each tooth. The tooth pose prediction data includes rotation parameter data and translation parameter data. The dental arch trajectory prediction module is used to perform center feature encoding and position enhancement processing based on the tooth centroid coordinate data to obtain dental arch prediction feature data, and then perform dental arch trajectory prediction processing to obtain dental arch curve parameter data. The iterative optimization module is used to perform tooth alignment constraint analysis based on the tooth pose prediction data and dental arch curve parameter data to obtain collision avoidance loss data and dental arch alignment loss data, and perform network iterative optimization processing to output the automatic tooth alignment result data corresponding to the target dentition.

[0133] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.

Claims

1. An automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction, characterized in that: Includes the following steps: Obtain the 3D mesh data of the target dentition and extract the tooth point coordinate data, tooth point normal vector data, and tooth centroid coordinate data; Based on the analysis of tooth point normal vector data, local geometric feature data of teeth are obtained; Position enhancement processing is performed based on tooth point coordinate data to obtain tooth position feature data. Feature fusion processing is then performed based on tooth local geometric feature data and tooth position feature data to obtain fused structural feature data. Global spatial dependency modeling is performed based on the fused structural feature data to obtain tooth pose prediction data for each tooth. The tooth pose prediction data includes rotation parameter data and translation parameter data. Based on the tooth centroid coordinate data, center feature encoding and position enhancement processing are performed to obtain dental arch prediction feature data. Then, dental arch trajectory prediction processing is performed to obtain dental arch curve parameter data. Based on tooth pose prediction data and dental arch curve parameter data, tooth alignment constraint analysis is performed to obtain collision avoidance loss data and dental arch alignment loss data. Then, network iterative optimization processing is performed to output the automatic tooth alignment result data corresponding to the target tooth row.

2. The automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction according to claim 1, characterized in that: The specific method for obtaining the local geometric feature data of the teeth is as follows: Obtain the mesh vertex data corresponding to each tooth in the target dentition, and perform adjacency analysis to obtain the set of adjacent vertices inside a single tooth, thereby constructing the topology graph structure of a single tooth; Based on the single-tooth topological graph structure, graph convolution aggregation is performed on the tooth point normal vector data to obtain the local topological feature data corresponding to each grid vertex. Channel fusion processing is performed based on the local topological feature data to obtain the local geometric feature data of the teeth.

3. The automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction according to claim 1, characterized in that: The specific method for obtaining the fused structural feature data is as follows: Local encoding is performed on the tooth point coordinate data to obtain point coordinate encoded feature data; corresponding dimension alignment processing is performed on the point coordinate encoded feature data and the tooth local geometric feature data to obtain feature alignment result data; Element-by-element fusion processing is performed based on the feature alignment result data to obtain fused structural feature data.

4. The automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction according to claim 1, characterized in that: The specific method for obtaining the tooth pose prediction data for each tooth is as follows: The fused structural feature data is input into a pre-defined visual transformation network model; Based on the multi-head self-attention mechanism in the visual transformation network model, tooth spatial association analysis is performed to obtain global tooth association feature data. Pose regression processing is performed based on global tooth association feature data to obtain rotational and translational parameter data for each tooth.

5. The automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction according to claim 1, characterized in that: The specific method for obtaining dental arch curve parameter data is as follows: Perform center feature encoding on the tooth centroid coordinate data to obtain centroid encoded feature data; Position enhancement processing is performed on the centroid encoded feature data based on the preset sine and cosine position encoding rules to obtain dental arch prediction feature data; Polynomial curve fitting is performed based on the predicted feature data of the dental arch to obtain the dental arch curve parameter data.

6. The automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction according to claim 5, characterized in that: The specific method for performing polynomial curve fitting processing based on dental arch prediction feature data is as follows: Generate corresponding dental arch curve coefficient data based on dental arch prediction feature data; Normalize and amplify the dental arch curve coefficient data to obtain the training dental arch parameter data; Based on the training dental arch parameter data, dental arch curve fitting is performed to obtain dental arch fitting result data; Inverse normalization recovery processing was performed on the dental arch fitting data to obtain the dental arch curve parameter data.

7. The automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction according to claim 1, characterized in that: The specific method for obtaining collision avoidance loss data is as follows: Rigid body transformation is performed on the 3D mesh data of teeth based on rotation parameter data and translation parameter data to obtain tooth pose transformation result data; Perform a preset projection plane mapping process on the tooth pose transformation result data to obtain two-dimensional projection data of the teeth; Based on the two-dimensional projection data of adjacent teeth, point distance analysis is performed to obtain tooth distance distribution data; Collision probability analysis is performed based on tooth distance distribution data and a preset Gaussian kernel function to obtain collision avoidance loss data. The method for calculating the collision loss is as follows: ; In the formula, This represents the numerical value of collision loss. Represents batch size. Indicates the batch number. Table A is a list of adjacent pairs of teeth, used to describe pairs of teeth that have spatial constraints in the dental arch topology. Its element form is: ; It refers to the number of teeth. Teeth The first on the projection plane A vector of point coordinates, Indicates the index of the point involved in the loss calculation. P represents the number of dots on each tooth. Teeth The first on the projection plane A vector of point coordinates, Indicates the index of the point involved in the loss calculation. σ represents the bandwidth of the Gaussian kernel function, which controls the degree of overlap sensitivity, and its value is a positive real number.

8. The automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction according to claim 1, characterized in that: The specific method for obtaining the dental arch alignment loss data is as follows: Predicted dental arch trajectory data is generated based on dental arch curve parameter data; Extract the centroid coordinates of each tooth; Based on the centroid coordinate data of each tooth and the predicted dental arch trajectory data, a preset direction deviation analysis is performed to obtain dental arch deviation data; The mean loss calculation is performed based on the dental arch deviation data to obtain the dental arch alignment loss data.

9. The automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction according to claim 1, characterized in that: The method for obtaining the automatic tooth alignment result data is as follows: A joint loss function is constructed based on collision avoidance loss data and dental arch alignment loss data; The automatic tooth alignment network is subjected to parameter iterative update processing based on the joint loss function to obtain the target network parameter data; Based on the target network parameter data, tooth pose prediction processing is performed on the target dentition to obtain the target tooth pose data; Based on the target tooth pose data, tooth alignment transformation is performed on the three-dimensional mesh data of the teeth to obtain automatic tooth alignment result data.

10. A system applying the automatic tooth alignment method based on multi-source feature fusion and dental arch trajectory prediction as described in any one of claims 1-9, characterized in that: Includes the following modules: The data acquisition module is used to acquire the three-dimensional mesh data of the target dentition and extract the tooth point coordinate data, tooth point normal vector data, and tooth centroid coordinate data. The local feature extraction module is used to analyze tooth point normal vector data to obtain local geometric feature data of teeth; The feature fusion analysis module is used to perform position enhancement processing based on tooth point coordinate data to obtain tooth position feature data, and to perform feature fusion processing based on tooth local geometric feature data and tooth position feature data to obtain fused structural feature data. The pose prediction module is used to perform global spatial dependency modeling based on fused structural feature data to obtain tooth pose prediction data for each tooth. The tooth pose prediction data includes rotation parameter data and translation parameter data. The dental arch trajectory prediction module is used to perform center feature encoding and position enhancement processing based on tooth centroid coordinate data to obtain dental arch prediction feature data, and then perform dental arch trajectory prediction processing to obtain dental arch curve parameter data. The iterative optimization module is used to perform tooth alignment constraint analysis based on tooth pose prediction data and dental arch curve parameter data, obtain collision avoidance loss data and dental arch alignment loss data, perform network iterative optimization processing, and output the automatic tooth alignment result data corresponding to the target dentition.