Face mouth shape coordination generation method based on geometric constraint

By constructing a mouth opening and closing spatial morphology matrix and an alveolar pose reference system, and combining geometric prior constraint tensors with texture fusion, the problem of inaccurate oral cavity region modeling in existing technologies is solved, achieving a high-precision and natural mouth shape synthesis effect.

CN120543764BActive Publication Date: 2026-03-03CLOUD ATTACK NETWORK TECH HEBEI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510876947.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2026-03-03
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Existing face modeling methods suffer from structural simplification when dealing with the oral cavity region, failing to accurately represent complex structures such as teeth, gums, and alveolar bone. This results in issues such as missing tooth texture, blurred lip-tooth boundaries, and 'hollowing out' of the mouth area when generating mouth images.

Method used

By constructing a mouth opening and closing spatial morphology matrix, an alveolar pose reference system, and a geometric prior constraint tensor, and combining multimodal feature transformation with confidence-weighted texture fusion, we can achieve fine modeling and dynamic control of the lip and tooth boundaries and the internal structure of the oral cavity, thereby improving the realism, geometric consistency, and temporal coherence of tooth textures.

Benefits of technology

It achieves high-precision, temporally continuous, and visually natural mouth shape synthesis effects, solves the problems of discontinuous lip and tooth structures and texture jumps, and improves the realism and stability of tooth texture details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543764B_ABST
    Figure CN120543764B_ABST
Patent Text Reader

Abstract

The application discloses a face mouth shape coordination generation method based on geometric constraints and relates to the technical field of image generation.The method comprises the following steps: step 1: synchronously generating a face vertex index table corresponding to a mouth opening and closing space form matrix; step 2: based on the face vertex index table, mapping a lip gingival joint boundary constraint set to a unified coordinate system to form a geometric prior constraint tensor; step 3: performing a multi-modal topological cascade reversible mixed transformation on the geometric prior constraint tensor to obtain a high-dimensional tooth texture candidate tensor and a constraint consistency confidence spectrum; and step 4: using the high-dimensional tooth texture candidate tensor and the constraint consistency confidence spectrum as input, completing anti-aliasing, illumination compensation and gamma correction on a unified rendering pipeline, and finally obtaining a mouth shape coordination synthesis frame sequence.The application effectively improves the realism, geometric consistency and time sequence coherence of tooth texture in the mouth shape synthesis process, and solves the problem of uncoordinated lips and teeth in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image generation technology, specifically to a method for generating facial lip shapes based on geometric constraints. Background Technology

[0002] Against the backdrop of the rapid development of facial modeling and expression-driven technologies, key issues surrounding lip generation, particularly the modeling and rendering accuracy of the geometry and texture representation within the oral cavity, have become common challenges in numerous application scenarios. With the increasing demand for virtual digital humans, voice-synchronized animation, video reconstruction, and post-production lip replacement in film and television, the realism and consistency of facial lip movements directly impact the usability and visual consistency of the entire facial driving system. However, existing technical solutions still suffer from a series of key technical shortcomings in the modeling and synthesis of the oral cavity region, particularly in areas such as multimodal alignment, anatomical constraints, and texture generation consistency.

[0003] Current mainstream face modeling methods typically rely on 3D deformation models (such as 3DMM or FLAME models) or deep learning-driven implicit shape representations for 3D fitting. These models can reconstruct the overall head shape, expression, and pose parameters relatively well, and ensure the accuracy of surface structures through keypoint alignment. However, they often suffer from structural simplification when dealing with the oral cavity region. Specifically, to ensure the uniformity of model training and convergence speed, many publicly available models adopt a uniform topology, where the interior of the oral cavity is abstracted as a simple opening and closing parameter or masking layer. While this modeling approach can be used to roughly control the opening and closing of the mouth, it cannot accurately represent the true spatial distribution of complex structures such as teeth, gums, and alveolar bone, especially under intense facial expressions such as speaking, smiling, or grinning, lacking the ability to resolve the junction between the lips and gums. This deficiency directly leads to frequent problems such as missing tooth texture, blurred lip and tooth boundaries, and "hollowing out" of the opening area when generating mouth shape images. Summary of the Invention

[0004] To address the aforementioned technical challenges, a fully geometrically constrained method for generating coordinated facial and lip shapes is provided. By constructing a mouth opening and closing spatial morphology matrix, an alveolar pose reference system, and a geometric prior constraint tensor, this method achieves precise modeling and dynamic control of lip and tooth boundaries and internal oral structures. Combined with multimodal feature transformation and confidence-weighted texture fusion, it effectively improves the realism, geometric consistency, and temporal coherence of tooth textures during lip shape synthesis. This method solves the problems of lip and tooth incoordination, texture jumps, and rendering distortion in existing methods, and has significant advantages such as high fidelity, structural stability, and strong adaptability.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A face-lip shape coordination generation method based on geometric constraints, the method comprising:

[0007] Step 1: Use a 3D deformation model to initialize the 3D shape of the face image to be processed, obtain the mouth opening and closing spatial shape matrix, and simultaneously generate a facial vertex index table that corresponds one-to-one with the mouth opening and closing spatial shape matrix.

[0008] Step 2: Based on the facial vertex index table, extract the lip contour, gingival boundary and alveolar pose, construct the lip-gingival joint boundary constraint set, and map the lip-gingival joint boundary constraint set to a unified coordinate system to form a geometric prior constraint tensor;

[0009] Step 3: Perform a multimodal topological concatenation invertible hybrid transformation on the geometric prior constraint tensor to obtain a high-dimensional tooth texture candidate tensor and a constraint consistency confidence spectrum;

[0010] Step 4: Using high-dimensional tooth texture candidate tensors and constraint consistency confidence spectra as input, the matrix is ​​mapped according to the mouth opening and closing spatial morphology matrix, and anti-aliasing, lighting compensation and gamma correction are completed in the unified rendering pipeline to finally obtain the mouth shape coordinated composite frame sequence.

[0011] Further, step 1 specifically includes: acquiring the face image to be processed, which is a static image or dynamic frame image sequence containing at least one frontal or near-frontal view of the face; performing face detection and facial key point localization operations on the face image to be processed, extracting the coordinates of two-dimensional facial feature points including eyebrows, eyes, nose wings, and lips, forming a two-dimensional facial feature point set; based on the two-dimensional facial feature point set, constructing an initial corresponding mapping relationship matching the three-dimensional deformation model, and calling the three-dimensional deformation model to perform three-dimensional fitting on the face image to be processed, outputting a three-dimensional morphological parameter set, including shape parameters, expression parameters, and pose parameters; and recovering the face image corresponding to the face image to be processed based on the three-dimensional morphological parameter set. A 3D deformation model with consistent resolution is processed for facial images. The 3D deformation model consists of facial vertices, each corresponding to a 3D coordinate value. In the 3D deformation model, for a subset of vertices related to the upper and lower boundaries of the lips and the opening and closing area of ​​the mouth, the corresponding 3D coordinate changes in each frame image are extracted, a continuous descriptive function for the degree of lip closure is established, and a mouth opening and closing spatial morphology matrix is ​​generated. The mouth opening and closing spatial morphology matrix is ​​used to represent the spatial configuration of the oral cavity at different time points or under different facial expressions. All vertices in the 3D deformation model are numbered according to their original order in the model initialization stage, and their spatial coordinate indices and topological connections are recorded to form a facial vertex index table.

[0012] Furthermore, the three-dimensional deformation model adopts a fixed-topology face triangular mesh, which contains 53,215 facial vertices and 105,054 edges, and each vertex retains its number throughout the lifecycle of the three-dimensional deformation model; the fixed topology covers the complete visible area from the hairline to the mandible, and deformable vertices are reserved at the inner edge of the lips to ensure the accuracy and consistency of the subsequent mouth opening and closing spatial morphology matrix.

[0013] Furthermore, in step 2, the predefined range of vertex numbers for the outer edge of the upper lip and the outer edge of the lower lip are retrieved from the facial vertex index table to obtain subsets of the outer edges of the upper and lower lip vertices, respectively. Based on the topological connection relationship of the 3D deformation model, the subsets of the outer edges of the upper and lower lip vertices are traversed ring by ring to obtain an ordered sequence of vertices that monotonically increases along the lip contour direction, forming a lip contour polygon chain. The range of vertex numbers for the inner edge of the gingiva is retrieved from the facial vertex index table to obtain a subset of the inner edge of the gingiva vertices. Based on the normal vector direction of the vertices and the local curvature threshold, redundant vertices connected to the hard palate are removed, and only vertices located in the gingival-alveolar junction region are retained to generate a purified gingival boundary sequence.

[0014] Furthermore, in step 2, the centroid of each vertex of the purified gingival boundary sequence is calculated and denoted as the alveolar center point; the right-handed coordinate system formed by the alveolar center point, the alveolar major axis direction, and the alveolar vertical direction is defined as the alveolar pose reference system; the lip contour polygon chain, the purified gingival boundary sequence, and the alveolar pose reference system are collectively encapsulated into a lip-gingival joint boundary constraint set; within the lip-gingival joint boundary constraint set, the original three-dimensional coordinate value, normal vector, curvature scalar, and boundary type label are recorded for each vertex.

[0015] Furthermore, in step 2, with the alveolar center point as the origin, the three-axis vector of the alveolar pose reference system is used as the basis vector to define the oral cavity unified coordinate system; a rigid coordinate transformation is performed on all vertices within the lip-gingival joint boundary constraint set to map the vertex 3D coordinates from the global coordinate system of the 3D deformation model to the oral cavity unified coordinate system; the mapped vertex 3D coordinates, normal vectors, curvature scalars, and their respective boundary type labels are tensorized and stored to form a geometric prior constraint tensor; a normalization operation is applied to the geometric prior constraint tensor so that the 3D coordinates are distributed in the interval -1 to 1, the normal vectors are distributed in the interval -1 to 1, and the curvature scalars are distributed in the interval 0 to 1.

[0016] Furthermore, in step 3, the geometric prior constraint tensor is separated according to four attribute dimensions: vertex 3D coordinates, normal vector, curvature scalar, and boundary type label, resulting in four unimodal tensor branches. For the 3D coordinates in each unimodal tensor branch, a stereo projection method is used to embed Euclidean space points into a hypersphere of radius one, forming a hypersphere embedding tensor. Linear dimension expansion is performed on the normal vector, curvature scalar, and boundary type label to make their dimensions consistent with the hypersphere embedding tensor. The topological folding kernel function is called in batches on the hypersphere embedding tensor. This function uses an invertible torus mapping to project the tensor onto the hypersphere. The radial amplitude on the surface is compressed into a hyperbolic amplitude, and absolute phase information is preserved at each turning point. The cascaded multiple access sampling function is invoked to perform multi-hop sampling with an exponentially decaying step size in the turning domain to obtain cross-scale feature representation. The topological turning kernel function and the cascaded multiple access sampling function alternate to form a complete iteration. After each iteration, the global topological energy index is calculated. If the index decreases by less than one-thousandth compared to the previous iteration, early convergence is triggered; otherwise, the next iteration continues. If the number of iterations reaches one hundred and convergence is still not achieved, the result of the iteration with the lowest energy is taken as the provisional output tensor.

[0017] Further, in step 3, the provisional output tensor is input into a greedy minimum topological energy path searcher; the searcher uses the global topological energy index as the cost function and performs a depth-first traversal on the high-dimensional energy graph of the tensor, evaluating the geometric consistency cost of each candidate path in real time; when the geometric consistency cost first falls below a preset threshold of 50%, the current path is marked as the optimal topological path and the search is terminated; a reversible hybrid reverse mapping is performed on the tensor on the optimal topological path, successively undoing the transformations of the cascaded multiple access sampling function and the topological foldback kernel function, to obtain the reconstructed mode tensor back to Euclidean space; the four reconstructed mode tensors are then arranged according to... The original modal order is reassembled to obtain a candidate 3D array of tooth textures. The first 128 elements of the candidate 3D array of tooth textures are taken and stacked according to their original sequence to form a high-dimensional candidate tooth texture tensor. For each candidate instance in the high-dimensional candidate tooth texture tensor, the Euclidean consistency distance between it and the geometric prior constraint tensor in the three attribute dimensions of vertex 3D coordinates, normal vector, and curvature scalar is calculated. The Euclidean consistency distance is mapped to a confidence level in the interval 0 to 1 using soft maximum operation to obtain the corresponding single instance confidence level value. The 128 single instance confidence levels are arranged in order to form a constraint consistency confidence spectrum of length 128.

[0018] Further, step 4 specifically includes: mapping the high-dimensional tooth texture candidate tensor to the constraint consistency confidence spectrum one-to-one according to the candidate sequence number dimension; applying the single instance confidence value corresponding to the constraint consistency confidence spectrum to each candidate texture instance in the high-dimensional tooth texture candidate tensor as a fusion weight to obtain a confidence-weighted tooth texture fusion tensor; performing a weighted summation on the confidence-weighted tooth texture fusion tensor along the candidate sequence number dimension to obtain a single fused tooth texture representation vector, which is used to represent the optimal tooth texture state corresponding to the current input sample; expanding the fused tooth texture matrix into a one-dimensional vector column-wise, and rearranging it into a two-dimensional texture mesh according to the unified UV layout table of the tooth region; within the two-dimensional texture mesh, performing bidirectional Laplacian interpolation to fill missing pixels and filling edge pixels... Gaussian weighted smoothing is performed to form an initial tooth texture map. Local histogram equalization is applied to the initial tooth texture map to enhance enamel highlights and neck shadow details, outputting an enhanced tooth texture map. The mouth opening and closing spatial morphology matrix corresponding to the enhanced tooth texture map is obtained. For each time frame in the mouth opening and closing spatial morphology matrix, the corresponding oral region vertex subset in the facial vertex index table is located. The two-dimensional texture coordinates of the enhanced tooth texture map are mapped to the oral region vertex subset in the three-dimensional deformation model according to the mouth opening and closing spatial morphology matrix of the current frame, generating a frame-level tooth texture vertex map dataset. Anti-aliasing, lighting compensation, and gamma correction are performed on each frame in the frame-level tooth texture vertex map dataset in a unified rendering pipeline, finally obtaining a mouth shape coordinated composite frame sequence.

[0019] Compared with existing technologies, the advantages of this invention are: it can achieve high-precision, temporally continuous, and visually natural mouth shape synthesis while maintaining the consistency between the real facial structure and the internal and external oral cavity. By introducing a three-dimensional deformation model and constructing a mouth opening and closing spatial morphology matrix, this invention achieves quantitative modeling of the dynamic geometric behavior of the oral cavity, effectively solving the problems of discontinuous lip and tooth structures and inaccurate control of mouth opening and closing states in traditional methods. Simultaneously, the alveolar pose reference system and unified oral cavity coordinate system defined in this invention ensure that the boundary expression of the lip margin and gingiva remains consistent across multiple frames, significantly improving the spatial mapping accuracy across time steps. In terms of texture generation, this invention guides the generation and selection of candidate textures through geometric prior constraint tensors and uses a confidence-weighted fusion strategy to avoid texture jumps and distortions caused by traditional hard matching, further improving the realism and stability of tooth texture details. Furthermore, the anti-aliasing, lighting compensation, and gamma correction mechanisms introduced in the rendering process ensure that the final synthesized frame sequence maintains a natural color and clear contrast visual effect under different lighting and display environments. Overall, this invention can balance anatomical consistency, texture continuity, and optical naturalness, and has broad application prospects in fields such as facial animation, speech-driven synthesis, virtual digital humans, and high-quality lip-syncing. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the method flow for the geometrically constrained face and lip shape coordination generation method proposed in this invention;

[0021] Figure 2 This is a schematic diagram of the generation of the three-dimensional deformation model and the mouth opening and closing spatial morphology matrix in an embodiment of the present invention;

[0022] Figure 3 This is a schematic diagram illustrating the extraction of lip contour and gingival boundary in an embodiment of the present invention.

[0023] Figure 4 This is a schematic diagram of the hyperspherical embedding in an embodiment of the present invention;

[0024] Figure 5 This is a schematic diagram of the topological return kernel function and the cascaded multiple access sampling function in an embodiment of the present invention. Detailed Implementation

[0025] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.

[0026] Reference Figure 1 As shown, a face-lip shape coordination generation method based on geometric constraints is described, the method comprising:

[0027] Step 1: Use a 3D deformation model to initialize the 3D shape of the face image to be processed, obtain the mouth opening and closing spatial shape matrix, and simultaneously generate a facial vertex index table that corresponds one-to-one with the mouth opening and closing spatial shape matrix.

[0028] Step 2: Based on the facial vertex index table, extract the lip contour, gingival boundary and alveolar pose, construct the lip-gingival joint boundary constraint set, and map the lip-gingival joint boundary constraint set to a unified coordinate system to form a geometric prior constraint tensor;

[0029] Step 3: Perform a multimodal topological concatenation invertible hybrid transformation on the geometric prior constraint tensor to obtain a high-dimensional tooth texture candidate tensor and a constraint consistency confidence spectrum;

[0030] Step 4: Using high-dimensional tooth texture candidate tensors and constraint consistency confidence spectra as input, the matrix is ​​mapped according to the mouth opening and closing spatial morphology matrix, and anti-aliasing, lighting compensation and gamma correction are completed in the unified rendering pipeline to finally obtain the mouth shape coordinated composite frame sequence.

[0031] This method treats the oral cavity, a highly variable local region constrained by anatomical structure, as being driven by two coupled continuous curves: "spatial morphology" and "texture state." The former defines the pose evolution of the dentition and labial margins over time in three-dimensional space, while the latter describes the optical distribution of enamel, gingiva, and even oral soft tissues in two-dimensional space. To simultaneously recover these two curves from any input facial image sequence and maintain their strict matching in each frame, the algorithm first performs morphological initialization using a three-dimensional deformation model. This step is not merely fitting two-dimensional pixels to a three-dimensional mesh, but rather synchronously mapping shape parameters, expression parameters, and pose parameters to a unified anatomical reference surface, further abstracted into a mouth opening and closing spatial morphology matrix, used to record time-varying opening and closing angles at the oral cavity entrance, the distance between the upper and lower lips, and boundary curvature, among other geometric measurements. Simultaneously, a synchronously generated facial vertex index table creates immutable global numbers for all vertices, ensuring that differences between subsequent frames are reflected only in coordinate values ​​rather than topological relationships, fundamentally guaranteeing temporal consistency.

[0032] Next, the algorithm extracts three types of geometric entities—lip contour, gingival boundary, and alveolar pose—based on this index table. These entities are located at the interfaces of soft tissue-air, mucosa-alveolar bone, and alveolar bone-oral cavity overall pose, respectively, and physically together constitute the closed boundary of the oral cavity entrance. After mapping these three types of entities to a unified coordinate system with the alveolar center point as the origin, a geometric prior constraint tensor with consistent scale and orientation can be obtained across different individuals and different expressions. The essence of this tensor is a set of high-dimensional vector fields, whose components correspond to the vertex 3D coordinates, normal vector, curvature scalar, and boundary type label, respectively. Since these components come from the common geometric features of soft and hard tissues, their changes on the time axis are much more stable than pixel-level textures, thus providing reliable morphological anchors for subsequent texture generation. The third step introduces a multimodal topological cascaded reversible hybrid transformation. The purpose is not simply compression or dimensionality reduction, but rather to stretch the inherent geodesic distances of each mode in Euclidean space into topological energy by repeatedly folding back and forth between the hypersphere and hyperbolic space. Then, multi-scale focused search is achieved by using the exponential decay strategy of the cascaded multiple access sampling function. Maintaining reversibility throughout the process means that no information is lost at any stage; the algorithm simply rearranges information in different geometric domains to expose potential cross-modal alignment relationships. When the global topological energy index reaches stability or the iteration limit, the searcher selects the path with the lowest cost and optimal geometric consistency on the energy graph, and maps it back to Euclidean space to generate a high-dimensional tooth texture candidate tensor. Simultaneously, by calculating the Euclidean consistency distance between the candidate instances and the geometric prior constraint tensor in the three dimensions of vertex 3D coordinates, normal vector, and curvature scalar, and normalizing it using a soft maximum function, the constraint consistency confidence spectrum is obtained.

[0033] This confidence spectrum can be seen as a statistical evaluation of the matching degree between candidate textures and real geometry. The higher the value, the better the candidate texture conforms to the morphological constraints at key details such as the crown edge and gingival sulcus. The final step is to fuse the high-dimensional tooth texture candidate tensor and the constraint consistency confidence spectrum one by one according to the sequence number to obtain a single fused tooth texture representation vector, which is then flattened into a two-dimensional texture mesh according to a unified UV layout. In order to eliminate sampling holes and smooth the boundaries, bidirectional Laplacian interpolation is used to fill missing pixels, and then Gaussian weighted kernel is used to correct the peripheral color level. Subsequently, local histogram equalization is used to enhance the contrast of enamel highlights and cervical shadows. At this point, a one-to-one correspondence has been established between the fused texture and the mouth opening and closing spatial morphological matrix. The system traverses each time frame in the matrix and projects the texture coordinates back to the oral vertex subset of the three-dimensional deformation model to obtain a frame-level tooth texture vertex texture dataset. The unified rendering pipeline performs anti-aliasing on each frame during the output stage to avoid pixelation at the junction of the lips and gums. A lighting compensation algorithm balances the energy of the lighting-sensitive oral cavity walls, ensuring consistent brightness between frames. Gamma correction further eliminates differences in color gradation response across different display devices. The resulting lip-synced composite frame sequence not only maintains a smooth visual transition between the opening and closing of the upper and lower lips, but also achieves high consistency with real images in terms of dentition occlusion, gum exposure, and even alveolar parallax details, realizing dynamic restoration driven by dual constraints of oral geometry and texture. This progressive design, from form to texture to optical rendering, enables the algorithm to output stable and reliable results even when dealing with rapid speech, exaggerated expressions, or drastic lighting changes, establishing a complete and rigorous technical chain for applications such as real-time expression-driven rendering, face video retargeting, and virtual digital human live streaming.

[0034] Furthermore, in the geometrically constrained face-lip shape coordination generation method, the core principle of step 1 can be understood as transforming the discrete facial observations in the two-dimensional pixel domain, which are susceptible to interference from lighting, pose, and occlusion, into a dynamic morphological description with complete topological structure and physical interpretability in three-dimensional Euclidean space, thereby providing deterministic boundary conditions for subsequent geometric constraint calculations. The system first requires a two-dimensional facial feature point locator covering key anatomical landmarks such as eyebrows, eyes, nose, and lips. Internally, this locator outputs precise pixel coordinates in the image coordinate system through a convolutional neural network. These coordinates themselves only contain local gray-level gradient information but lack depth and true metric attributes, thus failing to directly support the inference of oral cavity volume or lip curvature. To compensate for the lack of depth dimension, the algorithm introduces a three-dimensional deformation model, a linear subspace model whose shape parameters statistically encode facial differences in terms of race, gender, and age; expression parameters control the non-rigid deformation of soft tissue in the surface coordinate system; and pose parameters ensure that the global position and orientation of the head can freely change in the world coordinate system. To align the external 2D observations with the internal 3D substrate, the system predefines and initializes the corresponding mapping relationship. This mapping is generated based on a common training set using a joint method of average surface shape and thin-plate splines, allowing each 2D facial feature point to have a definite candidate vertex on the 3D deformation model. During the fitting phase, the optimizer minimizes projection error and regularization terms while maintaining the topological invariance of the 3D deformation model. This ensures that shape parameters do not overfit local noise, expression parameters do not excessively bend surfaces, and pose parameters guarantee depth consistency in monocular reconstruction. The projection function uses perspective geometry to project the 3D vertices back to the 2D image plane and uses partial derivatives to achieve gradient backpropagation, allowing shape, expression, and pose parameters to converge simultaneously in a single end-to-end backpropagation. This joint estimation based on differentiable projection ensures a natural separation of the coupling relationships between parameters: shape parameters explain long-term invariant geometric features across frames, expression parameters respond to snapshot-level muscle movements, and pose parameters only handle global rigid body transformations. The recovered 3D deformation model maintains a vertex density consistent with the input image resolution, enabling the model to accurately fit the original contours while avoiding the accumulation of resampling errors during subsequent texture mapping.

[0035] To measure the dynamic state of the oral cavity region, the algorithm extracts a subset of vertices related to the upper and lower boundaries of the lips and the opening and closing areas of the mouth from the vertex set of the 3D deformation model. The selection criteria are not simply filtered by spatial location, but based on the curvature extrema pairs corresponding to the lip margins, ensuring a topologically consistent vertex sequence across different facets. For each frame, the system records the relative changes in the 3D coordinates of these vertices and combines continuous measures such as the shortest Euclidean distance between the upper and lower lips, the angle between local normal vectors, and the cross-sectional area of ​​the oral cavity into a continuous descriptive function of the degree of mouth closure. Discretely sampling this function across all frames and stacking them in temporal order generates a spatial morphological matrix of mouth opening and closing. The matrix row indices map to frame timestamps, and the column indices correspond to specific geometric measures, thus comprehensively describing the evolutionary trajectory of the oral cavity at each instant during speech production, emotional expression, or spontaneous movements. Since the matrix elements are directly derived from 3D coordinates and contain no texture or color information, they are naturally immune to external factors such as lighting, occlusion, and camera exposure, providing a high-confidence reference for the subsequent construction of geometric prior constraint tensors.

[0036] The design principle of the facial vertex index table is to ensure consistent addressing capabilities across frames, stages, and even modules. During model initialization, the system assigns a globally unique number to all vertices in the original order of the vertex array, and archives the offset in memory, the face index to which each vertex belongs, and the linked list of adjacent vertices for each vertex. The static nature of this index table means that as long as the topology of the 3D deformation model remains unchanged throughout its lifecycle, any algorithmic step can use the same number to access the same anatomical location without repeated construction or spatial search inference. Furthermore, the topological connectivity recorded in the index table can be directly invoked later when calculating local tangent planes, discrete curvatures, or performing cycle traversal, greatly reducing time complexity. More importantly, the index table provides constant-time retrieval capability for subsets of vertices at the upper and lower boundaries of the lips. When the mouth opening and closing spatial morphology matrix is ​​updated at runtime, the corresponding rows and columns can be synchronously refreshed simply by using the vertex number, without needing to perform hashing or depth-first search on the entire mesh. Through this design, the system transforms the problem of unordered continuous frame deformation into the problem of evaluating an ordered vertex index sequence, fundamentally ensuring the robustness and interpretability of mouth shape coordination generation in the temporal dimension.

[0037] Furthermore, the fixed-topology face triangular mesh plays the role of the "geometric skeleton" in the geometrically constrained face and lip shape coordination generation method. Its vertices, edges, and faces remain intact throughout the algorithm's lifecycle and do not change with input samples, thus ensuring that all geometric processing is based on the same topological benchmark. This mesh consists of 53,215 facial vertices and 105,054 edges. The number of faces corresponds one-to-one with the number of vertices and edges via Eulerian relations, and each face is a triangle. This ensures that it can be directly accepted by the hardware rasterization pipeline during transmission over the GPU data channel without dynamic retriangulation. Vertex numbers are written to the model metadata file once during data preprocessing and reside in memory using sequential integer encoding. This facilitates fast addressing by downstream modules using offsets and makes the cross-frame vertex index mapping an identity relationship, completely eliminating potential topological ambiguities between multiple frames. To ensure that the fixed topology balances macroscopic integrity and local plasticity, the mesh covers the entire visible area from the hairline to the mandible, and a deformable vertex strip is reserved at the inner edge of the lips. Specifically, a ring of slightly denser vertices is inserted at the junction of the vermilion border and the oral cavity. These vertices are marked as "deformation degree of freedom enhancement" during the mesh generation stage, and their first and second adjacency weights are assigned lower Laplacian smoothing coefficients in the constraint matrix. This allows for a higher degree of deformability compared to other areas of the face during local expression-driven or mouth-opening / closing movements. Since the curvature change of the inner edge of the lips directly determines the accuracy of the mouth-opening / closing spatial morphology matrix, insufficient vertex density or overly rigid topological connections can lead to overstretching of the lip edge surface or the appearance of folding artifacts. This deformable vertex strip, however, can locally enhance geometric degrees of freedom while maintaining global topological invariance, resulting in a smoother and more refined change in the closure description function over continuous time.

[0038] To avoid redundancy in overall storage caused by high-density vertices in certain areas, the mesh uses a multi-level subdivision template in other regions: the cheeks and forehead are set to medium density, and the transition area between the hairline and the mandible is set to low density. A four-level detail level marker is also recorded in the mesh vertex attribute table. During rendering, different resolution vertex sets can be dynamically selected based on the viewing distance or real-world requirements, thus achieving a balance between performance and accuracy. At the data structure implementation level, all vertex coordinates, normal vectors, and their topological adjacency indices are laid out using structure arrays. The vertex index table is stored independently as a read-only constant in the read-only buffer of the video memory, ensuring no write conflicts occur during multi-threaded parallel access. When constructing the mouth opening and closing spatial morphology matrix, the algorithm only needs to query a subset of vertices related to the upper and lower boundaries of the lips and the opening and closing areas of the oral cavity. The numbering segments of these vertices are already written into the metadata list during model packaging, so the query operation can be completed in constant time. This maintains semantic consistency due to the unchanged global numbering while ensuring that the sampling resolution of local deformations is sufficient to capture extremely subtle lip dynamics.

[0039] Furthermore, in the geometrically constrained face-lip shape coordination generation method, the core task of step 2 is to transform the discrete and disordered set of lip and gingival vertices into a lip contour polygon chain that can be continuously indexed by arc length and a purified gingival boundary sequence. This process relies on the "numbering-topology-geometry" triple mapping provided by the facial vertex index table: the numbering determines the retrieval range, the topology defines the adjacency relationship, and the geometric attributes distinguish between soft and hard tissues in the local region. First, the system directly locks two lip subsets by retrieving the predefined range of upper lip outer edge vertex numbers and lower lip outer edge vertex numbers. The predefined numbers are based on the topological loop during model generation and are not affected by individual shape parameters and expression parameters, thus ensuring consistency across users and frames; the essence of this numbered retrieval is to reduce the complex spatial search to an integer interval query, significantly shortening the initialization time. Subsequently, the algorithm performs a loop-by-loop traversal of the upper lip outer edge vertex subset and the lower lip outer edge vertex subset based on the topological connection relationship of the 3D deformation model. Unlike simple adjacency list crawlers, cycle-by-cycle traversal requires maintaining consistent curvature signs and edge directions during the traversal process to obtain a monotonically increasing vertex order along the lip contour at each step. After traversal, the system repairs closed loops at the junctions of two ordered sequences using topological pruning, ultimately outputting a non-self-intersecting, end-to-end connected, and monotonic lip contour polygon chain. The significance of this chain lies in discretizing the original two-dimensional curve into a one-dimensional parameter domain. Subsequent texture mapping and curvature analysis only require linear interpolation within the chain index space, eliminating the need to return to the three-dimensional mesh for costly path searching.

[0040] The extraction approach for the gingival boundary is similar to that for the lip margin, but the challenge lies in the absence of a clear geometric line dividing the gingiva and hard palate. Instead, the boundary is formed by the direction of the normal vector and a gradual curvature transition. The algorithm first searches the facial vertex index table for the range of gingival inner edge vertex numbers, resulting in a subset of gingival inner edge vertices containing redundant hard palate vertices. To eliminate redundancy, the system calculates the normal vector direction of each vertex and performs an angle test with the normal reference of the alveolar pose reference system. Hard palate vertices face the oral cavity dome, and their normal vectors mostly point towards the cranial cavity, while gingival vertices face the outer wall of the alveolar bone, and their normal vectors tend towards the front of the oral cavity. By setting a threshold to distinguish between the two types of directions, the hard palate region can be initially filtered out. Next, the local curvature is estimated on the remaining vertices. The curvature exhibits a clear inflection point in the hard palate-gingival transition region. The system uses an increasing curvature window to find the inflection point and discards the hard palate vertices after it, ultimately retaining only the vertices located in the gingival-alveolar junction region and outputting a purified gingival boundary sequence. The purified sequence ensures surface continuity and aligns with the polar coordinate parameterization framework of the alveolar center point, facilitating subsequent rigid coordinate transformation and tensor storage. Through this multi-level screening process of "number retrieval - topology traversal - normal filtering - curvature elimination", step 2 compresses the key soft and hard tissue boundaries into two highly controllable one-dimensional ordered vertex sequences without destroying the global topology. This provides precise spatial anchor points for the geometric prior constraint tensor and lays the morphological benchmark for subsequent tooth texture candidate search and constraint consistency evaluation.

[0041] Furthermore, in the geometric prior construction process of the face-mouth shape coordination generation method based on geometric constraints, the purified gingival boundary sequence serves as a geometric bridge connecting soft and hard tissues. The system first takes the arithmetic mean of the three-dimensional coordinates of all vertices in the sequence in the world coordinate system to obtain a unique centroid, which is named the alveolar center point. The centroid is chosen as the origin because it simultaneously minimizes the squared distance to all gingival vertices, thus making the translation components in subsequent coordinate transformations statistically the most stable. Next, the algorithm performs principal axis analysis on the purified gingival boundary sequence, extracting the feature vector with the largest variation in the long direction from the boundary, denoted as the alveolar long axis direction. Then, it estimates the normal vector on the tangent plane of the locally fitted purified gingival boundary sequence, and obtains the alveolar vertical direction after orthogonalization with the alveolar long axis direction. Combining the alveolar center point, alveolar long axis direction, and alveolar vertical direction according to the right-hand rule forms the alveolar pose reference system. This local coordinate system not only conforms to the orientation of the individual alveolar bone in a three-dimensional geometric sense but is also decoupled from the global head pose parameters, thus providing a unified and stable pose reference across multiple frames or samples. Subsequently, the algorithm maps the lip contour polygon chain and the purified gingival boundary sequence together onto the alveolar pose reference frame, so that their three-dimensional coordinates, normal vectors, and curvature scalars all obtain normalized expressions independent of individual head pose. At this point, the relative positions between the three-dimensional point sets are no longer affected by the camera's perspective, but purely depend on the true geometric relationship between the lip margin and the gingiva.

[0042] To facilitate subsequent reversible topological folding operations in high-dimensional feature space, the system encapsulates the lip contour polygon chain, the purified gingival boundary sequence, and the alveolar pose reference frame into a lip-gingival joint boundary constraint set. This encapsulation is not merely a simple combination at the structural level; it also establishes explicit coupling across boundary types at the data semantic level: the algorithm simultaneously records the original 3D coordinates, normal vector, curvature scalar, and boundary type label for each vertex in the joint boundary constraint set. The original 3D coordinates preserve the true spatial position, the normal vector characterizes the local surface orientation, the curvature scalar measures the degree of surface bending, and the boundary type label uses integers to semantically group vertices, thus ensuring that the multimodal topological cascade reversible hybrid transformation can be dually aligned according to the attribute dimension and the boundary semantic dimension during subsequent processing. The same vertex has the same number in different time frames, but its coordinates, normal vector, and curvature change with facial expression evolution. This design integrates the joint boundary constraints into a composite data structure that includes both static numbers and dynamic geometric signals. Because the alveolar pose reference frame provides a locally orthogonal basis, the components of the normal vector under this basis directly reflect the actual tilt angle of the vertex relative to the alveolar bone, while the curvature scalar can be obtained through ellipsoidal fitting or the discrete Laplacian operator. The numerical distribution is limited to the interval between zero and one, so as to perform dimensionless comparisons between different vertices. Boundary type labeling strictly distinguishes between the lip contour and the gingival boundary, allowing subsequent soft maximum consistency mapping to adaptively assign weights according to the boundary category, avoiding the lip margin details being overwhelmed by the high curvature peaks of the gingiva.

[0043] Through the above steps, the algorithm not only transforms two originally independent boundaries into cooperative entities living together in the alveolar pose reference frame, but also prepares complete triple feature vectors of shape, orientation, and curvature for each vertex, as well as clear category labels. The joint boundary constraint set thus becomes the direct upstream input to the geometric prior constraint tensor, possessing complete geometric context information before entering the multimodal topological cascade invertible hybrid transformation. The entire process self-consistently solves three key problems: local coordinate unification, boundary combination, and attribute recording, providing reversible, interpretable, and high-quality morphological anchors that conform to individual anatomical characteristics for the subsequent generation of high-dimensional tooth texture candidate tensors.

[0044] Furthermore, in step 2, with the alveolar center point as the origin, the three-axis vector of the alveolar pose reference system is used as the basis vector to define the unified oral coordinate system; a rigid coordinate transformation is performed on all vertices within the lip-gingival joint boundary constraint set to map the vertex 3D coordinates from the global coordinate system of the 3D deformation model to the unified oral coordinate system; the mapped vertex 3D coordinates, normal vectors, curvature scalars, and their respective boundary type labels are tensorized and stored to form a geometric prior constraint tensor; a normalization operation is applied to the geometric prior constraint tensor so that the 3D coordinates are distributed in the interval -1 to 1, the normal vectors are distributed in the interval -1 to 1, and the curvature scalars are distributed in the interval 0 to 1.

[0045] For numerical stability and cross-sample comparability, the system applies normalization to the geometric prior constraint tensor. First, the extreme values ​​of the mapped 3D coordinates across the entire joint boundary constraint set are statistically analyzed. Using the maximum absolute value as the denominator, all coordinates are compressed to the range of -1 to 1. The normal vector, already a unit vector, is also linearly stretched on each component to ensure consistency with the coordinate distribution, so that its value falls within the range of -1 to 1. The curvature scalar is typically non-negative, and its numerical range can vary by an order of magnitude depending on individual surface shapes. The algorithm performs inverse scaling on its global maximum value, mapping the curvature distribution to the 0 to 1 range, while preserving the contribution of high curvature peaks to the sensitivity of subsequent features. Boundary type labels are not normalized because this field is a discrete semantic label and should not participate in continuous space scaling. After normalization, the tensor becomes a geometrically prior constraint tensor with a controlled numerical domain, clear attribute dimensions, and a consistent coordinate system. It maintains static topological consistency between vertices at the storage level and eliminates scale and pose differences at the numerical level, providing an input format with minimal redundancy and maximum information gain for subsequent multimodal topological cascade reversible hybrid transformations.

[0046] Furthermore, in the high-dimensional feature modeling stage of the geometrically constrained face-lip shape coordination generation method, the geometric prior constraint tensor, after entering step 3, is first separated according to four attribute dimensions: vertex 3D coordinates, normal vectors, curvature scalars, and boundary type labels. The purpose is to place fundamentally different geometric information in statistically homogeneous subspaces to reduce coupling interference during subsequent transformations. The 3D coordinate branch directly describes the spatial distribution of vertices in the unified oral cavity coordinate system and belongs to the dimension of significant Euclidean distance; the normal vector branch reflects the orientation of the local tangent plane, with a constant vector norm but complex directional distribution; the curvature scalar branch reflects second-order morphological changes and is sensitive to local concavity and convexity; the boundary type label branch is a discrete semantic label that only distinguishes between the lip margin and gingiva at the logical level. In order to process these four modalities within the same reversible geometric framework, the algorithm embeds the 3D coordinate subtensor into a hyperspherical space with a radius of one through stereo projection, thereby transitioning all points from linear Euclidean space to a Riemannian manifold with constant curvature. This can compress the scale differences between different coordinate dimensions, making distance measures into angle measures on the sphere.

[0047] To maintain tensor dimensionality consistency, the system linearly expands the normal vector, curvature scalar, and boundary type markers to the same vector length as the hyperspherical embedding tensor. The normal vector branch is directly copied and regularized, the curvature scalar branch is interpolated to fill continuous values ​​in the expanded dimension, and the boundary type markers are mapped to the same dimension using one-hot encoding. These four equal-length single-modal tensors are then fed into a topological folding kernel function. This function performs a reversible toroidal mapping on the hyperspherical embedding tensor, compressing the radial amplitude to the hyperbolic amplitude domain. The constant curvature transition to the negative curvature space significantly increases resolution in highly convex or concave regions, while recording absolute phase information at each folding point to ensure reversibility. Next, a cascaded multiple access sampling function is called to perform multiple hop sampling with exponentially decaying steps within the folding domain. This hop sampling strategy is equivalent to preserving high-frequency details locally while progressively downsampling regions far from the folding center, thus automatically balancing detail preservation with global context. The topological return kernel function and the cascaded multiple access sampling function are executed alternately to form a complete iteration. After each iteration, a global topological energy index is calculated. This index integrates the tensor's Bulman energy in the hyperbolic amplitude domain, return phase consistency, and cross-modal alignment error; the lower the value, the better the geometric consistency. If the energy of the current iteration decreases by less than one-thousandth compared to the previous iteration, it indicates that the marginal benefit of continuing the iteration is extremely low, and the algorithm immediately triggers early convergence to save computational resources. If the energy still decreases significantly, the iteration continues, running a maximum of one hundred times to avoid getting stuck in an infinite loop. If the convergence condition is not met before the iteration reaches the upper limit, the system selects the result of the iteration with the lowest energy as the provisional output tensor, because this result represents the optimal trade-off between geometric consistency and cross-modal balance in the searchable energy landscape. The entire process uses the reversible mapping from the hypersphere to hyperbolic space to bring the information of the four modes into the same topological domain. Then, through multi-scale sampling and energy monitoring, the most geometrically representative features are adaptively selected. This ensures that the search space of the subsequent high-dimensional tooth texture candidate tensor can fully encompass the local deformation of the labial margin and gingiva without being overwhelmed by high-dimensional noise, thus guaranteeing that the constraint consistency confidence spectrum can make a sensitive and stable response to the real geometric conditions.

[0048] Furthermore, in the geometrically constrained face-lip shape coordination generation method, the provisional output tensor is merely a set of candidate states found by the multimodal topological cascaded invertible hybrid transformation within a high-dimensional energy landscape, which has not yet undergone geometric consistency refinement. The location that truly represents the optimal coupling between tooth texture and lip-tooth geometry must be further filtered by a greedy minimum topological energy path searcher. This searcher first maps the provisional output tensor to a high-dimensional energy graph. Each node in the energy graph corresponds to a combination of values ​​for folded-back phase, hyperbolic amplitude, and cross-modal alignment error, while the edge weights are given by a global topological energy index. The global topological energy index aggregates the mutual information loss between the four single modes and the topological distortion within the tensor; therefore, a lower index indicates a higher degree of geometric matching. The searcher uses depth-first traversal rather than breadth-first traversal because the high-dimensional energy graph exhibits a clear hierarchical gradient structure, and connected branches along the energy descent direction tend to reach local troughs faster. During the traversal, the geometric consistency cost of the current path is calculated immediately each time the recursion deepens by one level. This cost is measured by the cumulative Euclidean difference between the provisional output tensor and the original geometric prior constraint tensor, representing the vertex's 3D coordinates, normal vector, and curvature scalar. The greedy strategy is implemented by immediately locking the path as the optimal topological path and terminating the traversal as soon as the geometric consistency cost of a path first falls below a preset threshold of 50%, thus avoiding exponential computational redundancy caused by continuing the search when the quality requirements are already met. This early stopping mechanism leverages the correlation between geometric consistency cost and energy metrics to achieve rapid convergence while ensuring that the optimal path maintains its low-energy characteristics.

[0049] After determining the optimal topological path, the system performs a reversible hybrid reverse mapping on the tensors along the path, strictly reversing the transformations of the cascaded multiple access sampling function and the topological return kernel function in the correct order. When reversing the cascaded multiple access sampling function, low-scale samples are inserted back to their corresponding high-scale positions according to the jump address log. Then, based on the return phase information, the torus mapping is reversed to re-unfold the hyperbolic amplitude back to the hyperspherical radial amplitude, ultimately returning to Euclidean space. This reversible step ensures that any optimal path obtained within the high-dimensional energy graph can be losslessly restored to a reconstructed modal tensor corresponding one-to-one with the original four modes. Subsequently, the algorithm reassembles the four reconstructed modal tensors according to the original modal order (vertex 3D coordinates, normal vectors, curvature scalars, boundary type markers) to obtain a 3D array of candidate tooth textures. The first 128 elements of the array are then stacked according to their original indices to form a high-dimensional candidate tooth texture tensor. This upper limit of 128 controls the complexity of subsequent distance calculations and ensures that the candidate set is sufficiently diverse to capture local extrema.

[0050] Next, the system processes each candidate instance in the high-dimensional tooth texture candidate tensor one by one, calculating the Euclidean consistency distance with the geometric prior constraint tensor in three attribute dimensions: vertex 3D coordinates, normal vector, and curvature scalar. A smaller distance indicates a closer fit between the candidate instance and the actual position of the lips and teeth in geometric space. However, to avoid sharp selection due to simple minimization and sacrifice of overall stability, the algorithm performs a soft maximization operation on the distance, mapping it to a confidence interval of zero to one. The soft maximization operation essentially involves exponentially weighting the negative distances of all candidate instances and then normalizing them to a probability distribution. Therefore, the confidence value of each single instance reflects both its own matching degree and implicitly implies its relative superiority or inferiority compared to other candidates. After all 128 single instance confidence values ​​have been calculated and arranged by element index, the system obtains a constraint consistency confidence spectrum of length 128. The curve shape of the confidence spectrum not only intuitively shows which sets of tooth textures best conform to geometric constraints, but also provides continuous weights for subsequent confidence-weighted fusion. This ensures that the final fusion result maintains a smooth transition to low-confidence textures when selecting high-confidence textures, preventing texture jumps due to the dominance of a single candidate. Through this continuous chain of greedy search, reversible inverse mapping, and soft maximum normalization, step three transforms the abstract energy minimum solution extracted from the multimodal return domain into a high-dimensional tooth texture candidate tensor and its reliability metric that simultaneously satisfies geometric consistency in physical space, orientation space, and curvature space. This lays the data and statistical foundation for subsequent steps to complete accurate mapping and optical correction in the unified rendering pipeline.

[0051] Furthermore, the system maps the high-dimensional tooth texture candidate tensor to the constraint consistency confidence spectrum one-to-one according to the candidate sequence number dimension, and regards the confidence value of each single instance in the confidence spectrum as the fusion weight of that candidate texture instance. Unlike the traditional hard selection of the best texture, a soft fusion strategy is adopted here: the weights are directly multiplied onto the high-dimensional tooth texture candidate tensor to obtain the confidence-weighted tooth texture fusion tensor. Weighted summation is performed along the candidate sequence number dimension, which is equivalent to performing probability density integration on the candidate space, and outputting a single fused tooth texture representation vector. Since the confidence is derived from the soft maximum mapping of geometric consistency distance, this approach ensures at the numerical level that the texture region most sensitive to geometric constraints receives a higher weight, while the contribution of candidates with large geometric errors is automatically attenuated, thereby avoiding texture jumps or local gaps caused by hard thresholding. Next, the system expands the fused tooth texture matrix into a one-dimensional vector column-wise, and then rearranges it into a two-dimensional texture mesh according to the unified UV layout table of the tooth region. The unified UV layout table is the result of pre-expanding tooth patches for a fixed topology 3D deformation model, ensuring that the positional correspondence between different frames or different individuals in the texture domain remains constant. Since the candidate fusion process may result in some pixels remaining null, the algorithm performs bidirectional Laplacian interpolation to fill the missing pixels within the 2D texture mesh. Bidirectional Laplacian interpolation takes into account second-order smoothness constraints in both the horizontal and vertical directions, eliminating holes and maintaining gradient continuity at boundaries. After interpolation, Gaussian weighted smoothing is applied to the mesh edge pixels to reduce high-frequency noise and avoid jagged artifacts in subsequent rendering. The resulting initial tooth texture map now has complete spatial coverage, but the tonal contrast is still limited by the average characteristics of the candidate textures. To improve visual quality, the system applies local histogram equalization to this map. This involves calculating and equalizing the grayscale distribution within each small region, then eliminating transition bands at region boundaries using bilinear interpolation. This enhances enamel highlights and neck shadow details, outputting an enhanced tooth texture map. Local equalization, compared to global processing, more accurately preserves the micro-texture of the tooth surface while avoiding overall exposure shift.

[0052] At this point, the system has obtained a high-quality texture map with an improved color dynamic range and UV layout consistent with the tooth region. To accurately match the dynamic mouth shape, time-dependent spatial pose information is needed. Therefore, the system obtains the mouth opening / closing spatial morphology matrix corresponding to the enhanced tooth texture map and, for each time frame in the matrix, locates a subset of oral region vertices in the facial vertex index table. Since the vertex numbers in the index table are constant throughout the model's lifecycle, and the corresponding set of triangular faces inside the oral cavity also remains consistent, the frame-level retrieval cost is very low. Subsequently, the algorithm maps the 2D texture coordinates of the enhanced tooth texture map to the subset of oral region vertices in the 3D deformable model according to the mouth opening / closing spatial morphology matrix of the current frame. The mapping process first expands the 3D vertex coordinates into polar coordinates in the unified oral coordinate system, then uses the index relationship in the UV layout table to obtain the corresponding texture coordinates, ultimately generating precise texture sampling points for each vertex, forming a frame-level tooth texture vertex map dataset. After receiving the vertex map dataset for each frame, the unified rendering pipeline sequentially performs anti-aliasing, lighting compensation, and gamma correction. The anti-aliasing stage employs a multi-sampling anti-aliasing algorithm to perform intra-pixel subsampling on the labial margin, gingival junction, and crown contour to reduce high-frequency jagged edges and ensure smooth edges. The lighting compensation stage utilizes ambient occlusion and radiance buffering to perform global light energy balancing on the oral cavity walls, preventing unnatural bright spots or shadows caused by changes in head posture. Finally, gamma correction performs color level transformation on the rendered result based on the standard gamma curve of the target display device, ensuring correct brightness and contrast across different hardware. As time frames advance, the mouth opening and closing spatial morphology matrix drives the continuous change in position and orientation of a subset of vertices in the oral cavity region, while the enhanced tooth texture map, projected via UV coordinates, remains closely attached to these vertices. The rendering pipeline dynamically adjusts lighting and hue in each frame, thereby outputting a sequence of coordinated lip-shape composite frames. The sequence visually presents a natural and coherent transition between pronunciation and facial expressions: the highlights on the tooth crowns move with changes in light, the details of the gingival margins remain seamless during mouth opening and closing, the occlusion relationship of the upper and lower lips is updated in real time, and all frames strictly conform to geometric prior constraints. Therefore, it can be directly used in scenarios sensitive to lip-shape consistency without post-processing, such as voice-driven virtual characters, lip-shape replacement in film and television post-production, or real-time facial driving for live-streaming virtual anchors. Step 4, through the progressive layering of confidence-weighted fusion, hole interpolation, local histogram equalization, and a unified rendering pipeline, successfully couples the high-dimensional tooth texture candidate tensor with time-varying geometric morphology, ensuring a high degree of consistency between lip movements and tooth details in both spatial and temporal dimensions. This provides a rigorous and high-fidelity conclusion to the entire geometrically constrained face-lip-shape coordination generation method.

[0053] The following example demonstrates the offline execution of the process on a workstation equipped with an Intel Core i9-12900K, an NVIDIA RTX 4090, and 64GB of RAM. The input is a 1920×1080 RGB video clip (approximately 2.5 seconds, 24fps) containing 60 frames, with the subject's face at a direct angle of less than 5°.

[0054] For each frame, a face detection and keypoint localization model (RetinaFace+PIPNet) is applied, outputting 68 facial feature points. Taking frame 12 as an example, the pixel coordinates of the outer corner of the left eye are (x, y) = (814, 472), the pixel coordinates of the tip of the nose are (946, 552), and the pixel coordinates of the center of the lower lip are (946, 706). A fixed topological triangular mesh with N vertices is used for the 3D deformation model. v =53215, number of edges N e =105054. Linear subspace dimension: shape parameter Expression parameters The attitude parameters are p = (θ, t), where θ = (α, β, γ) are the three Euler angles.

[0055] Initialize using frame 0 as the baseline. Minimize projection error. In the middle, Π represents perspective projection, and B... s B e These are the shape and expression base matrices, respectively. The mean vertex is θ. After 120 optimizations by LBFGS, the result converges to θ = (1.3°, -4.8°, 0.7°) and t = (-3.2, 4.6, 765.0) mm.

[0056] Oral cavity related vertex subset V m It contains 892 vertices. The distance between the upper and lower lips is calculated for each frame. d t Average curvature k of the lip margin t oral cavity cross-sectional area a t Combined into column vector o t =[d t ,k t ,a t ] T The mouth opening and closing spatial shape matrix obtained by stacking 60 frames. Example value o in frame 30 30 =[4.7,0.38,214.6]. Record the face vertex index table: vertices v0 to v 53214 Numbered according to the generator output order. Read the predefined numbering range of lip and gingiva: V up ={15230,…,15345},Vlow ={15346,…,15461},V gingiva ={29812,…,29955}; traversing the loops one by one yields an ordered sequence of lip contours with a length of 248. After removing 31 hard palate vertices from the inner edge of the gums, 113 vertices remain.

[0057] Calculate the centroid of the purified gingiva Principal component analysis yields the alveolar major axis direction as u1 = (0.92, 0.05, 0.38). Let u2 be the second eigenvector; its cross product with u1 gives u3 = u1 × u2. After orthogonal normalization of the three vectors, a rotation matrix R is formed. g = [u1,u2,u3]∈SO(3).

[0058] Perform a rigid transformation on all 361 vertices of the joint boundary constraint set. Will normal vector curvature k i and type tag t i ∈{0,1} serves as a four-way attribute, constituting the dimension. Divide the coordinate and normal components by m respectively xyz =7.5mm, resulting in a normalized interval [-1, 1]. The curvature normalization factor is m. k =1.6.

[0059] The four unimodal tensor branches are denoted as X, N, K, and L. Each row of coordinates x is projected onto the hypersphere: Linearly expand N, K, L to... Let these be denoted as N', K', L'. The return kernel function...

[0060] It applies the radial component ρ = arccos(u·s) in vector form. After the first-level backtracking, the energy decreases by ΔE0 = -0.014. The cascaded multiple access sampling function uses a step size sequence {1, 0.5, 0.25, 0.125}. After sampling, the energy decreases by ΔE1 = -0.009. Alternating backtracking and sampling is considered one iteration. After eight iterations, the global topological energy index E8 = 0.0261. E at the fourteenth iteration 14 =0.02417, If early convergence is triggered, a provisional tensor is output.

[0061] The number of nodes in the high-dimensional energy graph is M = 2205 (5 levels per foldback dimension, 5 in total for three dimensions). 3 The average branching factor b ≈ 4.1 for depth-first traversal. Path P * Achieving geometric consistency cost D(P) at the 37th depth unfolding *= 0.48 < 0.5. Reverse mapping cancels the wraparound and sampling, restoring the four-way tensor. Concatenates them into a candidate three-dimensional array of tooth texture dimensions. Calculate the Euclidean consensus distance, where candidate 7 d7 = 0.012, candidate 42 d7 = 0.012, and candidate 42 d7 = 0.012. 42 =0.037. Confidence level β = 120. For example, p7 = 0.086, p 42 =0.013. The confidence spectrum is obtained.

[0062] Weighted by serial number Transform Z into a column vector as follows Rearranged to a 64×64 UV mesh. Missing pixel rate 4.7%. Bidirectional Laplacian interpolation converged after 200 iterations, with a boundary Gaussian kernel σ = 1.3. Local histogram equalization window 8×8, pixel intensity range [0, 255]. Enhanced the average brightness of the crown highlights from 196 to 221. For each frame of oral vertex subset |V m | = 892 for texture sampling. The texture coordinates (u, v) = (0.37, 0.62) in frame 45 correspond to the 3D vertex v. 15034 The rendering stage uses MSAA4×, ambient occlusion factor 0.38, and gamma index 2.2. It generates a 60-frame lip-sync composition sequence with an average frame time of 13.4ms, and allows for real-time preview.

[0063] Figure 2 This diagram illustrates the overall process of generating a 3D deformation model and a mouth opening / closing spatial morphology matrix in the geometrically constrained face and mouth shape coordination generation method of this invention. The diagram details the specific implementation process of step 1, including the complete technical path from inputting the face image to finally generating the mouth opening / closing spatial morphology matrix and the facial vertex index table. Figure 2 As shown on the left, the present invention first acquires a face image to be processed, which is a static image or a sequence of dynamic frame images of a face containing at least one frontal or near-frontal view. In this example, the face image to be processed exhibits standard facial features, including clearly identifiable key facial regions such as eyebrows, eyes, nose, and lips. By performing face detection and facial landmark localization operations on the input image, the system can accurately extract the coordinates of two-dimensional facial feature points, including eyebrows, eyes, nose, and lips, thereby forming a two-dimensional facial feature point set. Figure 2The first image in the middle illustrates the construction process of the 3D deformation model. Based on a set of 2D facial feature points, the system constructs an initial mapping relationship that matches the 3D deformation model and then calls the 3D deformation model to perform 3D fitting on the face image to be processed. As shown in the figure, this 3D deformation model uses a fixed-topology facial triangular mesh, containing 53,215 facial vertices and 105,054 edges, and each vertex retains its number throughout the lifecycle of the 3D deformation model. The fixed topology covers the entire visible area from the hairline to the mandible, and deformable vertices are reserved at the inner edge of the lips to ensure the accuracy and consistency of the subsequent mouth opening and closing spatial morphology matrix. In the figure, black dots represent vertices distributed in various areas of the face, each vertex corresponding to a 3D coordinate value. In particular, the vertices in the lip area are more densely distributed to ensure accurate capture of mouth shape changes. Figure 2 The second image in the middle illustrates the generation process of the mouth opening and closing spatial morphology matrix. In the 3D deformation model, the system extracts the corresponding 3D coordinate changes of a subset of vertices related to the upper and lower boundaries of the lips and the opening and closing areas of the oral cavity in each frame, establishing a continuous descriptive function for the degree of lip closure. As shown in the matrix mesh structure, the black rectangular elements represent the spatial configuration data of the oral cavity at different time points or facial expressions. This mouth opening and closing spatial morphology matrix is ​​used to represent the spatial configuration of the oral cavity at different time points or facial expressions, providing a geometric basis for subsequent mouth shape coordination generation. Figure 1 The right side shows the generated facial vertex index table. The system numbers all vertices in the 3D deformable model according to their original order during model initialization, records their spatial coordinate indices and topological connectivity, forming a facial vertex index table. As shown in the figure, this index table stores the 3D coordinate information of each vertex in a structured manner, including the complete coordinate data (x1, y1, z1) to (xN, yN, zN) of vertices 1 to N, providing a precise vertex localization basis for subsequent lip contour extraction and gingival boundary recognition.

[0064] Figure 3The diagram illustrates the detailed processing steps for extracting the lip contour and gingival boundary. In the lip contour extraction stage, the system first retrieves the predefined range of vertex numbers for the outer edge of the upper lip and the outer edge of the lower lip from the facial vertex index table, obtaining subsets of the outer edges of the upper and lower lip vertices respectively. The upper lip outer edge vertex subset shown in the diagram contains key vertices distributed along the upper edge of the lip, while the lower lip outer edge vertex subset corresponds to the corresponding vertex positions along the lower edge of the lip. Based on the topological connectivity of the 3D deformation model, a cycle-by-cycle traversal operation is performed on these two vertex subsets to obtain a monotonically increasing ordered sequence of vertices along the lip contour direction, ultimately forming a complete lip contour polygon chain. This polygon chain is represented by dashed lines, accurately describing the geometric shape features of the lip boundary. In the gingival boundary extraction process, the system retrieves the range of vertex numbers for the inner edge of the gingiva from the facial vertex index table, obtaining a subset of the inner edge of the gingiva vertices. The key technical step lies in redundant vertex removal based on the vertex's normal vector direction and local curvature threshold. The redundant vertices marked with dashed circles in the figure represent irrelevant vertices connected to the hard palate. These vertices are effectively removed during the processing, and only the valid vertices located in the gingival-alveolar junction area are retained, thereby generating a purified gingival boundary sequence.

[0065] like Figure 4 As shown, this invention first separates the geometric prior constraint tensor according to four attribute dimensions: vertex 3D coordinates, normal vectors, curvature scalars, and boundary type labels, resulting in four unimodal tensor branches. For the 3D coordinates in each unimodal tensor branch, the system uses stereo projection to embed Euclidean space points into a hyperspherical space with a radius of one, forming a hyperspherical embedding tensor. As shown in the figure, multiple projection points are uniformly distributed on the hypersphere with a radius of one. These points establish a correspondence with vertices in the original Euclidean space through projection lines originating from the center of the sphere. This hyperspherical embedding transformation ensures the topological consistency of the 3D coordinate data in high-dimensional space, laying the geometric foundation for subsequent complex transformations. Simultaneously, linear dimensionality expansion is performed on the normal vectors, curvature scalars, and boundary type labels to make their dimensions consistent with the hyperspherical embedding tensor.

[0066] Figure 5 The left-middle figure illustrates the transformation process of the topological inversion kernel function. This function compresses the radial amplitude of the tensor on the hypersphere into a hyperbolic amplitude through a reversible toroidal mapping, preserving absolute phase information at each inversion point. As shown in the toroidal structure, the outer ring represents the original radial distribution, and the inner ring represents the compressed hyperbolic distribution. The inversion points are connected by phase-preserving lines, ensuring that key geometric information is not lost during the transformation. This transformation process achieves a reversible mapping from the radial coordinate system to the hyperbolic coordinate system, effectively reducing the dimensionality of the data without losing topological features. Figure 5The right-middle figure illustrates the processing mechanism of the cascaded multiple access sampling function. This function performs multiple skip sampling with exponentially decaying step sizes within the loopback domain to obtain cross-scale feature representations. As shown in the figure, the size of the sampling points exhibits an exponentially decaying distribution, gradually decreasing from the initial step size to a tiny step size, and the connecting lines represent the skip sampling path. This sampling strategy can capture the feature changes of the tensor at different scales, ensuring that both macroscopic topological structure and local detailed features are preserved. The topological loopback kernel function and the cascaded multiple access sampling function alternate to form a complete iteration, and the optimal transformation of the tensor is achieved through multiple iterations of optimization.

[0067] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. A method for face mouth shape coordination generation based on geometric constraints, characterized in that, The method comprises: Step 1: three-dimensional morphological initialization is performed on a to-be-processed human face image by using a three-dimensional morphing model, a mouth opening and closing spatial morphological matrix is obtained, and a face vertex index table corresponding to the mouth opening and closing spatial morphological matrix is synchronously generated; Step 2: based on the face vertex index table, a lip contour, a gingival boundary and a dental arch pose are extracted, a lip-gingival joint boundary constraint set is constructed, and the lip- gingival joint boundary constraint set is mapped to a unified coordinate system to form a geometric prior constraint tensor; Step 3: the geometric prior constraint tensor is separated according to four attribute dimensions of vertex three-dimensional coordinates, normal vectors, curvature scalars and boundary type markers to obtain four single-mode tensor branches; for the three-dimensional coordinates in each single-mode tensor branch, a stereographic projection method is used to embed the Euclidean space points into a hypersphere space with a radius of one to form a hypersphere embedding tensor; the normal vectors, the curvature scalars and the boundary type markers are subjected to linear dimension expansion so that the dimensions are consistent with the hypersphere embedding tensor; a topological folding kernel function is called in batches on the hypersphere embedding tensor, the function compresses the radial amplitude of the tensor on the hypersphere into hyperbolic amplitude through a reversible torus mapping, and absolute phase information is retained at each folding point; a cascaded multi-address sampling function is called to perform multiple address sampling with exponentially decaying steps in the folding domain to obtain a cross-scale feature representation; the topological folding kernel function and the cascaded multi-address sampling function alternately constitute a complete iteration; after each iteration, a global topological energy index is calculated, if the falling amplitude of the index compared with the last iteration is less than one thousandth, early convergence is triggered; otherwise, the next iteration is continued; if the number of iterations reaches one hundred and still does not converge, the result of the iteration with the lowest energy is taken as a tentative output tensor; the tentative output tensor is input into a greedy minimum topological energy path searcher; the searcher takes the global topological energy index as a cost function, performs a depth-first traversal on the high-dimensional energy graph of the tensor, and evaluates the geometric consistency cost of each candidate path in real time; when the geometric consistency cost is first lower than 50% of the preset threshold, the current path is marked as the optimal topological path and the search is terminated; reversible hybrid reverse mapping is performed on the tensor on the optimal topological path to sequentially undo the transformations of the cascaded multi-address sampling function and the topological folding kernel function, and a reconstructed modal tensor back to the Euclidean space is obtained; the four reconstructed modal tensors are reassembled in the original modal order to obtain a tooth texture candidate three-dimensional array; the first 128 elements of the tooth texture candidate three-dimensional array are stacked in the original order to form a high-dimensional tooth texture candidate tensor; for each candidate instance in the high-dimensional tooth texture candidate tensor, the Euclidean consistency distance of the vertex three-dimensional coordinates, the normal vectors and the curvature scalars of the candidate instance and the geometric prior constraint tensor is calculated; the Euclidean consistency distance is mapped to a confidence degree in the interval of 0 to 1 by using a soft maximum operation to obtain a corresponding single-instance confidence value; the 128 single-instance confidence values are sequentially arranged to form a constraint consistency confidence spectrum with a length of 128; Step 4: using the high-dimensional tooth texture candidate tensor and the constraint consistency confidence spectrum as inputs, mapping according to the mouth opening and closing space shape matrix, and completing anti-aliasing, illumination compensation and gamma correction on a unified rendering pipeline to finally obtain a sequence of lip shape coordinated synthesized frames.

2. The method for generating face mouth shape based on geometric constraints according to claim 1, wherein, Step 1 specifically comprises: obtaining a to-be-processed face image, which is a face static image or dynamic frame image sequence containing at least one face in a front or near-front view; performing face detection and facial key point positioning operations on the to-be-processed face image to extract two-dimensional facial feature point coordinates including eyebrows, eyes, nose wings, and lips to form a two-dimensional facial feature point set; based on the two-dimensional facial feature point set, an initial corresponding mapping relationship matching a three-dimensional morphable model is constructed, and the three-dimensional morphable model is called to perform three-dimensional fitting on the to-be-processed face image to output a three-dimensional shape parameter set including shape parameters, expression parameters, and posture parameters; according to the three-dimensional shape parameter set, a three-dimensional morphable model with a resolution consistent with the to-be-processed face image is restored; the three-dimensional morphable model is composed of facial vertices, each of which corresponds to a three-dimensional coordinate value; in the three-dimensional morphable model, for a vertex subset related to the upper and lower boundaries of the lips and the opening and closing area of the oral cavity, the three-dimensional coordinate changes thereof in each frame image are extracted to establish a continuous description function about the degree of lip closure to generate a mouth opening and closing space shape matrix; the mouth opening and closing space shape matrix is used to represent the spatial configuration of the oral cavity at different time points or expression states; all the vertices in the three-dimensional morphable model are numbered in their original order during the initialization stage of the model, and their spatial coordinate indexes and topological connection relationships are recorded to form a facial vertex index table.

3. The method for face mouth shape coordination generation based on geometric constraints according to claim 2, wherein, The three-dimensional morphable model adopts a set of fixed-topology face triangular meshes, which contains 53215 facial vertices and 105054 edges, and each vertex always keeps the same number in the life cycle of the three-dimensional morphable model; the fixed topology covers the complete visible area from the hairline to the mandible, and a deformable vertex is reserved on the inner edge of the lips to ensure the accuracy consistency of the subsequent mouth opening and closing space shape matrix.

4. The method for generating face mouth shape based on geometric constraints according to claim 3, wherein, In step 2, in the facial vertex index table, the pre-defined upper lip outer edge vertex number range and lower lip outer edge vertex number range are retrieved to obtain an upper lip outer edge vertex subset and a lower lip outer edge vertex subset, respectively; according to the topological connection relationship of the three-dimensional morphable model, the upper lip outer edge vertex subset and the lower lip outer edge vertex subset are traversed ring by ring to obtain a vertex ordered sequence that monotonically increases along the direction of the lip contour to form a lip contour polygon chain; in the facial vertex index table, the gingival inner edge vertex number range is retrieved to obtain a gingival inner edge vertex subset; based on the normal vector direction and the local curvature threshold of the vertex, redundant vertices connected to the hard palate are removed, and only vertices located in the gum-alveolar junction area are retained to generate a purified gingival boundary sequence.

5. The method for face mouth shape coordination generation based on geometric constraints according to claim 4, wherein, In step 2, the centroid of each vertex of the purified gingival boundary sequence is calculated, denoted as the alveolar center point; a right-handed coordinate system formed by the alveolar center point, the alveolar long axis direction and the alveolar vertical direction is defined as the alveolar posture reference system; the lip contour polygon chain, the purified gingival boundary sequence and the alveolar posture reference system are collectively encapsulated as the lip-gingival joint boundary constraint set; in the lip-gingival joint boundary constraint set, the original three-dimensional coordinate value, the normal vector, the curvature scalar and the boundary type marker to which each vertex belongs are recorded.

6. The method for generating face mouth shape based on geometric constraints according to claim 5, wherein, In step 2, the three-axis vectors of the alveolar posture reference system are defined as the basis vectors of the oral cavity unified coordinate system with the alveolar center point as the origin; rigid coordinate transformation is performed on all vertices in the lip-gingival joint boundary constraint set to map the three-dimensional coordinates of the vertices from the three-dimensional morphable model global coordinate system to the oral cavity unified coordinate system; rigid coordinate transformation is performed on all vertices in the lip-gingival joint boundary constraint set to map the three-dimensional coordinates of the vertices from the three-dimensional morphable model global coordinate system to the oral cavity unified coordinate system; The mapped three-dimensional coordinates of the vertices, the normal vectors, the curvature scalars and the boundary type markers are tensorized and stored to form a geometric prior constraint tensor; A normalization operation is applied to the geometric prior constraint tensor to distribute the three-dimensional coordinates in the interval of-1 to 1, the normal vectors in the interval of-1 to 1 and the curvature scalars in the interval of 0 to 1.

7. The method of claim 6, wherein, Step 4 specifically includes: one-to-one correspondence between the high-dimensional tooth texture candidate tensor and the constraint consistency confidence spectrum according to the candidate sequence number dimension; applying the single instance confidence value of each candidate texture instance in the high-dimensional tooth texture candidate tensor in the constraint consistency confidence spectrum as the fusion weight to obtain a confidence weighted tooth texture fusion tensor; performing weighted summation on the confidence weighted tooth texture fusion tensor along the candidate sequence number dimension to obtain a single fused tooth texture representation vector, which is used to represent the optimal tooth texture state corresponding to the current input sample; the fused tooth texture matrix is unfolded as a one-dimensional vector according to the column and rearranged as a two-dimensional texture grid according to the unified UV layout table of the tooth region; in the two-dimensional texture grid, bidirectional Laplacian interpolation filling is performed on the missing pixels and Gaussian weighted smoothing processing is performed on the edge pixels to form an initial tooth texture map; local histogram equalization is applied to the initial tooth texture map to enhance the enamel highlight and the shadow details of the tooth neck to output an enhanced tooth texture map; a mouth opening and closing space form matrix corresponding to the enhanced tooth texture map is obtained; for each time frame in the mouth opening and closing space form matrix, the corresponding oral region vertex subset in the face vertex index table is located; the two-dimensional texture coordinates of the enhanced tooth texture map are mapped to the oral region vertex subset in the three-dimensional morphable model according to the mouth opening and closing space form matrix of the current frame, and a frame-level tooth texture vertex map dataset is generated; anti-aliasing, illumination compensation and gamma correction are completed on each frame in the frame-level tooth texture vertex map dataset in the unified rendering pipeline to finally obtain a mouth shape coordinated synthesized frame sequence.

Citation Information

Patent Citations

  • Method for fast constructing high-accuracy personalized face model on basis of facial images

    CN102222363A

  • Facial expression capturing method and device, electronic equipment and storage medium

    CN118629074A