Method for predicting facial change in orthodontic operation from thick to thin
Through a coarse-to-fine prediction method combined with parametric statistical models and deep learning technology, the problem of predicting anatomical consistency and nonlinear interaction of facial changes in orthognathic surgery was solved, and high-precision postoperative facial morphology prediction was achieved.
Patent Information
- Application Number
- CN202510781929.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies make it difficult to maintain the anatomical rationality of the predicted results in orthognathic surgery while accurately characterizing the complex nonlinear interactive relationship between maxillary and mandibular displacement and facial surface deformation.
A coarse-to-fine prediction approach is adopted. First, a parametric statistical facial model is used to generate an initial estimate with strong anatomical priors. Then, a deep learning maxillofacial transformation module is used to optimize facial features and capture the nonlinear interaction between the jaw and the facial surface.
It significantly improves the accuracy and anatomical consistency of postoperative facial morphology prediction, surpassing the existing most advanced technology, especially in complex anatomical areas.
Smart Images

Figure CN120707615A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for predicting changes in facial soft tissue after orthognathic orthodontic surgery, which is used for preoperative planning of orthognathic orthodontic surgery. Background Art
[0002] Maxillofacial deformity is a disease that affects the occlusal function and facial appearance, causing patients to have clinical manifestations such as jaw asymmetry, maxillary and mandibular retrusion, and occlusal disorder. Maxillofacial surgery is a common treatment method for improving maxillofacial deformity. This surgery can improve the patient's maxillofacial occlusal function and facial appearance by adjusting the position and morphology of the jaw. Accurate prediction of postoperative facial changes is the key to orthognathic surgery planning. Advanced modeling technology can help doctors predict the appearance of postoperative soft tissue, thereby enhancing the clinical decision-making process and improving preoperative doctor-patient communication. However, due to the complex nonlinear coupling relationship between facial soft tissue and deep bone structure, how to establish an accurate mapping model between preoperative bone displacement and postoperative morphological changes still faces major challenges.
[0003] Traditionally, methods for predicting postoperative facial morphology have relied primarily on biomechanical models, including classic modeling methods such as finite element models (FEMs), mass-spring models (MSMs), and mass-tensor models (MTMs). Among them, the incremental finite element model with realistic lip sliding effect (FEM-RLSE) is recognized as the most accurate method for simulating facial deformation, especially in improving the prediction accuracy of lip soft tissue. However, despite the high prediction accuracy of FEM-RLSE, this method requires the manual construction of a high-quality patient-specific finite element mesh model containing detailed lip structures, a process that is time-consuming and labor-intensive. In addition, there is still uncertainty in the overall predictive effectiveness of this model for the entire facial tissue.
[0004] In recent years, data-driven deep learning (DL) methods have emerged as an effective alternative to finite element models, attracting widespread attention due to their higher computational speed and robustness. Based on paired pre- and post-operative facial scans, these methods directly learn the complex mapping relationship between bone displacement and soft tissue deformation through neural networks, thereby enabling the prediction of postoperative facial morphological changes. However, the lack of explicit anatomical modeling and constraints on skeletal structure in purely data-driven deep learning frameworks can lead to anatomical inconsistencies in predictions and poor prediction of changes in key structures.
[0005] In recent years, three-dimensional parametric facial models represented by FLAME and SCULPTOR have significantly improved the accuracy of human facial modeling by introducing structural constraint mechanisms. Of particular note is that the SCULPTOR parametric model simultaneously models the maxillary and mandibular bones and the external surface structure of the face, and innovatively proposes a bone-driven mechanism—directly regulating the external facial epidermal geometry through internal bone deformation. Thanks to the introduction of the jaw constraint mechanism, the model can provide more anatomically consistent and highly realistic facial modeling effects, which provides an important technical foundation for facial deformation prediction used in orthognathic surgery. However, SCULPTOR still has the inherent limitations of parametric models: it relies on the linear blend skinning (LBS) algorithm for facial soft tissue prediction, and therefore cannot fully represent the complex nonlinear interaction between the maxillary and mandibular bones and the external surface of the face. Summary of the Invention
[0006] The technical problem to be solved by the present invention is that it is difficult with existing technologies to effectively maintain the anatomical rationality of the prediction results (such as the physiological coupling relationship between facial soft tissue and bone structure) in the prediction of facial changes during orthognathic surgery, while accurately characterizing the complex nonlinear interactive relationship between the maxillary and mandibular displacements and the deformation of the facial surface.
[0007] In order to solve the above technical problems, the technical solution of the present invention is to disclose a method for predicting facial changes during orthognathic orthodontic surgery from coarse to fine, which is characterized by being divided into a coarse prediction stage and a refined prediction stage, wherein:
[0008] In the rough prediction stage, a parameterized statistical facial model is used to generate an initial estimate of the postoperative face with a strong anatomical prior, and the initial rough prediction result F of the postoperative face is obtained. coarse-post , including the following steps:
[0009] The template of parametric statistical facial model was used to analyze the maxillary and mandibular structures of each patient before surgery. ori-pre With the outer surface of the face F ori-pre Perform fitting to generate the fitted preoperative maxillary and mandibular S fit-pre and the outer surface of the face F fit-pre ;
[0010] The preoperative maxillary and mandibular bones S fit-pre The simulated postoperative maxillary and mandibular S obtained from the surgical planning sim-post Perform fitting and calculate the corresponding deformation offset;
[0011] Apply the deformation offset to the original pre-operative facial outer surface F ori-pre , generate the initial rough prediction result F of the postoperative face coarse-post ;
[0012] In the refined prediction stage, the initial rough prediction result F coarse-post Expand it into a two-dimensional UV feature map, and use the facial feature extractor to extract facial features V from the UV feature map coarse-post , and then the maxillofacial transformation module uses the cross attention mechanism to calculate the mandibular motion parameters S bony-move As a condition, learn facial features V coarse-post The deformation mapping relationship between the face and the bone structure is used to apply the learned adjustment based on bone deformation to the original UV feature map to generate a refined postoperative facial prediction result F. fine-post .
[0013] Preferably, for patient i, the deformation offset in the rough prediction stage is calculated using the following method:
[0014] Fix the shape parameter β of patient i i and attitude parameters θ i ;
[0015] Optimize surgical feature parameters γ i To minimize the following energy terms:
[0016] E sim (β i ,γ i ,θ i )=E d +λ γ E γ +λ lmk E lmk +λ p2pl E p2pl
[0017] Where: E p2pl represents the preoperative maxillary and mandibular S fitted from the parametric statistical facial model fit-pre Vertex to simulated postoperative maxillary and mandibular S sim-post The point-to-surface distance of the patch; E d is the geometric data item, E d =λ d CD(T p ,C p )+(1-λ d )CD n (T p ,C p ), CD(·) represents the chamfer distance between two grids, CD n (·) Calculate the angle between the normals of the corresponding vertices, C p is the CT scan data, T p is a personalized head mesh model; E γ and E β yes Regularization term; Elmk is the constraint term based on anatomical landmarks; γ ,λ lmk ,λ p2pl is the weight parameter;
[0018] The optimized surgical feature parameter γ i Substitute the geometric shapes of the facial surface and the upper and lower jaws Calculate the deformation offset corresponding to the simulated postoperative face:
[0019]
[0020] Where: LBS(·) represents the linear blend skinning function; W is the learned skin weight matrix; J p (β i ,γ i ) characterizes the individual-specific anatomical mandibular joint position; T p (β i ,γ i ,θ i ) corresponds to the personalized head mesh model.
[0021] Preferably, the universal header template Shape Key Surgical feature key and gesture keys Obtain a personalized head mesh model T through linear combination p (β i ,γ i ,θ i ).
[0022] Preferably, the personalized head mesh model T p (β i ,γ i ,θ i )Depend on express.
[0023] Preferably, the individual-specific mandibular anatomical joint position J p (β i ,γ i ) is expressed as:
[0024]
[0025] Where: is a sparse matrix used to derive joint positions by regressing individual jaw mesh vertices.
[0026] Preferably, the original face and maxillary and mandibular geometry C obtained from CT scan data p, the accurate modeling of the preoperative geometric shape is achieved by minimizing the energy term shown in the following formula, and the shape parameter β that can characterize patient i is obtained i , surgical characteristic parameter γ i and attitude parameters θ i :
[0027] E pre (β i ,γ i ,θ i )=E d +λ γ E γ +λ β E β +λ lmk E lmk
[0028] Where: E d is the geometric data item used to ensure the personalized head mesh model T of patient i p (β i ,γ i ,θ i ) and CT scan data C p Maintaining spatial registration consistency during deformation, E d =λ d CD(T p ,C p )+(1-λ d )CD n (T p ,C p ), CD(·) represents the chamfer distance between two grids, CD n (·) Calculate the angle between the normals of corresponding vertices;
[0029] E γ and E β yes Regularization term;
[0030] E lmk is a constraint item based on anatomical landmarks;
[0031] λ γ ,λ β ,λ lmk is the weight parameter.
[0032] Preferably, the maxillofacial transformation module learns the deformation mapping relationship between facial features and bone structures through the cross attention mechanism CA(·), and its mathematical expression is:
[0033] q=V coarse-post W q
[0034] k=Ssim-post W k
[0035] v=S bony-move W v
[0036] V facial-move =softmax(qk T )v
[0037] Among them, W q 、W k 、W v are parameters that can be learned by the network;
[0038] The obtained facial motion feature vector V facial-move The feature integration is performed through a one-dimensional convolution layer, and then normalized by the tanh activation function. It is then superimposed channel by channel with the original UV feature map to generate an optimized UV map. The coordinate index of the UV map is then mapped back to the three-dimensional grid space, and finally the refined postoperative facial mesh is obtained, which is recorded as the refined postoperative facial prediction result F fine-post .
[0039] Preferably, the facial feature extractor and the maxillofacial transformation module form a neural network f θ , the optimization goal is to minimize the predicted UV map The difference between the postoperative baseline real UV image i.
[0040] Preferably, the neural network f θ The loss function is defined as:
[0041]
[0042] Where FFT(·) represents the fast Fourier transform operation, and λ is a hyperparameter for balancing the joint training of frequency and spatial domains.
[0043] The present invention proposes a novel coarse-to-fine method for predicting postoperative facial soft tissue changes. The present invention innovatively combines the anatomical consistency advantages of parametric facial models with the efficient prediction capabilities of deep learning to construct a two-stage prediction paradigm. In the coarse prediction stage, the present invention uses the SCULPTOR model to generate an initial estimate of the postoperative face with a strong anatomical prior, and establishes a linear correlation between the maxillofacial region to compensate for the lack of anatomical structure constraints in pure data-driven deep learning methods. In the fine prediction stage, the present invention introduces deep learning technology to optimize the initial prediction, aiming to capture the complex nonlinear interactions between the jaw and the facial surface. Experimental results on patients undergoing orthognathic orthodontic surgery show that the method disclosed in the present invention is significantly superior to the existing state-of-the-art technology (SOTA) in terms of the accuracy of postoperative facial morphology prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 Figure 1 illustrates an overview of the coarse-to-fine framework proposed in this paper, including: (a) the coarse prediction stage, in which the SCULPTOR model generates an initial prediction of the postoperative face based on the linear relationship between the skeleton and the facial surface; (b) the fine prediction stage, in which the coarse initial prediction is mapped to the UV parameter space, and the nonlinear interaction between the jaw and the facial surface is learned by fusing the neural network of the Facial Feature Extractor (FFE) and the Maxillofacial Morphological Transformation Module (CTM);
[0045] Figure 2 Schematic diagram of facial subregion division for quantitative evaluation;
[0046] Figure 3 This is a heat map comparing the qualitative point-to-plane distance (P2PL) error of the method of the present invention with FSC-Net, ACMT-Net, and ablation test in the postoperative face prediction task (error value range: 0-5 mm). DETAILED DESCRIPTION
[0047] The present invention will be further described below in conjunction with specific examples. It should be understood that these examples are intended only to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art may make various changes or modifications to the present invention, and these equivalents fall within the scope limited by the appended claims of the present invention.
[0048] Combine Figure 1 The embodiment of the present invention discloses a method for predicting facial changes during orthognathic orthodontic surgery from coarse to fine, comprising the following steps:
[0049] Step 1: In the rough prediction phase, a parameterized statistical facial model, SCULPTOR, is used to generate an initial estimate of the post-operative face with a strong anatomical prior. SCULPTOR effectively captures the bidirectional linear relationship between the facial surface and the underlying skeletal structure during the modeling of surgical deformation. By establishing a linear relationship between the maxillofacial region and the face, it compensates for the lack of anatomical constraints in purely data-driven deep learning methods.
[0050] Specifically, if Figure 1 As shown in (a), step 1 further includes the following steps:
[0051] Step 101: Use the template of the parametric statistical facial model SCULPTOR to calculate the maxillary and mandibular bone S of each patient before surgery. ori-pre With the outer surface of the face F ori-pre Perform fitting (data is derived from preoperative CT scan data) to generate the fitted preoperative maxillary and mandibular S fit-preand the outer surface of the face F fit-pre This process enables the present invention to obtain a set of parameters that characterizes the individual anatomical features of the patient, including individual-specific shape, surgical characteristics, and posture parameters.
[0052] Based on the parametric statistical facial model SCULPTOR, the present invention combines the geometric morphology of the facial surface and the upper and lower jaws. Modeled as:
[0053]
[0054] Where β, γ, and θ represent the principal component analysis (PCA) coefficient vectors of shape, surgical features, and posture, respectively; LBS(·) represents the linear mixed skinning function; W is the learned skinning weight matrix; J p Characterizes the individual-specific anatomical mandibular joint position; T p Corresponding personalized head mesh model.
[0055] Specifically, the mandibular joint position J p The mathematical definition of is:
[0056]
[0057] Where: is a sparse matrix used to derive joint positions by regressing individual jaw mesh vertices; personalized head mesh model T p It is a universal head template that includes the facial surface and jaw geometry (including maxillary and mandibular bones) Shape Key B S , surgical feature key B D and gesture key B P The linear combination is obtained, Among them, surgical feature key B D A physiologically plausible transformation space encoding coupled changes in skeletal structure and facial morphology.
[0058] For each patient i, the shape, surgical features and posture parameters β in the parameterized statistical facial model SCULPTOR are optimized. i , γ i ,θ i To establish a parametric model of the face and jaw before surgery.
[0059] Based on the original face and maxillary and mandibular geometry obtained from CT scan data p , SCULPTOR achieves accurate modeling of preoperative geometry by minimizing the following energy terms:
[0060] E pre (β i ,γi ,θ i )=E d +λ γ E γ +λ β E β +λ lmk E lmk (3)
[0061] Where: E d is the geometric data item used to ensure the personalized head mesh model T of patient i p (β i ,γ i ,θ i ) and CT scan data C p Maintaining spatial registration consistency during deformation, E d =λ d CD(T p ,C p )+(1-λ d )CD n (T p ,C p ), CD(·) represents the chamfer distance between two grids, CD n (·) Calculate the angle between the normals of corresponding vertices;
[0062] E γ and E β yes Regularization term, used to prevent the SCULPTOR model from overfitting on facial surface and jaw data;
[0063] E lmk It is a constraint term based on anatomical landmarks. It further improves the anatomical rationality of the deformation results by optimizing the alignment accuracy of predefined anatomical landmarks between the original mesh and the SCULPTOR fitted mesh.
[0064] λ β ,λ β ,λ kmk is the weight parameter.
[0065] Step 102: Based on the mechanism that bone structure changes during orthognathic surgery will cause facial deformation, the maxillary and mandibular bones S fit-pre The simulated postoperative maxillary and mandibular S obtained from the surgical planning sim-post The fitting process is performed to calculate the corresponding deformation offset. This fitting process is achieved by adjusting the surgical feature parameters in the SCULPTOR model.
[0066] The present invention is to combine the upper and lower jaws S fit-preThe simulated maxillary and mandibular bone S obtained after surgical planning is presented to the surgeon. sim-post By fitting the deformation, the corresponding deformation offset can be derived. Since the surgical feature key can represent the physiologically reasonable anatomical transformation space, its corresponding surgical feature parameters can effectively guide the model to capture the skeletal and facial deformation laws with anatomical consistency and authenticity from preoperative to postoperative. Therefore, in this fitting process, for each patient i, the shape and posture parameters β calculated by formula (3) are fixed i ,θ i , by optimizing only the surgical feature parameter γ i To minimize the following energy terms:
[0067] E sim (β i ,γ i ,θ i )=E d +λ γ E γ +λ lmk E lmk +λ p2pl E p2pl (4)
[0068] Where: E p2pl The preoperative maxillary and mandibular bone S fitted from SCULPTOR fit-pre Vertex to simulated postoperative maxillary and mandibular S sim-post The point-to-surface distance of the patch; E d 、E γ With E lmk The definition of is consistent with the corresponding energy term in formula (3); γ ,λ lmk ,λ p2pl is the weight parameter. By i Substituting into formula (1), the corresponding deformation offset of the simulated postoperative face can be calculated.
[0069] Step 103: Apply the calculated offset to the original front face surface F ori-pre , you can generate the initial rough prediction result F of the postoperative face coarse-post .
[0070] Step 2: In the refined prediction stage, the initial rough prediction result F coarse-post The image is unfolded into a two-dimensional UV plane map, and FocalNet is used as the facial feature extractor (hereinafter referred to as "FFE") for feature extraction, and combined with the maxillofacial transformation module (hereinafter referred to as "CTM") to capture the complex nonlinear relationship between bones and facial surfaces.
[0071] In order to overcome the limitation of the SCULPTOR model in characterizing the complex nonlinear relationship between the maxillary and mandibular bones and the facial surface, the present invention proposes to use a neural network f θ For the initial rough prediction result F coarse-post Optimize.
[0072] like Figure 1 As shown in (b), the neural network f θ It is mainly composed of FFE and CTM, and the technical process is as follows:
[0073] The initial rough prediction result F coarse-post The map is expanded to the UV parameter space, and then the high-dimensional features are extracted by FFE. In order to effectively capture the nonlinear relationship between the maxillofacial region and the face, the present invention introduces the CTM module, which is based on the cross-attention mechanism and uses the mandibular motion parameters S bony-move As a condition, the nonlinear mapping law between bone displacement and facial deformation is learned. Finally, the learned adjustment based on bone deformation is applied to the initial UV feature map to generate the refined postoperative facial prediction result F fine-post .
[0074] The embodiment of the present invention adopts FocalNet as the backbone network of FFE to extract facial features V from UV feature map coarse-post In the CTM module, the deformation mapping relationship between facial features and bone structures is learned through the cross attention mechanism CA(·), and its mathematical expression is:
[0075]
[0076] Among them, W q 、W k 、W v are learnable parameters of the network.
[0077] The facial motion feature vector V obtained through this process facial-move , firstly, feature integration is performed through a one-dimensional convolutional layer, and then normalized by a tanh activation function. The optimized vector is superimposed channel by channel with the original UV map to generate an optimized UV map. By mapping the coordinate index of the UV map back to the three-dimensional grid space, the refined postoperative facial mesh is finally obtained, denoted as F fine-post .
[0078] In the embodiment of the present invention, the neural network f θ The optimization goal is to minimize the predicted UV map The difference between the postoperative baseline real UV image I. Specifically, the loss function is defined as:
[0079]
[0080] Where FFT(·) represents the Fast Fourier Transform operation, and λ is a hyperparameter for balancing the joint training of frequency domain and spatial domain, which is set to 0.1 based on experimental experience.
[0081] To verify the effectiveness of the disclosed method, a total of 84 pairs of preoperative and postoperative head CT image data and multi-view facial scan data were collected, of which 60 pairs were used as training sets, 12 pairs were used as test sets, and the remaining 12 pairs were used as validation sets. Three-dimensional maxillofacial CT images were acquired using a spiral CT scanner (LightSpeed 16; GE, UK) with a spatial resolution of 0.48 × 0.48 × 1 mm. 3 .
[0082] The prediction accuracy of the proposed coarse-to-fine framework is evaluated through quantitative and qualitative experiments, and compared with advanced methods including FSC-Net and ACMT-Net. Both FSC-Net and ACMT-Net are trained based on their official open source code and experimental details in their papers, and their performance is evaluated on the test set of this invention. In the framework proposed by this invention, the weight parameter of the energy term is set to: γ =0.1,λ β =0.1,λ d =0.5,λ lmk =1,λ p2pl = 1. In the fine prediction stage, the feature dimensions of FocalNet are configured as 128, 256, 512, 256, and 128, where the feature dimensions of q, k, and v in the cross attention mechanism are all set to 256. The network training uses the Adam optimizer with a batch size of 2 and an initial learning rate of 8×10 -4 After 200 training cycles, the linear decay is reduced to 1×10 -6 All experiments were performed on an NVIDIA RTX 4070Ti GPU (16GB video memory) platform.
[0083] In order to comprehensively evaluate the prediction performance, this paper uses point-to-plane distance (P2PL) and chamfer distance (CD) as quantitative indicators (unit: mm / mm). At the same time, a refined prediction accuracy analysis is performed on different anatomical regions of the face. Figure 2 Shown: (A) forehead area, (B) nose area, (C) lip area, (D) chin area, (E) left cheek area, and (F) right cheek area.
[0084] As shown in Table 1 below, the proposed method outperforms FSC-Net and ACMT-Net in most evaluation metrics and facial subregion prediction. Specifically, the proposed method achieves an overall point-to-plane distance (P2PL) error of 0.89mm, significantly lower than FSC-Net (1.53mm) and ACMT-Net (1.36mm). The chamfer distance (CD) is 2.88mm, also lower than FSC-Net (3.91mm) and ACMT-Net (3.70mm). It is worth noting that the method disclosed in this invention performs well in six facial sub-regions (forehead, nose, lips, chin and bilateral cheeks), especially in complex anatomical regions such as the nose (CD 1.24mm vs. FSC-Net 1.87mm / ACMT-Net1.65mm) and lips (CD 1.15mm vs. FSC-Net 1.92mm / ACMT-Net 1.73mm), verifying its ability to model fine facial structures and complex deformation patterns. Figure 3 The qualitative P2PL error heat map shows that the method disclosed in the present invention has the least high-error areas (red), especially in the perioral and chin areas. FSC-Net and ACMT-Net have obvious error clustering, while the method disclosed in the present invention presents a more uniform low-error distribution, highlighting its robustness in capturing nonlinear facial geometric deformation.
[0085]
[0086] Table 1: Quantitative performance of the proposed method, FSC-Net, ACMT-Net and ablation test in the postoperative face prediction task (unit: mm / mm), the best and second best results are respectively and Underline Mark
[0087] To further evaluate the contribution of each component module of this framework, we conduct two ablation experiments to evaluate: (1) the reliability of facial prediction results in the rough prediction stage; and (2) the effectiveness of the maxillofacial transformation module (CTM).
[0088] Table 1 above is consistent with Figure 3Ablation experiment results are presented. In the first ablation experiment, when only the coarse prediction results generated by the SCULPTOR model are evaluated, the overall point-to-plane distance (P2PL) error is 1.36mm, and the chamfer distance (CD) is 3.66mm, indicating that the SCULPTOR model provides a reasonable basis for postoperative face prediction. In the second ablation experiment, when the CTM module is removed and only FocalNet is used for optimization, the P2PL error is reduced to 1.27mm and the CD is reduced to 3.52mm. The most significant improvement is seen in the full framework that introduces the CTM module in the refinement stage: the overall P2PL error is reduced to 0.89mm, and the CD is further reduced to 2.88mm. These results verify the key role of the CTM module in refining the coarse prediction results and achieving high-precision postoperative face prediction.
Claims
1. A method for predicting facial changes during orthognathic orthodontic surgery from coarse to fine scale, characterized in that: It is divided into the rough prediction stage and the refined prediction stage, in which: In the rough prediction stage, a parameterized statistical facial model is used to generate an initial estimate of the postoperative face with a strong anatomical prior, and the initial rough prediction result F of the postoperative face is obtained. coarse-post , including the following steps: The template of parametric statistical facial model was used to analyze the maxillary and mandibular structures of each patient before surgery. ori-pre With the outer surface of the face F ori-pre Perform fitting to generate the fitted preoperative maxillary and mandibular S fit-pre and the outer surface of the face F fit-pre ; The preoperative maxillary and mandibular bones S fit-pre The simulated postoperative maxillary and mandibular S obtained from the surgical planning sim-post Perform fitting and calculate the corresponding deformation offset; Apply the deformation offset to the outer surface of the face F ori-pre , generate the initial rough prediction result F of the postoperative face coarse-post ; In the refined prediction stage, the initial rough prediction result F coarse-post Expand it into a two-dimensional UV feature map, and use the facial feature extractor to extract facial features F from the UV feature map coarse-post , and then the maxillofacial transformation module based on the cross attention mechanism is used to calculate the mandibular motion parameters F bony-move As a condition, learn facial features F coarse-post The deformation mapping relationship between the face and the bone structure is used to apply the learned adjustment based on bone deformation to the original UV feature map to generate a refined postoperative facial prediction result F. fine-post .
2. The method for predicting facial changes during orthognathic surgery from coarse to fine scale according to claim 1, characterized in that: For patient i, the deformation offset in the rough prediction stage is calculated using the following method: Fix the shape parameter β of patient i i and attitude parameters θ i ; Optimize surgical feature parameters γ i To minimize the following energy terms: E sim (b i ,c i ,i i )=E d +λ γ E γ +λ lmk E lmk +λ p2pl F p2pl Where: F p2pl represents the preoperative maxillary and mandibular S fitted from the parametric statistical facial model fit-pre Vertex to simulated postoperative maxillary and mandibular S sim-post The point-to-surface distance of the patch; E d is the geometric data item, E d =λ d CD(T p ,C p )+(1-λ d )CD n (T p ,C p ), CD(·) represents the chamfer distance between two grids, CD n (·) Calculate the angle between the normals of the corresponding vertices, C p is the CT scan data, T p is a personalized head mesh model; E γ and E β is the l2 regularization term; E lmk is the constraint term based on anatomical landmarks; γ ,λ lmk ,λ p2p1 is the weight parameter; The optimized surgical feature parameter γ i Substitute the geometric shapes of the facial surface and the upper and lower jaws Calculate the deformation offset corresponding to the simulated postoperative face: Where: LBS(·) represents the linear blend skinning function; W is the learned skin weight matrix; J p (β i ,γ i ) characterizes the individual-specific anatomical mandibular joint position; T p (β i ,γ i ,θ i ) corresponds to the personalized head mesh model.
3. The method for predicting facial changes during orthognathic surgery from coarse to fine scale according to claim 2, wherein: Universal Header Template Shape Key Surgical feature key and gesture keys Obtain a personalized head mesh model T through linear combination p (β i ,γ i ,θ i ).
4. The method for predicting facial changes during orthognathic surgery from coarse to fine scale according to claim 3, characterized in that: The personalized head mesh model T p (β i ,γ i ,θ i )Depend on express.
5. The method for predicting facial changes during orthognathic surgery from coarse to fine scale according to claim 2, wherein: The individual-specific mandibular anatomical joint position J p (γ i ,γ i ) is expressed as: Where: is a sparse matrix used to derive joint positions by regressing individual jaw mesh vertices.
6. The method for predicting facial changes during orthognathic surgery from coarse to fine scale according to claim 2, wherein: Based on the original face and maxillary and mandibular geometry obtained from CT scan data p , the accurate modeling of the preoperative geometric shape is achieved by minimizing the energy term shown in the following formula, and the shape parameter β that can characterize patient i is obtained i , surgical characteristic parameter γ i and attitude parameters θ u : E pre (b u ,c u ,i i )=E d +λ γ E γ +λ β E β +λ lmk E lmk Where: E d is the geometric data item used to ensure the personalized head mesh model T of patient i p (β i ,γ i ,θ i ) and CT scan data C p Maintaining spatial registration consistency during deformation, E d =λ d CD(T p ,C p )+(1-λ d )CD n (T p ,C p ), CD(·) represents the chamfer distance between two grids, CD n (·) Calculate the angle between the normals of corresponding vertices; E γ and E β is the l2 regularization term; E lmk is a constraint item based on anatomical landmarks; λ γ ,λ β ,λ lmk is the weight parameter.
7. The method for predicting facial changes during orthognathic surgery from coarse to fine scale according to claim 1, wherein: The maxillofacial transformation module learns the deformation mapping relationship between facial features and bone structures through the cross-attention mechanism CA(·), and its mathematical expression is: q=V coarse-post W q k=S sim-post W k v=S bony-move W v V facial-move =softmax(qk T )v Among them, W q 、W k 、W v are parameters that can be learned by the network; The obtained facial motion feature vector V facial-move The feature integration is performed through a one-dimensional convolution layer, and then normalized by the tanh activation function. It is then superimposed channel by channel with the original UV feature map to generate an optimized UV map. The coordinate index of the UV map is then mapped back to the three-dimensional grid space, and finally the refined postoperative facial mesh is obtained, which is recorded as the refined postoperative facial prediction result F fine-post .
8. The method for predicting facial changes during orthognathic surgery from coarse to fine scale according to claim 1, wherein: The facial feature extractor and the maxillofacial transformation module form a neural network f θ , the optimization goal is to minimize the predicted UV map The difference between the postoperative baseline real UV image I.
9. The method for predicting facial changes during orthognathic surgery from coarse to fine scale according to claim 8, characterized in that: The neural network f θ The loss function is defined as: Where FFT(·) represents the fast Fourier transform operation, and λ is a hyperparameter for balancing the joint training of frequency and spatial domains.