Lightweight orthodontic treatment effect prediction method and system based on semantics and three-dimensional optimization

By using LoRA fine-tuning and 3D optimization technology, the high data and computational costs of existing orthodontic efficacy prediction methods are solved, providing low-cost, controllable, and intuitive 3D visualization results to assist clinical decision-making.

CN122115362APending Publication Date: 2026-05-29SHANGHAI XUHUI DISTRICT DENTAL HOSPITAL (SHANGHAI XUHUI DISTRICT DENTAL DISEASE CONTROL & PREVENTION INSTITUTE)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI XUHUI DISTRICT DENTAL HOSPITAL (SHANGHAI XUHUI DISTRICT DENTAL DISEASE CONTROL & PREVENTION INSTITUTE)
Filing Date
2026-02-06
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for predicting orthodontic efficacy suffer from problems such as high data requirements, high computational costs, unintuitive results, and difficulty in widespread application in clinical settings.

Method used

By combining LoRA fine-tuning technology, structured medical semantic prompts, and 3D geometric consistency optimization, an efficient prediction model is trained with a very small amount of paired data, providing intuitive 3D visualization results.

Benefits of technology

It enables orthodontic efficacy prediction with low data and low computing power requirements, and generates results that are medically controllable and interpretable with high three-dimensional geometric consistency, supporting doctor-patient communication and treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115362A_ABST
    Figure CN122115362A_ABST
Patent Text Reader

Abstract

The application provides a lightweight orthodontic efficacy prediction method and system based on semantics and three-dimensional optimization, relates to the technical field of orthodontic efficacy prediction, and comprises the following steps: obtaining facial images and cephalometric data of historical orthodontic patients before and after orthodontic treatment to form a training set, constructing a structured orthodontic semantic prompt, using the pre-training diffusion model as the basis, using the training set images and the semantic prompt, and fine-tuning the prediction model through the LoRA technology; inputting the orthodontic pre-image and data of a target patient into the image to the prediction model, generating a two-dimensional facial image after orthodontic treatment, and performing three-dimensional reconstruction; optimizing the geometric consistency with the side image, and outputting a three-dimensional model for visual display. Through the combination of the LoRA fine-tuning technology, the structured medical semantic prompt and the three-dimensional geometric consistency optimization, the application can train an efficient prediction model under a small amount of paired data, reduce the calculation cost of model training and reasoning, adapt to the deployment of a clinical environment, and realize the controllability and interpretability of the generated results through the medical semantic prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of orthodontic efficacy prediction technology, and in particular to a lightweight orthodontic efficacy prediction method and system based on semantic and three-dimensional optimization. Background Technology

[0002] In the field of orthodontic treatment, post-treatment changes in facial soft tissues (such as the perioral, mandibular, and lateral facial areas) have a significant impact on patients' aesthetics and function. Demonstrating these potential changes in a visual and individualized manner before treatment is of significant clinical value in improving patient understanding and treatment adherence. Currently, several methods have attempted to predict post-orthodontic facial appearance, but significant limitations remain:

[0003] 1. Traditional statistical model methods: These methods construct average deformation templates based on limited population data, resulting in coarse predictions, a lack of individualized representation, and an inability to intuitively display detailed changes in soft tissue.

[0004] 2. Generative Adversarial Networks (GANs) or paired image transformation methods: While these can generate more realistic images, they rely on large-scale, strictly paired pre- and post-orthodontic image data. Such data is difficult to obtain in clinical settings, expensive, and easily affected by shooting conditions (such as lighting, pose, and background), resulting in poor model generalization ability.

[0005] 3. Reconstruction methods based on 3D scanning: Facial data is acquired through 3D scanning equipment and soft tissue simulation is performed. However, the hardware cost is high and the acquisition conditions are demanding, making it difficult to promote and apply on a large scale in clinical settings.

[0006] 4. Generation methods based on diffusion models: In recent years, diffusion models have shown powerful capabilities in image generation and editing. However, direct fine-tuning of all parameters of a large diffusion model still requires a large amount of data and computing power, which is not suitable for deployment in resource-limited clinical environments.

[0007] Therefore, there is an urgent need for a method to predict orthodontic outcomes that requires low data, has low computational cost, provides intuitive results, and is easy to deploy in clinical practice. Summary of the Invention

[0008] To address the problems of existing technologies, such as coarse and unintuitive prediction results, strong data dependence, high cost, and difficulty in clinical deployment, this invention aims to provide a lightweight orthodontic efficacy prediction method and system based on semantics and 3D optimization. By combining LoRA fine-tuning technology, structured medical semantic prompts, and 3D geometric consistency optimization, it enables the training of an efficient prediction model with a very small amount of paired data, reducing the computational cost of model training and inference, adapting to clinical deployment, and achieving controllability and interpretability of the generated results through medical semantic prompts. At the same time, it provides intuitive 3D visualization results to assist in doctor-patient communication and treatment decision-making.

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] The present invention provides a lightweight orthodontic efficacy prediction method based on semantics and 3D optimization in its first aspect, comprising: S1. acquiring frontal and lateral facial images and corresponding cephalometric data of historical orthodontic patients before and after orthodontic treatment, forming a training dataset; S2. performing standardized preprocessing on the facial images in the training dataset, and constructing a structured orthodontic semantic cue for each facial image, including its orthodontic stage, perspective attribute, age group, and semantic description based on cephalometric data conversion; S3. using a pre-trained diffusion model as the base model, and fine-tuning the base model using the standardized preprocessed facial images and their corresponding structured orthodontic semantic cue in the training dataset, employing LoRA technology to obtain an orthodontic efficacy prediction model; S4. receiving frontal and lateral facial images and corresponding cephalometric data of a target patient before orthodontic treatment; S5. S6. Perform standardized preprocessing on the facial image of the target patient in the same manner as in step S2, and construct corresponding structured orthodontic semantic prompts based on its cephalometric data, wherein the semantics of the orthodontic stage are set to the pre-orthodontic state; S7. Input the standardized preprocessed facial image of the target patient into the orthodontic efficacy prediction model, and generate the predicted two-dimensional facial images of the target patient from the front and side through local redrawing based on the structured orthodontic semantic prompts constructed in step S5; S8. Based on the predicted two-dimensional facial images of the target patient after orthodontic treatment, perform three-dimensional facial reconstruction, and use the predicted two-dimensional facial image of the side after orthodontic treatment as geometric constraints, optimize the geometric consistency of the reconstructed three-dimensional facial model through differentiable rendering technology, and output the predicted three-dimensional facial model after orthodontic treatment; S9. Visualize the three-dimensional facial model after orthodontic treatment.

[0011] In a first aspect, the present invention provides a preferred embodiment in which, in step S2, the orthodontic stage label includes before and after orthodontic treatment; the viewpoint attributes include frontal view, lateral view, and 45° view; the age group includes adolescents, young adults, adults, and middle-aged people as coarse-grained age group labels; the key measurements of the cephalometric lateral view are selected from the cephalometric measurements of the lateral view, including key indicators related to facial visual features, such as skeletal indicators, dental indicators, and soft tissue indicators; and the visual description is specifically generated by an image description model to describe facial features.

[0012] In a first aspect, the present invention provides a preferred embodiment in which, in steps S2 and S5, the structured orthodontic semantic prompt further includes: automatically generating descriptive text for the visual features of the perioral and mandibular regions in a facial image using an image description model.

[0013] In a first aspect, the present invention provides a preferred solution in which, in step S3, the fine-tuning of the base model using LoRA technology is specifically achieved by injecting a trainable low-rank matrix into the attention module weights of the base model.

[0014] The present invention provides a preferred embodiment in the first aspect, wherein in step S6, the local redrawing specifically involves: applying a mask to the soft tissue region where orthodontic changes are expected to occur on the standardized preprocessed facial image of the target patient, and generating a conditional image of the masked region by the orthodontic efficacy prediction model, while the content of the non-masked region remains unchanged.

[0015] The present invention provides a preferred embodiment in the first aspect. In step S7, the step of using the orthodontic predicted two-dimensional facial image of the side as a geometric constraint and optimizing the geometric consistency of the reconstructed three-dimensional facial model through differentiable rendering technology includes: rendering the reconstructed three-dimensional facial model to the same viewpoint as the orthodontic predicted two-dimensional facial image of the side to obtain a rendered side profile; calculating the difference between the rendered side profile and the actual profile of the side image, and iteratively optimizing the vertex positions of the three-dimensional facial model by combining a mesh smoothing regularization term with the goal of minimizing the difference.

[0016] In a first aspect, the present invention provides a preferred embodiment in which, in step S3, the structured orthodontic semantic prompt includes key clinical features describing the facial bony and soft tissue morphology, the key clinical features including one or more of the following: mandibular plane angle, chin projection, degree of lip protrusion, nasolabial angle, and soft tissue E-line offset.

[0017] The present invention provides a preferred embodiment in the first aspect, wherein the step of generating the image through local redrawing in step S6 specifically includes: applying a mask to the soft tissue region where orthodontic changes are expected to occur on the standardized preprocessed facial image of the target patient to construct a local missing area; the orthodontic efficacy prediction model, guided by the structured orthodontic semantic cues, performs image completion generation on the masked region while maintaining the identity features of the non-masked region unchanged, thereby obtaining a predicted two-dimensional facial image of the target patient after orthodontic treatment; wherein the soft tissue region includes at least the upper and lower lips, chin, and mandibular border.

[0018] In a first aspect, the present invention provides a preferred embodiment in which, in step S6, multiple candidate post-orthodontic prediction two-dimensional facial images are generated within a preset range by adjusting the denoising sampling intensity, wherein the preset range is 0.2 to 0.8.

[0019] In a second aspect, this invention provides a lightweight orthodontic efficacy prediction system based on semantics and 3D optimization, used to implement the method, comprising: a historical training dataset construction module, used to acquire frontal and lateral facial images and corresponding cephalometric data of historical orthodontic patients before and after orthodontics, constituting a training dataset; a first image preprocessing and semantic prompt construction module, used to perform standardized preprocessing on the facial images in the training dataset, and construct a structured orthodontic semantic prompt for each facial image, including its orthodontic stage, perspective attribute, age group, and semantic description based on cephalometric data conversion; an orthodontic efficacy prediction model construction module, used to fine-tune the basic model using LoRA technology with a pre-trained diffusion model as the base model, utilizing the standardized preprocessed facial images and their corresponding structured orthodontic semantic prompts in the training dataset, to obtain an orthodontic efficacy prediction model; and a target patient image and data acquisition module, used to receive frontal and lateral facial images and corresponding cephalometric data of the target patient before orthodontics. The system includes: a cephalometric measurement data module; a second image preprocessing and semantic prompting construction module, used to perform standardized preprocessing on the facial image of the target patient in the same manner as in step S2, and to construct corresponding structured orthodontic semantic prompts based on its cephalometric measurement data, wherein the semantics of the orthodontic stage are set to the pre-orthodontic state; a two-dimensional orthodontic efficacy prediction module, used to input the standardized preprocessed facial image of the target patient into the orthodontic efficacy prediction model, and, based on the structured orthodontic semantic prompts constructed in step S5, to generate predicted two-dimensional facial images of the target patient from the front and side through local redrawing; a three-dimensional orthodontic efficacy prediction module, used to perform three-dimensional facial reconstruction based on the predicted two-dimensional facial images of the target patient after orthodontic treatment, and using the predicted two-dimensional facial images of the side after orthodontic treatment as geometric constraints, to optimize the geometric consistency of the reconstructed three-dimensional facial model through differentiable rendering technology, and to output the predicted three-dimensional facial model after orthodontic treatment; and a visualization module, used to visualize the three-dimensional facial model after orthodontic treatment.

[0020] Compared with the prior art, the present invention has the following advantages:

[0021] 1. Low data and computing power requirements: Using LoRA fine-tuning, only a few hundred pairs of clinical data are needed for effective training, and the training and inference costs are far lower than full model fine-tuning, making it possible to deploy on ordinary clinical workstations.

[0022] 2. Predicted results are medically controllable and interpretable: By converting cephalometric parameters into semantic cues that the model can understand, the generation process is guided by clinical knowledge, which increases clinical interpretability and overcomes the "black box" defect of traditional generative models.

[0023] 3. The prediction results are highly realistic and consistent with the identity: By combining local redrawing strategies, the patient's identity features (such as facial features and skin color) are preserved during generation, and only the target soft tissue area is changed, which improves the naturalness and credibility of the prediction results.

[0024] 4. High consistency of 3D geometry: The innovative side view (side image) constraint optimization mechanism effectively corrects the side geometric error of single-view 3D reconstruction, so that the final 3D prediction model maintains a consistent and accurate shape in all views, providing spatial information far exceeding that of 2D images.

[0025] 5. The system is highly practical: It forms an integrated workflow from two-dimensional prediction to three-dimensional visualization, provides multiple hypotheses for doctors to choose from, and supports three-dimensional model display, which greatly facilitates doctor-patient communication and treatment plan design.

[0026] In summary, this invention presents a lightweight orthodontic efficacy prediction method and system based on semantic and 3D optimization. It utilizes low-rank adaptive (LoRA) technology to achieve low-cost, high-efficiency fine-tuning of the diffusion model, learning soft tissue change trends with minimal paired data. By constructing a structured semantic cueing system closely linked to cephalometric data, it achieves medical interpretability and clinical controllability of the generation process. Furthermore, it introduces 3D reconstruction and geometric consistency optimization based on side-view constraints to overcome the perspective inconsistency problem of pure 2D prediction and provides more intuitive and reliable 3D visualization results. Ultimately, it constructs a lightweight decision support tool that is easy to deploy and use clinically. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0028] Figure 1 This is a flowchart illustrating a lightweight orthodontic efficacy prediction method based on semantic and 3D optimization, provided in a specific embodiment of the present invention.

[0029] Figure 2 This is a schematic diagram of a module of a lightweight orthodontic efficacy prediction system based on semantic and three-dimensional optimization, provided as a specific embodiment of the present invention. Detailed Implementation

[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] It should be noted that the directional terms such as "front" and "side" and the ordinal numbers such as "first" and "second" used below are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance or the quantity implied. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0032] Example 1

[0033] Please refer to Figure 1 This embodiment provides a lightweight orthodontic efficacy prediction method based on semantics and 3D optimization, mainly implemented through steps S1 to S8:

[0034] S1. Obtain frontal and lateral facial images of historical orthodontic patients before and after orthodontic treatment, along with corresponding cephalometric data, to form a training dataset;

[0035] S2. Perform standardized preprocessing on the facial images in the training dataset, and construct a structured orthodontic semantic cue for each facial image, which includes its orthodontic stage, perspective attribute, age group, and semantic description based on cephalometric data conversion;

[0036] S3. Using the pre-trained diffusion model as the base model, and utilizing the standardized preprocessed facial images and their corresponding structured orthodontic semantic cues in the training dataset, the LoRA technique is used to fine-tune the base model to obtain an orthodontic efficacy prediction model.

[0037] S4. Receive the frontal and lateral facial images of the target patient before orthodontic treatment and the corresponding cephalometric data;

[0038] S5. Perform standardized preprocessing on the facial image of the target patient in the same way as in step S2, and construct a corresponding structured orthodontic semantic prompt based on its cephalometric data, wherein the semantics of the orthodontic stage are set to the state before orthodontics;

[0039] S6. Input the standardized preprocessed facial image of the target patient into the orthodontic efficacy prediction model, and generate the predicted two-dimensional facial images of the target patient's front and side views through local redrawing based on the structured orthodontic semantic prompts constructed in step S5.

[0040] S7. Based on the predicted two-dimensional facial image of the target patient after orthodontic treatment, perform three-dimensional facial reconstruction, and use the predicted two-dimensional facial image of the side after orthodontic treatment as a geometric constraint. Optimize the geometric consistency of the reconstructed three-dimensional facial model through differentiable rendering technology, and output the predicted three-dimensional facial model after orthodontic treatment.

[0041] S8. Visualize the three-dimensional facial model after orthodontic treatment.

[0042] The following is a detailed description of each step of the method in this embodiment:

[0043] 1. Construction of historical training dataset:

[0044] S1. Obtain frontal and lateral facial images and corresponding cephalometric measurements of historical orthodontic patients before and after orthodontic treatment to form a training dataset.

[0045] In practice, 150 patients with complete fixed orthodontic treatment and comprehensive records were selected from the hospital's orthodontic cohort between 2020 and 2025. For each patient, two sets of data were collected at two time points:

[0046] Images from T1 (pre-treatment) and T2 (post-treatment): Each period includes five clinical photographs taken under standardized conditions: a frontal view, a left 90° lateral view, a right 90° lateral view, a left 45° oblique lateral view, and a right 45° oblique lateral view. Patients were required to be in a resting jaw position with relaxed lip muscles and light biting contact to ensure consistent soft tissue condition during photography.

[0047] Corresponding cephalometric measurement data: Each period is accompanied by a digital lateral cephalometric X-ray taken at the same time. The measurements are determined by professional physicians using specialized software, and a series of cephalometric measurements (such as multiple skeletal, dental, and soft tissue measurements) are calculated to form a structured measurement data table.

[0048] All data collection and use are approved by the hospital's ethics committee, and images undergo strict anonymization processing to remove background and identifiable information. They are managed with unique anonymous codes to protect patient privacy.

[0049] The training dataset, or sample set, can be randomly divided into a training set and a test set at a ratio of 9:1 to ensure the sufficiency of model training and the objectivity of performance evaluation.

[0050] 2. Image preprocessing and structured semantic cue construction:

[0051] S2. Perform standardized preprocessing on the facial images in the training dataset, and construct a structured orthodontic semantic cue for each facial image, which includes its orthodontic stage, perspective attribute, age group, and semantic description based on cephalometric data conversion.

[0052] This step processes the data obtained in S1 to prepare high-quality image-text pairs for model training. The specific implementation process of this step is as follows:

[0053] S2.1 Image Standardization Preprocessing: A uniform processing procedure is performed on all collected facial images to eliminate composition and scale differences. It is understood that facial editing and appearance prediction tasks are highly sensitive to image composition and facial proportions. Unnormalized facial photographs can introduce unnecessary distributional differences during training. For example, variations in head size, composition tightness, and background area proportions from different patients or viewpoints may cause the model to incorrectly learn "compositional styles" rather than genuine soft tissue morphological changes, thus interfering with LoRA's effective learning in domain transfer before and after orthodontics. To reduce such irrelevant variance, this embodiment performs a uniform standardization preprocessing procedure on all images. Specifically, firstly, a high-precision face detector (or face detection model) is used to locate the face region. While maintaining the original facial geometry, tools such as the open-source library DeepFace are used to intelligently crop the face region, preserving the complete face and removing irrelevant background and contextual information. Subsequently, the cropped image is proportionally filled along its long side to obtain an approximately square composition, ensuring that the model receives a structurally consistent input frame across different cases. Finally, all images were scaled down to a uniform resolution of 1024×1024 pixels to align with the input requirements of the selected Stable Diffusion XL model (base model: a pre-trained diffusion model). This process ensures that the model focuses on facial morphological variations rather than learning irrelevant compositional styles.

[0054] This "detect-crop-fill-scale" process ensures consistency in the composition and proportions of the input images. This allows the model to focus its training on local morphological changes in facial soft tissues (such as lip protrusion, jaw position, and chin contour), ensuring the model concentrates on facial morphological variations rather than learning irrelevant layout changes or compositional styles. This standardized input not only improves the stability of training gradients but also reduces the risk of overfitting during subsequent LoRA fine-tuning, providing a cleaner and more consistent training environment for domain transfer before and after orthodontic treatment.

[0055] S2.2 Constructing structured orthodontic semantic cues: This is the key to achieving semantic control.

[0056] To ensure that the semantics accurately depict changes in the patient's soft tissue and dentofacial structure, this embodiment employs prompt annotation, strictly based on cephalometric parameters obtained from lateral cephalometric radiographs and related lateral radiographs, and has been revised and reviewed by a professional orthodontist. Unlike traditional methods that rely solely on visual observation or subjective description, this embodiment combines quantitative imaging measurements and facial visual features to construct a semantic description system that simultaneously possesses medical accuracy, visual expressiveness, and expert experience. Specifically, the medical terminology section covers key indicators reflecting the patient's facial shape and soft tissue morphology, such as the nasolabial angle, chin position, lip protrusion, the relationship between the upper and lower lips and the E-line, and the mandibular plane angle; the visual description section focuses on appearance features easily understood by the model, such as "lip protrusion," "mandibular retrusion," "mental muscle tension," and "convex facial shape." These two types of descriptions are combined using structured templates, ensuring that each image simultaneously possesses clinical interpretability and semantic cues usable by the generative model. This combined annotation method of quantitative image indicators and visual semantic cues enables the model to accurately capture the most critical soft tissue changes before and after orthodontic treatment during training, while ensuring that the generated predictions are consistent with clinical observations, thereby improving the controllability, medical effectiveness, and readability of the model output for clinicians.

[0057] In this step, a descriptive text is generated for each image in S1 (e.g., "Patient A-T1 stage - left lateral view"), with the following structure:

[0058] Orthodontic stage labeling: Add high-level semantic labels to each image, i.e., explicit annotation.<pre_ortho> (Before orthodontic treatment) or<post_ortho> (Post-orthodontic treatment). This label clearly distinguishes the treatment stage, allowing the model to focus on learning the morphological changes related to orthodontic intervention, without being disturbed by irrelevant factors such as lighting and clothing.

[0059] Viewpoint Attributes: Each image is labeled with viewpoint attributes, including frontal view, lateral view, and oblique clinical view (commonly used in clinical settings at 45°). Because multimodal encoders like CLIP are highly sensitive to viewpoint differences, different viewpoints often belong to completely different conceptual spaces in visual semantics. Therefore, explicitly describing the viewpoint can prevent the model from mistakenly classifying facial features from different angles into one category, thereby improving the stability of text-image alignment.

[0060] Age Groups: Patients were divided into four age groups based on their age: adolescent, young adult, adult, and middle-aged adult, serving as coarse-grained age (age group) labels. Age is closely related to bone mass, bone remodeling potential, and soft tissue, and is an important factor affecting the feasibility and effectiveness of orthodontics. By providing age group information, this embodiment injects a lightweight but effective biological prior into the model, helping it learn structural relationships that better conform to clinical patterns.

[0061] Key measurements from lateral cephalometric radiographs are converted into semantic descriptions: To establish a learnable correspondence between image semantics and medical measurements, this embodiment selects key indicators highly correlated with facial visual features from lateral cephalometric radiograph measurements, including but not limited to:

[0062] Bony: SNA, SNB, ANB, MP-SN, FMA (MP-SH)

[0063] Dental type: U1–NA, L1–NB, U1–SN, IMPA (L1–MP)

[0064] Soft tissue: UL–EP, LL–EP

[0065] For each measurement value, it is converted into a corresponding semantic phrase based on whether it is below / within / above the standard reference range. For example, consider SNB:

[0066] Below standard value: Mandibular Retrusion

[0067] Standard range: Normal Mandibular Position (Normal mandibular position)

[0068] Above standard value: Mandibular Protrusion

[0069] For example, if the SNB angle is 76° (below the normal range), it is converted to "Mandibular Retrusion". Semantic descriptions of multiple indicators will be written side by side in the prompt word.

[0070] This type of description allows the model to learn using both visual representation and medical structural information simultaneously. It should be noted that while lateral radiographs yield a vast number of cephalometric measurements, excessive stacking of semantic descriptions may distract the model and hinder the learning of stable mappings. Therefore, this embodiment only selects the core indicators most strongly correlated with visually perceptible morphology and possessing the greatest diagnostic significance for annotation, ensuring the simplicity and effectiveness of the text labels.

[0071] Visual Description: Facial feature description text generated by an image description model. Further, images can be input into a large-scale vision-language model such as LLaVA to generate a natural language description of the lower face morphology (e.g., "fulllips with slight protrusion"), which is then appended to the end of the prompt. Specifically, to improve the model's sensitivity to key visual features, in one optional implementation, an automated visual description generated by a large model is introduced for each image. When an image is input into an image description model, such as LLaVA, the model automatically provides a description of the current image. For example, a general description of the image might use the word "caption," while the labeled data used for model training typically uses prompt-like annotations. Compared to using only the original image, these captions explicitly indicate the areas and attributes the model should focus on, especially in facial images with multiple viewpoints and significant head position variations. By emphasizing visual elements related to the lower face and soft tissue (such as perioral soft tissue contours, lip shape, and jaw appearance) in the prompts, the model can obtain more stable semantic guidance during training, thereby more accurately capturing key features related to orthodontic variations. Furthermore, the introduction of text descriptions improves the usability and scalability of the dataset. A unified labeling system can easily integrate additional data (such as CT scans, facial scans, or clinical records) in the future, laying the foundation for multimodal models and significantly enhancing cross-domain generalization capabilities.

[0072] Furthermore, in a more preferred embodiment, the labeled text can be optimized by integrating the descriptive text from the previous stages, removing duplicate words, and merging words with similar meanings. Simultaneously, for LoRA training, if the labeled text is merely a measurement / diagnosis label, the model cannot directly see the corresponding visual features from the image, thus requiring secondary optimization. The results can be reviewed by a professional orthodontist to optimize key information into visually visible content.

[0073] 3. Training a lightweight orthodontic efficacy prediction model:

[0074] S3. Using the pre-trained diffusion model as the base model, and leveraging the standardized preprocessed facial images and their corresponding structured orthodontic semantic cues in the training dataset, the LoRA technique is used to fine-tune the base model to obtain an orthodontic efficacy prediction model.

[0075] (1) Basic model: Stable Diffusion XL 1.0 was used as the pre-trained basic model.

[0076] (2) Fine-tuning method: LoRA is used for lightweight adaptation.

[0077] In recent years, image generation methods based on diffusion models have shown significant advantages in medical imaging, facial reconstruction, and personalized prediction. However, conventional full fine-tuning is not only computationally expensive but also prone to catastrophic forgetting, making it unsuitable for small sample sizes and specific medical scenarios. To incorporate orthodontic priors while ensuring training efficiency and model stability, this embodiment employs LoRA (Low-Rank Adaptation) as the core lightweight fine-tuning technique.

[0078] The core idea of ​​LoRA is to restrict the updates of key weight matrices in the model to low-rank decomposition. By introducing a small set of low-rank parameters alongside the original model, a low-cost, highly stable, and controllable model can be achieved. Specifically, the LoRA module, in low-rank decomposition form, is inserted into the attention layer of all Transformer blocks in the SDXL model, acting on the projected weight matrices of the query, key, value, and output. The rank of LoRA is set to 8.

[0079] (3) Model training and inference strategy: The AdamW optimizer was used on the training configuration, with a learning rate of 1e-4. The batch size was set to 4. FP16 mixed precision training was adopted. The total number of training iterations was set to 4000. Training was completed on a workstation equipped with an NVIDIA GeForce RTX 4090 graphics card, which reflects the lightweight characteristics of the whole process.

[0080] The specific model training process and inference strategies will be further elaborated below:

[0081] a. Cue-driven model training: Enhance the model's understanding of maxillofacial structures through medical semantic cues.

[0082] Specifically, to enable the model to capture key visual features related to orthodontics, this embodiment first trains the model in a targeted manner based on the aforementioned orthodontic semantic prompts. The prompts used contain core descriptions of facial bony and soft tissue morphology, such as the mandibular plane angle, chin projection, degree of lip protrusion, nasolabial angle, and soft tissue E-line offset. These prompts are trained together with high-quality paired images before and after orthodontic treatment, allowing the model to internalize these medical semantics in visual space, thereby enabling it to automatically respond to changes in target morphology during image generation.

[0083] During training, the text encoder and model representation space are jointly optimized, enabling semantic cues to precisely control the generation process, thereby improving the model's sensitivity and controllability to subtle maxillofacial changes. This construction of a joint "language-vision" representation is particularly important for medical image prediction because it allows the model to understand structural changes in a unified latent space, rather than relying solely on pixel-level differences.

[0084] b. Patient consistency inference based on supplementary drawing (local redrawing): Local structure prediction is achieved by using a local redrawing strategy while maintaining patient identity.

[0085] Specifically, it mainly consists of the following key processes:

[0086] First, maintaining patient consistency: To strictly maintain consistency between patient identity and original visual features during the inference phase, this embodiment employs a local redrawing strategy for prediction. Specifically, clinicians or researchers only need to simply mask the local areas expected to change (mainly the upper and lower lips, chin, and mandibular border) on the original image, generating a mask. The model then completes the masked areas while maintaining global invariance. This inference method has several advantages:

[0087] ① Patient consistency: Since the model only modifies local areas, unrelated structures such as facial skeleton, skin texture, hairline, and nose are completely preserved, avoiding the identity drift problem that may be introduced by conventional generative models.

[0088] ②Regional controllability: By manually determining the occlusion area, the model can be clearly limited to changes within the clinically relevant area, such as the nasolabial fold area, mandibular retrusion, lip protrusion, and chin point position adjustment, thereby reducing unintended changes.

[0089] ③ Improve prediction reliability: Local supplementation reduces the model's pressure to reconstruct global features, which can significantly reduce the "free play" in the generation process, making the output more stable and more in line with medical semantics.

[0090] ④ Facilitates doctor interaction: The drawing operation is simple, allowing clinicians to quickly try different local adjustment methods and observe the generated results, providing visual assistance for orthodontic treatment plan discussions.

[0091] Secondly, the specific reasoning process:

[0092] During the inference phase, this embodiment deploys the trained LoRA model within a clinically-oriented workflow. The overall workflow begins with patient image acquisition and includes several steps such as viewpoint acquisition, automatic generation and manual revision of text prompts, local mask construction, and multi-view result generation.

[0093] First, one frontal and one lateral clinical photograph of the patient are acquired under standardized imaging conditions. For these two input images, a pre-trained image description model is used to automatically generate initial text prompts, which include information on changes in facial bony and soft tissue features, while preserving viewpoint-related information (frontal / lateral). At this stage, the original images are labeled with the corresponding stage label from the training process.<pre_ortho> During the prediction phase, the labels are explicitly rewritten as<post_ortho> This is to guide the model's transition from the "pre-treatment state" to the "post-treatment state" distribution.

[0094] Secondly, the system submits automatically generated prompts to orthodontists for optional manual fine-tuning. Doctors can add, delete, or adjust the weight of key semantics in the prompts based on past treatment experience and specific treatment plans. For example, they can strengthen the descriptions of "improvement of mandibular retrusion," "lip retraction," and "chin advancement," thereby achieving high-level control over the generated results without directly manipulating the model parameters. This mechanism combines the representational capabilities of the large model with the doctor's clinical prior knowledge, making the predicted results closer to the actual treatment goals.

[0095] Next, at the image input end, we manually smoothed out the lower half of the face region of the original image to construct a mask for locally missing areas. Editable space was only retained around the mouth, chin, and jawline, while strong constraints were applied to the upper half of the face and the background region to prevent identity shifts or deformations of irrelevant structures. Subsequently, the input image with locally missing areas and the stage labels were...<pre_ortho> Switch to<post_ortho> The text prompts and their corresponding binary masks are fed into the LoRA model and conditionally diffused within a set denoising range (e.g., 0.2–0.8).

[0096] Finally, this embodiment performs the aforementioned local redrawing inference process on the frontal and side input images respectively. Under unified semantic cues and stage label constraints, a predicted frontal view and a predicted side view are generated. The frontal view is mainly used to compare the geometric errors and perceptions of key points with the real frontal photo, while the side view emphasizes the coherence of contour lines, chin projection, and soft tissue contours, making it easier for doctors to assess whether the model has achieved the expected correction trend from a typical side profile perspective.

[0097] Through this reasoning path of "multi-perspective input - unified semantic control - multi-perspective prediction output", the model can provide valuable 2D prediction results with consistent 3D perception for clinical decision-making while maintaining identity consistency.

[0098] c. Multiple hypothesis generation for denoising control: Multiple possible prediction schemes are generated by adjusting the denoising intensity and combined with expert selection.

[0099] Specifically, to further enhance the controllability of inference, this embodiment introduces an adjustable denoising intensity during the generation stage, typically set in the range of 0.2–0.8. Lower denoising values ​​(e.g., 0.2–0.4) mean the model more strictly adheres to the original image structure, making only minor adjustments to local regions; while higher denoising values ​​(e.g., 0.6–0.8) allow the model to more boldly explore morphological variations in the latent space, thereby generating multiple potential prediction schemes. In practical applications, multiple candidate images are generated for the same patient using different denoising intensities, and orthodontic experts select the result that best matches clinical judgment. This strategy has the following advantages:

[0100] ① Avoid extreme results caused by a single random seed: No longer relying on a single inference output, significantly improving the stability of the results.

[0101] ② Forming "multiple hypotheses" predictions: Allowing for the presentation of multiple levels of facial changes, from conservative to radical, which facilitates clinical comparison.

[0102] ③ Introducing expert experience: Doctors select reasonable plans based on professional standards such as maxillofacial structure, soft tissue proportion, and facial balance, transforming the generation process from fully automated to collaborative judgment between models and experts, thereby improving clinical acceptability.

[0103] ④ Improve model robustness: By sampling in different denoising intervals, the randomness of the model is transformed into a controllable range of variation, avoiding the random deviation of a single output.

[0104] 4. Receive target patient data:

[0105] S4. Receive the frontal and lateral facial images of the target patient before orthodontic treatment and the corresponding cephalometric measurements.

[0106] Specifically, the pre-treatment clinical data of the patients to be predicted (target patients) are received, including: one standard frontal photograph, one standard lateral photograph, and the results of their pre-treatment cephalometric radiographs (i.e., cephalometric data table).

[0107] 5. Target patient data preprocessing and suggestion construction:

[0108] S5. Perform standardized preprocessing on the facial image of the target patient in the same manner as in step S2, and construct corresponding structured orthodontic semantic prompts based on their cephalometric data, wherein the semantics of the orthodontic stage are set to the state before orthodontics. For the frontal and side photos of the target patient received in S4, completely reproduce the standardized preprocessing process of S2.1 to obtain two standardized 1024×1024 images.

[0109] Based on the cephalometric data of the target patients, structured semantic cues were constructed strictly according to the method in S2.2. A key operation was that the treatment stage label must be manually set in the constructed cues.<post_ortho> This guides the model to generate the post-treatment state.

[0110] 6. Generate two-dimensional therapeutic effect prediction images

[0111] S6. Input the standardized preprocessed facial image of the target patient into the orthodontic efficacy prediction model, and generate the predicted two-dimensional facial images of the target patient's front and side views after orthodontic treatment through local redrawing based on the structured orthodontic semantic prompts constructed in step S5. The specific implementation process of this step is as follows:

[0112] (1) Defining the editing area and generating the mask: On the standardized frontal and side views of the target patient obtained in S5, the doctor or technician uses interactive tools to precisely paint over the soft tissue areas where changes need to be predicted. These areas mainly include: the upper and lower lips, the mentolabial sulcus, the chin, and the jawline contour. The painted area is called the "mask", and the rest of the image remains unchanged.

[0113] (2) Perform local redraw inference: Feed the following inputs into the LoRA model trained by S3: ① a masked image, ② the image constructed by S5 (stage is<post_ortho> ( ) semantic hints, ③ corresponding mask map. The model runs in local redraw mode.

[0114] (3) Multiple Hypothesis Prediction: Multiple possible results are generated by adjusting the "denoising intensity" parameter. Multiple values ​​(such as 0.3, 0.45, 0.6) are selected in the range of 0.2 to 0.8 for inference. The results are more conservative when the intensity is lower, and more significant when the intensity is higher.

[0115] (4) Output and Selection: The model outputs a set (e.g., three pairs) of predicted frontal and lateral images. Clinicians can select the set of predictions that best matches the treatment goals and biomechanical principles based on their professional experience, as the final two-dimensional prediction image. According to the key point error assessment, the average optimal denoising intensity is approximately 0.4255, and 0.4 is often chosen in practice.

[0116] 7. 3D facial reconstruction and geometric consistency optimization

[0117] S7. Based on the predicted two-dimensional facial image of the target patient after orthodontic treatment, perform three-dimensional facial reconstruction, and use the predicted two-dimensional facial image of the side after orthodontic treatment as a geometric constraint. Optimize the geometric consistency of the reconstructed three-dimensional facial model through differentiable rendering technology, and output the predicted three-dimensional facial model after orthodontic treatment.

[0118] After obtaining the patient's frontal and side facial images, this embodiment first constructs an initial 3D facial model based on the frontal image. This embodiment preferably uses the DECA (Detailed Expression Capture and Animation) framework for reconstruction. DECA, through 3DMM parametric regression, detailed texture encoding, and a lighting-based reconstruction module, can recover a complete 3D facial mesh containing mid-to-low frequency bony structures and high-frequency skin details from a single frontal image. This process provides a stable and detailed initial facial topology. However, uncertainties still exist in key lateral structures such as the mandibular border, chin projection, and nasal base retraction in a single frontal view. To further improve the geometric accuracy of the side view, this embodiment introduces differentiable rendering optimization based on nvdiffrast to fully utilize the geometric constraints in the side image. nvdiffrast is a high-performance differentiable rasterization framework from NVIDIA that supports end-to-end backpropagation of the rendering pipeline on the GPU, thereby allowing direct optimization of the vertex positions of the 3D mesh through image errors. Leveraging the differentiable rendering capabilities of nvdiffrast, this embodiment treats the mesh reconstructed by DECA as a learnable parameter, renders a predicted image with the same viewpoint as the real side view photograph, and corrects the side geometry based on the difference between the predicted image and the real side view image.

[0119] The specific process is as follows:

[0120] (1) Initial 3D Reconstruction: Select the final frontal prediction image determined by S6 and input it into the DECA (Detailed Expression Capture and Animation) 3D face reconstruction framework. DECA is based on the 3DMM parametric model and regresses the identity and expression parameters of the face from a single frontal image to generate an initial 3D mesh model.

[0121] (2) Differentiable rendering settings: Since the side profile reconstructed by DECA based solely on the front view is not accurate enough, it needs to be optimized. Using NVIDIA's nvdiffrast differentiable rendering library, the 3D mesh obtained in the previous step is rendered to a camera view that is completely consistent with the final side prediction image generated by S6, thus obtaining the rendered side profile (binary mask).

[0122] (3) Geometric consistency optimization iteration: Define the optimization loss function.

[0123] After obtaining the predicted post-orthodontic image, a 3D morphology of the patient's face is further constructed to verify geometric consistency from multiple perspectives and assist in clinical visualization. Traditional single-view-based inverse face reconstruction methods (such as DECA), while possessing excellent robustness and complete parametric representation capabilities, have inherent limitations: DECA's shape fitting primarily relies on frontal view information, lacking sufficient constraints on the depth structure of the profile. This leads to over-smoothing or incorrect regression of laterally dependent geometric features such as facial contours, chin prominence, and mandibular angle. This prevents the model from accurately reflecting key profile changes such as chin advancement and mandibular retrusion improvement in orthodontic prediction tasks. To overcome this deficiency, this embodiment explicitly introduces a side-view supervision and consistency correction mechanism during the 3D shape reconstruction process. During optimization, based on the initial 3DMM shape parameters of DECA, consistency adjustments are made to the shape parameters with the goal of reducing rendering errors between the frontal predicted image and the side input image. By combining forward-looking appearance constraints and lateral contour constraints in this correction method, the resulting 3D facial model maintains a coherent geometric structure across multiple perspectives. This reflects both the predicted facial change trend and the depth relationship of the actual profile. This integrated strategy significantly improves the lateral geometric accuracy of 3D reconstruction, enabling the model to provide more reliable and stable multi-view prediction results in clinical visualization.

[0124] In the differentiable rendering stage, the following loss can be further used to constrain the mesh update process for the chin local region vertices:

[0125] ① Contour Loss: Calculate the L2 norm distance between the rendered side contour mask and the side prediction contour mask, denoted as... ,in To render the contour mask, Contour mask for side profile prediction

[0126] ② Smoothing Regularization Loss: Calculates the Laplacian coordinate change of the mesh vertices, denoted as... This is to prevent unnatural local distortions caused by optimization. It ensures smooth local facial geometry changes and avoids noise that does not conform to physiological structure.

[0127] The total loss is: Set weight coefficients and By backpropagation, only the vertex coordinates (V) of the 3D mesh are optimized, iterating hundreds of times until the loss converges. This process forces the side profile of the 3D model to align with a high-precision 2D side profile prediction map.

[0128] (4) Output optimized 3D model: After optimization, a high-quality 3D facial mesh is obtained. The frontal shape of the model is generated by DECA based on the frontal prediction map, and the side shape is aligned with the side prediction map by the optimization process, achieving geometric consistency across viewpoints.

[0129] 8. Visualization of 3D prediction results

[0130] S8. Visualize the three-dimensional facial model after orthodontic treatment.

[0131] Specifically, the final 3D prediction model obtained by S7 is imported into 3D visualization software or the system's built-in rendering engine for display. Doctors and patients can interactively rotate and zoom the model to observe the prediction results from any angle.

[0132] In the results presentation section, one optional implementation first involves manually masking the area around the mouth and the region of interest in the original image to generate a local mask. This ensures that the editing scope during the inference stage is strictly limited to semantically relevant local regions. This strategy effectively avoids unnecessary modifications to unlabeled areas by the model, thereby maintaining higher stability in the prediction results in terms of identity consistency and overall composition.

[0133] Secondly, a systematic comparison was conducted on different denoising intensity parameters (0.2–0.8) to evaluate the impact of noise preservation on the model's generative degrees of freedom. When the denoising intensity is low (e.g., 0.2–0.4), the model tends to preserve the details and local structures of the original image, resulting in more stable predictions and a tendency towards slight corrections. However, as the denoising intensity increases (e.g., 0.6–0.8), the model has greater generative degrees of freedom and can make more significant shape adjustments, but it is also more prone to structural shifts, disrupting the consistency of the figures.

[0134] Finally, based on key evaluation, the noise reduction intensity parameter closest to the orthodontic result was selected, and the average value was calculated, which is approximately 0.4255. In this embodiment, the prediction graph with a value of 0.4 was selected for subsequent modeling. In addition, in practical applications, the clinical experience of professional doctors can be incorporated to select the version that best reflects the actual trend from multiple candidate results.

[0135] Example 2

[0136] Please refer to Figure 2This paper presents a lightweight orthodontic efficacy prediction system based on semantics and 3D optimization. This embodiment describes an integrated system for implementing the method of Embodiment 1. The system is deployed on a clinical workstation, and its module structure corresponds completely to the steps of Embodiment 1: It includes: a historical training dataset construction module 1, used to acquire pre- and post-orthodontic frontal and lateral facial images and corresponding cephalometric measurements of historical orthodontic patients to form a training dataset; a first image preprocessing and semantic prompt construction module 2, used to perform standardized preprocessing on the facial images in the training dataset and construct a structured orthodontic semantic prompt for each facial image, including its orthodontic stage, viewpoint attribute, age group, and semantic description based on cephalometric data conversion; an orthodontic efficacy prediction model construction module 3, used to fine-tune the basic model using LoRA technology with a pre-trained diffusion model as the base model, utilizing the standardized preprocessed facial images and their corresponding structured orthodontic semantic prompts in the training dataset; and a target patient image and data acquisition module 4, used to receive pre- orthodontic frontal and lateral facial images and corresponding cephalometric measurements of the target patient. The second image preprocessing and semantic prompting construction module 5 is used to perform standardized preprocessing on the facial image of the target patient in the same way as in step S2, and to construct corresponding structured orthodontic semantic prompts based on its cephalometric data, wherein the semantics of the orthodontic stage are set to the pre-orthodontic state. The two-dimensional orthodontic efficacy prediction module 6 is used to input the standardized preprocessed facial image of the target patient into the orthodontic efficacy prediction model, and, based on the structured orthodontic semantic prompts constructed in step S5, to generate predicted two-dimensional facial images of the target patient from the front and side through local redrawing. The three-dimensional orthodontic efficacy prediction module 7 is used to perform three-dimensional facial reconstruction based on the predicted post-orthodontic two-dimensional facial images of the target patient, and, using the side-view predicted post-orthodontic two-dimensional facial image as geometric constraints, to optimize the geometric consistency of the reconstructed three-dimensional facial model through differentiable rendering technology, and output the predicted post-orthodontic three-dimensional facial model. The visualization module 8 is used to visualize the post-orthodontic three-dimensional facial model.

[0137] Based on the specific embodiments described above, the method and system of the present invention can achieve the following beneficial effects:

[0138] (1) It achieves high-performance prediction with extremely low data and computing power requirements, and has excellent clinical deployment friendliness. Addressing the core bottleneck of massive data and huge computing power required for full parameter fine-tuning of existing diffusion models, this invention introduces LoRA (low-rank adaptive) technology to perform lightweight fine-tuning of pre-trained diffusion models. This method only needs to update a small number of low-rank parameters injected into the attention module, enabling the model to efficiently learn the deformation patterns of orthodontic soft tissue using hundreds of pairs of clinical case data. The training process can be completed quickly on a single consumer-grade GPU, fundamentally breaking down the high computing power barrier for the application of large AI models in medical scenarios, making the deployment and updating of high-performance prediction models on ordinary outpatient workstations a reality, and greatly improving the accessibility and practicality of the technology.

[0139] (2) A semantic control system deeply integrated with clinical knowledge was constructed, making the prediction process highly interpretable and controllable. Addressing the shortcomings of traditional generative models—such as "black box" operation and difficulty in interpreting and guiding results—this invention creatively designed a structured orthodontic semantic prompting system. This system automatically converts professional cephalometric parameters (such as SNA, SNB, and lip protrusion) into semantic descriptions understandable by the model according to clinical standards, and integrates them with information such as treatment stage, perspective, and age. This makes the generation process directly controlled by clinical diagnostic language. Doctors can intuitively explore the predictive effects under different treatment goals by adding, deleting, or modifying semantic elements in the prompts, realizing a paradigm shift from "data-driven" to "semantic-driven," significantly improving the clinical credibility of the prediction results and the ability to interactively customize treatment plans.

[0140] (3) Precise and natural editing of the target soft tissue region was achieved while strictly maintaining the consistency of the patient's identity. Addressing the problem that existing image generation methods easily lead to the drift of global identity features (such as facial features and skin texture) when editing local areas, this invention adopts a local redrawing inference strategy. During prediction, a mask is applied only to the soft tissue regions (such as the perioral and chin) in the input image that need to be changed. Guided by semantic prompts, the model only completes the masked regions. This mechanism ensures that the patient's facial identity features are perfectly preserved before and after prediction, resulting in natural and physiologically consistent image changes. It effectively avoids unreasonable deformations and significantly improves the realism and acceptability of the generated results.

[0141] (4) By optimizing the geometric consistency through side view constraints, the accuracy and consistency of the 2D prediction to 3D reconstruction results across multiple perspectives are ensured. Addressing the challenges of lacking stereoscopic information in pure 2D prediction and inaccurate side profiles in single-view 3D reconstruction, this invention proposes an innovative 2D-3D joint optimization framework. After initial 3D reconstruction using the frontal prediction image, differentiable rendering technology is introduced, using a high-precision side prediction image as a strong constraint target. The geometry of the 3D model is iteratively corrected by optimizing the contour loss function. This process forces the side profile of the 3D model to align with the 2D prediction, thereby outputting a 3D prediction model that maintains a high degree of morphological consistency from the front, side, and other arbitrary perspectives, providing reliable spatial morphological information far exceeding that of planar images.

[0142] (5) A complete digital clinical workflow covering data input to 3D display has been formed, significantly improving diagnostic and treatment efficiency and doctor-patient communication experience. The technical solution of this invention integrates multiple links such as data preprocessing, automated annotation, model reasoning, 3D reconstruction and visualization, providing an end-to-end solution. The system supports "multiple hypothesis" prediction by adjusting parameters to assist doctors in decision-making; the final 3D model can be interactively rotated and viewed, making the display of treatment effects more three-dimensional and intuitive. This not only frees doctors from tedious manual prediction work and improves the efficiency of treatment planning, but also builds an efficient doctor-patient communication bridge through vivid visualization results, which helps to establish consensus, improve patient satisfaction and treatment compliance.

[0143] In summary, this invention achieves breakthroughs in multiple dimensions, including data efficiency, result controllability, identity preservation, cross-perspective consistency, and clinical usability, by organically integrating key technologies such as lightweight adaptation, semantic control, and three-dimensional geometry optimization. It provides a practical and effective intelligent efficacy prediction tool for the field of orthodontics.

[0144] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, such as using different basic diffusion models, adjusting the injection module or rank size of LoRA, adopting other three-dimensional face reconstruction methods, or changing the specific form of the loss function, should all be covered within the scope of protection of the present invention.

Claims

1. A lightweight orthodontic efficacy prediction method based on semantics and 3D optimization, characterized in that, include: S1. Obtain frontal and lateral facial images of historical orthodontic patients before and after orthodontic treatment, along with corresponding cephalometric data, to form a training dataset; S2. Perform standardized preprocessing on the facial images in the training dataset, and construct a structured orthodontic semantic cue for each facial image, which includes its orthodontic stage, perspective attribute, age group, and semantic description based on cephalometric data conversion; S3. Using the pre-trained diffusion model as the base model, and utilizing the standardized preprocessed facial images and their corresponding structured orthodontic semantic cues in the training dataset, the LoRA technique is used to fine-tune the base model to obtain an orthodontic efficacy prediction model. S4. Receive the frontal and lateral facial images of the target patient before orthodontic treatment and the corresponding cephalometric data; S5. Perform standardized preprocessing on the facial image of the target patient in the same way as in step S2, and construct a corresponding structured orthodontic semantic prompt based on its cephalometric data, wherein the semantics of the orthodontic stage are set to the state before orthodontics. S6. Input the standardized preprocessed facial image of the target patient into the orthodontic efficacy prediction model, and generate the predicted two-dimensional facial images of the target patient's front and side views through local redrawing based on the structured orthodontic semantic prompts constructed in step S5. S7. Based on the predicted two-dimensional facial image of the target patient after orthodontic treatment, perform three-dimensional facial reconstruction, and use the predicted two-dimensional facial image of the side after orthodontic treatment as a geometric constraint. Optimize the geometric consistency of the reconstructed three-dimensional facial model through differentiable rendering technology, and output the predicted three-dimensional facial model after orthodontic treatment. S8. Visualize the three-dimensional facial model after orthodontic treatment.

2. The lightweight orthodontic efficacy prediction method based on semantics and 3D optimization according to claim 1, characterized in that, In step S2, the orthodontic stage label includes before and after orthodontic treatment; the viewpoint attributes include frontal view, lateral view, and 45° view; the age group includes adolescents, young adults, adults, and middle-aged people as coarse-grained age group labels; the key measurements of the cephalometric lateral view are selected from the cephalometric measurements of the lateral view, including key indicators related to facial visual features, such as skeletal indicators, dental indicators, and soft tissue indicators; the visual description is specifically facial feature description text generated by an image description model.

3. The lightweight orthodontic efficacy prediction method based on semantics and 3D optimization according to claim 1, characterized in that, In steps S2 and S5, the structured orthodontic semantic prompts further include: automatically generating descriptive text for the visual features of the perioral and mandibular regions in facial images using an image description model.

4. The lightweight orthodontic efficacy prediction method based on semantics and 3D optimization according to claim 1, characterized in that, In step S3, the fine-tuning of the base model using LoRA technology specifically involves injecting a trainable low-rank matrix into the attention module weights of the base model to achieve fine-tuning.

5. The lightweight orthodontic efficacy prediction method based on semantics and 3D optimization according to claim 1, characterized in that, In step S6, the local redrawing specifically involves: applying a mask to the soft tissue region where orthodontic changes are expected on the standardized preprocessed facial image of the target patient; generating a conditional image of the masked region by the orthodontic efficacy prediction model; and keeping the content of the non-masked region unchanged.

6. The lightweight orthodontic efficacy prediction method based on semantics and 3D optimization according to claim 1, characterized in that, In step S7, the step of using the orthodontic predicted 2D facial image of the side as a geometric constraint and optimizing the geometric consistency of the reconstructed 3D facial model through differentiable rendering technology includes: rendering the reconstructed 3D facial model to the same viewpoint as the orthodontic predicted 2D facial image of the side to obtain the rendered side profile; calculating the difference between the rendered side profile and the actual profile of the side image, and iteratively optimizing the vertex positions of the 3D facial model by combining a mesh smoothing regularization term with the goal of minimizing the difference.

7. The lightweight orthodontic efficacy prediction method based on semantics and 3D optimization according to claim 1, characterized in that, In step S3, the structured orthodontic semantic prompts include key clinical features describing facial bony and soft tissue morphology, including one or more of the following: mandibular plane angle, chin projection, degree of lip protrusion, nasolabial angle, and soft tissue E-line offset.

8. The lightweight orthodontic efficacy prediction method based on semantics and 3D optimization according to claim 1, characterized in that, In step S6, the step of generating the image through local redrawing specifically includes: applying a mask to the soft tissue region where orthodontic changes are expected to occur on the standardized preprocessed facial image of the target patient to construct a local missing area; the orthodontic efficacy prediction model, guided by the structured orthodontic semantic cues, performs image completion generation on the masked region while maintaining the identity features of the non-masked region unchanged, thereby obtaining the predicted two-dimensional facial image of the target patient after orthodontic treatment; wherein, the soft tissue region includes at least the upper and lower lips, chin, and mandibular border.

9. The lightweight orthodontic efficacy prediction method based on semantics and 3D optimization according to claim 1, characterized in that, In step S6, multiple candidate post-orthodontic prediction two-dimensional facial images are generated within a preset range by adjusting the denoising sampling intensity, wherein the preset range is 0.2 to 0.

8.

10. A lightweight orthodontic treatment efficacy prediction system based on semantics and 3D optimization, used to implement the method as described in any one of claims 1 to 9, characterized in that, include: The historical training dataset construction module is used to obtain frontal and lateral facial images of historical orthodontic patients before and after orthodontic treatment, as well as corresponding cephalometric data, to form the training dataset. The first image preprocessing and semantic prompting construction module is used to perform standardized preprocessing on the facial images in the training dataset, and to construct a structured orthodontic semantic prompt for each facial image, which includes its orthodontic stage, perspective attribute, age group and semantic description based on cephalometric data conversion. The orthodontic efficacy prediction model construction module is used to fine-tune the basic model using a pre-trained diffusion model as the base model, and using standardized preprocessed facial images and their corresponding structured orthodontic semantic prompts in the training dataset, and employing LoRA technology to obtain the orthodontic efficacy prediction model. The target patient image and data acquisition module is used to receive frontal and lateral facial images of the target patient before orthodontic treatment and corresponding cephalometric measurement data. The second image preprocessing and semantic prompting construction module is used to perform standardized preprocessing on the facial image of the target patient in the same way as in step S2, and to construct corresponding structured orthodontic semantic prompts based on its cephalometric data, wherein the semantics of the orthodontic stage are set to the state before orthodontics. The orthodontic efficacy two-dimensional prediction module is used to input the standardized preprocessed facial image of the target patient into the orthodontic efficacy prediction model, and generate the predicted two-dimensional facial images of the target patient's front and side views through local redrawing based on the structured orthodontic semantic prompts constructed in step S5. The orthodontic efficacy 3D prediction module is used to perform 3D facial reconstruction based on the predicted 2D facial image of the target patient after orthodontics, and to use the predicted 2D facial image of the side after orthodontics as a geometric constraint. Differentiable rendering technology is used to optimize the geometric consistency of the reconstructed 3D facial model and output the predicted 3D facial model after orthodontics. The visualization module is used to visualize the three-dimensional facial model after orthodontic treatment.