Method, device and equipment for generating orthodontic face and storage medium

By performing pixel-level semantic segmentation and tooth alignment prediction on the original facial image, combined with condition-guided feature generation technology, the problems of strong dependence on professional equipment, lack of prior clinical knowledge, and low generation efficiency in existing technologies are solved, and efficient and visually realistic post-orthodontic facial image generation is achieved.

CN121639486APending Publication Date: 2026-03-10SHENZHEN KEVIN PETER TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies in orthodontic clinical practice suffer from several problems, including strong reliance on specialized equipment, lack of prior clinical knowledge, weak ability to preserve details, low generation efficiency, and heavy reliance on data. As a result, the generated results lack medical basis and visual realism.

Method used

By acquiring the original facial images of the target patient and performing pixel-level semantic segmentation, tooth segmentation masks and depth feature maps are extracted. The facial images after orthodontic treatment are generated by combining tooth alignment prediction models and conditionally guided features. Clinical orthodontic knowledge and biomechanical laws are integrated, and a lightweight generation architecture is adopted to improve generation efficiency and visual realism.

Benefits of technology

It enables the automated and efficient generation of post-orthodontic facial images that are both medically plausible and visually realistic from a single 2D photograph, reducing reliance on specialized equipment and improving the clinical credibility and personalized expressiveness of the generated results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121639486A_ABST
    Figure CN121639486A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for generating an orthodontic face, equipment and a storage medium, and relates to the technical field of image processing. The method comprises the following steps: firstly, acquiring an original face image of a target patient, and then performing pixel-level semantic segmentation processing on the original face image to obtain a tooth segmentation mask and a depth feature map corresponding to the tooth segmentation mask; based on the tooth segmentation mask and the depth feature map, obtaining a tooth arrangement structure map after orthodontics and a structure feature map corresponding to the tooth arrangement structure map; and finally, generating an orthodontic face image of the target patient by using the tooth segmentation mask, the depth feature map, the tooth arrangement structure map and the structure feature map, thereby reducing dependence on professional equipment, and improving clinical credibility, personalized expressive force and real-time performance of a generated result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically, to a method, apparatus, device, and storage medium for generating an orthodontic face. Background Technology

[0002] In the clinical diagnosis and treatment of orthodontics, visualizing the post-treatment results is of great significance for doctor-patient communication, confirmation of treatment plans, and improvement of patient compliance.

[0003] Traditional facial design mainly relies on doctors drawing by hand or using simple image editing software (such as Photoshop) for simulation. This process is time-consuming, highly subjective, and makes it difficult to guarantee the consistency and professionalism of the results. Summary of the Invention

[0004] In view of this, the object of the present invention is to provide a method, apparatus, device and storage medium for generating an orthodontic face.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, the present invention provides a method for generating an orthodontic face, the method comprising: Obtain the original facial image of the target patient; The original facial image is subjected to pixel-level semantic segmentation to obtain a tooth segmentation mask and a depth feature map corresponding to the tooth segmentation mask; Based on the tooth segmentation mask and the depth feature map, a tooth arrangement structure map after orthodontic treatment and a structural feature map corresponding to the tooth arrangement structure map are obtained. Using the tooth segmentation mask, the depth feature map, the tooth arrangement structure map, and the structural feature map, a facial image of the target patient after orthodontic treatment is generated.

[0006] Optionally, the step of performing pixel-level semantic segmentation on the original facial image to obtain a tooth segmentation mask and a corresponding depth feature map includes: The original facial image is divided into multiple image blocks, and dynamic feature extraction is performed on the multiple image blocks to obtain multiple feature maps at different scales; The multiple feature maps are fused to obtain the tooth segmentation mask and the corresponding depth feature map.

[0007] Optionally, the step of obtaining the orthodontic tooth alignment structure map and the corresponding structural feature map based on the tooth segmentation mask and the depth feature map includes: The tooth segmentation mask and the depth feature map are input into the second network of the pre-trained tooth alignment prediction model. The second network is used to mimic the label distribution output by the first network of the tooth alignment prediction model to obtain the orthodontic tooth alignment structure map and the corresponding structural feature map. The first network was trained based on the point cloud data of dental molds of sample patients before and after orthodontic treatment.

[0008] Optionally, the step of generating the facial image of the target patient after orthodontic treatment using the tooth segmentation mask, the depth feature map, the tooth arrangement structure map, and the structural feature map includes: Based on the tooth segmentation mask, the depth feature map, the tooth arrangement structure map, and the structural feature map, conditional guided features are constructed; Based on the condition-guided features, a preliminary image is obtained; The preliminary image is enhanced with the depth feature map to obtain the facial image of the target patient after orthodontic treatment.

[0009] Optionally, the step of constructing condition-guided features based on the tooth segmentation mask, the depth feature map, the tooth arrangement structure map, and the structural feature map includes: Using the tooth arrangement structure diagram and the structural feature diagram, a first vector is constructed; A second vector is constructed using the background features corresponding to the original facial image, the tooth segmentation mask, and the depth feature map; Generate a weighted value vector based on the first vector and the second vector; The value vector is weighted and aggregated based on the weights to obtain the condition-guided features.

[0010] Optionally, the step of using the depth feature map to perform detail enhancement processing on the preliminary image to obtain the facial image of the target patient after orthodontic treatment includes: The preliminary image and the depth feature map are fused together using skip connections to obtain a fused image; The fused image is then subjected to color correction and edge smoothing to obtain the facial image of the target patient after orthodontic treatment.

[0011] Secondly, the present invention provides an apparatus for generating an orthodontic face, the apparatus comprising: The acquisition module is used to acquire the original facial image of the target patient; The processing module is used to perform pixel-level semantic segmentation processing on the original facial image to obtain a tooth segmentation mask and a depth feature map corresponding to the tooth segmentation mask; based on the tooth segmentation mask and the depth feature map, to obtain a tooth alignment structure map after orthodontic treatment and a structural feature map corresponding to the tooth alignment structure map; and to generate a facial image of the target patient after orthodontic treatment using the tooth segmentation mask, the depth feature map, the tooth alignment structure map, and the structural feature map.

[0012] Optionally, the processing module is specifically used to divide the original facial image into multiple image blocks, and to perform dynamic feature extraction on the multiple image blocks to obtain multiple feature maps of different scales; and to perform fusion processing on the multiple feature maps to obtain the tooth segmentation mask and the depth feature map corresponding to the tooth segmentation mask.

[0013] Thirdly, the present invention provides an electronic device including a processor and a memory, the memory storing machine-executable instructions executable by the processor, the processor executing the machine-executable instructions to implement the method for generating an orthodontic face as described in the first aspect above.

[0014] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating an orthodontic face as described in the first aspect above.

[0015] The method, apparatus, device, and storage medium for generating post-orthodontic facial images provided in this invention first acquire the original facial image of the target patient, then perform pixel-level semantic segmentation on the original facial image to obtain a tooth segmentation mask and a corresponding depth feature map; next, based on the tooth segmentation mask and the depth feature map, obtain a post-orthodontic tooth alignment structure map and a corresponding structural feature map; finally, use the tooth segmentation mask, depth feature map, tooth alignment structure map, and structural feature map to generate the post-orthodontic facial image of the target patient. Because this invention acquires the original facial image of the target patient and performs pixel-level semantic segmentation, accurately extracting the segmentation mask and deep features of the tooth region, and then integrating clinical orthodontic knowledge to predict the post-orthodontic tooth alignment structure that conforms to biomechanical principles, and combining multimodal features to guide the generation of highly realistic post-orthodontic facial images, it achieves automated and efficient generation of smile designs that combine medical rationality and visual realism using only a single 2D photograph, thereby reducing reliance on specialized equipment and improving the clinical credibility, personalized expressiveness, and real-time performance of the generated results.

[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This figure shows a schematic block diagram of an electronic device provided by an embodiment of the present invention; Figure 2 A flowchart illustrating a method for generating an orthodontic face according to an embodiment of the present invention is shown; Figure 3 This invention provides a comparison image of the face before and after orthodontic treatment. Figure 4 The diagram shows a functional block diagram of a device for generating an orthodontic face according to an embodiment of the present invention.

[0019] Icons: 100 - Electronic device; 110 - Memory; 120 - Processor; 130 - Communication module; 200 - Device for generating the orthodontic face; 201 - Acquisition module; 202 - Processing module. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0021] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0022] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0023] In the field of orthodontics, predicting aesthetic outcomes before treatment is a crucial step in doctor-patient communication and treatment plan development. With the development of computer vision and artificial intelligence technologies, image-based digital smile design (DSD) has gradually become a research hotspot. Currently, mainstream smile design techniques can be mainly divided into three categories: 3D model-based methods, pure 2D image generation methods, and traditional image processing and template replacement methods.

[0024] Category 1: Smile design method based on 3D model.

[0025] This type of method relies on high-precision intraoral scanning equipment to acquire a 3D point cloud or mesh model of the patient's teeth. Virtual orthodontic simulation is then performed using specialized software (such as iOrthoPredictor and 3Shape Smile Design), and the final facial smile image is rendered. This method can accurately reproduce the spatial alignment of teeth and has high clinical accuracy. However, its application is limited by expensive specialized acquisition equipment and complex operating procedures, making it difficult to widely use in general outpatient clinics or consumer-level scenarios (such as mobile phone photography). Furthermore, the 3D-to-2D rendering process still requires significant manual intervention to match appearance parameters such as skin tone and lighting, resulting in a low degree of automation.

[0026] The second category: pure 2D image generation methods based on deep learning.

[0027] In recent years, some studies have attempted to synthesize end-to-end smile design images directly from two-dimensional facial photographs using Generative Adversarial Networks (GANs) or diffusion models. For example, methods like OrthoAligner employ an encoder-decoder structure, training the model on paired datasets to learn the mapping relationship between "before treatment" and "after treatment." These methods do not require 3D input, lowering the hardware barrier and showing promising practical potential. However, they are essentially data-driven black-box modeling, lacking effective integration of oral biomechanics principles and clinical orthodontic knowledge. Because the training data often comes from aesthetic standards rather than actual orthodontic trajectories, the generated results often only satisfy visual aesthetics, ignoring crucial medical constraints such as the rationality of the dental arch morphology and the maintenance of interdental spaces, posing a risk of misleading patients.

[0028] The third category: traditional methods based on image processing and template replacement.

[0029] These methods typically involve first extracting the original tooth region through image segmentation, then deforming and aligning a pre-defined ideal tooth template before pasting it back into the original image. While simple to implement and computationally inexpensive, their generation mechanism is rigid and cannot adapt to significant individual differences in facial structure and lighting conditions. They are particularly poor at texture blending, often exhibiting unnatural boundaries, color distortion, and inconsistent highlights, resulting in unrealistic generated images that severely impact user experience and clinical credibility.

[0030] In summary, existing technologies generally suffer from the following common problems: 1. High dependence on specialized equipment: Most high-precision solutions rely on 3D scanners, which limits their application in low-resource environments; 2. Lack of prior clinical knowledge: The 2D generated model failed to effectively incorporate real tooth movement patterns and dental arch physiological characteristics, resulting in a lack of medical basis for the generated results; 3. Weak ability to preserve details: Traditional U-Net architecture is prone to losing edge information during downsampling, resulting in blurred tooth contours and unclear textures; 4. Low generation efficiency: The diffusion model-based method requires hundreds of iterations for denoising, which is difficult to meet the requirements of real-time interaction; 5. Heavy reliance on data: Supervised learning paradigms require large-scale paired "pre- and post-treatment" image data, which is costly to obtain and difficult to annotate in clinical practice.

[0031] To overcome the shortcomings of the prior art, embodiments of the present invention provide a method, apparatus, device, and storage medium for generating an orthodontic face, which will be described in detail below.

[0032] Please refer to Figure 1This is a block diagram of an electronic device 100. The electronic device 100 includes a memory 110, a processor 120, and a communication module 130. The memory 110, processor 120, and communication module 130 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0033] The memory 110 is used to store programs or data. The memory 110 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0034] The processor 120 is used to read / write data or programs stored in the memory 110 and to perform corresponding functions.

[0035] The communication module 130 is used to establish a communication connection between the electronic device 100 and other communication terminals through the network, and to send and receive data through the network.

[0036] It should be understood that, Figure 1 The structure shown is only a schematic diagram of the electronic device 100. The electronic device 100 may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.

[0037] Please refer to Figure 2 The method for generating the orthodontic face includes steps S101 to S104.

[0038] S101, Obtain the original facial image of the target patient.

[0039] In this embodiment of the invention, the original facial image of the target patient is first acquired. The original facial image is a single two-dimensional (2D) frontal or oblique side view photograph, which can be a high-definition digital image of a natural smiling state, with a resolution of not less than 1920×1080 pixels, uniform lighting, and no obvious occlusion or reflection. This image can be obtained by taking pictures with a smartphone, tablet device, or ordinary digital camera, without relying on a professional intraoral scanner or other three-dimensional imaging equipment, significantly reducing the threshold for use and improving the universality and scalability of the method.

[0040] The original facial images are stored in digital format on a local terminal or cloud server as the basic input data for subsequent image processing and model inference.

[0041] S102, perform pixel-level semantic segmentation on the original facial image to obtain a tooth segmentation mask and a corresponding depth feature map.

[0042] This step aims to accurately extract key anatomical structures of the oral cavity region from the original facial image, especially the boundary information of teeth and gums, to provide precise spatial localization and contextual feature support for subsequent structure alignment and image generation.

[0043] In a possible implementation, step S102 may include sub-steps S102-1 to S102-2.

[0044] S102-1 divides the original facial image into multiple image blocks and performs dynamic feature extraction on multiple image blocks to obtain multiple feature maps at different scales.

[0045] Original facial image The image is divided into several patches of a preset size (e.g., 16×16 pixels) and then converted into a one-dimensional feature sequence through a linear embedding layer. ,in For sequence length, Image patch size, This is the embedding dimension. The feature sequence is input into a state-space model-based Visual Mamba backbone network for encoding.

[0046] Visual Mamba employs a bidirectional scanning mechanism, traversing the image patch sequence from front to back and from back to front respectively, capturing long-distance dependencies during the hidden state update process. Its core state transition equation is expressed as follows:

[0047]

[0048] in, The parameter matrix is ​​the discretized matrix. This is the hidden state. Bidirectional scanning processes the sequence from front to back and from back to front respectively, ultimately fusing the output features from both directions. This enables global receptive field coverage and effectively models the long-distance dependencies between teeth and facial organs.

[0049] Furthermore, a Deformable Attention Module (DAM) is introduced to enhance the ability to focus on irregular edge regions at key levels. This module adaptively adjusts the sampling position through deformable convolution, focusing on high-frequency detail regions such as tooth contours; subsequently, a lightweight attention mechanism is used to calculate the correlation weights between any two pixels. :

[0050] in, is the feature vector, and Sim is the similarity function (such as the dot product). Weighted aggregation further enhances the representation ability of key regions. S102-2, multiple feature maps are fused to obtain a tooth segmentation mask and the corresponding depth feature map.

[0051] The constructed Pyramid Feature Fusion Network (PFFN) is used to fuse the aforementioned multi-scale feature maps step by step. PFFN restores the deep high-semantic feature maps to their original resolution through upsampling operations, and then adds or concatenates them with shallow high-resolution feature maps element by element through skip connections, thereby achieving effective integration of semantic information and spatial details.

[0052] The fused features are fed into the segmentation head, and the Softmax activation function outputs the probability distribution of the category to which each pixel belongs, generating pixel-level classification results. Finally, a binary or multi-class tooth segmentation mask is output, clearly identifying the regions of interest (ROIs) such as upper / lower teeth, gums, and lips.

[0053] That is, the multi-scale feature maps generated by the encoder at different depths. The data are simultaneously input into the PFFN, which uses upsampling and skip connections to progressively fuse deep semantic information with shallow spatial details, ultimately outputting a high-precision segmentation mask. This process can be represented as:

[0054] Meanwhile, the high-level feature map output from the encoder end is retained as a depth feature map. This feature map contains rich texture, lighting and skin color information, which will be used for realistic reconstruction in the subsequent generation stage.

[0055] S103, based on tooth segmentation mask and depth feature map, obtains the tooth arrangement structure map after orthodontic treatment and the corresponding structural feature map.

[0056] In a possible implementation, step S103 can be performed as follows: The tooth segmentation mask and depth feature map are input into the second network of a pre-trained tooth alignment prediction model. The second network then mimics the label distribution output by the first network of the tooth alignment prediction model to obtain the post-orthodontic tooth alignment structure map and the corresponding structural feature map. The first network is trained based on the point cloud data of the dental models of the sample patients before and after orthodontic treatment.

[0057] In other words, the tooth segmentation mask and depth feature map obtained in the aforementioned steps are input together into the second network (i.e., the student network) in the pre-trained tooth alignment prediction model. This network imitates the behavior of the first network (i.e., the teacher network) and outputs the orthodontic tooth alignment structure map and its corresponding structural feature map.

[0058] The first network is a 3D tooth alignment prediction model based on Diffusion Transformer (DiT), and its training data consists of point cloud data of dental models before and after orthodontic treatment of sample patients. This model learns clinical prior knowledge such as real tooth movement trajectories, changes in arch curvature, and occlusal relationship adjustments, and can output the ideally aligned 3D dental arch structure.

[0059] The second network, as a lightweight 2D mapping network, employs a knowledge distillation framework during the training phase, using the output distribution of the first network as "soft labels" to guide its learning of 3D correction rules. Its loss function is defined as:

[0060] in, This is the loss from the student network's own tasks (such as 2D keypoint prediction). It is the KL divergence, used to measure the probability distribution of student network outputs. Soft tags output by teachers' networks The difference between them. By minimizing Students can learn the tooth movement patterns in 3D orthodontic treatment plans without needing paired 3D-2D data.

[0061] In addition, to improve the model's generalization ability, a self-supervised contrastive learning strategy was introduced: positive sample pairs were generated by performing data augmentation operations such as rotation, cropping, and color jitter on the same patient images. This maximizes the similarity of the teeth in the feature space while minimizing the similarity with other patient samples, thereby enhancing the model's understanding of tooth alignment semantics.

[0062] In other words, the model learns to bring different augmented views of the same patient closer together in the feature space, while simultaneously pushing features from different patients further apart. Its contrast loss can be simplified as:

[0063] This enables the model to autonomously learn semantic features of tooth alignment from unlabeled data.

[0064] During the inference phase, the second network receives input from S102 and directly outputs a post-orthodontic tooth alignment structure map reflecting the ideal alignment state. This can be represented as a heatmap showing the target position of each tooth, or as a contour line depicting the new dental arch curve. Simultaneously, it outputs the corresponding structural feature map as a conditional guidance signal for subsequent generation modules.

[0065] To further ensure clinical rationality, differentiable geometric constraint loss functions can be introduced, including: arch smoothness constraint (penalizing the degree to which the predicted arch deviates from the ideal parabolic model) and adjacent tooth spacing consistency constraint (ensuring that physiological gaps are maintained between adjacent teeth to prevent penetration or excessive separation).

[0066] Geometric constraint loss function Satisfy the following formula:

[0067] This constraint, added as a regularization term to the total loss function, enhances the clinical rationality of the generated results.

[0068] These constraints are added as regularization terms to the total loss function during the training phase, making the prediction results closer to the actual correction logic.

[0069] S104 uses tooth segmentation mask, depth feature map, tooth arrangement structure map and structural feature map to generate facial image of the target patient after orthodontic treatment.

[0070] In a possible implementation, step S104 may include sub-steps S104-1 to S104-3.

[0071] S104-1 constructs condition-guided features based on tooth segmentation masks, depth feature maps, tooth arrangement structure maps, and structural feature maps.

[0072] In this embodiment of the invention, a first vector can be constructed using a tooth arrangement structure map and a structural feature map; a second vector can be constructed using background features, a tooth segmentation mask, and a depth feature map corresponding to the original facial image; a weighted pair vector can be generated based on the first and second vectors; and a weighted aggregation process can be performed on the weighted pair vector to obtain the conditional guided features.

[0073] In other words, firstly, a first vector (Query vector) is constructed using the tooth arrangement structure map and structural feature map to represent the position and shape of the target tooth; then, a second vector (Key and Value vector) is constructed using the background features extracted from the original facial image, the tooth segmentation mask, and the depth feature map to represent the skin color, lighting, and texture characteristics of the current oral environment.

[0074] Then, the correlation weights between the two are calculated using a cross-modal attention mechanism:

[0075]

[0076] in, Q It is a query vector for tooth structure features. K , V These are key-value pairs of features from the original image. This allows the generation process to precisely "implant" the new tooth into the original oral environment.

[0077] After weighted aggregation, a conditional guidance feature rich in semantic alignment information is obtained, which serves as the core control signal of the generator.

[0078] S104-2, Obtain a preliminary image based on condition-guided features.

[0079] A Consistency Model (CM) is adopted as the generative architecture to replace the traditional multi-step diffusion model. The Consistency Model learns to directly map any noisy samples on the Probabilistic Flow Ordinary Differential Equation (PFODE) trajectory back to the original clean data, supporting one-step or multi-step high-quality image generation. Its core loss function is:

[0080] in, d It is a distance metric function. It is a Consistency Model. and These are noisy samples from two different time steps on the same PFODE trajectory. This design improves inference speed by 5-10 times, meeting the needs of real-time applications.

[0081] At this stage, a preliminary low- or medium-resolution image is generated, which contains the aligned tooth shape and general color, but the details are not yet refined enough.

[0082] S104-3 uses depth feature maps to perform detail enhancement processing on the preliminary image to obtain the facial image of the target patient after orthodontic treatment.

[0083] To further restore the microtexture, gloss reflection, and edge sharpness of the tooth surface, a progressive refinement generation strategy is adopted.

[0084] In this embodiment of the invention, the preliminary image and the depth feature map can be fused through skip connections to obtain a fused image; the fused image is then subjected to color correction and edge smoothing to obtain the facial image of the target patient after orthodontic treatment.

[0085] In other words, the preliminary image output by S104-2 and the high-resolution depth feature map output by S102 are fused through skip connections to form a multi-scale input tensor; this tensor is then fed into a lightweight refiner network, which consists of several residual blocks and focuses on local detail restoration.

[0086] Furthermore, the following post-processing operations are performed on the fused image: color correction, which adjusts the color distribution of the new tooth area based on the original facial skin tone mean and variance to make it consistent with the surrounding tissue; edge smoothing, which uses an edge-aware filter (such as a guided filter) to eliminate artifacts at the synthesis boundary and improve the naturalness of the transition.

[0087] The final output is a high-fidelity image of the target patient's face after orthodontic treatment, such as... Figure 3 As shown, the teeth are neatly and beautifully arranged, the skin tone is naturally illuminated, and the texture is realistic and delicate, making them suitable for clinical demonstrations and doctor-patient communication.

[0088] In summary, this invention, through the construction of a three-level cascaded architecture of "segmentation-alignment-generation," achieves fully automated generation from a single 2D facial photograph to a clinically reasonable and visually realistic orthodontic effect image. The entire process requires no 3D scanning equipment, integrates real-world orthodontic knowledge, and balances generation efficiency, medical compliance, and visual quality, demonstrating significant practical value and industrialization prospects.

[0089] To perform the corresponding steps in the above embodiments and various possible methods, an implementation of a device 200 for generating an orthodontically corrected face is given below. Further, please refer to... Figure 4 , Figure 4 This is a functional block diagram of a device 200 for generating an orthodontic face according to an embodiment of the present invention. It should be noted that the basic principle and technical effects of the device 200 for generating an orthodontic face provided in this embodiment are the same as those in the above embodiments. For the sake of brevity, any parts not mentioned in this embodiment can be referred to the corresponding content in the above embodiments. The device 200 for generating an orthodontic face includes: The acquisition module 201 is used to acquire the original facial image of the target patient.

[0090] The processing module 202 is used to perform pixel-level semantic segmentation processing on the original facial image to obtain a tooth segmentation mask and a depth feature map corresponding to the tooth segmentation mask; based on the tooth segmentation mask and the depth feature map, to obtain a tooth alignment structure map after orthodontic treatment and a structural feature map corresponding to the tooth alignment structure map; and to generate a facial image of the target patient after orthodontic treatment using the tooth segmentation mask, the depth feature map, the tooth alignment structure map and the structural feature map.

[0091] Optionally, the processing module 202 is specifically used to divide the original facial image into multiple image blocks, and to perform dynamic feature extraction on the multiple image blocks to obtain multiple feature maps at different scales; and to perform fusion processing on the multiple feature maps to obtain a tooth segmentation mask and a depth feature map corresponding to the tooth segmentation mask.

[0092] Optionally, the above modules can be stored in the form of software or firmware. Figure 1 The memory 110 shown is either stored in or embedded in the operating system (OS) of the electronic device 100, and can be used by... Figure 1 The processor 120 executes the program. Meanwhile, the data and program code required to execute the above modules can be stored in the memory 110.

[0093] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0094] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0095] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0096] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method of generating a post-orthodontic face, characterized by, The method comprises: obtaining an original facial image of a target patient; performing pixel-level semantic segmentation processing on the original facial image to obtain a tooth segmentation mask and a depth feature map corresponding to the tooth segmentation mask; based on the tooth segmentation mask and the depth feature map, obtaining a tooth arrangement structure diagram after orthodontics and a structure feature map corresponding to the tooth arrangement structure diagram; using the tooth segmentation mask, the depth feature map, the tooth arrangement structure diagram and the structure feature map, generating a facial image of the target patient after orthodontics.

2. The method of generating a post-orthodontic face of claim 1, wherein, The step of performing pixel-level semantic segmentation processing on the original facial image to obtain a tooth segmentation mask and a depth feature map corresponding to the tooth segmentation mask comprises: dividing the original facial image into a plurality of image blocks, and performing dynamic feature extraction on the plurality of image blocks to obtain a plurality of feature maps of different scales; fuse the plurality of feature maps to obtain the tooth segmentation mask and the depth feature map corresponding to the tooth segmentation mask.

3. The method of generating a post-orthodontic face of claim 1, wherein, The step of obtaining a tooth arrangement structure diagram after orthodontics and a structure feature map corresponding to the tooth arrangement structure diagram based on the tooth segmentation mask and the depth feature map comprises: input the tooth segmentation mask and the depth feature map into a second network of a pre-trained tooth arrangement prediction model, use the second network to imitate the label distribution output by a first network of the tooth arrangement prediction model, and obtain a tooth arrangement structure diagram after orthodontics and a structure feature map corresponding to the tooth arrangement structure diagram; wherein the first network is trained based on the dental model point cloud data of the sample patient before and after orthodontics.

4. The method of generating a post-orthodontic face of claim 1, wherein, The step of using the tooth segmentation mask, the depth feature map, the tooth arrangement structure diagram and the structure feature map to generate a facial image of the target patient after orthodontics comprises: based on the tooth segmentation mask, the depth feature map, the tooth arrangement structure diagram and the structure feature map, constructing a conditional guide feature; obtaining a preliminary image according to the conditional guide feature; using the depth feature map to perform detail enhancement processing on the preliminary image to obtain the facial image of the target patient after orthodontics.

5. The method of generating a post-orthodontic face of claim 4, wherein, The step of constructing a conditional guide feature based on the tooth segmentation mask, the depth feature map, the tooth arrangement structure diagram and the structure feature map comprises: constructing a first vector using the tooth arrangement structure diagram and the structure feature map; constructing a second vector using the background feature corresponding to the original facial image, the tooth segmentation mask and the depth feature map; generating a weight pair value vector according to the first vector and the second vector; based on the weight pair value vector, performing weighted aggregation processing to obtain the conditional guide feature.

6. The method of generating a post-orthodontic face of claim 4, wherein, The step of using the depth feature map to perform detail enhancement processing on the preliminary image to obtain the facial image of the target patient after orthodontics comprises: fuse the preliminary image and the depth feature map through a skip connection to obtain a fused image; performing color correction processing and edge smoothing processing on the fused image to obtain the facial image of the target patient after orthodontics.

7. An apparatus for generating a post-orthodontic face, characterized by The device comprises: An acquisition module is configured to acquire an original facial image of a target patient; A processing module is configured to perform pixel-level semantic segmentation on the original facial image to obtain a tooth segmentation mask and a depth feature map corresponding to the tooth segmentation mask; based on the tooth segmentation mask and the depth feature map, obtain a tooth arrangement structure diagram after orthodontics and a structure feature map corresponding to the tooth arrangement structure diagram; and generate a facial image of the target patient after orthodontics by using the tooth segmentation mask, the depth feature map, the tooth arrangement structure diagram and the structure feature map.

8. The orthodontic posterior surface generating apparatus of claim 7, wherein, The processing module is specifically configured to divide the original facial image into a plurality of image blocks, and perform dynamic feature extraction on the plurality of image blocks to obtain a plurality of feature maps of different scales; and perform fusion processing on the plurality of feature maps to obtain the tooth segmentation mask and the depth feature map corresponding to the tooth segmentation mask.

9. An electronic device, comprising: A processor and a memory are included, the memory stores machine executable instructions which can be executed by the processor, and the processor can execute the machine executable instructions to implement the method for generating a facial image after orthodontics according to any one of claims 1-6.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by a processor, implements the method for generating a facial image after orthodontics according to any one of claims 1-6.