A high-definition tooth restoration method and device based on a GAN network
Through a high-definition tooth restoration method based on a GAN network, facial key point extraction and mask generation are utilized, combined with a coding and decoding structure and a fine-grained feature fusion module, the problems of tooth clarity and lip synchronization are solved, high-definition tooth generation and synchronization are achieved, and the quality and naturalness of voice-driven facial generation are improved.
Patent Information
- Application Number
- CN202411359898.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-09-27
AI Technical Summary
Existing speech-driven facial generation technology has deficiencies in the clarity and lip synchronization of the dental area, making it difficult to generate high-definition teeth and maintain high-precision synchronization with speech.
A high-definition tooth restoration method based on the GAN network is adopted. Through facial key point extraction and mask generation, combined with the encoding and decoding structure and fine-grained feature fusion module, the multi-loss function optimization model is used to generate high-definition teeth, and the generalization ability and synchronization of the model are improved through two-stage training.
The visual clarity and realism of the dental area are significantly improved, while maintaining the synchronization between lip shape and speech, improving the naturalness and smoothness of animation and user experience, adapting to different facial features and expressions, and having personalized customization capabilities and robustness.
Smart Images

Figure CN119251103B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of face image processing and enhancement combination, and particularly relates to a high-definition tooth repair method and device based on a GAN network. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, the technology of voice-driven face generation (TFG) has become an important research direction in the field of human-computer interaction. The TFG technology aims to generate facial animations synchronized with speech by analyzing audio signals and combining facial features, and is widely used in virtual assistants, video conferences, games, and film production scenarios. This technology can significantly improve the realism and affinity of computer-generated characters by simulating human facial expressions and lip movements.
[0003] Traditional digital human production methods first require complex 3D modeling, and then drive the movement of digital humans through motion capture or keyframe animation. This method not only takes a long time and costs a lot, but also is difficult to adapt to the rapidly changing market demand, especially in production environments that require a large amount of personalized content. In addition, traditional methods also have limitations in handling the naturalness of lip synchronization and facial expressions, often failing to achieve satisfactory results.
[0004] In recent years, with the development of deep learning technologies such as generative adversarial networks (GANs) and autoregressive models, TFG methods based on generative models have begun to emerge. These methods can automatically learn the rules of facial movements from a large amount of facial video and corresponding audio data, and then generate realistic facial animations. However, existing TFG methods still face challenges in generating high-quality visual content, especially in the clarity and detail performance of the teeth area.
[0005] Teeth, as an important part of the face, their clarity is crucial to improving the realism of face generation. However, in existing TFG research, the generation effect of the teeth area is often unsatisfactory, with problems such as blurring, distortion, or inconsistency with real tooth features. This is mainly due to the difficulty of capturing high-frequency detail information in the teeth area in facial images, combined with the complexity of dynamic changes in the teeth area during facial movements, making it difficult for existing models to accurately predict and generate high-quality tooth details.
[0006] In addition, the existing TFG method also has shortcomings in processing lip synchronization. The naturalness and realism of facial animation not only depend on the accuracy of facial expressions, but also rely on the precise matching of lip shapes and speech. However, in practical applications, due to errors in audio signal processing and facial feature extraction, as well as the limited learning ability of the model for complex speech information, the generated lip shapes often do not synchronize with the input speech.
[0007] To solve the above problems, researchers have tried various improvement strategies, such as introducing multi-scale feature fusion, using attention mechanisms, and optimizing model structures. These methods have improved the quality of facial generation to some extent, but there is still room for improvement in tooth clarity and lip synchronization. Therefore, developing a TFG method that can generate high-definition teeth and maintain lip synchronization is of great significance for promoting the development of virtual digital human technology. SUMMARY
[0008] In view of the above, the purpose of the present application is to provide a high-definition tooth repair method and device based on GAN network, which effectively improves the generation quality of TFG method in tooth area and ensures high-precision synchronization of lip shape and speech, to meet the demand of virtual reality, augmented reality, film and television production and other applications for high-quality facial generation.
[0009] To achieve the above-mentioned purpose of the application, the high-definition tooth repair method based on GAN network provided by the embodiment comprises the following steps:
[0010] The acquired facial image is subjected to facial key point extraction, the tooth contour and tooth mask are determined based on the facial key points, the tooth contour and tooth mask image are superimposed to obtain a tooth to-be-repaired image, the mouth region is segmented from the facial image as a reference image, and after image enhancement and fine processing of the mouth region reference image, the tooth to-be-repaired image and the mouth region reference image are subjected to normalization processing;
[0011] A high-definition tooth repair model is constructed based on a GAN network, wherein the generator contained in the GAN network adopts an encoding structure and a decoding structure, the encoding structure adopts two branches to extract and encode features of the normalized tooth to-be-repaired image and the mouth region reference image respectively, the decoding structure decodes the encoding results of the two branches to perform tooth repair and obtain a repaired tooth image, the discriminator contained in the GAN network is used to determine the authenticity of the repaired tooth image, and the generator after two-stage training serves as the high-definition tooth repair model;
[0012] High-definition tooth repair is performed using the high-definition tooth repair model.
[0013] Preferably, the MobileNetV3 network is used for face key point extraction, and the extracted face key points include mouth contour points and related feature points in the tooth region. The tooth contour and tooth mask are determined based on the tooth contour points, and the mouth region reference image M is segmented from the face image reference .
[0014] Preferably, when determining the tooth contour, first generate a tooth contour image M based on the face key point set P facial by a contour generation function C tooth_contour :
[0015] M tooth_contour =C(P facial )
[0016] Mask the face key point set P facial by a mask generation function M to obtain an initial tooth mask image M mask :
[0017] M mask =M(P facial )
[0018] XOR operation is performed on the initial tooth mask image and the mouth region reference image, so that the tooth region of the mouth region reference image forms a mask, and a final tooth mask image M is obtained
[0019] M tooth_mask =U(M mask ,M reference )
[0020] Wherein, U represents image XOR operation.
[0021] The tooth contour image and the tooth mask image are superimposed to obtain a tooth to be repaired image M tooth .
[0022] Preferably, the mouth region reference image is subjected to image enhancement and refinement processing, and the tooth to be repaired image and the mouth region reference image are subjected to normalization processing. When image enhancement is performed, contrast enhancement and sharpening processing are performed. When refinement processing is performed, edge sharpening and smoothing processing are performed. When normalization is performed, the pixel value of the image is normalized to [0, 1] or [-1, 1].
[0023] Preferably, the two branch structures adopted by the encoding structure are the same, each branch structure includes a feature extraction module and a fine-grained feature fusion module, two models are used for feature extraction and fine-grained feature fusion to realize the encoding process, and the fine-grained feature fusion module specifically includes channel fusion, HourGlass network, feature fusion, and feature refinement, wherein the channel fusion refers to gradually reducing the channel number of the feature map through convolution operation, while increasing the depth of the feature to extract finer features, the HourGlass network is used to extract multi-scale features, the feature fusion refers to fusing the output result of the channel fusion with the extracted multi-scale features, and the feature refinement refers to enhancing the details of the feature fusion result through convolution and activation operations as the encoding result.
[0024] Preferably, each fine-grained feature fusion module includes a convolution layer, a max-pooling layer, three channel fusions, an HourGlass network, a channel fusion, and feature refinement connected in sequence,
[0025] wherein the HourGlass network adopts two branches, one branch extracts fine features through channel fusion, and the other branch extracts multi-scale features through a combination of two pairs of max-pooling and channel fusion in sequence and upsampling, and the fine features and the multi-scale features are fused through feature fusion to obtain a feature fusion result;
[0026] The feature refinement adopts at least one depth separable convolution or a residual connection constructed between at least one depth separable convolution.
[0027] Preferably, the loss function L used when training the GAN network total includes a generator adversarial loss a generator reconstruction loss a generator perceptual loss L perc , a discriminator discriminative loss L D , and a discriminator R1 loss L R1 ,
[0028]
[0029] wherein λ1, λ2, λ3, and λ4 represent loss weights, λ is a hyperparameter for controlling the intensity of the R1 loss, is sampled from a distribution between a tooth reference image and a restored tooth image , is a discrimination result of the discriminator on , is a gradient of , ||·||2 represents an L2 norm for calculating the size of a gradient vector, E represents an expected value, and is calculated by sampling .
[0030] To achieve the above-mentioned object of the application, the embodiment of the application provides a high-definition tooth restoration device based on a GAN network, which comprises an image processing module, a model construction module and a tooth restoration module,
[0031] The image processing module is used for performing face key point extraction on an obtained face image, determining a tooth contour and a tooth mask based on the face key points, superimposing the tooth contour and the tooth mask image to obtain a tooth image to be restored, segmenting a mouth region from the face image as a reference image, and performing image enhancement and fine processing on the mouth region reference image, and then performing normalization processing on the tooth image to be restored and the mouth region reference image.
[0032] The model construction module is used for constructing a high-definition tooth restoration model based on a GAN network, wherein the generator contained in the GAN network adopts an encoding structure and a decoding structure, the encoding structure adopts two branches to perform feature extraction and encoding on the normalized tooth image to be restored and the mouth region reference image respectively, the decoding structure decodes the encoding results of the two branches to perform tooth restoration and obtain a restored tooth image, the discriminator contained in the GAN network is used for discriminating the authenticity of the restored tooth image, and the generator after two-stage training is used as the high-definition tooth restoration model.
[0033] The tooth restoration module is used for performing high-definition tooth restoration by using the high-definition tooth restoration model.
[0034] To achieve the above-mentioned object of the application, the embodiment further provides a computing device comprising a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the above-mentioned high-definition tooth restoration method based on the GAN network.
[0035] To achieve the above-mentioned object of the application, the embodiment further provides a computer readable storage medium having a program stored thereon, wherein the program is executed by a processor to implement the above-mentioned high-definition tooth restoration method based on the GAN network.
[0036] Compared with the prior art, the application has at least the following beneficial effects:
[0037] A more accurate and efficient face key point detection method: through the key point detection method of the MobileNetV3 network architecture, the key points of the face can be detected with high precision, real-time face tracking can be realized on ordinary consumer-grade hardware, and the method is suitable for application scenarios that require fast response. It is suitable for different lighting conditions and facial expressions, and can maintain high tracking accuracy even in complex environments.
[0038] High-definition detail generation: High-definition repair of details especially for tooth regions, through advanced fine-grained feature fusion (FGFF) module and optimized loss function in the generator of GAN network, significantly improving the visual clarity and realism of teeth.
[0039] Synchronicity and coherence: The invention maintains the synchronicity of mouth shape and voice while enhancing tooth clarity, ensuring natural and smooth animation, and improving user experience.
[0040] Generalization ability: Through the two-stage fine-tuning strategy, the invention can adapt to different facial features and expressions, enhancing the generalization ability of the model and making it work stably in various scenarios.
[0041] Computational efficiency: The invention focuses on computational efficiency in design, through optimized network structure and parallel processing technology, realizing fast image generation speed and meeting the needs of real-time applications.
[0042] Personalized customization: The invention allows personalized customization, through specific task fine-tuning, it can be optimized for specific characters or scenes to meet specific user needs.
[0043] Robustness: Through specific data preprocessing and enhancement techniques, the invention has high robustness to input data noise and variation, and can stably generate high-quality results on different quality data. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0045] Figure 1 is a flowchart of the high-definition tooth repair method based on GAN network provided by the embodiment;
[0046] Figure 2 is a structural schematic diagram of the generator provided by the embodiment;
[0047] Figure 3 is a structural schematic diagram of the high-definition tooth repair device based on GAN network provided by the embodiment. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific implementation described herein is only used to explain the present application and does not limit the protection scope of the present application.
[0049] The inventive concept of the present application is to solve the deficiencies of the prior art in tooth clarity and mouth shape synchronization. The present application provides a high-definition tooth repair method based on a GAN network, which focuses on significantly improving the visual quality of the tooth region in the speech-driven face generation (TFG) technology, i.e., the effect of high-definition repair of the tooth region, while ensuring high synchronization and temporal coherence with the speech. First, in the data preprocessing stage, through accurate extraction of facial key points and generation of a mask based on the key points, the method focuses on detail enhancement of the tooth region, avoids blurring of the lip edge, and provides standardized input data for the model. Then, in the pre-training stage, the model learns the general representation of facial features and movements on a large-scale face motion dataset, and uses multiple loss functions to optimize the generation quality and synchronization of the model. Most importantly, in the specific person fine-tuning stage, HDTR makes detailed adjustments to individualized features, uses data of a specific person to learn the unique tooth shape and facial features of the person, and further optimizes the loss function to adapt to individual differences. The fine-tuning in this stage significantly improves the performance of the model in mouth image reconstruction, and enhances the tooth clarity and the ability to preserve image details.
[0050] As shown in Figure 1 The embodiment provides a high-definition tooth repair method based on a GAN network, which includes the following steps:
[0051] S1, facial key points are extracted from the obtained facial image, the tooth contour and tooth mask are determined based on the facial key points, the tooth contour and tooth mask image are superimposed to obtain a tooth to-be-repaired image, the mouth region is segmented from the facial image as a reference image, and after image enhancement and fine processing of the mouth region reference image, the tooth to-be-repaired image and the mouth region reference image are normalized.
[0052] In the embodiment, facial images are obtained from a camera or other video sources, and these images usually contain RGB color channels. MobileNetV3 network is used to extract facial key points from the facial images, which specifically include three-dimensional coordinates of key points such as eyes, eyebrows, nose, and mouth.
[0053] Specifically, the open-source real-time face and facial landmark tracking library OpenSeeFace is used, which is based on the MobileNetV3 network architecture. By converting the model to ONNX format and using onnxruntime for inference, high-speed single face tracking at 30-60fps on CPU is achieved, significantly reducing the demand for hardware. OpenSeeFace adopts a unique 66-point facial landmark model that maintains stable tracking performance even under adverse conditions such as low light, high noise, and low resolution, and performs particularly well in head pose changes and mouth motion capture.
[0054] The network structure of MobileNetV3 is the latest achievement of the MobileNet series of lightweight deep learning models. The network structure can be divided into three main parts:
[0055] The starting part: It is a convolutional layer with a large kernel, used to capture the initial features of the input image. In MobileNetV3, the starting part usually contains a 3x3 convolutional layer, followed by a batch normalization (BatchNormalization, BN) layer and an activation function.
[0056] The middle part: This is the core of the network composed of multiple inverted residual structures (Inverted Residual Blocks). Each inverted residual structure can contain an expansion layer, a depth separable convolution layer, and an output layer. Depth separable convolution is the core feature of the MobileNet series, which decomposes the standard convolution into depth convolution (Depthwise Convolution) and pointwise convolution (Pointwise Convolution). In MobileNetV3, this structure is further optimized to improve efficiency. The general form of the inverted residual structure can be represented as:
[0057] out1 = Depthwise(Conv(Expand(input)))
[0058] out2 = Pointwise(out1)
[0059] out3 = Activation(out2)
[0060] Where Expand, Depthwise, and Pointwise represent the expansion layer, depth convolution layer, and pointwise convolution layer, respectively, and Activation represents the activation function.
[0061] The last part: In the last part of the network, one or more 1x1 convolutional layers are included to convert the feature map into the final output. In the classification task, this is usually a global average pooling layer followed by a 1x1 convolutional layer, and finally a fully connected layer outputs the class prediction. This process can be represented by the following formula:
[0062] out4 = GlobalAvgPool(Conv(out3))
[0063] out5 = FC(out4)
[0064] Where GlobalAvgPool represents the global average pooling layer, Conv represents the last 1x1 convolutional layer, and FC represents the fully connected layer.
[0065] In the inverted residual structure of MobileNetV3, a new activation function h-swish is used, and the sigmoid function in the SE module is approximated to improve computational efficiency. In addition, MobileNetV3 also introduces the NetAdapt algorithm to optimize part of the network layers, and redesigns the time-consuming structure in the network. OpenSeeFace converts the MobileNetV3 model into ONNX format and uses onnxruntime for inference to achieve efficient running on the CPU.
[0066] The face key points extracted based on the MobileNetV3 network include the mouth contour points and the related feature points in the tooth region. The embodiment determines the tooth contour and tooth mask based on the tooth contour points.
[0067] Tooth region segmentation and image enhancement directly affect the performance of the model and the final repair effect. The present application improves the traditional lip mask to a mask only for teeth through an innovative data processing method, effectively eliminating the phenomenon of blurred lip edges, and significantly enhancing the clarity of details in the tooth region. The key steps include accurate segmentation of the tooth region, appropriate application of image enhancement techniques, and data normalization processing, ensuring that the data set is both rich and suitable for model training. In addition, by finely processing the mouth region reference image, the image quality is further improved, the input data of the model is optimized, and the generalization ability and repair effect of the model are improved, providing a solid data foundation for generating high-definition and high-synchronization tooth regions.
[0068] In the embodiment, the tooth contour and tooth mask are determined based on the extracted face key points, and the mouth region reference image M reference is generated. Specifically, it includes: determining the tooth contour and tooth mask based on the face key point set P facialGenerate a tooth outline image:
[0069] M tooth_contour =M(P facial )
[0070] The tooth area is masked by the mask generation function M to obtain the initial tooth mask image M mask :
[0071] M mask =M(P facial )
[0072] Perform an XOR operation on the initial tooth mask image and the mouth area reference image so that the tooth area of the mouth area reference image forms a mask, and the final tooth mask image is obtained:
[0073] M tooth_mask =U(M mask ,M reference )
[0074] Wherein, U represents the image XOR operation.
[0075] The tooth contour image and the tooth mask image are superimposed to obtain the tooth image to be repaired M tooth .
[0076] Perform image enhancement and refinement on the mouth region reference image, and normalize the tooth restoration image and the mouth region reference image. Image enhancement involves contrast enhancement and sharpening, while refinement involves edge sharpening and smoothing. Normalization normalizes the image pixel values to between [0, 1] or [-1, 1]. The formula is:
[0077] I enhanced =ε(M reference ,α)
[0078] Among them, I enhanced is the enhanced reference image of the mouth area, M reference is the reference image of the mouth area, ε is the enhancement function, and α is the parameter that controls the degree of enhancement.
[0079] Refinement processing refers to the refinement of the reference image of the mouth area, including edge sharpening and smoothing to eliminate possible artifacts or unnatural effects. The formula is:
[0080] I refined =S(I enhance ,β)
[0081] Among them, I refined is the reference image of the mouth area after refinement, S is the refinement function, β is the parameter that controls the processing intensity, and I enhanceThe enhanced mouth region reference image is denoted as
[0082] The normalization processing is to adapt the input data to the model training, and the enhanced mouth region reference image and the tooth image to be repaired are normalized to make the pixel value range between [0, 1] or [-1, 1].
[0083] In the embodiment, the tooth image to be repaired and the mouth region reference image are taken as the input images of the high-definition tooth repair model.
[0084] S2, constructing a high-definition tooth repair model for repairing the tooth image to be repaired based on the GAN network.
[0085] In the embodiment, the generator contained in the GAN network adopts an encoding structure and a decoding structure, the encoding structure adopts two branches to extract and encode features of the tooth image to be repaired and the mouth region reference image respectively, the decoding structure decodes the encoding results of the two branches to perform tooth repair to obtain a repaired tooth image, the discriminator contained in the GAN network is used to judge the authenticity of the repaired tooth image, and the trained generator is taken as the high-definition tooth repair model.
[0086] Specifically, the two branches of the encoding structure are the same, each branch structure includes a feature extraction module and a fine-grained feature fusion module (FGFF), wherein the feature extraction module extracts features based on a convolutional neural network, and specifically extracts rich features related to teeth and surrounding areas, and the convolutional neural network specifically includes convolution operation, feature dimension reduction, pooling layer, residual connection, multi-scale feature fusion, and regularization technique.
[0087] Specifically, in a convolutional neural network, each convolutional layer implementing convolution operation uses a set of convolution kernels (or filters) to scan the input image, capture local features, and output feature maps. As the network level deepens, the spatial dimensions (i.e., height and width) of the feature maps gradually decrease, while the depth (i.e., the number of channels) of the features gradually increases. This design helps to extract higher-level semantic information from raw pixel-level information. Pooling layers (such as max pooling or average pooling) are used between consecutive convolutional layers to reduce the spatial size of the feature maps, thereby reducing the number of parameters and computational complexity, while making the feature detection more robust. To address the problem of gradient vanishing or gradient explosion in deep networks, residual connections are introduced between convolutional layers. Residual connections allow direct transmission of information from earlier layers to later layers. In the last stage of feature extraction, feature fusion techniques are used to merge feature maps at different levels to recover the loss of spatial resolution caused by pooling operations, while preserving multi-scale feature information. To prevent overfitting, regularization techniques such as dropout or batch normalization (Batch Normalization) may also be included in the feature extraction layer.
[0088] Through these carefully designed network structures and operations, the feature extraction layer can efficiently extract multi-level features that are crucial for dental restoration from facial images, laying a solid foundation for subsequent fine-grained feature fusion and high-definition tooth generation.
[0089] The fine-grained feature fusion module, as the core part, is used to extract and fuse high-frequency detailed features of teeth and their surrounding areas. The specific steps include channel fusion (CF), HourGlass network, feature fusion, and feature refinement. Channel fusion refers to gradually reducing the number of channels of the feature map and increasing the depth of the feature to extract more refined features through convolution operation. HourGlass network is used to extract multi-scale features to capture features from local to global. Feature fusion refers to the feature fusion of the output results of channel fusion and the extracted multi-scale features. Feature refinement is used to enhance the fine-grained features of the feature fusion results through convolution and activation operations as the encoding results.
[0090] In one embodiment, as Figure 2As shown, each fine-grained feature fusion module includes a convolutional layer, a max-pooling layer, three channel fusions, an HourGlass network, a channel fusion, and feature refinement connected in sequence. The HourGlass network adopts 2 branches, one branch extracts fine features through channel fusion for channel reduction, and the other branch extracts multi-scale features through a combination of 2 pairs of max-pooling and channel fusion and upsampling in sequence, and the fine features and multi-scale features are fused through feature fusion to obtain a feature fusion result. The feature refinement adopts at least one deep separable convolution or a residual connection constructed between at least one deep separable convolution.
[0091] Through these steps, the fine-grained feature fusion module can effectively extract and fuse the fine-grained features of the tooth region, providing rich feature support for the tooth restoration task. This fine-grained feature extraction and fusion strategy is the key to achieving high-definition tooth restoration.
[0092] In the embodiment, the tooth image to be repaired and the mouth region reference image are superimposed after the encoding results obtained by the two branches are obtained, input to the decoding structure, and the repaired tooth image is obtained after decoding.
[0093] In the embodiment, the discriminator is composed of a series of convolutional layers and activation functions, and the purpose is to provide a binary output to determine whether the input image is a real image or an image generated by the generator.
[0094] In training the above GAN, the loss function L total including the adversarial loss of the generator the reconstruction loss of the generator the perceptual loss L perc of the generator, the discriminant loss L D of the discriminator, and the R1 loss L R1 of the discriminator.
[0095] Adversarial loss is used to encourage the generator to produce images that are difficult for the discriminator to distinguish. Reconstruction loss ensures that the output of the generator is close to the real image at the pixel level. Perceptual loss L perc uses a pre-trained network (such as VGG) to measure the difference between the generated image and the real image at the feature level. Discriminant loss L D refers to training the discriminator to distinguish between real images and repaired tooth images generated by the generator.
[0096] R1 loss L R1is a regularization technique used to stabilize and improve the GAN training process. This loss function works by penalizing the gradients of the discriminator, especially when these gradients approach zero, meaning the discriminator is overly sensitive to small changes in the input images. The R1 loss helps avoid overfitting of the discriminator to the generated images and ensures that the generator has enough space to generate diverse images.
[0097]
[0098] where λ is a hyperparameter controlling the strength of the R1 loss, is sampled from the distribution intermediate between the tooth reference image and the post-repaired tooth image, is the discriminative result of the discriminator on , is the gradient of on , and ||·||2 denotes the L2 norm used to calculate the size of the gradient vector, E represents the expected value, which is calculated by sampling .
[0099] The total loss function L total is represented as:
[0100]
[0101] where λ1, λ2, λ3, and λ4 represent the loss weights, and by minimizing this total loss function L total , the GAN can generate realistic images while maintaining the stability and effectiveness of the discriminator. The design of this comprehensive loss function helps balance the generation quality and discrimination ability of the network, pushing the generator to produce more realistic and diverse outputs.
[0102] In the embodiment, a two-stage training strategy is adopted, including a pre-training phase in the first stage and fine-tuning in the second stage. In the first stage of high-definition tooth repair model, that is, the pre-training stage, the model is trained using a large-scale facial image dataset to learn general facial features and structures. The goal of this stage is to enable the model to capture key points and expression changes of the face, laying the foundation for high-definition repair of the tooth area. By using standardized training techniques, the model establishes a preliminary understanding of facial details in this stage, which includes general feature learning of tooth shape and surrounding tissue, preparing conditions for subsequent targeted fine-tuning. Pre-training does not involve specific task optimization, but focuses on acquiring broad features and enhancing model stability.
[0103] The model learns through a series of meticulous steps on a large-scale facial motion dataset, including initializing the generator and discriminator networks, inputting pre-processed keypoint sequences and corresponding video frames, and defining loss functions to optimize model performance. The generator is responsible for producing realistic dental images, while the discriminator evaluates the authenticity of the generated images. In the adversarial training, the GAN loss Guides the adversarial process between the generator and discriminator, reconstruction loss Ensures pixel-level consistency, perceptual loss Captures high-level feature differences through the VGG network. The autoregressive architecture enhances the temporal continuity of video frames, and the optical flow enhancement technique improves the authenticity of dynamic changes through a multi-scale dynamic discriminator. The texture adhesion improvement scheme uses non-aliasing convolution and global translation and rotation transformation modules to further improve the quality of the generated video. Backpropagation and optimization algorithms such as Adam are used to update network parameters, and performance evaluation is conducted through quantitative indicators (PSNR, SSIM) and qualitative analysis, with hyperparameter adjustment ensuring the efficiency and stability of model training. Specialized training of the multi-scale dynamic discriminator improves sensitivity to dynamic changes, and testing of model stability and generalization ability ensures the robustness of the model on diverse data. Finally, the integrated training of innovative modules and the preservation of pre-trained models lay the foundation for subsequent fine-tuning and application, making HDTR exhibit extensive application potential and outstanding performance in the field of digital human video generation.
[0104] The second phase of the high-definition dental repair model focuses on fine-tuning for specific tasks, which involves fine-tuning the model to ensure excellent performance in high-definition dental repair. A specific task dataset closely related to high-definition dental repair is selected, which may include high-definition facial images of specific individuals, especially those with high-definition dental regions. These samples contain specific facial features and dental details of specific individuals. Using these selected data, the model is fine-tuned on the basis of pre-training to optimize its performance on specific tasks. The fine-tuning process begins with fine adjustments to the model parameters to adapt to the characteristics of the new dataset. Then, the model is further improved in the clarity and realism of the dental region through iterative training using appropriate loss functions such as perceptual loss and reconstruction loss. The calculation of the loss function uses pixel-level differences and high-level feature differences to ensure the high quality of the generated images in terms of vision.
[0105] In the fine-tuning stage, the learning rate needs to be reduced, and optimizers such as Adam are used with their parameters adjusted to ensure that the model can finely adapt to the new dataset while avoiding excessive adjustment of the pre-trained parameters. Through small batch gradient descent, the model parameters are constantly updated to minimize the loss function and improve the repair effect. Through these fine adjustments, the second stage of fine-tuning ensures the professional performance of the model in high-definition tooth repair, improving the visual quality and detail clarity of the tooth area, and meeting the specific needs of high-definition tooth generation.
[0106] During the fine-tuning process, the performance of the model on the training set is monitored through the validation set to evaluate its generalization ability. The repair effect of the model is measured through quantitative indicators such as PSNR and SSIM, as well as qualitative evaluation such as visual inspection of the clarity and naturalness of the generated images. Ultimately, the fine-tuned model shows significant improvement in tooth details, mouth synchronization, and overall visual quality, enabling it to generate high-definition tooth video sequences that are highly consistent with specific facial movements.
[0107] S3, high-definition tooth repair using a high-definition tooth repair model.
[0108] After processing the face image through S1, the tooth image to be repaired and the mouth area reference image are obtained, and the tooth image to be repaired and the mouth area reference image are input into the high-definition tooth repair model to obtain the repaired tooth image through the encoder decoding.
[0109] As shown in Figure 3 The embodiment also provides a high-definition tooth repair device 30 based on a GAN network, which includes an image processing module 31, a model construction module 32, and a tooth repair module 33. The image processing module 31 is used to extract facial key points from the obtained face image, determine tooth contours and tooth masks based on the facial key points, superimpose the tooth contours and tooth mask images to obtain a tooth image to be repaired, segment the mouth area from the face image as a reference image, and perform image enhancement and fine processing on the mouth area reference image. After normalization processing of the tooth image to be repaired and the mouth area reference image, the model construction module 32 is used to construct a high-definition tooth repair model based on a GAN network, and the tooth repair module 33 is used to perform high-definition tooth repair using the high-definition tooth repair model.
[0110] It should be noted that the high-definition tooth repair device based on the GAN network provided in the above embodiment should be illustrated by the division of the above functional modules when performing high-definition tooth repair. The above functions can be completed by different functional modules as needed, that is, the internal structure of the terminal or the server is divided into different functional modules to complete all or part of the functions described above. In addition, the high-definition tooth repair device based on the GAN network provided in the above embodiment and the high-definition tooth repair method based on the GAN network embodiment belong to the same concept, and the specific implementation process is detailed in the high-definition tooth repair method based on the GAN network. Here, it will not be repeated.
[0111] Based on the same inventive concept, the embodiment also provides a computing device including a memory and one or more processors, the memory having stored therein executable code, the one or more processors, when executing the executable code, being configured to implement the high-definition tooth repair method based on the GAN network, specifically including the following steps:
[0112] S1, performing face key point extraction on the obtained face image, determining a tooth contour and a tooth mask based on the face key points, superimposing the tooth contour and the tooth mask image to obtain a tooth to-be-repaired image, segmenting a mouth region from the face image as a reference image, and performing image enhancement and fine processing on the mouth region reference image, and then performing normalization processing on the tooth to-be-repaired image and the mouth region reference image;
[0113] S2, constructing a high-definition tooth repair model for repairing the tooth to-be-repaired image based on a GAN network;
[0114] S3, performing high-definition tooth repair using the high-definition tooth repair model.
[0115] The computing device provided in the embodiment, in addition to including a processor and a memory, also includes an internal bus, a network interface, a memory, and other hardware required for business. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs to implement the high-definition tooth repair method based on the GAN network. Of course, in addition to the software implementation mode, the present application does not exclude other implementation modes, such as logic devices or a combination of software and hardware. That is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or a logic device.
[0116] Based on the same inventive concept, the embodiment also provides a computer readable storage medium having a program stored thereon, the program being executed by a processor to implement the high-definition tooth repair method based on the GAN network, specifically including the following steps:
[0117] S1, face key point extraction is performed on the obtained face image, a tooth contour and a tooth mask are determined based on the face key points, the tooth contour and the tooth mask image are superimposed to obtain a tooth to-be-repaired image, a mouth region is segmented from the face image as a reference image, image enhancement and fine processing are performed on the mouth region reference image, and then the tooth to-be-repaired image and the mouth region reference image are normalized;
[0118] S2, a high-definition tooth repair model for repairing the tooth to-be-repaired image is constructed based on a GAN network;
[0119] S3, high-definition tooth repair is performed by using the high-definition tooth repair model.
[0120] In the embodiments, the computer readable medium includes permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data.
[0121] The specific embodiments described above describe the technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the most preferred embodiment of the present application and is not intended to limit the present application. Any modification, supplement and equivalent replacement within the principle range of the present application should be included in the protection scope of the present application.
Claims
1. A high-definition tooth restoration method based on a GAN network, characterized in that: The following steps are involved: Extracting facial key points from the acquired facial image, including mouth contour points and relevant feature points of the tooth area, determining tooth contours and tooth masks based on the facial key points, superimposing the tooth contour and tooth mask images to obtain a tooth image to be restored, segmenting the mouth area from the facial image as a reference image, performing image enhancement and refinement on the mouth area reference image, and then normalizing the tooth image to be restored and the mouth area reference image; A high-definition tooth restoration model was constructed based on a GAN network. The generator contained in the GAN network used an encoding structure and a decoding structure. The encoding structure used two branches to extract and encode features from the normalized tooth image to be restored and the reference image of the mouth area, respectively. The decoding structure decoded the encoding results of the two branches to perform tooth restoration and obtain the restored tooth image. The discriminator contained in the GAN network was used to determine the authenticity of the restored tooth image. After the two-stage training, the generator was used as the high-definition tooth restoration model. Among them, the two branch structures used in the encoding structure are the same. Each branch structure includes a feature extraction module and a fine-grained feature fusion module. The two modules are used to extract features and perform fine-grained feature fusion to realize the encoding process. The fine-grained feature fusion module specifically includes channel fusion, HourGlass network, feature fusion, and feature refinement. Channel fusion refers to gradually reducing the number of channels of the feature map through convolution operations while increasing the depth of features to extract finer features. The HourGlass network is used to extract multi-scale features. Feature fusion refers to fusing the output results of the channel fusion pair with the extracted multi-scale features. Feature refinement refers to enhancing the details of the feature fusion result as the encoding result through convolution and activation operations. Each fine-grained feature fusion module includes a sequentially connected convolutional layer, a maximum pooling layer, three channel fusions, an HourGlass network, a feature fusion, and feature refinement. The HourGlass network uses two branches. One branch extracts fine features through channel fusion by channel indentation, and the other branch extracts multi-scale features through a combination of two pairs of maximum pooling and channel fusion, as well as upsampling. The fine features and multi-scale features are then fused to obtain a feature fusion result. Feature refinement uses at least one depth-wise separable convolution or at least one residual connection constructed between depth-wise separable convolutions. High-definition dental restorations are performed using high-definition dental restoration models.
2. The high-definition tooth restoration method based on the GAN network according to claim 1, characterized in that: The MobileNetV3 network is used to extract facial key points, and the tooth contour and tooth mask are determined based on the tooth contour points. The mouth area reference image M is segmented from the facial image. reference .
3. The high-definition tooth restoration method based on the GAN network according to claim 1 or 2, characterized in that: When determining the tooth contour, first use the contour generation function C based on the facial key point set P facial Generate tooth contour image M tooth_contour : M tooth_contour =C(P facial ) The facial key point set P is generated by the mask function M facial Perform masking to obtain the initial tooth mask image M mask : M mask =M(P facial ) Perform an XOR operation on the initial tooth mask image and the mouth area reference image so that the tooth area of the mouth area reference image forms a mask, and the final tooth mask image is obtained: M tooth_mask =U(M mask ,M reference ) Among them, U represents the image XOR operation; The tooth contour image and the tooth mask image are superimposed to obtain the tooth image to be repaired M tooth .
4. The high-definition tooth restoration method based on the GAN network according to claim 1, characterized in that: The reference image of the mouth area is enhanced and refined, and the image of the tooth to be repaired and the reference image of the mouth area are normalized. During image enhancement, contrast enhancement and sharpening are performed, and during refinement, edge sharpening and smoothing are performed. During normalization, the pixel values of the image are normalized to between [0, 1] or [-1, 1].
5. The high-definition tooth restoration method based on the GAN network according to claim 1, characterized in that: The loss function L used when training the GAN network total Including the adversarial loss of the generator Reconstruction loss of the generator Generator’s perceptual loss L perc , the discriminator's loss L D , and the discriminator's R1 loss L R1 , Among them, λ1, λ2, λ3, and λ4 represent the weights of each loss, and λ is a hyperparameter used to control the strength of R1 loss. is the distribution between the tooth reference image and the restored tooth image. The sampling is obtained, is the discriminator pair The judgment result of Yes The gradient of , ‖·‖2 represents the L2 norm, which is used to calculate the size of the gradient vector, and E represents the expected value, which is obtained by sampling To calculate.
6. A high-definition tooth restoration device based on a GAN network, characterized in that: include: Image processing module, model building module, tooth restoration module, The image processing module is used to extract facial key points from the acquired facial image, wherein the extracted facial key points include mouth contour points and relevant feature points of the tooth area, determine the tooth contour and tooth mask based on the facial key points, superimpose the tooth contour and tooth mask images to obtain the tooth image to be restored, segment the mouth area from the facial image as a reference image, perform image enhancement and refinement processing on the mouth area reference image, and then perform normalization processing on the tooth image to be restored and the mouth area reference image; The model construction module is used to construct a high-definition tooth restoration model based on a GAN network, wherein the generator included in the GAN network adopts an encoding structure and a decoding structure. The encoding structure uses two branches to extract and encode features of the normalized tooth to be restored image and the mouth area reference image respectively. The decoding structure decodes the encoding results of the two branches to perform tooth restoration to obtain a restored tooth image. The discriminator included in the GAN network is used to distinguish the authenticity of the restored tooth image. The generator after the two-stage training is used as the high-definition tooth restoration model; Among them, the two branch structures used in the encoding structure are the same. Each branch structure includes a feature extraction module and a fine-grained feature fusion module. The two modules are used to extract features and perform fine-grained feature fusion to realize the encoding process. The fine-grained feature fusion module specifically includes channel fusion, HourGlass network, feature fusion, and feature refinement. Channel fusion refers to gradually reducing the number of channels of the feature map through convolution operations while increasing the depth of features to extract finer features. The HourGlass network is used to extract multi-scale features. Feature fusion refers to fusing the output results of the channel fusion pair with the extracted multi-scale features. Feature refinement refers to enhancing the details of the feature fusion result as the encoding result through convolution and activation operations. Each fine-grained feature fusion module consists of a sequentially connected convolutional layer, a maximum pooling layer, three channel fusions, an HourGlass network, a feature fusion, and feature refinement. The HourGlass network uses two branches. One branch extracts fine features through channel fusion by channel indentation, and the other branch extracts multi-scale features through a combination of two pairs of maximum pooling and channel fusion, as well as upsampling. The fine features and multi-scale features are then fused to obtain the feature fusion result. Feature refinement uses at least one depth-wise separable convolution or at least one residual connection constructed between depth-wise separable convolutions; The tooth restoration module is used to perform high-definition tooth restoration using a high-definition tooth restoration model.
7. A computing device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, they are used to implement the high-definition tooth restoration method based on the GAN network according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that A program is stored thereon, and when the program is executed by the processor, the high-definition tooth restoration method based on the GAN network described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Restoration body image generation method and device, equipment and storage medium
CN113888615A
Unsupervised face forgery evaluation method
CN114267063A