Electronic device for performing face restoration on basis of deep learning model, and driving method therefor
The electronic device employs a multi-stage deep learning model with facial recognition thresholds to address the inefficiencies of existing face swapping technologies, ensuring rapid and high-clarity face restoration in film and video content creation.
Patent Information
- Application Number
- PCT/KR2024/007571
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-03
- Filing Date
- 2024-06-03
- Publication Date
- 2025-12-11
AI Technical Summary
Existing face swapping technologies using deep learning models are time-consuming and often produce blurry results due to insufficient learning time, leading to reduced production efficiency in film and video content creation.
An electronic device and method that utilize a multi-stage deep learning model architecture, including sub-deep learning models and a main deep learning model, with facial recognition thresholds to ensure high clarity face restoration by optimizing the input and learning process, and iteratively refining image data until desired clarity is achieved.
The solution enables rapid output of high-clarity face restoration images by minimizing time waste and maintaining high clarity through iterative refinement, enhancing production efficiency in content creation.
Smart Images

Figure KR2024007571_11122025_PF_FP_ABST
Abstract
Description
Electronic device for performing face restoration based on a deep learning model and its driving method
[0001] Various embodiments of the present invention relate to an electronic device for performing face restoration based on a deep learning model and a method for operating the same, and more particularly, to an electronic device for performing face restoration based on a deep learning model capable of easily outputting restored image data by inputting basic data into a deep learning model to restore the faces of past people, past animals, and future predicted people, and a method for operating the same.
[0002] Recently, deepfake, or face swapping, refers to creating new content by synthesizing or replacing the face of a subject in an original image or video with a face from another source. While issues such as synthesizing the faces of celebrities to create fake videos have arisen, the film and video industries are also using this technology to synthesize the faces of deceased individuals or to recreate the sophisticated appearances of younger actors. While existing techniques and processes exist for synthesizing faces using 3D models, such as digital doubles, face swapping has the advantage of significantly reducing costs, drawing significant attention within the industry.
[0003] However, in the past, the process of creating a deep learning model for converting a source face into a target face based on the source face data set and the target face data set required a lot of time, and there was a limitation that a blurry synthesis result was obtained if sufficient learning time was not reflected.
[0004] In particular, in the past, it took a lot of time to increase the clarity after implementing the initial face shape, which was a problem that lowered the production efficiency in producing movie / video content.
[0005] Accordingly, the present invention can provide an electronic device and a method of driving the same that can easily output face restoration image data while reducing waste of time while simultaneously inputting and learning face source data and face target data into a deep learning model and verifying the accuracy of the output values.
[0006] According to various embodiments, an electronic device for performing face restoration based on a deep learning model comprises: a communication interface; a memory; and a processor; , wherein the processor obtains face source data and face target data for restoration through the communication interface, inputs the face source data and the face target data into at least one first sub-deep learning model, respectively, to output a source data set and a target data set, inputs the source data set and the target data set into a second sub-deep learning model to output a mask data set, and if it is determined that the mask data set is greater than or equal to a preset first facial recognition threshold value, inputs the mask data set into a main deep learning model to output final restored image data, and if it is determined that the mask data set is less than or equal to the first facial recognition threshold value, inputs the mask data set into the main deep learning model while setting a setting value to the mask data set to output first restored image data, and if it is determined that the first restored image data is greater than or equal to a preset second facial recognition threshold value, inputs the first restored image data into the main deep learning model to output the final restored image data, and if the first restored image data is determined to be greater than or equal to the second If it is determined to be below the face recognition threshold, the process is set to repeat from the case where the first restored image data is below the first face recognition threshold to the case where it is determined to be above the second face recognition threshold, and the main deep learning model is trained based on a plurality of source data sets for a plurality of faces, a plurality of target data sets for a plurality of faces, a plurality of mask data sets for a plurality of faces, final restored image data for a plurality of restored faces, a plurality of first restored image data for a plurality of restored faces, a plurality of face restoration algorithms, and an image clarity evaluation algorithm.
[0007] According to various embodiments, the present embodiment has the advantage of being able to output face restoration image data with very high clarity in a short time by learning face source data and face target data respectively to form a primary data set, inputting these into a deep learning model based on the data, minimizing the output time while maintaining high clarity, and continuously verifying the same.
[0008] FIG. 1 illustrates a block diagram of an electronic device and a network according to various embodiments of the present invention.
[0009] FIG. 2 is a flowchart illustrating a method of operating an electronic device according to various embodiments of the present invention.
[0010] FIG. 3 is a flowchart specifically explaining an operation method in which an electronic device outputs output from a mask data set according to various embodiments of the present invention.
[0011] FIG. 4 is a flowchart specifically explaining an operation method in which an electronic device outputs first restored image data according to various embodiments of the present invention.
[0012] Hereinafter, various embodiments of the present document will be described with reference to the attached drawings. It should be understood that the embodiments and the terms used therein are not intended to limit the technology described in the present document to a specific embodiment, but rather include various modifications, equivalents, and / or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar components. The singular expression may include plural expressions unless the context clearly indicates otherwise. In this document, expressions such as "A or B" or "at least one of A and / or B" may include all possible combinations of the items listed together. Expressions such as "first," "second," "first," or "second," may modify the corresponding components regardless of order or importance, and are only used to distinguish one component from another, but do not limit the corresponding components. When it is said that a component (e.g., a first component) is “(functionally or communicatively) connected” or “connected” to another component (e.g., a second component), said component may be directly connected to said other component, or may be connected via another component (e.g., a third component).
[0013] In this document, "configured to" may be used interchangeably with, for example, "suitable for," "capable of," "modified to," "made to," "capable of," or "designed to," either in hardware or software. In some contexts, the phrase "a device configured to" may mean that the device is "capable of" doing something together with other devices or components. For example, the phrase "a processor configured to perform A, B, and C" may mean a dedicated processor (e.g., an embedded processor) for performing the operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform the operations by executing one or more software programs stored in a memory device.
[0014] An electronic device according to various embodiments of the present document may include, for example, at least one of a smartphone, a tablet PC, a desktop PC, a laptop PC, a netbook computer, a workstation, and a server.
[0015] Referring to FIG. 1, an electronic device (101) within a network environment (100) according to various embodiments is described. The electronic device (101) may include a bus (110), a processor (120), a memory (130), an input / output interface (140), a display (150), and a communication interface (160). In some embodiments, the electronic device (101) may omit at least one of the components or additionally include other components. The bus (110) may include a circuit that connects the components (110-170) to each other and transmits communication (e.g., control messages or data) between the components. The processor (120) may include one or more of a central processing unit, an application processor, or a communication processor (CP). The processor (120) may, for example, execute operations or data processing related to control and / or communication of at least one other component of the electronic device (101).
[0016] The memory (130) may include volatile and / or non-volatile memory. The memory (130) may store, for example, commands or data related to at least one other component of the electronic device (101). According to one embodiment, the memory (130) may store software and / or programs (140).
[0017] The input / output interface (140) can, for example, transmit commands or data input from a patient or other external device to other component(s) of the electronic device (101), or output commands or data received from other component(s) of the electronic device (101) to the patient or other external device.
[0018] The display (150) may include, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a micro electro mechanical systems (MEMS) display, or an electronic paper display. The display (150) may, for example, display various contents (e.g., text, images, videos, icons, and / or symbols) to the patient. The display (150) may include a touch screen and may receive, for example, touch, gesture, proximity, or hovering input using an electronic pen or a part of the patient's body. The communication interface (160) may, for example, establish communication between the electronic device (101) and an external device (e.g., the first external electronic device (102), the second external electronic device (104), or the server (108)). For example, the communication interface (160) can be connected to a network (162) via wireless communication or wired communication to communicate with an external device (e.g., a second external electronic device (104) or a server (108)).
[0019] The wireless communication may include, for example, cellular communication using at least one of LTE, LTE-A (LTE Advance), CDMA (code division multiple access), WCDMA (wideband CDMA), UMTS (universal mobile telecommunications system), WiBro (Wireless Broadband), or GSM (Global System for Mobile Communications). In one embodiment, the wireless communication may include, for example, at least one of WiFi (wireless fidelity), Bluetooth, Bluetooth low energy (BLE), Zigbee, near field communication (NFC), Magnetic Secure Transmission, radio frequency (RF), or body area network (BAN). In one embodiment, the wireless communication may include GNSS. The GNSS may be, for example, GPS (Global Positioning System), Glonass (Global Navigation Satellite System), Beidou Navigation Satellite System (hereinafter "Beidou"), or Galileo, the European global satellite-based navigation system. Hereinafter, in this document, "GPS" may be used interchangeably with "GNSS." Wired communication may include at least one of, for example, USB (universal serial bus), HDMI (high definition multimedia interface), RS-232 (recommended standard 232), power line communication, or POTS (plain old telephone service).The network (162) may include at least one of a telecommunications network, for example, a computer network (e.g., a LAN or WAN), the Internet, or a telephone network.
[0020] Each of the first and second external electronic devices (102, 104, 106) may be the same or a different type of device as the electronic device (101). According to various embodiments, all or part of the operations executed in the electronic device (101) may be executed in another one or more electronic devices (e.g., electronic devices (102, 104, 106), or server (108). According to one embodiment, when the electronic device (101) is to perform a certain function or service automatically or upon request, the electronic device (101) may request at least some functions related thereto from another device (e.g., electronic devices (102, 104, 106), or server (108)) instead of executing the function or service by itself or in addition. The other electronic device (e.g., electronic devices (102, 104, 106), or server (108)) may execute the requested function or additional function and transmit the result to the electronic device (101). The electronic device (101) may process the received result as is or additionally to provide the requested function or service. For this purpose, for example, cloud computing, distributed computing, or client-server computing technology may be used.
[0021]
[0022] FIG. 2 is a flowchart illustrating a method of operating an electronic device according to various embodiments of the present invention.
[0023]
[0024] In operation 201, the electronic device (101) (e.g., the processor (120) of FIG. 1) may obtain face source data and face target data for restoration via a communication interface. For example, the face for restoration may be composed of the face of a currently non-living person whose photo is not clear. In addition, the face for restoration may be composed of the faces of extinct animals other than humans. In addition, the face for restoration may be composed of the faces of the user's children in the future, as requested by the user's external electronic device (102, 104, 106), or may be composed of other faces. For example, the face source data may be composed of image data and / or video data including a source face extracted by transformation from among the faces for restoration. For example, the face target data may be composed of image data and / or video data including a target face to be changed from among the faces for restoration.
[0025] In operation 203, the electronic device (101) (e.g., the processor (120) of FIG. 1) may input face source data and face target data into at least one first sub-deep learning model, respectively, to output a source data set and a target data set. For example, the at least one first sub-deep learning model may include a first-first sub-deep learning model and a first-second sub-deep learning model. For example, the electronic device (101) (e.g., the processor (120) of FIG. 1) may input face source data into the first-first sub-deep learning model to output a source data set, and input face target data into the first-second sub-deep learning model to output a target data set.
[0026] For example, the first sub-deep learning model can be operated as a CNN-based deep learning model, and can be operated as a deep learning model to be input for each data. For example, it can be input to a CNN-based deep learning model. For example, the CNN-based deep learning model may include a deep neural network (DNN: Deep Neural Network). For example, the artificial intelligence model may include, but is not limited to, a CNN (Convolutional Neural Network), a DNN (Deep Neural Network), an RNN (Recurrent Neural Network), an RBM (Restricted Boltzmann Machine), a DBN (Deep Belief Network), a BRDNN (Bidirectional Recurrent Deep Neural Network), or a deep Q-network.
[0027] For example, a deep learning model for outputting a source data set and a target data set can be driven by an AI neural network model trained through unsupervised learning on basic data. The deep learning model can be configured to increase the ease of data collection and output various data values. The deep learning model has a structure that outputs 3D image data from line data and 2D image data, and at least one of BigScience's bloom and T0pp, EleutherAI's GPT series, Tsinghua UNIV's GLM series, GOOGLE's UL and T5 series, and META AI's OPT series can be used. According to one embodiment, the deep learning model can be implemented using the structure of Microsoft's Chat GPT, Google's BARD series, and NVIDIA's translation service-based Transformer model used as a multiple cloud foundation model. For example, the deep learning model can be trained as a multi-modal model based on line data and image data to implement various imaging. In one embodiment, this can reduce the amount of labeled task-specific training data compared to existing deep learning approaches, and once built, it can be trained on a variety of tasks with a small amount of training data, making data collection and labeling easier and improving accuracy.
[0028] For example, the electronic device (101) may perform a learning process of a deep learning model for outputting a source data set and a target data set by obtaining a result value (output data) using a deep learning model to which arbitrary weights are assigned, comparing the obtained result value with the labeled data of the learning data, and performing backpropagation according to the error to optimize the weights. Specifically, the learning of the deep learning model refers to a process of training the deep learning model based on learning data and labeled data or unlabeled data so that the deep learning model can determine output data for input data. In other words, the deep learning model forms rules for the data and makes a judgment. According to one embodiment, the electronic device (101) may use a plurality of learning algorithms among a plurality of learning algorithms that calculate a predicted value. For example, an ensemble method may be used for the deep learning model, and better prediction performance may be obtained compared to when learning algorithms are used separately. Training a deep learning model can mean adjusting the weights of the model. In some embodiments, various learning methods can be used, including supervised learning, unsupervised learning, reinforcement learning, imitation learning, and federated learning.
[0029] For example, the electronic device (101) may include an evaluation step for evaluating the performance of a deep learning model during the training process of the deep learning model. In the evaluation step, the deep learning model may be evaluated using an evaluation data set. The evaluation of the deep learning model may be a step of evaluating the deep learning model trained through the training step and evaluating new data using the deep learning model. Specifically, the evaluation step may be a step of measuring whether the trained deep learning model is capable of generalizing to new data.
[0030] In operation 205, the electronic device (101) (e.g., the processor (120) of FIG. 1) may input a source data set and a target data set into a second sub-deep learning model to output a mask data set. For example, the second sub-deep learning model may be trained based on a plurality of source data sets, a plurality of target data sets, and a plurality of mask data sets. For example, the second sub-deep learning model may be used with the same structure as the first sub-deep learning model and includes the same functions as those described above, so a detailed description thereof will be omitted.
[0031] In operation 207, the electronic device (101) (e.g., the processor (120) of FIG. 1) may determine whether a first face recognition threshold value is exceeded. For example, the first face recognition threshold value may include a preset first image sharpness percentage value and multiple face restoration algorithms. For example, the electronic device (101) may reconstruct a 3D face using a portrait photo database to use an auto-encoder stored in a main deep learning model, which will be described later, as a face restoration algorithm. Thereafter, the electronic device (101) may use the data for outputting final restored image data based on a mask data set. Specifically, the electronic device (101) may select an algorithm with the best matching among the multiple face restoration algorithms, and then output the final restored image data using the selected face restoration algorithm and the main deep learning model together with the mask data set. In addition, the electronic device (101) is set to provide an optimal algorithm for the facial restoration algorithm, but when it obtains command data for preferentially selecting an algorithm from the user's external electronic device (102, 104, 106), it can set the facial restoration algorithm based on the obtained data.
[0032] For example, the first image sharpness percentage value may be configured as a value (e.g., sharpness 85%) for comparing the image sharpness of at least one of the mask data set and / or the first restored image data, the second restored image data, and the third restored image data described below. For example, the electronic device (101) may compare the first image sharpness percentage value with the sharpness percentage value of the mask data set to determine whether the first face recognition threshold value is exceeded.
[0033] In operation 209, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the first facial recognition threshold value has been exceeded, the electronic device may input the mask data set into the main deep learning model to output final restored image data. Specifically, if the sharpness percentage value of the mask data set (e.g., sharpness 94%) is such that the first facial recognition threshold value has been exceeded, the mask data set may be input into the main deep learning model based on the sharpness percentage value to output final restored image data. This has the advantage of preventing time waste because the final restored image data can be output immediately without a separate algorithm setting because the image data of the mask data set already has very good sharpness. For example, the main deep learning model may be trained based on multiple source data sets for multiple faces, multiple target data sets for multiple faces, multiple mask data sets for multiple faces, final reconstructed image data for multiple reconstructed faces, multiple first reconstructed image data for multiple reconstructed faces, and multiple face reconstructed algorithms and image sharpness evaluation algorithms. For example, the main deep learning model may be a CNN-based deep learning model trained to input at least one of the mask data sets, the first reconstructed image data, the second reconstructed image data, and the third reconstructed image data and output final reconstructed image data. Specifically, the main deep learning model may use at least one of AlexNet, LENET-5, NIN, VGGNet, ResNet, WideResnet, GoogleNet, FractaNet, DenseNet, FitNet, RitResNet, HighwayNet, MobileNet, and DeeplySupervisedNet as a CNN structure.More specifically, LeNet-5 is the most recent LeNet model, developed in the 1990s by Yann LeCun's lab, and can be used to recognize zip codes and numbers. Furthermore, the key point is that LeNet's architecture is not significantly different from today's CNNs, using convolution and subsampling, and connecting feature maps using full connections that flatten them. AlexNet, the winner of the ILSVRC 2012 competition, can be seen as a revolution in deep learning at the time. This is because AlexNet's CNN architecture significantly reduces the top-5 error rate of the past. AlexNet then marked the beginning of the application of CNN techniques to image neural networks, and this architecture is unique in that instead of applying multiple filters simultaneously, it splits them into two halves, enabling analysis using two GPUs. ZFNet is almost identical in structure to AlexNet, and in fact, it only changed some of the parameters used in AlexNet, but its performance was improved. This means that to train a CNN well, you can see how the filters in the intermediate layers are learned. As another example, GoogLeNet is the model that won LSVRC 2014. Compared to AlexNet, it has a deeper depth and wider width, but you can see that the number of parameters is greatly reduced. This is because the concept of the inception module was introduced in GoogLeNet, and the inception module was applied based on the concept that more information could be found when the filter operation is performed nonlinearly, even though the filter operation is a linear operation.The structure of GoogLeNet is a network-like layer within a network structure, and it can be formed into a Network-In-Network (NIN) structure, which can be nonlinear. In addition, the main deep learning model applies the structure of a CNN, and examples of this are not limited. For example, the image sharpness assessment algorithm may include the Laplacian algorithm, the Laplacian of Gaussian (LOG) algorithm, the difference of Gaussians (DOF) algorithm, the Sobel algorithm, the Robert algorithm, and the Prewitt algorithm.
[0034] Meanwhile, in operation 207, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the first facial recognition threshold value is not exceeded, in operation 211, the electronic device (101) (e.g., the processor (120) of FIG. 1) may input the first restored image data into the main deep learning model while setting a setting value in the mask data set. Specifically, the electronic device (101) may determine that the first facial recognition threshold value is not exceeded because the sharpness value of the mask data set does not exceed the first image sharpness percentage value as a result of comparing the first image sharpness percentage value (e.g., sharpness 85%) with the sharpness percentage value of the mask data set (e.g., sharpness 76%). Thereafter, the electronic device (101) can specifically set the setting values of the mask data set as described below, and input the mask data set into the main deep learning model according to the set values to output the first restored image data.
[0035] In operation 213, the electronic device (101) (e.g., the processor (120) of FIG. 1) may determine whether a second facial recognition threshold value is exceeded. Specifically, the second facial recognition threshold value may include a first image sharpness percentage value and a plurality of image sharpness evaluation algorithms. For example, the first image sharpness percentage value may be configured as a value (e.g., sharpness 85%) for comparing image sharpness of at least one of a mask data set and / or first restored image data, second restored image data, and third restored image data described below. For example, the electronic device (101) may compare the first image sharpness percentage value with a sharpness percentage value of the mask data set to determine whether the second facial recognition threshold value is exceeded. To be more specific, the electronic device (101) can determine whether the sharpness percentage value of the first restored image data is greater than or equal to the second image sharpness percentage value.
[0036] Meanwhile, in operation 213, the electronic device (101) (e.g., the processor (120) of FIG. 1) may be configured to repeat the process from determining that the first restored image data is below the first facial recognition threshold to determining that the first restored image data is above the second facial recognition threshold when the first restored image data is below the first facial recognition threshold. The specific details thereof will be described in detail with reference to FIGS. 3 and 4.
[0037] In operation 215, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the first facial recognition threshold value has been exceeded, the electronic device (101) may input the first restored image data into the main deep learning model to output final restored image data. For example, specifically, if the first image sharpness percentage value (e.g., sharpness 98%) is determined to have exceeded the first facial recognition threshold value, the electronic device (101) may input the first restored image data into the main deep learning model to output final restored image data.
[0038]
[0039] FIG. 3 is a flowchart specifically explaining an operation method in which an electronic device outputs output from a mask data set according to various embodiments of the present invention.
[0040] In operation 211a, the electronic device (101) (e.g., the processor (120) of FIG. 1) may extract a sharpness percentage value of the output first restored image data when inputting a mask data set into the main deep learning model. As an example, the electronic device (101) inputs the mask data set into the main deep learning model and outputs the first restored image data. Thereafter, the electronic device (101) may extract a sharpness percentage value of the output first restored image data.
[0041] In operation 213, the electronic device (101) (e.g., the processor (120) of FIG. 1) can determine whether the sharpness percentage value of the first restored image data exceeds the second facial recognition threshold value. The electronic device (101) is operated in the same manner as operation 213 described above, and if the first restored image data exceeds the second facial recognition threshold value, it is operated in the same manner while continuing with operation 215, so that specific details can be omitted. In addition, the reason why the operation is included in the present embodiment is that since the sharpness value of the first restored image data output from the main deep learning model of the mask data set has increased, there is an advantage in that a separate correction process can be omitted by first checking whether it exceeds the second facial recognition threshold value.
[0042] Meanwhile, in operation 213, if it is determined that the sharpness percentage value of the first restored image data does not exceed the first image sharpness percentage value, then in operation 211b, the electronic device (101) (e.g., the processor (120) of FIG. 1 ) may determine whether the sharpness percentage value of the first restored image data is equal to or greater than the second image sharpness percentage value. Specifically, if the electronic device (101) determines that the sharpness percentage value of the first restored image data (e.g., sharpness 79%) does not exceed the first image sharpness percentage value (e.g., sharpness 85%), it may determine that the second facial recognition threshold value is not exceeded, and determine whether the first restored image data exceeds the second image sharpness percentage value (e.g., sharpness 70%). This may be set so that the present embodiment can increase the sharpness step by step before retraining the main deep learning model.
[0043] In operation 211c, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the first restored image data is greater than or equal to the second image sharpness percentage value, the electronic device (101) may input the first restored image data to a main deep learning model while setting it to a fast value and output the corrected second restored image data. Specifically, the electronic device (101) may determine that the sharpness percentage value of the first restored image data (e.g., sharpness 79%) exceeds the second image sharpness percentage value (e.g., sharpness 70%) and may set it to a fast value so as to minimize the correction time thereof, thereby reducing the correction time of the main deep learning model while increasing the sharpness. Additionally, the electronic device (101) may input the first restored image data into the main deep learning model to extract the second restored image data and the sharpness percentage value of the second restored image data, and return to operation 213 to determine whether the second restored image data exceeds the second facial recognition threshold value.
[0044] Meanwhile, in operation 211b, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the first restored image data is less than or equal to the second image sharpness percentage value, then in operation 211d, the electronic device (101) (e.g., the processor (120) of FIG. 1) may determine whether the sharpness percentage value of the first restored image data is greater than or equal to the third image sharpness percentage value. Specifically, if the electronic device (101) determines that the sharpness percentage value of the first restored image data (e.g., sharpness 65%) does not exceed the second image sharpness percentage value (e.g., sharpness 70%), then the electronic device may determine that the second image sharpness percentage value does not exceed the second image sharpness percentage value, and may determine whether the first restored image data exceeds the third image sharpness percentage value (e.g., sharpness 60%). This may be set so that the present embodiment can increase the sharpness step by step before retraining the main deep learning model.
[0045] In operation 211e, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the first restored image data is greater than or equal to the third image sharpness percentage value, the electronic device (101) may set the first restored image data to a normal value and input it to the main deep learning model to output the corrected second restored image data. Specifically, the electronic device (101) may determine that the sharpness percentage value of the first restored image data (e.g., sharpness 65%) exceeds the third image sharpness percentage value (e.g., sharpness 60%), and may minimize the correction time thereof, but unlike operation 211c, may set it to a value with a certain amount of time, thereby reducing the correction time of the main deep learning model while increasing the sharpness. Additionally, the electronic device (101) may input the first restored image data into the main deep learning model to extract the second restored image data and the sharpness percentage value of the second restored image data, and return to operation 213 to determine whether the second restored image data exceeds the second facial recognition threshold value.
[0046] Meanwhile, in operation 211d, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the first restored image data is less than or equal to the third image sharpness percentage value, in operation 211f, the electronic device (101) (e.g., the processor (120) of FIG. 1) may input the first restored image data to the main deep learning model and output the corrected second restored image data while setting the first restored image data to a slow value. Specifically, the electronic device (101) may determine that the sharpness percentage value of the first restored image data (e.g., sharpness 49%) exceeds the third image sharpness percentage value (e.g., sharpness 60%), and may minimize the correction time thereof, but unlike operation 211e, may set the value to a value that allocates more time, thereby sufficiently providing the correction time of the main deep learning model, thereby increasing the sharpness of the second restored image data. Additionally, the electronic device (101) may input the first restored image data into the main deep learning model to extract the second restored image data and the sharpness percentage value of the second restored image data, and return to operation 213 to determine whether the second restored image data exceeds the second facial recognition threshold value.
[0047] In addition, the electronic device (101) can automatically set the learning time of the first restoration image data based on the 211c operation, the 211e operation, and the 211f operation, but when command data for preferentially selecting an algorithm selection and a time selection is obtained from the user's external electronic device (102, 104, 106), the face restoration algorithm can be set based on the obtained data.
[0048]
[0049] FIG. 4 is a flowchart specifically explaining an operation method in which an electronic device outputs first restored image data according to various embodiments of the present invention.
[0050] In operation 211a, the electronic device (101) (e.g., the processor (120) of FIG. 1) may extract a sharpness percentage value of the output second restored image data when inputting the first restored image data into the main deep learning model. For example, the electronic device (101) inputs the first restored image data into the main deep learning model and outputs the second restored image data. Thereafter, the electronic device (101) may extract a sharpness percentage value of the output second restored image data.
[0051] In operation 213, the electronic device (101) (e.g., the processor (120) of FIG. 1) can determine whether the sharpness percentage value of the second restored image data exceeds the second facial recognition threshold value. The electronic device (101) is operated in the same manner as operation 213 described above, and if the second restored image data exceeds the second facial recognition threshold value, it is operated in the same manner while continuing with operation 215, so that specific details can be omitted. In addition, the reason why the operation is included in the present embodiment is that since the sharpness value of the second restored image data output from the main deep learning model of the first restored image data has increased, there is an advantage in that a separate correction process can be omitted by first checking whether it exceeds the second facial recognition threshold value.
[0052] Meanwhile, in operation 213, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the second restored image data is less than or equal to the second image sharpness percentage value, then in operation 211b', the electronic device (101) (e.g., the processor (120) of FIG. 1) may determine whether the sharpness percentage value of the second restored image data is greater than or equal to the third image sharpness percentage value. Specifically, if the electronic device (101) determines that the sharpness percentage value of the second restored image data (e.g., sharpness 73%) does not exceed the first image sharpness percentage value (e.g., sharpness 85%), then the electronic device may determine that the second facial recognition threshold value is not exceeded, and determine whether the second restored image data exceeds the second image sharpness percentage value (e.g., sharpness 70%). This may be set so that the present embodiment can increase the sharpness step by step before retraining the main deep learning model.
[0053] In operation 211c, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the second restored image data is greater than or equal to the second image sharpness percentage value, the electronic device (101) may input the second restored image data to a main deep learning model while setting it to a fast value and output the corrected third restored image data. Specifically, the electronic device (101) may determine that the sharpness percentage value of the second restored image data (e.g., sharpness 73%) exceeds the third image sharpness percentage value (e.g., sharpness 70%) and may set it to a fast value so as to minimize the correction time thereof, thereby reducing the correction time of the main deep learning model while increasing the sharpness. Additionally, the electronic device (101) may input the first restored image data into the main deep learning model to extract the second restored image data and the sharpness percentage value of the second restored image data, and return to operation 213 to determine whether the second restored image data exceeds the second facial recognition threshold value.
[0054] Meanwhile, in operation 211b', if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the first restored image data is less than or equal to the second image sharpness percentage value, in operation 211d', the electronic device (101) (e.g., the processor (120) of FIG. 1) may input the second restored image data to a normal value and output the corrected third restored image data to the main deep learning model. Specifically, the electronic device (101) may determine that the sharpness percentage value of the second restored image data (e.g., sharpness 61%) exceeds the second image sharpness percentage value (e.g., sharpness 70%), and may minimize the correction time thereof, but unlike operation 211c, may set the value to a value with a certain amount of time, thereby reducing the correction time of the main deep learning model while increasing the sharpness. Additionally, the electronic device (101) may input the second restored image data into the main deep learning model to extract the third restored image data and the sharpness percentage value of the third restored image data, and return to operation 213 to determine whether the third restored image data exceeds the second facial recognition threshold value.
[0055] This means that in the case of the contents of Fig. 4, the first restored image data has already been learned once compared to the mask data set, so there is no need to separately set the slowness, thereby preventing waste of time in advance, and based on this, the main deep learning model also has the advantage of learning multiple second restored image data and multiple third restored image data, thereby further reducing the learning and output time.
[0056] In addition, the electronic device (101) can automatically set the learning time of the second restoration image data based on the 211c operation, the 211e operation, and the 211f operation, but when command data for preferentially selecting an algorithm selection and a time selection is obtained from the user's external electronic device (102, 104, 106), the face restoration algorithm can be set based on the obtained data.
[0057]
[0058] According to various embodiments, an electronic device for performing face restoration based on a deep learning model comprises: a communication interface; a memory; and a processor; , wherein the processor obtains face source data and face target data for restoration through the communication interface, inputs the face source data and the face target data into at least one first sub-deep learning model, respectively, to output a source data set and a target data set, inputs the source data set and the target data set into a second sub-deep learning model to output a mask data set, and if it is determined that the mask data set is greater than or equal to a preset first facial recognition threshold value, inputs the mask data set into a main deep learning model to output final restored image data, and if it is determined that the mask data set is less than or equal to the first facial recognition threshold value, inputs the mask data set into the main deep learning model while setting a setting value to the mask data set to output first restored image data, and if it is determined that the first restored image data is greater than or equal to a preset second facial recognition threshold value, inputs the first restored image data into the main deep learning model to output the final restored image data, and if the first restored image data is determined to be greater than or equal to the second If it is determined to be below the face recognition threshold, the process is set to repeat from the case where the first restored image data is below the first face recognition threshold to the case where it is determined to be above the second face recognition threshold, and the main deep learning model is trained based on a plurality of source data sets for a plurality of faces, a plurality of target data sets for a plurality of faces, a plurality of mask data sets for a plurality of faces, final restored image data for a plurality of restored faces, a plurality of first restored image data for a plurality of restored faces, a plurality of face restoration algorithms, and an image clarity evaluation algorithm.
[0059] According to various embodiments, the at least one first sub-deep learning model includes a 1-1 sub-deep learning model and a 1-2 sub-deep learning model, and the processor is configured to input the facial source data into the 1-1 sub-deep learning model to output the source data set, and to input the facial target data into the 1-2 sub-deep learning model to output the target data set.
[0060] According to various embodiments, the first-first sub-deep learning model is trained based on a plurality of face source data and a plurality of the source data sets, and the first-second sub-deep learning model is trained based on a plurality of face target data and a plurality of the target data sets.
[0061] According to various embodiments, the first face recognition threshold value includes a preset first image sharpness percentage value and the plurality of face restoration algorithms, the second face recognition threshold value includes the first image sharpness percentage value and the plurality of image sharpness evaluation algorithms, and the image sharpness evaluation algorithms include a Laplacian algorithm, a laplacian of gaussian (LOG) algorithm, a difference of gaussian (DOF) algorithm, a Sobel algorithm, a Robert algorithm, and a Prewitt algorithm.
[0062] According to various embodiments, when the mask data set that does not exceed the first facial recognition threshold value is input to the main deep learning model, the processor extracts a sharpness percentage value of the output first restored image data, and if it is determined that the sharpness percentage value of the first restored image data is greater than or equal to a second image sharpness percentage value that is set smaller than the first image sharpness percentage value, the processor outputs second restored image data that has been corrected by setting the first restored image data to a fast value and inputting it to the main deep learning model, and if it is determined that the sharpness percentage value of the first restored image data is greater than or equal to a third image sharpness percentage value that is set smaller than the second image sharpness percentage value, the processor outputs second restored image data that has been corrected by setting the first restored image data to a normal value and inputting it to the main deep learning model, and if it is determined that the sharpness percentage value of the first restored image data is less than or equal to the third image sharpness percentage value, the processor outputs the first restored image data to the main deep learning model, while setting the first restored image data to a slow value. It is set to output the second restored image data that has been corrected by inputting it.
[0063] According to various embodiments, the processor is configured to, when inputting the first restored image data that does not exceed the first facial recognition threshold value and / or the second facial recognition threshold value into the main deep learning model, extract a sharpness percentage value of the output second restored image data, and if it is determined that the sharpness percentage value of the second restored image data is greater than or equal to a second image sharpness percentage value that is set to be smaller than the first image sharpness percentage value, input the second restored image data to the main deep learning model while setting it to a fast value and output the corrected third restored image data, and if it is determined that the sharpness percentage value of the second restored image data is less than or equal to the second image sharpness percentage value, input the second restored image data to the main deep learning model while setting it to a normal value and output the corrected third restored image data.
[0064]
[0065] The term "module" or "part" used in this document includes a unit composed of hardware, software, or firmware, and can be used interchangeably with terms such as logic, logic block, component, or circuit, for example. The "module" or "part" can be an integrally configured component or a minimum unit or a part thereof that performs one or more functions. The "module" or "part" can be implemented mechanically or electronically, and can include, for example, an ASIC (application-specific integrated circuit) chip, FPGAs (field-programmable gate arrays), or a programmable logic device, known or to be developed in the future, that performs certain operations, and can be executed by the processor (120). At least a part of the device (e.g., modules or functions thereof) or method (e.g., operations) according to various embodiments can be implemented as instructions stored in a computer-readable storage medium (e.g., memory (130)) in the form of a program module. When the above command is executed by a processor (e.g., processor (120)), the processor can perform a function corresponding to the command. The computer-readable recording medium may include a hard disk, a floppy disk, a magnetic medium (e.g., a magnetic tape), an optical recording medium (e.g., a CD-ROM, a DVD, a magneto-optical medium (e.g., a floptical disk), a built-in memory, etc. The command may include a code generated by a compiler or a code executable by an interpreter. A module or program module according to various embodiments may include at least one or more of the above-described components, some of which may be omitted, or other components may be further included. Operations performed by a module, a program module, or other components according to various embodiments may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
[0066] The embodiments disclosed in this document are presented for the purpose of explaining and understanding the disclosed technical content, and do not limit the scope of the present disclosure. Therefore, the scope of the present disclosure should be interpreted to include all modifications or various other embodiments based on the technical concepts of the present disclosure.
Claims
1. An electronic device for performing face restoration based on a deep learning model, communication interface; memory; and Processor; including, The above processor, Through the above communication interface, facial source data and facial target data for restoration are obtained, Inputting the above face source data and the above face target data into at least one first sub-deep learning model to output a source data set and a target data set, The above source data set and the above target data set are input into the second sub-deep learning model to output a mask data set, If it is determined that the above mask data set is greater than or equal to the preset first facial recognition threshold, the above mask data set is input into the main deep learning model to output the final restored image data. If it is determined that the above mask data set is below the first facial recognition threshold, the first restored image data is output by inputting the main deep learning model while setting the setting value to the above mask data set, If it is determined that the first restored image data is greater than or equal to the preset second facial recognition threshold value, the first restored image data is input into the main deep learning model to output the final restored image data, and If the first restored image data is determined to be below the second facial recognition threshold, the process is set to repeat from the case where the first restored image data is below the first facial recognition threshold until it is determined to be above the second facial recognition threshold. The above main deep learning model is, A method for training a plurality of face restoration algorithms and an image sharpness evaluation algorithm, the method comprising: learning a plurality of source data sets for a plurality of faces, a plurality of target data sets for a plurality of faces, a plurality of mask data sets for a plurality of faces, a plurality of final restored image data for a plurality of restored faces, a plurality of first restored image data for a plurality of restored faces, and a plurality of face restoration algorithms and an image sharpness evaluation algorithm. Electronic devices.
2. In paragraph 1, The at least one first sub-deep learning model, Includes the 1-1 sub-deep learning model and the 1-2 sub-deep learning model, The above processor, Inputting the above facial source data into the 1-1 sub-deep learning model to output the above source data set, and The above facial target data is input into the first-second sub-deep learning model and set to output the target data set. Electronic devices.
3. In paragraph 2, The above 1-1 sub-deep learning model is, It is trained based on multiple face source data and multiple above-mentioned source data sets, The above 1-2 sub-deep learning models are, Learned based on multiple face target data and multiple sets of said target data, Electronic devices.
4. In paragraph 1, The above first facial recognition threshold value is, comprising a preset first image sharpness percentage value and the plurality of face restoration algorithms, The above second facial recognition threshold value is, comprising the first image sharpness percentage value and a plurality of the image sharpness evaluation algorithms, The above image clarity evaluation algorithm is, Including the Laplacian algorithm, the laplacian of Gaussian (LOG) algorithm, the difference of Gaussian (DOF) algorithm, the Sobel algorithm, the Robert algorithm, and the Prewitt algorithm. Electronic devices.
5. In paragraph 4, The above processor, When the mask data set that does not exceed the first facial recognition threshold is input to the main deep learning model, the sharpness percentage value of the output first restored image data is extracted, If it is determined that the sharpness percentage value of the first restored image data is greater than or equal to the second image sharpness percentage value set to be smaller than the first image sharpness percentage value, the first restored image data is input to the main deep learning model while setting it to a fast value, and the corrected second restored image data is output. If it is determined that the sharpness percentage value of the first restored image data is greater than or equal to a third image sharpness percentage value set smaller than the second image sharpness percentage value, the first restored image data is input to the main deep learning model while setting it to a normal value, and the second restored image data is output as corrected, and If it is determined that the sharpness percentage value of the first restored image data is less than or equal to the third image sharpness percentage value, the first restored image data is set to a slow value and input to the main deep learning model to output the corrected second restored image data. Electronic devices.
6. In paragraph 4, The above processor, When the first restored image data that does not exceed the first facial recognition threshold value and / or the second facial recognition threshold value is input to the main deep learning model, a sharpness percentage value of the output second restored image data is extracted, If it is determined that the sharpness percentage value of the second restored image data is greater than or equal to the second image sharpness percentage value set smaller than the first image sharpness percentage value, the second restored image data is input to the main deep learning model while setting it to a fast value, and the corrected third restored image data is output, and If it is determined that the sharpness percentage value of the second restored image data is less than or equal to the second image sharpness percentage value, the second restored image data is set to be input to the main deep learning model while setting it to a normal value to output the corrected third restored image data. Electronic devices.
Citation Information
Patent Citations
Pharmaceutical composition for treating ischemic cerebrovascular disease
KR1020240125820A
Toilet bowl cleaner with replacement cycle notification function
KR1020250138516A
Video conferencing based on adaptive face re-enactment and face restoration
US20220398692A1
Light estimation method for three-dimensional (3D) rendered objects
US20230419599A1
KR20220131807A