Electronic device for performing cacial reconstruction based on deep learning model and method operation thereof

KR103013341B1Active Publication Date: 2026-09-02STUDIO METAK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
KR1020240072432
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-03
Publication Date
2026-09-02
Estimated Expiration
2044-06-03

Smart Images

  • Figure R1020240072432_ABST
    Figure R1020240072432_ABST
Patent Text Reader

Abstract

According to various embodiments, an electronic device for performing face reconstruction based on a deep learning model comprises: a communication interface; a memory; and a processor; The method includes, wherein the processor acquires face source data and face target data for restoration through the communication interface, inputs the face source data and the face target data into at least one first sub-deep learning model to output a source data set and a target data set, inputs the source data set and the target data set into a second sub-deep learning model to output a mask data set, and if it is determined that the mask data set is greater than or equal to a preset first face recognition threshold, inputs the mask data set into a main deep learning model to output final restored image data, and if it is determined that the mask data set is less than or equal to the first face recognition threshold, inputs the mask data set into the main deep learning model while setting a value to output first restored image data, and if it is determined that the first restored image data is greater than or equal to a preset second face recognition threshold, inputs the first restored image data into the main deep learning model to output the final restored image data, and the first restored image data is equal to the second face recognition threshold If it is determined that the value is less than or equal to the first face recognition threshold, the process is set to repeat from the case where the first restored image data is less than or equal to the first face recognition threshold until it is determined to be greater than or equal to the second face recognition threshold, and the main deep learning model is trained based on a plurality of source data sets for a plurality of faces, a plurality of target data sets for a plurality of faces, a plurality of mask data sets for a plurality of faces, final restored image data for a plurality of restored faces, a plurality of first restored image data for a plurality of restored faces, a plurality of face restoration algorithms, and an image clarity evaluation algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Various embodiments of the present invention relate to an electronic device for performing face restoration based on a deep learning model and a method for operating the same, and more specifically, to an electronic device for performing face restoration based on a deep learning model and a method for operating the same that can easily output restored image data by inputting basic data into a deep learning model to restore the faces of past people, past animals, and predicted future people. Background Technology

[0002] Recently, deepfake, or face transformation, refers to the creation of new content by synthesizing or replacing the face of a subject appearing in an original image or video with a face from another image or video. While there are issues such as the creation of fake videos by synthesizing the faces of celebrities, the film and video industry also uses this technology to synthesize the faces of the deceased or to meticulously recreate the appearance of an actor in their younger years. Although processes and technologies involving the creation and synthesis of 3D models, such as digital doubles, already existed, the industry is showing great interest in face transformation due to the advantage of significantly reducing costs.

[0003] However, conventionally, the process of creating a deep learning model to transform a source face into a target face based on the source face dataset and the target face dataset required a lot of time, and there was a limitation in that if sufficient training time was not reflected, a blurry synthesis result was obtained.

[0004] In particular, conventionally, a significant amount of time is required to enhance clarity after the initial facial shape is realized, which has been a problem that reduces production efficiency in film and video content production. Prior art literature

[0005] Korean Patent Publication No. 10-2547630 (June 21, 2023) Korean Patent Publication No. 10-2353837 (January 17, 2022) The problem to be solved

[0006] Accordingly, the present invention can provide an electronic device and a method for operating the same that can easily output face restoration image data while reducing wasted time and simultaneously inputting and learning face source data and face target data into a deep learning model and verifying the accuracy of the output value. means of solving the problem

[0007] According to various embodiments, an electronic device for performing face reconstruction based on a deep learning model comprises: a communication interface; a memory; and a processor; The method includes, wherein the processor acquires face source data and face target data for restoration through the communication interface, inputs the face source data and the face target data into at least one first sub-deep learning model to output a source data set and a target data set, inputs the source data set and the target data set into a second sub-deep learning model to output a mask data set, and if it is determined that the mask data set is greater than or equal to a preset first face recognition threshold, inputs the mask data set into a main deep learning model to output final restored image data, and if it is determined that the mask data set is less than or equal to the first face recognition threshold, inputs the mask data set into the main deep learning model while setting a value to output first restored image data, and if it is determined that the first restored image data is greater than or equal to a preset second face recognition threshold, inputs the first restored image data into the main deep learning model to output the final restored image data, and the first restored image data is equal to the second face recognition threshold If it is determined that the value is less than or equal to the first face recognition threshold, the process is set to repeat from the case where the first restored image data is less than or equal to the first face recognition threshold until it is determined to be greater than or equal to the second face recognition threshold, and the main deep learning model is trained based on a plurality of source data sets for a plurality of faces, a plurality of target data sets for a plurality of faces, a plurality of mask data sets for a plurality of faces, final restored image data for a plurality of restored faces, a plurality of first restored image data for a plurality of restored faces, a plurality of face restoration algorithms, and an image clarity evaluation algorithm. Effects of the invention

[0008] According to various embodiments, the present embodiment has the advantage of being able to output highly clear face restoration image data in a short time by learning face source data and face target data to form a primary data set, inputting them into a deep learning model based on them to minimize output time while maintaining high clarity, and continuously verifying this through a step. Brief explanation of the drawing

[0009] FIG. 1 illustrates a block diagram of an electronic device and network according to various embodiments of the present invention. FIG. 2 is a flowchart illustrating a method of operation of an electronic device according to various embodiments of the present invention. FIG. 3 is a flowchart for specifically explaining a method of operation in which an electronic device outputs from a mask data set according to various embodiments of the present invention. FIG. 4 is a flowchart for specifically explaining a method of operation in which an electronic device outputs from a first restored image data according to various embodiments of the present invention. Specific details for implementing the invention

[0010] Hereinafter, various embodiments of this document are described with reference to the accompanying drawings. The embodiments and the terms used therein are not intended to limit the technology described in this document to specific embodiments and should be understood to include various modifications, equivalents, and / or substitutions of said embodiments. In relation to the description of the drawings, similar reference numerals may be used for similar components. A singular expression may include a plural expression unless the context clearly indicates otherwise. In this document, expressions such as "A or B" or "at least one of A and / or B" may include all possible combinations of items listed together. Expressions such as "first," "second," "first," or "second" may modify said components regardless of order or importance and are used only to distinguish one component from another and do not limit said components. When it is mentioned that a certain (e.g., 1st) component is "(functionally or telecommunicationally) connected" or "connected" to another (e.g., 2nd) component, said certain component may be directly connected to said other component or connected through another component (e.g., 3rd component).

[0011] In this document, "configured to" may be used interchangeably with, depending on the context, for example, hardware- or software-wise, "suitable for," "capable of," "modified to," "made to," "capable of," or "designed to." In some cases, the expression "device configured to" may mean that the device is "capable of" in conjunction with other devices or components. For example, the phrase "processor configured to perform A, B, and C" may mean a dedicated processor for performing the corresponding operations (e.g., an embedded processor), or a general-purpose processor capable of performing the corresponding operations by executing one or more software programs stored in a memory device (e.g., a CPU or application processor).

[0012] An electronic device according to various embodiments of the present document may include, for example, at least one of a smartphone, a tablet PC, a desktop PC, a laptop PC, a netbook computer, a workstation, and a server.

[0013] Referring to FIG. 1, an electronic device (101) within a network environment (100) in various embodiments is described. The electronic device (101) may include a bus (110), a processor (120), a memory (130), an input / output interface (140), a display (150), and a communication interface (160). In some embodiments, the electronic device (101) may omit at least one of the components or additionally include other components. The bus (110) may include a circuit that connects the components (110-170) to each other and transmits communication (e.g., control messages or data) between the components. The processor (120) may include one or more of a central processing unit, an application processor, or a communication processor (CP). The processor (120) may, for example, perform operations or data processing regarding the control and / or communication of at least one other component of the electronic device (101).

[0014] The memory (130) may include volatile and / or non-volatile memory. The memory (130) may store instructions or data related to at least one other component of the electronic device (101), for example. According to one embodiment, the memory (130) may store software and / or a program (140).

[0015] The input / output interface (140) can, for example, transmit commands or data input from a patient or other external device to other component(s) of the electronic device (101), or output commands or data received from other component(s) of the electronic device (101) to the patient or other external device.

[0016] The display (150) may include, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a microelectromechanical system (MEMS) display, or an electronic paper display. The display (150) may display various content (e.g., text, images, videos, icons, and / or symbols, etc.) to a patient, for example. The display (150) may include a touch screen and may receive touch, gesture, proximity, or hovering input using, for example, an electronic pen or a part of the patient's body. The communication interface (160) may establish communication between, for example, the electronic device (101) and an external device (e.g., a first external electronic device (102), a second external electronic device (104), or a server (108)). For example, the communication interface (160) can be connected to a network (162) via wireless or wired communication to communicate with an external device (e.g., a second external electronic device (104) or a server (108)).

[0017] Wireless communication may include cellular communication using at least one of, for example, LTE, LTE-A (LTE Advance), CDMA (code division multiple access), WCDMA (wideband CDMA), UMTS (universal mobile telecommunications system), WiBro (Wireless Broadband), or GSM (Global System for Mobile Communications). According to one embodiment, wireless communication may include at least one of, for example, WiFi (wireless fidelity), Bluetooth, Bluetooth Low Energy (BLE), Zigbee, NFC (near field communication), Magnetic Secure Transmission, Radio Frequency (RF), or Body Area Network (BAN). According to one embodiment, wireless communication may include GNSS. GNSS may be, for example, GPS (Global Positioning System), Glonass (Global Navigation Satellite System), Beidou Navigation Satellite System (hereinafter "Beidou"), or Galileo, the European global satellite-based navigation system. Hereinafter, in this document, "GPS" may be used interchangeably with "GNSS". Wired communication may include at least one of, for example, USB (universal serial bus), HDMI (high definition multimedia interface), RS-232 (recommended standard 232), power line communication, or POTS (plain old telephone service).The network (162) may include at least one of a telecommunications network, for example, a computer network (e.g., LAN or WAN), the Internet, or a telephone network.

[0018] Each of the first and second external electronic devices (102, 104, 106) may be the same or a different type of device as the electronic device (101). According to various embodiments, all or part of the operations performed on the electronic device (101) may be performed on one or more other electronic devices (e.g., electronic devices (102, 104, 106), or a server (108). According to one embodiment, when the electronic device (101) needs to perform a function or service automatically or upon request, the electronic device (101) may request at least some of the associated functions from another device (e.g., electronic devices (102, 104, 106), or a server (108)) instead of performing the function or service itself or additionally. The other electronic device (e.g., electronic devices (102, 104, 106), or a server (108)) may perform the requested function or additional functions and transmit the result to the electronic device (101). The electronic device (101) may provide the requested function or service by processing the received result as is or additionally. For this purpose, for example, cloud computing, distributed computing, or client-server computing technologies may be used.

[0020] FIG. 2 is a flowchart illustrating a method of operation of an electronic device according to various embodiments of the present invention.

[0022] In operation 201, an electronic device (101) (e.g., the processor (120) of FIG. 1) may acquire face source data and face target data for restoration through a communication interface. As an example, the face to be restored may consist of the face of a person who is no longer alive and whose photograph is unclear. Additionally, the face to be restored may consist of the faces of extinct animals other than humans. Additionally, the face to be restored may consist of the faces of the user's children requested from the user's external electronic device (102, 104, 106) in the future, and may consist of other faces. As an example, the face source data may consist of image data and / or video data containing the source face extracted by transformation among the faces to be restored. As an example, the face target data may consist of image data and / or video data containing the target face to be modified among the faces to be restored.

[0023] In operation 203, the electronic device (101) (e.g., the processor (120) of FIG. 1) may input face source data and face target data into at least one first sub-deep learning model, respectively, to output a source data set and a target data set. As an example, the at least one first sub-deep learning model may include a first-1 sub-deep learning model and a first-2 sub-deep learning model. As an example, the electronic device (101) (e.g., the processor (120) of FIG. 1) may input face source data into the first-1 sub-deep learning model to output a source data set, and input face target data into the first-2 sub-deep learning model to output a target data set.

[0024] For example, the first sub-deep learning model can be operated as a CNN-based deep learning model and can be operated as a deep learning model to be input for each data. For example, it can be input into a CNN-based deep learning model. For example, the CNN-based deep learning model may include a Deep Neural Network (DNN). For example, the artificial intelligence model may include, but is not limited to, CNN (Convolutional Neural Network), DNN (Deep Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), BRDNN (Bidirectional Recurrent Deep Neural Network), or Deep Q-Networks.

[0025] As an example, a deep learning model for outputting source data sets and target data sets can be driven by an AI neural network model trained through unsupervised learning on basic data. The deep learning model can be configured to increase the ease of data collection and to provide various output values ​​for the data. The deep learning model is structured to output 3D image data from line data and 2D image data, and at least one of BigScience’s bloom and T0pp, EleutherAI’s GPT series, Tsinghua UNIV’s GLM series, GOOGLE’s UL and T5 series, and META AI’s OPT series may be used. According to one embodiment, the deep learning model can be implemented using a plurality of Cloud Foundation models, utilizing the structure of Microsoft’s Chat GPT, Google’s BARD series, and NVIDIA’s translation service-based Transformer model. For example, the deep learning model can be trained as a multi-modal based model based on line data and image data to implement various image renderings. In one embodiment, this can reduce the amount of labeled training data per task compared to existing deep learning methods, and once established, various training can be performed with a small amount of training data, making data collection and labeling easier and improving accuracy.

[0026] As an example, the electronic device (101) may be used as a training process for a deep learning model to output a source data set and a target data set, by obtaining a result value (output data) using a deep learning model to which arbitrary weights are assigned, comparing the obtained result value with the labeled data of the training data, and performing backpropagation according to the error to optimize the weights. Specifically, training of the deep learning model refers to a process of training the deep learning model based on training data and labeled data or unlabeled data so that the deep learning model can determine output data for the input data. That is, the deep learning model forms rules and makes judgments regarding the data. According to one embodiment, the electronic device (101) may use a plurality of learning algorithms among a plurality of learning algorithms that calculate predicted values. For example, an ensemble method may be used in the deep learning model, and better prediction performance can be obtained compared to when the learning algorithms are used separately. Training a deep learning model may mean adjusting the weights of the model. According to one embodiment, as a learning method, various methods such as supervised learning, unsupervised learning, reinforcement learning, imitation learning, and federated learning may be used.

[0027] As an example, the electronic device (101) may include an evaluation step for evaluating the performance of the deep learning model in the learning process of the deep learning model. In the evaluation step, the deep learning model may be evaluated using an evaluation data set. The evaluation of the deep learning model may be a step of evaluating the deep learning model learned by the learning step and evaluating new data using the deep learning model. Specifically, the evaluation step may be a step of measuring whether the learned deep learning model is capable of generalization to new data.

[0028] In operation 205, the electronic device (101) (e.g., the processor (120) of FIG. 1) can input a source data set and a target data set into a second sub-deep learning model to output a mask data set. As an example, the second sub-deep learning model can be trained based on a plurality of source data sets, a plurality of target data sets, and a plurality of mask data sets. As an example, the second sub-deep learning model can be used with the same structure as the first sub-deep learning model and includes the same functions as described above, so a detailed description is omitted.

[0029] In operation 207, the electronic device (101) (e.g., the processor (120) of FIG. 1) can determine whether a first face recognition threshold is exceeded. As an example, the first face recognition threshold may include a preset first image clarity percentage value and a plurality of face restoration algorithms. As an example, the electronic device (101) can reconstruct a 3D face using a portrait database to use an autoencoder stored in the main deep learning model described later for the face restoration algorithm. Subsequently, the electronic device (101) can use the mask data set as data to output final restored image data. Specifically, the electronic device (101) can select the algorithm with the best match among the plurality of face restoration algorithms, and then output the final restored image data using the selected face restoration algorithm along with the mask data set and the main deep learning model. Additionally, the electronic device (101) is configured to provide an optimal algorithm for the face restoration algorithm, but if it obtains command data for prioritizing the selection of an algorithm from the user's external electronic device (102, 104, 106), it can set the face restoration algorithm based on the obtained data.

[0030] As an example, the first image clarity percentage value may be composed of a value for comparing the image clarity of at least one of the mask data set and / or the first reconstructed image data, the second reconstructed image data, and the third reconstructed image data to be described later (e.g., clarity 85%). As an example, the electronic device (101) may determine whether the first image clarity percentage value exceeds the clarity percentage value of the mask data set by comparing the first image clarity percentage value with the clarity percentage value of the mask data set.

[0031] In operation 209, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the first face recognition threshold has been exceeded, it can input the mask data set into the main deep learning model to output the final reconstructed image data. Specifically, if the sharpness percentage value of the mask data set (e.g., sharpness 94%) is high, it can be determined that the first face recognition threshold has been exceeded, and based on this, the final reconstructed image data can be output by inputting it into the main deep learning model. This has the advantage of preventing wasted time because the final reconstructed image data can be output immediately without a separate algorithm setting since the image data of the mask data set already has very good sharpness. As an example, the main deep learning model can be trained based on a plurality of source data sets for a plurality of faces, a plurality of target data sets for a plurality of faces, a plurality of mask data sets for a plurality of faces, final reconstructed image data for a plurality of reconstructed faces, a plurality of first reconstructed image data for a plurality of reconstructed faces, a plurality of face reconstruction algorithms, and an image sharpness evaluation algorithm. As an example, the main deep learning model may be a CNN-based deep learning model trained to receive at least one image data among a mask dataset, a first reconstructed image data, a second reconstructed image data, and a third reconstructed image data as input and output a final reconstructed image data. Specifically, as a CNN structure, the main deep learning model may utilize at least one of AlexNet, LENET-5, NIN, VGGNet, ResNet, WideResNet, GoogleNet, FractaNet, DenseNet, FitNet, RitResNet, HighwayNet, MobileNet, and DeeplySupervisedNet.More specifically, LeNet-5 is the most recent model among the LeNet models, developed in the 1990s at Yann LeCun's lab, and can be utilized for recognizing zip codes or numbers. Furthermore, the key point is that the LeNet structure is not significantly different from current CNNs; it employs convolution and subsampling, and allows for full-connection by flattening feature maps into a straight line. Next, AlexNet is the model that won ILSVRC 2012, and it can be considered to have revolutionized deep learning at that time. This is because AlexNet, with its CNN structure, can significantly reduce the top 5 errors of the past. Subsequently, AlexNet marked the beginning of the application of CNN techniques in image neural networks. While it proceeds with this structure, a unique aspect is that instead of applying multiple filters at once, it splits the process on both sides, allowing for analysis to be performed using two GPUs. ZFNet is very similar in structure to AlexNet, and its performance was improved with only minor modifications to the parameters used in AlexNet; this demonstrates how filters are learned in the intermediate layers when training a CNN effectively. As another example, GoogLeNet is the winning model from LSVRC 2014; compared to AlexNet, it has greater depth and width, yet the number of parameters has been significantly reduced. This is because the concept of the Inception Module was introduced in GoogLeNet. It was applied based on the idea that performing filter operations non-linearly, rather than linearly, might allow for the discovery of more information.The structure of GoogLeNet consists of a network-like layer within the network structure, which forms a Network In Network (NIN) structure and can be non-linear. Additionally, the main deep learning model applies the structure of a CNN, and examples thereof are not limited. As examples, image sharpness evaluation algorithms may include the Laplacian algorithm, the Laplacian of Gaussian (LOG) algorithm, the Difference of Gaussian (DOF) algorithm, the Sobel algorithm, the Robert algorithm, and the Prewitt algorithm.

[0032] Meanwhile, in operation 207, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the first face recognition threshold has not been exceeded, in operation 211, the electronic device (101) (e.g., the processor (120) of FIG. 1) can input the first reconstructed image data into the main deep learning model while setting a value in the mask data set. Specifically, the electronic device (101) can determine that the first face recognition threshold has not been exceeded because, as a result of comparing the first image clarity percentage value (e.g., clarity 85%) with the clarity percentage value of the mask data set (e.g., clarity 76%), the clarity value of the mask data set has not exceeded the first image clarity percentage value. Afterwards, the electronic device (101) can specifically set the setting values ​​of the mask data set as described below, and input the mask data set into the main deep learning model according to the set values ​​to output the first restored image data.

[0033] In operation 213, the electronic device (101) (e.g., the processor (120) of FIG. 1) can determine whether a second facial recognition threshold is exceeded. Specifically, the second facial recognition threshold may include a first image clarity percentage value and a plurality of image clarity evaluation algorithms. As an example, the first image clarity percentage value may be composed of a mask data set and / or a value for comparing the image clarity of at least one of the first reconstructed image data, the second reconstructed image data, and the third reconstructed image data to be described later (e.g., clarity 85%). As an example, the electronic device (101) can determine whether the second facial recognition threshold is exceeded by comparing the first image clarity percentage value with the clarity percentage value of the mask data set. To describe this specifically, the electronic device (101) can determine whether the clarity percentage value of the first restored image data is greater than or equal to the clarity percentage value of the second image.

[0034] Meanwhile, in operation 213, the electronic device (101) (e.g., the processor (120) of FIG. 1) may be configured to repeat the process from when the first reconstructed image data is below the first facial recognition threshold until when the first reconstructed image data is above the second facial recognition threshold. The specific details thereof will be explained in detail with reference to FIG. 3 and FIG. 4.

[0035] In operation 215, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the first facial recognition threshold has been exceeded, it can input the first reconstructed image data into the main deep learning model and output the final reconstructed image data. As an example, specifically, if the first image clarity percentage value (e.g., clarity 98%) is the first, it can be determined that the first facial recognition threshold has been exceeded. Afterward, the electronic device (101) can input the first reconstructed image data into the main deep learning model and output the final reconstructed image data.

[0037] FIG. 3 is a flowchart for specifically explaining a method of operation in which an electronic device outputs from a mask data set according to various embodiments of the present invention.

[0038] In operation 211a, the electronic device (101) (e.g., the processor (120) of FIG. 1) can extract a sharpness percentage value of the output first reconstructed image data when the mask data set is input into the main deep learning model. As an example, the electronic device (101) inputs the mask data set into the main deep learning model and outputs the first reconstructed image data. Afterward, the electronic device (101) can extract a sharpness percentage value of the output first reconstructed image data.

[0039] In operation 213, the electronic device (101) (e.g., the processor (120) of FIG. 1) can determine whether the sharpness percentage value of the first reconstructed image data exceeds the second face recognition threshold. The electronic device (101) operates in the same manner as the operation 213 described above, and if the first reconstructed image data exceeds the second face recognition threshold, it proceeds to operation 215 and operates in the same manner, so specific details can be omitted. Furthermore, the reason the operation is included in this embodiment is that since the sharpness value of the first reconstructed image data output from the main deep learning model has increased in the mask data set, there is an advantage in that a separate correction process can be omitted by first checking whether it exceeds the second face recognition threshold.

[0040] Meanwhile, in operation 213, if it is determined that the sharpness percentage value of the first reconstructed image data does not exceed the first image sharpness percentage value, in operation 211b, the electronic device (101) (e.g., the processor (120) of FIG. 1) can determine whether the sharpness percentage value of the first reconstructed image data is greater than or equal to the second image sharpness percentage value. Specifically, if the electronic device (101) determines that the sharpness percentage value of the first reconstructed image data (e.g., sharpness 79%) does not exceed the first image sharpness percentage value (e.g., sharpness 85%), it determines that the second face recognition threshold is not exceeded and can determine whether the first reconstructed image data exceeds the second image sharpness percentage value (e.g., sharpness 70%). This allows the present embodiment to be configured to increase sharpness step by step before retraining the main deep learning model.

[0041] In operation 211c, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the first restored image data is greater than or equal to the sharpness percentage value of the second image, it can input the first restored image data into the main deep learning model while setting it to a fast value and output the corrected second restored image data. Specifically, the electronic device (101) determines that the sharpness percentage value of the first restored image data (e.g., sharpness 79%) exceeds the sharpness percentage value of the second image (e.g., sharpness 70%), and sets it to a fast value to minimize the correction time, thereby reducing the correction time of the main deep learning model while increasing the sharpness. Additionally, the electronic device (101) can input the first restored image data into the main deep learning model to extract the second restored image data and the sharpness percentage value of the second restored image data, return to operation 213, and again determine whether the second restored image data exceeds the second facial recognition threshold value.

[0042] Meanwhile, in operation 211b, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the first restored image data is less than or equal to the second image sharpness percentage value, in operation 211d, the electronic device (101) (e.g., the processor (120) of FIG. 1) can determine that the sharpness percentage value of the first restored image data is greater than or equal to the third image sharpness percentage value. Specifically, if the electronic device (101) determines that the sharpness percentage value of the first restored image data (e.g., sharpness 65%) does not exceed the second image sharpness percentage value (e.g., sharpness 70%), it determines that the second image sharpness percentage value is not exceeded and can determine whether the first restored image data exceeds the third image sharpness percentage value (e.g., sharpness 60%). This allows the present embodiment to be configured to increase sharpness step by step before retraining the main deep learning model.

[0043] In operation 211e, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the first restored image data is greater than or equal to the sharpness percentage value of the third image, it can input the first restored image data into the main deep learning model while setting it to a normal value and output the corrected second restored image data. Specifically, the electronic device (101) can determine that the sharpness percentage value of the first restored image data (e.g., sharpness 65%) exceeds the sharpness percentage value of the third image (e.g., sharpness 60%), and the correction time is minimized, but unlike operation 211c, it can be set to a value with a certain amount of time, thereby reducing the correction time of the main deep learning model while increasing sharpness. Additionally, the electronic device (101) can input the first restored image data into the main deep learning model to extract the second restored image data and the sharpness percentage value of the second restored image data, and return to operation 213 to determine again whether the second restored image data exceeds the second facial recognition threshold value.

[0044] Meanwhile, in operation 211d, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the first restored image data is less than or equal to the sharpness percentage value of the third image, in operation 211f, the electronic device (101) (e.g., the processor (120) of FIG. 1) can input the first restored image data into the main deep learning model while setting it to a slow value and output the corrected second restored image data. Specifically, the electronic device (101) can determine that the sharpness percentage value of the first restored image data (e.g., sharpness 49%) exceeds the sharpness percentage value of the third image (e.g., sharpness 60%), and can increase the sharpness of the second restored image data by minimizing the correction time but setting it to a value that allocates more time than in operation 211e, thereby providing sufficient correction time for the main deep learning model. Additionally, the electronic device (101) can input the first restored image data into the main deep learning model to extract the second restored image data and the sharpness percentage value of the second restored image data, and return to operation 213 to determine again whether the second restored image data exceeds the second facial recognition threshold value.

[0045] Additionally, the electronic device (101) can automatically set the learning time of the first restored image data based on the 211c operation, the 211e operation, and the 211f operation, but if it obtains command data that prioritizes algorithm selection and time selection from the user's external electronic device (102, 104, 106), it can set the face restoration algorithm based on the obtained data.

[0047] FIG. 4 is a flowchart for specifically explaining a method of operation in which an electronic device outputs from a first restored image data according to various embodiments of the present invention.

[0048] In operation 211a', the electronic device (101) (e.g., the processor (120) of FIG. 1) can extract a sharpness percentage value of the output second reconstructed image data when the first reconstructed image data is input into the main deep learning model. As an example, the electronic device (101) inputs the first reconstructed image data into the main deep learning model and outputs the second reconstructed image data. Afterward, the electronic device (101) can extract a sharpness percentage value of the output second reconstructed image data.

[0049] In operation 213, the electronic device (101) (e.g., the processor (120) of FIG. 1) can determine whether the sharpness percentage value of the second reconstructed image data exceeds the second facial recognition threshold. The electronic device (101) operates in the same manner as the operation 213 described above, and if the second reconstructed image data exceeds the second facial recognition threshold, it proceeds to operation 215 and operates in the same manner, so specific details can be omitted. Furthermore, the reason the operation is included in this embodiment is that since the sharpness value of the second reconstructed image data output from the main deep learning model has increased, there is an advantage in that a separate correction process can be omitted by first checking whether the second facial recognition threshold has been exceeded.

[0050] Meanwhile, in operation 213, if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the second reconstructed image data is less than or equal to the second image sharpness percentage value, in operation 211b', the electronic device (101) (e.g., the processor (120) of FIG. 1) can determine that the sharpness percentage value of the second reconstructed image data is greater than or equal to the third image sharpness percentage value. Specifically, if the electronic device (101) determines that the sharpness percentage value of the second reconstructed image data (e.g., sharpness 73%) does not exceed the first image sharpness percentage value (e.g., sharpness 85%), it determines that the second face recognition threshold has not been exceeded, and can determine whether the second reconstructed image data has exceeded the second image sharpness percentage value (e.g., sharpness 70%). This allows the present embodiment to be configured to increase sharpness step by step before retraining the main deep learning model.

[0051] In operation 211c', if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the second restored image data is greater than or equal to the sharpness percentage value of the second image, it can input the second restored image data into the main deep learning model while setting it to a fast value and output the corrected third restored image data. Specifically, the electronic device (101) determines that the sharpness percentage value of the second restored image data (e.g., sharpness 73%) exceeds the sharpness percentage value of the third image (e.g., sharpness 70%), and sets it to a fast value so as to minimize the correction time, thereby reducing the correction time of the main deep learning model while increasing the sharpness. Additionally, the electronic device (101) can input the first restored image data into the main deep learning model to extract the second restored image data and the sharpness percentage value of the second restored image data, and return to operation 213 to determine again whether the second restored image data exceeds the second facial recognition threshold value.

[0052] Meanwhile, in operation 211b', if the electronic device (101) (e.g., the processor (120) of FIG. 1) determines that the sharpness percentage value of the first restored image data is less than or equal to the sharpness percentage value of the second image, in operation 211d', the electronic device (101) (e.g., the processor (120) of FIG. 1) can input the second restored image data into the main deep learning model while setting it to a normal value and output the corrected third restored image data. Specifically, the electronic device (101) can determine that the sharpness percentage value of the second restored image data (e.g., sharpness 61%) exceeds the sharpness percentage value of the second image (e.g., sharpness 70%), and the correction time is minimized, but unlike operation 211c, the sharpness can be increased while reducing the correction time of the main deep learning model by setting it to a value with a certain amount of time. Additionally, the electronic device (101) can input the second restored image data into the main deep learning model to extract the third restored image data and the sharpness percentage value of the third restored image data, and return to operation 213 to determine again whether the third restored image data exceeds the second facial recognition threshold value.

[0053] In the case of the content of Fig. 4, the first reconstructed image data has already been trained once relative to the mask data set, so there is no need to separately set the slowness, which can prevent waste of time in advance. Based on this, the main deep learning model can also train multiple second reconstructed image data and multiple third reconstructed image data, thereby having the advantage of further reducing training and output time.

[0054] Additionally, the electronic device (101) can automatically set the learning time of the second restored image data based on the 211c operation, the 211e operation, and the 211f operation, but if it obtains command data that prioritizes algorithm selection and time selection from the user's external electronic device (102, 104, 106), it can set the face restoration algorithm based on the obtained data.

[0056] According to various embodiments, an electronic device for performing face reconstruction based on a deep learning model comprises: a communication interface; a memory; and a processor; The method includes, wherein the processor acquires face source data and face target data for restoration through the communication interface, inputs the face source data and the face target data into at least one first sub-deep learning model to output a source data set and a target data set, inputs the source data set and the target data set into a second sub-deep learning model to output a mask data set, and if it is determined that the mask data set is greater than or equal to a preset first face recognition threshold, inputs the mask data set into a main deep learning model to output final restored image data, and if it is determined that the mask data set is less than or equal to the first face recognition threshold, inputs the mask data set into the main deep learning model while setting a value to output first restored image data, and if it is determined that the first restored image data is greater than or equal to a preset second face recognition threshold, inputs the first restored image data into the main deep learning model to output the final restored image data, and the first restored image data is equal to the second face recognition threshold If it is determined that the value is less than or equal to the first face recognition threshold, the process is set to repeat from the case where the first restored image data is less than or equal to the first face recognition threshold until it is determined to be greater than or equal to the second face recognition threshold, and the main deep learning model is trained based on a plurality of source data sets for a plurality of faces, a plurality of target data sets for a plurality of faces, a plurality of mask data sets for a plurality of faces, final restored image data for a plurality of restored faces, a plurality of first restored image data for a plurality of restored faces, a plurality of face restoration algorithms, and an image clarity evaluation algorithm.

[0057] According to various embodiments, the at least one first sub-deep learning model includes a first-1 sub-deep learning model and a first-2 sub-deep learning model, and the processor is configured to input the face source data into the first-1 sub-deep learning model to output the source data set, and input the face target data into the first-2 sub-deep learning model to output the target data set.

[0058] According to various embodiments, the first-1 sub-deep learning model is trained based on a plurality of face source data and a plurality of said source data sets, and the first-2 sub-deep learning model is trained based on a plurality of face target data and a plurality of said target data sets.

[0059] According to various embodiments, the first face recognition threshold includes a preset first image clarity percentage value and the plurality of face restoration algorithms, and the second face recognition threshold includes the first image clarity percentage value and the plurality of image clarity evaluation algorithms, and the image clarity evaluation algorithms include a Laplacian algorithm, a Laplacian of Gaussian (LOG) algorithm, a difference of Gaussian (DOF) algorithm, a Sobel algorithm, a Robert algorithm, and a Prewitt algorithm.

[0060] According to various embodiments, when the mask data set that does not exceed the first face recognition threshold is input to the main deep learning model, the processor extracts the clarity percentage value of the output first reconstructed image data; if it is determined that the clarity percentage value of the first reconstructed image data is greater than or equal to a second image clarity percentage value set to be smaller than the first image clarity percentage value, the processor inputs the first reconstructed image data into the main deep learning model while setting it to a fast value and outputs the corrected second reconstructed image data; if it is determined that the clarity percentage value of the first reconstructed image data is greater than or equal to a third image clarity percentage value set to be smaller than the second image clarity percentage value, the processor inputs the first reconstructed image data into the main deep learning model while setting it to a normal value and outputs the corrected second reconstructed image data; and if it is determined that the clarity percentage value of the first reconstructed image data is less than or equal to the third image clarity percentage value, the processor inputs the first reconstructed image data into the main deep learning model while setting it to a slow value. It is set to output the corrected second restored image data.

[0061] According to various embodiments, the processor is configured to input the first restored image data, which does not exceed the first face recognition threshold and / or the second face recognition threshold, into the main deep learning model, extract a sharpness percentage value of the output second restored image data, and if it is determined that the sharpness percentage value of the second restored image data is greater than or equal to a second image sharpness percentage value set to be smaller than the first image sharpness percentage value, input the second restored image data into the main deep learning model while setting it to a fast value and output a corrected third restored image data, and if it is determined that the sharpness percentage value of the second restored image data is less than or equal to the second image sharpness percentage value, input the second restored image data into the main deep learning model while setting it to a normal value and output the corrected third restored image data.

[0063] As used in this document, the terms “module” or “part” include a unit composed of hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. “Module” or “part” may be a component formed integrally or a minimum unit or part thereof that performs one or more functions. “Module” or “part” may be implemented mechanically or electronically and may include, for example, an application-specific integrated circuit (ASIC) chip, field-programmable gate arrays (FPGAs), or programmable logic device known or to be developed that performs certain operations, and may be executed by a processor (120). At least part of the device (e.g., modules or functions thereof) or method (e.g., operations) according to various embodiments may be implemented as instructions stored in a computer-readable storage medium (e.g., memory (130)) in the form of a program module. When the above instruction is executed by a processor (e.g., processor (120)), the processor may perform a function corresponding to the above instruction. Computer-readable recording media may include a hard disk, a floppy disk, a magnetic medium (e.g., magnetic tape), an optical recording medium (e.g., CD-ROM, DVD, magneto-optical medium (e.g., floptical disk), built-in memory, etc. Instructions may include code generated by a compiler or code that can be executed by an interpreter. A module or program module according to various embodiments may include at least one of the aforementioned components, some of which may be omitted, or additionally include other components. Operations performed by a module, program module, or other components according to various embodiments may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.

[0064] Furthermore, the embodiments disclosed in this document are presented for the purpose of explaining and understanding the disclosed technical content and are not intended to limit the scope of this disclosure. Accordingly, the scope of this disclosure should be interpreted to include all modifications or various other embodiments based on the technical concept of this disclosure.

Claims

Claim 1 An electronic device for performing face reconstruction based on a deep learning model, comprising: a communication interface; a memory; and a processor;The method includes, wherein the processor acquires face source data and face target data for restoration through the communication interface, inputs the face source data and the face target data into at least one first sub-deep learning model respectively to output a source data set and a target data set, inputs the source data set and the target data set into a second sub-deep learning model to output a mask data set, if it is determined that the mask data set exceeds a preset first face recognition threshold, inputs the mask data set into a main deep learning model to output final restored image data, if it is determined that the mask data set is less than or equal to the first face recognition threshold, inputs the mask data set into the main deep learning model while setting a value to output first restored image data, if it is determined that the first restored image data exceeds a preset second face recognition threshold, inputs the first restored image data into the main deep learning model to output the final restored image data, and the first restored image data is less than or equal to the second face recognition threshold An electronic device configured to output the first restored image data by inputting it into the main deep learning model while setting the setting value of the mask data set from the case where the mask data set is below the first face recognition threshold until the first restored image data is determined to exceed the second face recognition threshold when it is determined to be below the value, wherein the main deep learning model is trained based on a plurality of source data sets for a plurality of faces, a plurality of target data sets for a plurality of faces, a plurality of mask data sets for a plurality of faces, final restored image data for a plurality of restored faces, a plurality of first restored image data for a plurality of restored faces, a plurality of face restoration algorithms, and an image clarity evaluation algorithm. Claim 2 An electronic device according to claim 1, wherein the at least one first sub-deep learning model comprises a first-1 sub-deep learning model and a first-2 sub-deep learning model, and the processor is configured to input the face source data into the first-1 sub-deep learning model to output the source data set and input the face target data into the first-2 sub-deep learning model to output the target data set. Claim 3 An electronic device according to claim 2, wherein the first-1 sub-deep learning model is learned based on a plurality of face source data and a plurality of said source data sets, and the first-2 sub-deep learning model is learned based on a plurality of face target data and a plurality of said target data sets. Claim 4 An electronic device according to claim 1, wherein the first face recognition threshold is a preset first image clarity percentage value, the second face recognition threshold is the first image clarity percentage value, and the image clarity evaluation algorithm comprises a Laplacian algorithm, a Laplacian of Gaussian (LOG) algorithm, a difference of Gaussian (DOF) algorithm, a Sobel algorithm, a Robert algorithm, and a Prewitt algorithm. Claim 5 In claim 4, the processor, when inputting the mask data set that does not exceed the first face recognition threshold into the main deep learning model, extracts the clarity percentage value of the output first reconstructed image data; if it is determined that the clarity percentage value of the first reconstructed image data is greater than or equal to a second image clarity percentage value set to be smaller than the first image clarity percentage value, inputs the first reconstructed image data into the main deep learning model while setting it to a fast value and outputs the corrected second reconstructed image data; if it is determined that the clarity percentage value of the first reconstructed image data is greater than or equal to a third image clarity percentage value set to be smaller than the second image clarity percentage value, inputs the first reconstructed image data into the main deep learning model while setting it to a normal value and outputs the corrected second reconstructed image data; and if it is determined that the clarity percentage value of the first reconstructed image data is less than or equal to the third image clarity percentage value, inputs the first reconstructed image data into the main deep learning model while setting it to a slow value An electronic device configured to output the corrected second restored image data. Claim 6 An electronic device according to claim 4, wherein the processor is configured to extract a sharpness percentage value of the output second reconstructed image data when the first reconstructed image data that does not exceed the first face recognition threshold and / or the second face recognition threshold is input to the main deep learning model, and if it is determined that the sharpness percentage value of the second reconstructed image data is greater than or equal to a second image sharpness percentage value set to be smaller than the first image sharpness percentage value, the processor inputs the second reconstructed image data into the main deep learning model while setting it to a fast value and outputs the corrected third reconstructed image data, and if it is determined that the sharpness percentage value of the second reconstructed image data is less than or equal to the second image sharpness percentage value, the processor inputs the second reconstructed image data into the main deep learning model while setting it to a normal value and outputs the corrected third reconstructed image data.

Citation Information

Patent Citations

  • Portrait restoration method and device, electronic equipment and computer storage medium

    CN112330574A

  • Methods, devices, electronic devices, storage media and program products for restoring human image

    KR1020230054432A

  • Method for restoring a masked face image by using the neural network model

    KR102547630B1