Electronic device for processing image, and operating method thereof
The electronic device enhances photo quality by processing multiple images to replace unsatisfactory facial regions with better-captured alternatives, addressing user dissatisfaction with mobile device photography.
Patent Information
- Application Number
- PCT/KR2024/096854
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-18
- Filing Date
- 2024-12-12
- Publication Date
- 2025-09-25
AI Technical Summary
Users are dissatisfied with the quality of photos taken on mobile devices, particularly when individuals in the photos have closed eyes, obscured faces, or unfavorable poses, necessitating retakes.
An electronic device processes multiple images to determine a base image and a source image based on completion level and swap compatibility, extracting and synthesizing facial regions to generate a corrected image.
The solution allows for improved photo quality by replacing unsatisfactory facial regions with better-captured alternatives, enhancing user satisfaction.
Smart Images

Figure KR2024096854_25092025_PF_FP_ABST
Abstract
Description
Electronic device for processing images and method of operation thereof
[0001] The present invention relates to an electronic device for processing images and an operating method thereof, and more particularly, to an electronic device for combining valid image fragments from among a plurality of images acquired through continuous shooting and an operating method thereof.
[0002] In recent years, rapid advancements in communication technology have led to a gradual expansion of the functionality of mobile devices. This, in turn, has led to the development of more diverse user interfaces (UIs) and their associated features. To enhance the utility of these devices and satisfy the diverse needs of users, a variety of applications capable of running on these devices are being developed.
[0003] In particular, with the growing user interest in photography and video, most mobile devices now offer digital camera functionality. At the same time, users demand high-quality photos taken with digital cameras and want to capture themselves in a pleasing way. Methods to satisfy these preferences, such as correcting hand shake or retouching photos, are being introduced. However, these are merely supplementary methods for retouching photos, and if the user is not satisfied with the expression or pose, the photo will ultimately need to be retaken.
[0004] A method disclosed as a technical means for achieving a technical task may include the steps of acquiring a plurality of images including a plurality of persons, determining one of the plurality of images as a base image, determining a source image from the plurality of images based on a completion level of shooting of each of the plurality of images and a swap compatibility of each of the plurality of images, extracting a facial area of one of the plurality of persons from the source image, and generating a corrected image by synthesizing the extracted facial area onto the base image. For each image from the plurality of images, the completion level of shooting may include a completion level of shooting a facial area of the person in each image, and the swap compatibility may include a swap compatibility between a facial area of the person in each image and a facial area of the person in the base image.
[0005] An electronic device disclosed as a technical means for achieving a technical task includes an input / output interface for receiving a user input requesting processing of an image, outputting a processed image according to the user input, a memory storing commands for processing the image, and at least one processor, wherein the at least one processor executes a program or at least one instruction stored in the memory, thereby obtaining a plurality of images including a plurality of persons, determining one of the plurality of images as a base image, determining a source image among the plurality of images based on a completion level of shooting of each of the plurality of images and a swap compatibility of each of the plurality of images, extracting a face area of one of the plurality of persons from the source image, and synthesizing the extracted face area onto the base image, thereby generating a corrected image. For each image among the plurality of images, the completion level of shooting may include a completion level of shooting the face area of the person in each image, and the swap compatibility may include a swap compatibility between the face area of the person in each image and the face area of the person in the base image.
[0006] A computer-readable recording medium disclosed as a technical means for achieving a technical task may have stored thereon a program for executing at least one of the embodiments of the disclosed method on a computer.
[0007] A computer program disclosed as a technical means for achieving a technical task may be stored on a medium for performing at least one of the embodiments of the disclosed method on a computer.
[0008] The above and other aspects, features and advantages of specific embodiments of the present disclosure will become apparent from the following description taken with reference to the accompanying drawings.
[0009] FIG. 1 is a conceptual diagram illustrating a method for generating a correction image according to an embodiment of the present disclosure.
[0010] FIG. 2 is a flowchart illustrating a method for generating a correction image according to an embodiment of the present disclosure.
[0011] FIG. 3 is a conceptual diagram illustrating a method for evaluating multiple images to determine a base image or a source image among multiple images according to one embodiment of the present disclosure.
[0012] FIG. 4 is a flowchart illustrating a method for determining a base image among a plurality of images according to an embodiment of the present disclosure.
[0013] FIG. 5 is a flowchart illustrating a method for determining a target image among a plurality of images according to one embodiment of the present disclosure.
[0014] FIG. 6 is a flowchart illustrating a method of considering a user's preference to determine a target image according to one embodiment of the present disclosure.
[0015] FIG. 7 is a flowchart illustrating a method for considering a user's pose to determine a target image according to one embodiment of the present disclosure.
[0016] FIG. 8 is a conceptual diagram illustrating a method of filtering a plurality of images to remove a blurry image and then generating a corrected image based on the filtered plurality of images according to one embodiment of the present disclosure.
[0017] FIG. 9 is a conceptual diagram illustrating an operation of generating a 3D face model based on an image including a first person according to one embodiment of the present disclosure.
[0018] FIG. 10 is a flowchart illustrating a method for correcting a facial region of a base image based on a 3D facial model according to an embodiment of the present disclosure.
[0019] FIG. 11 is a conceptual diagram illustrating a method for rotating a pose of a face region based on a 3D face model according to one embodiment of the present disclosure.
[0020] FIG. 12 is a conceptual diagram illustrating a method for changing illumination of a facial region based on a 3D facial model according to one embodiment of the present disclosure.
[0021] FIG. 13 is a conceptual diagram illustrating a method for correcting an occluded area within a facial region based on a 3D facial model according to an embodiment of the present disclosure.
[0022] FIG. 14 is a conceptual diagram illustrating a method for reflecting an object on a final corrected image when there is an object in a face area of a base image according to one embodiment of the present disclosure.
[0023] FIG. 15 is a block diagram illustrating a configuration of an electronic device according to an embodiment of the present disclosure.
[0024] In describing this disclosure, descriptions of technical details that are well-known in the technical field to which this disclosure pertains and are not directly related to this disclosure will be omitted. Furthermore, the terms described below are defined based on their functions in this disclosure and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the overall content of this specification.
[0025] For the same reason, some components in the attached drawings are exaggerated, omitted, or schematically depicted. Furthermore, the dimensions of each component do not entirely reflect its actual size. Identical or corresponding components in each drawing are assigned the same reference numbers.
[0026] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail together with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. The disclosed embodiments are provided to ensure that the disclosure is complete and to fully inform those skilled in the art of the present disclosure of the scope of the disclosure. An embodiment of the present disclosure may be defined according to the claims. Like reference numerals denote like elements throughout the specification. In addition, when describing an embodiment of the present disclosure, if a detailed description of a related function or configuration is determined to unnecessarily obscure the gist of the present disclosure, the detailed description thereof will be omitted.
[0027] Terms such as “unit,” “module,” “member,” and “block” may be implemented in hardware or software. As used herein, multiple “units,” “modules,” “members,” and “blocks” may be implemented as a single component, and a single “unit,” “module,” “member,” and “block” may include multiple components.
[0028] When an element is referred to as being “connected” to another element, it is understood that it can be connected to the other element either directly or indirectly, where an indirect connection can include “connecting via a wireless communications network.”
[0029] Additionally, if a part includes (comprises) a certain element, unless there is a special statement to the contrary, that part does not exclude other elements and may further include other elements.
[0030] Throughout this specification, when a configuration is said to be “on” another configuration, this includes not only when the configuration is in contact with the other configuration, but also when there is another configuration between the two configurations.
[0031] As used herein, the expressions “at least one of A, B or C” and “at least one of A, B and C” refer to “A”, “B”, “C”, “A and B”, “A and C”, “B and C” and “A, B and C”.
[0032] Although terms such as "first," "second," and "third" may be used herein to describe various elements, it will be understood that the present disclosure is not limited by these terms. These terms are used solely to distinguish one element from another.
[0033] The singular expressions used in this specification may include plural expressions unless the context clearly indicates otherwise.
[0034] In connection with any method or process described herein, drawing symbols may be used for convenience of explanation, but are not intended to illustrate the order of each step or operation. Unless otherwise specified in the context, each step or operation may be implemented in a different order than the illustrated order. Unless otherwise specified in the context of the present disclosure, one or more steps or operations may be omitted.
[0035] In one embodiment, each block of each flowchart diagram and combinations of flowchart diagrams can be performed by computer program instructions. The computer program instructions can be installed on a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, and the instructions, when executed by the processor of the computer or other programmable data processing apparatus, can create means for performing the functions described in the flowchart block(s). The computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing apparatus to implement the functions in a particular manner, and the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). The computer program instructions can also be installed on a computer or other programmable data processing apparatus.
[0036] Additionally, each block in the flowchart diagram may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). In one embodiment, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may be executed substantially simultaneously or, depending on the function, may be executed in reverse order.
[0037] The term '~ unit' used in one embodiment of the present disclosure may represent software or a hardware component such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC), and the '~ unit' may perform a specific role. Meanwhile, the '~ unit' is not limited to software or hardware. The '~ unit' may be configured to be on an addressable storage medium and may be configured to play one or more processors. In one embodiment, the '~ unit' may include components such as software components, object-oriented software components, class components, and task components, processes, functions, properties, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided through a specific component or a specific '~ unit' may be combined to reduce the number of components or separated into additional components. In addition, in one embodiment, the '~ unit' may include one or more processors.
[0038] Below, the meanings of terms used in this disclosure are explained.
[0039] A corrected image can be generated from a plurality of images according to one or more of the embodiments described below.
[0040] In the present disclosure, a base image may be an image selected as a base from among a plurality of images. The base image may be selected as an image requiring minimal correction from among the plurality of images. For example, a method according to an embodiment of the present disclosure may replace a portion of a base image with an area of another image based on the base image. The term "base image" is used only to mean an image used as a basis for image processing, and may be replaced with various terms such as best image, best take, best cut, best shot, reference image, and base image.
[0041] In the present disclosure, a source image may be an image including a region among a plurality of images in which a portion of a base image is to be replaced. The source image includes a region corresponding to a portion of the base image, and the corresponding region of the source image is more advantageous in satisfying the user than the corresponding portion of the base image in terms of various factors such as aesthetics, the user's posture preference, and clarity. The term "source image" is used only in the sense of a source for extracting a facial region, and may be replaced with various terms such as "source image," "swap image," "substitute image," and "substitute image."
[0042] In this disclosure, the completion level of shooting refers to the degree to which an image is well captured. From the perspective of the subject of the shooting, the completion level of shooting may refer to the degree to which a person included in the image is well captured. For example, the completion level of shooting may be determined for one person among multiple people included in the image. For one image, multiple completion levels of shooting may be determined, such as the completion level of shooting for a first person, the completion level of shooting for a second person, and the completion level of shooting for a third person. Furthermore, the completion level of shooting may be determined by considering factors related to how the person included in the image was captured. For example, the completion level of shooting may be determined not only by considering the person's facial expression, but also by considering aesthetic aspects such as whether the person included in the image has their eyes closed, whether the person's gaze is directed toward the camera, whether the person's face is shadowed, or whether the person's face is obscured by an object.
[0043] In this disclosure, the aesthetic score refers to a numerical score that quantifies the degree of completion of an image's photography. Similar to the degree of completion, the aesthetic score may be determined for one of multiple people included in an image, and multiple aesthetic scores may be determined for a single image, such as an aesthetic score for a first person, an aesthetic score for a second person, and an aesthetic score for a third person.
[0044] In the present disclosure, the aesthetic score is described as being divided into a first aesthetic score and a second aesthetic score. The first aesthetic score refers to an aesthetic score used in a method of determining a base image. The first aesthetic score refers to a score that quantifies the degree of completion in capturing the facial region of a plurality of people for each of a plurality of images. In other words, the first aesthetic score refers to a score determined for an arbitrary person for each of the plurality of images. The first aesthetic score may be a score determined for each of all people included in the plurality of images. The second aesthetic score refers to an aesthetic score used in a method of determining a source image for replacing the facial region of the first person. The second aesthetic score refers to a score that quantifies the degree of completion in capturing the facial region of the first person for each of the plurality of images. In other words, the second aesthetic score refers to a score determined for a first person selected for each of the plurality of images. The second aesthetic score refers to a score determined for the first person included in the plurality of images.
[0045] However, the first aesthetic score and the second aesthetic score may be suitability scores that can be calculated in the same manner, and are merely distinct terms for convenience of explanation.
[0046] In the present disclosure, swap compatibility refers to the swap compatibility between facial regions of the same person included in two images. Here, swap refers to a correction in which, in a first image and a second image containing a first person, the facial region of the first person in the first image is replaced with the facial region of the first person in the second image. Swap compatibility may refer to the degree to which the swap can be performed appropriately based on the compatibility between the swapped facial region and its surrounding area.
[0047] From the perspective of the subject of the photograph, substitution suitability may refer to the degree to which a substitution between facial regions can be appropriately performed for a person included in an image. For example, substitution suitability may be determined for one of multiple persons included in an image. Regarding whether a facial region included in an image can be substituted for a facial region included in a base image, multiple substitution suitability values, such as substitution suitability for a first person, substitution suitability for a second person, and substitution suitability for a third person, may be determined.
[0048] In this disclosure, a compatibility score refers to a score that quantifies the substitution suitability of an image. Similar to substitution suitability, a compatibility score may be determined for one of multiple individuals included in an image, and multiple compatibility scores may be determined for a single image, such as a compatibility score for a first individual, a compatibility score for a second individual, and a compatibility score for a third individual.
[0049] In the present disclosure, the suitability score can be explained by dividing it into a first suitability score and a second suitability score. The first suitability score refers to a suitability score used in a method of determining a base image. The first suitability score refers to a score that quantifies the substitution suitability between face region pairs of target persons included in an image pair consisting of two images among a plurality of images. In other words, the first suitability score refers to a score determined for each individual person in the image pair. The first suitability score may be a score determined for each individual person included in the plurality of images. The second suitability score refers to a suitability score used in a method of determining a source image for substituting the face region of the first individual. The second suitability score refers to a score that quantifies the substitution suitability between face region pairs of the first individual included in the image pair. In other words, the second suitability score refers to a score determined for the first individual in the image pair. The second suitability score refers to a score determined for the first individual included in the plurality of images.
[0050] However, the first suitability score and the second suitability score may be suitability scores that can be calculated in the same manner, and are merely distinct terms for convenience of explanation.
[0051] In the present disclosure, a keypoint (feature point) refers to a point within an image that is distinguishable or identifiable from the surrounding background, and corresponds to a key point of the body. For example, a keypoint for a hand may include points corresponding to multiple joints within the hand. A keypoint may be expressed as a three-dimensional position coordinate value, which is position information about the x-axis, y-axis, and z-axis of a key point of the body.
[0052] In this disclosure, feature points are described as base feature points and source feature points. Base feature points refer to feature points extracted from a base image. Base feature points refer to feature points for a person within the base image. Source feature points refer to feature points extracted from the source image. Source feature points refer to feature points for a person within the source image.
[0053] In the present disclosure, blurriness may refer to the degree of blurring, whereby details in an image are smoothly blurred or indistinct. Blurry may occur due to reasons such as lack of focus (defocus), rapid movement of the subject during shooting (motion blur), or low resolution or excessive compression of the image (low-resolution blur).
[0054] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings so that those skilled in the art can practice the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In addition, in order to clearly describe the present disclosure in the drawings, parts that are not related to the description are omitted, and similar parts are designated with similar reference numerals throughout the specification. In addition, the reference numerals used in each drawing are only for describing each drawing, and different reference numerals used in different drawings do not indicate different elements. The present disclosure will be described in detail below with reference to the attached drawings.
[0055] FIG. 1 is a conceptual diagram illustrating a method for generating a correction image according to an embodiment of the present disclosure.
[0056] Referring to FIG. 1, the electronic device can acquire a plurality of images (110; 111, 112, 113, 114).
[0057] The plurality of images (110) may be images captured continuously over a set period of time. The plurality of images (110) may include an image sequence, which refers to a set of images arranged in order over time. For example, the plurality of images (110) may include at least one of a set of images captured over a set period of time, a motional image composed of images captured over a set period of time, and a video.
[0058] In one embodiment, the plurality of images (110) may each represent a group photo including a plurality of individuals. The plurality of images (110) may each be images including facial areas of the plurality of individuals. The plurality of images (110) may be images including facial areas of four individuals, and the number of individuals does not limit the technical concept of the present disclosure.
[0059] In one embodiment, the plurality of images (110) may include a group photo including not only a person but also a plurality of living things. The plurality of images (110) may each include images including facial regions of the plurality of living things. For example, at least one of the plurality of images (110) may be an image including the face of a dog, and an image processing method according to one embodiment may perform an operation of replacing the facial region of the dog in the base image with the facial region of the dog in the source image. However, the present invention will be described focusing on a method of replacing the facial region of a person.
[0060] In one embodiment, the plurality of images (110) may be images capturing the appearance of a plurality of people as they change over time. Over time, some of the plurality of images (110) may be images capturing at least one of the plurality of people. For example, one image of the plurality of images (110) may include all of the plurality of people, while another image of the plurality of images (110) may include two of the plurality of people.
[0061] Each person included in the plurality of images (110) may be captured differently for each image, such as by taking different poses or making different facial expressions over time. In particular, with respect to the face area, each person included in the plurality of images (110) may be captured with their eyes closed, obscured by external objects, or looking in an unexpected direction. This may cause the user to be dissatisfied with the captured images.
[0062] In one embodiment, the electronic device may determine any one of the plurality of images (110) as the base image. For example, the electronic device may determine the first image (111) as the base image.
[0063] For example, unlike as shown in FIG. 1, the electronic device may determine the fourth image (114) acquired last among the plurality of images (110) as the base image. As another example, the electronic device may determine the first image (111) with the highest degree of completion among the plurality of images (110) as the base image. The electronic device may consider the completion level of shooting and swap compatibility, which will be described in the following embodiments, to determine the degree of completion, and a description thereof will be provided below using FIG. 3.
[0064] In one embodiment, the electronic device may determine any one of the plurality of images (110) as a source image to replace the face area (A1_2) of the first person in the base image.
[0065] The base image may be a group photo containing multiple individuals. The base image may be an image containing the facial areas of multiple individuals. The multiple individuals included in the base image may include a first individual. The first individual may be any individual selected for convenience of explanation.
[0066] In one embodiment, the electronic device may determine a second image (112) from among a plurality of images (110) including a first person as a source image for the first person. The source image for the first person may refer to a source image selected to replace the face area (A1_2) of the first person in the base image.
[0067] In one embodiment, the electronic device may extract a facial region (A1_1) of a first person included in a source image of the first person. For example, the electronic device may extract the facial region (A1_1) of the first person from a second image (112).
[0068] In one embodiment, the electronic device can replace the face area (A1_2) of the first person in the base image with the face area (A1_1) of the first person in the source image. The face area (A1_2) of the first person in the base image can be replaced with the face area (A1_1) of the first person in the source image. The face area (A1_1) of the first person in the source image can also be composited onto the face area (A1_2) of the first person in the base image. The face area (A1_1) of the first person in the source image can also be composited by overlapping it with the face area (A1_2) of the first person in the base image. The method of replacing the face area (A1_2) of the first person in the base image with the face area (A1_1) of the first person in the source image does not limit the technical idea of the present disclosure.
[0069] For reference, only the method of replacing the face area of a person in the base image with the face area of a person in the source image based on the first person has been described, and the method of replacing the face area of a person in the base image with the face area of a person in the source image for other people among multiple images except the first person is also the same, so any redundant explanation will be brief or omitted.
[0070] For example, the electronic device may determine a third image (113) among a plurality of images (110) including a second person as a source image for the second person. The electronic device may extract a facial area (A2_1) of the second person included in the third image (113), which is the source image for the second person. The electronic device may replace the facial area (A2_2) of the second person in the base image with the facial area (A2_1) of the second person in the source image.
[0071] In one embodiment, a first image (111) among a plurality of images (110) including a fourth person may be determined as a source image for the fourth person. As illustrated in FIG. 1, since the first image (111) is determined as a base image and a source image for the fourth person, the electronic device may maintain a face area (A4_2) of the fourth person within the base image. The face area (A4_2) of the fourth person within the base image may be identical to the face area (A4_1) of the fourth person within the first image (111), which is a source image for the fourth person. That is, when a base image among the plurality of images (110) is determined as a source image for the fourth person, an operation of replacing the face area of the fourth person is unnecessary and may not be performed.
[0072] In one embodiment, the electronic device may determine a source image for each person, such as a source image for a first person, a source image for a second person, and a source image for a third person, among a plurality of images (110). The source image for the first person and the source image for the second person may be different from each other. For example, as illustrated in FIG. 1, the source image for the first person may be determined as the second image (112), and the source image for the second person may be determined as the third image (113).
[0073] In one embodiment, the electronic device can generate a corrected image (120) by replacing face areas (A1_2, A2_2, A3_2) of a plurality of persons in a base image with face areas (A1_1, A2_1, A3_1) in a source image for each person.
[0074] For example, the generated corrected image (120) may be an image in which the face area of each included person is substituted based on the first image (111), which is the base image. The corrected image (120) may be an image in which the face area (A1_2) of the first person is substituted with the face area (A1_1) of the first person extracted from the second image (112), which is the source image for the first person. The corrected image (120) may be an image in which the face area (A2_2) of the second person is substituted with the face area (A2_1) of the second person extracted from the third image (113), which is the source image for the second person. The corrected image (120) may be an image in which the face area (A3_2) of the third person is substituted with the face area (A3_1) of the third person extracted from the second image (112), which is the source image for the third person. The face area (A4_2) of the fourth person in the correction image (120) may be an image in which the face area (A4_1) of the fourth person in the first image (111) is maintained because the first image (111), which is the source image for the fourth person, is also a base image.
[0075] FIG. 2 is a flowchart illustrating a method for generating a correction image according to an embodiment of the present disclosure.
[0076] Referring to FIG. 2, in step S210, the electronic device can acquire multiple images including multiple people.
[0077] In one embodiment, an electronic device may include a camera module and may capture multiple images including multiple people using the camera module. The electronic device may capture a sequence of images by capturing the multiple images over a set period of time. The electronic device may capture, for example, a video or a motion image.
[0078] In one embodiment, the electronic device can acquire multiple images through communication with a separate server. The multiple images may be multiple images captured continuously over a set period of time.
[0079] In step S220, the electronic device can determine one of the plurality of images as a base image.
[0080] In one embodiment, the electronic device may determine the last image captured among a plurality of sequentially captured images as the base image. Alternatively, the electronic device may determine the first image captured among a plurality of sequentially captured images as the base image. The electronic device may determine one image among the plurality of images as the base image.
[0081] In step S230, the electronic device can determine a source image among a plurality of images based on the shooting completeness and substitution suitability.
[0082] In one embodiment, the completion level of shooting may refer to the completion level of shooting of the facial regions of a first person included in multiple images. The completion level of shooting may refer to the degree to which the face of the first person has been well captured. The completion level of shooting may be evaluated for one person among multiple people included in a single image, and the electronic device may obtain multiple completion levels of shooting for the multiple images.
[0083] The quality of a shot can be determined by considering factors related to how the subject is captured. For example, quality of a shot can be determined not only by considering the subject's facial expression, but also by considering aesthetic considerations such as whether the subject's eyes are closed, whether the subject's gaze is directed toward the camera, whether the subject's face is shadowed, or whether the subject's face is obscured by an object.
[0084] In one embodiment, swap compatibility may refer to an evaluation of the degree to which the face of a first person included in a plurality of images can be naturally substituted for the face region of the first person included in the base image. Swap compatibility may be evaluated for one person among the plurality of people included in one image, and the electronic device may obtain multiple swap compatibility values for the plurality of images and the plurality of people.
[0085] To determine the source image, substitution suitability can be determined by comparing the face of a person in one image with the face of the same person in the base image. For example, substitution suitability can be determined by considering the orientation, angle, and perspective of the person's face.
[0086] In one embodiment, the substitution suitability for determining the source image may be determined by comparing a person contained in one image with the same person contained in the base image. For example, the substitution suitability may be determined by considering the orientation, angle, and perspective of the person's face, as well as the pose, orientation, angle, and perspective of the body associated with the person's face.
[0087] In one embodiment, the electronic device may determine a source image from among a plurality of images based on the degree of completion of the shooting and the suitability for replacement. The electronic device may determine a source image for the first person from among the plurality of images based on the degree of completion of the shooting and the suitability for replacement of the plurality of images for the first person. The source image for the first person may refer to a source image for replacing the facial area of the first person in the base image.
[0088] In step S240, the electronic device can extract a facial region of a first person from the source image. In step S250, the electronic device can generate a corrected image by synthesizing the extracted facial region onto a base image.
[0089] In one embodiment, an electronic device can determine a source image for a first person. The electronic device can extract a facial region of the first person from the source image for the first person. The electronic device can synthesize the extracted facial region of the first person onto a facial region of the first person in a base image. The facial region of the first person in the base image can be replaced by the facial region of the first person extracted from the source image for the first person. By replacing the facial region of the first person in the base image, the electronic device can obtain a corrected image.
[0090] FIG. 3 is a conceptual diagram illustrating a method for evaluating multiple images to determine a base image or a source image among multiple images according to one embodiment of the present disclosure.
[0091] For convenience of explanation, parts that overlap with those described using Figure 1 are simplified or omitted.
[0092] Referring to FIG. 3, an electronic device can acquire a plurality of images. An arbitrary j-th image frame among the plurality of images can be selected. The arbitrary j-th image frame is selected merely for convenience of explanation, and the electronic device can associate each of the plurality of images with the arbitrary j-th image frame and acquire an aesthetic score and a suitability score for each based on the arbitrary j-th image frame.
[0093] In step S310, the electronic device may determine one of the multiple images as a base image. The electronic device may determine the base image based on an aesthetic score and a suitability score.
[0094] 1. Regarding aesthetic score calculation
[0095] In one embodiment, the electronic device can obtain an aesthetic score, which is a numerical score indicating the degree of completion in capturing the face region of an arbitrary i-th person within an arbitrary j-th image frame. The aesthetic score, which is a numerical score indicating the degree of completion in capturing the face region of an arbitrary i-th person within an arbitrary j-th image frame, can be expressed, for example, by mathematical expression 1.
[0096]
[0097] In mathematical expression 1, j can be a natural number that distinguishes multiple images acquired by an electronic device. For example, the aesthetic score for the i-th person in the 1st image frame is , and the aesthetic score for the i-th person in the second image frame is can be expressed as
[0098] In mathematical expression 1, i can be a natural number that distinguishes multiple people included in multiple images acquired by an electronic device. For example, the aesthetic score for the first person in the jth image frame is , and the aesthetic score for the second person in the j-th image frame is can be expressed as
[0099] In one embodiment, the aesthetic score may refer to a numerical score for the degree of completion of the photographic completion of the facial region of each of the plurality of individuals for each of the plurality of images. The aesthetic score may refer to a numerical score evaluating whether the facial region has been captured well.
[0100] The aesthetic score can be determined by evaluating factors such as whether the subject's eyes are closed, whether the motion is blurred due to large movements, whether the face is covered by hands or hair, whether the expression is appropriate, and whether the face's pose is facing forward.
[0101] For example, if the subject in the photo has their eyes closed, the aesthetic score may be evaluated as low. If the subject's movement is large and the photo is blurred, the aesthetic score may be evaluated as low. If the face is covered by hands or hair, the aesthetic score may be evaluated as low. If the subject is smiling, the aesthetic score may be evaluated as high. If the user prefers a neutral expression based on personal preference, the aesthetic score may be set to be high when the subject has a neutral expression. If the subject's face is facing forward, the aesthetic score may be evaluated as high. If the user prefers a side view of the face based on personal preference, the aesthetic score may be set to be high when the subject faces to the side.
[0102] In one embodiment, the aesthetic score may refer to a numerical score for the degree of completion of the photographic completion of the facial areas of multiple individuals in each of the multiple images, based on the relationships between the multiple individuals. The multiple images may be group photos of multiple individuals, and the score may refer to a numerical score evaluating whether the faces of the multiple individuals are well-photographed and match well with each other.
[0103] For example, aesthetic scores may be determined based on the gaze directions of multiple characters in multiple images, with the emphasis placed on the fact that multiple characters have identical or similar gaze directions. An aesthetic score may be higher when multiple characters have identical gazes directed toward the camera.
[0104] As another example, aesthetic scores may be determined based on the facial expressions or movements of multiple characters in multiple images, with the expressions or movements of multiple characters being highly valued for being identical or similar. When multiple characters share the same smiling or crying expressions, or when they share similar movements, the aesthetic score may be evaluated highly.
[0105] In one embodiment, the electronic device may obtain an aesthetic score for the ith person in the jth image frame, taking into account the individual's preference.
[0106] In one embodiment, an electronic device may acquire an image of an i-th person from a plurality of images and store the acquired images in memory. The memory may include a cluster that stores data with similar attributes. The electronic device may extract an image of the i-th person from the plurality of images and store the group of extracted images in the cluster.
[0107] In one embodiment, the electronic device may acquire a personal preference for an i-th person based on a group of images of the i-th person stored in a cluster. The personal preference may be determined based on a frequency or rate of appearance of the i-th person based on the group of images of the i-th person.
[0108] For example, from a group of images of the ith person, the higher the proportion of the ith person resting his chin on his hand, the more likely it is that the ith person prefers the image of resting his chin on his hand. As another example, from a group of images of the ith person, the higher the proportion of the ith person looking at the sky, the more likely it is that the ith person prefers the image of looking at the sky. As another example, from a group of images of the ith person, the lower the proportion of the ith person making a frowning expression, the more likely it is that the ith person does not prefer the image of making a frowning expression.
[0109] In one embodiment, the electronic device may additionally acquire an external image including an ith person, and the external image including the ith person may be stored in a cluster. The external image may include images other than a plurality of images continuously captured over a set period of time. For example, the electronic device may acquire the external image from a server. The electronic device may evaluate a personal preference from the external image including the ith person, and based on the evaluated personal preference, may obtain an aesthetic score for the ith person in the jth image frame.
[0110] 2. Regarding the calculation of suitability scores
[0111] In one embodiment, the electronic device can obtain a suitability score, which is a score quantifying the substitution suitability between the facial region of the ith person included in any jth image frame and the facial region of the ith person included in the Mth image frame. The suitability score, which is a score quantifying the substitution suitability between the facial region of the ith person included in any jth image frame and the facial region of the ith person included in the Mth image frame, can be expressed, for example, by mathematical expression 2.
[0112]
[0113] In mathematical expression 2, j and M can be natural numbers that distinguish multiple images acquired by an electronic device. For example, the compatibility score between the facial region of the i-th person included in the 1st image frame and the facial region of the i-th person included in the M-th image frame is , and the compatibility score between the facial region of the i-th person included in the second image frame and the facial region of the i-th person included in the M-th image frame is can be expressed as
[0114] Since two image frames are compared to calculate the suitability score, the two image frames are simply divided into the j-th image frame and the M-th image frame, and M can be a variable with the same meaning as j.
[0115] However, M may mean the number of multiple images acquired by the electronic device, and j may mean one of the natural numbers between 1 and M. That is, the j-th image frame may mean any image frame selected between the 1st image frame and the M-th image frame.
[0116] In mathematical expression 2, i can be a natural number that distinguishes multiple people included in multiple images acquired by an electronic device. For example, the compatibility score between the face region of the first person included in the j-th image frame and the face region of the first person included in the M-th image frame is , and the compatibility score between the facial region of the second person included in the j-th image frame and the facial region of the second person included in the M-th image frame is can be expressed as
[0117] In one embodiment, the suitability score may be a score that quantifies the substitution suitability between face region pairs of multiple people included in an image pair consisting of two images among a plurality of images. The suitability score may mean a score that quantifies an evaluation of whether the face region of a first person included in a reference image can be naturally substituted with the face region of a first person included in another image. The suitability score may be determined by considering the relative substitution suitability between the face regions included in two images of the image pair.
[0118] In one embodiment, the suitability score may be determined by considering the pose of the target person. The pose of the target person may include at least one of the pose of the face of the target person and the pose of the body of the target person. The pose of the face of the target person may be determined by considering the direction, angle, perspective, etc. of the face. The pose of the body of the target person may be determined based on the direction, angle, movement of the arms and legs, perspective, etc. The electronic device may acquire the pose of the target person through an accelerometer or a gyroscope, but the type of sensor does not limit the technical idea of the present disclosure.
[0119] For example, the more similar the face orientation of the first person included in the reference image is to the face orientation of the first person included in another image, the higher the suitability score can be evaluated. The more similar the face size of the first person included in the reference image is to the face size of the first person included in another image in terms of perspective, the higher the suitability score can be evaluated. The more similar the body orientation of the first person included in the reference image is to the body orientation of the first person included in another image, the higher the suitability score can be evaluated. The more similar the neck arrangement of the first person included in the reference image is to the neck arrangement of the first person included in the image, the higher the suitability score can be evaluated. Of course, the suitability score can be calculated by comprehensively evaluating the facial posture and body posture of the target person.
[0120] In one embodiment, the electronic device can obtain a compatibility score between a facial region of an i-th person included in a j-th image frame and a facial region of an i-th person included in an M-th image frame, based on feature points regarding key locations of a body of the target person.
[0121] In one embodiment, the electronic device can extract a first feature point related to a key location of the body of the ith person from the jth image frame. The electronic device can extract a second feature point related to a key location of the body of the ith person from the Mth image frame. By comparing the first feature point and the second feature point, the electronic device can obtain a suitability score regarding whether the face of the ith person in the jth image frame can be suitably substituted with the face of the ith person in the Mth image frame. Consequently, the electronic device can obtain suitability scores regarding whether the faces of any persons in each image can be substituted with each other, for an image pair composed of two images among a plurality of images.
[0122] In one embodiment, the feature points may include coordinate value data corresponding to major locations of the body of the target person. The electronic device may determine the direction, angle, perspective, etc. of the face of the target person by considering the feature points. Based on the determined direction, angle, and perspective of the face, the electronic device may calculate a suitability score regarding whether the face of the ith person in the jth image frame can be suitably replaced with the face of the ith person in the Mth image frame. Based on the determined direction, angle, and perspective of the face, the electronic device may calculate the suitability score by comparing the first feature point corresponding to the face of the ith person in the jth image frame with the second feature point corresponding to the face of the ith person in the Mth image frame.
[0123] In one embodiment, the feature points may correspond to key locations within the face of the target person's body. For example, key locations of the feature points may include the eyes, nose, mouth, chin, and cheekbone protrusions.
[0124] For example, the electronic device can compare a first feature point corresponding to the nose of the face of the ith person in the jth image frame with a second feature point corresponding to the nose of the face of the ith person in the Mth image frame, thereby comparing the position of the face of the ith person. The electronic device can further compare a first feature point corresponding to the mouth of the face of the ith person in the jth image frame with a second feature point corresponding to the mouth of the face of the ith person in the Mth image frame, thereby comparing the position and perspective of the face of the ith person. The electronic device can further compare a first feature point corresponding to the eye of the face of the ith person in the jth image frame with a second feature point corresponding to the eye of the face of the ith person in the Mth image frame, thereby comparing the position, direction, angle, and perspective of the face of the ith person.
[0125] The electronic device can compare the position, direction, angle, and perspective of the face of the i-th person by comparing a plurality of first feature points corresponding to the i-th person in the j-th image frame with a plurality of second feature points corresponding to the i-th person in the M-th image frame, and can calculate a suitability score.
[0126] In one embodiment, feature points may correspond to key locations within the subject's body. For example, key feature locations may include the eyes, nose, mouth, chin, and cheekbones within the face, as well as multiple locations corresponding to the neck connected to the face, the chest, stomach, and collarbone for determining the orientation of the upper body, and the hands, elbows, and shoulders for determining the posture of the upper body.
[0127] For example, the electronic device can compare the first feature point corresponding to the connection point of the neck and face of the ith person in the jth image frame with the second feature point corresponding to the connection point of the neck and face of the ith person in the Mth image frame, and compare how the connection points of the neck and face of the ith person are arranged. The connection point of the neck and face may correspond to a single feature point, but the part where the neck and face are connected may be defined as a line or a plane and may correspond to multiple feature points. The electronic device can determine whether the face of the ith person in the jth image frame and the face of the ith person in the Mth image frame can be naturally substituted by comparing the first feature point and the second feature point. Specifically, the more similar the arrangements of the first feature point and the second feature point corresponding to the connection point of the neck and face are, the easier it is to substitute the face of the ith person between the jth image frame and the Mth image frame, and the higher the suitability score can be calculated.
[0128] As another example, the electronic device may compare a first feature point corresponding to the upper body of the ith person in the jth image frame with a second feature point corresponding to the upper body of the ith person in the Mth image frame, thereby comparing the direction of the upper body of the ith person. The feature point corresponding to the upper body may include a plurality of feature points. By comparing the first feature point and the second feature point, the electronic device may determine whether the face of the ith person in the jth image frame and the face of the ith person in the Mth image frame can be naturally substituted. Specifically, the more similar the directions of the upper bodies determined based on the first feature point and the second feature point are, the easier it is to substitute the face of the ith person between the jth image frame and the Mth image frame, and a high suitability score may be calculated.
[0129] The electronic device can compare the position, direction, angle, and perspective of the face of the i-th person, as well as whether the connection between the face and the neck of the i-th person is natural and whether the direction of the face and the direction of the body are natural, by comparing at least one first feature point corresponding to the i-th person in the j-th image frame and at least one second feature point corresponding to the i-th person in the M-th image frame. The electronic device can calculate a suitability score for the i-th person between the j-th image frame and the M-th image frame based on the comparison result of the first feature point and the second feature point.
[0130] 3. Regarding determining the base image
[0131] In one embodiment, the electronic device may determine a base image based on an aesthetic score and a suitability score. In step S310, the electronic device may determine a base image based on the aesthetic score and the suitability score. For example, the base image may be determined by mathematical expression 3 calculated based on the aesthetic score and the suitability score.
[0132]
[0133] In mathematical expression 3, is identical to mathematical formula 2, and redundant explanations are omitted.
[0134] In mathematical expression 3, is identical to mathematical formula 1, and redundant explanations are omitted.
[0135] In mathematical expression 3, It can mean a mathematical formula that comprehensively considers the substitution suitability and shooting completion for the i-th person between the j-th image frame and the M-th image frame.
[0136] Mathematical expression 3 determines the M-th image frame with the highest overall substitution suitability and shooting completion for the i-th person in the j-th image frame, and for each person in the j-th image frame It may mean a mathematical formula for obtaining a value by adding all values and deriving the j-th image frame with the highest obtained value. Mathematical formula 3 may be a mathematical formula designed to select the best image among a plurality of images by considering the substitution suitability between a plurality of people included in the j-th image frame and each person included in another image frame and the degree of completion of the photography of a plurality of people included in the j-th image frame. However, this is only an example, and the technical idea of the present disclosure for determining a base image is not limited to Mathematical Formula 3.
[0137] 4. Regarding determining the source image
[0138] In one embodiment, the electronic device may determine a source image based on an aesthetic score and a suitability score. The electronic device may determine a source image among a plurality of images based on an aesthetic score and a suitability score of a target person included in the plurality of images. In step S320, the electronic device may determine the source image based on the aesthetic score and the suitability score. For example, the source image may be determined by mathematical expression 4 calculated based on the aesthetic score and the suitability score.
[0139]
[0140] In mathematical equation 4, is identical to mathematical formula 2, and redundant explanations are omitted.
[0141] In mathematical equation 4, is identical to mathematical formula 1, and redundant explanations are omitted.
[0142] In mathematical equation 4, It can mean a mathematical formula that comprehensively considers the substitution suitability and shooting completion for the i-th person between the j-th image frame and the M-th image frame.
[0143] Mathematical expression 4 may refer to a mathematical expression for determining the M-th image frame with the highest substitution suitability and shooting completion for the i-th person in the j-th image frame. Mathematical expression 4 may be a mathematical expression designed to select the best image in which the target person is best captured among a plurality of images by considering the substitution suitability between the target person included in the j-th image frame and the target person included in another image frame and the shooting completion of the target person included in the j-th image frame. However, this is merely an example, and the technical idea of the present disclosure for determining a source image is not limited to Mathematical expression 4.
[0144] In one embodiment, the source image may be determined from separately stored external images. In the present disclosure, the source image has been described as being determined from among a plurality of images acquired by the electronic device, but the source image may also be determined from received external images. The electronic device may receive external images of a first person from a separate server or database, and determine a source image for replacing the facial area of the first person in the base image from among the external images.
[0145] FIG. 4 is a flowchart illustrating a method for determining a base image among a plurality of images according to an embodiment of the present disclosure.
[0146] For convenience of explanation, parts that overlap with those described using Figure 2 are simplified or omitted.
[0147] Referring to FIG. 4, step S220 of FIG. 2 may include steps S410, S420, and S430 of FIG. 4.
[0148] In step S410, the electronic device can obtain a first aesthetic score for each of the plurality of images.
[0149] In one embodiment, the first aesthetic score may be a numerical score indicating the degree of completion of the photographic completion of the facial region of multiple people for each of the multiple images. The first aesthetic score may refer to an aesthetic score determined for any one of the multiple people included in the multiple images.
[0150] For example, the electronic device may obtain, for each of a plurality of images, a first aesthetic score determined for a first person among a plurality of people included in the plurality of images. The electronic device may obtain, for each of the plurality of images, a first aesthetic score determined for a second person among a plurality of people included in the plurality of images. The first aesthetic score may include aesthetic scores determined for each of the plurality of images, and for each of the plurality of people included in each of the plurality of images.
[0151] In one embodiment, an electronic device can detect a facial region of a person included in a plurality of images. The electronic device can obtain a first aesthetic score of the person based on the detected facial region. The facial region detection can generally be performed through object detection, but the technical concepts of the present disclosure are not limited thereto.
[0152] In one embodiment, the first aesthetic score may be determined based on whether the face of the subject included in the image is well captured. For example, the first aesthetic score may be determined based on preferences extracted from external images of the subject stored in a database.
[0153] In step S420, the electronic device can obtain a first suitability score for an image pair consisting of two images among a plurality of images.
[0154] In one embodiment, the first suitability score may be a score that quantifies the substitution suitability between pairs of facial regions of multiple people included in an image pair consisting of two images from among a plurality of images. The first suitability score may refer to a suitability score determined for any person among the multiple people included in the plurality of images.
[0155] For example, the electronic device may obtain a first suitability score determined for a first person among a plurality of people included in one target image and one comparison image among a plurality of images. The electronic device may obtain a first suitability score determined for a second person among a plurality of people included in one target image and one comparison image among a plurality of images. The first suitability score may include suitability scores determined for each of a plurality of people between one target image and one comparison image among a plurality of images. The plurality of people may be people included in both one target image and one comparison image.
[0156] In one embodiment, an electronic device can detect a facial region of a person included in a plurality of images. The electronic device can obtain a first suitability score of the person based on the detected facial region. The facial region detection can generally be performed through object detection, but the technical concept of the present disclosure is not limited thereto.
[0157] In one embodiment, an electronic device can detect a body region associated with a face of a person included in multiple images. The electronic device can obtain a first suitability score for the person based on the detected face and body regions. For example, the electronic device can determine the suitability score based on the degree of similarity between the face and the body pose associated with the face.
[0158] In step S430, the electronic device may determine one of the plurality of images as a base image based on the first aesthetic score and the first suitability score.
[0159] In one embodiment, the electronic device may evaluate a plurality of images based on a first aesthetic score and a first suitability score. For example, the equation for evaluating the plurality of images may include Mathematical Equation 3 described with reference to FIG. 3. The electronic device may evaluate the plurality of images according to Mathematical Equation 3, and determine the image with the highest score among the plurality of images as the base image. However, Mathematical Equation 3 for evaluating the plurality of images is merely an example, and the technical concept of the present disclosure is not limited to Mathematical Equation 3.
[0160] FIG. 5 is a flowchart illustrating a method for determining a target image among a plurality of images according to one embodiment of the present disclosure.
[0161] For convenience of explanation, parts that overlap with those described using Figures 2 and 4 are simplified or omitted.
[0162] In one embodiment, step S230 of FIG. 2 may include steps S510, S520, and S530 of FIG. 5.
[0163] In step S510, the electronic device can obtain a second aesthetic score for each of the plurality of images.
[0164] In one embodiment, the second aesthetic score may be a numerical score indicating the degree of completion of the photographic completion of the facial region of multiple people for each of the multiple images. The second aesthetic score may refer to an aesthetic score determined for a first person among the multiple people included in the multiple images.
[0165] For example, an electronic device may obtain a second aesthetic score for a first person to replace the facial region of the first person. Accordingly, the second aesthetic score for each of the multiple images is determined based on the first person included in the multiple images. The image with the highest second aesthetic score among the multiple images may indicate that the image best captures the facial region of the first person.
[0166] In step S520, the electronic device can obtain a second suitability score for each of the plurality of images.
[0167] In one embodiment, the second suitability score may be a score that quantifies the substitution suitability between a pair of facial regions of a first person included in an image pair consisting of two images among a plurality of images. The second suitability score may refer to a suitability score determined for the first person included in the plurality of images.
[0168] For example, the electronic device may obtain a second suitability score determined for a first person among a plurality of persons included in one target image and one comparison image among a plurality of images. Accordingly, the second suitability score between one target image and one comparison image is determined for the first person included in both images. A highest second aesthetic score between one target image and one comparison image may mean that the facial region of the first person in one target image is suitable for substituting the facial region of the first person in one comparison image.
[0169] In step S530, the electronic device may determine a source image for extracting a facial region of a first person from among a plurality of images based on the second aesthetic score and the second suitability score.
[0170] In one embodiment, the electronic device may evaluate a relationship between two images among a plurality of images for replacing a facial region of a first person based on a second aesthetic score and a second suitability score. For example, an equation for evaluating a relationship between two images among a plurality of images may include Mathematical Expression 4 described using FIG. 3. The electronic device may evaluate a relationship between one target image and one comparison image according to Mathematical Expression 4, and may determine a comparison image with the highest evaluation among the plurality of images based on one target image as a source image. However, Mathematical Expression 4 is merely an example, and the technical idea of the present disclosure is not limited to Mathematical Expression 4.
[0171] FIG. 6 is a flowchart illustrating a method of considering a user's preference to determine a target image according to one embodiment of the present disclosure.
[0172] For convenience of explanation, parts that overlap with those described using FIGS. 2 and 5 are simplified or omitted.
[0173] In one embodiment, step S510 of FIG. 5 may include steps S610, S620, and S630 of FIG. 6.
[0174] In step S610, the electronic device can acquire a plurality of first person images including the first person.
[0175] In one embodiment, the electronic device may obtain a plurality of first person images via a communication unit. For example, the electronic device may obtain a plurality of first person images via a server.
[0176] In one embodiment, the electronic device may obtain first person images from a memory. The memory may include a cluster that stores data of the same attribute, and the electronic device may obtain a plurality of first person images from the cluster for the first person.
[0177] In step S620, the electronic device can extract a preference for the first person from a plurality of first person images.
[0178] In one embodiment, the electronic device can extract a preference for a first person from the posture of the first person included in a plurality of first person images. The preference for the first person can be determined based on at least one of the first person's facial expression, clothing, hairstyle, eye blinking, head pose, hand pose, and whether or not the first person is covered included in the first person images.
[0179] For example, the preference for a first person may be determined based on the ratio of the first person's appearance included in a plurality of first person images. The first person's appearance may refer to facial expressions, clothing, hairstyle, blinking, head poses, hand poses, whether the first person is covered, etc. As a specific example, if the ratio of smiling expressions among the first person's appearances included in a plurality of first person images is high, the preference for smiling expressions may be determined to be high.
[0180] In step S630, the electronic device may determine a second aesthetic score based on the preference for the first person.
[0181] In one embodiment, the electronic device may determine a higher second aesthetic score for the first person among a plurality of images, as the image includes elements with a higher preference for the first person. The electronic device may determine a lower second aesthetic score for the first person among a plurality of images, as the image includes elements with a lower preference for the first person.
[0182] FIG. 7 is a flowchart illustrating a method for considering a user's pose to determine a target image according to one embodiment of the present disclosure.
[0183] For convenience of explanation, parts that overlap with those described using FIGS. 2 and 5 are simplified or omitted.
[0184] In one embodiment, step S520 of FIG. 5 may include steps S710, S720, and S730 of FIG. 7.
[0185] In step S710, the electronic device can extract base feature points of the first person from the base image. In step S720, the electronic device can extract target feature points of the first person from a plurality of images.
[0186] In one embodiment, the processor of the electronic device may include an artificial intelligence algorithm or an artificial intelligence network for obtaining feature points from a base image. The artificial intelligence algorithm or the artificial intelligence network may be an artificial intelligence model including a feature point extraction algorithm.
[0187] In one embodiment, the electronic device may extract base features of a first person from a base image using an artificial intelligence model. In one embodiment, the electronic device may extract target features of the first person from a plurality of images using an artificial intelligence model.
[0188] In step S730, the electronic device can determine a suitability score for replacing a face area of a first person in a plurality of images onto a base image by comparing base feature points and target feature points.
[0189] In one embodiment, the base feature points and the target feature points may each include position values of points corresponding to key locations on the body of the first person. The electronic device may compare the poses of the first person included in the base image and one of the multiple images by comparing the base feature points and the target feature points corresponding to the same location. The electronic device may determine a suitability score based on the comparison results.
[0190] For example, the electronic device may determine a degree of similarity between the facial poses of the first person included in each of the base image and one of the plurality of images by comparing base feature points and target feature points corresponding to key locations of the face of the first person, respectively. The electronic device may determine a suitability score for replacing a facial area of the first person in one of the plurality of images onto the base image by considering the degree of similarity between the facial poses of the first person included in each of the base image and one of the plurality of images.
[0191] As another example, the electronic device may determine a degree of similarity between body poses of the first person included in each of the base image and one of the plurality of images by comparing base feature points and target feature points corresponding to major locations of the body of the first person, respectively. The electronic device may determine a suitability score for replacing a face area of the first person in one of the plurality of images onto the base image by considering the degree of similarity between body poses of the first person included in each of the base image and one of the plurality of images.
[0192] FIG. 8 is a conceptual diagram illustrating a method of filtering a plurality of images to remove a blurry image and then generating a corrected image based on the filtered plurality of images according to one embodiment of the present disclosure.
[0193] For convenience of explanation, parts that overlap with those described using Figure 1 are simplified or omitted.
[0194] Referring to FIG. 8, in one embodiment, the electronic device may include a plurality of images (110; 111, 112, 113, 114, 115, 116). The plurality of images (110) may be images captured continuously over a set period of time.
[0195] In one embodiment, the plurality of images (110) may have different blurrinesses. Blurriness may refer to the degree of blurring, whereby details in the image are smoothly blurred or obscured.
[0196] For example, blurriness can be determined by various factors, such as incorrect focus setting of the camera lens, movement of the subject during shooting, shutter speed when shooting (for example, a slow shutter speed can cause blurriness in the image due to movement or hand shake), lens quality, and information loss that occurs when compressing and storing the image. Of course, the technical concept of the present disclosure does not limit the factors that affect blurriness.
[0197] In one embodiment, the electronic device can set a threshold. The electronic device can determine an image among a plurality of images (110) whose blurriness exceeds the threshold as a blurry image. The electronic device can determine an image among a plurality of images (110) whose blurriness does not exceed the threshold as a clear image.
[0198] For example, as illustrated in FIG. 8, the electronic device can determine the blurriness of each of the plurality of images (110). The electronic device can determine that the blurriness of the first image (111) and the sixth image (116) each exceeds a threshold value. The electronic device can determine the first image (111) and the sixth image (116) as blurry images. The electronic device can determine that the blurriness of the second to fifth images (112, 113, 114, 115) each does not exceed a threshold value. The electronic device can determine the second to fifth images (112, 113, 114, 115) as clear images.
[0199] In one embodiment, the electronic device can select a clear image by filtering a plurality of images (110). The electronic device can select the second to fifth images (112, 113, 114, 115) as clear images among the plurality of images (110). The electronic device can perform the method for generating a corrected image described in FIG. 1 based on the clear images. This will be briefly described below.
[0200] In one embodiment, the electronic device can obtain a plurality of sharp images, namely, second image to fifth image (112, 113, 114, 115), by filtering a plurality of images (110).
[0201] The electronic device may determine any one of the second to fifth images (112, 113, 114, 115) as the base image. For example, the electronic device may determine the second image (112) as the base image.
[0202] In one embodiment, the electronic device may determine any one of the sharp images as a source image to replace the face area (A1_2) of the first person in the base image. The electronic device may determine the fourth image (114) among the second to fifth images (112, 113, 114, 115) including the first person as the source image for the first person. The source image for the first person may refer to a source image selected to replace the face area (A1_2) of the first person in the base image.
[0203] In one embodiment, the electronic device may extract a facial region (A1_1) of a first person included in a source image of the first person. For example, the electronic device may extract a facial region (A1_1) of the first person from a fourth image (114).
[0204] In one embodiment, the electronic device may replace the face area (A1_2) of the first person in the base image with the face area (A1_1) of the first person in the source image.
[0205] FIG. 9 is a conceptual diagram illustrating an operation of generating a 3D face model based on an image including a first person according to one embodiment of the present disclosure.
[0206] In one embodiment, the electronic device can generate a 3D face model (950) based on a plurality of images (910). The electronic device can correct the image based on the 3D face model (950), and a specific embodiment of correcting the image based on the 3D face model (950) will be described below using FIGS. 11 to 13.
[0207] In one embodiment, the electronic device may acquire a plurality of images (910) of a first person. The plurality of images (910) may be images including a facial area of the first person, and may be images taken from various sides of the facial area of the first person.
[0208] The plurality of images (910) may be images captured continuously over a set period of time. The plurality of images (910) may include an image sequence, which refers to a set of images arranged in order over time. For example, the plurality of images (910) may refer to images including a first person among the plurality of images (110) described using FIG. 1. The plurality of images (910) may be separately stored images. The electronic device may obtain the plurality of images (910) of the first person from a cluster for the first person. The cluster for the first person may be configured to collect and store images of the first person. Of course, the electronic device may also obtain the plurality of images (910) through a separate server.
[0209] In one embodiment, the electronic device can generate a 3D face model (950) by 3D modeling a plurality of images (910). The electronic device can include a processor, an artificial intelligence algorithm or an artificial intelligence network for obtaining a 3D face model (950) of a target person from a plurality of 2D images including the target person. The artificial intelligence algorithm or the artificial intelligence network can be an artificial intelligence model including an algorithm for generating the 3D face model (950).
[0210] In one embodiment, the plurality of images (910) may correspond to the plurality of images (110) of FIG. 1. For example, the electronic device may acquire a plurality of images (110) that are continuously captured, as described in FIG. 1. The captured plurality of images (110) may be images including a target person. The electronic device may acquire a plurality of images (110) captured, for example, through a motion photo function. The electronic device may acquire a plurality of images (110) that are stored in a video format. The electronic device may generate a 3D face model (950) regarding the face of the target person based on the acquired plurality of images (110). For example, the electronic device may generate a 3D face model regarding the face of the first person based on the acquired plurality of images (110) including the first person.
[0211] In one embodiment, the 3D face model (950) may include an image or simulation data that three-dimensionally represents the head of a target person. The 3D face model (950) may represent the curvature, parts, texture, etc. of the face of the target person. The 3D face model (950) mainly represents the face of the target person, but is not limited to the face. For example, the 3D face model (950) may also include an image or simulation data that further includes the neck and parts of the upper body connected to the face.
[0212] FIG. 10 is a flowchart illustrating a method for correcting a facial region of a base image based on a 3D facial model according to an embodiment of the present disclosure.
[0213] For convenience of explanation, parts that overlap with those described using FIG. 2 and FIG. 9 are simplified or omitted.
[0214] Referring to FIG. 10, step S250 of FIG. 2 may include steps S1010, S1020, and S1030 of FIG. 10.
[0215] In step S1010, the electronic device can generate a 3D face model of the first person based on a plurality of first person images including the first person.
[0216] In one embodiment, an electronic device may acquire first person images. The first person images may be stored in a cluster for the first person or may be stored via a separate server. The electronic device may acquire the first person images from the cluster or the separate server using a communication unit.
[0217] In one embodiment, an electronic device may generate a 3D face model based on a plurality of first person images. The electronic device may generate the 3D face model by 3D modeling the plurality of first images. The electronic device may include a processor, an artificial intelligence algorithm or an artificial intelligence network for obtaining a 3D face model of the first person from a plurality of first person images including the first person. The artificial intelligence algorithm or the artificial intelligence network may be an artificial intelligence model including an algorithm for generating a 3D face model.
[0218] In step S1020, the electronic device can correct the extracted facial region based on the 3D facial model.
[0219] In one embodiment, the electronic device may extract a facial region of a first person within a source image. The extracted facial region may be an image corresponding to the facial region of the first person within the source image for replacing the face of the first person. The electronic device may correct the facial region based on a 3D facial model.
[0220] For example, if the face orientation of a face region extracted from a source image does not match the face orientation of a first person in the base image, the electronic device may correct the face orientation of the extracted face region to match the face orientation of the first person in the base image. The electronic device may correct the face region by changing the face orientation of the extracted face region based on a 3D shape and displaying the face region with the changed face orientation again on a 2D image. An embodiment related to this will be described in detail below using FIG. 11.
[0221] As another example, if the lighting for the extracted facial region does not match the lighting for the first person in the base image, the electronic device may correct the lighting for the extracted facial region to match the lighting for the first person in the base image. The electronic device may correct the facial region by changing the lighting for the extracted facial region based on the 3D shape and displaying the face region with the changed lighting again on the 2D image. An embodiment related to this will be described in detail below using FIG. 12.
[0222] As another example, if the extracted face region contains a region that obscures the face, the electronic device can extract a face fragment corresponding to the obscured face region from the 3D face model. The electronic device can then combine the face fragment extracted from the 3D face model with the face region extracted from the source image based on the 3D face model. The electronic device can then correct the face region by displaying the combined face region again on a 2D image. An embodiment related to this will be described in detail below with reference to FIG. 13.
[0223] In step S1030, the electronic device can generate a corrected image by synthesizing the corrected face area onto a base image.
[0224] In one embodiment, the electronic device can correct the base image by replacing the facial area of the first person in the base image with the facial area of the first person in the source image. The description of step S1030 is not different from the description of step S250, and is therefore omitted.
[0225] FIG. 11 is a conceptual diagram illustrating a method for rotating a pose of a face region based on a 3D face model according to one embodiment of the present disclosure.
[0226] For convenience of explanation, parts that overlap with those described using Figures 1 to 10 are simplified or omitted.
[0227] Referring to FIG. 11, in one embodiment, an electronic device may determine a base image (411) from among a plurality of images and a source image (311) for replacing a face area of a first person. The electronic device may obtain head pose information of the first person included in the source image (311). The head pose information may include information on at least one of a direction and an angle of the face of the first person.
[0228] For example, the direction and angle of the face of the first person can be determined using the xyz coordinate system. The electronic device can obtain head pose information of the first person included in the source image (311), and the direction and angle of the face of the first person can be determined using three axes of x1, y1, and z1.
[0229] In one embodiment, the electronic device can obtain head pose information of a first person included in a base image (411). For example, the direction and angle of the face of the first person included in the base image (411) can be determined using three axes: x2, y2, and z2.
[0230] In one embodiment, the electronic device can obtain a 3D face model (950). The electronic device can determine a head pose of the 3D face model (950) based on head pose information of the first person included in the base image (411). As illustrated in FIG. 11, the electronic device can rotate the face direction and angle of the 3D face model (950) based on the head pose information of the first person included in the base image (411). The electronic device can change the face direction and angle of the 3D face model (950) to match the head pose information of the first person included in the base image (411). The face direction and angle of the 3D face model (950) can be changed to be determined using three axes of x2, y2, and z2.
[0231] In one embodiment, the electronic device may rotate the face region of the first person included in the source image (311) based on the face orientation and angle of the changed 3D face model (950). The electronic device may change the face region of the first person included in the source image (311) to match the face orientation and angle of the changed 3D face model (950).
[0232] For example, the electronic device can overlap a face area of a first person included in the source image (311) on a 3D face model (950) with a changed face direction and angle. The electronic device can change the face area of the first person included in the source image (311) by converting the 3D face model (950) with the overlapped face area of the first person into a 2D image. The face area of the first person included in the source image (311) can be changed from a face direction and angle determined by three axes of x1, y1, and z1 to a face direction and angle determined by three axes of x2, y2, and z2.
[0233] The electronic device can obtain a source image (511) with a changed face direction and angle of the first person.
[0234] FIG. 12 is a conceptual diagram illustrating a method for changing illumination of a facial region based on a 3D facial model according to one embodiment of the present disclosure.
[0235] For convenience of explanation, parts that overlap with those described using Figures 1 to 11 are simplified or omitted.
[0236] Referring to FIG. 12, in one embodiment, an electronic device may determine a base image (412) from among a plurality of images and a source image (312) for replacing a facial area of a first person. The electronic device may obtain lighting information about the first person included in the source image (312). The electronic device may obtain lighting information about the first person included in the base image (412).
[0237] Lighting information for the first person may include information about the brightness and darkness of the face as light shines on the face of the first person, taking into account the curvature of the face.
[0238] In one embodiment, the electronic device may acquire a 3D face model (950). The electronic device may determine brightness considering facial curvature on the 3D face model (950) based on lighting information for the first person included in the base image (412). As illustrated in FIG. 12, the electronic device may change the brightness of the 3D face model (950) based on lighting information for the first person included in the base image (412). The electronic device may change the brightness of the 3D face model (950) to match the lighting information for the first person included in the base image (412).
[0239] In one embodiment, the electronic device may change the brightness of the face of the first person included in the source image (312) based on the brightness of the changed 3D face model (950). The electronic device may change the face area of the first person included in the source image (312) to match the brightness of the changed 3D face model (950).
[0240] For example, the electronic device can overlap a facial area of a first person included in the source image (312) on a 3D facial model (950) in which the brightness of the face has been changed. The electronic device can change the facial area of the first person included in the source image (312) by converting the 3D facial model (950) in which the facial area of the first person has been overlapped into a 2D image.
[0241] The electronic device can obtain a source image (511) with changed contrast toward the first person.
[0242] FIG. 13 is a conceptual diagram illustrating a method for correcting an occluded area within a facial region based on a 3D facial model according to an embodiment of the present disclosure.
[0243] For convenience of explanation, parts that overlap with those described using Figures 1 to 12 are simplified or omitted.
[0244] Referring to FIG. 13, in one embodiment, an electronic device may determine a source image (313) for replacing a base image and a face area of a first person among a plurality of images. The electronic device may determine whether an occlusion area (315) is formed on the face of the first person included in the source image (313).
[0245] In one embodiment, the electronic device can generate a corrected image by replacing a face area in the base image with a face area in the source image according to the embodiment described using FIGS. 1 to 12 when no occlusion area (315) is formed on the face of the first person included in the source image (313).
[0246] In one embodiment, when a occlusion area (315) is formed on the face of a first person included in the source image (313), the electronic device may obtain a first face image excluding the occlusion area (315) from the source image (313). The first face image may be an image from the source image (313) from which the occlusion area (315) is excluded.
[0247] In one embodiment, the electronic device can obtain a 3D face model (950). The electronic device can obtain the 3D face model (950) based on face images of a first person. In one embodiment, the electronic device can obtain an image corresponding to an occlusion area on the face of the first person included in the source image (313) based on a three-dimensional shape of the 3D face model (950). The electronic device can obtain a second face image (955) corresponding to an occlusion area (315) on the face of the first person included in the source image (313).
[0248] In one embodiment, the electronic device can remove the occlusion area (315) on the face of the first person by combining the first face image and the second face image (955). The electronic device can obtain a third face image (513) from which the occlusion area (315) has been removed. The third face image (513) may be an image combining the first face image and the second face image (955), and may be a face image of the first person from which the occlusion area (315) has been removed.
[0249] For example, the electronic device can adjust the 3D face model (950) to match the face direction of the first person included in the source image (313). The electronic device can combine the first face image and the second face image (955) by overlapping the first face image on the 3D face model (950). The device can obtain a third face image (513) from which the occlusion area (315) is removed by converting the 3D face model (950) with the first face image overlapped thereon into a 2D image.
[0250] FIG. 14 is a conceptual diagram illustrating a method for reflecting an object on a final corrected image when there is an object in a face area of a base image according to one embodiment of the present disclosure.
[0251] For convenience of explanation, parts that overlap with those described using Figures 1 to 13 are simplified or omitted.
[0252] Referring to FIG. 14, the electronic device may determine a base image (414) from among a plurality of images and a source image (314) for replacing the facial area of the first person. The electronic device may detect an object (O1) located within the facial area of the first person within the base image (414). The detected object (O1) may be something worn by the first person within the base image (414).
[0253] The detected object (O1) may refer to something captured together with the first person in the base image (414). For example, the object (O1) may refer to accessories worn by the first person, and may specifically include glasses, a hat, earrings, a mask, etc.
[0254] In one embodiment, the object (O1) may be located within the face of the first person. In one embodiment, the electronic device may detect the object (O1) located within the face of the first person in the base image (414) and insert the detected object (O1) into the same or corresponding position in the second corrected image (614). The area occupied by the detected object (O1) in the base image (414) may replace the corresponding area in the second corrected image (614).
[0255] In one embodiment, the processor of the electronic device may include an artificial intelligence algorithm or an artificial intelligence network for detecting or classifying objects from an image. The artificial intelligence algorithm or the artificial intelligence network may be an artificial intelligence model including a feature extraction algorithm. The electronic device may use the artificial intelligence model to detect or classify objects within a facial region.
[0256] For example, the electronic device can determine a base image (414) from among a plurality of images. The electronic device can detect an object (O1) located within the face of a first person within the base image (414).
[0257] In one embodiment, the electronic device may obtain a face region of a first person from the source image (314). In step S1410, the electronic device may synthesize the face region of the first person of the source image (314) onto the face region of the first person of the base image (414). The electronic device may replace the face region of the first person of the base image (414) with the face region of the first person of the source image (314).
[0258] As a result, the electronic device can obtain a first corrected image (514). The first corrected image (514) may be an image in which the face area of the first person in the base image (414) is replaced with the face area of the first person in the source image (314), based on the base image (414).
[0259] In one embodiment, in step S1420, the electronic device may extract an image of an object (O1) located within the face of the first person from the base image (414). The electronic device may synthesize the image of the extracted object (O1) onto the first corrected image (514). As a result, the electronic device may obtain a second corrected image (614). The second corrected image (614) may be an image in which the image of the object (O1) extracted from the base image (414) is added based on the first corrected image (514). Accordingly, the second corrected image (614) may be an image in which the face area of the first person in the base image (414) is replaced with the face area of the first person in the source image (314), based on the base image (414), and the image of the object (O1) worn by the first person in the base image (414) is synthesized.
[0260] Hereinafter, with reference to FIG. 15, the configuration of an electronic device for performing the image processing operations described so far will be described. FIG. 15 is a block diagram illustrating the configuration of an electronic device according to an embodiment of the present disclosure.
[0261] For convenience of explanation, parts that overlap with those described using Figures 1 to 14 are simplified or omitted.
[0262] Referring to FIG. 15, an electronic device (1000) according to an embodiment may include an input / output interface (1100), a memory (1200), and a processor (1300). However, the components of the electronic device (1000) are not limited to the above-described examples, and the electronic device (1000) may include more or fewer components than the above-described components. In an embodiment, some or all of the input / output interface (1100), the memory (1200), and the processor (1300) may be implemented in the form of a single chip, and the processor (1300) may include one or more processors.
[0263] The input / output interface (1100) may include an input interface (e.g., touch screen, hard button, microphone, etc.) for receiving control commands or information from a user, and an output interface (e.g., display panel, speaker, etc.) for displaying the results of execution of an operation according to the user's control or the status of the electronic device (1000).
[0264] For example, the electronic device (1000) can acquire a plurality of images based on a user's image capturing command obtained through the input / output interface (1100). The processor (1300) of the electronic device (1000) can determine a base image and a source image based on the plurality of images, and perform the image processing operations described using FIGS. 1 to 14.
[0265] The memory (1200) is a configuration for storing various programs or data, and may be configured as a storage medium such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD, or a combination of storage media. The memory (1200) may not exist separately and may be configured to be included in the processor (1300). The memory (1200) may be configured as a volatile memory, a non-volatile memory, or a combination of volatile memory and non-volatile memory. Programs or instructions for performing operations according to the embodiments described with reference to FIGS. 1 to 14 may be stored in the memory (1200). The memory (1200) may also provide stored data to the processor (1300) upon request of the processor (1300).
[0266] The processor (1300) controls a series of processes so that the electronic device (1000) operates according to the embodiments described with reference to FIGS. 1 to 14, and may be composed of one or more processors. One or more processors included in the processor (1300) may be circuitry such as a System on Chip (SoC), an Integrated Circuit (IC), etc. In this case, one or more processors may be a general-purpose processor such as a CPU, an AP, a Digital Signal Processor (DSP), a graphics-only processor such as a GPU, a Vision Processing Unit (VPU), or an artificial intelligence-only processor such as an NPU. For example, when one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0267] The processor (1300) can write data to the memory (1200) or read data stored in the memory (1200), and in particular, process data according to predefined operation rules or artificial intelligence models by executing a program or at least one instruction stored in the memory (1200). Accordingly, the processor (1300) can perform the operations described in the embodiments described above, and the operations described as being performed by the electronic device (1000) in the embodiments described above can be regarded as being performed by the processor (1300) unless otherwise specifically described.
[0268] A method according to one embodiment may include the step of acquiring a plurality of images, each including a plurality of persons. The method may include the step of determining one of the plurality of images as a base image. The method may include the step of determining a source image from the plurality of images based on a completion level of shooting of each of the plurality of images and a swap compatibility of each of the plurality of images. The method may include the step of extracting a facial region of one of the plurality of persons from the source image. The method may include the step of generating a corrected image by synthesizing the extracted facial region onto a base image. For each image from the plurality of images, the completion level of shooting may include the completion level of shooting a facial region of a person in each image, and the swap compatibility may include a swap compatibility between a facial region of a person in each image and a facial region of a person in a base image.
[0269] In one embodiment, the plurality of images may include a plurality of images captured continuously over a set period of time.
[0270] In one embodiment, the step of determining one of the plurality of images as the base image may include the step of obtaining, for each of the plurality of images, a first aesthetic score, which is a score that quantifies the degree of completion of photographing a facial region of each of the plurality of people. The step of determining one of the plurality of images as the base image may include the step of obtaining, for an image pair composed of two images among the plurality of images, a first suitability score, which is a score that quantifies the degree of substitution suitability between a facial region of one person included in each of the two images. The step of determining one of the plurality of images as the base image may include the step of determining one of the plurality of images as the base image based on the first aesthetic score and the first suitability score.
[0271] In one embodiment, the step of determining the source image may include a step of obtaining, for each of the plurality of images, a second aesthetic score, which is a score quantifying the degree of completion of photographing facial regions of the first person. The step of determining the source image may include a step of obtaining, for each of the plurality of images, a second suitability score, which is a score quantifying the suitability of substitution for the facial region of the first person. The step of determining the source image may include a step of determining the source image based on the second aesthetic score and the second suitability score.
[0272] In one embodiment, the step of obtaining a second aesthetic score may include the step of obtaining a plurality of first person images including the first person. The step of obtaining the second aesthetic score may include the step of extracting a preference for the first person from the plurality of first person images. The step of obtaining the second aesthetic score may include the step of determining the second aesthetic score based on the preference for the first person.
[0273] In one embodiment, the preference of the first person may be determined based on at least one of the first person's facial expression, clothing, hairstyle, eye blinking, head pose, hand pose, and whether or not the first person is covered included in the first person image.
[0274] In one embodiment, the step of obtaining the second suitability score may include the step of extracting a base feature point related to a location of a body of the first person from a base image. The step of obtaining the second suitability score may include the step of extracting a target feature point related to a location of a body of the first person from a plurality of images. The step of obtaining the second suitability score may include the step of determining a second suitability score for each of the plurality of images by comparing the base feature point and the target feature point.
[0275] In one embodiment, the corrected image may include an image in which a face region of a first person in a base image is replaced with an extracted face region.
[0276] In one embodiment, the method may further include a step of obtaining a degree of blurriness of each of the plurality of images. The method may further include a step of selecting at least one image from among the plurality of images, the blurriness of which does not exceed a threshold. The step of determining a base image may be determining a base image from among the at least one selected image. The step of determining a source image may be determining a source image from among the at least one selected image.
[0277] In one embodiment, the step of generating a corrected image may include the step of generating a 3D facial model of the person based on a plurality of first facial images including the person. The step of generating the corrected image may include the step of correcting an extracted facial region based on the 3D facial model. The step of generating the corrected image may include the step of generating the corrected image by synthesizing the corrected facial region onto a base image.
[0278] An electronic device according to an embodiment may include an input / output interface for receiving a user input requesting image processing and outputting an image processed according to the user input, at least one memory storing one or more commands for processing the image, and at least one processor. The electronic device may acquire a plurality of images including a plurality of people by having the at least one processor execute a program or at least one command stored in the memory. The electronic device may determine one of the plurality of images as a base image. The electronic device may determine a source image among the plurality of images based on a completion level of shooting of each of the plurality of images and a swap compatibility of each of the plurality of images. The electronic device may extract a facial region of one of the plurality of people from the source image. The electronic device may generate a corrected image by synthesizing the extracted facial region onto the base image. For each image among the plurality of images, the shooting completeness may include the completeness of shooting the facial area of the person in each image, and the substitution suitability may include the substitution suitability between the facial area of the person in each image and the facial area of the person in the base image.
[0279] In one embodiment, the plurality of images may include a plurality of images captured continuously over a set period of time.
[0280] In one embodiment, the electronic device may obtain a first aesthetic score, which is a score that quantifies the degree of completion of capturing a facial region of each of the plurality of people, for each of the plurality of images. The electronic device may obtain a first suitability score, which is a score that quantifies the degree of substitution suitability between the facial regions of one person included in each of the two images, for an image pair composed of two images among the plurality of images. The electronic device may determine one of the plurality of images as a base image based on the first aesthetic score and the first suitability score.
[0281] In one embodiment, when determining a source image, the electronic device may obtain a second aesthetic score, which is a numerical score indicating the degree of completion of photographing facial regions of a first person, for each of the plurality of images. The electronic device may obtain a second suitability score, which is a numerical score indicating the suitability of substitution for the facial region of the first person, for each of the plurality of images. The electronic device may determine the source image based on the second aesthetic score and the second suitability score.
[0282] In one embodiment, the electronic device may obtain a plurality of first person images including a first person to obtain an aesthetic score. The electronic device may extract a preference for the first person from the plurality of first person images. The electronic device may determine a second aesthetic score based on the preference for the first person.
[0283] In one embodiment, the electronic device may, in obtaining a suitability score, extract a base feature point related to a body location of a first person from a base image. The electronic device may extract a target feature point related to a body location of the first person from a plurality of images. The electronic device may determine a second suitability score for each of the plurality of images by comparing the base feature point and the target feature point.
[0284] In one embodiment, the correction image may be an image in which the face region of the first person in the base image is replaced with the extracted face region.
[0285] In one embodiment, an electronic device can obtain the degree of blurriness of each of a plurality of images. The electronic device can select at least one image from among the plurality of images, the degree of blurriness of which does not exceed a threshold. The electronic device can determine a base image from among the at least one selected image. The electronic device can determine a source image from among the at least one selected image.
[0286] In one embodiment, the electronic device may generate a 3D face model of the first person based on a plurality of first person images including the first person, when generating a corrected image. The electronic device may correct an extracted face region based on the 3D face model. The electronic device may generate the corrected image by synthesizing the corrected face region onto a base image.
[0287] A non-transitory computer-readable recording medium having recorded thereon a program for performing any one of the methods according to one embodiment of the present disclosure on a computer may be provided.
[0288] Various embodiments of the present disclosure may be implemented or supported by one or more computer programs, and the computer programs may be formed from computer-readable program code and embodied in a computer-readable medium. In the present disclosure, "application" and "program" may refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in computer-readable program code. "Computer-readable program code" may include various types of computer code, including source code, object code, and executable code. "Computer-readable medium" may include various types of media that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), a hard disk drive (HDD), a compact disc (CD), a digital video disc (DVD), or various types of memory.
[0289] Additionally, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, a 'non-transitory storage medium' is a tangible device and may exclude wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Meanwhile, this 'non-transitory storage medium' does not distinguish between cases where data is permanently stored in the storage medium and cases where it is temporarily stored. For example, a 'non-transitory storage medium' may include a buffer where data is temporarily stored. A computer-readable medium may be any available medium that can be accessed by a computer, and may include both volatile and non-volatile media, and removable and non-removable media. A computer-readable medium includes a medium on which data can be permanently stored and a medium on which data can be stored and later overwritten, such as a rewritable optical disk or an erasable memory device.
[0290] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0291] The above description of the present disclosure is for illustrative purposes only, and those skilled in the art will appreciate that the present disclosure can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present disclosure. For example, suitable results can be achieved even if the described techniques are performed in a different order than the described method, and / or components of the systems, structures, devices, circuits, etc. described are combined or combined in a different form than the described method, or are replaced or substituted by other components or equivalents. Therefore, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. For example, each component described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined form.
[0292] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.
Claims
1. A step of obtaining multiple images each containing multiple characters; A step of determining one of the plurality of images as a base image; A step of determining a source image among the plurality of images based on the completion level of shooting of each of the plurality of images and the swap compatibility of each of the plurality of images; A step of extracting a facial region of one of the plurality of persons from the source image; and A step of generating a correction image by synthesizing the extracted facial area onto the base image, A method wherein, for each image among the plurality of images, the shooting completeness includes the completeness of shooting the face area of the person in each image, and the substitution suitability includes the substitution suitability between the face area of the person in each image and the face area of the person in the base image.
2. In paragraph 1, A method wherein the plurality of images include a plurality of images continuously captured over a set period of time.
3. In either of paragraphs 1 and 2, The step of determining one of the above multiple images as a base image is: For each of the plurality of images, a step of obtaining a first aesthetic score, which is a score that quantifies the degree of completion of the photographing of the facial area of each of the plurality of people; A step of obtaining a first suitability score, which is a score that quantifies the substitution suitability between the facial regions of one person included in each of the two images, for an image pair composed of two images among the plurality of images; and A method comprising the step of determining one of the plurality of images as the base image based on the first aesthetic score and the first suitability score.
4. In any one of the clauses 1 to 3, The step of determining the above source image is: For each of the above multiple images, a step of obtaining a second aesthetic score, which is a score that quantifies the degree of completion of the photographing of the facial areas of the first person; For each of the plurality of images, a step of obtaining a second suitability score, which is a score that quantifies the suitability for substitution for the face area of the first person; and A method comprising the step of determining the source image based on the second aesthetic score and the second suitability score.
5. In paragraph 4, The step of obtaining the second suitability score is: A step of extracting a base feature point regarding the position of the body of the first person from the base image; A step of extracting a target feature point related to the location of the body of the first person from the plurality of images; and A method comprising the step of determining the second suitability score for each of the plurality of images by comparing the base feature points with the target feature points.
6. In any one of paragraphs 1 to 5, A method wherein the above-mentioned correction image includes an image in which the face area of the first person in the base image is replaced with the extracted face area.
7. In any one of paragraphs 1 to 6, A step of obtaining the degree of blurriness of each of the plurality of images; Further comprising a step of selecting at least one image among the plurality of images, the blurriness of which does not exceed a threshold, The step of determining the base image is to determine the base image among at least one of the selected images, A method, wherein the step of determining the source image is to determine the source image from among at least one of the selected images.
8. In any one of paragraphs 1 to 7, The step of generating the above correction image is: A step of generating a 3D face model of the person based on a plurality of first person images including the person; A step of correcting the extracted facial region based on the 3D facial model; and A method comprising the step of generating the corrected image by synthesizing the corrected face area onto the base image.
9. An input / output interface for receiving user input requesting image processing and outputting an image processed according to the user input; At least one memory storing one or more commands for processing an image; and Contains at least one processor, The electronic device, by causing at least one processor to execute a program or at least one instruction stored in the memory, Obtain multiple images containing multiple characters, Decide on one of the above multiple images as a base image, Based on the completion level of shooting of each of the plurality of images and the swap compatibility of each of the plurality of images, a source image is determined among the plurality of images, Extracting a face region of one of the plurality of people from the source image, A corrected image is created by synthesizing the extracted facial area onto the base image, An electronic device, wherein, for each image among the plurality of images, the shooting completeness includes the completeness of shooting the face area of the person in each image, and the substitution suitability includes the substitution suitability between the face area of the person in each image and the face area of the person in the base image.
10. In paragraph 9, An electronic device wherein the plurality of images include a plurality of images captured continuously over a set period of time.
11. In any one of the 9th and 10th clauses, The above electronic device, For each of the plurality of images, a first aesthetic score is obtained, which is a numerical score indicating the degree of completion of the shooting of the facial area of each of the plurality of people, For an image pair consisting of two images among the plurality of images, a first suitability score is obtained, which is a score that quantifies the substitution suitability between the facial regions of one person included in each of the two images, An electronic device that determines one of the plurality of images as the base image based on the first aesthetic score and the first suitability score.
12. In any one of the clauses 9 to 11, The electronic device determines the source image, For each of the above multiple images, a second aesthetic score is obtained, which is a score that quantifies the degree of completion of the shooting of the facial areas of the first person, For each of the plurality of images, a second suitability score is obtained, which is a score that quantifies the suitability for substitution for the face area of the first person, An electronic device that determines the source image based on the second aesthetic score and the second suitability score.
13. In paragraph 12, The electronic device obtains the suitability score, From the above base image, a base feature point is extracted regarding the position of the body of the first person, From the above plurality of images, target feature points are extracted regarding the location of the body of the first person, An electronic device that determines the second suitability score for each of the plurality of images by comparing the base feature points with the target feature points.
14. In any one of the clauses 9 to 13, The above electronic device, Obtain the degree of blurriness of each of the above multiple images, Among the plurality of images, at least one image is selected whose blurriness does not exceed a threshold, determining the base image from at least one of the selected images; An electronic device that determines the source image from among at least one of the selected images.
15. A computer-readable recording medium having recorded thereon a program for performing the method of any one of clauses 1 to 8 on a computer.
Citation Information
Patent Citations
3D human model generation device and program
JP2022033237A
Method and system for recognizing face of person included in digital data by using feature data
KR100840021B1
Apparatus and method for taking a picture continously
KR101906827B1
Method and device for determining facial image quality, electronic device and computer storage medium
KR102320649B1
KR20240017665A