Generating images with improved head pose or face region

A computer-implemented method using head and face editors generates composite images with aligned head positions and favorable facial expressions, addressing the challenge of inconsistent group photographs by integrating source and target images effectively.

JP2026500595AActive Publication Date: 2026-01-08GOOGLE LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025519836
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-03
Filing Date
2024-10-03
Publication Date
2026-01-08
Estimated Expiration
2044-10-03

AI Technical Summary

Technical Problem

Generating high-quality group photographs where all individuals have favorable facial expressions and aligned head positions is challenging due to variations in facial angles and head orientations.

Method used

A computer-implemented method that uses a head editor and/or face editor to generate composite images by replacing or adjusting facial features and head positions based on source images, employing techniques like interpolation, inpainting, and machine learning models to maintain realism.

Benefits of technology

The method seamlessly integrates source and target images, ensuring all subjects in the composite image have favorable facial expressions and aligned head positions, enhancing the quality of group photographs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026500595000001_ABST
    Figure 2026500595000001_ABST
Patent Text Reader

Abstract

The media application receives a set of images including a source image and a target image. The source image and the target image include at least a subject. The media application determines whether to use one or more editors selected from the group of a head editor, a face editor, or a combination thereof. In response to determining to use the head editor, the media application generates a composite image by replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from the source head of the subject in the source image, and replacing neck pixels associated with the target neck and shoulder pixels associated with the target shoulders, including a region between the target head and the target torso, with interpolated regions generated from interpolation between the source image and the target image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a non-provisional application claiming priority under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 542,283, entitled "Generating a Group Photo with Head Pose and Facial Recognition Improvements," filed October 3, 2023, the contents of which are incorporated herein by reference in their entirety. [Background technology]

[0002] Group photographs are a common way to commemorate an event. Getting an image in which everyone is smiling and looking toward the camera can be difficult because the more people in a photo, the greater the likelihood that at least one person will not have the most favorable facial expression. For example, some people may have their mouths open, others may have their eyes closed, and others may not be looking at the camera. Also, some people may have their head tilted in a different direction than the others in the image, be at an angle to the camera, or otherwise be posed in a way that is undesirable for a high-quality photograph.

[0003] The background art provided herein is intended to generally present the context of the present disclosure. To the extent described in this background art section, the work of the inventors named herein, and aspects of the present disclosure that may not qualify as prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure. Summary of the Invention

[0004] A computer-implemented method includes receiving a set of images including a source image and a target image, where the source image and the target image include at least a subject. The method further includes determining, based on the set of images, whether to use one or more editors selected from the group consisting of a head editor, a face editor, or a combination thereof. In response to determining to use the head editor, the method further includes generating a composite image by replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from the source head of the subject in the source image, and replacing neck pixels associated with the target neck and shoulder pixels associated with the target shoulders, including a region between the target head and the target torso, with interpolated regions generated from interpolation between the source image and the target image.

[0005] In some embodiments, the method further includes, in response to determining to use the face editor, adjusting at least a portion of the target facial features in the target image based on facial pixels from the source facial features in the source image. In some embodiments, adjusting at least a portion of the target facial features in the target image based on facial pixels from the source facial features in the source image includes extracting the target head in an initial pose and the source head, aligning the target head to a standard pose, encoding the aligned target head as a target vector and the source head as a source vector in latent space, copying one or more components from the source vector to the target vector, rendering a modified target vector including the one or more components from the encoded source head, re-aligning the rendered target head to the initial pose, and blending the re-aligned target head with the source image.

[0006] In some embodiments, the determination to use the face editor is based on an angular difference between a first angle of the target head and a second angle of the source head. In some embodiments, the determination to use the head editor is based on a bounding box surrounding the target head or target face and a distance between the bounding box and a bounding box associated with one or more other objects in the target image. In some embodiments, generating the composite image further includes inpainting remaining target pixels in the target image associated with the target head and not associated with the source head in response to identifying remaining target pixels in the target image associated with the target head and not associated with the source head. In some embodiments, the method further includes determining target head occlusion or source head occlusion based on determining a difference in color histograms between the target image and the source image, and the determination to use the head editor is based on the target head occlusion or source head occlusion.

[0007] In some embodiments, the method further includes capturing a set of images with a camera before determining whether to use the one or more editors based on the set of images, and providing a user interface to a user including a target image and an option to select a source head from the set of source images, the set of source images including a source image, and the method further includes receiving a selection of the source image from the user. In some embodiments, at least one subject in the source image is a person or an animal.

[0008] The system includes one or more processors and one or more computer-readable media having instructions stored on the one or more computer-readable media and, when executed by the one or more processors, causing the one or more processors to perform operations including receiving a set of images including source images and target images, the source images and target images including at least a subject, determining, based on the set of images, whether to use one or more editors selected from the group of a head editor, a face editor, or a combination thereof, and, in response to determining to use the head editor, generating a composite image by replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from the source head of the subject in the source image, and replacing neck pixels associated with the target neck and shoulder pixels associated with the target shoulders, including a region between the target head and the target torso, with interpolated regions generated from interpolation between the source image and the target image.

[0009] In some embodiments, the operations further include, in response to determining to use the face editor, adjusting at least a portion of the target facial features in the target image based on facial pixels from the source facial features in the source image. In some embodiments, adjusting at least a portion of the target facial features in the target image based on facial pixels from the source facial features in the source image includes extracting the target head in an initial pose and the source head, aligning the target head to a standard pose, encoding the aligned target head as a target vector and the source head as a source vector in latent space, copying one or more components from the source vector to the target vector, rendering a modified target vector including one or more components from the encoded source head, realigning the rendered target head to the initial pose, and blending the realigned target head with the source image. In some embodiments, the determination to use the face editor is based on an angular difference between a first angle of the target head and a second angle of the source head. In some embodiments, the determination to use the head editor is based on a bounding box surrounding the target head or face and a distance between the bounding box and bounding boxes associated with one or more other subjects in the target image. In some embodiments, generating the composite image further includes inpainting remaining target pixels in the target image that are associated with the target head and that are not associated with the source head, in response to identifying remaining target pixels in the target image that are associated with the target head and that are not associated with the source head.

[0010] A non-transitory computer-readable medium having instructions stored thereon, the instructions, in response to execution by one or more processing devices, causing the one or more processing devices to perform operations, including receiving a set of images including source images and target images, the source images and target images including at least a subject, the operations further including determining, based on the set of images, whether to use one or more editors selected from the group of a head editor, a face editor, or a combination thereof, and, in response to determining to use the head editor, generating a composite image by replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from a source head of the subject in the source image, and replacing neck pixels associated with the target neck and shoulder pixels associated with the target shoulders, including a region between the target head and the target torso, with interpolated regions generated from interpolation between the source image and the target image.

[0011] In some embodiments, the operations further include, in response to determining to use the face editor, adjusting at least a portion of the target facial features in the target image based on facial pixels from the source facial features in the source image. In some embodiments, adjusting at least a portion of the target facial features in the target image based on facial pixels from the source facial features in the source image includes extracting the target head in an initial pose and the source head, aligning the target head to a standard pose, encoding the aligned target head as a target vector and the source head as a source vector in latent space, copying one or more components from the source vector to the target vector, rendering a modified target vector including the one or more components from the encoded source head, realigning the rendered target head to the initial pose, and blending the realigned target head with the source image. In some embodiments, determining to use the face editor is based on an angle difference between a first angle of the target head and a second angle of the source head. In some embodiments, the decision to use the head editor is based on a bounding box surrounding the target head or face and the distance between this bounding box and bounding boxes associated with one or more other objects in the target image. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram of an exemplary network environment, according to some embodiments described herein. [Figure 2] FIG. 1 is a block diagram of an exemplary computing device according to some embodiments described herein. [Figure 3A] 1 is an exemplary user interface including options for a user to specify a target image from a set of images, according to some embodiments described herein. [Figure 3B]1 is an exemplary user interface including options for a user to specify one or more source images from a set of images, according to some embodiments described herein. [Figure 4A] 1 is an exemplary target image for a head editor, according to certain embodiments described herein. [Figure 4B] 1 is an exemplary composite image generated by a head editor according to certain embodiments described herein. [Figure 5] 1 is an exemplary target image according to some embodiments described herein. [Figure 6A] 1 is an exemplary first target head separated from a target image according to certain embodiments described herein. [Figure 6B] 10 is an exemplary second target head separated from the target image according to certain embodiments described herein. [Figure 7A] 1 is an exemplary first source head separated from a source image, according to certain embodiments described herein. [Figure 7B] 10 is an exemplary second source head separated from a source image, according to certain embodiments described herein. [Figure 8A] 1 is an exemplary first source head aligned with a first target head, according to certain embodiments described herein. [Figure 8B] 10 is an exemplary second source head aligned with a second target head, according to certain embodiments described herein. [Figure 9] 1 is an exemplary composite image analyzed for inpainting, according to certain embodiments described herein. [Figure 10] 6 is an exemplary composite image in which the target head of FIG. 5 is replaced with a source head, according to certain embodiments described herein. [Figure 11]FIG. 1 is an example block diagram of a machine learning model for generating synthetic images, according to some embodiments described herein. [Figure 12A] 1 is an exemplary target image for a face editor according to some embodiments described herein. [Figure 12B] 1 is an exemplary composite image generated by a face editor according to some embodiments described herein. [Figure 13] 10 illustrates an exemplary user interface for selecting a best take, according to some embodiments described herein. [Figure 14] 1 shows a flowchart of an example method for generating a composite image, according to some embodiments described herein. DETAILED DESCRIPTION OF THE INVENTION

[0013] overview A media application generates a composite image, and one or more of the subjects in the composite image have heads and / or faces from the source images. Previous attempts to combine portions of the images can result in unrealistic composite images, such as visible seams, pixels from the original object visible where the replacement object is not aligned with the original object, and artifacts caused by occluding objects.

[0014] In some embodiments, a media application receives a source image and a target image and determines whether to use a head editor and / or a face editor to generate a composite image. For example, the media application may select a head editor based on the heads of two subjects in the target image being sufficiently separated or not occluded by an object. In another example, the media application may select a face editor based on the difference in pose of the face angle between the source image and the target image being close enough to allow a portion of the source image to be added to the target image.

[0015] The head editor can replace a head in the target image with a head in the source image. The face editor can adjust facial features, such as the eyes and mouth, from the target image based on facial features from the source image. For example, the face editor may use embeddings to calculate pixel values ​​for facial features in the target image. The media application generates a composite image from a combination of the target image and the source image.

[0016] Generating the composite image may include analyzing the image data to determine occlusion and transforming the image data, and may perform inpainting to avoid visible seams or other imperfections that may cause the composite image to appear unrealistic. In other words, the techniques described herein for generating a composite image from one or more source images may seamlessly maintain the realism of the one or more source images.

[0017] Example Environment FIG. 1 shows a block diagram of an exemplary environment 100. In some embodiments, environment 100 includes a media server 101, user device 115a, and user device 115n coupled to a network 105. Users 125a, 125n may be associated with respective user devices 115a, 115n. In some embodiments, environment 100 may include other servers or devices not shown in FIG. 1. In FIG. 1 and the remaining figures, a letter following a reference number, e.g., "115a," denotes a reference to the element with that specific reference number. A reference number in the text without a following letter, e.g., "115," denotes a general reference to an embodiment of the element bearing that reference number.

[0018] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively coupled to a network 105 via signal line 102. The signal line 102 may be a wired connection, such as Ethernet, coaxial cable, or fiber optic cable, or a wireless connection, such as Wi-Fi, Bluetooth, or other wireless technology. In some embodiments, the media server 101 transmits data to and receives data from one or more of the user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.

[0019] The database 199 may store machine learning models, training data sets, images, etc. The database 199 may also store social network data associated with the user 125, the user preferences of the user 125, etc.

[0020] The user device 115 may be a computing device that includes a memory coupled to a hardware processor. For example, the user device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or any other electronic device that can access the network 105.

[0021] In the illustrated implementation, user device 115a is coupled to network 105 via signal line 108, and user device 115n is coupled to network 105 via signal line 110. Media application 103 may be stored as media application 103b on user device 115a and / or as media application 103c on user device 115n. Signal lines 108 and 110 may be wired connections, such as Ethernet, coaxial cable, or fiber optic cable, or wireless connections, such as Wi-Fi, Bluetooth, or other wireless technologies. User devices 115a and 115n are accessed by users 125a and 125n, respectively. User devices 115a and 115n in FIG. 1 are used as an example. While FIG. 1 shows two user devices, 115a and 115n, this disclosure applies to system architectures having one or more user devices 115.

[0022] The media application 103 may be stored on the media server 101 or the user device 115. In some embodiments, the operations described herein are executed on the media server 101 or the user device 115. In some embodiments, some operations may be executed on the media server 101 and some may be executed on the user device 115. Execution of operations is subject to user settings. For example, the user 125a may specify that operations be executed on each device 115a and not on the media server 101. Such settings result in operations described herein being executed entirely on the user device 115a and not on the media server 101. Furthermore, the user 125a may specify that user images and / or other data be stored only locally on the user device 115a and not on the media server 101. Such settings result in user data not being transmitted to or stored on the media server 101. The transmission of user data to the media server 101, any temporary or permanent storage of such data by the media server 101, and the performance of operations on such data by the media server 101 are performed only if the user consents to the transmission, storage, and performance of operations by the media server 101. The user is provided with the option to change settings at any time, for example, so that the user can enable or disable use of the media server 101.

[0023] Machine learning models (e.g., neural networks or other types of models) are stored and used locally on the user device 115 with specific user permission when utilized for one or more operations. Server-side models are used only with user permission. Additionally, trained models may be provided for use on the user device 115. During such use, on-device training of the model may be performed if permitted by the user 125. Updated model parameters may be sent to the media server 101, for example, to enable federated learning, if permitted by the user 125. The model parameters do not include any user data.

[0024] The media application 103 receives a set of images including a source image and a target image, the source image and the target image including at least a subject. The media application 103 determines whether to use a head editor and / or a face editor based on the set of images. In response to determining to use the head editor, the media application 103 generates a composite image by replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from the source head of the subject in the source image, and replacing neck pixels associated with the target neck and shoulder pixels associated with the target shoulders, including a region between the target head and the target torso, with interpolated regions generated from interpolation between the source image and the target image.

[0025] In some embodiments, the media application 103 may be executed using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, the media application 103a may be executed using a combination of hardware and software.

[0026] Exemplary Computing Device 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. Computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, computing device 200 is a media server 101 used to run media application 103a. In another example, computing device 200 is a user device 115.

[0027] In some embodiments, computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, a display 241, a camera 243, and a storage device 245, all coupled via bus 218. Processor 235 may be coupled to bus 218 via signal line 222, memory 237 may be coupled to bus 218 via signal line 224, I / O interface 239 may be coupled to bus 218 via signal line 226, display 241 may be coupled to bus 218 via signal line 228, camera 243 may be coupled to bus 218 via signal line 230, and storage device 245 may be coupled to bus 218 via signal line 232.

[0028] Processor 235 may be one or more processors and / or processing circuits that execute program code and control basic operations of computing device 200. A “processor” includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) with one or more cores (e.g., a single-core, dual-core, or multi-core configuration), a system having multiple processing units (e.g., a multiprocessor configuration), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuitry for implementing functionality, a dedicated processor for performing processing based on neural network models, neural circuitry, a processor optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, processor 235 may include one or more coprocessors that perform neural network processing. In some embodiments, processor 235 may be a processor that processes data to generate a probabilistic output; for example, the output generated by processor 235 may be inaccurate or accurate within a range from an expected output. Processing need not be limited to a particular geographic location or have time limitations. For example, a processor may perform its functions in real time, offline, in batch mode, etc. Portions of processing may be performed at different times and in different locations by different (or the same) processing systems. A computer may be any processor in communication with a memory.

[0029] Memory 237 is typically provided in computing device 200 for access by processor 235 and may be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc., located separately from and / or integrated with processor 235, suitable for storing instructions for execution by the processor or set of processors. Memory 237 may store software operated on computing device 200 by processor 235, including media application 103.

[0030] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, etc. One or more methods disclosed herein may operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application (“app”) running on a mobile computing device, etc.

[0031] Application data 266 may be data generated by other applications 264 or by hardware of computing device 200. For example, application data 266 may include images used by an image library application, user actions identified by other applications 264 (e.g., social networking applications), etc.

[0032] I / O interface 239 may provide functionality that allows computing device 200 to interface with other systems and devices. The interfaced devices may be included as part of computing device 200 or may be separate and in communication with computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 245), and input / output devices may communicate through I / O interface 239. In some embodiments, I / O interface 239 may connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, monitor, etc.).

[0033] Some examples of interface devices that can be connected to I / O interface 239 can include display 241, which can be used to display content, e.g., images, video, and / or user interfaces of output applications described herein, and to receive touch (or gesture) input from a user. For example, display 241 can be utilized to display a user interface, including a graphical guide, on a viewfinder. Display 241 can include any suitable display device, such as a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device. For example, display 241 can be a flat display screen provided on a mobile device, multiple display screens embedded in a glasses form factor or headset device, or a monitor screen of a computing device.

[0034] Camera 243 may be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video that I / O interface 239 sends to media application 103.

[0035] The storage device 245 stores data related to the media application 103. For example, the storage device 245 can store training data sets including labeled images, machine learning models, output from machine learning models, etc.

[0036] FIG. 2 illustrates an exemplary media application 103 stored in memory 237, including an image module 202, a head editor 204, and a face editor 206.

[0037] The image module 202 generates graphical data for displaying a user interface including a set of images. The set of images may be received from the camera 243 of the computing device 200 and / or from the media server 101 via the I / O interface 239. For example, the set of images may include images from a burst of images captured by the camera 243. The burst of images may include multiple photographs captured in quick succession over a short period of time. The set of images includes one or more source images and a target image, each containing one or more objects. The one or more objects may be people, animals, etc.

[0038] The image module 202 obtains permission from the user to modify any image in the set of images. Controls may be provided to the user that allow the user to choose both whether and when the systems, programs, or features described herein may enable the collection or use of user information (e.g., the user's identification in the image, the user's preferences, or the user's current location), as well as whether the user is sent content or communications from the server. Additionally, certain data may be processed in one or more ways to remove personally identifiable information before being stored or used. For example, the user's identification information may be processed so that the user's personally identifiable information cannot be determined, or, if location information is obtained (e.g., to the city, zip code, or state level), the user's geographic location may be generalized so that the user's specific location cannot be determined. Thus, the user may control what information is collected about them, how that information is used, and what information is provided to them.

[0039] In some embodiments, the image module 202 selects one or more source images and a target image to be used to generate a composite image. For example, the image module 202 may automatically generate a composite score for each image in the set of images based on composite factors such as the type of object in the image, the position of the subject, and lighting. For example, the target image may be selected based on the best landscape composition and well-positioned subjects in the group. In another example, the target image may be selected based on the highest number of subjects looking at the camera, smiling, with their eyes open, etc. The image with the highest overall composite score may be used for the target image. In some embodiments, a user may select a particular image in the set of images as the target image, and one or more other images in the set of images may be presented as source images.

[0040] The image module 202 can select one or more source images based on the objects in the images. For example, the image module 202 can generate a face score for each object in the image based on a quality metric. For example, each face can be scored based on whether the face is characterized by a smile, open eyes, whether it is fully or partially occluded in the photo by another person or object, and / or whether it is motionless (e.g., not blurred). The image module 202 can generate a head score based on the position of the object's head, e.g., generate a higher head score for a head that is in a vertical plane (e.g., facing the camera) rather than at another angle (e.g., tilted, rotated away from the camera, etc.). Such scoring can be performed using image detection and recognition techniques, such as a trained machine learning model that can detect faces in photos and apply quality scoring criteria.

[0041] Based on the scores, the image module 202 may rank the faces in the images. For example, the image module 202 may determine that multiple particular faces in different images of the set of images belong to the same subject (person, animal, etc.), and each such subject may be associated with a ranked list of faces from the image. The image module 202 may automatically select the target image and one or more source images, or the image module 202 may suggest the top-ranked images as suggestions to the user for the user to confirm that a particular image has been selected as the target image.

[0042] In some embodiments, the image module 202 generates graphical data for displaying a user interface that provides a user with the option to select a target image and one or more source images from a set of images, which may be a burst of captured images, images captured during a particular time period (e.g., the last 24 hours, images captured at a particular location, etc.).

[0043] 3A is an exemplary user interface 300 that includes options for a user to specify a target image from a set of images. In this example, the user selects an image 305 to be used as the target image. In some embodiments, a particular image from the set of images may be highlighted (e.g., with a border, an icon, etc.) and suggested as the target image. The user may select the suggested image or any other image as the target image.

[0044] 3B is an exemplary user interface 325 that includes options for a user to specify one or more source images from a set of images. In this example, the user selects a first image 330 to be used as a source image for a first object and a second image 335 to be used as a source image for a second object. In some embodiments, the image module 202 can generate a user interface that includes options for selecting a particular object within an image. For example, the computing device 200 can receive user input in the form of a double-click on the object, a circle drawn around the object's head, etc.

[0045] In some embodiments, a single image may be selected as a source image for two or more subjects. In some embodiments, each subject may be associated with a different source image. In some embodiments, source image suggestions may be provided for one or more subjects. For example, if a target image has a first subject tilting their head and a second subject with their eyes closed, a first source image in which the first subject faces the camera without tilting their head and a second source image in which the second subject has their eyes open may be suggested as respective source images. In some embodiments, other factors such as lighting, the duration between capture of the target image and a particular source image, and the distance between the subject's position in the source image and the target image may be used in selecting a recommended particular source image.

[0046] In some embodiments, the user interface may include an option to search for other source images. For example, the user interface may include an option for the user to scroll (or otherwise browse) images in a camera roll and select a source image. In some embodiments, because a source image from a camera roll may have been captured with different lighting conditions, different shadows, etc., the image module 202 may modify the source image to have colors (and / or other image attributes, such as brightness, white balance, contrast, etc.) that match corresponding attributes of the target image.

[0047] The image module 202 determines whether to use the head editor 204, the face editor 206, or both to generate the composite image. The image module 202 can make its decision using different criteria, such as the proximity of the head in the target or source image, the angle of the head in the target or source image, and the occlusion of the head in the target or source image.

[0048] In some embodiments, the image module 202 decides whether to use the head editor 204 or the face editor 206 based on different factors evaluated by different classifiers. In some embodiments, the image module 202 performs a weighted sum or logistic regression of different factors so that the decision is based on an ensemble of factors rather than a single factor. In some embodiments, some factors may be determining orientation, such as when one of the heads is 70% occluded by an object.

[0049] In some embodiments, the classifier is trained from evaluations of head edits from head editor 204 and face edits from face editor 206 in a training image set. For example, the training data may include target images, source images, composite images generated by head editor 204 and / or face editor 206, and corresponding evaluations of the composite images, which may be provided by a human or a quality algorithm. As a result of receiving the evaluations, the image module 202 may recalculate the weights of the classifier to improve the quality of the composite images generated by head editor 204 and face editor 206.

[0050] In some embodiments, the image module 202 generates a bounding box around each target head / face in the target image and determines whether to use the head editor 204 to replace the target head with a source head based on the distance between the bounding boxes in the target image. For example, if there is overlap between the bounding boxes, this indicates that the subject heads are close to each other, and the head editor 204 may not be used due to variations in head position when multiple subjects are close together, which could result in large sections of background being inpainted in the composite image, potentially overlapping the subjects, and complicating the alignment of the subjects in the target image and one or more source images. In some embodiments, the image module 202 determines whether to use the head editor 204 based on a continuous function. In some embodiments, the continuous function is based on a weighted distance between the bounding boxes and / or a weighted overlap percentage between the bounding boxes, and the weight values ​​may be learned during training of the classifier used by the image module 202.

[0051] In some embodiments, the image module 202 determines whether to use the head editor 204 based on the angle difference between the first angle of the target head and the second angle of the source head. The face editor 206 may generate an unsatisfactory or low-quality composite image if the head angle of the source image is greater than a threshold difference from the head angle of the target image.

[0052] The image module 202 may determine to use the face editor 206 based on applying a non-linear function to the difference between the first angle and the second angle. In some embodiments, the image module 202 applies a cosine to the first angle and a cosine to the second angle to determine the difference between the first angle and the second angle. In some embodiments, the image module 202 uses a threshold angle difference to determine whether to use the face editor 206, such that if the angle difference exceeds the threshold angle difference, the image module 202 determines to use the head editor 204 rather than the face editor 206.

[0053] In some embodiments, the image module 202 determines whether to use the head editor 204 based on the occlusion of the target head or the occlusion of the source head. When the image module 202 receives either an occluded target image or a source image, the resulting composite image has a higher failure rate and / or is of lower quality if the source image is occluded. In some embodiments, the image module 202 can use a color histogram of the image to determine whether the source head or the target head is occluded. For example, the image module 202 can use the average value of a particular channel in the histogram to identify occlusion.

[0054] In some embodiments, the image module 202 determines whether to use the head editor 204, the face editor 206, or both the head editor 204 and the face editor 206 based on the distance from the subject to the image boundary. For example, if the head is close to the image boundary and the head editor 204 changes the head pose near the image boundary, it may result in part of the head being cropped at the image boundary (and thus providing a less than satisfactory composite image because the subject's head is not completely within the image).

[0055] The head editor 204 replaces the target head with the source head. In some embodiments, the head editor 204 replaces target pixels associated with the target head in the target image with source pixels from the source head in the source image. The head editor 204 can generate a head mask including the subject's hair, segment the head mask from the source image, and apply the pixels in the head mask to the target image. The head editor 204 can modify the position and scale of the source head to match the dimensions of the target head.

[0056] The head editor 204 performs inpainting when replacing the target head with the source head results in portions of the target head not overlapping with the source head. This can occur when the target head and the source head are associated at different angles. This can also occur when the angles of the target head and the source head in the target image and the source image are different, and the head editor 204 aligns the angle of the target head with the angle of the source head before replacement. Aligning the target head with the source head can identify portions of the target head where the source head does not overlap. The head editor 204 can perform inpainting of the remaining target pixels by identifying remaining pixels in the target image that are associated with the target head but not with the source head and replacing them with the remaining target pixels. Inpainting pixels can include determining the distance between the source pixel and the remaining pixel and generating a replacement pixel based on the similarity and distance to the source pixel.

[0057] Due to different angles between the target head and the source head, the position of the subject's neck region may be different in the two images. The head editor 204 renders a smooth transition between the target torso and the source head by generating interpolated regions for the neck and shoulder regions. In some embodiments, the head editor 204 generates interpolated regions that include regions between the target head and the target torso (e.g., regions described as the target neck and target shoulders) and replaces target pixels associated with the target neck and target shoulders with interpolated regions that are interpolations of the target and source pixels of the corresponding neck and shoulder regions. In some embodiments, the head editor 204 uses multiple image frames, such as a set of image frames generated from a burst of images captured by the camera 243, and generates the interpolated regions from the multiple image frames.

[0058] 4A, an exemplary target image 400 for a head editor is shown, according to some embodiments described herein. The subject 405 on the right is looking up and has smooth shoulders.

[0059] 4B is an example composite image 425 generated by head editor 204. Head editor 204 replaced the target head with the source head in subject 430 on the right. In composite image 425, the interpolated region 430 over the neck and upper shoulders is the result of the neck being in a different orientation and the shoulders containing more wrinkled fabric compared to the shoulders of subject 405 in target image 400.

[0060] In some embodiments, the head editor 204 includes a machine learning model that receives a target image and one or more source images as input and outputs a synthetic image. The trained machine learning model may include one or more model types or structures. For example, the model types or structures may include any type of neural network, such as a linear network, a deep learning neural network running multiple layers (e.g., a linear network with "hidden layers" between the input and output layers), a convolutional neural network (e.g., a network that divides or partitions input data into multiple portions or tiles, processes each tile separately using one or more neural network layers, and aggregates the results of processing each tile), a sequence-to-sequence neural network (e.g., a network that receives sequential data as input, such as words in a sentence or frames of a video, and outputs a sequence of results), etc.

[0061] The model format or structure may specify the connectivity between various nodes and their organization into layers. For example, nodes in a first layer (e.g., an input layer) may receive data as input data or application data. Such data may include, for example, one or more pixels per node, for example, when the trained model is used to analyze, for example, a target image and one or more source images. Subsequent intermediate layers may receive as input the output of nodes in the previous layer according to the connectivity specified in the model format or structure. These layers may also be referred to as hidden layers. The final layer (e.g., an output layer) generates the output of the machine learning model. For example, the output layer may output a synthetic image. In some embodiments, the model format or structure also specifies the number and / or type of nodes in each layer.

[0062] In another embodiment, the trained model may include one or more models. One or more of the models may include multiple nodes arranged in layers according to a model structure or format. In some embodiments, a node may be a computational node without memory configured, for example, to process a unit of input and generate a unit of output. The computation performed by the node may include, for example, multiplying each of multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to generate the node output. In some embodiments, the computation performed by the node may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computation may include operations such as matrix multiplication. In some embodiments, computation by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, individual processing units of a graphics processing unit (GPU), or dedicated neural circuitry. In some embodiments, a node may include memory, e.g., capable of storing and using one or more previous inputs when processing a subsequent input. For example, a node with memory may include a long short-term memory (LSTM) node. LSTM nodes may use memory to maintain "state," which allows the node to behave like a finite state machine (FSM).

[0063] In some embodiments, the trained model may include embeddings or weights for individual nodes. For example, the model may begin as multiple nodes organized into layers, as specified by the model format or model structure. At initialization, a respective weight may be applied to the connection between each pair of nodes connected according to the model format, e.g., nodes in successive layers of a neural network. For example, each weight may be randomly assigned or initialized to a default value. The model may then be trained, e.g., using training data, to produce results.

[0064] The training may include applying a supervised learning method. In supervised learning, the training data may include multiple inputs (e.g., target and source images, segmentation masks, etc.) and corresponding ground truth outputs for each input (e.g., ground truth masks that accurately identify portions of a subject, such as the subject's face, in each image, synthetic images, etc.). Based on a comparison of the model's outputs and the ground truth outputs, the values ​​of the weights are automatically adjusted, e.g., in a manner that increases the probability that the model will produce the ground truth output for the synthetic image.

[0065] In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In some embodiments, the trained model may include a fixed set of weights, for example, downloaded from a server that provides the weights. In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In embodiments where data is omitted, the head editor 204 may generate a trained model based on prior training, for example, by a developer of the head editor 204, by a third party, etc. In some embodiments, the trained model may include a fixed set of weights, for example, downloaded from a server that provides the weights.

[0066] In some embodiments, the trained machine learning model receives a target image and one or more source images having one or more objects. The machine learning model can generate one or more segmentation masks that identify pixels in the one or more source images that correspond to one or more heads, including hair, in the one or more source images. For each object, the machine learning model replaces head pixels from the target image with head pixels from the one or more source images that have corresponding positions, scales, and angles that match the source images. The machine learning model blends the head pixels along the edges of the head. In some embodiments, background pixels exposed as a result of differences between the source and target heads are inpainted. In some embodiments, the machine learning model blends interpolated regions into the target image that correspond to the neck and shoulders of the object. In some embodiments, the trained machine learning model outputs a composite image that incorporates these changes.

[0067] In some embodiments, the machine learning model outputs a confidence value for each synthetic image output by the trained machine learning model. The confidence value may be expressed as a percentage, a number between 0 and 1, etc. For example, the machine learning model outputs a confidence value of 85% for the confidence that the synthetic image correctly replaces the target head with the source head and does not contain pixels from other people or objects. In some embodiments, if the confidence value exceeds a confidence threshold, the synthetic image is provided to the user. In some embodiments, the confidence value is pre-calculated, and the synthetic image is not generated unless the confidence value exceeds a threshold confidence value.

[0068] In some embodiments, head editor 204 includes multiple machine learning models that perform different functions in the steps used to generate the composite image. For example, a first machine learning model can replace target pixels associated with a target head with source pixels from a source head, a second machine learning model aligns the source head with the target head, and a third machine learning model generates interpolated regions that replace target pixels associated with the target neck and target shoulders.

[0069] 5 is an exemplary target image 500. The boy's head 505 and the girl's head 510 in the target image 500 are replaced by heads from the source image.

[0070] Figure 6A is an exemplary first target head 600 isolated from the target image 500 of Figure 5. Figure 6B is an exemplary second target head 650 isolated from the target image 500 of Figure 5. The first target head 600 can be isolated and selected for replacement because the boy's nose is pointing up and his mouth is open. The second target head 650 can be isolated and selected for replacement because the girl is not looking directly at the camera.

[0071] Figure 7A is an example first source head 700 separated from a source image. Figure 7B is an example second source head 750 separated from the same source image containing the first source head 700 or from a different source image. The head editor 204 replaces at least a portion of the target pixels associated with the first target head 600 and at least a portion of the target pixels associated with the second target head 650 with source pixels from the first source head 700 and the second source head 750, respectively.

[0072] Figure 8A is an example of a first source head aligned with a first target head 800. Figure 8B is an example of a second source head aligned with a second target head 850. Although the images are aligned, there are some places where the images do not overlap, such as the location of the boy's cap 805.

[0073] 9 is an example composite image 900 that is analyzed for inpainting based on the alignment of a first source head that exposes areas of the background previously covered by a first target head. After the head editor 204 performs the inpainting of the target head, the composite image appears as if the source head fits into the composite image without seams or other imperfections.

[0074] Figure 10 is an example composite image 1000 in which the source heads have been replaced with the target heads. The boy's head 505 in Figure 5 has been replaced with the boy's head 1005 in Figure 10, and the girl's head 510 in Figure 5 has been replaced with the girl's head 1010 in Figure 10.

[0075] The face editor 206 transfers facial features (e.g., smile, open eyes, mouth shape, etc.) from the source image to the target image. The face editor 206 does not change the pose of the target face. In some embodiments, the face editor 206 adjusts at least a portion of the target pixel values ​​associated with the target facial feature in the target image based on source pixels from the source facial feature in the source image.

[0076] In some embodiments, face editor 206 extracts target and source faces of the same subject using a face matching algorithm run with specific user permission. Face editor 206 aligns the target face from an initial pose to a standard pose, where the standard pose is one of the poses that the machine learning model is trained to use.

[0077] In some embodiments, face editor 206 uses a machine learning model selected from one of the model types described above with respect to head editor 204. For example, the machine learning model may include an encoder and a convolutional neural network (CNN).

[0078] In some embodiments, not all of the target face is replaced with the source because some facial components may be sensitive to small changes in angle and pose. For example, adding the nose of an upturned head to a front-facing face may make the person's nose look unnatural. As a result, the machine learning model's encoder encodes the source head components as a source vector using a facial feature mask (e.g., an eye / mouth mask). The encoder also encodes the target head as a target vector. The machine learning model copies corresponding components from the source vector to the target vector.

[0079] In some embodiments, the CNN includes multiple layers that each generate a version of the synthetic image by rendering a target vector containing source vector components. The CNN generates the synthetic image at increasing levels of resolution with each layer. The high-resolution output is blended with existing lower-resolution layers until the final synthetic image is output. In some embodiments, the machine learning model realigns the rendered target head back to its initial pose and blends the realigned target head with the source image, which is output as the synthetic image.

[0080] 11 is an example block diagram 1100 of an on-device machine learning model for generating a synthetic image. A source image 1105 and a target image 1110 are received by both an offline module 1115 and an on-device module 1120, such as a media application 103 stored on a user device 115 shown in FIG. 1. The on-device module 1120 outputs a predicted image 1130 in which facial expressions and / or head poses from the source image 1105 are fused into the target image 1110. The on-device module 1120 is implemented such that it can run on a user device, e.g., a smartphone with or without machine learning circuitry, within latency and memory usage budgets.

[0081] The offline module 1115 can provide a supervision for on-device model training, which advantageously enables a lightweight on-device module 1120 to generate a composite image 1125. The composite image 1125 can be a higher quality version of the predicted image 1130 because the offline module 1115 has a larger budget for latency and memory usage.

[0082] In some embodiments, the on-device module 1120 may use warping and restoration, in which a dense face mesh warps the source image 1105 to a target position, and may use a neural network (or other suitable technique) to remove artifacts. Given two face images and their face meshes, the on-device module 1120 aligns the source face pose to the target face mesh and then warps the source face to the target image. The warped face has the source expression but has artifacts caused by the alignment and warping, which may be considered degradations. The on-device module 1120 may include a supervised restoration model from the offline module 1115 that can remove or repair degradations.

[0083] In some embodiments, the on-device module 1120 can modify the face to reflect a specified facial expression. The on-device module 1120 may include two encoders: a source encoder that extracts expression codes from a source image, and a target encoder that extracts head pose and appearance information from a target image. The on-device module 1120 can insert expression codes at one or more levels of the decoder (e.g., at any level). A skim connection from the target encoder can also be added to ensure that the on-device module 1120 edits the face and not other parts of the image.

[0084] FIG. 12A is an exemplary target image 1200 in which a second subject 1205 has an expression that is replaced by the face editor 206.

[0085] 12B is an exemplary composite image 1250 generated by face editor 206 according to the techniques described above. As can be seen, the open-mouthed smile of subject 1205 on the right in target image 1200 is replaced with the closed-mouthed smile of subject 1255 in composite image 1250.

[0086] 13 shows exemplary user interfaces 1300, 1350, 1375 for selecting a best take, according to some embodiments described herein. The first user interface 1300 includes a target image 1305 of two subjects 1306, 1307. A user can select a "Best Take" button 1310 to begin the process for obtaining a composite image.

[0087] The second user interface 1350 includes a target image 1355 and icons 1357 and 1359 for a first and second subject, respectively. The user can select one of the icons 1357 and 1359 to select a different face of the subject. In this example, the user has selected the second subject icon 1359. The user interface 1350 also includes three options 1361, 1363, and 1365 for the second subject. The first option 1361 and the third option 1365 are from the source image, and the second option 1363 is from the target image 1355. Once the user selects one of the three options 1361, 1363, and 1365, the user can select a done button 1367 to view a composite image of the selected option on the target image 1355, or may select a reset button 1369 to restart the process. In this example, the user selects the first option 1361 and the done button 1367.

[0088] The third user interface 1375 includes a composite image 1377 generated from the target image 1355 and the source face 1379 of the second user interface 1350. The user may press a save copy button 1385 to save a copy of the image.

[0089] Exemplary Flowchart 14 shows a flowchart of an example method 1400 for segmenting an image according to some embodiments described herein. Method 1400 may be performed by computing device 200 of FIG. 2. In some embodiments, method 1400 is performed by user device 115, media server 101, or partially on user device 115 and partially on media server 101.

[0090] 14 may begin at block 1402. At block 1402, a set of images is received, including a source image and a target image, where the source image and the target image include one or more objects. In some embodiments, before block 1402, method 1400 further includes capturing the set of images with a camera and providing a user interface including the target image and an option to select a source image from the set of source images, where the set of source images includes the source image, and method 1400 further includes receiving a selection of the source image from the user. The object may include a person or an animal. Block 1402 may be followed by block 1404.

[0091] At block 1404, it is determined whether permission to modify the source and target images has been granted. If permission has not been granted, block 1404 may be followed by block 1406, where method 1400 ends. If permission has been granted, block 1404 may be followed by block 1408.

[0092] At block 1408, it is determined whether to use one or more editors selected from the group of a head editor, a face editor, or a combination thereof. The decision to use the face editor may be based on an angle difference between a first angle of the target head and a second angle of the source head, a bounding box surrounding the target head or target face and a distance between the bounding box and bounding boxes associated with other objects in the target image, or target head occlusion or source head occlusion. The decision to use the target head occlusion or source head occlusion may be based on determining a difference in color histograms.

[0093] If a head editor is selected, block 1408 may be followed by block 1410. In block 1410, the composite image is generated by replacing at least a portion of head pixels associated with the target head of the subject in the target image with head pixels from the source head of the subject in the source image, and replacing neck pixels associated with the target neck and shoulder pixels associated with the target shoulders, including the region between the target head and the target torso, with interpolated regions generated from interpolation between the source and target images. In some embodiments, the head editor aligns the angle of the target head with the angle of the source head before replacing at least a portion of the target pixels associated with the target head in the target image. If a face editor is selected, block 1408 may be followed by block 1412.

[0094] In block 1412, the composite image is generated by adjusting at least a portion of the target facial features in the target image based on facial pixels from the source facial features in the source image. In some embodiments, adjusting at least a portion of the target facial features in the target image based on facial pixels from the source facial features in the source image includes extracting the target head in an initial pose and the source head, aligning the target head to a standard pose, encoding the aligned target head as a target vector and the source head as a source vector in latent space, copying one or more components from the source vector to the target vector, rendering a modified target vector including the one or more components from the encoded source head, re-aligning the rendered target head to the initial pose, and blending the re-aligned target head with the source image.

[0095] In the preceding description, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of the specification. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In some instances, structures and devices are shown in block diagram form to avoid obscuring the specification. For example, embodiments may be described above primarily with reference to user interfaces and specific hardware. However, embodiments may apply to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.

[0096] A reference herein to "some embodiments" or "some instances" means that a particular feature, structure, or characteristic described in connection with an embodiment or instance may be included in at least one implementation herein. The appearances of the phrase "in some embodiments" in various places in this specification are not necessarily all referring to the same embodiments.

[0097] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. These steps require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It is sometimes convenient, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0098] It should be recognized, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise specified, as will be apparent from the discussion that follows, throughout this specification, discussions utilizing terms including "processing," "calculating," "figuring out," "determining," or "displaying," etc., will be understood to refer to the actions and processes of a computer system or similar electronic computing device that manipulates and converts data represented as physical (electronic) quantities in the computer system's registers and memory into other data that is also represented as physical quantities in the computer system's memory or registers, or other such information storage, transmission, or display device.

[0099]

[0013] Embodiments herein may also relate to a processor for performing one or more steps of the above-described methods. The processor may be a special-purpose processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored on a non-transitory computer-readable storage medium, including, but not limited to, an optical disk, a ROM, a CD-ROM, a magnetic disk, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory including a USB key having non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0100] The specification may take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements, hi some embodiments, the specification is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.

[0101] Furthermore, this specification may take the form of a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For purposes of this description, a computer-usable or computer-readable medium may be any apparatus that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, instruction execution apparatus, or instruction execution device.

[0102] A data processing system suitable for storing or executing program code will include at least one processor coupled directly or indirectly to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage of at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution.

Claims

1. 1. A computer-implemented method comprising: receiving a set of images including a source image and a target image, the source image and the target image including at least a subject, the method further comprising: determining whether to use one or more editors selected from the group of a head editor, a face editor, or a combination thereof based on the set of images; generating a composite image in response to determining to use the head editor, said generating including: replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from a source head of the subject in the source image; replacing neck pixels associated with a target neck and shoulder pixels associated with a target shoulder, including a region between the target head and target torso, with an interpolated region generated from interpolation between the source image and the target image; A method carried out by

2. 10. The method of claim 1, further comprising, in response to determining to use the face editor, adjusting at least a portion of a target facial feature in the target image based on facial pixels from a source facial feature in the source image.

3. adjusting at least a portion of a target facial feature in the target image based on facial pixels from the source facial feature of the source image; Extracting the target head and the source head in an initial pose; aligning the target head to a standard posture; encoding the aligned target head as a target vector and the source head as a source vector in a latent space; copying one or more components from the source vector to the target vector; rendering a modified target vector that includes the one or more components from the encoded source head; realigning the rendered target head to the initial pose; blending the realigned target head with the source image; The method of claim 2 , comprising:

4. The method of claim 2 , wherein the decision to use the face editor is based on an angular difference between a first angle of the target head and a second angle of the source head.

5. 10. The method of claim 1, wherein the decision to use the head editor is based on a bounding box surrounding the target head or face and a distance between the bounding box and bounding boxes associated with one or more other objects in the target image.

6. 2. The method of claim 1, wherein generating the composite image further comprises inpainting remaining target pixels in the target image that are associated with the target head and not associated with the source head in response to identifying the remaining target pixels in the target image that are associated with the target head and not associated with the source head.

7. determining the occlusion of the target head or the occlusion of the source head based on determining a difference in color histograms between the target image and the source image; The method of claim 1 , wherein the decision to use the head editor is based on the occlusion of the target head or the occlusion of the source head.

8. Before determining whether to use the one or more editors based on the set of images, the method may further comprise: capturing said set of images with a camera; providing a user with a user interface including the target image and an option to select the source head from a set of source images, the set of source images including the source image, the method further comprising: The method of claim 1 , comprising receiving a selection of the source image from the user.

9. The method of claim 1 , wherein the at least one object in the source image is a person or an animal.

10. 1. A system comprising: one or more processors; and one or more computer-readable media having instructions stored thereon, the instructions, when executed by the one or more processors, causing the one or more processors to perform operations, the operations including: receiving a set of images including a source image and a target image, the source image and the target image including at least a subject, the operations further comprising: determining whether to use one or more editors selected from the group of a head editor, a face editor, or a combination thereof based on the set of images; generating a composite image in response to determining to use the head editor, said generating including: replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from a source head of the subject in the source image; replacing neck pixels associated with a target neck and shoulder pixels associated with a target shoulder, including a region between the target head and target torso, with an interpolated region generated from interpolation between the source image and the target image; The system is carried out by

11. The operation is 11. The system of claim 10, further comprising, in response to determining to use the face editor, adjusting at least a portion of a target facial feature in the target image based on facial pixels from a source facial feature in the source image.

12. adjusting at least a portion of a target facial feature in the target image based on facial pixels from the source facial feature of the source image; Extracting the target head and the source head in an initial pose; aligning the target head to a standard posture; encoding the aligned target head as a target vector and the source head as a source vector in a latent space; copying one or more components from the source vector to the target vector; rendering a modified target vector that includes the one or more components from the encoded source head; realigning the rendered target head to the initial pose; blending the realigned target head with the source image; The system of claim 11 , comprising:

13. The system of claim 11 , wherein the decision to use the face editor is based on an angular difference between a first angle of the target head and a second angle of the source head.

14. 11. The system of claim 10, wherein the decision to use the head editor is based on a bounding box surrounding the target head or face and a distance between the bounding box and bounding boxes associated with one or more other objects in the target image.

15. 11. The system of claim 10, wherein generating the composite image further comprises inpainting remaining target pixels in the target image that are associated with the target head and not associated with the source head in response to identifying the remaining target pixels in the target image that are associated with the target head and not associated with the source head.

16. A non-transitory computer-readable medium having stored thereon instructions that, in response to execution by one or more processing devices, cause the one or more processing devices to perform operations, the operations including: receiving a set of images including a source image and a target image, the source image and the target image including at least a subject, the operations further comprising: determining whether to use one or more editors selected from the group of a head editor, a face editor, or a combination thereof based on the set of images; generating a composite image in response to determining to use the head editor, said generating including: replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from a source head of the subject in the source image; replacing neck pixels associated with a target neck and shoulder pixels associated with a target shoulder, including a region between the target head and target torso, with an interpolated region generated from interpolation between the source image and the target image; A non-transitory computer-readable medium performed by

17. The operation is 17. The non-transitory computer-readable medium of claim 16, further comprising, in response to determining to use the face editor, adjusting at least a portion of a target facial feature in the target image based on facial pixels from a source facial feature in the source image.

18. adjusting at least a portion of the target facial feature in the target image based on a target facial feature in the target image having facial pixels from the source facial feature of the source image; Extracting the target head and the source head in an initial pose; aligning the target head to a standard posture; encoding the aligned target head as a target vector and the source head as a source vector in a latent space; copying one or more components from the source vector to the target vector; rendering a modified target vector that includes the one or more components from the encoded source head; realigning the rendered target head to the initial pose; blending the realigned target head with the source image; 20. The non-transitory computer-readable medium of claim 17, comprising:

19. 20. The non-transitory computer-readable medium of claim 17, wherein the decision to use the face editor is based on an angular difference between a first angle of the target head and a second angle of the source head.

20. 17. The non-transitory computer-readable medium of claim 16, wherein the decision to use the head editor is based on a bounding box surrounding the target head or face and a distance between the bounding box and bounding boxes associated with one or more other objects in the target image.

Citation Information

Patent Citations

  • Image synthesis system, image synthesis method, and program of the method

    JP2006330800A

  • Image processor and image processing method

    JP2008198062A

  • Image compositing apparatus and program

    JP2010239440A

  • Information processor and recording medium

    JP2014123261A

  • Image synthesizing device, and image synthesizing method and program

    JP2014225870A