Generation of images with improved head posture or facial region

The system addresses the challenge of capturing desirable expressions and alignments in group photos by using head and face editors to align and interpolate features, resulting in high-quality composite images.

JP7852158B2Active Publication Date: 2026-04-27GOOGLE LLC
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GOOGLE LLC
Filing Date
2024-10-03
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Generating high-quality group photos becomes challenging as individuals often have undesirable expressions or head angles, making it difficult to capture everyone smiling and facing the camera simultaneously.

Method used

A method and system that utilize head and face editors to replace and adjust facial features in images, aligning head poses and interpolating neck and shoulder regions to generate composite images with improved realism by inpainting to avoid visible seams.

Benefits of technology

The system effectively generates composite images with improved head and facial alignment, ensuring all individuals appear desirable and aligned, enhancing the quality of group photos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007852158000001
    Figure 0007852158000001
  • Figure 0007852158000002
    Figure 0007852158000002
  • Figure 0007852158000003
    Figure 0007852158000003
Patent Text Reader

Abstract

The media application receives a set of images including a source image and a target image. The source image and the target image include at least a subject. The media application determines whether to use one or more editors selected from the group of a head editor, a face editor, or a combination thereof. In response to determining to use the head editor, the media application generates a composite image by replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from the source head of the subject in the source image, and replacing neck pixels associated with the target neck and shoulder pixels associated with the target shoulders, including a region between the target head and the target torso, with interpolated regions generated from interpolation between the source image and the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application is a non - provisional application claiming priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 63 / 542,283, filed on October 3, 2023, entitled "Generating a Group Photo with Head Pose and Facial Recognition Improvements", the content of which is hereby incorporated by reference in its entirety.

Background Art

[0002] Group photos are a common way to commemorate an event. As the number of people in a photo increases, it becomes more difficult to obtain an image where everyone is smiling and facing the camera because at least one person is likely not to have their most preferred expression. For example, one person may have their mouth open, another may have their eyes closed, and another may not be looking at the camera. Also, a person may be tilting their head in a different direction from the others in the image, or at an oblique angle to the camera, or taking a pose that is otherwise undesirable as a high - quality photo.

[0003] The background art provided herein is for the purpose of generally presenting the context of the present disclosure. Within the scope described in the background art section, the achievements of the inventors named herein, as well as aspects of this document that may not meet the requirements of prior art at the time of filing, are not to be recognized as prior art to the present disclosure, either explicitly or implicitly.

Summary of the Invention

[0004] A method performed by a computer includes receiving a set of images, including a source image and a target image, where the source image and target image include at least a subject. The method further includes determining, based on the set of images, whether to use one or more editors selected from a group of head editors, face editors, or combinations thereof. The method further includes, in response to the decision to use a head editor, generating a composite image, which is done by replacing at least some of the head pixels associated with the target head of the subject in the target image with head pixels from the source head of the subject in the source image, and by replacing the neck pixels associated with the target neck and the shoulder pixels associated with the target shoulder, including the region between the target head and the target torso, with an interpolated region generated from interpolation between the source image and the target image.

[0005] In some embodiments, the method further includes adjusting at least a portion of the target facial features in the target image based on face pixels from the source facial features in the source image, in response to a decision made to use a face editor. In some embodiments, adjusting at least a portion of the target facial features in the target image based on face pixels from the source facial features in the source image includes: extracting a target head in an initial pose and a source head; aligning the target head to a standard pose; encoding the aligned target head as a target vector and the source head as a source vector in latent space; copying one or more components from the source vector to the target vector; rendering a modified target vector containing one or more components from the encoded source head; realigning the rendered target head to an initial pose; and blending the realigned target head with the source image.

[0006] In some embodiments, the determination made using the face editor is based on the angular difference between a first angle of the target head and a second angle of the source head. In some embodiments, the determination made using the head editor is based on the distance between a bounding box surrounding the target head or target face, and bounding boxes associated with one or more other subjects in the target image. In some embodiments, generating a composite image further includes inpainting remaining target pixels in the target image that are related to the target head but not to the source head, in response to identifying those remaining target pixels. In some embodiments, the method further includes determining occlusion of the target head or source head based on determining the difference in color histograms between the target image and the source image, and the determination made using the head editor is based on occlusion of the target head or source head.

[0007] In some embodiments, the method further includes capturing a set of images with a camera before deciding whether to use one or more editors based on the set of images, and providing the user with a user interface that includes a target image and a choice for selecting a source head from a set of source images, the set of source images including source images, and the method further includes receiving a selection of source images from the user. In some embodiments, at least one subject in the source images is a person or an animal.

[0008] The system includes one or more processors and one or more computer-readable media containing instructions, the instructions being stored in the one or more computer-readable media and, when executed by one or more processors, causing one or more processors to perform an action. This action includes receiving a set of images, including a source image and a target image, the source image and the target image including at least a subject, and further includes deciding whether to use one or more editors selected from a group of head editors, face editors, or combinations thereof, based on the set of images, and, in response to deciding to use a head editor, generating a composite image, which is done by replacing at least some of the head pixels associated with the target head of the subject in the target image with head pixels from the source head of the subject in the source image, and replacing the neck pixels associated with the target neck and the shoulder pixels associated with the target shoulder, including the region between the target head and the target torso, with an interpolated region generated from interpolation between the source image and the target image.

[0009] In some embodiments, this operation further includes adjusting at least a portion of the target facial features in the target image based on face pixels from the source facial features in the source image, in response to a decision made using the face editor. In some embodiments, adjusting at least a portion of the target facial features in the target image based on face pixels from the source facial features in the source image includes: extracting the target head in an initial pose and the source head; aligning the target head to a standard pose; encoding the aligned target head as a target vector and the source head as a source vector in latent space; copying one or more components from the source vector to the target vector; rendering a modified target vector containing one or more components from the encoded source head; realigning the rendered target head to an initial pose; and blending the realigned target head with the source image. In some embodiments, the decision made using the face editor is based on the angular difference between a first angle of the target head and a second angle of the source head. In some embodiments, the decision made using the head editor is based on a bounding box surrounding the target head or target face, and the distance between this bounding box and bounding boxes associated with one or more other subjects in the target image. In some embodiments, generating a composite image further includes inpainting the remaining target pixels in the target image in response to identifying the remaining target pixels in the target image that are related to the target head but not to the source head.

[0010] A non-temporary computer-readable medium on which instructions are stored, the instructions causing one or more processing devices to perform an action in response to execution by one or more processing devices. The action includes receiving a set of images, including a source image and a target image, the source image and the target image including at least a subject, and further includes deciding whether to use one or more editors selected from a group of head editors, face editors, or combinations thereof, based on the set of images, and generating a composite image in response to deciding to use a head editor, the generation being done by replacing at least some of the head pixels associated with the target head of the subject in the target image with head pixels from the source head of the subject in the source image, and replacing the neck pixels associated with the target neck and the shoulder pixels associated with the target shoulder, including the region between the target head and the target torso, with an interpolated region generated from interpolation between the source image and the target image.

[0011] In some embodiments, this operation further includes adjusting at least a portion of the target facial features in the target image based on face pixels from the source facial features in the source image, in response to a decision made using the face editor. In some embodiments, adjusting at least a portion of the target facial features in the target image based on face pixels from the source facial features in the source image includes: extracting the target head in an initial pose and the source head; aligning the target head to a standard pose; encoding the aligned target head as a target vector and the source head as a source vector in latent space; copying one or more components from the source vector to the target vector; rendering a modified target vector containing one or more components from the encoded source head; realigning the rendered target head to an initial pose; and blending the realigned target head with the source image. In some embodiments, the decision made using the face editor is based on the angular difference between a first angle of the target head and a second angle of the source head. In some embodiments, the decision made using the head editor is based on the distance between a bounding box surrounding the target head or target face, and the bounding boxes associated with one or more other subjects in the target image. [Brief explanation of the drawing]

[0012] [Figure 1] This is a block diagram of an exemplary network environment according to some embodiments described herein. [Figure 2] This is a block diagram of an exemplary computing device according to some embodiments described herein. [Figure 3A] This specification describes an exemplary user interface that includes options for the user to specify a target image from a set of images, according to several embodiments described herein. [Figure 3B]This specification describes an exemplary user interface that includes options for the user to specify one or more source images from a set of images, according to several embodiments described herein. [Figure 4A] These are exemplary target images for a head editor according to some embodiments described herein. [Figure 4B] These are exemplary composite images generated by a head editor according to some embodiments described herein. [Figure 5] These are exemplary target images according to some embodiments described herein. [Figure 6A] An exemplary first target head separated from a target image according to some embodiments described herein. [Figure 6B] An exemplary second target head separated from a target image, according to some embodiments described herein. [Figure 7A] An exemplary first source head separated from a source image according to some embodiments described herein. [Figure 7B] An exemplary second source head separated from a source image according to some embodiments described herein. [Figure 8A] An exemplary first source head aligned with a first target head, according to some embodiments described herein. [Figure 8B] An exemplary second source head aligned with a second target head, according to some embodiments described herein. [Figure 9] These are exemplary composite images analyzed for inpainting according to some embodiments described herein. [Figure 10] This is an exemplary composite image in which the target head in Figure 5 is replaced with a source head, according to some embodiments described herein. [Figure 11]An exemplary block diagram of a machine learning model for generating a synthetic image, according to some embodiments described herein. [Figure 12A] An exemplary target image for a face editor, according to some embodiments described herein. [Figure 12B] An exemplary synthetic image generated by a face editor, according to some embodiments described herein. [Figure 13] An exemplary user interface for selecting a best take, according to some embodiments described herein. [Figure 14] A flowchart of an exemplary method for generating a synthetic image, according to some embodiments described herein. [[ID=**14**]] **DETAILED DESCRIPTION**

[0013] **Overview** Media applications generate synthetic images, where one or more of the subjects in the synthetic image have a head and / or face from a source image. Previous attempts to combine parts of images have resulted in unrealistic synthetic images, such as visible seams, visible pixels from the original object in parts where the replacement object is not aligned with the original object, and artifacts caused by occluding objects.

[0014] In some embodiments, a media application receives a source image and a target image and determines whether to use a head editor and / or a face editor to generate a synthetic image. For example, the media application may select a head editor based on the heads of two subjects in the target image being sufficiently far apart or not being occluded by an object. In other examples, the media application may consider the angle of the face between the source image and the target image or Note: There seems to be a formatting issue with the original text where line breaks are a bit irregular. The translation attempts to maintain the overall structure as closely as possible while making the English text more grammatically correct and clear. Also, the "**DETAILED DESCRIPTION**" and "**Overview**" headings are added in a way to make the English more in line with typical patent text structure, as the original text doesn't have a clear indication of these being headings in the given format. The "**14**" in the translation of line ID 14 is just the preservation of the original tag " ".The face editor may be selected based on the pose difference being close enough to allow a portion of the source image to be added to the target image.

[0015] The head editor can replace the head in the target image with the head in the source image. The face editor can adjust portions of the face, such as the eyes and mouth, from the target image based on the face portion from the source image. For example, the face editor may use an embedding to calculate pixel values for portions of the face in the target image. The media application generates a composite image from a combination of the target image and the source image.

[0016] Generating a composite image includes analyzing the image data to determine occlusions and transforming the image data, and may perform inpainting to avoid visible seams or other defects that could cause the composite image to look unrealistic. In other words, the techniques described herein for generating a composite image from one or more source images can seamlessly maintain the realism of the one or more source images.

[0017] Exemplary Environment FIG. 1 shows a block diagram of an exemplary environment 100. In some embodiments, environment 100 includes a media server 101, a user device 115a, and a user device 115n coupled to a network 105. Users 125a, 125n may be associated with their respective user devices 115a, 115n. In some embodiments, environment 100 may include other servers or devices not shown in FIG. 1. In FIG. 1 and the remaining figures, a letter following a reference number, such as “115a,” represents a reference to an element having that particular reference number. A reference number in the text without a subsequent letter, such as “115,” represents a general reference to an embodiment of the element bearing that reference number.

[0018] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicably coupled to the network 105 via signal lines 102. The signal lines 102 may be a wired connection such as Ethernet®, coaxial cable, or fiber optic cable, or a wireless connection such as Wi-Fi®, Bluetooth®, or other wireless technology. In some embodiments, the media server 101 sends and receives data to and from one or more user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.

[0019] Database 199 may store machine learning models, training datasets, images, etc. Database 199 may also store social network data associated with user 125, user preferences of user 125, etc.

[0020] The user device 115 may be a computing device that includes memory coupled to a hardware processor. For example, the user device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or other electronic device that can access the network 105.

[0021] In the illustrated embodiment, user device 115a is connected to network 105 via signal line 108, and user device 115n is connected to network 105 via signal line 110. Media application 103 can be stored on user device 115a as media application 103b and / or on user device 115n as media application 103c. Signal lines 108 and 110 may be wired connections such as Ethernet, coaxial cable, or fiber optic cable, or wireless connections such as Wi-Fi®, Bluetooth®, or other wireless technologies. User devices 115a and 115n are accessed by users 125a and 125n, respectively. User devices 115a and 115n in Figure 1 are used as examples. Although Figure 1 shows two user devices, 115a and 115n, this disclosure applies to system architectures having one or more user devices 115.

[0022] The media application 103 may be stored on the media server 101 or the user device 115. In some embodiments, the operations described herein are performed on the media server 101 or the user device 115. In some embodiments, some operations may be performed on the media server 101 and some on the user device 115. The execution of operations is subject to user settings. For example, user 125a may specify that operations are performed on each device 115a and not on the media server 101. Such a setting would cause the operations described herein to be performed entirely on the user device 115a and not on the media server 101. Furthermore, user 125a may specify that user images and / or other data be stored only locally on the user device 115a and not on the media server 101. Such a setting would cause user data not to be sent to or stored on the media server 101. The transmission of user data to the media server 101, the media server 101's arbitrary temporary or permanent storage of such data, and the media server 101's execution of actions on such data are performed only if the user consents to the transmission, storage, and execution of actions by the media server 101. The user is provided with the option to change settings at any time, for example, so that the user can enable or disable the use of the media server 101.

[0023] Machine learning models (e.g., neural networks or other types of models), when used for one or more operations, are stored and used locally on the user device 115 with the permission of a specific user. Server-side models are used only with the user's permission. Furthermore, trained models may be provided for use on the user device 115. During such use, on-device training of the model may be performed if permitted by user 125. If permitted by user 125, updated model parameters may be sent to the media server 101, for example, to enable federative learning. Model parameters do not contain any user data.

[0024] Media application 103 receives a set of images, including a source image and a target image, the source image and target image including at least a subject. Based on the set of images, media application 103 decides whether to use a head editor and / or a face editor. In response to deciding to use a head editor, media application 103 includes generating a composite image, which is done by replacing at least some of the head pixels associated with the target head of the subject in the target image with head pixels from the source head of the subject in the source image, and by replacing the neck pixels associated with the target neck and the shoulder pixels associated with the target shoulder, including the region between the target head and the target torso, with an interpolated region generated from interpolation between the source image and the target image.

[0025] In some embodiments, the media application 103 can run using hardware including a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, the media application 103a can run using a combination of hardware and software.

[0026] Exemplary computing device Figure 2 is a block diagram of an exemplary computing device 200 that may be used to perform one or more of the features described herein. The computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, the computing device 200 is a media server 101 used to run a media application 103a. In another example, the computing device 200 is a user device 115.

[0027] In some embodiments, the computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, display 241, camera 243, and storage device 245, all coupled via a bus 218. The processor 235 may be coupled to the bus 218 via a signal line 222, the memory 237 may be coupled to the bus 218 via a signal line 224, the I / O interface 239 may be coupled to the bus 218 via a signal line 226, the display 241 may be coupled to the bus 218 via a signal line 228, the camera 243 may be coupled to the bus 218 via a signal line 230, and the storage device 245 may be coupled to the bus 218 via a signal line 232.

[0028] The processor 235 may be one or more processors and / or processing circuits that execute program code and control the basic operation of the computing device 200. "Processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., single-core, dual-core, or multi-core configuration), multiple processing units (e.g., a multi-processor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a composite programmable logic device (CPLD), a dedicated circuit for realizing a function, a dedicated processor for performing processing based on a neural network model, a neural circuit, a processor optimized for matrix calculations (e.g., matrix multiplication), or a system having other systems. In some embodiments, the processor 235 may include one or more coprocessors that perform neural network processing. In some embodiments, the processor 235 may be a processor that processes data to produce a probabilistic output, for example, the output produced by the processor 235 may be inaccurate or accurate within a range from an expected output. Processing does not need to be limited to a specific geographical location or have temporal constraints. For example, a processor can perform its functions in real time, offline, or batch mode. Parts of the processing can be executed by different (or the same) processing systems at different times and in different locations. A computer can be any processor that communicates with memory.

[0029] Memory 237 may be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), or flash memory, typically located separately from and / or integrated with the processor 235, provided in the computing device 200 for access by the processor 235 and suitable for storing instructions for execution by the processor or a set of processors. Memory 237 can store software running on the computing device 200 by the processor 235, including media applications 103.

[0030] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, and the like. One or more methods disclosed herein may operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application ("App") that runs on a mobile computing device, and the like.

[0031] Application data 266 may be data generated by other applications 264 or the hardware of computing device 200. For example, application data 266 may include images used by an image library application and user actions identified by other applications 264 (e.g., a social networking application).

[0032] The I / O interface 239 can provide functionality that enables the computing device 200 to interface with other systems and devices. Interfaced devices may be included as part of the computing device 200, or they may be separate and communicate with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 245), and input / output devices can communicate via the I / O interface 239. In some embodiments, the I / O interface 239 can connect to interface devices such as input devices (keyboards, pointing devices, touchscreens, microphones, scanners, sensors, etc.) and / or output devices (display devices, speaker devices, printers, monitors, etc.).

[0033] Some examples of interface devices that can be connected to the I / O interface 239 include a display 241 that can be used to display content, such as images, videos, and / or user interfaces of output applications described herein, and to receive touch (or gesture) input from the user. For example, the display 241 may be used to display a user interface, including graphical guides, on a viewfinder. The display 241 may include any suitable display device such as a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, cathode ray tube (CRT), television, monitor, touchscreen, three-dimensional display screen, or other visual display device. For example, the display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in a glasses form factor or headset device, or a monitor screen on a computer device.

[0034] Camera 243 can be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video that the I / O interface 239 sends to the media application 103.

[0035] The storage device 245 stores data related to the media application 103. For example, the storage device 245 can store a training dataset that includes labeled images, machine learning models, and outputs from the machine learning models.

[0036] Figure 2 shows an exemplary media application 103 stored in memory 237, which includes an image module 202, a head editor 204, and a face editor 206.

[0037] The image module 202 generates graphic data for displaying a user interface that includes a set of images. The set of images may be received from the camera 243 of the computing device 200 and / or from the media server 101 via the I / O interface 239. For example, the set of images may include images from a burst of images captured by the camera 243. A burst of images may include multiple photographs captured quickly and consecutively over a short period of time. The set of images includes one or more source images and target images, each containing one or more subjects. The one or more subjects may be people, animals, etc.

[0038] Image module 202 obtains permission from the user to modify any image within the set of images. The user may be provided with controls that allow the user to choose whether and when the systems, programs, or functions described herein may enable the collection or use of user information (e.g., the user's identification in the images, user preferences, or user's current location), as well as whether the user receives content or communications from the server. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, user identification information may be processed so that personally identifiable information cannot be determined, or the user's geographical location may be generalized so that the user's specific location cannot be determined, if location information is available (e.g., city, zip code, or state level). Thus, the user can control what information is collected about them, how that information is used, and what information is provided to them.

[0039] In some embodiments, the image module 202 selects one or more source images and a target image to be used to generate a composite image. For example, the image module 202 may automatically generate a composite score for each image in the set of images based on composite factors such as the type of objects in the image, the position of the subjects, and the lighting. For example, the target image may be selected based on having the best landscape composition and the appropriate positioning of the subjects within the group. In another example, the target image may be selected based on having the most subjects looking at the camera, smiling, or with their eyes open. An image with a high overall composite score may be used as the target image. In some embodiments, the user may select a specific image in the set of images as the target image, and one or more other images in the set of images may be presented as source images.

[0040] Image module 202 can select one or more source images based on subjects in the image. For example, image module 202 can generate a face score for each subject in the image based on quality metrics. For example, each face can be scored based on whether the face is characterized by being smiling, having open eyes, being completely or partially occluded in the photograph by other people or objects, and / or being motionless (e.g., not blurred). Image module 202 can generate a head score based on the position of the subject's head, for example, it may generate a higher head score for a head that is in a vertical plane (e.g., facing the camera) rather than at other angles (e.g., tilted, rotated away from the camera, etc.). Such scoring can be performed using image detection and recognition techniques, such as a trained machine learning model that can detect faces in a photograph and apply quality scoring criteria.

[0041] Based on the score, the image module 202 can rank faces within an image. For example, the image module 202 can determine that multiple specific faces in different images within a set of images belong to the same subject (people, animals, etc.), and each such subject can be associated with a ranked list of faces from the images. The image module 202 can automatically select a target image and one or more source images, or it can suggest higher-ranked images as a suggestion to the user to confirm that a particular image has been selected as the target image.

[0042] In some embodiments, the image module 202 generates graphic data for displaying a user interface that provides the user with the option to select a target image and one or more source images from a set of images. The set of images included in the user interface may be a burst of captured images, images captured over a specific period of time (e.g., images captured over the past 24 hours, at a specific location, etc.).

[0043] Figure 3A shows an exemplary user interface 300 that includes options for the user to specify a target image from a set of images. In this example, the user selects an image 305 to be used as the target image. In some embodiments, a particular image from the set of images may be highlighted (e.g., with a border, an icon, etc.) and suggested as the target image. The user can select the suggested image or any other image as the target image.

[0044] Figure 3B shows an exemplary user interface 325 that includes choices for the user to specify one or more source images from a set of images. In this example, the user selects a first image 330 to be used as the source image for a first subject, and a second image 335 to be used as the source image for a second subject. In some embodiments, the image module 202 can generate a user interface that includes choices for selecting a particular subject within an image. For example, the computing device 200 may receive user input in the form of a double-click on a subject, a circle drawn around the head of a subject, or the like.

[0045] In some embodiments, a single image may be selected as the source image for two or more subjects. In some embodiments, each subject may be associated with a different source image. In some embodiments, a suggestion of source images may be provided for one or more subjects. For example, if the target image has a first subject with its head tilted and a second subject with its eyes closed, a first source image in which the first subject is facing the camera without tilting its head, and a second source image in which the second subject has its eyes open may be suggested as the respective source images. In some embodiments, other factors such as lighting, the duration between the capture of the target image and a particular source image, and the distance between the position of the subject in the source image and the target image may be used when selecting a particular source image to recommend.

[0046] In some embodiments, the user interface may include options for searching for other source images. For example, the user interface may include options for the user to scroll (or otherwise browse) through images in the camera roll and select a source image. In some embodiments, since source images from the camera roll may be captured with different lighting conditions, different shadows, etc., the image module 202 may modify the source image to have colors (and / or other image attributes such as brightness, white balance, contrast, etc.) that match the corresponding attributes of the target image.

[0047] Image module 202 determines whether to use head editor 204, face editor 206, or both to generate a composite image. Image module 202 can make this decision using different criteria, such as the proximity of heads in the target image or source image, the angle of heads in the target image or source image, and the occlusion of heads in the target image or source image.

[0048] In some embodiments, the image module 202 determines whether to use the head editor 204 or the face editor 206 based on different factors evaluated by different classifiers. In some embodiments, the image module 202 performs a weighted sum or logistic regression of different factors so that the decision is based on the set of factors rather than a single factor. In some embodiments, some factors may be orientation-determining, such as whether one of the heads is 70% occluded by an object.

[0049] In some embodiments, the classifier is trained from evaluations of head edits from the head editor 204 and face edits from the face editor 206 in a training image set. For example, the training data may include a target image, a source image, a composite image generated by the head editor 204 and / or the face editor 206, and a corresponding evaluation of the composite image, which may be provided by a human or a quality algorithm. As a result of receiving the evaluation, the image module 202 can recalculate the classifier weights to improve the quality of the composite image generated by the head editor 204 and the face editor 206.

[0050] In some embodiments, the image module 202 generates bounding boxes around each target head / target face in the target image and uses the head editor 204 to determine whether to replace the target head with a source head based on the distance between the bounding boxes in the target image. For example, if there is overlap between the bounding boxes, this indicates that the heads of the subjects are close to each other, and the head editor 204 may not be used due to variations in head position if multiple subjects are close together, resulting in a larger section of background being inpainted in the composite image, potentially causing subjects to overlap, and making the alignment of subjects in the target image and one or more source images complex. In some embodiments, the image module 202 determines whether to use the head editor 204 based on a continuous function. In some embodiments, the continuous function is based on a weighted distance between bounding boxes and / or a weighted overlap percentage between bounding boxes, the weight values ​​may be learned during training of the classifier used by the image module 202.

[0051] In some embodiments, the image module 202 determines whether to use the head editor 204 based on the angle difference between a first angle of the target head and a second angle of the source head. The face editor 206 may produce an unsatisfactory or low-quality composite image if the angle of the head in the source image is greater than a threshold difference with the angle of the head in the target image.

[0052] The image module 202 may decide to use the face editor 206 based on applying a nonlinear function to the difference between a first angle and a second angle. In some embodiments, the image module 202 applies a cosine function to the first angle, applies a cosine function to the second angle, and determines the difference between the first and second angles. In some embodiments, the image module 202 uses a threshold angle difference to decide whether to use the face editor 206, and if the angle difference exceeds the threshold angle difference, the image module 202 decides to use the head editor 204 instead of the face editor 206.

[0053] In some embodiments, the image module 202 determines whether to use the head editor 204 based on the occlusion of the target head or the source head. When the image module 202 receives either the target image or the source image to be occluded, the resulting composite image will have a higher failure rate and / or lower quality if the source image is occluded. In some embodiments, the image module 202 can use the color histogram of the image to determine whether the source head or the target head is occluded. For example, the image module 202 may use the mean value of a particular channel in the histogram to identify occlusion.

[0054] In some embodiments, the image module 202 determines whether to use the head editor 204, the face editor 206, or both the head editor 204 and the face editor 206 based on the distance from the subject to the image boundary. For example, if the head is close to the image boundary and the head editor 204 modifies the head's pose near the image boundary, this may result in part of the head being cropped at the image boundary (and therefore the subject's head is not fully in the image, resulting in an unsatisfactory composite image).

[0055] The head editor 204 replaces the target head with the source head. In some embodiments, the head editor 204 replaces the target pixels associated with the target head in the target image with source pixels from the source head in the source image. The head editor 204 can generate a head mask including the subject's hair, segment the head mask from the source image, and apply the pixels in the head mask to the target image. The head editor 204 can adjust the position and scale of the source head to match the dimensions of the target head.

[0056] The head editor 204 performs inpainting when replacing the target head with the source head results in areas where the target head and source head do not overlap. This can occur if the target head and source head are related at different angles. This can also occur if the angles of the target head and source head in the target image and source image are different, and the head editor 204 aligns the angle of the target head with the angle of the source head before replacement. By aligning the target head with the source head, it is possible to identify the areas of the target head that do not overlap with the source head. The head editor 204 may perform inpainting of remaining target pixels by identifying the remaining pixels in the target image that are related to the target head but not to the source head, and replacing the remaining target pixels. Inpainting pixels may include determining the distance between the source pixels and the remaining pixels, and generating replacement pixels based on similarity to the source pixels and the distance.

[0057] Because the angles between the target head and the source head are different, the position of the subject's neck region may differ in the two images. The head editor 204 renders a smooth transition between the target torso and the source head by generating interpolated regions for the neck and shoulder regions. In some embodiments, the head editor 204 generates interpolated regions that include the region between the target head and the target torso (e.g., the region described as the target neck and target shoulder), and replaces the target pixels associated with the target neck and target shoulder with the interpolated regions, which are interpolations of the target and source pixels of the corresponding neck and shoulder regions. In some embodiments, the head editor 204 uses multiple image frames, such as a set of image frames generated from a burst of images captured by the camera 243, and generates interpolated regions from multiple image frames.

[0058] Referring to Figure 4A, an exemplary target image 400 for a head editor is shown according to several embodiments described herein. The subject 405 on the right is looking upward and has smooth shoulders.

[0059] Figure 4B is an exemplary composite image 425 generated by the head editor 204. The head editor 204 replaced the target head with the source head in the subject 430 on the right. In composite image 425, the interpolated region 430 above the neck and shoulders is a result of the neck being oriented differently and the shoulders containing more wrinkled fabric compared to the shoulders of subject 405 in the target image 400.

[0060] In some embodiments, the head editor 204 includes a machine learning model that takes a target image and one or more source images as input and outputs a composite image. The trained machine learning model may include one or more model forms or structures. For example, the model form or structure may include any type of neural network, such as a linear network, a deep learning neural network that runs on multiple layers (with a "hidden layer" between the input and output layers, where each layer is a linear network), a convolutional neural network (a network that divides or partitions input data into multiple parts or tiles, processes each tile individually using one or more neural network layers, and aggregates the results of processing each tile), or a sequence-to-sequence neural network (a network that takes sequential data such as words in a sentence or frames of a video as input and outputs a resulting sequence).

[0061] The model format or structure may specify the connectivity between various nodes and the organization of the nodes into layers. For example, the nodes in the first layer (e.g., the input layer) may receive data as input data or application data. Such data may include, for example, one or more pixels per node, when a trained model is used to analyze, for example, a target image and one or more source images. Subsequent intermediate layers may receive the outputs of the nodes in the previous layer as input, according to the connectivity specified in the model format or structure. These layers may also be referred to as hidden layers. The final layer (e.g., the output layer) produces the output of the machine learning model. For example, the output layer may output a composite image. In some embodiments, the model format or structure also specifies the number and / or type of nodes in each layer.

[0062] In another embodiment, the trained model may include one or more models. One or more of the models may include multiple nodes arranged in layers according to the model structure or form. In some embodiments, a node may be a memoryless computation node configured to process one unit of input and produce one unit of output. The computation performed by the node may include, for example, multiplying each of the multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce a node output. In some embodiments, the computation performed by the node may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computation may include operations such as matrix multiplication. In some embodiments, computations by multiple nodes can be performed in parallel, for example, using multiple processor cores of a multicore processor, individual processing units of a graphics processing unit (GPU), or dedicated neural circuits. In some embodiments, a node may include memory, for example, capable of storing and using one or more previous inputs when processing subsequent inputs. For example, a node with memory may include a Long Short-Term Memory (LSTM) node. LSTM nodes can use memory to maintain "states" that allow the node to behave like a finite state machine (FSM).

[0063] In some embodiments, the trained model may include embeddings or weights for individual nodes. For example, the model may start as a group of nodes organized into layers, as specified by the model format or model structure. In initialization, each weight may be applied to the connections between each pair of nodes connected according to the model format, for example, to nodes in a series of layers of a neural network. For example, each weight may be assigned randomly or initialized to a default value. The model may then be trained, for example, using training data to produce results.

[0064] Training may involve applying supervised learning methods. In supervised learning, training data may include multiple inputs (e.g., target and source images, segmentation masks, etc.) and corresponding ground truth outputs for each input (e.g., ground truth masks that accurately identify parts of a subject, such as faces, in each image, a composite image, etc.). Based on a comparison of the model's output with the ground truth outputs, the weight values ​​are automatically adjusted, for example, in a way that increases the probability that the model will produce the ground truth output for the composite image.

[0065] In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In some embodiments, the trained model may include a fixed set of weights downloaded from, for example, a server that provides weights. In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In embodiments where data is omitted, the head editor 204 may generate a trained model based on prior training, for example, by the developer of the head editor 204, by a third party, etc. In some embodiments, the trained model may include a fixed set of weights downloaded from, for example, a server that provides weights.

[0066] In some embodiments, a trained machine learning model receives a target image and one or more source images, each containing one or more subjects. The machine learning model can generate one or more segmentation masks that identify pixels in one or more source images that correspond to one or more heads, including hair, in the one or more source images. For each subject, the machine learning model replaces head pixels from the target image with head pixels from one or more source images that have corresponding positions, scales, and angles that fit the source images. The machine learning model blends the head pixels along the edges of the head. In some embodiments, background pixels exposed as a result of the difference between the source heads and the target heads are inpainted. In some embodiments, the machine learning model blends interpolated regions corresponding to the neck and shoulders of the subjects into the target image. In some embodiments, the trained machine learning model outputs a composite image that incorporates these changes.

[0067] In some embodiments, the machine learning model outputs a confidence value for each composite image produced by the trained machine learning model. The confidence value may be expressed as a percentage, a number between 0 and 1, etc. For example, the machine learning model outputs an 85% confidence value for the confidence that the composite image correctly replaces the target head with the source head and does not contain pixels from other people or objects. In some embodiments, the composite image is provided to the user if the confidence value exceeds a confidence threshold. In some embodiments, the confidence value is calculated in advance, and the composite image is not generated unless the confidence value exceeds a threshold confidence value.

[0068] In some embodiments, the head editor 204 includes multiple machine learning models that perform different functions in the steps used to generate a composite image. For example, a first machine learning model can replace target pixels associated with a target head with source pixels from a source head, a second machine learning model aligns the source head with the target head, and a third machine learning model generates interpolation regions that replace target pixels associated with a target neck and target shoulders.

[0069] Figure 5 is an exemplary target image 500. The boy's head 505 and the girl's head 510 in the target image 500 are replaced with heads from the source image.

[0070] Figure 6A is an exemplary first target head 600 separated from the target image 500 in Figure 5. Figure 6B is an exemplary second target head 650 separated from the target image 500 in Figure 5. The first target head 600 may be separated and selected for replacement because the boy's nose is pointing upwards and his mouth is open. The second target head 650 may be separated and selected for replacement because the girl is not looking directly at the camera.

[0071] Figure 7A shows an exemplary first source head 700 separated from the source image. Figure 7B shows an exemplary second source head 750 separated from the same source image containing the first source head 700, or from a different source image. The head editor 204 replaces at least some of the target pixels associated with the first target head 600 and at least some of the target pixels associated with the second target head 650 with source pixels from the first source head 700 and the second source head 750, respectively.

[0072] Figure 8A shows an example of a first source head aligned with a first target head 800. Figure 8B shows an example of a second source head aligned with a second target head 850. The images are aligned, but there are some areas where the images do not overlap, such as the position of the boy's cap 805.

[0073] Figure 9 is an exemplary composite image 900 analyzed for inpainting based on the alignment of a first source head that exposes areas of background previously covered by the first target head. After the head editor 204 performs inpainting of the target head, the composite image appears as if the source head has been incorporated into the composite image without seams or other defects.

[0074] Figure 10 is an exemplary composite image 1000 in which the source head is replaced with the target head. The boy's head 505 in Figure 5 is replaced with the boy's head 1005 in Figure 10, and the girl's head 510 in Figure 5 is replaced with the girl's head 1010 in Figure 10.

[0075] The face editor 206 transfers facial features (e.g., smile, open eyes, mouth shape, etc.) from the source image to the target image. The face editor 206 does not change the pose of the target face. In some embodiments, the face editor 206 adjusts at least some of the target pixel values ​​associated with the target facial features in the target image based on the source pixels from the source facial features in the source image.

[0076] In some embodiments, the face editor 206 extracts target and source faces of the same subject using a face matching algorithm that runs with the permission of a specific user. The face editor 206 aligns the target face from an initial pose to a standard pose, where the standard pose is one of the poses trained for use by a machine learning model.

[0077] In some embodiments, the face editor 206 uses a machine learning model selected from one of the model types described above with respect to the head editor 204. For example, the machine learning model may include an encoder and a convolutional neural network (CNN).

[0078] In some embodiments, not all of the target face can be replaced with the source, as some facial features may be susceptible to small changes in angle and pose. For example, adding a nose from an upward-tilted head to a front-facing face may make the person's nose appear unnatural. As a result, the machine learning model's encoder uses a facial feature mask (e.g., an eye / mouth mask) to encode the components of the source head as a source vector. The encoder also encodes the target head as a target vector. The machine learning model then copies the corresponding components from the source vector to the target vector.

[0079] In some embodiments, the CNN includes multiple layers, each generating a version of the composite image by rendering a target vector containing source vector components. The CNN generates the composite image by progressively increasing the resolution level in each layer. The high-resolution outputs are blended with the existing low-resolution layers until the final composite image is output. In some embodiments, the machine learning model realigns the rendered target head back to its initial pose, blends the realigned target head with the source image, and outputs the composite image.

[0080] Figure 11 is an exemplary block diagram 1100 of an on-device machine learning model that generates a composite image. The source image 1105 and target image 1110 are received by both an offline module 1115 and an on-device module 1120, such as a media application 103 stored in a user device 115 shown in Figure 1. The on-device module 1120 outputs a predicted image 1130 such that facial expressions and / or head poses from the source image 1105 are fused into the target image 1110. The on-device module 1120 is implemented so that it can run on a user device, e.g., a smartphone with or without machine learning circuitry, within a latency and memory usage budget.

[0081] The offline module 1115 can provide training for the on-device model, which favorably enables the lightweight on-device module 1120 to generate the composite image 1125. Because the offline module 1115 has a larger budget for latency and memory usage, the composite image 1125 may be a higher quality version than the predicted image 1130.

[0082] In some embodiments, the on-device module 1120 may use warping and reconstruction in which a high-density face mesh warps the source image 1105 to a target position, or it may use a neural network (or other suitable technique) to remove artifacts. Given two face images and their face meshes, the on-device module 1120 aligns the pose of the source face to the target face mesh, and then warps the source face to the target image. The warped face has the expression of the source but has artifacts caused by the alignment and warping, which may be considered degradation. The on-device module 1120 may include a managed reconstruction model from an offline module 1115 that can remove or repair the degradation.

[0083] In some embodiments, the on-device module 1120 can modify the face to reflect a specified expression. The on-device module 1120 may include two encoders: a source encoder that extracts expression codes from a source image, and a target encoder that extracts head pose and appearance information from a target image. The on-device module 1120 can insert expression codes at one or more levels of the decoder (e.g., at any level). Additionally, a skim connection from the target encoder can be added to ensure that the on-device module 1120 edits the face but not other parts of the image.

[0084] Figure 12A is an exemplary target image 1200 in which a second subject 1205 has an expression that is replaced by the face editor 206.

[0085] Figure 12B is an exemplary composite image 1250 generated by the face editor 206 according to the technique described above. As can be seen from the figure, the open-mouthed, teeth-showing smile of subject 1205 on the right in the target image 1200 is replaced with a closed-mouthed smile of subject 1255 in the composite image 1250.

[0086] Figure 13 shows exemplary user interfaces 1300, 1350, and 1375 for selecting the best take according to several embodiments described herein. The first user interface 1300 includes target images 1305 of two subjects 1306 and 1307. The user can select the “Best Take” button 1310 to start the process of obtaining a composite image.

[0087] The second user interface 1350 includes a target image 1355 and icons 1357 and 1359 for the first and second subjects, respectively. The user can select a different face of the subject by selecting one of the icons 1357 and 1359. In this example, the user has selected the icon 1359 for the second subject. The user interface 1350 includes three options 1361, 1363, and 1365 for the second subject. The first option 1361 and the third option 1365 are from the source image, and the second option 1363 is from the target image 1355. Once the user has selected one of the three options 1361, 1363, and 1365, the user can select the complete button 1367 to see a composite image of the selected option on the target image 1355, or select the reset button 1369 to restart the process. In this example, the user has selected the first option 1361 and the complete button 1367.

[0088] The third user interface 1375 includes a composite image 1377 generated from the target image 1355 and source face 1379 of the second user interface 1350. The user can save a copy of the image by pressing the copy save button 1385.

[0089] Exemplary flowchart Figure 14 shows a flowchart of an exemplary method 1400 for segmenting an image, according to several embodiments described herein. Method 1400 may be performed by the computing device 200 of Figure 2. In some embodiments, Method 1400 is performed by a user device 115, a media server 101, or partially on the user device 115 and partially on the media server 101.

[0090] Method 1400 in Figure 14 may begin with block 1402. In block 1402, a set of images is received, including a source image and a target image, where the source image and target image include one or more subjects. In some embodiments, prior to block 1402, Method 1400 further includes capturing a set of images with a camera and providing a user interface that includes a target image and a choice to select a source head from a set of source images, where the set of source images includes source images, and Method 1400 further includes receiving a selection of source images from the user. The subjects may include people or animals. Block 1402 may be followed by block 1404.

[0091] Block 1404 determines whether permission has been granted to modify the source image and the target image. If permission is not granted, block 1406 may follow block 1404, and method 1400 terminates. If permission is granted, block 1408 may follow block 1404.

[0092] In block 1408, it is determined whether to use one or more editors selected from the group consisting of a head editor, a face editor, or a combination thereof. The decision to use the face editor may be based on the angle difference between a first angle of the target head and a second angle of the source head, on the distance between the bounding box surrounding the target head or target face and the bounding box relating to other subjects in the target image, or on the occlusion of the target head or the source head. The decision to use the occlusion of the target head or the source head may be based on determining the difference in color histograms.

[0093] If the head editor is selected, block 1410 may follow block 1408. In block 1410, the composite image is created by replacing at least some of the head pixels associated with the target head of the subject in the target image with head pixels from the source head of the subject in the source image, and by replacing the neck pixels associated with the target neck and the shoulder pixels associated with the target shoulder, including the region between the target head and the target torso, with an interpolated region generated from the interpolation between the source image and the target image. In some embodiments, the head editor aligns the angle of the target head with the angle of the source head before replacing at least some of the target pixels associated with the target head in the target image. If the face editor is selected, block 1412 may follow block 1408.

[0094] In block 1412, the composite image is generated by adjusting at least a portion of the target facial features in the target image based on the facial pixels from the source facial features in the source image. In some embodiments, adjusting at least a portion of the target facial features in the target image based on the facial pixels from the source facial features in the source image includes: extracting the target head in an initial pose and the source head; aligning the target head to a standard pose; encoding the aligned target head as a target vector and the source head as a source vector in latent space; copying one or more components from the source vector to the target vector; rendering the modified target vector, which includes one or more components from the encoded source head; realigning the rendered target head to an initial pose; and blending the realigned target head with the source image.

[0095] In the above description, many specific details have been given for illustrative purposes to provide a complete understanding of this specification. However, it will be apparent to those skilled in the art that this disclosure can be implemented without these specific details. In some cases, structures and devices are shown in block diagram form to avoid obscuring this specification. For example, embodiments can be described above with reference primarily to user interfaces and specific hardware. However, embodiments can be applied to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.

[0096] Any reference in this specification to “some embodiments” or “some examples” means that certain features, structures, or characteristics described in relation to the embodiments or examples may be included in at least one embodiment of this specification. The phrase “in some embodiments” appearing in various places in this specification does not necessarily refer to the same embodiments.

[0097] Some parts of the detailed explanation above are presented in terms of algorithms and symbolic representations of operations on data bits in computer memory. These descriptions and representations of algorithms are means used by those skilled in the art to most effectively communicate the content of their work to others skilled in the art. Here, and also generally, an algorithm is considered to be a self-consistent set of steps that lead to a desired result. These steps are steps that require the physical manipulation of physical quantities. Usually, though not always necessary, these quantities take the form of electrical or magnetic data that can be stored, transferred, combined, compared, and otherwise manipulated. For reasons of general use, it is sometimes convenient to refer to these data as bits, values, elements, symbols, characters, terms, numbers, etc.

[0098] However, it should be recognized that all these terms and similar terms should correspond to appropriate physical quantities and are merely convenient labels applied to those quantities. Unless otherwise specified, as will become clear from the following discussions, throughout this specification, discussions using terms such as “process,” “calculate,” “calculate,” “determine,” or “display” refer to the actions and processes of a computer system or similar electronic computing device that manipulate data represented as physical (electronic) quantities in the registers and memory of the computer system and convert it into other data similarly represented as physical quantities in the memory or registers of the computer system, or in other such information storage devices, transmission devices, or display devices.

[0099] Embodiments of this specification may also relate to a processor for performing one or more steps of the methods described above. The processor may be a dedicated processor that is selectively started or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-temporary computer-readable storage medium, including but not limited to optical discs, ROMs, CD-ROMs, magnetic disks, RAMs, EPROMs, EEPROMs, magnetic or optical cards, flash memory including USB keys with non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0100] The specification may take the form of several entirely hardware embodiments, several entirely software embodiments, or several embodiments that include both hardware and software elements. In some embodiments, the specification is implemented with software including, but not limited to, firmware, resident software, and microcode.

[0101] Furthermore, this specification may take the form of a computer program product accessible from a computer-enabled or computer-readable medium, which provides program code for use by or in connection with a computer or any instruction execution system. For the purposes of this description, a computer-enabled or computer-readable medium may be any device that can store, communicate, propagate, or carry a program for use by or in connection with an instruction execution system, instruction execution unit, or instruction execution device.

[0102] A data processing system suitable for storing or executing program code would include at least one processor directly or indirectly coupled to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, bulk storage, and cache memory providing temporary storage for at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution.

Claims

1. A method that is performed on a computer, The method includes receiving a set of images, the source image and the target image, the source image and the target image each include at least one subject, and the method further includes: Based on the set of images, it is determined whether to use one or more editors selected from a group of head editors, face editors, or combinations thereof. Using the head editor, in response to a decision, includes generating a composite image, and said generation is Replacing at least a portion of the head pixels associated with the target head of the at least one subject in the target image with head pixels from the source head of the at least one subject in the source image, Replacing neck pixels associated with the target neck and shoulder pixels associated with the target shoulder, including the region between the target head and the target torso, with interpolated regions generated from interpolation between the source image and the target image, Tested by, method.

2. The method according to claim 1, further comprising adjusting at least a portion of the target face features in the target image based on face pixels from the source face features in the source image, in response to a determination made using the face editor.

3. Adjusting at least a portion of the target face features in the target image based on the face pixels from the source face features of the source image is: Extracting the target head and the source head in their initial positions, Aligning the target head to a standard posture, In the latent space, the aligned target head is encoded as a target vector, and the source head is encoded as a source vector, Copying one or more components from the source vector to the target vector, Rendering a modified target vector including the one or more components from the encoded source head, Realigning the rendered target head to the initial pose, Blending the realigned target head with the source image, The method according to claim 2, including the method described in claim 2.

4. The method according to claim 2, wherein the determination made using the face editor is based on the angle difference between the first angle of the target head and the second angle of the source head.

5. The method according to claim 1, wherein the determination using the head editor is based on the distance between a bounding box surrounding the target head or target face and a bounding box associated with one or more other subjects in the target image.

6. The method according to claim 1, wherein generating the composite image further comprises inpainting the remaining target pixels in the target image in response to identifying the remaining target pixels in the target image that are related to the target head and not related to the source head.

7. The method further includes determining the occlusion of the target head or the source head based on determining the difference in color histograms between the target image and the source image. The method according to claim 1, wherein the determination made using the head editor is based on the occlusion of the target head or the occlusion of the source head.

8. Before deciding whether to use the one or more editors based on the set of images, the method further: The camera captures the aforementioned set of images, The method includes providing a user interface to the user that includes the target image and a selection of source heads from a set of source images, wherein the set of source images includes the source images, and the method further includes The method according to claim 1, comprising receiving the selection of the source image from the user.

9. The method according to claim 1, wherein the at least one subject in the source image is a person or an animal.

10. It is a system, One or more processors, The system includes one or more computer-readable media on which instructions are stored, and when an instruction is executed by the one or more processors, it causes the one or more processors to perform an operation, and the operation is The operation includes receiving a set of images including a source image and a target image, wherein the source image and the target image each include at least one subject, and the operation further includes: Based on the set of images, it is determined whether to use one or more editors selected from a group of head editors, face editors, or combinations thereof. Using the head editor, in response to a decision, includes generating a composite image, and said generation is Replacing at least a portion of the head pixels associated with the target head of the at least one subject in the target image with head pixels from the source head of the at least one subject in the source image, Replacing neck pixels associated with the target neck and shoulder pixels associated with the target shoulder, including the region between the target head and the target torso, with interpolated regions generated from interpolation between the source image and the target image, A system that is carried out by [someone / something].

11. The aforementioned operation is, The system according to claim 10, further comprising adjusting at least a portion of the target face features in the target image based on face pixels from the source face features in the source image, in response to a decision made using the face editor.

12. Adjusting at least a portion of the target face features in the target image based on the face pixels from the source face features of the source image is: Extracting the target head and the source head in their initial positions, Aligning the target head to a standard posture, In the latent space, the aligned target head is encoded as a target vector, and the source head is encoded as a source vector, Copying one or more components from the source vector to the target vector, Rendering a modified target vector including the one or more components from the encoded source head, Realigning the rendered target head to the initial pose, Blending the realigned target head with the source image, The system according to claim 11, including the following:

13. The system according to claim 11, wherein the determination to use the face editor is based on the angle difference between the first angle of the target head and the second angle of the source head.

14. The system according to claim 10, wherein the determination made using the head editor is based on the distance between a bounding box surrounding the target head or target face and a bounding box associated with one or more other subjects in the target image.

15. The system according to claim 10, wherein generating the composite image further comprises inpainting the remaining target pixels in the target image in response to identifying the remaining target pixels in the target image that are related to the target head and not related to the source head.

16. The system includes an instruction, the instruction causing one or more processing devices to perform an action in response to execution by one or more processing devices, and the action is: The operation includes receiving a set of images including a source image and a target image, wherein the source image and the target image each include at least one subject, and the operation further includes: Based on the set of images, it is determined whether to use one or more editors selected from a group of head editors, face editors, or combinations thereof. Using the head editor, in response to a decision, includes generating a composite image, and said generation is Replacing at least a portion of the head pixels associated with the target head of the at least one subject in the target image with head pixels from the source head of the at least one subject in the source image, Replacing neck pixels associated with the target neck and shoulder pixels associated with the target shoulder, including the region between the target head and the target torso, with interpolated regions generated from interpolation between the source image and the target image, A computer program performed by [a specific entity / organization].

17. The aforementioned operation is, The computer program according to claim 16, further comprising adjusting at least a portion of the target facial features in the target image based on facial pixels from the source facial features in the source image, in response to a decision made using the face editor.

18. Adjusting at least a portion of the target face features in the target image based on the face pixels from the source face features of the source image is: Extracting the target head and the source head in their initial positions, Aligning the target head to a standard posture, In the latent space, the aligned target head is encoded as a target vector, and the source head is encoded as a source vector, Copying one or more components from the source vector to the target vector, Rendering a modified target vector including the one or more components from the encoded source head, Realigning the rendered target head to the initial pose, Blending the realigned target head with the source image, The computer program according to claim 17, including the computer program described in claim 17.

19. The computer program according to claim 17, wherein the determination to use the face editor is based on the angle difference between the first angle of the target head and the second angle of the source head.

20. The computer program according to claim 16, wherein the determination using the head editor is based on the distance between a bounding box surrounding the target head or target face, and a bounding box associated with one or more other subjects in the target image.

Citation Information

Patent Citations

  • Image synthesis system, image synthesis method, and program of the method

    JP2006330800A

  • Image processor and image processing method

    JP2008198062A

  • Image compositing apparatus and program

    JP2010239440A

  • Information processor and recording medium

    JP2014123261A

  • Image synthesizing device, and image synthesizing method and program

    JP2014225870A