Generating images with head pose or facial region improvements
By using head and face editor technology in group photos, replacing and adjusting head and facial features in the image, the problem of poor expressions or postures of individuals during shooting is solved, and high-quality group photos are achieved.
Patent Information
- Application Number
- CN202480004275.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-03
- Filing Date
- 2024-10-03
- Publication Date
- 2025-06-10
AI Technical Summary
When taking group photos, it is difficult to ensure that everyone smiles and looks at the camera, especially when there are many people, some people have poor facial expressions or head postures, resulting in low quality of the photos taken.
By receiving an image collection of source and target images, determine whether to use a head editor or a face editor to generate a composite image. The head editor adjusts the head and torso areas in the target image by replacing the head pixels and interpolated areas, while the face editor adjusts the facial features in the target image based on the facial features in the source image.
Generate high-quality group photos, ensuring that everyone looks at the camera in the photo and has a good expression, avoiding unreal synthetic seams and artifacts that may appear in the photo.
Smart Images

Figure CN120129925A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application is a non - provisional application claiming priority under 35 U.S.C.§ 119(e) to U.S. Provisional Patent Application No. 63 / 542,283, filed on October 3, 2023, entitled "Generating a Group Photo with Head Pose and Facial Recognition Improvements", the content of which is hereby incorporated by reference in its entirety. Background Art
[0002] A group photo is a popular way to commemorate an event. It is difficult to obtain an image in which people in a group are all smiling and looking at the camera because the more people there are in the image, the greater the likelihood that at least one of the people's faces is not their best representation. For example, one person may have their mouth open, another may have their eyes closed, another may not be looking at the camera, and so on. Additionally, a person may tilt their head in a different way from the others in the picture, be at an angle to the camera, or otherwise be in a position that is not conducive to taking a high - quality photo.
[0003] The background description provided herein is for the purpose of generally presenting the context of the disclosure. The work of the currently named inventors (to the extent it is described in this background art section) and aspects of this specification that may not have been prior art at the time of filing are neither expressly nor impliedly admitted to be prior art to the disclosure. Summary of the Invention
[0004] A computer - implemented method includes: receiving an image set including a source image and a target image, the source image and the target image including at least one subject. The method further includes: determining, based on the image set, whether to use one or more editors selected from the group consisting of a head editor, a face editor, or a combination thereof. The method further includes: in response to determining to use the head editor, generating a synthetic image by: replacing at least a portion of the head pixels associated with a target head of a subject in the target image with head pixels of a source head of the subject from the source image; replacing the neck pixels associated with the target neck and the shoulder pixels associated with the target shoulders, including the area between the target head and the target torso, with an interpolated region generated from an interpolation of the source image and the target image.
[0005] In some embodiments, the method further comprises: in response to determining that a face editor is used, adjusting at least a portion of a target facial feature in a target image based on facial pixels of a source facial feature in a source image. In some embodiments, adjusting at least a portion of a target facial feature in a target image based on facial pixels of a source facial feature in a source image comprises: extracting a target head and a source head in an initial pose; aligning the target head to a canonical pose; encoding the aligned target head as a target vector and the source head as a source vector in a latent space; copying one or more components from the source vector to the target vector; rendering a modified target vector including one or more components from the encoded source head; realigning the rendered target head to the initial pose; and blending the realigned target head with the source image.
[0006] In some embodiments, determining that a face editor is used is based on an angular difference between a first angle of the target head and a second angle of the source head. In some embodiments, determining that a head editor is used is based on a bounding box surrounding the target head or the target face and a distance between the bounding box and one or more other bounding boxes associated with other subjects in the target image. In some embodiments, generating a synthetic image further comprises: in response to identifying remaining target pixels in the target image associated with the target head rather than the source head, repairing the remaining target pixels. In some embodiments, the method further comprises: determining an occlusion of the target head or an occlusion of the source head based on a difference in color histograms of the target image and the source image, wherein determining that a head editor is used is based on the occlusion of the target head or the occlusion of the source head.
[0007] In some embodiments, the method further comprises: before determining whether to use one or more editors based on an image set, the method further comprises: capturing an image set using a camera; providing a user interface to the user including the target image and an option to select a source head from a source image set, the source image set including the source image; and receiving a selection of the source image from the user. In some embodiments, at least one subject in the source image is a human or an animal.
[0008] A system includes: one or more processors; and one or more computer-readable storage media having instructions stored thereon that, when executed by the one or more processors, cause the one or more processors to perform operations. The operations include: receiving an image set including a source image and a target image, the source image and the target image including at least one subject; determining, based on the image set, whether to use one or more editors selected from the group consisting of a head editor, a face editor, or a combination thereof; and in response to determining to use the head editor, generating a synthetic image by: replacing at least a portion of the head pixels associated with the target head of the subject in the target image with head pixels of the source head of the subject from the source image; replacing the neck pixels associated with the target neck and the shoulder pixels associated with the target shoulders, including the region between the target head and the target torso, with an interpolated region generated from an interpolation of the source image and the target image.
[0009] In some embodiments, the operations further include: in response to determining to use the face editor, adjusting at least a portion of the target face features in the target image based on face pixels of the source face features from the source image. In some embodiments, adjusting at least a portion of the target face features in the target image based on face pixels of the source face features from the source image includes: extracting the target head and the source head in an initial pose; aligning the target head to a canonical pose; encoding the aligned target head as a target vector and the source head as a source vector in a latent space; copying one or more components from the source vector to the target vector; rendering the modified target vector including one or more components from the encoded source head; realigning the rendered target head to the initial pose; and blending the realigned target head with the source image. In some embodiments, determining to use the face editor is based on an angular difference between a first angle of the target head and a second angle of the source head. In some embodiments, determining to use the head editor is based on a bounding box surrounding the target head or the target face and a distance between the bounding box and bounding boxes associated with one or more other subjects in the target image. In some embodiments, generating the synthetic image further includes: in response to identifying remaining target pixels in the target image associated with the target head rather than the source head, repairing the remaining target pixels.
[0010] A non - transitory computer - readable medium storing instructions that, when executed by one or more processing devices, cause the one or more processing devices to perform operations. The operations include: receiving an image set including a source image and a target image, the source image and the target image including at least one subject; determining whether to use one or more editors selected from the group consisting of a head editor, a face editor, or a combination thereof based on the image set; and in response to determining to use the head editor, generating a synthetic image by: replacing at least a portion of the head pixels associated with the target head of the subject in the target image with head pixels from the source head of the subject in the source image; replacing the neck pixels associated with the target neck and the shoulder pixels associated with the target shoulders, including the region between the target head and the target torso, with an interpolated region generated from an interpolation of the source image and the target image.
[0011] In some embodiments, the operations further include: in response to determining to use the face editor, adjusting at least a portion of the target face features in the target image based on face pixels from the source face features in the source image. In some embodiments, adjusting at least a portion of the target face features in the target image based on face pixels from the source face features in the source image includes: extracting the target head and the source head in an initial pose; aligning the target head to a canonical pose; encoding the aligned target head as a target vector and the source head as a source vector in a latent space; copying one or more components from the source vector to the target vector; rendering the modified target vector including one or more components from the encoded source head; realigning the rendered target head to the initial pose; and blending the realigned target head with the source image. In some embodiments, determining to use the face editor is based on the angular difference between a first angle of the target head and a second angle of the source head. In some embodiments, determining to use the head editor is based on a bounding box surrounding the target head or target face and the distance between the bounding box and bounding boxes associated with one or more other subjects in the target image. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a block diagram of an example network environment in accordance with some embodiments described herein.
[0013] Figure 2 is a block diagram of an example computing device in accordance with some embodiments described herein.
[0014] Figure 3A is an example user interface including an option for a user to specify a target image from an image set in accordance with some embodiments described herein.
[0015] Figure 3Bis an example user interface according to some embodiments described herein that includes an option for a user to specify one or more source images from an image collection.
[0016] Figure 4A is an example target image for a head editor according to some embodiments described herein.
[0017] Figure 4B is an example composite image generated by a head editor according to some embodiments described herein.
[0018] Figure 5 is an example target image according to some embodiments described herein.
[0019] Figure 6A is an example first target head separated from a target image according to some embodiments described herein.
[0020] Figure 6B is an example second target head separated from a target image according to some embodiments described herein.
[0021] Figure 7A is an example first source head separated from a source image according to some embodiments described herein.
[0022] Figure 7B is an example second source head separated from a source image according to some embodiments described herein.
[0023] Figure 8A is an example first source head aligned with a first target head according to some embodiments described herein.
[0024] Figure 8B is an example second source head aligned with a second target head according to some embodiments described herein.
[0025] Figure 9 is an example composite image to be analyzed for repair according to some embodiments described herein.
[0026] Figure 10 is according to some embodiments described herein where a source head replaces Figure 5 the target head in an example composite image.
[0027] Figure 11 is an example block diagram of a machine learning model for generating a composite image according to some embodiments described herein.
[0028] Figure 12A is an example target image for a face editor according to some embodiments described herein.
[0029] Figure 12B An example synthetic image generated by a face editor according to some embodiments described herein.
[0030] Figure 13 Illustrates an example user interface for selecting a best take according to some embodiments described herein.
[0031] Figure 14 A flowchart illustrating an example method for generating a synthetic image according to some embodiments described herein. DETAILED DESCRIPTION OVERVIEW
[0032] A media application generates a synthetic image, where one or more of the subjects in the synthetic image have a head and / or face from a source image. Previous attempts to combine parts of images may result in unrealistic synthetic images where seams are visible, pixels from the original object are visible where the replacement object does not align with the original object, occluding objects cause artifacts, and so on.
[0033] In some embodiments, the media application receives a source image and a target image and determines whether to use a head editor and / or a face editor to generate the synthetic image. For example, the media application may select the head editor based on two subjects in the target image having heads that are separated far enough apart or the heads not being occluded by an object. In another example, the media application may select the face editor based on the pose of the face in the source image and the target image having a difference in angles that is close enough such that parts of the source image can be added to the target image.
[0034] The head editor may replace the head in the target image with the head in the source image. The face editor may adjust parts of the face in the target image, such as the eyes and mouth, based on parts of the face from the source image. For example, the face editor may use an embedding to calculate the pixel values of the parts of the face in the target image. The media application generates a synthetic image from the combination of the target image and the source image.
[0035] Generating the synthetic image may include analyzing the image data to determine occlusion and transforming the image data, and performing inpainting to avoid visible seams or other defects that may cause the synthetic image to look unrealistic. In other words, the techniques described herein for generating a synthetic image from one or more source images can seamlessly preserve the realism of the one or more source images. EXAMPLE ENVIRONMENT
[0036] Figure 1Block diagram of an exemplary environment 100. In some embodiments, environment 100 includes a media server 101, user devices 115a and 115n coupled to a network 105. Users 125a, 125n may be associated with the respective user devices 115a, 115n. In some embodiments, environment 100 may include Figure 1 other servers or devices not shown. In Figure 1 and the remaining figures, a letter following a reference numeral (e.g., "115a") refers to a reference to an element having that particular reference numeral. A reference numeral in the text without a following letter (e.g., "115") refers to a general reference to an embodiment of an element carrying that reference numeral.
[0037] The media server 101 may include a processor, a memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively coupled to the network 105 via a signal line 102. The signal line 102 may be a wired connection (such as Ethernet, coaxial cable, fiber optic cable, etc.) or a wireless connection (such as Wi-Fi®, Bluetooth®, or other wireless technologies). In some embodiments, the media server 101 sends data to and receives data from one or more of the user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.
[0038] The database 199 may store machine learning models, training data sets, images, etc. The database 199 may also store social network data associated with the user 125, user preferences of the user 125, etc.
[0039] The user device 115 may be a computing device including a memory coupled to a hardware processor. For example, the user device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device capable of accessing the network 105.
[0040] In the illustrated implementation, user device 115a is coupled to network 105 via signal line 108, and user device 115n is coupled to network 105 via signal line 110. Media application 103 may be stored on user device 115a as media application 103b and / or on user device 115n as media application 103c. Signal lines 108 and 110 may be wired connections (such as Ethernet, coaxial cable, fiber optic cable, etc.) or wireless connections (such as Wi-Fi®, Bluetooth® or other wireless technologies). User devices 115a, 115n are accessed by users 125a, 125n respectively. Figure 1 User devices 115a, 115n in Figure 1 are used only as examples. Although
[0041] two user devices, 115a and 115n, are illustrated, the present disclosure is applicable to system architectures having one or more user devices 115.
[0042] A machine learning model (e.g., a neural network or other type of model) is locally stored and utilized on user device 115 for one or more operations with specific user permissions. The server-side model is used only with user permission. Further, a trained model can be provided for use on user device 115. During such use, on-device training of the model can be performed if user 125 permits. If user 125 permits, the updated model parameters can be sent to media server 101, e.g., to enable federated learning. The model parameters do not include any user data.
[0043] Media application 103 receives an image set including a source image and a target image, where the source image and the target image include at least one subject. Media application 103 determines whether to use a head editor and / or a face editor based on the image set. In response to determining to use the head editor, media application 103 generates a synthetic image by: replacing at least a portion of the head pixels associated with the target head of the subject in the target image with head pixels of the source head of the subject from the source image; and replacing the neck pixels associated with the target neck and the shoulder pixels associated with the target shoulder, which include the region between the target head and the target torso, with an interpolated region generated from the source image and the target image.
[0044] In some embodiments, media application 103 can be implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, media application 103a can be implemented using a combination of hardware and software. Example computing device
[0045] Figure 2 is a block diagram of an example computing device 200 that can be used to implement one or more features described herein. Computing device 200 can be any suitable computer system, server, or other electronic or hardware device. In one example, computing device 200 is media server 101 for implementing media application 103a. In another example, computing device 200 is user device 115.
[0046] In some embodiments, computing device 200 includes a processor 235, a memory 237, an input / output (I / O) interface 239, a display 241, a camera 243, and a storage device 245, all coupled via a bus 218. The processor 235 may be coupled to the bus 218 via a signal line 222, the memory 237 may be coupled to the bus 218 via a signal line 224, the I / O interface 239 may be coupled to the bus 218 via a signal line 226, the display 241 may be coupled to the bus 218 via a signal line 228, the camera 243 may be coupled to the bus 218 via a signal line 230, and the storage device 245 may be coupled to the bus 218 via a signal line 232.
[0047] The processor 235 may be one or more processors and / or processing circuits configured to execute program code and control the basic operations of the computing device 200. A "processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. The processor may include a system having: a general-purpose central processing unit (CPU) with one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multi-processor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), a dedicated circuit system for implementing functionality, a dedicated processor for implementing processing based on a neural network model, a neural circuit, a processor optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, the processor 235 may include one or more coprocessors for implementing neural network processing. In some embodiments, the processor 235 may be a processor that processes data to produce a probabilistic output; for example, the output produced by the processor 235 may be imprecise, or may be accurate within a range from the expected output. Processing need not be limited to a particular geographical location or have a time limit. For example, the processor may perform its functions in real time, offline, in batch mode, etc. Portions of the processing may be performed at different times and in different locations by different (or the same) processing systems. A computer may be any processor that communicates with a memory.
[0048] The memory 237 is typically provided in the computing device 200 for access by the processor 235 and may be any suitable processor-readable storage medium suitable for storing instructions for execution by the processor or set of processors and located separately from and / or integrated with the processor 235, such as random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc. The memory 237 may store software operated by the processor 235 on the computing device 200, which includes media application 103.
[0049] The memory 237 may include an operating system 262, other applications 264, and application data 266. The other applications 264 may include, for example, an image gallery application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, etc. One or more methods disclosed herein may operate in several environments and platforms, such as a stand-alone computer program that can run on any type of computing device, a web application with web pages, a mobile application (“app”) that runs on a mobile computing device, and the like.
[0050] The application data 266 may be data generated by the other applications 264 or the hardware of the computing device 200. For example, the application data 266 may include images used by an image gallery application and user actions identified by other applications 264 (e.g., a social networking application).
[0051] The I / O interface 239 may provide functions that enable the computing device 200 to interface with other systems and devices. The interfaced devices may be included as part of the computing device 200 or may be separate and communicate with the computing device 200. For example, a network communication device, a storage device (e.g., the memory 237 and / or the storage device 245), and input / output devices may communicate via the I / O interface 239. In some embodiments, the I / O interface 239 may be connected to interface devices, such as input devices (keyboard, pointing device, touch screen, microphone, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, monitor, etc.).
[0052] Some examples of the interfaced devices that may be connected to the I / O interface 239 may include a display 241, which may be used to display content (e.g., images, videos, and / or user interfaces of output applications as described herein) and receive touch (or gesture) input from a user. For example, the display 241 may be used to display a user interface including graphical guidance on a viewfinder. The display 241 may include any suitable display device, such as a liquid crystal display (LCD), a light emitting diode (LED), or a plasma display panel, a cathode ray tube (CRT), a television, a monitor, a touch screen, a three-dimensional display, or other visual display devices. For example, the display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in a glasses-shaped or head-mounted device, or a monitor screen of a computer device.
[0053] The camera 243 may be any type of image capture device capable of capturing images and / or videos. In some embodiments, the camera 243 captures an image or video, and the I / O interface 239 sends the image or video to the media application 103.
[0054] The storage device 245 stores data related to the media application 103. For example, the storage device 245 may store a training data set including tagged images, a machine learning model, an output from the machine learning model, and the like.
[0055] Figure 2 Illustrative example media application 103 stored in the memory 237, the example media application including an image module 202, a header editor 204, and a face editor 206.
[0056] The image module 202 generates graphic data for displaying a user interface including an image set. The image set may be received from the camera 243 of the computing device 200 and / or from the media server 101 via the I / O interface 239. For example, the image set may include a burst of images from the images captured by the camera 243. A burst of images may include a plurality of photos captured rapidly and continuously within a short time period. The image set includes one or more source images and target images including one or more subjects. The one or more subjects may be human, animal, etc.
[0057] The image module 202 obtains permission from the user to modify any of the images in the image set. Controls may be provided to the user that allow the user to make a choice regarding whether and when the systems, programs, or features described herein may enable the collection or use of user information (e.g., identification of the user in the image, user preferences, or the user's current location), and whether to send content or communications to the user from the server. Additionally, before storing or using certain data, it may be processed in one or more ways such that personally identifiable information is removed. For example, the identity of the user may be processed such that the user's personally identifiable information cannot be determined, or the user's geographical location may be generalized (such as to the city, zip code, or state level) if location information is obtained such that the user's specific location cannot be determined. Thus, the user can control what information about the user is collected, how that information is used, and what information is provided to the user.
[0058] In some embodiments, the image module 202 selects one or more source images and a target image for generating a synthetic image. For example, the image module 202 can automatically generate a composite score for each image in an image set based on composite factors such as the type of object in the image, the positioning of the subject, lighting, etc. For example, the target image can be selected based on having the best landscape composition and the subject in the group being well positioned. In another example, the target image can be selected based on having the largest number of subjects looking at the camera, smiling, eyes open, etc. Images with an overall high composite score can be used as the target image. In some embodiments, the user can select a specific image in the image set as the target image, and one or more other images in the image set can be presented as source images.
[0059] The image module 202 can select one or more source images based on the subjects in the image. For example, the image module 202 can generate a face score for each subject in the image based on a quality metric. For example, each face can be scored based on whether the face has a smile, open eyes, is completely or partially occluded by another person or object in the photo, and / or whether there is no movement (e.g., is not blurry). The image module 202 can generate a head score based on the positioning of the subject's head, such as generating a higher head score for a head in a vertical plane (e.g., directly facing the camera) rather than at other angles (e.g., tilted, rotated away from the camera, etc.). This scoring can be performed using image detection and recognition techniques (such as a trained machine learning model that can detect faces within a photo and apply quality scoring criteria).
[0060] Based on the scores, the image module 202 can rank the faces in the image. For example, the image module 202 can determine that multiple specific faces in different images in the image set belong to the same subject (person, animal, etc.), and each such subject can be associated with a ranked list of faces from the images. The image module 202 can automatically select the target image and one or more source images, or the image module 202 can suggest the top-ranked images to the user as a recommendation for the user to confirm the selection of a specific image as the target image.
[0061] In some embodiments, the image module 202 generates graphical data for displaying a user interface that provides options for the user to select a target image and one or more source images from an image set. The image set included in the user interface can be from a burst of captured images, images captured within a specific time period (e.g., the past 24 hours, images captured at a specific location, etc.).
[0062] Figure 3AAn example user interface 300 includes an option for a user to specify a target image from an image collection. In this example, the user selects image 305 to be used as the target image. In some embodiments, a particular image from the image collection may be suggested as the target image, e.g., highlighted (e.g., with a border, an icon, etc.). The user may select the suggested image or any other image as the target image.
[0063] Figure 3B An example user interface 325 includes an option for a user to specify one or more source images from an image collection. In this example, the user selects a first image 330 to be used as the source image for a first subject and a second image 335 to be used as the source image for a second subject. In some embodiments, the image module 202 may generate a user interface that includes an option for selecting a particular subject in an image. For example, the computing device 200 may receive user input in the form of a double-tap on a subject, a circle around the head of a subject, etc.
[0064] In some embodiments, a single image may be selected as the source image for two or more subjects. In some embodiments, each subject may be associated with a different source image. In some embodiments, suggestions for source images may be provided for one or more subjects. For example, if the target image has a first subject with a tilted head and a second subject with closed eyes, a first source image in which the first subject has a head facing the camera without tilting and a second source image in which the second subject has open eyes may be suggested as the corresponding source images. In some embodiments, other factors such as lighting, the duration between the capture of the target image and a particular source image, the distance between the positioning of the source image and the subject in the target image, etc. may be used to select the particular source image to be recommended.
[0065] In some embodiments, the user interface may include an option to search for another source image. For example, the user interface may include an option for the user to scroll through (or otherwise browse) the images in the camera roll and select a source image. In some embodiments, since the source images from the camera roll may have been captured under different lighting conditions, different shadows, etc., the image module 202 may modify the source images to have colors (and / or other image attributes such as brightness, white balance, contrast, etc.) that are consistent with the corresponding attributes of the target image.
[0066] The image module 202 determines whether to use the head editor 204, the face editor 206, or both to generate a composite image. The image module 202 may use different criteria to make the determination, such as the proximity of the head within the target image or source image, the angle of the head within the target image or source image, and the occlusion of the head in the target image or source image.
[0067] In some embodiments, the image module 202 determines whether to use the head editor 204 or the face editor 206 based on different factors evaluated by different classifiers. In some embodiments, the image module 202 performs a weighted sum or a logistic regression of the different factors such that the determination is made based on the totality of the factors rather than a single factor. In some embodiments, some of the factors can be decisive, such as if one of the heads is occluded by an object by 70%.
[0068] In some embodiments, the classifier is trained based on the ratings of the head edits from the head editor 204 and the face edits from the face editor 206 in the training set of images. For example, the training data can include target images, source images, synthetic images generated by the head editor 204 and / or the face editor 206, and corresponding ratings of the synthetic images, where the corresponding ratings can be provided by a human or a quality algorithm. As a result of receiving the ratings, the image module 202 can recalculate the classifier weights to improve the quality of the synthetic images generated by the head editor 204 and the face editor 206.
[0069] In some embodiments, the image module 202 generates bounding boxes around each target head / target face in the target image and determines whether to use the head editor 204 to replace the target head with the source head based on the distances between the bounding boxes in the target image. For example, if there is an overlap between the bounding boxes, which indicates that the heads of the subjects are close together, then when the close proximity of multiple subjects results in a larger section of the background to be repaired in the synthetic image, it may not be possible to use the head editor 204 due to the change in head positioning, and the likelihood of overlap among the subjects complicates the alignment of the subjects in the target image and one or more source images. In some embodiments, the image module 202 determines whether to use the head editor 204 based on a continuous function. In some embodiments, the continuous function is based on the weighted distance between the bounding boxes and / or the weighted percentage of overlap between the bounding boxes, where the values of the weights can be learned during the training of the classifier used by the image module 202.
[0070] In some embodiments, the image module 202 determines whether to use the head editor 204 based on the angular difference between a first angle of the target head and a second angle of the source head. If the angle of the head in the source image is greater than a threshold difference from the angle of the head in the target image, the face editor 206 may generate an unsatisfactory or low-quality synthetic image.
[0071] The image module 202 can determine to use the face editor 206 based on applying a non-linear function to the difference between the first angle and the second angle. In some embodiments, the image module 202 applies a cosine to the first angle, applies a cosine to the second angle, and determines the difference between the first angle and the second angle. In some embodiments, the image module 202 uses a threshold angular difference to determine whether to use the face editor 206 such that if the difference between the angles exceeds the threshold angular difference, the image module 202 determines to use the head editor 204 instead of the face editor 206.
[0072] In some embodiments, the image module 202 determines whether to use the head editor 204 based on occlusion of the target head or occlusion of the source head. When the image module 202 receives an occluded target image or source image, the resulting composite image has a higher failure rate and / or lower quality when the source image is occluded. In some embodiments, the image module 202 can use the color histogram of the image to determine whether the source head or the target head is occluded. For example, the image module 202 can use the average value of a specific channel in the histogram to identify the occlusion.
[0073] In some embodiments, the image module 202 determines whether to use the head editor 204, the face editor 206, or both the head editor 204 and the face editor 206 based on the distance from the subject to the image boundary. For example, if the head is close to the image boundary and the head editor 204 changes the pose of the head close to the image boundary, it may cause a part of the head to be cropped out of the image boundary (thereby providing an unsatisfactory composite image because the subject's head is not entirely within the image).
[0074] The head editor 204 replaces the target head with the source head. In some embodiments, the head editor 204 replaces the target pixels associated with the target head in the target image with the source pixels of the source head from the source image. The head editor 204 can generate a head mask including the subject's hair, segment the head mask from the source image, and apply the pixels within the head mask to the target image. The head editor 204 can modify the positioning and scale of the source head to match the size of the target head.
[0075] In the case where replacing the target head with the source head results in portions where the target head and the source head do not overlap, the head editor 204 performs a repair. This can occur when the target head and the source head are associated with different angles. This can also occur when the angles of the target head and the source head in the target image and the source image are different and the head editor 204 aligns the angle of the target head with the angle of the source head before replacement. Aligning the target head with the source head can identify portions of the target head where the source head does not overlap. The head editor 204 can identify the remaining pixels in the target image that are associated with the target head rather than the source head, and perform a repair of the remaining target pixels by replacing the remaining target pixels. Repairing the pixels can include determining the distance between the source pixels and the remaining pixels, and generating replacement pixels based on the similarity and distance to the source pixels.
[0076] Because the target head and the source head are at different angles, the neck region of the body can be positioned differently in the two images. The head editor 204 renders a smooth transition between the target torso and the source head by generating an interpolated region of the neck and shoulder regions. In some embodiments, the head editor 204 generates an interpolated region that includes the region between the target head and the target torso (e.g., the region described as the target neck and target shoulders), and replaces the target pixels associated with the target neck and target shoulders with the interpolated region, where the interpolated region is an interpolation of the target pixels and the source pixels of the corresponding neck and shoulder regions. In some embodiments, the head editor 204 uses multiple image frames (such as a set of image frames generated on the fly from images captured by the camera 243), and generates the interpolated region from the multiple image frames.
[0077] Go to Figure 4A , an example target image 400 for a head editor according to some embodiments described herein. The subject 405 on the right is looking up and has smooth shoulders.
[0078] Figure 4B is an example composite image 425 generated by the head editor 204. The head editor 204 replaces the target head with the source head for the subject 430 on the right. In the composite image 425, the interpolated region 430 for the neck and the top of the shoulders is the result of the neck being in a different orientation and the shoulders including more wrinkled fabric compared to the shoulders on the subject 405 in the target image 400.
[0079] In some embodiments, the head editor 204 includes a machine learning model that receives a target image and one or more source images as inputs and outputs a synthesized image. The trained machine learning model can include one or more model forms or architectures. For example, the model form or architecture can include any type of neural network, such as a linear network, a deep learning neural network that implements multiple layers (e.g., "hidden layers" between an input layer and an output layer, where each layer is a linear network), a convolutional neural network (e.g., a network that splits or partitions input data into multiple parts or tiles, uses one or more neural network layers to process each tile individually, and aggregates the results of the processing from each tile), a sequence-to-sequence neural network (e.g., a network that receives sequential data, such as words in a sentence, frames in a video, etc., as inputs and produces a result sequence as output), and so on.
[0080] The model form or architecture can specify the connectivity between various nodes and the organization of nodes into layers. For example, the nodes of the first layer (e.g., the input layer) can receive data as input data or applied data. For example, when the trained model is used for analysis (e.g., of an initial image), such data can include, for example, one or more pixels per node. Subsequent intermediate layers can receive the outputs of the nodes of the previous layer as inputs according to the connectivity specified in the model form or architecture. These layers can also be referred to as hidden layers. The final layer (e.g., the output layer) produces the output of the machine learning model. For example, the output layer can output a synthesized image. In some embodiments, the model form or architecture also specifies the number and / or type of nodes in each layer.
[0081] In various embodiments, the trained model may include one or more models. One or more of the models may include a plurality of nodes arranged in layers according to a model structure or form. In some embodiments, the nodes may be computational nodes without memory, which are configured, for example, to process one unit of input to produce one unit of output. The computation performed by a node may include, for example, multiplying each node input among a plurality of node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum using a bias or intercept value to produce a node output. In some embodiments, the computation performed by a node may further include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a non-linear function. In various embodiments, such computation may include operations such as matrix multiplication. In some embodiments, the computation performed by the plurality of nodes may be executed in parallel, for example, using multiple processor cores of a multi-core processor, individual processing units of a graphics processing unit (GPU), or dedicated neural circuitry. In some embodiments, the nodes may include memory, for example, may be capable of storing one or more earlier inputs and using the one or more earlier inputs to process subsequent inputs. For example, a node with memory may include a long short-term memory (LSTM) node. The LSTM node may use the memory to maintain a "state" that permits the node to act like a finite state machine (FSM).
[0082] In some embodiments, the trained model may include embeddings or weights for individual nodes. For example, the model may be initiated as a plurality of nodes organized into layers as specified by the model form or structure. At initialization, corresponding weights may be applied to the connections between each pair of nodes (e.g., nodes in consecutive layers of a neural network) connected according to the model form. For example, the corresponding weights may be randomly assigned or initialized to default values. Then, the model may be trained (e.g., using training data) to produce results.
[0083] Training may include applying supervised learning techniques. In supervised learning, the training data may include a plurality of inputs (e.g., target images and source images, segmentation masks, etc.) and corresponding groundtruth outputs for each input (e.g., a groundtruth mask correctly identifying a part of a subject (such as the face of the subject) in each image, synthetic image, etc.). Based on the comparison of the model's output with the groundtruth output, the values of the weights are automatically adjusted, for example, in a manner that increases the probability that the model produces a groundtruth output for a synthetic image.
[0084] In various embodiments, a trained model includes a set of weights or embeddings corresponding to the model architecture. In some embodiments, the trained model can include, for example, a fixed set of weights downloaded from a server providing the weights. In various embodiments, a trained model includes a set of weights or embeddings corresponding to the model architecture. In embodiments where data is omitted, the head editor 204 can generate a trained model that is based on, for example, prior training performed by the developer of the head editor 204, by a third party, etc. In some embodiments, the trained model can include, for example, a fixed set of weights downloaded from a server providing the weights.
[0085] In some embodiments, a trained machine learning model receives a target image having one or more subjects and one or more source images. The machine learning model can generate one or more segmentation masks that identify pixels in the one or more source images corresponding to one or more heads including hair in the one or more source images. For each subject, the machine learning model replaces the head pixels from the target image with head pixels from the one or more source images at corresponding locations, scales, and angles that match the source images. The machine learning model blends the head pixels along the edges of the head. In some embodiments, background pixels that are revealed due to differences between the source head and the target head are repaired. In some embodiments, the machine learning model blends interpolated regions into the target image corresponding to the neck and shoulders of the subject. In some embodiments, the trained machine learning model outputs a composite image incorporating these changes.
[0086] In some embodiments, the machine learning model outputs a confidence value for each composite image output by the trained machine learning model. The confidence value can be expressed as a percentage, a number from 0 to 1, etc. For example, the machine learning model outputs a confidence value of 85% for a composite image in which the source head correctly replaces the target head and does not include pixels from another person or object. In some embodiments, if the confidence value exceeds a threshold confidence value, the composite image is provided to the user. In some embodiments, the confidence value is pre - calculated and the composite image is not generated unless the confidence value exceeds the threshold confidence value.
[0087] In some embodiments, the head editor 204 includes multiple machine learning models that perform different functions in the steps for generating the composite image. For example, a first machine learning model can replace target pixels associated with the target head with source pixels from the source head, a second machine learning model aligns the source head with the target head, and a third machine learning model generates an interpolated region that replaces target pixels associated with the target neck and target shoulders.
[0088] Figure 5This is the example target image 500. The head 505 of the boy and the head 510 of the girl in the target image 500 will be replaced by heads from the source image.
[0089] Figure 6A is the example first target head 600 separated from Figure 5 the target image 500. Figure 6B is the example second target head 650 separated from Figure 5 the target image 500. The first target head 600 can be separated and selected for replacement because the boy's nose is upturned and his mouth is open. The second target head 650 can be separated and selected for replacement because the girl is not looking directly at the camera.
[0090] Figure 7A is the example first source head 700 separated from the source image. Figure 7B is the example second source head 750 separated from the same source image including the first source head 700 or from a different source image. The head editor 204 replaces at least a portion of the target pixels associated with the first target head 600 and at least a portion of the target pixels associated with the second target head 650 with source pixels from the first source head 700 and the second source head 750, respectively.
[0091] Figure 8A is the example 800 first source head aligned with the first target head. Figure 8B is the example 850 second source head aligned with the second target head. Although the images are aligned, there are still some places where the images do not overlap, such as the positioning of the boy's hat 805.
[0092] Figure 9 is the example composite image 900, which is analyzed for repair based on the alignment of the first source head exposing the area of the background previously covered by the first target head. After the head editor 204 performs the repair of the target head, the composite image has an appearance such that the source head fits into the composite image without visible seams or other defects.
[0093] Figure 10 is the example composite image 1000 in which the source head has replaced the target head. Figure 5 The head 505 of the boy in Figure 10 is replaced by the head 1005 of the boy in Figure 5 and the head 510 of the girl in Figure 10 is replaced by the head 1010 of the girl in
[0094] The face editor 206 transfers facial features (e.g., smile, open eyes, mouth shape, etc.) from a source image to a target image. The face editor 206 does not change the pose of the target face. In some embodiments, the face editor 206 adjusts at least a portion of the target pixel values associated with the target facial features in the target image based on source pixels from the source facial features in the source image.
[0095] In some embodiments, the face editor 206 uses a face matching algorithm implemented under a specific user license to extract the target face and the source face of the same subject. The face editor 206 aligns the target face from an initial pose to a canonical pose, where the canonical pose is one of the poses that the machine learning model is trained to use.
[0096] In some embodiments, the face editor 206 uses a machine learning model selected from one of the model types described above with reference to the head editor 204. For example, the machine learning model can include an encoder and a convolutional neural network (CNN).
[0097] In some embodiments, not all of the target face is replaced with the source because some components of the face may be more susceptible to small changes in angle and pose. For example, if the nose from an upward-tilted head is added to a face looking straight ahead, the person's nose may look inappropriate. Thus, the encoder of the machine learning model uses a facial feature mask (e.g., an eye / mouth mask) to encode the components of the source head as a source vector. The encoder also encodes the target head as a target vector. The machine learning model copies the corresponding components from the source vector to the target vector.
[0098] In some embodiments, the CNN includes multiple layers that each generate a version of the synthetic image by rendering the target vector including the source vector components. The CNN generates synthetic images with increasingly high resolution levels in each layer. The high-resolution output is blended with the existing low-resolution layers until the final synthetic image is output. In some embodiments, the machine learning model realigns the rendered target head back to the initial pose and blends the realigned target head with the source image, which is output as the synthetic image.
[0099] Figure 11 is an example block diagram 1100 of a machine learning model on a device for generating a synthetic image. The source image 1105 and the target image 1110 are processed by an offline module 1115 and an on-device module 1120 (such as Figure 1Both the media application 103) stored on the user device 115 as illustrated are received. The on-device module 1120 outputs a predicted image 1130 such that the facial expression and / or head pose from the source image 1105 are blended into the target image 1110. The on-device module 1120 is implemented such that the on-device module can operate on a user device (e.g., a smartphone with or without machine learning circuitry) within the budget of latency and memory usage.
[0100] The offline module 1115 can provide supervision for on-device model training, which advantageously enables the lightweight on-device module 1120 to generate a synthetic image 1125. The synthetic image 1125 can be a higher quality version compared to the predicted image 1130 because the offline module 1115 has a larger budget for latency and memory usage.
[0101] In some embodiments, the on-device module 1120 can use a warp and recovery where a dense facial mesh warps the source image 1105 to the target location and uses a neural network (or other suitable techniques) to remove artifacts. Given two facial images and their facial meshes, the on-device module 1120 aligns the pose of the source face to the target facial mesh and then warps the source face to the target image. The warped face has the source expression but has artifacts caused by the alignment and warping, which can be regarded as degradation. The on-device module 1120 can include a recovery model with supervision from the offline module 1115, which can remove or remedy the degradation.
[0102] In some embodiments, the on-device module 1120 can modify the face to reflect a specified expression. The on-device module 1120 can include two encoders: a source encoder for extracting an expression code from the source image; and a target encoder for extracting head pose and appearance information from the target image. The on-device module 1120 can insert the expression code at one or more levels (e.g., at each level) of the decoder. Skip connections from the target encoder can also be added to ensure that the on-device module 1120 edits the face rather than other parts of the image.
[0103] Figure 12A Is an example target image 1200 where the second subject 1205 has an expression replaced by the face editor 206.
[0104] Figure 12B Is an example synthetic image 1250 generated by the face editor 206 according to the techniques described above. As can be seen, the open-mouthed smile of the right subject 1205 in the target image 1200 is replaced by the closed-mouth smile of the subject 1255 in the synthetic image 1250.
[0105] Figure 13Illustrate example user interfaces 1300, 1350, 1375 for selecting the best shot according to some embodiments described herein. The first user interface 1300 includes a target image 1305 of two subjects 1306, 1307. The user can select the "Best Shot" button 1310 to start the process for obtaining a composite image.
[0106] The second user interface 1350 includes a target image 1355 and icons 1357, 1359 of a first subject and a second subject respectively. The user can select one of the icons 1357, 1359 to select a different face for the subject. In this example, the user has selected the icon 1359 of the second subject. The user interface 1350 includes three options 1361, 1363, and 1365 for the second subject. The first option 1361 and the third option 1365 are from the source images, while the second option 1363 is from the target image 1355. Once the user selects one of the three options 1361, 1363, and 1365, the user can select the Done button 1367 to view the composite image of the selected option on the target image 1355, or select the Reset button 1369 to restart the process. In this example, the user selects the first option 1361 and selects the Done button 1367.
[0107] The third user interface 1375 includes a composite image 1377 that is generated from the target image 1355 of the second user interface 1350 and the source face 1379. The user can press the Save Copy button 1385 to save a copy of the image. Example flowchart
[0108] Figure 14 A flowchart illustrating an example method 1400 for segmenting an image according to some embodiments described herein. The method 1400 can be executed by Figure 2 the computing device 200 in. In some embodiments, the method 1400 is executed by the user device 115, the media server 101, or partially on the user device 115 and partially on the media server 101.
[0109] Figure 14 The method 1400 can start at block 1402. At block 1402, an image set including a source image and a target image is received, the source image and the target image including one or more subjects. In some embodiments, before block 1402, the method 1400 further includes: capturing the image set using a camera; providing a user interface to the user including the target image and an option for selecting a source head from a set of source images, the set of source images including the source image; and receiving a selection of the source image from the user. The subject can include a human or an animal. Block 1402 can be followed by block 1404.
[0110] At block 1404, it is determined whether permission to modify the source image and the target image is obtained. If permission is not obtained, block 1404 may be followed by block 1406, where method 1400 ends. If permission is obtained, block 1404 may be followed by block 1408.
[0111] At block 1408, it is determined whether to use one or more editors selected from the group consisting of a head editor, a face editor, or a combination thereof. Determining to use the face editor may be based on: the angular difference between a first angle of the target head and a second angle of the source head; the bounding box surrounding the target head or the target face and the distance between the bounding box and the bounding boxes associated with other subjects in the target image; or the occlusion of the target head or the occlusion of the source head. Determining the occlusion of the target head or the occlusion of the source head may be based on determining the difference in color histograms.
[0112] If the head editor is selected, block 1408 may be followed by block 1410. At block 1410, a synthetic image is generated by: replacing at least a portion of the head pixels associated with the target head of the subject in the target image with head pixels of the source head of the subject from the source image; and replacing the neck pixels associated with the target neck and the shoulder pixels associated with the target shoulders, including the region between the target head and the target torso, with an interpolated region generated from an interpolation of the source image and the target image. In some embodiments, the head editor aligns the angle of the target head with the angle of the source head before replacing at least a portion of the target pixels associated with the target head in the target image. If the face editor is selected, block 1408 may be followed by block 1412.
[0113] At block 1412, a synthetic image is generated by adjusting at least a portion of the target facial features in the target image based on facial pixels from the source facial features in the source image. In some embodiments, adjusting at least a portion of the target facial features in the target image based on facial pixels from the source facial features in the source image includes: extracting the target head and the source head in an initial pose; aligning the target head to a canonical pose; encoding the aligned target head as a target vector and the source head as a source vector in a latent space; copying one or more components from the source vector to the target vector; rendering the modified target vector including one or more components from the encoded source head; realigning the rendered target head to the initial pose; and blending the realigned target head with the source image.
[0114] In the foregoing description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present specification. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In some instances, structures and devices are shown in block diagram form in order to avoid obscuring the description. For example, the above may primarily describe embodiments with reference to a user interface and specific hardware. However, the embodiments may be applied to any type of computing device that can receive data and commands, as well as any peripheral device that provides services.
[0115] References in this specification to "some embodiments" or "some examples" mean that a particular feature, structure, or characteristic described in connection with the embodiments or examples may be included in at least one implementation of the description. The phrase "in some embodiments" appearing in various places in this specification does not necessarily all refer to the same embodiment.
[0116] Some portions of the detailed descriptions above are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulation of physical quantities. Although not necessarily, these quantities typically take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, etc.
[0117] However, it should be borne in mind that all of these and similar terms are to be associated with appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise, as will be apparent from the following discussion, it should be understood that throughout the description, discussions using terms including "processing" or "computing" or "calculating" or "determining" or "displaying" etc. refer to the actions and processes of a computer system, or similar electronic computing device, that manipulates data represented as physical (electronic) quantities within the registers and memories of the computer system and transforms that data into other data similarly represented as physical quantities within the memories or registers or other such information storage, transmission, or display devices of the computer system.
[0118] Embodiments of this specification may also relate to a processor for performing one or more steps of the above method. The processor may be a dedicated processor selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, which includes but is not limited to: any type of disk (including optical disks), ROM, CD-ROM, magnetic disks, RAM, EPROM, EEPROM, magnetic cards or optical cards, flash memory (including USB keys with non-volatile memory), or any type of medium suitable for storing electronic instructions, each medium being coupled to a computer system bus.
[0119] This specification may take the form of some fully hardware embodiments, some fully software embodiments, or some embodiments that include both hardware elements and software elements. In some embodiments, this specification is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.
[0120] In addition, this specification may take the form of a computer program product, which can be accessed from a computer-usable or computer-readable medium that provides program code for use by or in conjunction with a computer or any instruction execution system. For the purposes of this description, a computer-usable or computer-readable medium can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0121] A data processing system suitable for storing or executing program code will include at least one processor directly or indirectly coupled to a memory element through a system bus. The memory element may include local memory, mass storage devices, and cache memory employed during the actual execution of the program code, and the cache memory provides temporary storage of at least some of the program code to reduce the number of times code must be retrieved from the mass storage device during execution.
Claims
1. A computer-implemented method comprising: receiving an image set including a source image and a target image, the source image and the target image including at least one subject; determining whether to use one or more editors selected from the group of a head editor, a face editor, or a combination thereof based on the set of images; and In response to determining to use the header editor, generating a composite image by: replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from a source head of the subject in the source image; as well as Neck pixels associated with a target neck and shoulder pixels associated with a target shoulder are replaced with an interpolated region generated from interpolation of the source image and the target image.
2. The method of claim 1, further comprising: In response to determining to use the facial editor, at least a portion of a target facial feature in the target image is adjusted based on facial pixels from a source facial feature in the source image.
3. The method of claim 2, wherein adjusting at least a portion of a target facial feature in the target image based on facial pixels from the source facial feature in the source image comprises: extracting the target head and the source head in an initial posture; aligning the target head to a standard posture; encoding the aligned target head as a target vector and encoding the source head as a source vector in latent space; copying one or more components from the source vector to the destination vector; rendering a modified target vector including the one or more components from the encoded source header; realigning the rendered target head to the initial pose; as well as The realigned target head is blended with the source image.
4. The method of claim 2, wherein determining to use the facial editor is based on an angular difference between a first angle of the target head and a second angle of the source head.
5. The method of claim 1, wherein determining to use the head editor is based on a bounding box surrounding the target head or target face and a distance between the bounding box and bounding boxes associated with one or more other subjects in the target image.
6. The method of claim 1, wherein generating the composite image further comprises: In response to identifying remaining target pixels in the target image that are associated with the target header but not the source header, the remaining target pixels are inpainted.
7. The method of claim 1, further comprising: Determining the occlusion of the target head or the occlusion of the source head based on determining a difference in color histograms of the target image and the source image; Wherein determining to use the head editor is based on the occlusion of the target head or the occlusion of the source head.
8. The method of claim 1, before determining whether to use the one or more editors based on the image collection, the method further comprising: capturing the set of images using a camera; as well as providing a user with a user interface including the target image and an option to select the source header from a set of source images, the set of source images including the source image; and A selection of the source image is received from the user.
9. The method of claim 1, wherein the at least one subject in the source image is a human or an animal.
10. A system comprising: one or more processors; as well as One or more computer-readable media having stored thereon instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: receiving an image set including a source image and a target image, the source image and the target image including at least one subject; determining whether to use one or more editors selected from the group of a head editor, a face editor, or a combination thereof based on the set of images; and In response to determining to use the header editor, generating a composite image by: replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from a source head of the subject in the source image; and Neck pixels associated with a target neck and shoulder pixels associated with a target shoulder are replaced with an interpolated region generated from interpolation of the source image and the target image.
11. The system of claim 10, wherein the operations further comprise: In response to determining to use the facial editor, at least a portion of a target facial feature in the target image is adjusted based on facial pixels from a source facial feature in the source image.
12. The system of claim 11, wherein adjusting at least a portion of a target facial feature in the target image based on facial pixels from the source facial feature in the source image comprises: extracting the target head and the source head in an initial posture; aligning the target head to a standard posture; encoding the aligned target head as a target vector and encoding the source head as a source vector in latent space; copying one or more components from the source vector to the destination vector; rendering a modified target vector including the one or more components from the encoded source header; realigning the rendered target head to the initial pose; as well as The realigned target head is blended with the source image.
13. The system of claim 11, wherein determining to use the facial editor is based on an angular difference between a first angle of the target head and a second angle of the source head.
14. The system of claim 10, wherein determining to use the head editor is based on a bounding box surrounding the target head or target face and a distance between the bounding box and bounding boxes associated with one or more other subjects in the target image.
15. The system of claim 10, wherein generating the composite image further comprises: In response to identifying remaining target pixels in the target image that are associated with the target header but not the source header, the remaining target pixels are inpainted.
16. A non-transitory computer-readable medium having stored thereon instructions that, in response to being executed by one or more processing devices, cause the one or more processing devices to perform operations comprising: receiving an image set including a source image and a target image, the source image and the target image including at least one subject; determining whether to use one or more editors selected from the group of a head editor, a face editor, or a combination thereof based on the set of images; and In response to determining to use the header editor, generating a composite image by: replacing at least a portion of head pixels associated with a target head of the subject in the target image with head pixels from a source head of the subject in the source image; as well as Neck pixels associated with a target neck and shoulder pixels associated with a target shoulder are replaced with an interpolated region generated from interpolation of the source image and the target image.
17. The non-transitory computer readable medium of claim 16, wherein the operations further comprise: In response to determining to use the facial editor, at least a portion of a target facial feature in the target image is adjusted based on facial pixels from a source facial feature in the source image.
18. The non-transitory computer-readable medium of claim 17, wherein adjusting at least a portion of a target facial feature in the target image using facial pixels from the source facial feature in the source image based on the target facial feature in the target image comprises: extracting the target head and the source head in an initial posture; aligning the target head to a standard posture; encoding the aligned target head as a target vector and encoding the source head as a source vector in latent space; copying one or more components from the source vector to the destination vector; rendering a modified target vector including the one or more components from the encoded source header; realigning the rendered target head to the initial pose; as well as The realigned target head is blended with the source image.
19. The non-transitory computer readable medium of claim 17, wherein determining to use the facial editor is based on an angular difference between a first angle of the target head and a second angle of the source head.
20. The non-transitory computer-readable medium of claim 16, wherein determining to use the head editor is based on a bounding box surrounding the target head or target face and a distance between the bounding box and bounding boxes associated with one or more other subjects in the target image.