Computer-implemented method and system for transferring a style from at least two images to another image

By decomposing and aligning images into multiple energy levels and using gain maps to transfer style features from multiple references, the method effectively edits headshot portraits to match desired stylistic appearances without altering the person's identity.

DE102016011998B4Active Publication Date: 2025-07-03ADOBE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102016011998
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2015-11-19
Filing Date
2016-10-06
Publication Date
2025-07-03
Estimated Expiration
2036-10-06

AI Technical Summary

Technical Problem

Existing image editing algorithms struggle to automatically transfer the stylistic appearance of headshot portraits from multiple reference images to an input image without changing the person's identity, as they often apply global modifications or ignore local changes, leading to undesirable results.

Method used

The technique decomposes the input and reference images into multiple energy levels, calculates local energy signatures, and uses gain maps to transfer style features from multiple reference images at each level, ensuring the output image matches the best-matching reference images for each detail while preserving the input image's identity.

Benefits of technology

The method produces an output image that accurately incorporates the stylistic features of multiple reference images, maintaining the person's identity and appearance, while avoiding undesirable compromises common in single-reference image techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented method for transferring a style from at least two images to another image, the method comprising: receiving, by a computer processor (1030), input data for representing each of an input image (130), a first reference image (132), and a second reference image (132), wherein the first reference image (132) and the second reference image (132) are different from each other; decomposing, by the computer processor (1030), based on the input data, the input image (130) into data representing a first level of detail and a second level of detail and a corresponding first energy level and second energy level; converting, by the computer processor (1030), the first level of detail of the input image (130) based on a pre-calculated first energy level of the first reference image (132); converting the second level of detail of the input image (130) based on a pre-calculated second energy level of the second reference image (132) by the computer processor (1030); and generating output data (134) for representing an output image (134) by the computer processor (1030) by merging the converted first level of detail of the input image (130) and the converted second level of detail of the input image (130), the method further comprising: receiving further input data by a computer processor (1030) to display a third reference image (132); decomposing each of the input image (130) and the third reference image (132) into a residue based on the further input data by the computer processor (1030); and converting, by the computer processor (1030), at least a portion of the residue in the input image (130) based on the residue in the third reference image (132), wherein generating the output data (134) further comprises merging the converted remainder of the input image (130), and the method further comprising: separating, by the computer processor (1030), a foreground region of the input image (130) and a background region of the input image (130) based on the input data; calculating, by the computer processor (1030), a mean and a standard deviation of a change in the foreground region based on the output data (134); and converting, by the computer processor (1030), the background region by the mean and standard deviation of the change in the foreground region.
Need to check novelty before this filing date? Find Prior Art

Description

Area of Revelation

[0001] The present disclosure relates generally to the field of digital image processing and, more particularly, to techniques for automatically transferring a style of at least one image, such as a headshot portrait, to another image. background

[0002] Headshot portraits are a popular subject of photography. Professional photographers devote considerable time and effort to editing headshot photos and achieving a convincing style. Different styles convey different moods. For example, a high-contrast black and white portrait can convey dignity, while a bright and colorful portrait can convey an uplifting atmosphere. However, the editing process to create such renditions requires advanced skills, as features such as the eyes, eyebrows, skin, mouth, and hair each require special treatment. In many cases, editing an image to achieve a convincing result involves maintaining a visually pleasing appearance while making significant adjustments to the original image.However, the margin for error is low, as even small adjustments to a photograph can lead to undesirable results. Therefore, advanced editing skills beyond the capabilities of most casual photographers are required to produce such reproductions.

[0003] SHIH, YiChang, [et al.]: Style transfer for headshot portraits. In: ACM Trans. 2014, Graph. 33, 4, Article 148, describes a method for transferring the style of one sample portrait image to another. The method uses a multi-scale technique to transfer the local statistics of the sample portrait photo to the other. The corresponding source code is disclosed under SHIH, YiChang: style_transfer.m. In: Code, Style Transfer for Headshot Portraits (SIGGRAPH), 2014, http: / / groups.csail.mit.edu / graphics / face / tracker / release / zipfiles / code.zip, archived at https: / / archive.org. Summary

[0004] The above and other objects are achieved by the subject matter of the independent claims. Additional embodiments are set forth in the dependent claims. Short description of the drawing Fig. 1 shows an exemplary system for automatically transferring a portrait style of a set of one or more images to another image according to an embodiment of the present disclosure. Fig. 2 shows an example of multiple images and decompositions of each image at different energy levels as well as the process of finding match hits at each level according to an embodiment of the present disclosure. Fig. 3 is a pictorial representation of an exemplary methodology for matching and transferring a portrait style of a set of images to another image according to an embodiment of the present disclosure. Fig. 4 is a pictorial representation of an exemplary methodology for transferring a background region style of one set of images to another image according to an embodiment of the present disclosure. Fig. 5 is a flowchart of an exemplary methodology for transferring a portrait style of one set of images to another image according to an embodiment of the present disclosure. Fig. 6 is a block diagram illustrating an exemplary computing device that may be used to perform any of the techniques described at various points in this disclosure, according to an embodiment of the present disclosure. Detailed description

[0005] There are cases where it is desired to edit a person's portrait so that the stylistic appearance of the image resembles that of a set of one or more reference images, such as a set of professionally produced images, without changing the person's identity. The stylistic appearance can include, for example, lighting, contrast, texture, color, and background effects, which in various combinations add an artistic touch to the portrait. From a technical perspective, however, editing headshots is challenging because edits are typically applied locally. For example, hair does not receive the same treatment as skin, and skin may be treated differently in different areas, such as the forehead, cheeks, and chin.Furthermore, lighting is critical to the appearance of a person's face. Point light sources, for example, can create a very different appearance than diffuse light sources. Similarly, frontal lighting can create a very different appearance than side lighting. Existing algorithms that automate the processing of generic photographs often produce poor results for headshots because they apply modifications to the image globally or otherwise ignore the specifics of headshot retouching while neglecting the effects of local changes to the image. Furthermore, such algorithms attempt to use a single reference image to provide the best match, which can lead to undesirable compromises in cases where the single reference image does not match in all aspects.

[0006] To this end, according to one embodiment of the present disclosure, techniques are disclosed for automatically transferring a style from at least two images, such as two or more reference images, to another image, such as an input image. The reference images contain one or more style features, such as lighting, contrast, color, and texture features, to be reproduced in the input image. The visual appearance of the input image can generally be altered by decomposing the input image into multiple scales, manipulating the decomposition at each scale using the reference images, and then recombining the decompositions to produce an output image.In detail, this means that the input image and each of the reference images are decomposed into different levels, also known as detail levels, by filtering each image using a series of low-pass filters. At each detail level, the local energy is calculated. Local energy is an estimate of how much the signal varies locally at a given level. The calculated energies at the different detail levels are referred to as energy levels in the present disclosure. Each energy level represents the visual appearance or style of the image at the corresponding detail level. Finer detail levels can capture, for example, skin texture, while coarser detail levels can capture lighting and shadows.The final level of decomposition is a residual component representing overall style features not captured at subsequent energy levels, such as overall color and overall intensity. Next, a gain map is computed from the energy level of the input image and the energy level of one of the reference images that best matches the input image at that level. A style transfer operation uses the gain map to transfer the decomposition of the input image such that the energy level of the input image locally approximates the energy level of the reference image. In this way, a first reference image that best matches the input image at one energy level can be used to transfer, for example, fine details.Similarly, a second reference image that best matches the input image at a further energy level can be used to transfer, for example, coarser details. Additionally, the residual component of the input image is transformed, for example, using a histogram transfer or a statistical transfer of a residue of a reference image that best matches the rest of the input image. The transformed decomposition at each detail level and the transformed residue of the input image are then combined to produce an output image that exhibits various stylistic features of the reference images.

[0007] Note that each reference image can be used to modify a different aspect of the optical appearance of the input image. The style transfer process is performed at each level of detail using the best-matching reference image at that energy level to account for a wide range of appearances exhibited by a face, from fine-grained skin texture to larger signal variations caused by the eyes, lips, and nose. The resulting transformation of the input image matches the various optical styles of the reference images without changing the identity of the person in the input image. For example, the transformed portrait depicts the same person as the input, in the same pose and expression, while the color, texture distribution, and overall illumination closely match those of the reference portraits.In this way, the resulting output image may include the best-matching matching features of the multiple reference images. In some cases, the above-described transformations are performed on the foreground region of the input image, which typically includes the person in the portrait. In some cases, according to some embodiments, the background region of the input image may be transformed to an extent proportional to the merged foreground region transformations.

[0008] Embodiments of the present disclosure are to be distinguished from techniques that use a single reference image to transfer the style of a given input image. One difficulty with such single reference image techniques is that a given reference image may have different properties at different levels, making it unlikely that the best-matching reference image will match all properties of the input image (e.g., skin texture, bone structure, facial hair, glasses, hair, skin tone, and lighting). As a result, these techniques often produce unacceptable transformations of an input image by attempting to match a reference image that is very similar in some aspects but quite dissimilar in some others.In contrast to such techniques, embodiments of the present disclosure provide techniques for mapping the styles of multiple different reference images to the input image. The reference images that best match a given energy level of the input image are used to achieve different styling effects. For example, various embodiments of the present disclosure map low-frequency lighting effects of one reference image to the input image and high-frequency texture effects of another reference image to the input image based on how well a given energy level of the respective reference image matches the corresponding energy level of the input image.Because each match between the input image and one of the reference images is constrained to a specific energy level, a good match for a particular style aspect is much more likely than using a single reference image for all style aspects, thereby producing more pleasing results. Various embodiments of the present disclosure also differ from techniques in how the residual and background portions of the input image are rendered, which will be described in more detail below. Numerous embodiments and variations will become apparent in light of the present disclosure. Exemplary system

[0009] Fig. 1 shows an exemplary system 100 for automatically transferring a style from at least two images to another image according to an embodiment of the present disclosure. The system 100 includes a computing device 110 configured to execute a style transfer application 120. The style transfer application 120 includes a pose calculation and image decomposition module 122, a signature calculation module 124, and an image matching and style transfer module 126. The style transfer application 120 is configured to receive data representing an input image 130 and two or more reference images 132 and to generate data representing an output image 134.For example, the input image 130, each of the reference images 132, and the output image 134 may each include image data representing a headshot portrait of a person having a foreground region and a background region, with the person's face forming the foreground. The input image 130, each of the reference images 132, and the output image 134 are all distinct from one another. For example, the input image 130 may include a portrait of a person taken by a user, while each of the reference images 132 may include portraits of other people taken or edited by someone other than the user, such as a professional photographer or an artist who modified the image.Each of the reference images 132 may include one or more stylistic features related to lighting, contrast, texture, color, background, or other optical effects that modify the image from an original, unretouched state. For example, the output image 132 may include data representing the input image as transformed by the stylistic features extracted from the reference images, such that the output image depicts the same person as the input image, but modified to reflect at least some of the stylistic features of the reference images (e.g., increased contrast, altered colors, dimmer lighting, smoother textures, or combinations of features) at various locations.

[0010] In some embodiments, the system 100 includes a data communications network 160 (e.g., the Internet or an intranet) and a file server 150 that communicates with the computing device 110 over the network 160. The file server 150 may host a database for storing pre-computed reference image data 152 (e.g., energy levels, residuals, canonical pose information, and signature vectors), which will be described in more detail below. In some embodiments, the computing device 110 is configured to retrieve the pre-computed reference image data 152 from the file server 150 over the network 160. The pre-computed reference image data 152 may be used instead of, or in addition to, the reference images 132. In some embodiments, the system 100 includes a server 170 having a style transfer application 120 as described herein.In such embodiments, computing device 110 may include a client user interface connected to server 170 via network 160, wherein at least some of the computations may be performed on server 170 (e.g., by style transfer application 120), and images and other data may be transferred back and forth between computing device 110 and server 170.

[0011] Fig. 2 shows an example of an input image 130, a set of reference images 132, and various energy levels of decomposition 202 thereof, according to an embodiment of the present disclosure. Each of the reference images 132 may belong to a set of images representing one or more particular style features. The reference images 132 in such a set may, for example, have similar lighting, contrast, or texture features, even though the images are different. The pose calculation and image decomposition module 122 of Fig. 1 is designed to decompose the input image 130 and each of the reference images 132 into data 140 representing a multi-level decomposition of the images into multiple levels of detail corresponding to energy levels and a residue, generally designated 202. It should be appreciated that the images may be decomposed into any number of levels of detail and corresponding energy levels (e.g., two levels, three levels, four levels, five levels, and so on). As in Fig. 2, each of the images can be decomposed, for example, using a two-dimensional normalized Gaussian kernel of standard deviation σ in the following way: Ll[l]=[I−I⊗G(2)if l=0I⊗G(2l)−I⊗G(2l+1)if l>0 R[I]=I⊗G(2n)

[0012] Here, L I[I] the decomposition levels of the input image 130, where G is a Gaussian function, I represents the energy level (level), and R[I] represents the residue. The energy levels can be calculated independently for each level of the decomposed input 130 and the reference images 132, for example, by averaging the square of the energy level coefficients. This generally provides an estimate of how much the signal varies locally at each energy level. The energy level of a given image can be calculated, for example, as follows: Sl[I]=Ll2(I)⊗G(2l+1) Here S I [I] the energy level at level I. These equations can also be used to calculate the decomposition and energy levels of each of the reference images 132 and the output image 134.

[0013] As in Fig. 1 and Fig. 2, the signature computation module 124 is configured to align canonical feature poses of the input image 130 to the reference images 132 as captured by the pose computation and image decomposition module 122 and compute local energy signatures 142, also referred to as signature vectors, for each of the aligned input 130 and the reference images 132. The reference images 132, which define a particular stylistic feature, may be associated with a predefined canonical feature pose. Each feature pose represents the location of a facial feature, such as the eyes, nose, mouth, and chin. The canonical feature pose may be computed by averaging the feature poses of all the reference images 132. The signature computation module can align the features of an image with the canonical pose by computing an image morph that maps the location of each feature to the canonical location.Image morphing can be performed on the energy levels to align the energy levels accordingly. Once the features are aligned, a signature vector can be created for each energy level or the remainder by concatenating the data into a single linear vector.

[0014] In some embodiments, the data for the reference images 132, such as the energy levels, residues, the canonical pose, and the signature vectors, may be pre-computed to form a data set (e.g., the pre-computed data 152 of Fig. 1) that represents the characteristics of a particular style. The data set can be stored, for example, locally in a database or on a network-connected file server (for example, the file server 150 of Fig. 1) for subsequent retrieval. This way, the reference image data does not need to be calculated each time the dataset is used for a style transfer operation. In some embodiments, different datasets may be used to represent reference images with multiple different portrait styles in various combinations with the input image 130, allowing a user to see various possible styles.

[0015] Fig. 3 shows a pictorial example of a methodical approach for matching the reference images 132 with the foreground area of the input image 130 and for transferring the style of the reference images to the input image, thereby generating the output image 134, according to one embodiment. The top row of Fig. Figure 3 shows the input image 130, along with four energy levels and the residue 202 from the decomposition of the input image. The next five rows show examples of the desired style that best matches the input image 130 at the energy level and residue 202. The bottom row shows gain maps and a new residue applied to the decomposition detail levels and the residue of the input image 130 and merged to obtain the output image 134. The resulting output image 134 has the style features transferred from the respective energy levels of the reference images 132.

[0016] The image matching and style transfer module 126 of Fig. 1 is configured to transfer one or more style features of the reference images 132 to the input image 130 using the signatures 142 of the aligned images generated by the signature calculation module 124. Although the signatures are calculated and aligned with the general pose of the reference images 132, the style transfer can be achieved by aligning the energy levels of the reference images 132 to the pose of the input 130 while maintaining the integrity of the input. The result of the transfer is the output image 134, which is the input image 130 transformed by the optical style(s) of the reference images 132 without changing the identity of the person in the input image.In other words, the output image 134 represents the same person as the input image 130 in the same pose and with the same expression, but with the color and texture distribution as well as the overall illumination adjusted to one or more of the reference images 132.

[0017] As in Fig. 3, the operations of the image matching and style transfer module 126 can be performed at each of the energy levels to accommodate a wide range of appearances exhibited by a face, from the fine-grained skin texture to the larger signal variations caused by the eyes, lips, and nose. Furthermore, as shown in Fig. 3, the style transfer process may use different reference images 132 at each energy level (e.g., level 0, example of level 0, level 1, example of level 1, and the like) as well as for the rest. Note that in some cases, the same reference image 132 may be used for different energy levels (e.g., levels 1 and 2 are separate examples of Fig. 3 both decompositions of the same reference image). The transfer process for a given level of detail of a given image (e.g., the input image 130 and the reference images 132) can be calculated, for example, as follows: Ll[O]=Ll(I)×gain map Gain map=Sl[E]Sl[I]+ε

[0018] Here, ε is a small number to avoid division by zero (for example, 0.01 2). Each transfer may be performed independently in the pose of the input image 130, with the energy level signatures of the reference images 132 being morphed to the pose of the input image 130. A mask may be applied to the input image 130 to limit the transfer operations to the foreground regions, including, for example, the subject's face, so that the background region is not modified at this stage in the process.

[0019] A histogram transfer can be performed for the residual L channel (intensity). The histogram of intensity values captures the overall illumination, overall darkness, and overall contrast in an image. The histogram of the reference image residual is calculated, and a histogram matching algorithm is applied to the input image residual to match the histogram of the input image residual. Histogram matching adjusts the overall illumination and contrast in the input image to match the reference image, while preserving the spatial location of the highlights and shadows in the input image.

[0020] A statistical transfer is performed on the residual a, b channels (color). The mean and standard deviation of the a and b channels of the reference image residual are calculated, and a linear affine conversion is performed on the a and b channels of the input image residual so that the a and b channels of the input image residual have the same mean and standard deviation as the a and b channels of the reference image residual. This results in the input image having the same average color and average range of color variation as the example, but the spatial distribution of the colors remains unaffected. It should be understood that the techniques described in various places in connection with the image matching and style transfer operations are merely examples, and that other techniques can be used to match images and transfer the style features.

[0021] Fig. 4 shows a pictorial example of a methodical procedure for transferring the background area of the reference images 132 to the input image 130 according to one embodiment. In summary, this means that for a given input image, the style of one or more reference images can be transferred to the foreground area of the output image 134, as described above with reference to Fig. 1 to 3. For the background area of the input image 130, the image matching and style transfer module 126 of Fig. 1 is configured to separate the background region from the foreground region of the input image 130, calculate the net statistical change (i.e., the change in mean and standard deviation) between the foreground region of the input image 130 and the foreground region of the output image 134, and perform a linear affine transformation on the background region of the input image 130 such that the mean and standard deviation of the background changes change by the same factor as the mean and standard deviation of the foreground. In this way, the original background texture and detail are preserved, but the overall color and illumination correspond to the transformed foreground. In some embodiments, the user is given the option of selecting the background of any of the reference images 132, the original background of the input image 130, the transformed background of the output image 134, a fixed orconstant (solid) average color of any of these or any mixture or other combination of these options. Exemplary methodical approach to style transfer

[0022] Fig. 5 is a flowchart of an exemplary method 500 for transferring a style from at least two images to another image according to an embodiment of the present disclosure. Various aspects of the method 500 may be implemented, for example, in the computing device 110 of Fig. 1. The method 500 begins with receiving 502 input data for representing an input image (for example, the input image 130 of Fig. 1), a first reference image and a second reference image (for example, the reference images 132 of Fig. 1). It should be understood that any number of reference images may be received and used. The reference images may, for example, form a set representing a particular style or a set of style features that the user wishes to transfer to the input image. The input image and each of the reference images may be different from each other, as explained above with reference to Fig. 2 and Fig. 3. The method 500 continues with the decomposition 506 of each of the images (e.g., the input image, the first reference image, the second reference image, and any additional reference images) into at least a first decomposition level (level of detail), a second decomposition level, and a corresponding first energy level and second energy level, as well as a residue (e.g., the level of detail, the energy level, and the residue data 140 of Fig. 1). It should be understood that the images can be divided into any number of levels of detail and energy, such as Fig. 2 (e.g., two levels, three levels, four levels, five levels, and so on). In some embodiments, data representing the decomposed reference images (e.g., energy levels, residuals, canonical pose, and signature vectors) may be pre-computed and stored in a dataset. In such embodiments, the method may utilize the pre-computed reference image data instead of calculating it for each input image.

[0023] In some embodiments, the method 500 includes calculating 508 a first energy signature for the first energy level of the input image, the first reference image, and the second reference image, as well as for any additional reference images (e.g., the energy level data 142 of Fig. 1). Similarly, a second energy signature for the second energy level, a third energy signature for a third energy level (if present), and so on, may be calculated for each of the energy levels in each of the decomposed images, as well as for the residues of each image. As mentioned above, in some embodiments, the signatures for the reference images may be pre-calculated. The method 500 may further include determining 510 whether the first energy signature of the input image is more similar to the first energy signature of the first reference image compared to the first energy signature of the second reference image, as well as determining whether the second energy signature of the input image is more similar to the second energy signature of the second reference image compared to the second energy signature of the first reference image, which continues for each of the energy levels and each of the residues.In this way, the reference image that is most similar to the input image at each energy level and residue can be used, since different reference images may have more similar signatures to the input image at some energy levels than at other energy levels compared to other reference images. As shown, for example, in . Fig. 3, one reference image (level 0 example) may be most similar to the input image at level 0, while another reference image may be most similar to the input image at levels 1 and 2 (level 1 example and level 2 example).

[0024] Once the signature similarities have been determined, method 500 continues by transforming 512 the first decomposition level of the input image based on the first energy or the first energy level of the input image and the first energy level of the first reference image, and transforming the second decomposition level of the input image based on the second energy level of the input image and the second energy level of the second reference image. If additional energy levels are present, the process is similar to transforming each corresponding level of the input image. In some embodiments, each level of the input image is transformed independently of one or more other energy levels in the input image.In some embodiments, transforming 512 the input image includes calculating a gain map for each of the energy levels of the input image based on the local signal variation for the respective levels, wherein transforming the decomposition levels of the input image is a function of the gain map as described above. The method 500 continues by transforming 514 the residue of the input image based on the residue of a third reference image, which may be the same as the first and second reference images or different depending on the residue signature similarities. In some embodiments, transforming 514 the input image includes transforming at least a portion of the residue in the input image using histogram equalization (e.g., for the intensity channel), a linear affine conversion (e.g., for the a, b color channels), or both.

[0025] The method 500 is continued by generating 516 output data (for example, the output image 134 of Fig. 1) to represent an output image by merging the converted first decomposition level of the input image and the converted second decomposition level of the output image, as well as any additional decomposition levels of the input image. In some embodiments, method 500 includes transforming 518 a background region of the output image by separating a foreground region of the input image from the background region of the input image, calculating a mean and standard deviation of a change in the foreground region (e.g., a change between the original input image and the converted output image), and transforming the background region of the output image by the mean and standard deviation of the change in the foreground region. Exemplary computing device

[0026] Fig. 6 is a block diagram illustrating an exemplary computing device 1000 that may be used to perform any of the techniques described above at various points in this disclosure. The system 100 of Fig. 1 or any sections thereof and the methodological procedures of Fig.5 or any portions thereof may, for example, be implemented in computing device 1000. Computing device 1000 may, for example, be any computer system, such as a workstation, a desktop computer, a server, a laptop, a handheld computer, a tablet computer (e.g., the iPad™ tablet computer), a mobile computing or communications device (e.g., the iPhone™ mobile communications device, the Android™ mobile communications device, and the like), or any other form of computing or telecommunications device capable of communication and having sufficient processing power and memory capacity to perform the operations described in the present disclosure. A distributed computing system comprising a plurality of such computing devices may be contemplated.

[0027] Computing device 1000 includes one or more storage devices 1010 and / or non-transitory computer-readable media 1020 encoded with one or more computer-readable instructions or software for implementing techniques described above at various points in the present disclosure. Storage devices 1010 may include computer system memory or random access memory, such as persistent disk storage (which may include any suitable optical or magnetic persistent storage devices, such as RAM, ROM, flash, USB drive, or any other semiconductor-based storage medium), a hard disk drive, CD-ROM, or other computer-readable media for storing data and computer-readable instructions and / or software implementing various embodiments according to the teachings of the present disclosure.The storage device 1010 may also include other types of storage or combinations thereof. The storage device 1010 may be provided on the computing device 1000 or may be separate or remote from the computing device 1000. The non-transitory computer-readable media 1020 may include, but is not limited to, one or more types of hardware storage, non-transitory physical media (e.g., one or more magnetic storage disks, one or more optical disks, one or more USB flash drives), and the like. The non-transitory computer-readable media 1020 included in the computing device 1000 may store computer-readable and computer-executable instructions or software for implementing various embodiments. The computer-readable media 1020 may be provided on the computing device 1000 or may be separate or remote from the computing device 1000.

[0028] Computing device 1000 also includes at least one processor 1030 for executing computer-readable and computer-executable instructions or software stored in storage device 1010 and / or non-transitory computer-readable media 1020 and other programs for controlling system hardware. Virtualization may be employed in computing device 1000 such that infrastructure and resources within computing device 1000 may be dynamically shared. For example, a virtual machine may be provided to accomplish a process running on multiple processors in such a way that the process appears to be running on only one resource rather than multiple computing resources. Multiple virtual machines may also be used with one processor.

[0029] A user may interact with the computing device 1000 through an output device 1040, such as a screen or monitor, which may display one or more user interfaces provided in accordance with some embodiments. The output device 1040 may also display other aspects, elements, and / or information or data associated with some embodiments. The computing device 1000 may include other I / O devices 1050 for receiving input from a user, such as a keyboard, a joystick, a game controller, a pointing device (e.g., a mouse or a user's finger directly interfacing with a display device), or any other suitable user interface. The computing device 1000 may include other suitable conventional I / O peripherals.Computing device 1000 may include and / or be operatively coupled to various suitable devices for performing one or more of the aspects described at various points in the present disclosure, such as digital cameras for capturing digital images and video displays for displaying digital images.

[0030] The computing device 100 may run on any operating system, such as any version of the Microsoft operating systems ® Windows ® , any of the different versions of the Unix and Linux operating systems, any version of MacOS ®for Macintosh computers, any embedded operating system, a real-time operating system, any open-source operating system, any proprietary operating system, any operating system for mobile computing devices, and any other operating system capable of running on the computing device 100 and performing the operations described in the present disclosure. In one embodiment, the operating system may run on one or more cloud machine instances.

[0031] In some embodiments, the functional components / modules may be implemented with hardware, such as gate-level logic (e.g., FPGA) or a dedicated semiconductor (e.g., ASIC). Other embodiments may be implemented with a microcontroller implementing a number of input / output ports for receiving and outputting data and a number of embedded routines for performing the functionality described in the present disclosure. More generally, any combination of hardware, software, and firmware may be used, as will be apparent.

[0032] As will be appreciated in light of the present disclosure, the various modules and components of the system, such as the style transfer application 120, the pose calculation and image decomposition module 122, the signature calculation module 124, the image matching and style transfer module 126, or any combination thereof, may be implemented in software, such as a set of instructions (e.g., HTML, XML, C, C++, object-oriented C, JavaScript, Java, Basic, and the like) that may be encoded on any computer-readable medium or computer program product (e.g., a hard disk, a server, a floppy disk, or other suitable non-temporary storage or set of memories) and that, when executed by one or more processors, cause various methodological procedures provided in the present disclosure to be carried out.It should be appreciated that in some embodiments, various functions and data transformations performed by the user computing system as described in the present disclosure may be performed by similar processors and / or databases in various configurations and arrangements, and the described embodiments are not intended to be limiting. Various components of the exemplary embodiment, including computing device 1000, may be incorporated into, for example, one or more desktop or laptop computers, workstations, tablets, smartphones, game consoles, set-top boxes, or other such computing devices.Other component parts and modules typical of a computing system, such as processors (e.g., a central processing unit and a coprocessor, a graphics processor, and the like), input devices (e.g., a keyboard, a mouse, a touchpad, a touchscreen, and the like), and an operating system, are not shown but are readily apparent.

[0033] Numerous embodiments are apparent in light of the present disclosure, and the features described herein may be combined in any number of configurations. An exemplary embodiment provides a computer-implemented method for transferring a style from at least two images to another image. The method includes receiving, by a computer processor, input data for representing an input image,a first reference image and a second reference image; a decomposition of the input image into data representing a first level of detail and a second level of detail and a corresponding first energy level and second energy level by the computer processor based on the input data; a conversion of the first level of detail of the input image based on a pre-calculated first energy level of the first reference image by the computer processor; a conversion of the second level of detail of the input image based on a pre-calculated second energy level of the second reference image by the computer processor; and a generation of output data representing an output image by merging the converted first level of detail of the input image and the converted second level of detail of the input image. The input image,the first reference image and the second reference image may be different from each other. In some cases, the first and second levels of detail and energy levels of the first and second reference images are calculated based on decompositions of those images, which may be pre-calculated. In some cases, the first level of detail of the input image is converted independently of the second level of detail of the input image. In some cases, the method includes, by the computer processor, calculating a first energy signature for the first energy level of the input image; and, by the computer processor, calculating a second energy signature for the second energy level of the input image. In some of these cases, the method includes, by the computer processor, determining, prior to generating the output data,that the first energy signature of the input image is more similar to a pre-calculated first energy signature of the first reference image compared to a pre-calculated first energy signature of the second reference image. In some further of these cases, the method includes determining, by the computer processor prior to generating the output data, that the second energy signature of the input image is more similar to a pre-calculated second energy signature of the second reference image compared to a predetermined second energy signature of the first reference image. In some cases, converting the first level of detail of the input image and converting the second level of detail of the input image each further include calculating, by the computer processor, energy levels based on a local signal variation within each of the levels of detail of each of the input images.the first reference image and the second reference image; and calculating, by the computer processor, a gain map for each of the first and second levels of detail of the input image based on the energy levels for the respective levels, wherein transforming the first and second levels of detail of the input image is a function of the gain map. In some cases, the method includes receiving, by a computer processor, further input data representing a third reference image; decomposing, by the computer processor, each of the input image and the third reference image into a residue based on the further input data; and transforming, by the computer processor, at least a portion of the residue in the input image based on the residue in the third reference image,wherein generating the output data further comprises merging the converted residue of the input image. In some cases, converting at least the portion of the residue in the input image uses histogram matching or a linear affine conversion. In some cases, the method includes, by the computer processor, separating a foreground region of the input image and a background region of the input image based on the input data;, by the computer processor, calculating a mean and a standard deviation of a change in the foreground region based on the output data; and, by the computer processor, converting the background region by the mean and standard deviation of the change in the foreground region.

[0034] Another embodiment provides a system including a memory and a computer processor operatively coupled to the memory. The computer processor is configured to execute instructions stored in the memory that, when executed, cause the computer processor to perform a process. The process includes receiving input data representing each of an input image, a first reference image, and a second reference image, wherein the input image, the first reference image, and the second reference image are different from one another; decomposing each of the input images based on the input data,the first reference image and the second reference image into data representing a first energy level and a second energy level and a corresponding first energy level and second energy level; converting the first level of detail of the input image based on the first energy level of the first reference image; converting the second level of detail of the input image based on the second energy level of the second reference image; and generating output data representing an output image by merging the converted first level of detail of the input image and the converted second level of detail of the input image. In some cases, the first level of detail of the input image is converted independently of the second level of detail of the input image. In some cases, the process includes calculating a first energy signature for each of the first energy level of the input image,the first energy level of the first reference image and the first energy level of the second reference image; and calculating a second energy signature for each of the second energy level of the input image, the second energy level of the first reference image, and the second energy level of the second reference image. In some such cases, the process includes determining, prior to generating the output data, that the first energy signature of the input image is more similar to the first energy signature of the first reference image compared to the first energy signature of the second reference image. In some further such cases, the process includes determining, prior to generating the output data,that the second energy signature of the input image is more similar to the second energy signature of the second reference image compared to the second energy signature of the first reference image. In some cases, transforming the first energy level of the input image and transforming the second energy level of the input image each further include calculating the first and second energy levels based on a local signal variation within each of the detail levels of each of the input image, the first reference image, and the second reference image; and calculating a gain map for each of the first and second detail levels of the input image based on the respective first and second energy levels,wherein converting the first and second detail levels of the input image is a function of the gain map. In some cases, the process includes receiving further input data to represent a third reference image; decomposing each of the input image and the third reference image into a residue based on the further input data; and converting at least a portion of the residue in the input image based on the residue in the third reference image.wherein generating the output data further comprises merging the converted residue of the input image. In some such cases, converting at least the portion of the residue in the input image uses a histogram adjustment or a linear affine conversion. In some cases, the process includes separating a foreground region of the input image and a background region of the input image based on the input data; calculating a mean and a standard deviation of a change in the foreground based on the output data; and converting the background region by the mean and the standard deviation of the change in the foreground region. Another exemplary embodiment provides a non-temporary computer program product encoded with instructions that, when executed by one or more processors, causethat a process is carried out to perform one or more of the aspects described at various points in this paragraph.

[0035] The foregoing description and drawings of the various embodiments are provided purely by way of example. The examples are not intended to be exhaustive or to limit the invention to the precise details disclosed. Alterations, modifications, and variations will become apparent in light of the present disclosure and are intended to be included within the scope of the invention as defined in the claims.

Claims

[1] A computer-implemented method for transferring a style of at least two images to another image, the method comprising: receiving, by a computer processor (1030), input data for representing each of an input image (130), a first reference image (132), and a second reference image (132), wherein the first reference image (132) and the second reference image (132) are different from each other; decomposing, by the computer processor (1030), based on the input data, the input image (130) into data representing a first level of detail and a second level of detail and a corresponding first energy level and second energy level; converting, by the computer processor (1030), the first level of detail of the input image (130) based on a pre-calculated first energy level of the first reference image (132); converting the second level of detail of the input image (130) based on a pre-calculated second energy level of the second reference image (132) by the computer processor (1030); and generating output data (134) for representing an output image (134) by the computer processor (1030) by merging the converted first level of detail of the input image (130) and the converted second level of detail of the input image (130), the method further comprising: receiving further input data by a computer processor (1030) to display a third reference image (132); decomposing each of the input image (130) and the third reference image (132) into a residue based on the further input data by the computer processor (1030); and converting, by the computer processor (1030), at least a portion of the residue in the input image (130) based on the residue in the third reference image (132), wherein generating the output data (134) further comprises merging the converted remainder of the input image (130), and the method further comprising: separating, by the computer processor (1030), a foreground region of the input image (130) and a background region of the input image (130) based on the input data; calculating, by the computer processor (1030), a mean and a standard deviation of a change in the foreground region based on the output data (134); and converting, by the computer processor (1030), the background region by the mean and standard deviation of the change in the foreground region. [2] The method of claim 1, wherein the first level of detail of the input image (130) is converted independently of the second level of detail of the input image (130). [3] The method according to any one of claims 1 or 2, further comprising: calculating, by the computer processor (1030), a first energy signature for the first energy level of the input image (130); and calculating, by the computer processor (1030), a second energy signature for the second energy level of the input image (130). [4] The method of claim 3, further comprising: determining, by the computer processor (1030) prior to generating the output data (134), that the first energy signature of the input image (130) is more similar to a pre-calculated first energy signature of the first reference image (132) compared to a pre-calculated first energy signature of the second reference image (132). [5] A method according to any one of claims 3 or 4, further comprising: determining, by the computer processor (1030) prior to generating the output data (134), that the second energy signature of the input image (130) is more similar to a pre-calculated second energy signature of the second reference image (132) compared to a pre-calculated second energy signature of the first reference image (132). [6] The method of any preceding claim, wherein converting the first level of detail of the input image (130) and converting the second level of detail of the input image (130) each further comprise: calculating, by the computer processor (1030), energy levels based on a local signal variation within each of the detail levels of each of the input image (130), the first reference image (132), and the second reference image (132); and calculating, by the computer processor (1030), a gain map for each of the first and second detail levels of the input image (130) based on the energy levels for the respective levels, wherein transforming the first and second levels of detail of the input image (130) is a function of the gain mapping. [7] The method of any preceding claim, wherein transforming at least the portion of the residue in the input image (130) uses at least one of a histogram adjustment and a linear affine transformation. [8] A system in a digital image processing environment, comprising: a memory; and a computer processor (1030) operatively coupled to the memory, the computer processor (1030) being configured to execute instructions stored in the memory that, when executed, cause the computer processor (1030) to perform a process comprising: Receiving input data to represent each of an input image (130), a first reference image (132), and a second reference image (132), wherein the first reference image (132) and the second reference image (132) are different from each other; based on the input data, decomposing each of the input image (130), the first reference image (132) and the second reference image (132) into data representing a first energy level and a second energy level and a corresponding first energy level and second energy level; Converting the first level of detail of the input image (130) based on the first energy level of the first reference image (132); Converting the second level of detail of the input image (130) based on the second energy level of the second reference image (132); and Generating output data (134) for representing an output image (134) by merging the converted first level of detail of the input image (130) and the converted second level of detail of the input image (130), the process further includes: Receiving further input data for displaying a third reference image (132); Decomposing each of the input image (130) and the third reference image (132) into a residue based on the further input data; and Converting at least a portion of the residue in the input image (130) based on the residue in the third reference image (132), wherein generating the output data (134) further comprises merging the converted remainder of the input image (130), and the process further includes: Separating a foreground area of the input image (130) and a background area of the input image (130) based on the input data; Calculating a mean and a standard deviation of a change in the foreground region based on the output data (134); and Transform the background area by the mean and standard deviation of the change in the foreground area. [9] The system of claim 8, wherein the first level of detail of the input image (130) is converted independently of the second level of detail of the input image (130). [10] The system of any one of claims 8 or 9, wherein the process further comprises: Calculating a first energy signature for each of the first energy level of the input image (130), the first energy level of the first reference image (132), and the first energy level of the second reference image (132); and calculating a second energy signature for each of the second energy level of the input image (130), the second energy level of the first reference image (132), and the second energy level of the second reference image (132). [11] The system of claim 10, wherein the process further comprises: determining, prior to generating the output data (134), that the first energy signature of the input image (130) is more similar to the first energy signature of the first reference image (132) compared to the first energy signature of the second reference image (132). [12] The system of any one of claims 10 or 11, wherein the process further comprises: determining, prior to generating the output data (134), that the second energy signature of the input image (130) is more similar to the second energy signature of the second reference image (132) compared to the second energy signature of the first reference image (132). [13] The system of any one of claims 8 to 12, wherein converting the first energy level of the input image (130) and converting the second energy level of the input image (130) each further comprise: Calculating the first and second energy levels based on a local signal variation within each of the detail levels of each of the input image (130), the first reference image (132), and the second reference image (132); and Calculating a gain map for each of the first and second detail levels of the input image (130) based on the respective first and second energy levels, wherein transforming the first and second levels of detail of the input image (130) is a function of the gain mapping. [14] The system of any of claims 8 to 13, wherein transforming at least the portion of the residue in the input image (130) uses at least one of a histogram adjustment and a linear affine transformation. [15] A non-temporary computer program product having encoded thereon instructions that, when executed by one or more computer processors (1030), cause the one or more computer processors (1030) to perform a process comprising: Receiving input data to represent each of an input image (130), a first reference image (132), and a second reference image (132), wherein the first reference image (132) and the second reference image (132) are different from each other; based on the input data, decomposing each of the input image (130), the first reference image (132) and the second reference image (132) into a first level of detail and a second level of detail and a corresponding first energy level and second energy level; Converting the first level of detail of the input image (130) based on the first energy level of the first reference image (132); Converting the second level of detail of the input image (130) based on the second energy level of the second reference image (132); and Generating output data (134) for representing an output image (134) by merging the converted first level of detail of the input image (130) and the converted second level of detail of the input image (130), the process further includes: Receiving further input data for displaying a third reference image (132); Decomposing each of the input image (130) and the third reference image (132) into a residue based on the further input data; and Converting at least a portion of the residue in the input image (130) based on the residue in the third reference image (132), wherein generating the output data (134) further comprises merging the converted remainder of the input image (130), and the process further includes: Separating a foreground area of the input image (130) and a background area of the input image (130) based on the input data; Calculating a mean and a standard deviation of a change in the foreground region based on the output data (134); and Transform the background area by the mean and standard deviation of the change in the foreground area. [16] The non-temporary computer program product of claim 15, wherein converting the first level of detail of the input image (130) and converting the second level of detail of the input image (130) each further comprise: calculating by the computer processor (1030) the first and second energy levels based on a local signal variation within each of the detail levels of each of the input image (130), the first reference image (132), and the second reference image (132); and calculating, by the computer processor (1030), a gain map for each of the first and second detail levels of the input image (130) based on the energy levels for the respective levels, wherein transforming the first and second levels of detail of the input image (130) is a function of the gain mapping.