Adaptive Convolution in Neural Networks
By using adaptive convolution in style transfer technology, the global and local features of the style image are migrated to the content image, and the problem of poor local feature migration in the existing technology is solved, achieving a more efficient style transfer effect.
Patent Information
- Application Number
- CN202111355035.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-06
- Filing Date
- 2021-11-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-11-16
AI Technical Summary
Existing style transfer technologies are difficult to effectively migrate local features in style images, resulting in the output image lacking lower-level features such as edges and lines of style images.
By applying the neural network layer to the latent representation of the style samples to generate the convolution kernel and convolve the latent representation of the content samples with these convolution kernels, the convolution output is generated and the decoder layer is applied to generate the style transfer result.
Effective migration of global and local features of style images is realized, and the generated style transfer results capture the style of style samples more accurately, reducing resource consumption and manual processing time.
Smart Images

Figure CN114511440B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 114,504, filed on November 16, 2020, titled "ADAPTIVE CONVOLUTIONS FOR STYLE TRANSFER". The subject matter of this related application is incorporated herein by reference. Field of the Invention
[0003] Fields of Various Embodiments
[0004] Embodiments of the present disclosure generally relate to convolutional neural networks, and more particularly to adaptive convolutions in neural networks. Background Art
[0005] Description of the Related Art
[0006] Style transfer refers to the technique of transferring the "style" of a first image to a second image without modifying the content of the second image. For example, the color, pattern, and / or other style - based attributes of the first image can be transferred to one or more faces, buildings, bridges, and / or other objects in the second image without removing the objects from the second image or adding new objects to the second image.
[0007] Existing style transfer methods typically use convolutional neural networks to learn or characterize the "global" statistics of a style image and transfer the statistics to a content image. For example, an encoder network can be used to generate feature maps for the content and style images. Means and standard deviations can be calculated for one or more portions of the feature map of the style image, and the corresponding portions of the feature map of the content image can be normalized to have the same means and standard deviations. Then, a decoder network can be used to convert the normalized feature maps into an output image that combines the style of the style image with the content of the content image.
[0008] On the other hand, existing style transfer techniques cannot identify the "local" features in a style image or transfer them to the content image. Continuing the above example, the output image can capture the overall style of the style image but lacks the edges, lines, and / or other lower - level features of the style image.
[0009] As previously explained, what is needed in the art are techniques for improving the transfer of global and local features of a style image to a content image during style transfer. Summary of the Invention
[0010] One embodiment describes a technique for performing style transfer between a content sample and a style sample. The technique includes applying one or more neural network layers to a first latent representation of the style sample to generate one or more convolutional kernels. The technique also includes generating a convolutional output by convolving a second latent representation of the content sample with the one or more convolutional kernels. The technique further includes applying one or more decoder layers to the convolutional output to produce a style transfer result that includes one or more content-based attributes of the content sample and one or more style-based attributes of the style sample.
[0011] One technical advantage of the disclosed technique is reduced overhead and / or resource consumption compared to existing techniques for generating content in a certain style. For example, traditional techniques for adapting images, videos, and / or other content to a new style may involve a user manually capturing, creating, editing, and / or re-rendering the content to reflect the new style. Drawing, modeling, editing, and / or other tools that a user uses to create, update, and store content can consume significant amounts of computing, memory, storage, network, and / or other resources. In contrast, the disclosed technique can perform batch processing that uses a style transfer model to automatically transfer the style to the content, which consumes less time and / or resources than the manual creation or modification of content performed in traditional techniques. Thus, by automating the transfer of different styles to content, the disclosed embodiments provide a technical improvement to computer systems, applications, frameworks, and / or techniques for generating content and / or performing style transfer. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To understand the above features of the various embodiments in detail, reference may be made to the more specific description of the inventive concept briefly summarized above, where some embodiments are illustrated in the drawings. However, it should be noted that the drawings only illustrate typical embodiments of the inventive concept and should not be considered to limit the scope in any way, and there are other equally effective embodiments.
[0013] Figure 1 Illustrates a system configured to implement one or more aspects of the various embodiments.
[0014] Figure 2 For the Figure 1 training engine and estimation engine according to the various embodiments.
[0015] Figure 3 Is a flowchart of method steps for training a style transfer model according to the various embodiments.
[0016] Figure 4 Is a flowchart of method steps for performing style transfer according to the various embodiments.
[0017] Figure 5Flowchart of method steps for performing adaptive convolution in a neural network according to various embodiments. Detailed Description
[0018] In the following description, numerous specific details are set forth to provide a more thorough understanding of the various embodiments. However, it will be apparent to one of ordinary skill in the art that the inventive concept may be practiced without one or more of these specific details.
[0019] System Overview
[0020] Figure 1 Illustrated is a computing device 100 configured to implement one or more aspects of the various embodiments. In one embodiment, the computing device 100 may be a desktop computer, laptop computer, smart phone, personal digital assistant (PDA), tablet computer, or any other type of computing device configured to receive input, process data, and optionally display images, and suitable for practicing one or more embodiments. The computing device 100 is configured to run a training engine 122 and an execution engine 124 residing in a memory 116. It should be noted that the computing device described herein is illustrative, and any other technically feasible configuration falls within the scope of the present disclosure. For example, multiple instances of the training engine 122 and the execution engine 124 may be executed on a set of nodes in a distributed system to implement the functions of the computing device 100.
[0021] In one embodiment, the computing device 100 includes, but is not limited to, an interconnect (bus) 112 connecting one or more processors 102, an input / output (I / O) device interface 104 coupled to one or more input / output (I / O) devices 108, a memory 116, a storage device 114, and a network interface 106. The processor 102 may be any suitable processor implemented as a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), an artificial intelligence (AI) accelerator, any other type of processing unit, or a combination of different processing units, such as a CPU configured to operate in conjunction with a GPU. Generally, the processor 102 may be any technically feasible hardware unit capable of processing data and / or executing software applications. Additionally, in the context of the present disclosure, the computing elements shown in the computing device 100 may correspond to a physical computing system (e.g., a system in a data center), or may be virtual computing instances executing within a computing cloud.
[0022] In one embodiment, the I / O device 108 includes devices capable of providing input, such as a keyboard, a mouse, a touch screen, etc., and devices capable of providing output, such as a display device. Additionally, the I / O device 108 may include devices capable of receiving input and providing output, such as a touch screen, a Universal Serial Bus (USB) port, etc. The I / O device 108 may be configured to receive various types of input from an end user ( For example , designer) of the computing device 100, and also provide various types of output to the end user of the computing device 100, such as a displayed digital image or digital video or text. In some embodiments, one or more of the I / O devices 108 are configured to couple the computing device 100 to the network 110.
[0023] In one embodiment, the network 110 is any technically feasible type of communication network that allows for the exchange of data between the computing device 100 and an external entity or device, such as a network server or another networked computing device. For example, the network 110 may include a Wide Area Network (WAN), a Local Area Network (LAN), a wireless (WiFi) network, and / or the Internet, etc.
[0024] In one embodiment, the storage device 114 includes non-volatile storage for applications and data, and may include fixed or removable disk drives, flash devices, and CD-ROM, DVD-ROM, Blu-Ray, HD-DVD, or other magnetic, optical, or solid-state storage devices. The training engine 122 and the execution engine 124 may be stored in the storage device 114 and loaded into the memory 116 when executed.
[0025] In one embodiment, the memory 116 includes Random Access Memory (RAM) modules, flash memory cells, or any other type of storage cells or a combination thereof. The processor 102, the I / O device interface 104, and the network interface 106 are configured to read data from and write data to the memory 116. The memory 116 includes various software programs executable by the processor 102 and application data associated with the software programs, including the training engine 122 and the execution engine 124.
[0026] The training engine 122 includes the function of training a style transfer model, and the execution engine 124 includes using the style transfer model to generate an output including an input style sample ( For example, the function of the style transfer result of the style of the image) and the content of the input content sample. As described in more detail below, the style transfer model can learn the features of the style sample at different granularities and / or resolutions. Then, the features can be combined with the content of the content sample to generate a style transfer result that "adapts" the content in the content sample to the style of the style sample. Therefore, the style transfer model can produce an output that more accurately captures the style of the style sample than existing style transfer techniques.
[0027] Adaptive Convolution for Style Transfer
[0028] Figure 2 For various embodiments of Figure 1 A more detailed description of the training engine 122 and the execution engine 124. As described above, the training engine 122 and the execution engine 124 operate to train and execute a style transfer model 200 that generates a style transfer result 236 from a content sample 226 and a style sample 230.
[0029] The content sample 226 includes a visual representation and / or model of one or more content-based attributes 240. For example, the content sample 226 may include one or more images, meshes, and / or other two-dimensional (2D) or three-dimensional (3D) depictions and / or abstract shapes of objects ( For example , faces, buildings, vehicles, animals, plants, roads, water, etc.), and / or other two-dimensional (2D) or three-dimensional (3D) depictions and / or abstract shapes ( For example , straight lines, squares, circles, curves, polygons, etc.). The content-based attributes 240 of the content sample 226 may include visual or physical attributes, hierarchies, or arrangements that distinguish these objects and / or shapes ( For example , a face is an object that includes a recognizable arrangement of eyes, ears, nose, mouth, hair, and / or other objects, and each object within the face is represented by a recognizable arrangement of lines, angles, polygons, and / or other abstract shapes).
[0030] The style sample 230 includes a visual representation and / or model of one or more style-based attributes 238. For example, the style sample 230 may include a drawing, painting, sketch, rendering, photograph, and / or another 2D or 3D depiction that is different from the content sample 226. The style-based attributes 238 in the style sample 230 may include, but are not limited to, brushstrokes, lines, edges, patterns, colors, bokeh, and / or other artistic or naturally occurring attributes that define the way the content is depicted.
[0031] In one or more embodiments, the execution engine 124 combines the content-based attributes 240 of the content sample 226 and the style-based attributes 238 of the style sample 230 into a style transfer result 236. More specifically, the execution engine 124 can provide the content sample 226 and the style sample 230 as inputs to the style transfer model 200 being trained, and the style transfer model 200 can extract the content-based attributes 240 from the content sample 226 and the style-based attributes 238 from the style sample 230. Then, the style transfer model 200 can generate a style transfer result 236 to have a predefined and / or user-controlled mixture or balance of the content-based attributes 240 from the content sample 226 and the style-based attributes 238 from the style sample 230.
[0032] As shown, the style transfer model 200 includes one or more encoders 202, 204, a kernel predictor 220, and a decoder 206. For a given content sample ( For example , the content sample 226), the encoder 202 can generate a latent representation 216 of the content sample. For a given style sample ( For example , the style sample 230), the encoder 204 can generate a latent representation 218 of the style sample. For example, each of the encoders 202, 204 can convert pixels, voxels, points, textures, and / or other information in the input sample ( For example , the style and / or content sample) into a plurality of vectors and / or matrices in a lower-dimensional latent space. Generally, the encoders 202 and 204 can be implemented as the same encoder or different encoders.
[0033] In some embodiments, the encoders 202, 204 include one or more parts of one or more pre-trained convolutional neural networks (CNNs). These pre-trained CNNs can include, but are not limited to, VGG, ResNet, Inception, MobileNet, DarkNet, AlexNet, GoogLeNet, and / or another type of deep CNN trained to perform image classification, object detection, and / or other tasks related to the content in large image datasets.
[0034] The encoders 202, 204 can include one or more layers from the same and / or different pre-trained CNNs. For example, each of the encoders 202, 204 can use the same set of layers from a pre-trained CNN to generate feature embeddings F c and F s . Each feature embedding can include multiple channels of a matrix of a certain size ( For example , 16×16, 8×8, etc.) For example, 512). In another example, the encoders 202, 204 may use different CNNs and / or layers to convert different types of data ( For example , 2D image data and 3D mesh data) into the feature embeddings F c and F s and / or generate feature embeddings with different sizes and / or numbers of channels from the corresponding content and style samples.
[0035] Each of the encoders 202, 204 optionally includes additional layers that further convert the output of the corresponding pre-trained CNN into a latent representation of the corresponding input sample ( For example , latent representations 216, 218). For example, encoder 202 may include one or more neural network layers that generate the latent representation 216 from the feature embedding F c as the normalized feature embedding F c N ( For example , by scaling and shifting the values in F c to have a specific mean and standard deviation). In another example, encoder 204 may include one or more neural network layers that generate the latent representation 218 by compressing the feature embedding F s into a vector W s in a d-dimensional "latent style space" associated with the corresponding style sample.
[0036] The kernel predictor 220 generates a plurality of convolutional kernels 222 from the latent representation 218 output by the encoder 204 from a given style sample. For example, the kernel predictor 220 may convert the latent representation 218 ( For example , vector W s ) into a plurality of n×n ( For example , 3×3) convolutional kernels 222 K s . The normalized feature embedding F c N generated by the encoder 202 from a given content sample and / or another latent representation 216 are convolved with K s to transfer the statistical and structural characteristics of the style sample to the latent representation 216 of the content sample. In some embodiments, the statistical characteristics include one or more statistical values associated with the visual attributes of the style sample, such as the mean and standard deviation of color, brightness, and / or sharpness in the style sample, regardless of where these attributes appear in the style sample. In some embodiments, the structural characteristics include the "spatial distribution" of patterns, geometric shapes, and / or other features in the style sample, which may be captured by some or all of the convolutional kernels 222.
[0037] In some embodiments, the kernel predictor 220 additionally generates a scalar bias for each output channel from each convolutional kernel. The bias can be added to the convolutional output produced by convolving a given input with the corresponding convolutional kernel included in the convolutional kernel 222.
[0038] In some embodiments, the kernel predictor 220 generates a plurality of convolutional kernels 222 that are applied in a varying resolution order to convey features at different levels of detail and / or granularity from the style sample. For example, the kernel predictor 220 may generate a first series of convolutional kernels 222 that produce a convolutional output at a first resolution. The latent representation 216 can be input into the first convolutional kernel in the first series to generate a convolutional output at the first resolution ( For example , a resolution higher than the latent representation 216), and the output of each kernel in the first series is used as the input into the next kernel in the first series to produce an additional convolutional output at the first resolution. The kernel predictor 220 may also generate a second series of convolutional kernels 222 that produce a convolutional output at a second resolution higher than the first resolution. The output of the last kernel in the first series is used as the input into the first kernel in the second series to produce a convolutional output at the second resolution, and the output of each kernel in the second series is used as the input into the next kernel in the second series to produce an additional convolutional output at the second resolution. Prior to performing the convolution with subsequent convolutional kernels, additional non-linear activations, fixed convolutional boxes, upsampling operations, and / or other types of layers or operations may be applied to the convolutional output of a given convolutional kernel. The additional series of convolutional kernels 222 are optionally generated from the latent representation 216 and convolved with the output from the previous convolutional kernels 222 to further increase the resolution of the convolutional output and / or to apply features associated with the style sample to the latent representation 216 of the content sample at an increased resolution. Thus, the kernel predictor 220 can "adapt" the convolutional kernels 222 to reflect the multi-level features in the style sample rather than using the same static set of convolutional kernels to perform the convolutions in the style transfer model 200.
[0039] The decoder 206 converts the convolutional output from the last convolutional kernel in K s into a visual representation and / or model of the content and / or style represented by the convolutional output. For example, the decoder 206 may include a CNN that applies additional convolutions and / or upsampling to the convolutional output to generate the decoder output 210 that includes an image, a mesh, and / or another 2D or 3D representation.
[0040] In one or more embodiments, some or all of the convolutions involving the latent representation 216 and the convolutional kernels 222 are integrated into the decoder 206. For example, during the conversion of the convolutional output into the decoder output 210, the decoder 206 may convolve the convolutional output generated from the latent representation 216 by one convolutional kernel or a series of convolutional kernels 222 with one or more additional series of convolutional kernels 222. Alternatively, all of the convolutional kernels 222 may be used in the layers of the decoder 206 to convert the latent representation 216 into the decoder output 210. Using the decoder 206 to perform some or all of the convolutions involving the latent representation 216 and the convolutional kernels 222 allows these convolutions to be performed at varying ( For example , increased) resolutions. In other words, after the convolutional kernels 222 have been generated from the latent representation 218 by the kernel predictor 220, the convolutional kernels 222 may be used by any component or layer of the style transfer model 200.
[0041] The training engine 122 trains the style transfer model 200 to perform style transfer between paired training content samples 224 and training style samples 228 in a set of training data 214. For example, the training engine 122 may generate each pair of samples by randomly selecting a training content sample from a set of training content samples 224 in the training data 214 and a training style sample from a set of training style samples 228 in the training data 214.
[0042] For each training content sample-training style sample pair selected from the training data 214, the training engine 122 inputs the training content sample into the encoder 202 and the training style sample into the encoder 204. Next, the training engine 122 inputs the latent representation 218 of the training style sample into the kernel predictor 220 to generate convolutional kernels 222 that reflect the feature maps associated with the training style sample, and convolves the latent representation 218 with the convolutional kernels 222 to produce a convolutional output. Then, the training engine 122 inputs the convolutional output into the decoder 206 to produce the decoder output 210 from the convolutional output. The training engine 122 also or alternatively uses some or all of the convolutional kernels 222 in one or more layers of the decoder 206 to convert the latent representation 216 and / or the convolutional output from the previous convolutional kernels 222 into the decoder output 210.
[0043] The training engine 122 updates the parameters of one or more components of the style transfer model 200 based on an objective function 212 that includes a style loss 232 and a content loss 234. As shown, the style loss 232 and the content loss 234 can be determined using the latent representations 216, 218, and the latent representation 242 generated by the encoder 208 from the decoder output 210. For example, the encoder 208 can include the same pre-trained CNN layers as the encoders 202 and / or 204. As a result, the encoder 208 can output the latent representation 242 in the same latent space and / or a similar latent space as the feature embeddings F c and F s and output the latent representation 242 in the same latent space and / or a similar latent space as the feature embeddings F
[0044] In one or more embodiments, the style loss 232 represents the difference between the latent representation 242 and the latent representation 218, and the content loss 234 represents the difference between the latent representation 242 and the latent representation 216. For example, the style loss 232 can be calculated as a measure of the distance between the latent representations 218 and 242 ( For example , cosine similarity, Euclidean distance, etc.), and the content loss 234 can be calculated as a measure of the distance between the latent representations 216 and 242.
[0045] The objective function 212 can thus include a weighted sum and / or another combination of the style loss 232 and the content loss 234. For example, the objective function 212 can be a loss function that includes the sum of the style loss 232 multiplied by one coefficient and the content loss 234 multiplied by another coefficient. The sum of the coefficients can be 1, and each coefficient can be selected to increase or decrease the presence of the style-based attributes 238 and the content-based attributes 240 in the decoder output 210.
[0046] In some embodiments, the style loss 232 and / or the content loss 234 are calculated using the features output by the respective layers of the encoders 202 and 204 and / or the decoder 206. For example, the style loss 232 and / or the content loss 234 can include a measure of the distance between the features generated by the earlier layers of the encoders 202 and 204 and / or the decoder 206, which capture the smaller features in the corresponding input ( For example , details, textures, edges, etc.). The style loss 232 and / or the content loss 234 can also or alternatively include a measure of the distance between the features generated by the subsequent layers of the encoders 202 and 204 and / or the decoder 206, which capture the more global features in the corresponding input ( For example , the overall shape of the object, parts of the object, etc.).
[0047] When the style loss 232 and / or the content loss 234 include a distance ( For exampleWhen calculating multiple metrics of the distance between features generated by different encoder layers, the objective function 212 can assign different weights to each metric. For example, the style loss 232 can include a higher weight or coefficient for the distance between the lower-level features generated from the decoder output 210 by earlier layers of the encoder 208 and the features generated from the style sample by the corresponding layers of the encoder 204, to increase the presence of "local" style-based attributes 238 such as lines, edges, strokes, colors, and / or patterns. Conversely, the content loss 234 can include a higher weight for the distance between the higher-level "global" features generated from the decoder output 210 by subsequent layers of the encoder 208 and the features generated from the content sample by the corresponding layers of the encoder 202 at a higher resolution, to increase the presence of all content-based attributes 238 such as the recognizable features or shapes of the object.
[0048] After calculating the style loss 232, the content loss 234, and the objective function 212 for one or more pairs of training content samples 224 and training style samples 228 in the training data 214, the training engine 122 updates the parameters of one or more components of the style transfer model 200 based on the objective function 212. For example, the training engine 122 can use training techniques ( For example , gradient descent and backpropagation) and / or one or more hyperparameters to iteratively update the weights of the kernel predictor 220 and / or the decoder 206 in a way that reduces the loss function ( For example , the objective function 212) associated with the style loss 232 and the content loss 234. In some embodiments, the hyperparameters define the higher-level characteristics of the style transfer model 200 and / or are used to control the training of the style transfer model 200. For example, the hyperparameters of the style transfer model 200 can include, but are not limited to, batch size, learning rate, number of iterations, number and size of the convolutional kernels 222 output by the kernel predictor 220, number of layers in each of the encoders 202 and 204 and the decoder 206, and / or a threshold for pruning weights in the neural network layers. Then, the decoder output 210 generated for subsequent pairs of training content samples 224 and training style samples 228 can include a proportion of style-based attributes 238 and content-based attributes 240, which reflects the weights and / or coefficients associated with the style loss 232 and the content loss 234 in the loss function.
[0049] After the training engine 122 has completed the training of the style transfer model 200, the execution engine 124 can execute the trained style transfer model 200 to generate a style transfer result 236 from new content samples 226 and style samples 230. For example, the execution engine 124 can take a content image ( For example , a facial image) and a style image ( For example, which need not be an artistic depiction of a face or scene) is input into the style transfer model 200, and a style transfer image including one or more style-based attributes 238 of the style image (independent of the content in the style image) and one or more content-based attributes 240 of the content image (independent of the style of the content image) is obtained as the output from the style transfer model 200. Thus, if the content image includes a face and the style image includes colors, edges, brushstrokes, lines, and / or other patterns representing a specific artistic style, the style transfer image may include shapes representing eyes, nose, mouth, ears, hair, face shape, accessories, and / or clothing associated with the face. These shapes may be drawn or rendered using the colors, edges, brushstrokes, lines, and / or patterns found in the style image, thereby transferring the "style" of the style image to the content of the content image.
[0050] In another example, the execution engine 124 may select a 3D mesh as the content sample 226 and select a different 3D mesh or 2D image as the style sample 230. After the content sample 226 and the style sample 230 are input into the style transfer model 200, the execution engine 124 may obtain a 3D mesh having a shape similar to the 3D mesh in the content sample 226 and a texture obtained from the 3D mesh or 2D image in the style sample 230 as the style transfer result 236. Then, the style transfer result 236 may be rendered into a 2D image representing a view of the 3D mesh textured with the 2D image.
[0051] The execution engine 124 may additionally include the function of generating style transfer results 236 for a series of related content samples and / or style samples. For example, the content samples may include a series of frames in a first 2D or 3D movie or animation, and the style samples may include one or more frames from a second 2D or 3D movie or animation. The execution engine 124 may use the style transfer model 200 to combine each frame in the content samples with a given artistic style in the style samples into a new series of frames, which includes the content from the first movie or animation and the style from the second movie or animation. This type of style transfer can be used to apply the style of a given movie to a related movie ( For example , prequels, sequels, etc.) and / or jump between different styles within the same movie ( For example , by combining scenes in the movie with different style samples). Thus, the style transfer model 200 may allow 2D or 3D content to adapt to different and / or new styles without the need to manually recreate or modify the content to reflect the desired style.
[0052] Figure 3 A flowchart of method steps for training a style transfer model according to various embodiments. Although combined with Figure 1-2method steps of system description, but those skilled in the art will understand that any system configured to execute method steps in any order falls within the scope of the present disclosure.
[0053] As shown in the figure, in operation 302, the training engine 122 selects a training style sample and a training content sample from a set of training data for the style transfer model. For example, the training engine 122 may randomly select a training style sample from a set of training style samples in the training data. The training engine 122 may also randomly select a training content sample from a set of training content samples in the training data.
[0054] Next, in operation 304, the training engine 122 applies the style transfer model to the training style sample and the training content sample to generate a style transfer result. For example, the training engine 122 may use one or more encoder networks to convert the training style sample and the training content sample into latent representations. Next, the training engine 122 may use one or more layers of the kernel predictor to generate a series of convolutional kernels from the latent representation of the training style sample. Then, the training engine 122 may convolve the latent representation of the training content sample with the convolutional kernels to generate a convolutional output, and use a decoder network to convert the convolutional output into a style transfer result.
[0055] In operation 306, the training engine 122 also updates one or more sets of weights in the style transfer model based on one or more losses calculated between the style transfer result and the training content sample and / or the training style sample. For example, the training engine 122 may calculate a style loss between the style transfer result and the latent representation of the training style sample and a content loss between the style transfer result and the latent representation of the training content sample. Then, the training engine 122 may calculate the total loss as a weighted sum of the style loss and the content loss, and use gradient descent and backpropagation to update the parameters of the kernel predictor and the decoder network in a manner that reduces the total loss.
[0056] After operations 302, 304, and 306 are completed, the training engine 122 may evaluate a condition 308 indicating whether the training of the style transfer model is complete. For example, the condition 308 may include, but is not limited to, convergence of the parameters of the style transfer model, reduction of the style and / or content loss below a threshold, and / or execution of a certain number of training steps, iterations, batches, and / or epochs. If the condition 308 is not met, then the training engine 122 may continue to select pairs of training style samples and training content samples from the training data (operation 302), input the training style samples and the training content samples into the style transfer model to generate a style transfer result (operation 304), and update the weights of one or more neural networks and / or neural network layers in the style transfer model (operation 306). If the condition 308 is met, then the training engine 122 ends the process of training the style transfer model.
[0057] Figure 4 A flowchart of method steps for performing style transfer according to various embodiments. Although the method steps are described in connection with Figure 1-2 a system, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present disclosure.
[0058] As shown, in operation 402, the execution engine 124 applies an encoder network and / or one or more additional neural network layers to a style sample and a content sample to generate a first latent representation of the style sample and a second latent representation of the content sample. For example, the content and style samples may include images, meshes, and / or other 2D or 3D representations of objects, textures, or scenes. The execution engine 124 may use a pre-trained encoder, such as VGG, ImageNet, ResNet, GoogLeNet, and / or Inception, to convert the style sample and the content sample into two independent feature maps. The execution engine 124 may use a multi-layer perceptron to compress the feature map for the style sample into a latent style vector and use the latent style vector as the first latent representation of the style sample. The execution engine 124 may normalize the feature map for the content sample and use the normalized feature map as the second latent representation of the content sample.
[0059] Next, in operation 404, the execution engine 124 applies one or more neural network layers in a kernel predictor to the first latent representation to generate one or more convolutional kernels. For example, the execution engine 124 may use the kernel predictor to generate one or more series of convolutional kernels, where each series of convolutional kernels is used to produce an output at a corresponding resolution. The execution engine 124 may also generate one or more biases as additional outputs of one or more neural network layers to be applied after some or all of the convolutional kernels.
[0060] In operation 406, the execution engine 124 generates a convolutional output by convolving the second latent representation of the content sample with the convolutional kernels. For example, the execution engine 124 may convolve the second latent representation with a first kernel to produce a first output matrix at a first resolution. The execution engine 124 may apply one or more additional layers and / or operations to the first output matrix to produce a modified output matrix and then convolve the modified output matrix with one or more additional convolutional kernels to produce a second output matrix at a second resolution higher than the first resolution. As a result, the execution engine 124 may apply features at different resolutions extracted from the style sample to the second latent representation of the content sample.
[0061] In operation 408, the execution engine 124 applies one or more decoder layers to the convolutional output to produce a style transfer result that includes one or more content-based attributes of the content samples and one or more style-based attributes of the style samples. For example, the execution engine 124 may use convolutional and / or upsampling layers in the decoding network to convert the convolutional output into an image, a grid, and / or another 2D or 3D representation. The representation may include the shape and / or other discriminative attributes of the object in the content samples, as well as the color, pattern, stroke, line, edge, and / or other depictions of the style in the style samples.
[0062] As described above, some or all of the convolutions performed in operation 406 may be integrated into operation 408. For example, during the conversion of the convolutional output into the style transfer result, some or all of the decoder layers may be used to convolve the convolutional output generated by one convolutional kernel or series of convolutional kernels with one or more additional series of convolutional kernels. Additionally, all convolutional kernels may be used in the decoder layer to convert a second latent representation of the content samples into the style transfer result. Thus, after the convolutional kernels have been generated from the first latent representation of the style samples, the convolutional kernels may be used by any component, layer, or operation.
[0063] Adaptive Convolution in Neural Networks
[0064] Although the adaptive convolution technique has been described above for style transfer, the convolutional kernel 222 may be used by the training engine 122, the execution engine 124, and / or other components in various applications related to decoding operations in the neural network. In the following, reference Figure 5 describes the general use of the adaptive convolutional kernel 222 in neural network decoding operations and additional applications of the adaptive convolutional kernel 222 in neural network decoding operations.
[0065] Figure 5 FIG. is a flowchart of method steps for performing adaptive convolution in a neural network according to various embodiments. Although the method steps are described in conjunction with Figure 1-2 the system, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present disclosure.
[0066] As shown in the figure, in operation 502, the execution engine 124 and / or another component apply one or more neural network layers in the kernel predictor to the first input to generate one or more convolutional kernels. For example, the first input can be generated by an encoder network and / or another type of neural network as a feature map, embedding, encoding, and / or other representation of the first set of data. The component can use the kernel predictor to generate one or more series of convolutional kernels from the first input, where each series of convolutional kernels is used to produce an output at a corresponding resolution. The component can also generate one or more biases as additional outputs of one or more neural network layers to be applied after some or all of the convolutional kernels.
[0067] Next, in operation 504, the component generates a convolutional output by convolving the second input with the convolutional kernels. For example, the component can use the convolutional kernels to apply features extracted from the first input at different resolutions to the second input.
[0068] In operation 506, the component applies one or more decoder layers to the convolutional output to produce a decoded result that includes one or more attributes associated with the first input and one or more attributes associated with the second input. For example, the component can apply the decoder layers to the convolutional output generated after all convolutional kernels have been convolved with the second input, or the component can use some or all of the convolutional kernels in the decoder layers to generate the decoded result from the second input.
[0069] In one or more embodiments, operations 502, 504, and 506 are performed in the context of a generative model such as a generative adversarial network (GAN) to control and / or adjust the generation of images, text, audio, and / or other types of outputs of the generative model. For example, the GAN can include a Style Generative Adversarial Network (StyleGAN) or a StyleGAN2 model, and the first input can include a latent code w generated by the mapping network in the StyleGAN or StyleGAN2 model based on a sample z from the distribution of latent variables learned by the mapping network. Within the StyleGAN model, each learned affine transformation "A" and adaptive instance normalization (AdaIN) block (which performs AdaIN on the output of each convolutional layer in the synthesis network g using the latent code) can be replaced with a corresponding kernel predictor and multiple convolutional kernels generated by the kernel predictor from the latent code ( For example , depthwise 3×3 convolution, pointwise 1×1 convolution, and per-channel bias). Similarly, within the StyleGAN2 model, each weight demodulation block in the synthesis network (which converts the latent code into a demodulation operation applied to the corresponding 3×3 convolution) can be replaced with a kernel predictor and the corresponding convolutional kernels generated by the kernel predictor from the latent code ( For example , depthwise 3×3 convolution, pointwise 1×1 convolution, and per-channel bias).
[0070] Continuing with the above example, standard techniques can be used to train a StyleGAN or StyleGAN2 model, and operation 502 can be performed to generate convolutional kernels from the latent code at each layer of the synthesis network. In operations 504 and 506, the convolutional kernels can be applied within the synthesis network to a second input that includes a constant input c, an upsampled input from the previous layer in the synthesis network, and / or a Gaussian noise input that includes a per-channel scaling factor "B" applied to the upsampled input. At each resolution level of the synthesis network, a 1×1 convolution can be used to convert the output of the last layer into RGB to produce an image that is added to the upsampled RGB result of the previous layer; this gives the decoded result at the current resolution.
[0071] Operations 502, 504, and 506 can also or alternatively be performed in the context of generating or modifying a 2D or 3D scene. For example, operation 502 can be performed to generate one or more series of convolutional kernels from a first input that includes camera parameters ( For example , camera model, camera pose, focal length, etc.), lighting parameters ( For example , light source, lighting interaction, lighting model, shadows, etc.), and / or an embedded and / or encoded representation of other types of parameters that affect scene rendering or appearance. In operation 504 and / or 506, the convolutional kernels can be applied to a second input that includes points, pixels, textures, feature embeddings, and / or other representations of the scene. The convolutional kernels can be applied before decoding in operation 506, or some or all of the convolutional kernels can be applied by one or more decoder layers during operation 506. The output of the decoder layer can include a representation of a 2D or 3D scene ( For example , image, mesh, point cloud, etc.). This representation can include objects, shapes, and / or structures from the second input that are depicted in a way that reflects the camera, lighting, and / or other types of parameters from the first input.
[0072] In summary, the disclosed techniques utilize deep learning and adaptive convolutions and decoding operations in neural networks, such as decoding operations that perform style transfer between content samples and style samples. The content samples and style samples can include (but are not limited to) one or more images, meshes, and / or other depictions or models of objects, scenes, or concepts. An encoder network can be used to convert the content samples and style samples into latent representations in a lower-dimensional space. A kernel predictor generates multiple convolutional kernels from the latent representation of the style sample such that the convolutional kernels are "adapted" to capture features at varying resolutions or granularities in the style sample. Then, the latent representation of the content sample is convolved with the convolutional kernels to produce convolutional outputs at different resolutions, and some or all of the convolutional outputs are converted into a style transfer result that incorporates the content of the content sample and the style of the style sample.
[0073] Advantageously, by identifying features at varying resolutions in the style sample and transferring those features to the content sample ( For example , by convolving the features with a latent representation of the content sample), the disclosed techniques allow both low-level and high-level style attributes in the style sample to be included in the style transfer result. Accordingly, compared to traditional style transfer results that only incorporate the global statistics of the style sample into the content of the content sample, the style transfer result can include a better depiction of the style in the style sample. Compared to existing techniques for generating content in a certain style, the disclosed techniques provide additional improvements in terms of overhead and / or resource consumption. For example, traditional techniques for adapting an image, video, and / or other content to a new style can involve the user manually capturing, creating, editing, and / or rendering the content in the new style. The drawing, modeling, editing, and / or other tools that the user uses to create, update, and store the content can consume a large amount of computing, memory, storage, network, and / or other resources. In contrast, the disclosed techniques can perform batch processing that uses a style transfer model to automatically transfer the style to the content, which consumes less time and / or resources than the manual creation or modification of the content performed in traditional techniques. Accordingly, by automating the transfer of style to content and improving the comprehensiveness and accuracy of style transfer, the disclosed embodiments provide a technical improvement to computer systems, applications, frameworks, and / or techniques for generating content and / or performing style transfer.
[0074] 1. In some embodiments, a method for performing style transfer between a content sample and a style sample includes applying one or more neural network layers to a first latent representation of the style sample to generate one or more convolutional kernels, generating a convolutional output by convolving a second latent representation of the content sample with the one or more convolutional kernels, and applying one or more decoder layers to the convolutional output to produce a style transfer result that includes one or more content-based attributes of the content sample and one or more style-based attributes of the style sample.
[0075] 2. The method according to clause 1, further comprising updating a first set of weights in the one or more neural network layers and a second set of weights in the one or more decoder layers based on one or more losses computed between the style transfer result and at least one of the content sample or the style sample.
[0076] 3. The method according to clause 1 or 2, wherein the one or more losses include a style loss between a third latent representation of the style transfer result and the first latent representation of the style sample.
[0077] 4. The method according to any one of clauses 1-3, wherein the one or more losses include a content loss between a third latent representation of the style transfer result and the second latent representation of the content sample.
[0078] 5. The method according to any one of clauses 1-4, wherein the one or more losses include a weighted sum of a first loss between the style transfer result and the style sample and a second loss between the style transfer result and the content sample.
[0079] 6. The method according to any one of clauses 1-5, further comprising applying an encoder network to the content sample to produce the second latent representation as a feature embedding of the content sample.
[0080] 7. The method according to any one of clauses 1-6, further comprising generating one or more biases to be applied after the one or more convolutional kernels as an additional output of the one or more neural network layers.
[0081] 8. The method according to any one of clauses 1-7, further comprising applying an encoder network to the style sample to produce a feature embedding of the style sample, and inputting the feature embedding into one or more additional neural network layers to produce the first latent representation as a latent style vector.
[0082] 9. The method according to any one of clauses 1-8, wherein generating the convolutional output includes convolving the second latent representation with a first kernel to produce a first output matrix at a first resolution, applying one or more additional neural network layers to the first output matrix to produce a modified output matrix, and convolving the modified output matrix with one or more additional convolutional kernels to produce a second output matrix at a second resolution higher than the first resolution.
[0083] 10. The method according to any one of clauses 1-9, wherein at least a part of the convolutional output is generated using the one or more decoder layers.
[0084] 11. The method according to any one of clauses 1-10, wherein the content sample and the style sample include at least one of an image or a mesh.
[0085] 12. In some embodiments, a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform the following steps: applying one or more neural network layers to a first latent representation of a style sample to generate one or more convolutional kernels, generating a convolutional output by convolving a second latent representation of a content sample with the one or more convolutional kernels, and applying one or more decoder layers to the convolutional output to produce a style transfer result that includes one or more content-based attributes of the content sample and one or more style-based attributes of the style sample.
[0086] 13. The non-transitory computer-readable medium according to clause 12, wherein when executed by the processor, the instructions further cause the processor to perform the following steps: updating a first set of weights in the one or more neural network layers and a second set of weights in the one or more decoder layers based on one or more losses computed between the style transfer result and at least one of the content sample or the style sample.
[0087] 14. The non-transitory computer-readable medium according to clause 12 or 13, wherein the one or more losses include a weighted sum of a style loss between a third latent representation of the style transfer result and the first latent representation of the style sample and a content loss between the third latent representation of the style transfer result and the second latent representation of the content sample.
[0088] 15. The non-transitory computer-readable medium according to any one of clauses 12-14, wherein when executed by the processor, the instructions further cause the processor to perform the following steps: applying an encoder network to the style sample to produce a first feature embedding of the style sample, and inputting the first feature embedding into one or more additional neural network layers to produce the first latent representation as a latent style vector.
[0089] 16. The non-transitory computer-readable medium according to any one of clauses 12-15, wherein when executed by the processor, the instructions further cause the processor to perform the following steps: applying the encoder network to the content sample to produce the second latent representation as a second feature embedding of the content sample, and normalizing the second latent representation before generating the convolutional output.
[0090] 17. The non-transitory computer-readable medium according to any one of clauses 12-16, wherein generating the convolutional output includes convolving the second latent representation with a first kernel to produce a first output matrix at a first resolution, applying one or more additional neural network layers to the first output matrix to produce a modified output matrix, and convolving the modified output matrix with one or more additional convolutional kernels to produce a second output matrix at a second resolution higher than the first resolution.
[0091] 18. The non-transitory computer-readable medium according to any one of clauses 12-17, wherein the one or more content-based attributes include a recognizable arrangement representing an abstract shape of an object in the content image.
[0092] 19. The non-transitory computer-readable medium according to any one of clauses 12-18, wherein the one or more style-based attributes include at least one of lines, edges, strokes, colors, or patterns in the style image.
[0093] 20. In some embodiments, a system includes: a memory that stores instructions, and a processor that is coupled to the memory and, when executing the instructions, is configured to apply an encoder network to a style image and a content image to generate a first latent representation of the style image and a second latent representation of the content image, apply one or more neural network layers to the first latent representation to generate one or more convolutional kernels, generate a convolutional output by convolving the second latent representation with the one or more convolutional kernels, and apply one or more decoder layers to the convolutional output to produce a style transfer result that includes one or more content-based attributes of the content sample and one or more style-based attributes of the style sample.
[0094] 21. In some embodiments, a method for performing convolution within a neural network includes applying one or more neural network layers to a first input to generate one or more convolutional kernels, generating a convolutional output by convolving a second input with the one or more convolutional kernels, and applying one or more decoder layers to the convolutional output to produce a decoded result, wherein the decoded result includes one or more first attributes of the first input and one or more second attributes of the second input.
[0095] 22. The method according to clause 21, wherein the first input includes one or more samples from a latent distribution associated with a generator network, and the second input includes one or more noise samples from one or more noise distributions.
[0096] 23. The method according to clause 21 or 22, wherein the one or more convolutional kernels comprise depthwise convolution, pointwise convolution, and per-channel bias.
[0097] 24. The method according to any one of clauses 21-23, wherein the second input comprises a representation of a scene, and the first input comprises one or more parameters for controlling the depiction of the scene.
[0098] 25. The method according to any one of clauses 21-24, wherein the one or more parameters comprise at least one of an illumination parameter and a camera parameter.
[0099] Any and all combinations of any claim elements recited in any claim and / or any elements described in any way in this application fall within the intended scope of the invention and protection.
[0100] For purposes of illustration, descriptions of various embodiments have been presented, but these descriptions are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to a person of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
[0101] Aspects of the present embodiment may be implemented as a system, method, or computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may generally be referred to herein as a "module", "system", or "computer". Additionally, any hardware and / or software technologies, processes, functions, components, engines, modules, or systems described in the present disclosure may be implemented as one circuit or a set of circuits. Further, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.
[0102] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0103] The foregoing has described, with reference to flowcharts illustrations and / or block diagrams, aspects of the present disclosure of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / acts specified in the flowchart and / or one or more block diagrams. Such a processor may be, but is not limited to, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable gate array.
[0104] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a special purpose hardware-based system that performs the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0105] Although the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, and the scope of the present disclosure is determined by the appended claims.
Claims
1. A method for performing style transfer between a content sample and a style sample, comprising: Applying one or more neural network layers to a first latent representation of the style sample to generate one or more convolutional kernels; Generating a convolutional output by convolving a second latent representation of the content sample with the one or more convolutional kernels; and Applying one or more decoder layers to the convolutional output to produce a style transfer result that includes one or more content-based attributes of the content sample and one or more style-based attributes of the style sample.
2. The method according to claim 1, further comprising updating a first set of weights in the one or more neural network layers and a second set of weights in the one or more decoder layers based on one or more losses calculated between the style transfer result and at least one of the content sample or the style sample.
3. The method according to claim 2, wherein the one or more losses include a style loss between a third latent representation of the style transfer result and the first latent representation of the style sample.
4. The method according to claim 2, wherein the one or more losses include a content loss between a third latent representation of the style transfer result and the second latent representation of the content sample.
5. The method according to claim 2, wherein the one or more losses include a weighted sum of a first loss between the style transfer result and the style sample and a second loss between the style transfer result and the content sample.
6. The method according to claim 1, further comprising applying an encoder network to the content sample to produce the second latent representation as a feature embedding of the content sample.
7. The method according to claim 6, further comprising generating one or more biases to be applied after the one or more convolutional kernels as an additional output of the one or more neural network layers.
8. The method according to claim 1, further comprising: Applying an encoder network to the style sample to produce a feature embedding of the style sample; and Inputting the feature embedding into one or more additional neural network layers to produce the first latent representation as a latent style vector.
9. The method according to claim 1, wherein generating the convolutional output includes: Convolving the second latent representation with a first kernel to produce a first output matrix at a first resolution; Applying one or more additional neural network layers to the first output matrix to produce a modified output matrix; and Convolving the modified output matrix with one or more additional convolutional kernels to produce a second output matrix at a second resolution higher than the first resolution.
10. The method according to claim 1, wherein the one or more decoder layers are used to generate at least a portion of the convolutional output.
11. The method according to claim 1, wherein the content sample and the style sample include at least one of an image or a mesh.
12. The method according to claim 1, wherein the one or more convolutional kernels comprise depthwise convolution, pointwise convolution, and per-channel bias.
13. The method according to claim 1, wherein the content sample comprises a representation of a scene, and the style sample comprises one or more parameters for controlling the depiction of the scene.
14. The method according to claim 13, wherein the one or more parameters comprise at least one of an illumination parameter and a camera parameter.
15. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform the following steps: Apply one or more neural network layers to a first latent representation of a style sample to generate one or more convolutional kernels; Generate a convolutional output by convolving a second latent representation of a content sample with the one or more convolutional kernels; and Apply one or more decoder layers to the convolutional output to produce a style transfer result comprising one or more content-based attributes of the content sample and one or more style-based attributes of the style sample.
16. The non-transitory computer-readable medium according to claim 15, wherein when executed by the processor, the instructions further cause the processor to perform the following steps: Update a first set of weights in the one or more neural network layers and a second set of weights in the one or more decoder layers based on one or more losses computed between the style transfer result and at least one of the content sample or the style sample.
17. The non-transitory computer-readable medium according to claim 16, wherein the one or more losses comprise a weighted sum of a style loss between a third latent representation of the style transfer result and the first latent representation of the style sample and a content loss between the third latent representation of the style transfer result and the second latent representation of the content sample.
18. The non-transitory computer-readable medium according to claim 15, wherein when executed by the processor, the instructions further cause the processor to perform the following steps: Apply an encoder network to the style sample to produce a first feature embedding of the style sample; and Input the first feature embedding into one or more additional neural network layers to produce the first latent representation as a latent style vector.
19. The non-transitory computer-readable medium according to claim 18, wherein when executed by the processor, the instructions further cause the processor to perform the following steps: Apply the encoder network to the content sample to produce the second latent representation as a second feature embedding of the content sample; and Normalize the second latent representation before generating the convolutional output.
20. The non-transitory computer-readable medium according to claim 15, wherein generating the convolutional output comprises: Convolving the second latent representation with a first kernel to produce a first output matrix at a first resolution; Apply one or more additional neural network layers to the first output matrix to produce a modified output matrix; and Convolve the modified output matrix with one or more additional convolutional kernels to produce a second output matrix at a second resolution higher than the first resolution.
21. The non-transitory computer-readable medium according to claim 15, wherein the one or more content-based attributes include recognizable arrangements representing the abstract shapes of the objects in the content sample.
22. The non-transitory computer-readable medium according to claim 15, wherein the one or more style-based attributes include at least one of lines, edges, strokes, colors, or patterns in the style sample.
23. A system, comprising:[[]] A memory that stores instructions, and A processor coupled to the memory and configured to, when executing the instructions:[[]] Apply an encoder network to a style image and a content image to generate a first latent representation of the style image and a second latent representation of the content image; Apply one or more neural network layers to the first latent representation to generate one or more convolutional kernels; Generate a convolutional output by convolving the second latent representation with the one or more convolutional kernels; and Apply one or more decoder layers to the convolutional output to produce a style transfer result that includes one or more content-based attributes of the content image and one or more style-based attributes of the style image.
Citation Information
Patent Citations
A training method of a convolutional neural network for image style migration and an image style migration method
CN109766895A
Image encoding method and apparatus and image decoding method and apparatus
US20200327701A1