Temporal consistent semantic color transfer from multiple multi-modality style references

The semantic color transfer system addresses labor-intensive manual grading by using deep convolutional networks and whitening transforms to achieve consistent color and brightness across multiple images/videos, ensuring temporal stability and reducing artifacts.

WO2025221551A1PCT designated stage Publication Date: 2025-10-23DOLBY LABORATORIES LICENSING CORP

Patent Information

Application Number
PCT/US2025/023936
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-02
Filing Date
2025-04-09
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing color consistency and manipulation across multiple image/video footages rely heavily on labor-intensive manual grading, making it difficult to achieve high-quality results, especially for local color grading.

Method used

A semantic color transfer system using deep convolutional networks and whitening and coloring transforms to automatically transfer color and brightness from selected regions of style images/videos to content images/videos, while maintaining temporal stability and reducing spatial artifacts.

Benefits of technology

Enables efficient, automatic color and brightness transfer across multiple perspectives, ensuring consistent color and brightness levels across different viewpoints and reducing flickering artifacts, with a lightweight architecture that maintains image details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000010_0001
    Figure IMGF000010_0001
  • Figure IMGF000011_0001
    Figure IMGF000011_0001
  • Figure IMGF000012_0001
    Figure IMGF000012_0001
Patent Text Reader

Abstract

Correspondence relationships are established between first semantic classes and second semantic classes. A first convolutional layer initialized with first pretrained weights is used to extract specific content features in the first semantic regions of content images and to extract specific style features in the second semantic regions of style images. Based on (a) the correspondence relationships and the specific content and style features, a semantic style transform is applied to transfer a specific color and brightness look associated with the second semantic regions of the style images to the specific content features in the first semantic regions of the content images. Based on the specific content features with the transferred specific color and brightness look, a decoder generates semantic content transferred images.
Need to check novelty before this filing date? Find Prior Art

Description

TEMPORAL CONSISTENT SEMANTIC COLOR TRANSFER FROM MULTIPLE MULTI-MODALITY STYLE REFERENCES CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from Indian Patent Application No. 202411030324, filed on 15 April 2024, and European Patent Application No.24185929.7, each of which is incorporated by reference herein in its entirety. TECHNICAL FIELD

[0002] The present invention relates generally to video coding and more particularly to semantic color transfer. BACKGROUND

[0003] Existing solutions for color consistency and color manipulation cross multiple image / video footages heavily rely on time consuming and labor intensive manual grading using color grading tools. When local color grading / transfer is desired, the tasks become even more difficult to carry out to high quality completion.

[0004] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Similarly, issues identified with respect to one or more approaches should not assume to have been recognized in any prior art on the basis of this section, unless otherwise indicated.

[0005] BRIEF DESCRIPTION OF DRAWINGS

[0006] The present invention is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings and in which like reference numerals refer to similar elements and in which:

[0007] FIG.1A and FIG.1B illustrate example images generated with color and brightness transfer;

[0008] FIG.2A illustrates an example VGG architecture; FIG.2B illustrates an example color and brightness transfer deep learning neural network; FIG.2C illustrates an example architecture for decoder training; FIG.2D illustrates example online training of a segmentation mask refiner network; FIG.2E illustrates an example online trained segmentation mask refiner network;

[0009] FIG.3A through FIG.3D illustrate example configuration for transferring semantic color and brightness looks from style image(s) / video(s) to content image(s) / video(s);

[0010] FIG.4 illustrates an example process flow; and

[0011] FIG.5 illustrates an example hardware platform on which a computer or a computing device as described herein may be implemented. DESCRIPTION OF EXAMPLE EMBODIMENTS

[0010] Example embodiments, which relate to semantic color transfer, are described herein. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail, in order to avoid unnecessarily occluding, obscuring, or obfuscating the present invention.

[0011] Example embodiments are described herein according to the following outline: 1. GENERAL OVERVIEW 2. VGG ARCHITECTURE 3. WHITENING AND COLORING TRANSFORM 4. SEMANTIC SEGMENTATION 5. GUIDED FILTER 6. COLOR TRANSFER NETWORK 7. PRE- AND POST-PROCESSING 8. COLOR TRANSFER NETWORK TRAINING 9. ONLINE TRAINING FOR FINETUNING SEGMENTATION MASKS 10. ACCELERATING INFERENCING AND ONLINE TRAINING 11. SEMANTIC TRANSFER APPLICATIONS 12. SINGLE STYLE IMAGE AND SINGLE CONTENT IMAGE 13. SINGLE STYLE IMAGE AND CONTENT VIDEO 14. SINGLE STYLE VIDEO AND CONTENT VIDEO 15. MULTIPLE STYLE IMAGES AND CONTENT VIDEO 16. EXAMPLE PROCESS FLOWS 17. IMPLEMENTATION MECHANISMS – HARDWARE OVERVIEW 18. EQUIVALENTS, EXTENSIONS, ALTERNATIVES ANDMISCELLANEOUS 1. GENERAL OVERVIEW

[0012] This overview presents a basic description of some aspects of an example embodiment of the present invention. It should be noted that this overview is not an extensive or exhaustive summary of aspects of the example embodiment. Moreover, it should be noted that this overview is not intended to be understood as identifying any particularly significant aspects or elements of the example embodiment, nor as delineating any scope of the example embodiment in particular, nor the invention in general. This overview merely presents some concepts that relate to the example embodiment in a condensed and simplified format, and should be understood as merely a conceptual prelude to a more detailed description of example embodiments that follows below.

[0013] Techniques as described herein can be used to implement or provide a semantic color transfer system or framework that allows users to prepare preferred or chosen color semantics from selected style images / videos and specify correspondence or paired-up (color transfer) semantics for input or original image / video content of different color semantics to target image / video content of the preferred or chosen color semantics with corresponding color (schemes) automatically transferred. These techniques can be applied to a wide range of applications or operational scenarios using or incorporating color scheme transfer tools as described herein. For example, some or all of these techniques or color scheme transfer tools may be used in or incorporated by color grading tools, video editing tools, content creation tools, content enhancement systems or components, content consumption applications for customized user preferred color themes in consumer devices, etc. Additionally, optionally or alternatively, these techniques or tools can be applied in multiple perspective audio and video scenarios, for example with multiple preferred or chosen color schemes.

[0014] Under the techniques as described herein, a novel algorithm or method may be implemented or performed to effectuate color / brightness transfer (without distorting the textures or semantic image details) from selected or identified semantic regions of (reference) style image(s) / video(s) to selected or identified semantic regions of source (or resultant target) image(s) / video(s).

[0015] As used herein, “a semantic region” refers to specific selected or identified – e.g., automatically by trained and / or finetuned artificial intelligence (AI) or machine learning (ML) models or artificial neural networks or ANNs – spatial portion(s) of an image or video that belong to the same group of semantic categories, like tree, sky, sea, etc. Hence, the semantic region refers to visually perceptible image details such as visual objects, visual characters, visualscene elements, etc., - which may be delineated with irregular and visually perceptible spatial shapes, contours, boundaries, etc., rather than delineated with artificial or regular shapes or visually non-perceptible shapes, contours, boundaries, etc. – to a human user.

[0016] The semantic color / brightness transfer refers to transferring to or changing source semantic regions (e.g., sky / sea / tree, etc.) of source image / video content with the same or similar color / brightness look of a reference semantic region of a reference image / video content, where the source semantic regions have a correspondence relationship – established automatically by a semantic color / brightness transfer tool and / or with user input – to the reference semantic region.

[0017] FIG.1A and FIG.1B illustrate example images generated by a color transfer algorithm, method or process flow as described herein. Under other approaches, style transfer mainly focuses on transferring the overall perceptual style from the style image to the content image. In comparison, the color / brightness transfer techniques as described herein can be used to transfer styles at semantic content level.

[0018] More specifically, FIG.1A illustrates a first target image (the lower right image of FIG.1A) generated from semantic color / brightness transfer operations that transfer or change source semantic regions of a first source image (the lower left image of FIG.1A) with the same or similar color / brightness look of a reference semantic region (in a reference image) in the same semantic class as that of the source semantic regions of the source image.

[0019] As shown, there are two reference or style images: a first reference or style image at the upper left of FIG.1A and a second reference or style image at the upper right of FIG.1A. From the first reference or style image in the upper left of FIG.1A, the color / brightness look of the sky is transferred to the sky of the first source (content) image in the lower left of FIG.1A to result in the sky of the first target (content) image in the lower right of FIG.1A. From the second reference or style image in the upper right of FIG.1A, the color / brightness look of the water is transferred to the water of the first source (content) image in the lower left of FIG.1A to result in the water of the first target (content) image in the lower right of FIG.1A.

[0020] The semantic color / brightness transfer as illustrated in FIG.1A may be referred to as “same semantic class wise color transfer” or simply “same semantic class transfer.”

[0021] In comparison, FIG.1B illustrates a second target image generated from semantic color / brightness transfer that transfers or changes source semantic regions of a second source image with the same or similar color / brightness look of a reference semantic region (in a reference image) in a different semantic class from that of the source semantic regions of the source image.

[0022] As shown, there are also two reference or style images: a third reference or styleimage at the upper left of FIG.1B and a fourth reference or style image at the upper right of FIG.1B. From the third reference or style image in the upper left of FIG.1B, the color / brightness look of the sky is transferred to the water of the second source (content) image in the lower left of FIG.1B to result in the water of the second target (content) image in the lower right of FIG.1B. From the fourth reference or style image in the upper right of FIG.1B, the color / brightness look of the water is transferred to the sky of the second source (content) image in the lower left of FIG.1B to result in the sky of the second target (content) image in the lower right of FIG.1B.

[0023] The semantic color / brightness transfer as illustrated in FIG.1B may be referred to as “cross semantic class color transfer” or simply “cross sematic class transfer.”

[0024] There is a wide range of applications or operational scenarios for the color / brightness transfer system or framework as described herein. With the wide deployment of mobile devices, it is very often that multiple users capture their videos and audios in the same event (such as concerts and sports); or a user uses multiple cameras to shoot from multiple angles under the same scene (such as cooking shows, teaching video to repair cars, etc.). In those application scenarios, there are multi-perspective audios and videos (MPAV). MPAV has many useful applications, such as crowd event reconstructions, influencers channels in social media, and educational purposes (surgery, house repair), and can provide a more interactive and intuitive way to engage the end users. In those scenarios, the color and brightness level of the same (e.g., semantic, etc.) content may change or vary visually due to content capturing from multiple perspectives to the same event(s) or scene(s) and there may also be differences in rendering or production of the same scene(s) from each camera’s image / video processing and / or delivery pipeline.

[0025] Another relatively common scenario is outdoor video shooting where the lighting conditions may change or vary depending on the environmental or weather or ambient conditions which produce varying color and brightness levels in the captured footage or image / video content. To have a good immersive experience of perceiving a content, the color and brightness level may be made relatively consistent across different perspectives under techniques as described herein. The switching / transition between different perspectives may be made relatively smooth with little or no visually perceptible differences (in color / brightness looks). There can also be many application or operational scenarios relating to multimedia entertainment in which the color and brightness levels between multiple scenes may be made relatively consistent to provide or produce the same or similar perceptual experience. Another use case is to transform or transfer the relatively modern or contemporary look of sports contentto older archival sports content such as colorization of existing gray-level content to result in enhanced archival sports contents with the relatively modern or contemporary look. In all these application or operational scenarios, the semantic color and brightness transfer techniques as described herein can be used to play a significant or enabling role.

[0026] The color / brightness transfer system or framework may include or operate with very deep convolutional networks for large-scale image recognition (denoted as VGG), which are described in Simonyan, Karen, and Andrew Zisserman, “Very deep convolutional networks for large-scale image recognition,” ICLR, 2015 (available at https: / / arxiv.org / abs / 1409.1556; accessed on April 2, 2024), the contents of which are incorporated by reference herein. Additionally, optionally or alternatively, the system or framework may include or operate with whitening coloring transform (denoted as WCT), which are described in Li, Yijun, Chen Fang, Jimei Yang, Zhaowen Wang, Xin Lu, and Ming-Hsuan Yang, “Universal style transfer via feature transforms,” Advances in neural information processing systems 30 (2017), the contents of which are incorporated by reference herein.

[0027] VGG may be incorporated or included in the system or framework as a simple 4- layer network along with one (e.g., single, etc.) WCT block to achieve a relatively high performance color / brightness transfer or transformation. Additionally, optionally or alternatively, in some operational scenarios, deeper network (than 4-layer) and / or multiple WCT blocks may be incorporated or included in the system or framework.

[0028] The system or framework can be used to support or implement semantic transfer on multiple style / modalities of references with semantic-wise color and brightness transfer algorithms, methods or process flows on images / videos. These algorithms, methods or process flows can take multiple images / videos as input. The system and / or users can select which semantic region of a source content image / video for style transfer or modification by using the specific color and brightness (look) from a corresponding semantic region selected in a reference style image / video. Once these semantic region correspondence relationships are established automatically by the system or in part with user input, the algorithms, methods or process flows can perform specific color / brightness transfer or transformation of the semantic region in the source content image / video.

[0029] The system or framework can be used to support or implement flexible transfer via modular design. The system or users may group some or all semantic classes into a super class and / or can divide / partition a semantic class into semantic sub-classes and cause (e.g., new, existing, etc.) system- or user-defined semantic class(es) or sub-class(es) to be used for color or brightness transfer or transformation. Different segmentation algorithms or methods may beused or applied in defining, grouping or partitioning semantic classes and / or in delineating semantic regions in images / videos.

[0030] The system or framework can be used to support or implement temporal stability via online training. Under some approaches, temporal inconsistency may exist in segmentation masks – for example, pixels visually in the same semantic classes in a time domain or along a temporal direction are classified to different semantic classes) in videos or sequences of images, thereby creating or generating flickering artifacts in the output videos or sequences of images. In comparison, under techniques as described herein, training (e.g., online training, etc.) may be used to train some or all of the AI / ML models in the system or framework to handle or deal with temporal inconsistency in the segmentation masks and / or to prevent or reduce flickering artifacts in the output videos or sequences of images.

[0031] The system or framework can be used to prevent spatial boundary artifact via edge- aware filter. Many style transfer approaches suffer from relatively severe boundary visual artifacts such as halo, color leakage near object boundaries, etc. In comparison, under techniques as described herein, edge-aware filter(s) may be deployed to remove or reduce spatial domain boundary visual artifacts.

[0032] As compared with other approaches in which color transfer is deployed with relatively large models, the system or framework may be implemented with a relatively light- weight or shallow architecture for transferring color and brightness.

[0033] Example embodiments described herein relate to image processing operations. One or more class correspondence relationships are established between one or more first semantic classes to which one or more first semantic regions of one or more content images belong and one or more second semantic classes to which one or more second semantic regions of one or more style images belong. A first convolutional layer initialized with first pretrained weights of a relatively deep convolutional network (denoted as VGG) for relatively large-scale image recognition is used to extract specific content features in the one or more first semantic regions of one or more content images and to extract specific style features in the one or more second semantic regions of one or more style images. Based at least in part on (a) the one or more class correspondence relationships, (b) the specific content features and (c) the specific style features, a whitening and coloring transform (WCT) is applied to transfer a specific color and brightness look associated with the one or more second semantic regions of the one or more style images to the specific content features in the one or more first semantic regions of the one or more content images. A second convolutional layer initialized with second pretrained weights of the VGG is used to process the specific content features with the transferred specific color and brightnesslook into second specific content features. Based at least in part on the second specific content features, a decoder generates one or more semantic content transferred images corresponding to the one or more content images.

[0034] In some example embodiments, mechanisms as described herein form a part of a media processing system, including but not limited to any of: a wearable device, a handheld computing device, game machine, television, laptop computer, netbook computer, tablet computer, desktop computer, computer workstation, computer kiosk, or various other kinds of computing devices and media processing units.

[0035] Various modifications to the preferred embodiments and the generic principles and features described herein will be readily apparent to those skilled in the art. Thus, the disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features described herein. 2. VGG ARCHITECTURE

[0036] A system as described herein may include a (e.g., relatively small, consisting of only two convolutional layers, proper, etc.) subset of convolutional layers specifically selected from a relatively (e.g., very, etc.) deep neural network architecture such as a VGG architecture as illustrated in FIG.2A. Pretrained weights for all convolutional layers the VGG architecture (e.g., VGG-19, VGG-16, etc.) may be accessible from one or more tools or API sets (e.g., Keras 3, MATLAB, etc.).

[0037] The VGG architecture from which the pretrained weights for the subset of convolutional layers originate may include a plurality of blocks such as Block-1 to Block-6 as shown in FIG.2A.

[0038] Among these blocks, Block-1 to Block-5 include numerous (e.g., 19, 16, etc.) convolution layers – e.g., low-level convolutional layers (denoted as “Conv1” and “Conv2”) each with each comprising 3x3 kernels extracting 64 features, mid-level convolutional layers (denoted as “Conv3”, “Conv4”, “Conv5”, “Conv6”, “Conv7” and “Conv8”) each comprising 3x3 kernels extracting 128 or 256 features, high-level convolutional layers (denoted as “Conv9”, “Conv10”, “Conv11”, “Conv12”, “Conv13”, “Conv14”, “Conv15” and “Conv16”) each comprising 3x3 kernels extracting 512 features, etc. – each of which convolution layers may be followed by a respective ReLU (activation function) non-linear layer of a plurality of ReLU non- linear layers.

[0039] The VGG architecture may also include max-pooling operators (denoted as “Pool / 2”) after each block to reduce (output) spatial dimensions and redundant information. The VGG architecture further includes Block-6 that consists of consecutive fully connected layers (denotedas FC), which maps flattened features outputted from Block-5 into numerous (e.g., 1000, from 1000 output neurons, etc.) different image category classes. Here, each neuron in a subsequent layer in the consecutive fully connected layers is fully connected with all outputs of all previous neurons such as 4096 neurons in an immediately preceding layer in the consecutive fully connected layers.

[0040] This VGG architecture extracts different kinds or types of features in multiple hierarchical levels. At initial convolutional layers such as the low-level convolutional layers, it extracts low-level features that are mainly influenced by color and brightness (as opposed to spatial texture or frequencies) of input image(s). Going deeper into a feature space formed by features extracted with successive convolutional layers such as mid-level features extracted with the mid-level convolutional layers, the (deeper such as the mid-level) features are influenced more and more by textures or spatial frequencies present in the image(s). High-level features extracted with the high-level convolutional layers are more related to (e.g., semantic, viewer perceptible or depicted objects or characters or background / foreground, etc.) object classes present in the image(s), as compared with relatively low level features.

[0041] The semantic color transfer system or framework as described herein may use or re- use pretrained VGG weights (e.g., VGG19, etc.) of relatively low level convolutional layers including but not necessarily limited to only those in Block-1. In some operational scenarios, only a few lowest level convolutional layers (e.g., CONV1 and CONV2, a proper subset, etc.) are selected to be included in the semantic color transfer system or frameworks; some or all of the rest of the numerous convolutional layers and other constructs in the VGG architecture may not be included or used in the semantic color transfer system or framework as described herein.

[0042] As a result of using pretrained weights and only a limited subset of convolutional layers in the VGG architecture, time and costs for training the AI / ML models in the semantic color transfer system or framework can be much reduced. In addition, time and costs for applying or using the AI / ML models in the semantic color transfer system or framework for inferencing or performing semantic color (and brightness) transfer operations between style images and content images can also be much reduced. 3. WHITENING AND COLORING TRANSFORM

[0043] The semantic color transfer system or framework as described herein includes or uses a Whitening and Coloring Transform (WCT), which may be implemented in two stages or steps: a whitening stage and a coloring stage.

[0044] Given a pair of images, one may be a style image denoted as ^ ^^ ∈ ℝ ^×^^×^, and theother may be a content image denoted as ^ ^^ ∈ ℝ ^×^^×^.

[0045] Extracted (style) feature maps generated or derived from ^^using a VGG encoder – e.g., a convolutional layer followed by a nonlinear (e.g., ReLU, etc.) activation layer, one or more convolutional layers followed by one or more nonlinear (e.g., ReLU, etc.) activationlayers, etc. – may be denoted as ^^ ∈ ℝ^^×^^×^. Extracted (content) feature maps generated orderived from ^ using ^^×^^×^^ the same or a different VGG encoder may be denoted as ^^ ∈ ℝ ,where ^^and ^^represent the height and width of the style image; ^^and ^^represent the height and width of the content image; ^ represent the total number of feature channels (or types) supported or generated by the VGG encoder from a given input image (e.g., the content image, the style image, etc.).

[0046] Before further processing, spatial dimensions of the style and content features represented in these extracted feature maps may be flattened and / or re-arranged into, forexample, style features ^^ ∈ ℝ^×^^ and content features ^ ^×^^ ∈ ℝ ^, where ^^ = ^^ × ^^ and^^ = ^^ × ^^.

[0047] The WCT can be used by the semantic color transfer system or framework to transfer a color and brightness look of the style image as represented in ^ to ^^, for example with matching covariance matrixes.

[0048] In the whitening stage / step as mentioned earlier, the WCT can apply a whitening transform as follows. The whitening transform first calculates a mean denoted as ^^of thecontent features for each feature channel denoted as ^, where ^ ^×^^ ∈ ℝ . Subtract the mean^^to center the content features ^^as follows: ^^ = ^^ − ^^ (1)

[0049] content features ^^^ , such that the transform features become uncorrelated, as follows:^^ ^^^ = ^^^^^^^^^^(2)

[0050] As ^^^ is uncorrelated, ^^^ . ^^ ^^ = ^, where T denotes matrix transposition, and here Idenotes the identity matrix with all ones in diagonal matrix elements and all zeros in off- diagonal matrix elements.

[0051] Let ^^be a diagonal matrix with eigen values of the covariance matrix ^^^^^∈ℝ^×^. Let ^ be a corresponding orthogonal matrix of eigen vector ^ ^^ s such that ^^^^ = ^^^^^^

[0052] In the coloring stage / step as mentioned earlier, the WCT applies a coloring transform or an inverse coloring transform as follows. The coloring transform first calculates a meandenoted as ^ ^×^^ of the style features for each feature channel, where ^^ ∈ ℝ . Subtract themean ^^to center the style features ^^, as follows: ^^ = ^^ − ^^ (3)

[0053] The (e.g., inverse, etc.) coloring transform can be applied to the converted contentfeatures ^^^ to generate transformed content features ^^^ such that the covariance matrix of thecontent transformed feature ^^^matches with the covariance matrix of ^^, as follows: ^. ^ ^^^ = ^ ^^^ ^ . ^^ (4)

[0054] The content transformed features ^^^in expression (4) can be calculated or derived as follows: ^ ^^^ = ^^^^ ^^ ^^ ^^^ (5)where ^ represents a diagonal matrix with eigen value ^ ^×^^ s of the covariance matrix ^^^^ ∈ ℝ ,and ^^represents a corresponding orthogonal matrix of eigen vectors.

[0055] Now re-center the content transformed features ^^^by adding the means of the style features to get or obtain the final output for the content transformed features ^^^, as follows: ^^^ = ^^^ + ^^ (6)

[0056] The foregoing process may be mathematically expressed or represented as follows: ^^^ = ^^"#^^ , ^^% (7)where ^^" denotes an overall operator representing the WCT process used in the semantic color transfer system or framework. 4. SEMANTIC SEGMENTATION

[0057] The semantic color transfer system or framework as described herein includes or uses a (e.g., relatively, etc.) deep semantic segmentation model – capable of recognizing viewer perceptible visual objects or characters or other semantic visual elements, for example with visually perceptible contours and boundaries of irregular spatial shapes – to generate a semantic segmentation map from a received input image. The semantic segmentation map shows or specifies probabilistic memberships or belongings of each pixel into (e.g., N, etc.) different semantic classes or categories. These different semantic classes or categories may be (e.g., uniquely, distinctly, with integer or discrete values, etc.) identified or indexed with different semantic class (or category) labels.

[0058] More specifically, for a given input image denoted as ^ ∈ ℝ^×^×^, the semanticsegmentation model performs segmentation and / or objectdenoted as ^^&'to generate or output a corresponding semantic segmentation map denoted as ( ∈ ℝ^×^×), asfollows: (= ^^&'#^% (8)where ^ and ^ represent the width and height of the input image I; * represents the total number of the different semantic classes or categories.

[0059] At each pixel location denoted as m represented in the semantic segmentation map M, * different probability values denoted as (+(where , is from 1 to N) corresponding to * different semantic classes or categories represent semantic mask values at the pixel location. Thesum of these probability values or values across the third dimension representing the * different semantic classes or categories is one, as follows: ∑) ^×^+.^ (+ = ^ , where ^ ∈ ℝ (9)where ^ matrixelements and all zeros in off-diagonal matrix elements).

[0060] From the semantic segmentation map M, the system or framework can generate or obtain a pixel-wise classification map denoted as / by selecting a specific semantic class / category label – among the different semantic class / category labels of the different semantic classes or categories – whose probability value is the maximum among the N probability values at each pixel location represented in the segmentation map, as follows: / = (01_30456#(% (10)where (01_30456 represents an operator that generates the pixel-wise classification map B from the semantic segmentation map M. 5. GUIDED FILTER

[0061] Guided filtering may be performed with the semantic color transfer system or framework based at least in part on an assumption that there exists a linear correlation (locally in a pixel neighborhood around each pixel) between an input image denoted as 7+and an guidance image denoted as 8+.

[0062] Given a local pixel neighborhood denoted as 9+– centering at pixel i (e.g., with twocorresponding pixel i in the input and guidance– with a / × / pixel neighborhoodsize, a linear function with coefficients (:+, ;+) can be constructed based on this linear correlation assumption between a predicted image (whose pixel may be denoted as 7+̂) of the input image and the guidance images, as follows: 7+̂ = :+ ∙ 8+ + ;+ (11)

[0063] may be modelled or derived as noise (>+), as follows: 7+̂ = 7+ + >+ (12)

[0064] Optimized values (denoted as :?@A ?@A+ , ;+ ) for the linear function coefficients (:+, ;+)for the local linear model represented in expression (11) above may be obtained or generated by minimizing the difference between >B+and >CB+, as follows: D:?@A+ , ;?@A+ E = argmin ∑PSTM #‖7P̂ − 7P‖Q + R#:+%Q%#LM ,NM%

[0065] size) may,but is not, examplevalue for a weight or regularization parameter (denoted as R) may, but is not necessarily limitedto only, set as follows: R = 10^^. In (e.g., image / frame, etc.) boundary regions of the inputimage and / or the guidance image, reflection padding may be used to provide pixel neighborhoods beyond (e.g., image / frame, etc.) boundaries.

[0066] The optimized solution of the linear model / function coefficients D:?@A ?@A+ , ;+ E can beobtained or computed as follows: ^∑Z\] #YZ∙[ % ^^ ∙[̅:?@AX^ M Z M M+=#`M%^aS(14-1) 2)theaverage of the individual linear models for individual pixels in the local neighborhood may be taken to ensure the smoothness, as follows: :e+= ^c^ ∑PSTM #:P% (15-1)2)

[0068] from the input image and the guided image based on the linear model with the average optimized coefficient values :e+, ;̅+is given as follows: 7+̂ =∙ 8+ + ;̅+ (16)

[0069] For simplicity, the foregoing guided filter operation may be represented or denoted as follows: 7+̂ = fg^^^#{7+i, {8+i, / % (17)where fg^^^#… % represents the guided filter operator or operation.6. COLOR TRANSFER NETWORK

[0070] FIG.2B illustrates an example color and brightness transfer deep learning neural network (which may be referred to as “a color transfer network” for simplicity) that may beincluded or used in the semantic color transfer system or framework as described herein. Some or all modules or subnetworks used in the color transfer network or the semantic color transfer system or framework – including but not limited to the core transfer network / component, pre- processing modules / operations, post-processing modules / operations, etc. – may be pretrained or trained. The color transfer network may be used, implemented, trained, or applied to process different modalities of inputs to the pre-processing modules / operations, core transfer network, post-processing modules / operations, etc.

[0071] The core (networks / modules) of the color transfer network comprises or includes several (e.g., main, etc.) components, namely a VGG encoder, a WCT module, a decoder, etc. Pre-processing modules / operations may be used or implemented to pre-process original image / video input of content image(s) and style image(s) and to generate VGG encoder inputs are not shown in FIG.2B. Similarly, post-processing modules / operations may be used or implemented to receive outputs from the decoder of FIG.2B and to post-process the decoder outputs are not shown in FIG.2B. The pre-processing and / or post-processing modules / operations will be discussed later in more details.

[0072] The VGG encoder includes or comprises a subset of (VGG) convolutional layers – e.g., consisting of two convolution layers selected from (the lowest level) convolutional layers of the pretrained (and frozen in the color transfer network) VGG architecture. The weights in the convolutional layers of the VGG encoder are initialized by pretrained weights (and then frozen) for the selected convolutional layers in the VGG architecture (e.g., VGG19, VGG16, etc.). A ReLU non-linear layer may be used after each of the convolution layers in the VGG encoder.

[0073] The VGG encoder (denoted as ^&kl) includes two VGG encoder parts or portions respectively denoted as ^&kl^and ^&kl^. Hence, ^&kl^represents the first VGG encoder part or portion VGG Encoder-1, whereas ^&kl^represents the second VGG encoder part or portion VGG Encoder-2.

[0074] The semantic informative WCT module in between VGG Encoder-1 and VGG Encoder-2 is used for semantic region-wise color transfer. The WCT module receives and works on extracted features of the first convolution layer of the VGG encoder or the first VGG encoder part or portion VGG Encoder-1 and transfers to the content features with color and brightness looks of the style features based at least in part on semantic class correspondence relationships and semantic segmentation maps. After that, output image(s) can be reconstructed by the decoder (denoted as ^m&l) from further content features derived or extracted from the content features with the transferred color and brightness looks. The decoder may contain the same total number (e.g., two, etc.) of convolutional layers as the VGG encoder. The convolutional layers inthe decoder are randomly initialized and then trained using a training framework or dataset.

[0075] More specifically, the WCT module becomes semantic informative by way of semantic segmentation maps generated from the original (e.g., before pre-processing operations are performed, etc.) content and style images and by way of semantic class (or category) correspondence relationships automatically generated by the system and / or (e.g., partly, etc.) manually defined or specified with user input.

[0076] The first VGG encoder part or portion VGG Encoder-1 extracts (e.g., the lowest level, etc.) content features ^^and style features ^^from the received inputs such as pre- processed content image(s) ^^and pre-processed style image(s) ^^, respectively.

[0077] These extracted features from VGG Encoder-1 can be provided to the WCT module, which uses the segmentation maps and the semantic class correspondence relationships as well as the extracted content and style features to transfer color and brightness looks of style features of one or more first semantic regions of one or more first semantic classes in the style image(s) to content features of one or more second semantic regions of one or more second semantic classes in the content image(s). For example, features – belonging to the same semantic class or the same system- and / or user-defined semantic class correspondence relationship – may be selected from both ^^and ^^in the style image and the content image. The WCT module can then be used to transfer the color and brightness look from the selected features ^^to the selected features ^^and / or transform the selected content features ^^to selected color transferred features ^^^.

[0078] A segment region in a style or content image as described herein may refer to a spatial region specifically delineated or demarcated or identified in a semantic segmentation map as belonging to or a member of a specific semantic class or category among the different semantic classes or categories.

[0079] The (e.g., lowest level, etc.) content features – e.g., the selected color transferred features ^^^, etc. – with the transferred color and brightness looks of the style features in the corresponding semantic classes represented in the style image(s) may be received by VGG Encoder-2 to extract (e.g., the next level, etc.) content features – which the transferred color and brightness looks – of the content image(s).

[0080] The extracted content features from VGG Encoder-2 may be provided as input to the decoder in the color transfer network of FIG.2B for the purpose of reconstructing or generating semantic transferred content images corresponding to the original content images. These semantic transferred content images contain the same or similar texture or image details but with the transferred color and brightness looks of the style image(s) in some or all semantic regions.

[0081] Unlike other approaches that rely on deeper VGG features from numerous VGG layers in the VGG architecture, under techniques as described therein, only a (relatively small and relatively low level) limited or proper subset of convolutional layers in the VGG architecture such as the very first two convolutional layers of the VGG architecture are used to extract features to which color and brightness looks of extracted style features are transferred. These features extracted with the limited subset of convolutional layers are sufficient for the purpose of transfer the color and brightness information at the semantic level from the style image(s) to the content image(s).

[0082] In comparison, if the features from the deeper or higher-level convolutional layers of the VGG architecture were used, then those features would contain specific texture information of the style image(s) along with the color and brightness information of the style image(s). As a result, the WCT module would try or attempt to transfer the specific style textural information from the style image(s) into the resultant transformed content image(s), which is undesirable as compared with the techniques as described herein. 7. PRE- AND POST-PROCESSING

[0083] A number of subsystems or modules may operate with the core network / model of the semantic color transfer system or framework to pre-process inputs into, or to post-process outputs from, the core network / model as discussed herein.

[0084] An example of pre-processing operations as described herein may include, but are not necessarily limited to only, operations to generate semantic segmentation maps (also referred to as segmentation masks), which serve as inputs to the WCT in the core network / model.

[0085] In some operational scenarios, one or more (e.g., off-the-shelf, proprietary, off-the- shelf with proprietary enhancement, etc.) segmentation tools may be used in the system or framework to obtain or classify pixels of a given input image into, or as belonging to, different semantics such as different semantic classes or categories identified or indexed with respective semantic class / category labels. The segmentation tools can output specific semantic class labels and probability values of respective semantic classes or categories for each pixel in the input image.

[0086] The segmentation tools may be applied to both content image(s) / video(s) and style image(s) / video(s). Additional functions or operations may be implemented, invoked or performed in the system or framework to enable relatively flexible high-level or semantic color and brightness transfer in connection with the content and style images / videos.

[0087] In some operational scenarios, (e.g., two, etc.) additional modules or operations attendant to the segmentation tools may be included or performed in the semantic color andbrightness transfer network or pre-processing operations performed therewith to assist the segmentation task performed by the segmentation tools.

[0088] For example, the segmentation tools may include (e.g., state of the art, up to date, etc.) learning based segmentation models trained on standard (e.g., training image, etc.) datasets to generate multiple (e.g., segmentation, visual object, semantic, candidate, initial, 250, a relatively high total number of, etc.) classes.

[0089] An additional module or operation attendant to the segmentation tools may be included in the system or framework to merge a relatively high total number of classes into a relatively small total number of super-classes (or super-categories). By way of illustration but not limitation, 250 (e.g., initial, candidate, etc.) classes from the segmentation tools may be merged into seven (7) super-classes or final classes (or simply classes): stationary man-made outdoor objects; non-stationary man-made outdoor objects; indoor objects; sky; trees; natural stationary objects (e.g., Earth, mountain, terrestrial field, physical ground, etc.); waterbodies (e.g., ocean, lake, river, etc.); and so forth. It should be noted that, in various operational scenarios, more or fewer or different total number of (e.g., initial, final, pre-merged, merged, semantic, etc.) classes, super-classes, etc., may be used or supported by a system or modules / components / subsystems as described herein.

[0090] As discussed herein, the WCT module in the core network / model of the semantic color transfer system or framework may receive semantic class correspondence relationships as a part of inputs. Some or all of these correspondence relationships may be generated, mapped, defined and / or specified by the system automatically, by one or more (designated) users manually, by the system with user input from the users, etc.

[0091] A semantic class correspondence relationship (also referred to as “class mapping” for simplicity) as described herein specifies or defines or identifies (e.g., a first semantic label for, etc.) a first semantic class / category in the style image(s) / video(s) as corresponding to (e.g., a second semantic label for, etc.) a second semantic class / category in the content image(s) / video(s). Given the correspondence relationship, the system or framework or the WCT module therein effectuates or carries out a transfer operation from the (or a first) color and brightness look of a first semantic region of the first semantic class / category in the style image(s) / video(s) to a second semantic region of the second semantic class / category in the content image(s) / video(s).

[0092] By way of illustration but not limitation, to generate a user-defined class mapping, a user may input – through a user interface supported by the system – a specific semantic class label of a semantic class found or detected in a semantic region of a style image / video and aspecific semantic class label of a semantic class found or detected in semantic region of a content image / video. Given the user-defined class mapping, the specific color and brightness of the semantic region of the style image / video will be transferred by the system or the WCT module therein to, or will replace the color or brightness look of, the semantic region of the content image / video.

[0093] The user may be provided or given freedom to choose or select a specific semantic region of interest in the style image / video and a specific semantic region of interest in the content image / video. The semantic region of the style image / video does not have to have the same semantic class (or semantic class label) as that of the semantic region of the content image / video. For example, the user can select to transfer the color (and brightness) from the sky in a style image to the color (and brightness) of a river in a content image by way of a corresponding user-defined class mapping.

[0094] If no user input is given or received for some or all class mappings by the system, the system may automatically generate these class mappings (referred to as system-defined class mappings). In some operational scenarios, the system-defined class mappings may map semantic labels (or corresponding semantic classes) in the style image(s) / video(s) to the same semantic labels (or corresponding semantic classes) in the content image(s) / video(s). These automatically generated class mappings may cause the color and brightness looks of semantic regions of the same semantic labels to be matched or the same – region-wise or matched at a semantic region level – between the style image(s) / video(s) and the content image(s) / video(s). will be matched between the style image and the content image, and color will be transferred semantic region- wise. Hence, the sky from a style image may be matched with or transferred to the sky in a content image; the same for other semantic regions of the same semantic label or class. If there is a semantic label in the content image that is not present in the style image, no color (and brightness) transfer will be performed for the semantic label (or a corresponding semantic region) in the content image.

[0095] An additional module or operation attendant to the segmentation tools may be included in the system or framework to pre-process individual segmentation masks (or semantic segmentation maps) – which may be temporal unstable – generated for individual images in a video (or a corresponding sequence of consecutive images, for example for a visual scene) into an overall temporally consistent (video) segmentation mask.

[0096] In some operational scenarios, this module or operation may be used or performed to transfer the color and brightness look to semantic regions of a content video instead of a single content image. As the semantic segmentation is performed on each video image / frame in thecontent video independently to generate the individual segmentation masks, these masks may be inconsistent in the temporal domain or along the playback time direction represented or covered in the sequence of consecutive images in the content video, thereby creating visually perceptible flickering artifacts in the rendered or output content video eventually.

[0097] The additional module or mechanism attendant to the segmentation tools can be implemented or used to finetune the individual segmentation masks relating to the individual images / frames of the content video, for example to the (predicted or estimated) overall segmentation mask, to improve temporal consistency. The overall segmentation mask may be applied to the individual images / frames of the content video in operations to transfer color and brightness looks of semantic classes or regions between style image(s) and the content video. As a result, this additional module or mechanism can operate with the segmentation tools to significantly reduce or remove visually perceptible flickering artifacts from the rendered or output content video.

[0098] In some operational scenarios, the finetuning of the individual segmentation masks into the overall temporally consistent segmentation mask is performed via online training or optimization using the individual segmentation masks as input, as follows: (^ = ^[&'^nA o^^ , (p^ ; r^stuvwxy (18)where ^^ (p^representsthe finetuned mask with no or little temporal inconsistency.

[0099] During online training, the content video z^or the individual images / frames ^^therein as well as the individual temporally inconsistent segmentation masks (p^ may be used totrain or optimize operational parameters denoted as r^stuvwxin the finetuning operations (denoted as ^[&'^nA) to improve temporal consistency and reduce / remove visually perceptible flickering artifacts from the (e.g., final, etc.) output content video comprising a sequence of consecutive semantic color (and brightness) transferred images generated by the system or framework. The online training or optimization of the temporally inconsistent segmentation masks into the overall temporally consistent segmentation mask will be discussed later in further detail.

[0100] It should be noted that theoretically or conceptually temporal consistent finetuning of segmentation masks may be performed for images / frames in a style video as well. In some operational scenarios, the final performance in connection with the removal or reduction of visually perceptible flickering artifacts may not be affected even without applying the finetuningto the individual segmentation masks for the images / frames in the style video. Hence, in some operational scenarios, the fine tuning of segmentation masks may be performed on the content video, rather than the style video.

[0101] In addition to the pre-processing modules or operations, post processing modules or operations may be included or used in the semantic color and brightness transfer system or framework, for example to improve visual quality of output content image(s) / video(s) with transferred color and brightness looks of style image(s) / videos.

[0102] For example, semantic segmentation maps generated for images / videos as described herein may not be accurate in (e.g., visual object, image / frame, etc.) boundary regions such as boundary regions / portions of semantic regions. Visually perceptible ghosting artifacts may exist in these boundary regions in the output content video. In some operational scenarios, guided filtering may be applied or used to remove these artifacts from the boundary regions. The Guided Filtering (GF) may be defined or given as follows: ^^^ = f^#^p^^ , ^^% (19)where ^p^^ represents an style transferred image with visually perceptible ghosting or boundaryartifacts; ^^represents an input or pre-guided-filtered content image giving rise to the style transferred image; ^^^represents a guided filtered content or output image – corresponding tothe style transferred image ^p^^ but without visually perceptible ghosting or boundary artifacts.

[0103] Additionally, optionally or alternatively, some or all of the modules or operations in the semantic color and brightness transfer system or framework may operate on subsampled or resampled images / frames, for example to improve speed performance or reduce processing time and computing resource usages. For example, inputs to some of these modules or operations may be applied or pre-processed with spatial-temporal resampling or subsampling operations, as will be discussed in further detail. 8. COLOR TRANSFER NETWORK TRAINING

[0104] As noted, the semantic color and brightness transfer system or framework may include a number of (e.g., three or more, etc.) artificial or deep learning neural networks. For example, a semantic (or image) segmentation tool as described herein may be implemented with AI / ML model(s) such as an artificial neural network (referred to as “semantic segmentation network”). Similarly, a color transfer network as illustrated in FIG.2B may also be implemented with AI / ML model(s) such as an artificial neural network. A segmentation mask refinement module / operation may likewise be implemented with AI / ML model(s) such as an artificial neural network (referred to as “segmentation mask refinement network”).

[0105] The semantic segmentation network can be a pretrained AI / ML segmentation or image tool using a pretrained (e.g., available, state-of-the-art, open source, commercially available, etc.) AI / ML model or architecture. Pretrained weights of the AI / ML model or architecture may be frozen or used in semantic segmentation operations performed in the semantic color and brightness transfer system or framework.

[0106] The color transfer network and segmentation mask refinement network may be trained or optimized as follows.

[0107] FIG.2C illustrates an example architecture for training the decoder of FIG.2B. As shown, the semantic informative WCT module is not used or included during training, but rather used during inferencing only. VGG encoder weights – e.g., weights used in the convolutional layers of the VGG encoder of the network of FIG.2C without the WCT module or of the color transfer network of FIG. 2B with the WCT module – are initialized and frozen with pretrained weights of a selected VGG architecture (e.g., VGG19, etc.). In comparison, decoder weights – e.g., weights used in the convolutional layers of the decoder of the network of FIG.2C or of the color transfer network of FIG.2B are randomly initialized weights (e.g., weights are initialized with random values, etc.) and updated / modified during training, for example through back- propagating prediction errors as measured with a loss (or error / objective) function between prediction results (or network output) generated by the network of FIG.2C from a population of training images in a training dataset and ground truths (e.g., the same training images used as the ground truths, etc.).

[0108] Hence, the weights of the VGG encoder may not be modified during training or in model inferencing / application. Trainable network parameters such as the decoder weights can be trained or optimized using any natural images as some or all of the training images. The training objective is to minimize the (e.g., prediction, etc.) errors between the input or training images and the predicted images or network output.

[0109] As illustrated in FIG.2A, the pretrained VGG (model) architecture is trained for image classification and hierarchically extracts different types / kinds / levels of image features from input images to classify the input images into multiple image classes or categories.

[0110] The semantic color and brightness transfer system or framework as described herein may use or re-use the first lowest (e.g., two, etc.) convolutional layers of the VGG architecture to form the VGG encoder ^&klof FIG.2B or FIG.2C, which extracts mainly relatively low-level features like color and brightness from the input images.

[0111] The semantic color and brightness transfer system or framework further adds or includes a decoder ^m<o reconstruct output images the same as or relatively closelyapproximating the input images with transferred color and brightness looks but keeping or maintaining semantic content (e.g., visual objects, characters, visually perceptible image details other than the color and brightness looks, etc.) of the input images in the reconstructed images or the network output.

[0112] The purpose of the (trained) decoder is to reconstruct image data or content to be included in the reconstructed images from the features (e.g., in a feature domain, latent arrays / matrixes, latent data, etc.) such as mainly the color and brightness features generated by the VGG encoder ^&klwith the pretrained weights.

[0113] After the decoder is trained, during model testing, the WCT module can be incorporated or used within the VGG encoder to perform color and brightness transfer or transformation operations with the extracted image features in the feature domain and to validate the operations of the trained decoder, for example in terms of visual qualities of the reconstructed image. Hence, during model testing or inferencing, the trained decoder can be used to reconstruct the color and brightness transformed images form a color and brightness transformed feature space. Here, the color and brightness transformed feature space includes some or all of the color and brightness transformed features generated by the VGG encoder with the WCT module.

[0114] Denote the network of FIG. 2C as ^: {^&kl , ^m&li with trainable parameters r. Thesetrainable parameters r may include some or all operational parameters or weights / biases in non- linear layers following convolutional layers and the convolutional layers in the decoder.

[0115] Denote a training (input and / or ground truth) image as ^. The network output or a corresponding reconstructed image may be obtained from the network of FIG.2C as follows: ^?|A = ^#^; r% (20)

[0116] The prediction (or reconstruction) errors between ^?|Aand ^ can be calculated and used to optimize the parameters r. For the purpose of illustration only, the mean-squared (3Q) errors between the input image and the reconstructed image may be used as a loss function. The Adam optimizer may be used to update some or all of the parameters r such as the weights in the decoder. A learning rate during training may be set to a value such as 10^}. The network or model of FIG.2C may be trained for a plurality of (e.g., 250, etc.) epochs and each of the epochs may constitute or include a plurality of (e.g., 1000, etc.) batch updates.

[0117] For the purpose of illustration only, it has been described that, in some operational scenarios, the VGG or the two lowest convolutional layers or a subset of the lowest convolutional layers therein with pretrained weights – without some or all higher convolutionallayers of the VGG – may be used to implement a color and brightness transfer network as described herein.

[0118] It should be noted that, in other operational scenarios, one or more other (image feature) deep learning networks other than the VGG or artificial neural network layers therein – which, for example, have been pretrained to learn and predict image features relating to color and brightness of input images – may be used to implement a color and brightness transfer network.

[0119] For the purpose of illustration only, it has been described that, in some operational scenarios, the WCT may be used to implement a color and brightness transfer network as described herein.

[0120] It should be noted that, in other operational scenarios, one or more other color and brightness transforms (or transfer neural networks) other than the WCT may be used to implement a color and brightness transfer network.

[0121] For the purpose of illustration only, it has been described that, in some operational scenarios, a color and brightness transfer network as described herein may include a first encoder convolutional layer, followed by a color and brightness transform, and further followed by a second encoder convolutional layer and two decoder convolutional layers, including but not limited to non-linear activation layer(s) followed each of some or all of the convolutional layers.

[0122] It should be noted that, in other operational scenarios, one or more other neural network architectures or structures may be used to implement a color and brightness transfer network. For example, in some operational scenarios, a general encoder-decoder architecture – which, for example, comprises a first encoder (e.g., the lowest convolutional layer of the VGG, etc.) and a second decoder (e.g., the second lowest convolutional layer of the VGG in combination of two decoder convolutional layers, etc.) – may be used to implement a color and brightness transfer network.

[0123] In some operational scenarios, the convolutional network for image recognition such as a VGG network as described herein includes a first convolutional layer that receives the content images and the style images as input and generates relatively low-level image features. The relatively low-level image features serve as input to one or more subsequent convolutional layers of the convolutional network for image recognition to generate relatively deeper image features from the relatively low-level image features. The specific content features and the specific style features are extracted by the first convolutional layer only of the convolutional network for image recognition. The specific content features and the specific style features are free of relatively deep content features and of relatively deep style features extractable by theone or more subsequent convolutional layers of the convolutional network for image recognition. 9. ONLINE TRAINING FOR FINETUNING SEGMENTATION MASKS

[0124] FIG.2D illustrates example online training of a segmentation mask refiner network for a content video. Denote a sequence of (e.g., RGB, etc.) video images / frames in the content video as z. Denote a pre-trained semantic segmentation tool (or model / network) as ^[&'with pretrained parameters denoted as rnstu. ^ and ^ represent the height and width of each of the video frames. ^ represents the total number of the images / frames in the content video. Hencez ∈ ℝ^×^×^×^.

[0125] The pretrained segmentation tool or model receives z as input and generateindividual segmentation masks / maps (^ as follows:(^ = ^[&' oz; rnstuy (21)where (^ ×^, and * represents the number of different semantic classes or categoriesfor semantic regions demarcated in the segmentation masks / maps (^ .

[0126] As the pretrained segmentation tool or model ^[&'performs semantic segmentation on each image / frame independently, temporal inconsistency (or instability) may be observed inoperational scenarios in which the individual segmentation masks / maps (^ are used directly inthe color and brightness transfer algorithm, method or operation. Under techniques as describedherein, the individual segmentation masks / maps (^ may be improved for temporal consistency.

[0127] More specifically, an additional (e.g., training and inferencing, etc.) module orframework may be used to finetune (^ for the color and brightness transfer with the (e.g., single,etc.) content video. The training of the additional module / framework may be performed during the color and brightness transfer, and may be referred to as online training.

[0128] The additional module / framework may be implemented with another neural network denoted as ^[&'^nAwith trainable or learnable parameters rnstuvwxrandomly initialized (or initialized with random values) to learn (e.g., hidden,between the sequence ofvideo images / frames z in the content video and the individual segmentation maps / masks (^generated for and corresponding to the video images / frames V of the content video.

[0129] A forward pass or path through the neural network ^[&'^nAto generate predicted temporal consistent maps / masks denoted as (^may be defined or specified as follows: (^ = ^[&'^nA oz; (22)

[0130] Losses or errors computed between (′ and (^ may be minimized to optimize theoperational parameters rnstuvwxof the neural network ^[&'^nA– for example through back propagating the losses / errors and thereby training the neural network ^[&'^nA. A loss / error as described herein may be computed using the mean-squared error as a loss function and the Adam optimizer with a learning rate such as 10^}as a weight update rule.

[0131] If the neural network ^[&'^nAis trained for relatively long epochs, then ^[&'^nAmayexactly learn (^ and largely duplicate (^ in its output (′. As a result, the estimated / predictedsegmentation maps / masks (′ may also contain flickering and relatively high-frequencycomponents as captured or represented in (^ .

[0132] In some operational scenarios, the neural network ^[&'^nAmay be trained for relatively short epochs instead of relatively long epochs. Training for the neural network ^[&'^nAmay be performed with a relatively small number of epochs. When the neural network ^[&'^nAis trained with a training dataset for the relatively short epochs and the relatively small number ofepochs, ^[&'^nA will learn the mapping relationships between z and (^ using relatively low-frequency and / or mid-frequency details during the short (or initial learning) epochs and the small number of epochs, rather than higher frequency details acquired with longer training or larger number of epochs. As a result, high-frequency details that could cause flickering are not captured in (′. Therefore, training the neural network ^[&'^nA(or the deep learning model) for a relatively few numbers of relatively short epochs can help achieve temporal consistency in (′ with no or little flickering.

[0133] By way of example but not limitation the U-Net architecture – which is described in Olaf Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation,” (available at https: / / arxiv.org / abs / 1505.04597; accessed on April 8, 2024), the contents of which are incorporated by reference herein – may be adopted as the neural network architecture for ^[&'^nA.

[0134] However, it should be noted that, in contrast with other approaches, the U-Net architecture adopted here for ^[&'^nAmay not use or include a batch normalization layer. This is because the batch normalization layer would otherwise help converge and overfit relatively fast thereby capturing or introducing flickering in the output segmentation maps / masks (′.

[0135] The segmentation mask refinement model or neural network ^[&'^nAmay be trained to remove flickering while keeping segmentation information loss at a minimum. For example, for a content video with 300 frames, the U-Net implementing the segmentation mask refinementmodel or neural network ^[&'^nAmay be trained for a relatively few number of (e.g., 30, etc.) epochs.

[0136] FIG.2E illustrates example testing or inferencing operations of the segmentation mask refinement model or network for a content video with a sequence of (e.g., RGB, etc.) video images / frames. As shown, after training, the trained model or network ^[&'^nAas illustrated in FIG.2E may be used to generate the (relatively temporally stable) segmentation (probability) masks without temporal inconsistency, as follows: (= ^[&'^nA oz; rnstuvwxy (22)

[0137] masks a a[&'^nAor parameters therein are content dependent, for example having been trained or optimized using a subset of images / frames in the content video and / or a subset of segmentation masks generated from the content video, etc. In some operational scenarios, the network ^[&'^nAmay be initialized with random values and separately trained given each new pair of a content video and a style video. Hence, the network ^[&'^nArepresents an online training module. 10. ACCELERATING INFERENCING AND ONLINE TRAINING

[0138] A number of optimizations may be implemented to accelerate inferencing and online training with no or little loss of perceptual visual quality and to improve output content video quality.

[0139] For example, semantic segmentation may be performed on style and content videos. Resultant segmentation masks may be inaccurate in boundary regions.

[0140] Under techniques as described herein, a segmentation model for mask estimation including but not limited to temporal stability refinement may be implemented along with guided filtering to support spatial resolution resampling or downsampling or downscaling while maintaining a relatively high visual quality in output content video.

[0141] More specifically, for high-spatial-resolution images, the segmentation mask estimation may take relatively long inference time. To accelerate inferencing and online training, spatial resolutions of images in the style and content videos may be downsampled or resampled. The segmentation mask estimation may be performed with downsampled or resampled images with relatively low spatial resolutions as input instead of the high-spatial resolution images. Segmentation masks corresponding to or derived from the downsampled or resampled images as generated by the segmentation mask estimation may subsequently be upsampled (or upscaled), for example using interpolation techniques. This may help reduce overall inferencing and / oronline training time but may increase inaccuracy in (e.g., segmentation, visual object, etc.) boundaries or boundary regions between different semantic regions.

[0142] The inaccuracies in the boundary regions after the semantic segmentation can be efficiently handled or alleviated by guided filtering, which results in no or little visible performance or quality difference between an output content video generated using a segmentation mask from downscaled video and an output content video generated without downscaling, downsampling or resampling.

[0143] In another example, inferencing and online training may be accelerated through temporal resampling or downsampling, for example in image processing operations such as transferring color and brightness looks from a style video to a content video.

[0144] A few intermediate images / frames in a sequence of consecutive images / frames in the style video and / or the content video may be skipped to reduce overall computational overheads. The consecutive images / frames may contain a relatively large amount of redundancy or almost identical or highly similar image data, which provides no advantage or benefit to consider all the individual images / frames of the style video and / or the content video. The temporal resampling or downsampling can help reduce the overall computational overheads and burdens and improve overall performance in the image processing operations as described herein.

[0145] In a further example, the color and brightness transfer network (e.g., as illustrated in FIG.2B, etc.) may be implemented to support using relatively low spatial resolution style reference data representing color and brightness looks in semantic regions of a style image / video.

[0146] Downscaled or downsampled style video images / frames may be used with no or little visual or perceptual quality change. In some operational scenarios, the color and brightness transfer network may still use or process full or actual spatial resolutions of content video images / frames with no loss in content textural information even when the style video images / frames are downsampled or downscaled. The downsampled style video images / frames may contain relatively intact color and brightness information for the semantic regions in the original style video images / frames for the WCT module to operate in the VGG feature domain. As a result, computational costs or burdens can be reduced with no or little significant performance drop or decrease.

[0147] In some operational scenarios, online training of the segmentation mask refinement model or network may be performed with relatively low spatial resolutions of downsampled or downscaled content images / frames in a content video, to help accelerate segmentation mask refinement operations performed by the model / network. The segmentation mask refinementmodel can be trained and tested on both the downscaled images / frames and resultant segmentation masks without causing any perceptual change in the final output content video.

[0148] Additionally, optionally or alternatively, guided filtering can be sped up by subsampling an image N times and using the subsampled image to calculate, derive or optimize guided filter parameters. Relatively fast guided filtering using guided filter parameters derived from subsampled images is described in Kaiming He et al., “ Fast Guided Filter,” (available at https: / / arxiv.org / pdf / 1505.00996.pdf; accessed on April 8, 2024). 11. SEMANTIC TRANSFER APPLICATIONS

[0149] The semantic color and brightness transfer system or framework as described herein can be implemented or used to support a wide variety of applications and / or operational scenarios. An arbitrary number of inputs such as images / videos can be received and processed.

[0150] A user can provide user input through user interface(s) of the system to select or specify which images / videos are content images / videos and which images / videos are style images / videos.

[0151] Additionally, optionally or alternatively, the user and / or the system can choose or select which semantic regions of interest or semantic classes / categories in the content images / videos correspond to which semantic regions of interest or semantic classes / categories in the style images / videos in semantic class correspondence relationships, which can be used to effectuate transferring color and brightness looks of the semantic regions in the style images / videos to the corresponding semantic regions in the content images / videos.

[0152] If the user does not provide any preference or input for semantic class correspondence relationships, the system can automatically establish (e.g., default, etc.) semantic class correspondence relationships, which may be used to effectuate transferring color and brightness looks of semantic regions of one or more semantic classes in the style images / videos to semantic regions of the same semantic classes, respectively, in the content images / videos.

[0153] In some operational scenarios, the semantic color and brightness transfer system or framework may include or implement semantic transfer algorithms, methods or operations for two different categories or solutions.

[0154] First, the system may implement first semantic transfer algorithms, methods or operations for a first category or a first unified solution using a single style image / video input. Under this first solution, semantic regions in an input content video can be transformed or transferred with color (or brightness) looks of semantic regions of the single style image / video input based at least in part on semantic class correspondence relationships.

[0155] Features in the single style image may be extracted, for example by a VGG encoder or convolutional layer(s) therein, from semantic or pixel regions in each of which pixels carry the same corresponding semantic (class) label for a respective semantic class / category.

[0156] Based on these features extracted from the semantic regions of the single style image / video, respective color and brightness looks of the semantic regions of the single style image / video may be transferred to semantic regions of the content video as specified by the semantic class correspondence relationships.

[0157] There may be a scenario in which semantic (class) labels present in some semantic regions of the content video may not present in any semantic regions of a single style image. Hence, a style video comprising multiple style images may be used to provide some or all of the semantic labels that may be missing or absent in a single style image. in this case.

[0158] Given a style video, semantic regions of the same semantic class in some or all images or frames in the style video may be grouped before features for the same semantic class are extracted from these semantic regions. The (group) features that belong to the same semantic class and are extracted from the grouped semantic regions of the images / frames of the style video may be used as a basis in transferring a color and brightness look of the grouped semantic regions of the style video into semantic regions of the same semantic class in the content video.

[0159] Second, the system may implement second semantic transfer algorithms, methods or operations for a second category or a second unified solution using multiple style image / video inputs. Under this second solution, both single and multiple style inputs may be considered or derived from the multiple style image / video inputs. In a multiple style inputs scenario, user intervention or input may be used. The user may provide user input to select a specific semantic region in the multiple style image / video inputs. From the specific semantic region, a style input is to be derived and transferred to a specific semantic region in the content video.

[0160] The user is allowed by the system to choose which specific semantic region in the multiple style image / video inputs and which specific semantic region in the content video. For example, if there are two style images in the multiple style image / video inputs, the user can select that (a first color and brightness look of) the sky from the first style image is to be mapped or transferred to the river of the content video and (a second color and brightness look of) the river of the second style image is to be mapped or transferred to the sky of the content video. 12. SINGLE STYLE IMAGE AND SINGLE CONTENT IMAGE

[0161] The semantic color and brightness transfer system or framework as described herein can be implemented or used to support different combinations of style and content inputs.

[0162] FIG.3A illustrates an example process flow, method or algorithm (denoted as Algorithm 1) for transferring color and brightness from a single style image denoted as ^^to a single content image denoted as ^^. As shown, the system receives the single style image ^^and the single content image ^^as input.

[0163] The system can transfer a color and brightness look of a semantic region / object of a semantic class / category in ^^to semantic region(s) / object(s) of the same (or a corresponding) semantic class / category present in ^^to generate or reconstruct a style transferred content imagedenoted as ^ ^^^ as output. The single style image ^^ ∈ ℝ ^×^^×^ and content image ^^ ∈ℝ^^×^^×^may be two input color images. ^^and ^^represent the height and width of the style image. ^^and ^^represent the height and width of the content image.

[0164] Algorithm 1 (also illustrated in TABLE 1 below) implements a pipeline with three main stages: semantic segmentation with a semantic segmentation model, color and brightness transfer from the style image to the content image with a semantic informed WCT in a color transfer network further including a VGG encoder (VGG Encoder-1 and VGG Encoder-2) and a decoder, and post-processing.

[0165] More specifically, at the first stage, the semantic segmentation is performed on both^^ and ^^. / ^ ∈ ℝ^^×^^ denotes the output (e.g., map, mask, arrays, etc.) of the semanticsegmentation performed on the style image. This semantic segmentation output includes, for each pixel location of some or all pixels in the style image, a specific semantic class label to which a pixel at the pixel location belongs.

[0166] The output (e.g., map, mask, arrays, etc.) of the semantic segmentation performed onthe content image – represented or denoted as ( ^^ ∈ ℝ ^×^^×) – includes, for each pixellocation of some or all pixels in the content image, semantic class labels to which a pixel at the pixel location may belong along with their respective probability values. Each pixel of thesemantic segmentation output of the content image in an ^^ × ^^ matrix has * different values.These * different values indicate specific probabilities of each pixel’s belongingness to or memberships of * different semantic class labels. The sum of all the probability values at eachpixel location is one: ∑)+.^ (^ = ^, ^ ∈ ℝ^^×^^ , where ^ is an all-one matrix (or the matrix ofall ones includingmatrix elements).

[0167] At the second stage, the color transfer process or operation is performed or started. The first convolutional layer ^&kl^(followed by a corresponding nonlinear activation function layer) of the VGG encoder may receive the input content image and the style image and extractcontent features ^ ∈ ℝ^^×^^×^ from the c ^^×^^×^^ ontent image, and style features ^^ ∈ ℝ fromthe style image, where ^ represents the total number of extracted image or spatial features of these images.

[0168] The style features ^^extracted from the style image may be filtered or specifically selected to generate filtered (or selected) features that belong to a specific semantic class present in / ^. These filtered features may represent a specific color and brightness look for the specific semantic class present in the feature domain of the style image. Based on these filtered features filtered from ^^, the WCT transform may be applied or invoked to transfer the specific color and brightness look of the specific semantic class in the style image to the content features ^^of a semantic class corresponding to the specific semantic class in the style image.

[0169] For a semantic class label (corresponding to a specific semantic class) denoted as ^, its transformed features for the specific semantic class ^ as generated with the WCT transform may be denoted as ^^^Z. For any class out of the * different semantic class labels not present in the style image, the same ^^features (not semantically transferred) may be kept or used. After completing this WCT assisted process for all the * different classes, * different transformed content features are generated. The previously generated segmentation output (^shows or indicates respective probability values of memberships or belongingness of each pixel into each of the N semantic class labels or classes for each pixel. These * different probability values from (^can be used to multiply with * different transformed features to generate a weightedsum (denoted as ^p^^ ) of transformed features ^^^Z, where ^ = 1, 2, … . , *, as follows:^p^^ = ∑)P.^ #^^^Z × (^[: , : , ^]% (23)

[0170] After that, the second convolutional layer ^&kl^(followed by a respective non-linear activation layer) of the VGG encoder and the decoder ^m&lcan be used to reconstruct a semantictransformed content image ^p^^ .

[0171] At the third stage, the post processing including but not limited to guided filtering may be applied or used to remove or reduce boundary artifacts that might be caused by inaccurate segmentation masks or outputs in (e.g., visual objects, between different semantic regions, etc.) boundary regions to generate or produce the final output (semantic transferredcontent image) denoted as ^^^ = f^#^p^^ , ^^%.TABLE 1 Algorithm – 1: Transferring Color from Single Style Image to Content Image Input: style image ^^, content image^^^^ ∈ ℝ^^×^^×^, ^^ ∈ℝ^^×^^×^Output: semantic color transferred image ^^= num^^ber of semantic labels1. Semantic Segmentation a. ^^= ^^^_^^^^^ ^^^^^ o^^; ^^^^^y^ , ^^ ∈ℝ^^×^^; b. ^^ = ^^^^ o^^; ^^^^^ y , ^^ ∈ℝ^^×^^×^; 2. Color and Brightness Transfer (WCT – style mage to content image or I2I) a. ^^ = ^^^^^#^^%, ^^ =^^^^^#^^%; / / ^^ ∈ ℝ^^×^^×^, ^^ ∈ℝ^^×^^×^, ^= feature sizeb. for semantic labels, ^ in ^^do c. ^^^ = ^^[^^ == ^];d. ^^^^ = ^^ #^^, ^^^%;e. end f. for semantic labels, ^ not in ^^do g. ^^^^ = ^^;h. end i. ^p^^ = ∑^^.^ #^^^^ ×^^[: , : , ^]%;j. ^p^^ = ^¡^^#^^^^¢D^p^^ E%;3. Post-processing a. ^^^ = £^#^p^^ , ^^%;13. SINGLE STYLE IMAGE AND CONTENT VIDEO

[0172] FIG.3B illustrates an example process flow, method or algorithm (denoted as Algorithm 2) for transferring color and brightness from a single style image denoted as ^^to a content video denoted as z^(with the total number of images / frames denoted as Pc). As shown, the system receives the single style image ^^and the content video z^as input.

[0173] The system can transfer a color and brightness look of a semantic region / object of a semantic class / category in ^^to semantic region(s) / object(s) of the same (or a corresponding) semantic class / category present in z^to generate or reconstruct a style transferred content videodenoted as z as output. The sing ^^×^^×^^^ le style image ^^ ∈ ℝ and content video ¤^ ∈may include input color images. ^^and ^^represent the height and width of the style image. ^^and ^^represent the height and width of images / frames in the content video. Video decoding operations may be performed – for example as part of the pre-processingoperations – on the content video to generate or obtain content video images / frames denoted as^ @^ = z,¥5¦^5§¦¥58#¤^%, where ¨ = 1, 2, … ^^.

[0174] Algorithm 2 (also illustrated in TABLE 2 below) implements a pipeline with four main stages: semantic segmentation with a semantic segmentation model, finetuning resultant segmentation masks of content video frames, color and brightness transfer from the style image to the content image with a semantic informed WCT in a color transfer network further including a VGG encoder (VGG Encoder-1 and VGG Encoder-2) and a decoder, and post- processing.

[0175] More specifically, at the first stage, the semantic segmentation is performed on both^^ and ^ @^ . / ^ ∈ ℝ^^×^^ denotes the output (e.g., map, mask, arrays, etc.) of the semanticsegmentation performed on the style image. This semantic segmentation output includes, for each pixel location of some or all pixels in the style image, a specific semantic class label to which a pixel at the pixel location belongs.

[0176] The output (e.g., map, mask, arrays, etc.) of the semantic segmentation performed onthe content video or the images / frames therein – represented or denoted as (p ^^ ∈ ℝ ^×^^×)×^^(where the superscript p for the (each) p-th image / frame of the content video is omitted for simplicity) – includes, for each pixel location of some or all pixels in each image / frame of the content video, semantic class labels to which a pixel at the pixel location may belong along with their respective probability values. Each pixel of the semantic segmentation output of theimage / frame of the content video in an ^^ × ^^ matrix has * different values. These *different values indicate specific probabilities of each pixel’s belongingness to or memberships of * different semantic class labels. The sum of all the probability values at each pixel locationis one: ∑)+.^ (p^ = ^, ^ ∈ ℝ^^×^^ , for each frame p, where ^ is an all-one matrix (or the matrixof all including both diagonal and off-diagonal matrix elements).

[0177] At the second stage, an online training mechanism may be used to finetune thecontent segmentation masks (p^ to remove or prevent flickering caused by temporallyinconsistent segmentation (or semantic class) labels from the final output or semantic transferred content video and to generate temporally stable segmentation (or semantic class) masks denoted as (^@for the (each) p-th image / frame of the content video.

[0178] At the third stage, the color transfer process or operation is performed or started. The first convolutional layer ^&kl^(followed by a corresponding nonlinear activation function layer)of the VGG encoder may receive the style image and extract style features ^ ^^ ∈ ℝ ^×^^×^ fromthe style image, where ^ represents the total number of extracted image or spatial features of these images. Also, the first convolutional layer ^&kl^(followed by a corresponding nonlinearactivation function layer) of the VGG encoder may receive the (e.g., decoded, etc.) images / frames of the input content video, process all the images / frames of the content video independently, and extract content features ^^@from each image / frame ^^@(which is the p-th image / frame) of the content video.

[0179] The style features ^^extracted from the style image may be filtered or specifically selected to generate filtered (or selected) features that belong to a specific semantic class present in / ^. These filtered features may represent a specific color and brightness look for the specific semantic class present in the feature domain of the style image. Based on these filtered features filtered from ^^, the WCT transform may be applied or invoked to transfer the specific color and brightness look of the specific semantic class in the style image to the content features ^^@of a semantic class corresponding to the specific semantic class in the style image.

[0180] For a semantic class label (corresponding to a specific semantic class) denoted as ^, its transformed features for the specific semantic class ^ as generated with the WCT transform may be denoted as ^^^Z@. For any class out of the * different semantic class labels not present in the style image, the same ^^@features (not semantically transferred) may be kept or used. After completing this WCT assisted process for all the * different classes, * different transformed content features are generated. The previously generated segmentation output (^@for the (each) p-th image / frame of the content video shows or indicates respective probability values of memberships or belongingness of each pixel into each of the N semantic class labels or classes for each pixel. These * different probability values from (^@can be used to multiply with * different transformed features of the (each) p-th image / frame of the content video to generate aweighted sum (denoted as ^p @ @^^ ) of transformed features ^^^Z , where ^ = 1, 2, … . , *, asfollows: ^p @ @^^ = ∑)P.^ #^^^Z × ( @^ [: , : , ^]% (24)

[0181] non- activation layer) of the VGG encoder and the decoder ^m&lcan be used to reconstruct a semantictransformed content image ^p @^^ .

[0182] At the fourth stage, the post processing including but not limited to guided filtering may be applied or used to remove or reduce boundary artifacts that might be caused by inaccurate segmentation masks or outputs in (e.g., visual objects, between different semantic regions, etc.) boundary regions to generate or produce the final output (semantic transferredcontent image) denoted as ^ @ = f^#^p @, @^^ ^^ ^^ % to be included in the final output (semantictransferred content video ¤^^.TABLE 2 Algorithm – 2: Transferring Color from Single Style Image to Content Video Input: style image ^^, content video ¤^^^ ∈ ℝ^^×^^×^, ¤^ ∈ℝ^^×^^×^ש^Output: semantic color and brightnesstransferred video ¤^ = number of sem^a^ntic labels1. Video Decode a. ^ ª^ =«¬¡^­®^^­¡^¯#¤^% °±^¯^ ª =^, ¢, … ©^2. Semantic Segmentation a. ^^=^^^_^^^^^ ^^^^^ o^^; ^^^^^y^ , ^^ ∈ℝ^^×^^; b. ^p^ =^^^^ o^ ª^; ^^^^^y °±^¯^ ª =^, ¢, … ©^, ^p^ ∈ ℝ^^×^^×^ש^;3. Finetune the Segmentation Masks of Content Frames a. ^^=^^^^^^² o^ ª^, ^p^ ; ^^^^^v^² y °±^¯^ ª =^, ¢, … ©^;4. Color and Brightness Transfer (WCT – style image to content video or I2V) a. ^^ = ^^^^^#^^%b. for fra ³me, ³ in all © ³^frames do c. ^^ = ^^^^ #^^ %;³ ^d. ^^ = ^^[: , : , : , ³]e. for semantic label, ^ in ^^do f. ^^^ = ^^[^^ == ^];g. ^ ³ ³^^^ = ^^ #^^ , ^^^%;h. end i. for semantic label, ^ not in ^^do j. ^ ³ ³^^^ = ^^ ;k. end l. ^p ³^^ = ∑^^.^ #^ ³^^ ׳ ^^^ [: , : , ^]%;m. ^p ³= ^ # p³^^ ¡^^ ^^^^¢#^^^%%; n. end 5. Post-processing a. for frame, ³ in all frames d b. ^^^ = £^#^p ³o ³³^^ , ^^ %;c. end 6. Video Encode a. ¤^^=«¬¡^­´^^­¡^¯#^ ^ ¢ ©^^ , ^^^ , … … ^^^ ^%14. SINGLE STYLE VIDEO AND CONTENT VIDEO

[0183] FIG.3C illustrates an example process flow, method or algorithm (denoted as Algorithm 3) for transferring color and brightness from a single style video denoted as z[(with the total number of images / frames denoted as Ps) to a content video denoted as z^(with the total number of images / frames denoted as Pc). As shown, the system receives the style z[and the single content video z^as input.

[0184] The system can transfer a color and brightness look of a semantic region / object of a semantic class / category in z^to semantic region(s) / object(s) of the same (or a corresponding) semantic class / category present in z^to generate or reconstruct a style transferred content videodenoted as z^^ as output. The single style video ¤ ^^ ∈ ℝ ^×^^×^×^^ and content video ¤^ ∈ℝ^^×^^×^×^^may include input color images. ^^and ^^represent the height and width of (style) images / frames of the style video. ^^and ^^represent the height and width of (content) images / frames in the content video. Video decoding operations may be performed – for example, as part of the pre-processing operations – on the style video to generate or obtain stylevideo images / frames denotes as ^ B^ = z,¥5¦^5§¦¥58#¤^% µℎ585 · = 1, 2, … ^^, and may alsobe performed on the content video to generate or obtain content video images / frames denoted as^ @^ = z,¥5¦^5§¦¥58#¤^%, where ¨ = 1, 2, … ^^.

[0185] Algorithm 3 (also illustrated in TABLE 3 below) implements a pipeline with four main stages: semantic segmentation with a semantic segmentation model, finetuning resultant segmentation masks of content video frames, color and brightness transfer from the style video to the content image with a semantic informed WCT in a color transfer network further including a VGG encoder (VGG Encoder-1 and VGG Encoder-2) and a decoder, and post- processing.

[0186] More specifically, at the first stage, the semantic segmentation is performed on both^ B^ and ^ @^ . / ^ ∈ ℝ^^×^^×^^ denotes the output (e.g., map, mask, arrays, etc.) of the semanticsegmentation performed on the style video. This semantic segmentation output includes, for each pixel location of some or all pixels in a specific (e.g., each, etc.) image / frame of the style video, a specific semantic class label to which a pixel at the pixel location belongs.

[0187] The output (e.g., map, mask, arrays, etc.) of the semantic segmentation performed onthe content video or the images / frames therein – represented or denoted as (p ^^ ∈ ℝ ^×^^×)×^^(where the superscript p for the (each) p-th image / frame of the content video is omitted for simplicity) – includes, for each pixel location of some or all pixels in each image / frame of thecontent video, semantic class labels to which a pixel at the pixel location may belong along with their respective probability values. Each pixel of the semantic segmentation output of theimage / frame of the content video in an ^^ × ^^ matrix has * different values. These *different values indicate specific probabilities of each pixel’s belongingness to or memberships of * different semantic class labels. The sum of all the probability values at each pixel locationis one: ∑)+.^ (p^ = ^, ^ ∈ ℝ^^×^^ , for each frame p, where ^ is an all-one matrix (or the matrixof all ones both and off-diagonal matrix elements).

[0188] At the training mechanism may be used to finetune thecontent segmentation masks (p^ to remove or prevent flickering caused by temporallyinconsistent segmentation (or semantic class) labels from the final output or semantic transferred content video and to generate temporally stable segmentation (or semantic class) masks denoted as (^@for the (each) p-th image / frame of the content video.

[0189] At the third stage, the color transfer process or operation is performed or started. The first convolutional layer ^&kl^(followed by a corresponding nonlinear activation function layer) of the VGG encoder may receive all the style images / frames in the style video and extract stylefeatures ^ ^ ×^ ×^×^ ∈ ℝ ^ ^ ^s from all the style images / frames, where ^ represents the totalnumber of extracted image or spatial features of these images. Also, the first convolutional layer ^&kl^(followed by a corresponding nonlinear activation function layer) of the VGG encoder may receive the (e.g., decoded, etc.) images / frames of the input content video, process all the images / frames of the content video independently, and extract content features ^^@from each image / frame ^^@(which is the p-th image / frame) of the content video.

[0190] The style features ^^extracted from the style images / frames of the style video may be filtered or specifically selected to generate filtered (or selected) features that belong to a specific semantic class present in / ^. These filtered features may represent a specific color and brightness look for the specific semantic class present in the feature domain of the style images / frames of the style video. Based on these filtered features filtered from ^^, the WCT transform may be applied or invoked to transfer the specific color and brightness look of the specific semantic class in the style images / frames of the style video to the content features ^^@of a semantic class corresponding to the specific semantic class in the style images / frames of the style video.

[0191] For a semantic class label (corresponding to a specific semantic class) denoted as ^, its transformed features for the specific semantic class ^ as generated with the WCT transform may be denoted as ^^^Z@. For any class out of the * different semantic class labels not present in the style images / frames of the style video, the same ^^@features (not semantically transferred)may be kept or used. After completing this WCT assisted process for all the * different classes, * different transformed content features are generated. The previously generated segmentation output (^@for the (each) p-th image / frame of the content video shows or indicates respective probability values of memberships or belongingness of each pixel into each of the N semantic class labels or classes for each pixel. These * different probability values from (^@can be used to multiply with * different transformed features of the (each) p-th image / frame of the contentvideo to generate a weighted sum (denoted as ^p @^^ ) of transformed features ^ @^^Z , where ^ =1, 2, … . , *, as follows:^p @= ∑) # @ @^^ P.^ ^^^Z × (^ [: , : , ^]% (25)

[0192] After that, the second convolutional layer ^&kl^(followed by a respective non-linear activation layer) of the VGG encoder and the decoder ^m&lcan be used to reconstruct a semantictransformed content image ^p @^^ .

[0193] At the fourth stage, the post processing including but not limited to guided filtering may be applied or used to remove or reduce boundary artifacts that might be caused by inaccurate segmentation masks or outputs in (e.g., visual objects, between different semantic regions, etc.) boundary regions to generate or produce the final output (semantic transferredcontent image) denoted as ^ @^^ = f^#^p @^^ , ^ @^ % to be included in the final output (semantictransferred content video ¤^^. TABLE 3 Algorithm – 3: Transferring Color froma. ^^= ^^^^^^² o^ ª^, ^p^ ; ^^^^^v^² y °±^¯^ ª =^, ¢, … ©^;4. Color and Brightness Transfer (WCT – style video to content video or V2V) a. for frame, ³ in all © fr b. ^ ³ ³^ames do ^= ^^^^^#^^ %;c. end d. ^^ = [^ ^^ , ^ ¢^ , … . , ^ ©^ ^ ]e. for fra ³me, ³ in all ©^frames do f. ^^ = ^^^^^#^ ³^ %;g. ^ ³^ = ^^[: , : , : , ³]h. for semantic label, ^ in ^^do i. ^^^ = ^^[^^ == ^];j. ^ ³^ = ³^^ ^^ #^^ , ^^^%;k. end l. for seman ³tic labe ³l, ^ not in ^^do m. ^^^^ = ^^ ;n. end o. ^p ³= ∑^ ³^^ ^.^ #^^^ ׳ ^^^ [: , : , ^]%;p. ^p ³= ^^ ^^^¢ p³^^ ^¡ ^^ o^^^y^; q. end 5. Post-processing a. for frame, ³ in all © frames b. ^ = p ³^do ³³^^ £^#^^^ , ^^ %;c. end 6. Video Encode a. ¤^^= «¬¡^­´^^­¡^¯#^ ^ ¢ ©^^ , ^^^ , … … ^^^ ^%15. MULTIPLE STYLE IMAGES AND CONTENT VIDEO

[0194] FIG.3D illustrates an example process flow, method or algorithm (denoted as Algorithm 4) for transferring color and brightness from multiple style images denoted as^^¸ where · = 1, 2, … " (with the total number of style images denoted as T) with respectivesemantic label information – for example, semantic class labels denoted as 3^=¹3^^ , 3^^ , … … , 3^º» as provided by the user in user input to the system as part of the pre-operations – specific to each of the style images ^^¸to a content video denoted as z^(with the total number of images / frames denoted as Pc). As shown, the system receives the style images ^^¸and the single content video z^as input.

[0195] The system can transfer a color and brightness look of a semantic region / object of a semantic class / category in 3^to semantic region(s) / object(s) of the same (or a corresponding) semantic class / category present in z^to generate or reconstruct a style transferred content videodenoted as z as output. The style ima ^^×^^×^^^ ges ^^¸ ∈ ℝ and content video ¤^ ∈ℝ^^×^^×^×^^may include input color images. ^^and ^^represent the height and width of style images ^^¸. ^^and ^^represent the height and width of (content) images / frames in the content video. Video decoding operations may be performed – for example, as part of the pre- processing operations – on the content video to generate or obtain content video images / framesdenoted as ^ @^ = z,¥5¦^5§¦¥58#¤^%, where ¨ = 1, 2, … ^^ .

[0196] Algorithm 4 (also illustrated in TABLE 4 below) implements a pipeline with four main stages: semantic segmentation with a semantic segmentation model, finetuning resultant segmentation masks of content video frames, color and brightness transfer from the style images to the content image with a semantic informed WCT in a color transfer network further including a VGG encoder (VGG Encoder-1 and VGG Encoder-2) and a decoder, and post- processing.

[0197] As part of the pre-processing operations, the user can provide user input based at least in part on which semantic class correspondence relationships may be generated, mapped, specified and / or defined between the style images and the content video. For example, as noted, the user input provided by the user may be used to assign respective semantic class labels 3^=¹3^^ , 3^^ , … … , 3^º» to the style images ^^¸ where · = 1, 2, … ". The user input provided by theuser may also generate the semantic class correspondence relationships in which input classlabels denoted as 3^ = ¹3^^ , 3^^ , … … , 3^º» may be used or specified to identify a specific styleimage (e.g., the j-th, etc.) among all the style images. For example, a semantic class correspondence relationship as described herein may specify a semantic class in the (j-th) specific style image – for example, as identified with a semantic class label 3^¸in the semantic class correspondence relationship – to which a semantic class that may be present or detected in the content video corresponds.

[0198] At the first stage, the semantic segmentation is performed on both ^^¸and ^^@. / ^¸∈ℝ^^×^^ where · = 1, 2, … " denotes the output (e.g., map, mask, arrays, etc.) of the semanticsegmentationon the style images. This semantic segmentation output includes, for each pixel location of some or all pixels in a specific (e.g., each, etc.) style image of the style images, a specific semantic class label to which a pixel at the pixel location belongs.

[0199] The output (e.g., map, mask, arrays, etc.) of the semantic segmentation performed onthe content video or the images / frames therein – represented or denoted as (p ^^ ∈ ℝ ^×^^×)×^^(where the superscript p for the (each) p-th image / frame of the content video is omitted for simplicity) – includes, for each pixel location of some or all pixels in each image / frame of thecontent video, semantic class labels to which a pixel at the pixel location may belong along with their respective probability values. Each pixel of the semantic segmentation output of theimage / frame of the content video in an ^^ × ^^ matrix has * different values. These *different values indicates specific probabilities of each pixel’s belongingness to or memberships of * different semantic class labels. The sum of all the probability values at each pixel locationis one: ∑)+.^ (p^ = ^, ^ ∈ ℝ^^×^^ , for each frame, where ^ is an all-one matrix (or the matrix ofall ones including both diagonal and off-diagonal matrix elements).

[0200] At the second stage, an online training mechanism may be used to finetune thecontent segmentation masks (p^ to remove or prevent flickering caused by temporallyinconsistent segmentation (or semantic class) labels from the final output or semantic transferred content video and to generate temporally stable segmentation (or semantic class) masks denoted as (^@for the (each) p-th image / frame of the content video.

[0201] At the third stage, the color transfer process or operation is performed or started. The first convolutional layer ^&kl^(followed by a corresponding nonlinear activation function layer) of the VGG encoder may receive all the style images and extract style features ^^¸∈ℝ^^×^^×^ where · = 1, 2, … " from all the style images, where ^ represents the total number ofextracted image orfeatures of these images. Also, the first convolutional layer ^&kl^(followed by a corresponding nonlinear activation function layer) of the VGG encoder may receive the (e.g., decoded, etc.) images / frames of the input content video, process all the images / frames of the content video independently, and extract content features ^^@from each image / frame ^^@(which is the p-th image / frame) of the content video.

[0202] The style features ^^¸extracted from the style images may be filtered or specifically selected to generate filtered (or selected) features that belong to a specific semantic class – for example, as labeled with 3^¸– present in / ^. These filtered features may represent a specific color and brightness look for the specific semantic class – as labeled with 3^¸– present in the feature domain of the style images such as the j-th style image. Based on these filtered features filtered from ^^¸, the WCT transform may be applied or invoked to transfer the specific color and brightness look of the specific semantic class – as labeled with 3^¸– in the j-th style image to the content features ^^@of a semantic class corresponding to the specific semantic class in the j-th style image.

[0203] For a semantic class label (corresponding to a specific semantic class) denoted as ^, its transformed features for the specific semantic class ^ as generated with the WCT transform may be denoted as ^^^Z@. For any class out of the * different semantic class labels not present inthe style images, the same ^^@features (not semantically transferred) may be kept or used. After completing this WCT assisted process for all the * different classes, * different transformed content features are generated. The previously generated segmentation output (^@for the (each) p-th image / frame of the content video shows or indicates respective probability values of memberships or belongingness of each pixel into each of the N semantic class labels or classes for each pixel. These * different probability values from (^@can be used to multiply with * different transformed features of the (each) p-th image / frame of the content video to generate aweighted sum (denoted as ^p @ @^^ ) of transformed features ^^^Z , where ^ = 1, 2, … . , *, asfollows: ^p @= ^p @+ ^ @ × ( @[ @^^ ^^ ^^Z ^ : , : , 3^¸[^]]. ^^^Z (26)

[0204] non-linearactivation layer) of the VGG encoder and the decoder ^m&lcan be used to reconstruct a semantictransformed content image ^p @^^ .

[0205] At the fourth stage, the post processing including but not limited to guided filtering may be applied or used to remove or reduce boundary artifacts that might be caused by inaccurate segmentation masks or outputs in (e.g., visual objects, between different semantic regions, etc.) boundary regions to generate or produce the final output (semantic transferredcontent image) denoted as ^ @ = f^#^p @, @^^ ^^ ^^ % to be included in the final output (semantictransferred content video ¤^^. TABLE 4 Algorithm – 4: Transferring Color from Multiple Style Images to Content Video Input: style images ^^ª °±^¯^ ª = ^, ¢, …   ,content video ¤^^^ ^ ×^ ×^ª ∈ ℝ ^ ^ , ¤^ ∈ℝ^^×^^×^ש^^^ª: semantic labels to be selected from style image ^^ª °±^¯^ ª = ^, ¢, …  ^^ = ¹^^^ , ^^¢ , … … , ^^ »^^ª: corresponding semantic labels of each style label ^^ªto the semantic label in content video ¤^ °±^¯^ ª =^, ¢, …  ^^ = ¹^^^ , ^^¢ , … … , ^^ »Output: semantic color and brightness transferred video ¤^^^ = number of semantic labels,   = numberof style images 1. Video Decodea. ^ ª^ =«¬¡^­®^^­¡^¯#¤^% °±^¯^ ª =^, ¢, … ©^2. Semantic Segmentation a. ^^ª=^^^_^^^^^ ^^^^^ o^^ª; ^^^^^ y^ °±^¯^ ª =^, ¢, …  , ^^ ^ ×^ª ∈ ℝ ^ ^;b. ^p^ = ^^^^ o^ ª^; ^^^^^y °±^¯^ ª =^, ¢, … ©^, ^p^ ∈ ℝ^^×^^×^ש^;3. Finetune the Segmentation Masks of Content Frames a. ^^=^^^^^^² o^ ª^, ^p^ ; ^^^^^v^² y °±^¯^ ª =^, ¢, … ©^;4. Color and Brightness Transfer (WCT – multiple style images to content video or MI2V) a. ^^ª = ^^^^^ o^^ªy °±^¯^ ª =^, ¢, …  ;b. for frame, ³ in all ©^frames do c. ^ ³^ = ^^^^^#^ ³^ %;d. ^ ³^. ^ ³ = ^^[: , : , : , ³];e p^^ = ¼;f. for ª = ^, ¢, …   dog. for all ^²±semantic label in ^^ªdo h. ^^^ª = ^^ª[^^ª ==^^ª[^]]; i. ^ ³ = ^^ # ³^^^ ^^ , ^^^ª%;j. ^p ³^^ = ^p ³^^ + ^ ³^^ ׳ ^^^ [: , : , ^^ª[^]];k. end l. end m. for sema ³ntic label, ^ not in ^^do n. ^ = ^ ³^^^ ^ ;o. ^p ³^^ = ^p ³^^ + ^ ³^^^ ×^ ³^ [: , : , ^];p. end q. ^p ³^^ = ^¡^^ ^^^^^¢ o^p³^^y^; r. end 5. Post-processing a. for frame, ³ in all © b. ^ = £^#^p ³^frames do ³, ^ ³^^ ^^ ^ %;c. end 6. Video Encode a. ¤^^=«¬¡^­´^^­¡^¯#^ ^ ¢ ©^^ , ^^^ , … … ^^^ ^%;16. EXAMPLE PROCESS FLOWS

[0206] FIG.4 illustrates an example process flow according to an embodiment. In some embodiments, one or more computing devices or components (e.g., a semantic transfer system, a semantic color and brightness transfer system, one or more video codecs, an encoding device / module, a transcoding device / module, a coding device / module, a volumetric or immersive video server system, etc.) may perform this process flow. In block 402, an image processing system establishes one or more class correspondence relationships between one or more first semantic classes to which one or more first semantic regions of one or more content images belong and one or more second semantic classes to which one or more second semantic regions of one or more style images belong.

[0207] In block 404, the system uses a first convolutional layer initialized with first pretrained weights of a relatively deep convolutional network (denoted as VGG) for relatively large-scale image recognition to extract specific content features in the one or more first semantic regions of one or more content images and to extract specific style features in the one or more second semantic regions of one or more style images.

[0208] Based at least in part on (a) the one or more class correspondence relationships, (b) the specific content features and (c) the specific style features, in block 406, the system applies a whitening and coloring transform (WCT) to transfer a specific color and brightness look associated with the one or more second semantic regions of the one or more style images to the specific content features in the one or more first semantic regions of the one or more content images.

[0209] In block 406, the system uses a second convolutional layer initialized with second pretrained weights of the VGG to process the specific content features with the transferred specific color and brightness look into second specific content features.

[0210] In block 408, the system or a decoder therein generates, based at least in part on the second specific content features, one or more semantic content transferred images corresponding to the one or more content images.

[0211] In an embodiment, the system further performs: applying one or more segmentation tools to generate one or more segmentation maps for the one or more content images, respectively, each of the one or more segmentation maps specifying, for each pixel location in a respective content image of the one or more content images, a plurality of first semantic classes along with a plurality of probability values; applying the one or more segmentation tools to segment the one or more style images into the one or more second semantic regions, each second semantic region of the one or more second semantic regions belonging to a secondrespective semantic class of one or more second semantic classes.

[0212] In an embodiment, the system further performs: filtering the specific style features for each second semantic class of the one or more second semantic classes present in the one or more style images to generate respective per-class filtered style features for the second semantic classes; applying the WCT to transfer the respective per-class filtered style features for the second semantic classes to the specific content features based at least in part on the one or more class correspondence relationships to generate a plurality of per-class transformed content features for the plurality of the first semantic classes; generating, based on the plurality of probability values, for the pixel location in the respective content image of the one or more content images, a weighted sum of the plurality of the per-class transformed content features, wherein the weight sum is included in the specific content features with the transferred specific color and brightness look.

[0213] In an embodiment, the one or more content images represent a content video; wherein the one or more segmentation maps for the one or more content images are generated with a learning model by finetuning one or more initial segmentation maps directly created from the one or more segmentation tools; wherein the learning model is trained online with a relatively small set of content images in the content video.

[0214] In an embodiment, the system further performs post-processing operations on the one or more semantic content transferred images to generate one or more post-processed semantic content transferred images.

[0215] In an embodiment, the post-processing operations include a guided filtering operation.

[0216] In an embodiment, the one or more first semantic regions of the one or more content images are delineated with one or more segmentation masks; the WCT is applied based further on the one or more segmentation masks.

[0217] In an embodiment, the one or more segmentation masks are generated by fine-tuning one or more temporally unstable segmentation masks of the one or more content images with a trained deep learning model; the one or more temporally unstable segmentation masks are generated from the one or more content images by one or more segmentation tools.

[0218] In an embodiment, the one or more content images originate from a content video that includes the one or more content images as a sequence of consecutive images.

[0219] In an embodiment, the one or more style images originate from a style video that includes the one or more style images as a sequence of consecutive images.

[0220] In an embodiment, the one or more style images are independent of one another anddo not form a sequence of consecutive images in a style video.

[0221] In an embodiment, the decoder includes two convolutional layers with respective weights initialized to random values and trained offline with a training dataset comprising a population of training content images.

[0222] In an embodiment, the one or more first semantic classes are assigned with one or more first semantic labels, respectively; the one or more second semantic classes are assigned with one or more second semantic labels, respectively.

[0223] In an embodiment, the one or more second semantic classes include at least two second semantic classes from at least two different style images; the one or more second semantic labels include at least two second semantic labels designated to fully distinguish among the at least two different style images; the one or more semantic class correspondence relationships include a semantic class correspondence relationship between one of the one or more first semantic classes and one of the at least two second semantic classes with one of the at least two second semantic labels.

[0224] In an embodiment, the one or more second semantic classes include at least two second semantic classes present in at least two independent style images, respectively; the one or more second semantic labels include at least two second semantic labels designated to fully distinguish the at least two independent style images, respectively; the one or more semantic class correspondence relationships identify a correspondence relationship between (a) one of the at least two second semantic classes present in one of the at least two independent style images based on one of the at least two second semantic labels, and (b) one of the one or more first semantic classes present in one of the one or more content images.

[0225] In an embodiment, the one or more content images is composed of a single content image; wherein the one or more style images is composed of a single style image.

[0226] In an embodiment, the one or more content images are decoded from a single content video; wherein the one or more style images is composed of a single style image.

[0227] In an embodiment, the one or more content images are decoded from a single content video; wherein the one or more style images are decoded from a single style video.

[0228] In an embodiment, the one or more content images are decoded from a single content video; wherein the one or more style images include a plurality of independent style images.

[0229] In an embodiment, the first convolutional layer operates with zero or more additional first convolutional layers to extract the specific content features; the second convolutional layer operates with zero or more additional second convolutional layers to generate the second specific content features.

[0230] In an embodiment, the first and second convolutional layers are selected from among the lowest level convolutional layers of the VGG with publicly accessible pretrained weights.

[0231] In an embodiment, a computing device such as a display device, a mobile device, a set-top box, a multimedia device, etc., is configured to perform any of the foregoing methods. In an embodiment, an apparatus comprises a processor and is configured to perform any of the foregoing methods. In an embodiment, a non-transitory computer readable storage medium, storing software instructions, which when executed by one or more processors cause performance of any of the foregoing methods.

[0232] In an embodiment, a computing device comprising one or more processors and one or more storage media storing a set of instructions which, when executed by the one or more processors, cause performance of any of the foregoing methods.

[0233] Note that, although separate embodiments are discussed herein, any combination of embodiments and / or partial embodiments discussed herein may be combined to form further embodiments. 17. IMPLEMENTATION MECHANISMS – HARDWARE OVERVIEW

[0234] Embodiments of the present invention may be implemented with a computer system, systems configured in electronic circuitry and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA), or another configurable or programmable logic device (PLD), a discrete time or digital signal processor (DSP), an application specific IC (ASIC), and / or apparatus that includes one or more of such systems, devices or components. The computer and / or IC may perform, control, or execute instructions relating to the adaptive perceptual quantization of images with enhanced dynamic range, such as those described herein. The computer and / or IC may compute any of a variety of parameters or values that relate to the adaptive perceptual quantization processes described herein. The image and video embodiments may be implemented in hardware, software, firmware and various combinations thereof.

[0235] Certain implementations of the invention comprise computer processors which execute software instructions which cause the processors to perform a method of the disclosure. For example, one or more processors in a display, an encoder, a set top box, a transcoder or the like may implement methods as described above by executing software instructions in a program memory accessible to the processors. Embodiments of the invention may also be provided in the form of a program product. The program product may comprise any non-transitory medium which carries a set of computer-readable signals comprising instructions which, when executedby a data processor, cause the data processor to execute a method of an embodiment of the invention. Program products according to embodiments of the invention may be in any of a wide variety of forms. The program product may comprise, for example, physical media such as magnetic data storage media including floppy diskettes, hard disk drives, optical data storage media including CD ROMs, DVDs, electronic data storage media including ROMs, flash RAM, or the like. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0236] Where a component (e.g. a software module, processor, assembly, device, circuit, etc.) is referred to above, unless otherwise indicated, reference to that component (including a reference to a "means") should be interpreted as including as equivalents of that component any component which performs the function of the described component (e.g., that is functionally equivalent), including components which are not structurally equivalent to the disclosed structure which performs the function in the illustrated example embodiments of the invention.

[0237] According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing devices may be hard-wired to perform the techniques, or may include digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) that are persistently programmed to perform the techniques, or may include one or more general purpose hardware processors programmed to perform the techniques pursuant to program instructions in firmware, memory, other storage, or a combination. Such special- purpose computing devices may also combine custom hard-wired logic, ASICs, or FPGAs with custom programming to accomplish the techniques. The special-purpose computing devices may be desktop computer systems, portable computer systems, handheld devices, networking devices or any other device that incorporates hard-wired and / or program logic to implement the techniques.

[0238] For example, FIG.5 is a block diagram that illustrates a computer system 500 upon which an embodiment of the invention may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for communicating information, and a hardware processor 504 coupled with bus 502 for processing information. Hardware processor 504 may be, for example, a general purpose microprocessor.

[0239] Computer system 500 also includes a main memory 506, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 also may be used for storing temporary variables or other intermediate information during execution of instructions to beexecuted by processor 504. Such instructions, when stored in non-transitory storage media accessible to processor 504, render computer system 500 into a special-purpose machine that is customized to perform the operations specified in the instructions.

[0240] Computer system 500 further includes a read only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504. A storage device 510, such as a magnetic disk or optical disk, is provided and coupled to bus 502 for storing information and instructions.

[0241] Computer system 500 may be coupled via bus 502 to a display 512, such as a liquid crystal display, for displaying information to a computer user. An input device 514, including alphanumeric and other keys, is coupled to bus 502 for communicating information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, a trackball, or cursor direction keys for communicating direction information and command selections to processor 504 and for controlling cursor movement on display 512. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allows the device to specify positions in a plane.

[0242] Computer system 500 may implement the techniques described herein using customized hard-wired logic, one or more ASICs or FPGAs, firmware and / or program logic which in combination with the computer system causes or programs computer system 500 to be a special-purpose machine. According to one embodiment, the techniques as described herein are performed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequences of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions.

[0243] The term “storage media” as used herein refers to any non-transitory media that store data and / or instructions that cause a machine to operation in a specific fashion. Such storage media may comprise non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 510. Volatile media includes dynamic memory, such as main memory 506. Common forms of storage media include, for example, a floppy disk, a flexible disk, hard disk, solid state drive, magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH-EPROM, NVRAM, any other memory chip or cartridge.

[0244] Storage media is distinct from but may be used in conjunction with transmission media. Transmission media participates in transferring information between storage media. For example, transmission media includes coaxial cables, copper wire and fiber optics, including the wires that comprise bus 502. Transmission media can also take the form of acoustic or light waves, such as those generated during radio-wave and infra-red data communications.

[0245] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 504 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 500 can receive the data on the telephone line and use an infra-red transmitter to convert the data to an infra-red signal. An infra-red detector can receive the data carried in the infra-red signal and appropriate circuitry can place the data on bus 502. Bus 502 carries the data to main memory 506, from which processor 504 retrieves and executes the instructions. The instructions received by main memory 506 may optionally be stored on storage device 510 either before or after execution by processor 504.

[0246] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides a two-way data communication coupling to a network link 520 that is connected to a local network 522. For example, communication interface 518 may be an integrated services digital network (ISDN) card, cable modem, satellite modem, or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 518 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. Wireless links may also be implemented. In any such implementation, communication interface 518 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.

[0247] Network link 520 typically provides data communication through one or more networks to other data devices. For example, network link 520 may provide a connection through local network 522 to a host computer 524 or to data equipment operated by an Internet Service Provider (ISP) 526. ISP 526 in turn provides data communication services through the world wide packet data communication network now commonly referred to as the “Internet” 528. Local network 522 and Internet 528 both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 520 and through communication interface 518, which carry the digital data to and from computer system 500, are example forms of transmission media.

[0248] Computer system 500 can send messages and receive data, including program code, through the network(s), network link 520 and communication interface 518. In the Internet example, a server 530 might transmit a requested code for an application program through Internet 528, ISP 526, local network 522 and communication interface 518.

[0249] The received code may be executed by processor 504 as it is received, and / or stored in storage device 510, or other non-volatile storage for later execution. 18. EQUIVALENTS, EXTENSIONS, ALTERNATIVES AND MISCELLANEOUS

[0250] In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Thus, the sole and exclusive indicator of what is claimed embodiments of the invention, and is intended by the applicants to be claimed embodiments of the invention, is the set of claims that issue from this application, in the specific form in which such claims issue, including any subsequent correction. Any definitions expressly set forth herein for terms contained in such claims shall govern the meaning of such terms as used in the claims. Hence, no limitation, element, property, feature, advantage or attribute that is not expressly recited in a claim should limit the scope of such claim in any way. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense. Enumerated Exemplary Embodiments

[0251] The invention may be embodied in any of the forms described herein, including, but not limited to the following Enumerated Example Embodiments (EEEs) which describe structure, features, and functionality of some portions of embodiments of the present invention.

[0252] EEE1. A method, comprising: establishing one or more class correspondence relationships between one or more first semantic classes to which one or more first semantic regions of one or more content images belong and one or more second semantic classes to which one or more second semantic regions of one or more style images belong; using a first convolutional layer initialized with a convolutional network for image recognition to extract specific content features in the one or more first semantic regions of one or more content images and to extract specific style features in the one or more second semantic regions of one or more style images;based at least in part on (a) the one or more class correspondence relationships, (b) the specific content features and (c) the specific style features, applying a semantic style transform to transfer a specific color and brightness look associated with the one or more second semantic regions of the one or more style images to the specific content features in the one or more first semantic regions of the one or more content images; generating, by a decoder based at least in part on the specific content features with the transferred specific color and brightness look, one or more semantic content transferred images corresponding to the one or more content images.

[0253] EEE2. The method as recited in EEE1, wherein the decoder generates the one or more semantic content transferred images by performing: using a second convolutional layer initialized with second pretrained weights of the convolutional network for image recognition to process the specific content features with the transferred specific color and brightness look into second specific content features; generating, by decoder convolutional layers based at least in part on the second specific content features, the one or more semantic content transferred images corresponding to the one or more content images.

[0254] EEE3. The method as recited in EEE1 or EEE2, wherein the convolutional network for image recognition is a VGG network.

[0255] EEE4. The method as recited in any of EEE1-EEE3, wherein the semantic style transform is implemented with a whitening and color transform (WCT).

[0256] EEE5. The method as recited in any preceding EEE, wherein the convolutional network for image recognition includes the first convolutional layer which receives the content images and the style images as input and generates relatively low-level image features; wherein the relatively low-level image features serve as input to one or more subsequent convolutional layers of the convolutional network for image recognition to generate relatively deeper image features from the relatively low-level image features; wherein the specific content features and the specific style features are extracted by the first convolutional layer only of the convolutional network for image recognition; wherein the specific content features and the specific style features are free of relatively deep content features and of relatively deep style features extractable by the one or more subsequent convolutional layers of the convolutional network for image recognition.

[0257] EEE6. The method as recited in any preceding EEE, the first convolutional layer is initialized with pretrained weights of the convolutional network for image recognition.

[0258] EEE7. The method as recited in any of EEE1 to EEE5, the first convolutionallayer operates with re-trained weights of the convolutional network for image recognition.

[0259] EEE8. The method as recited in any preceding EEE, further comprising: applying one or more segmentation tools to generate one or more segmentation maps for the one or more content images, respectively, wherein each of the one or more segmentation maps specifies, for each pixel location in a respective content image of the one or more content images, a plurality of first semantic classes along with a plurality of probability values; applying the one or more segmentation tools to segment the one or more style images into the one or more second semantic regions, each second semantic region of the one or more second semantic regions belonging to a second respective semantic class of one or more second semantic classes.

[0260] EEE9. The method as recited in EEE8, further comprising: filtering the specific style features for each second semantic class of the one or more second semantic classes present in the one or more style images to generate respective per-class filtered style features for the second semantic classes; applying the semantic style transform to transfer the respective per-class filtered style features for the second semantic classes to the specific content features based at least in part on the one or more class correspondence relationships to generate a plurality of per-class transformed content features for the plurality of the first semantic classes; generating, based on the plurality of probability values, for the pixel location in the respective content image of the one or more content images, a weighted sum of the plurality of the per-class transformed content features, wherein the weight sum is included in the specific content features with the transferred specific color and brightness look.

[0261] EEE10. The method as recited in EEE8 or EEE9, wherein the one or more content images represent a content video; wherein the one or more segmentation maps for the one or more content images are generated with a learning model by finetuning one or more initial segmentation maps directly created from the one or more segmentation tools; wherein the learning model is trained online with a relatively small set of content images in the content video.

[0262] EEE11. The method as recited in any preceding EEE, further comprising: performing post-processing operations on the one or more semantic content transferred images to generate one or more post-processed semantic content transferred images.

[0263] EEE12. The method as recited in EEE11, wherein the post-processing operations include a guided filtering operation.

[0264] EEE13. The method as recited in any preceding EEE, wherein the one or morefirst semantic regions of the one or more content images are delineated with one or more segmentation masks; wherein the semantic style transform is applied based further on the one or more segmentation masks.

[0265] EEE14. The method as recited in EEE13, wherein the one or more segmentation masks are generated by fine-tuning one or more temporally unstable segmentation masks of the one or more content images with a trained deep learning model; wherein the one or more temporally unstable segmentation masks are generated from the one or more content images by one or more segmentation tools.

[0266] EEE15. The method as recited in any preceding EEE, wherein the one or more content images originate from a content video that includes the one or more content images as a sequence of consecutive images.

[0267] EEE16. The method as recited in any preceding EEE, wherein the one or more style images originate from a style video that includes the one or more style images as a sequence of consecutive images.

[0268] EEE17. The method as recited in any preceding EEE, wherein the one or more style images are independent of one another and do not form a sequence of consecutive images in a style video.

[0269] EEE18. The method as recited in any preceding EEE, wherein the decoder includes two convolutional layers with respective weights initialized to random values and trained offline with a training dataset comprising a population of training content images.

[0270] EEE19. The method as recited in any preceding EEE, wherein the one or more first semantic classes are assigned with one or more first semantic labels, respectively; wherein the one or more second semantic classes are assigned with one or more second semantic labels, respectively.

[0271] EEE20. The method as recited in EEE19, wherein the one or more second semantic classes include at least two second semantic classes from at least two different style images; wherein the one or more second semantic labels include at least two second semantic labels designated to fully distinguish among the at least two different style images; wherein the one or more semantic class correspondence relationships include a semantic class correspondence relationship between one of the one or more first semantic classes and one of the at least two second semantic classes with one of the at least two second semantic labels.

[0272] EEE21. The method as recited in EEE20, wherein the one or more second semantic classes include at least two second semantic classes present in at least two independent style images, respectively; wherein the one or more second semantic labels include at least twosecond semantic labels designated to fully distinguish the at least two independent style images, respectively; wherein the one or more semantic class correspondence relationships identify a correspondence relationship between (a) one of the at least two second semantic classes present in one of the at least two independent style images based on one of the at least two second semantic labels, and (b) one of the one or more first semantic classes present in one of the one or more content images.

[0273] EEE22. The method as recited in any preceding EEE, wherein the one or more content images is composed of a single content image; wherein the one or more style images is composed of a single style image.

[0274] EEE23. The method as recited in any preceding EEE, wherein the one or more content images are decoded from a single content video; wherein the one or more style images is composed of a single style image.

[0275] EEE24. The method as recited in any preceding EEE, wherein the one or more content images are decoded from a single content video; wherein the one or more style images are decoded from a single style video.

[0276] EEE25. The method as recited in any preceding EEE, wherein the one or more content images are decoded from a single content video; wherein the one or more style images include a plurality of independent style images.

[0277] EEE26. The method as recited in any preceding EEE, wherein the first convolutional layer operates with zero or more additional first convolutional layers to extract the specific content features; wherein the second convolutional layer operates with zero or more additional second convolutional layers to generate the second specific content features.

[0278] EEE27. The method as recited in any preceding EEE, wherein the first and second convolutional layers are selected from among the lowest level convolutional layers of the convolutional network for image recognition with publicly accessible pretrained weights.

[0279] EEE28. An apparatus performing any of the methods as recited in EEE1-EEE27.

[0280] EEE29. A non-transitory computer readable medium, storing software instructions, which when executed by one or more processors cause performance of the steps of any of the methods as recited in EEE1-EEE27. EEE30. A computing device comprising one or more processors and one or more storage media, storing a set of instructions, which when executed by one or more processors cause performance of the method recited in any of EEE1-EEE27.

Claims

CLAIMS What is claimed is:

1. A method, comprising: establishing one or more class correspondence relationships between one or more first semantic classes to which one or more first semantic regions of one or more content images belong and one or more second semantic classes to which one or more second semantic regions of one or more style images belong; using a first convolutional layer initialized with a convolutional network for image recognition to extract specific content features in the one or more first semantic regions of one or more content images and to extract specific style features in the one or more second semantic regions of one or more style images; based at least in part on (a) the one or more class correspondence relationships, (b) the specific content features and (c) the specific style features, applying a semantic style transform to transfer a specific color and brightness look associated with the one or more second semantic regions of the one or more style images to the specific content features in the one or more first semantic regions of the one or more content images; using a second convolutional layer initialized with second pretrained weights of the convolutional network for image recognition to process the specific content features with the transferred specific color and brightness look into second specific content features; and generating, by a decoder based at least in part on the second specific content features one or more semantic content transferred images corresponding to the one or more content images.

2. The method as recited in Claim 1, wherein the convolutional network for image recognition is a VGG network.

3. The method as recited in any preceding claim, wherein the semantic style transform is implemented with a whitening and color transform (WCT).

4. The method as recited in any preceding claim, wherein the convolutional network for image recognition includes the first convolutional layer which receives the content images and the style images as input and generates relatively low-level image features; wherein the relatively low-level image features serve as input to one or more subsequent convolutional layers of the convolutional network for image recognition to generate relatively deeper image features from the relatively low-level image features; wherein the specific content features and the specific style features are extracted by the first convolutional layer only of the convolutional network for image recognition; wherein the specific content features and the specific style features are free of relatively deep content features and of relatively deep style features extractable by the one or more subsequent convolutional layers of the convolutional network for image recognition.

5. The method as recited in any preceding claim, wherein the first convolutional layer is initialized with pretrained weights of the convolutional network for image recognition.

6. The method as recited in any of claims 1 to 4, wherein the first convolutional layer operates with re-trained weights of the convolutional network for image recognition.

7. The method as recited in any preceding claim, further comprising: applying one or more segmentation tools to generate one or more segmentation maps for the one or more content images, respectively, wherein each of the one or more segmentation maps specifies, for each pixel location in a respective content image of the one or more content images, a plurality of first semantic classes along with a plurality of probability values; applying the one or more segmentation tools to segment the one or more style images into the one or more second semantic regions, each second semantic region of the one or more second semantic regions belonging to a second respective semantic class of one or more second semantic classes.

8. The method as recited in Claim 7, further comprising: filtering the specific style features for each second semantic class of the one or more second semantic classes present in the one or more style images to generate respective per-class filtered style features for the second semantic classes; applying the semantic style transform to transfer the respective per-class filtered stylefeatures for the second semantic classes to the specific content features based at least in part on the one or more class correspondence relationships to generate a plurality of per-class transformed content features for the plurality of the first semantic classes; generating, based on the plurality of probability values, for the pixel location in the respective content image of the one or more content images, a weighted sum of the plurality of the per-class transformed content features, wherein the weight sum is included in the specific content features with the transferred specific color and brightness look.

9. The method as recited in Claim 7 or 8, wherein the one or more content images represent a content video; wherein the one or more segmentation maps for the one or more content images are generated with a learning model by finetuning one or more initial segmentation maps directly created from the one or more segmentation tools; wherein the learning model is trained online with a relatively small set of content images in the content video.

10. The method as recited in any preceding claim, further comprising: performing post- processing operations on the one or more semantic content transferred images to generate one or more post-processed semantic content transferred images.

11. The method as recited in Claim 10, wherein the post-processing operations include a guided filtering operation.

12. The method as recited in any preceding claim, wherein the one or more first semantic regions of the one or more content images are delineated with one or more segmentation masks; wherein the semantic style transform is applied based further on the one or more segmentation masks.

13. The method as recited in Claim 12, wherein the one or more segmentation masks are generated by fine-tuning one or more temporally unstable segmentation masks of the one or more content images with a trained deep learning model; wherein the one or more temporally unstable segmentation masks are generated from the one or more content images by one or more segmentation tools.

14. The method as recited in any preceding claim, wherein the one or more content images originate from a content video that includes the one or more content images as a sequence of consecutive images.

15. The method as recited in any preceding claim, wherein the one or more style images originate from a style video that includes the one or more style images as a sequence of consecutive images.

16. The method as recited in any preceding claim, wherein the one or more style images are independent of one another and do not form a sequence of consecutive images in a style video.

17. The method as recited in any preceding claim, wherein the decoder includes two convolutional layers with respective weights initialized to random values and trained offline with a training dataset comprising a population of training content images.

18. The method as recited in any preceding claim, wherein the one or more first semantic classes are assigned with one or more first semantic labels, respectively; wherein the one or more second semantic classes are assigned with one or more second semantic labels, respectively.

19. The method as recited in Claim 18, wherein the one or more second semantic classes include at least two second semantic classes from at least two different style images; wherein the one or more second semantic labels include at least two second semantic labels, each second semantic label specific to a respective one of the at least two different style images; wherein the one or more semantic class correspondence relationships include a semantic class correspondence relationship between one of the one or more first semantic classes and one of the at least two second semantic classes with one of the at least two second semantic labels.

20. The method as recited in Claim 19, wherein the one or more second semantic classes include at least two second semantic classes present in at least two independent style images, respectively; wherein the one or more second semantic labels include at leasttwo second semantic labels each second semantic label specific to a respective one of the at least two independent style images; wherein the one or more semantic class correspondence relationships identify a correspondence relationship between (a) one of the at least two second semantic classes present in one of the at least two independent style images based on one of the at least two second semantic labels, and (b) one of the one or more first semantic classes present in one of the one or more content images.

21. The method as recited in any preceding claim, wherein the one or more content images is composed of a single content image; wherein the one or more style images is composed of a single style image.

22. The method as recited in any preceding claim, wherein the one or more content images are decoded from a single content video; wherein the one or more style images is composed of a single style image.

23. The method as recited in any preceding claim, wherein the one or more content images are decoded from a single content video; wherein the one or more style images are decoded from a single style video.

24. The method as recited in any preceding claim, wherein the one or more content images are decoded from a single content video; wherein the one or more style images include a plurality of independent style images.

25. The method as recited in any preceding claim, wherein the first convolutional layer operates with zero or more additional first convolutional layers to extract the specific content features; wherein the second convolutional layer operates with zero or more additional second convolutional layers to generate the second specific content features.

26. The method as recited in any preceding claim, wherein the first and second convolutional layers are selected from among the lowest level convolutional layers of the convolutional network for image recognition with publicly accessible pretrained weights.

27. An apparatus performing any of the methods as recited in claims 1 to 26.

28. A non-transitory computer readable medium storing software instructions, which, when executed by one or more processors cause performance of the steps of any of the methods as recited in claims 1 to 26.

29. A computing device comprising one or more processors and one or more storage media, storing a set of instructions, which when executed by one or more processors cause performance of the method recited in any of claims 1 to 26.

Citation Information

Patent Citations

  • Domain Stylization Using a Neural Network Model

    US20190244060A1

Cited By

  • Game image generation system and method based on artificial intelligence enabling

    CN121366215A