Video-controlled hybrid-style video generation through generative ai and gaussian splatting

The hybrid-style video generation using generative AI and Gaussian splatting addresses the challenges of controlling scene and object styles in video generation, achieving high-quality, spatial-temporal consistent results for various industries.

WO2026089974A1PCT designated stage Publication Date: 2026-04-30FUTUREWEI TECHNOLOGIES INC
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
FUTUREWEI TECHNOLOGIES INC
Filing Date
2025-10-16
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing video generation technologies face challenges in accurately controlling scene composition, scene motion, visual details, and achieving spatial-temporal consistency in generated videos, particularly in text-to-video generation engines like Sora and CustomVideo, which struggle to describe detailed aspects of image content and maintain geometric accuracy.

Method used

A hybrid-style video generation method using generative AI and Gaussian splatting, where a guiding video controls semantic scene content and image examples control object styles, employing a multi-style diffusion model and 3D Gaussian splatting for real-time, photorealistic rendering from arbitrary viewpoints with spatial-temporal consistency.

Benefits of technology

Enables high-quality, controlled video generation with precise object styling and scene integrity, supporting applications in entertainment, fashion, gaming, augmented reality, and marketing by allowing real-time, photorealistic rendering and seamless integration of personalized content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025051297_30042026_PF_FP_ABST
    Figure US2025051297_30042026_PF_FP_ABST
Patent Text Reader

Abstract

A method comprising receiving a guiding video, a set of image examples, a set of region masks associated with the image examples; generating a set of hybrid inputs based on the guiding video, the set of image examples, and the set of region masks; generating a set of multi-style latent features based on the hybrid inputs using a multi-style diffusion model; generating, using a restoration model, a sequence of multi-style output frames based on the multi-style latent features and the hybrid inputs; constructing, using a semantic-aware multi-style three-dimensional (3D) Gaussian splat (GS) model, a multi-style Gaussian model representing a 3D scene based on the sequence of multi-style output frames, the set of region masks, the camera parameters, and the camera views; rendering a synthesized image of the 3D scene from a target camera view using the multi-style Gaussian model; and generating a hybrid-style video from the synthesized image.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Docket No. 4502-84201 (6000729PCT02)Video-Controlled Hybrid-Style Video Generation through Generative Al and Gaussian SplattingCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U. S. Provisional Application No. 63 / 711,535 filed on October 24, 2024, which is hereby incorporated by reference.TECHNICAL FIELD

[0002] The present disclosure describes techniques for generating video content. More specifically, this disclosure describes techniques for generating video content using generative artificial intelligence (GAI) and Gaussian splatting (GS).BACKGROUND

[0003] Generative Artificial Intelligence (Al) is a form of Al that learns patterns from massive datasets and then uses this knowledge to create new, original content like text, images, audio, video, and even code. Powered by deep learning models, generative Al works by encoding information into a mathematical representation and then decoding it to generate novel outputs when prompted by a user.SUMMARY

[0004] The disclosed embodiments provide techniques for generating a hybrid-style free-view video with controls over both the semantic scene content and desired styles of various objects. In an embodiment, the semantic scene content is controlled using a guiding video input, and the desired styles of specific objects are controlled by image examples of the objects. In an embodiment, GS is employed as an efficient scene representation and rendering technique to improve the quality and performance of free-view video synthesis, enabling real-time, photorealistic rendering from arbitrary viewpoints while preserving spatial-temporal consistency.

[0005] A first aspect relates to a method implemented by a computing device, comprising: receiving a guiding video, a set of image examples, and a set of region masks associated with the image examples, wherein the guiding video comprises a sequence of image frames, and wherein each image frame comprises one or more target objects; generating a set of hybrid inputs based on the guiding video, the set of image examples, and the set of region masks; generating a set of multiAtty. Docket No. 4502-84201 (6000729PCT02)style latent features based on the hybrid inputs using a multi-style diffusion model; generating, using a restoration model, a sequence of multi-style output frames based on the set of multi-style latent features and the set of hybrid inputs; obtaining camera parameters and camera views corresponding to the guiding video; constructing, using a semantic-aware multi-style three-dimensional (3D) Gaussian splat (GS) model, a multi-style Gaussian model representing a 3D scene based on the sequence of multi-style output frames, the set of region masks, the camera parameters, and the camera views; rendering a synthesized image of the 3D scene from a target camera view using the multi-style Gaussian model; generating a hybrid-style video from the synthesized image, wherein the hybrid-style video maintains an original style of the guiding video for non-target regions and applies visual styles from the set of image examples to the one or more target objects with spatial and temporal consistency; and displaying the hybrid-style video on a display device.

[0006] Optionally, in any of the preceding aspects, another implementation of the aspect provides receiving one or more textual descriptions associated with the target objects; and generating the set of multi-style latent features based on the hybrid inputs and the textual descriptions using the multi-style diffusion model.

[0007] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the multi-style Gaussian model comprises a set of 3D Gaussians representing the 3D scene.

[0008] Optionally, in any of the preceding aspects, another implementation of the aspect provides that each 3D Gaussian comprises a position vector, a scale vector, a rotation quaternion, an opacity value, a color vector, or a semantic mask vector for splatting-based rendering.

[0009] Optionally, in any of the preceding aspects, another implementation of the aspect provides rendering the synthesized image comprises performing point-based alpha by projecting each 3D Gaussian into a 2D Gaussian in the target camera view.

[0010] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the target camera view is a novel view not present in the guiding video.

[0011] Optionally, in any of the preceding aspects, another implementation of the aspect provides generating a semantic-aware 3D GS model based on the guiding video to stabilize the multi-style Gaussian model.

[0012] Optionally, in any of the preceding aspects, another implementation of the aspect provides generating the semantic-aware 3D GS model comprises: receiving the guiding video, theAtty. Docket No. 4502-84201 (6000729PCT02)region masks, and the camera views; generating an initial 3D GS model using a Gaussian splatting method; and assigning semantic labels to Gaussians in the initial 3D GS model using an inverse semantic projection module to generate the semantic-aware 3D GS model.

[0013] Optionally, in any of the preceding aspects, another implementation of the aspect provides updating the semantic-aware 3D GS model by refining the Gaussians and reapplying the inverse semantic projection.

[0014] Optionally, in any of the preceding aspects, another implementation of the aspect provides generating the semantic-aware multi-style 3D GS model by receiving the multi-style output frames, the region masks, the camera views, the camera parameters; generating an initial multi-style 3D GS model using Gaussian splatting; and assigning semantic labels to multi-style Gaussians in the initial multi-style 3D GS model using an inverse semantic projection module to generate the semantic-aware multi-style 3D GS model.

[0015] Optionally, in any of the preceding aspects, another implementation of the aspect provides generating the semantic-aware multi-style 3D GS model by: receiving the multi-style output frames, the region masks, the camera views, the camera parameters, and the semantic-aware 3D GS model, wherein the semantic-aware 3D GS model comprising Gaussians with first semantic masks; generating an initial multi-style 3D GS model using Gaussian splatting; and assigning semantic labels to multi-style Gaussians in the initial multi-style 3D GS model using an inverse semantic projection module to generate the semantic-aware multi-style 3D GS model, wherein the semantic-aware multi-style 3D GS model comprising Gaussians with second semantic masks.

[0016] Optionally, in any of the preceding aspects, another implementation of the aspect provides computing a semantic regularization loss based on differences between the first semantic masks and the second sematic masks; computing a multi-style loss for each synthesized image; updating the semantic-aware multi-style GS model based on the semantic regularization loss and the multi-style loss; and backpropagating the multi-style loss to update parameters of the multi-style diffusion module and the restoration module used to generate the multi-style output images.

[0017] Optionally, in any of the preceding aspects, another implementation of the aspect provides the multi-style diffusion model comprises a multi-style embedding model and / or a reverse diffusion model, and wherein training the multi-style diffusion model comprises: generating a training masked hybrid input based on a training guiding image, a training target image, a training region mask, and a training text description; computing, using the multi-style embedding model, anAtty. Docket No. 4502-84201 (6000729PCT02)embedded training joint latent feature based on the training masked hybrid input; generating, using the reverse diffusion model, a training multi-style latent feature based on the embedded training joint latent feature and a random noise; processing, using the restoration model, the training multi-style latent feature and the training masked hybrid input to obtain a training multi-style output; and updating the multi-style embedding model and / or reverse diffusion model based on the training multi-style output.

[0018] Optionally, in any of the preceding aspects, another implementation of the aspect provides receiving, for each of one or more target objects in the 3D scene, the image examples illustrating a desired appearance style, a textual description of the desired appearance style, and the synthesized image; and computing, using a multi-style editing module, an edited output image by modifying the synthesized image to match the desired appearance style for each target object.

[0019] Optionally, in any of the preceding aspects, another implementation of the aspect provides generating the edited output image comprises identifying, for each target object, a best matching image example from the image examples, wherein the best matching image example has a viewpoint most similar to that of the synthesized image; computing, using an alignment module, an aligned synthesized image based on the viewpoint; inverting, using a multi-style inversion module, the aligned synthesized image into a Gaussian noise representation based on 3D structure information of a target object and the best matching image example; and generating, using a reverse diffusion module, the edited output image based on the Gaussian noise representation, the best matching image example, the synthesized image, and the 3D structure information of the target object.

[0020] Optionally, in any of the preceding aspects, another implementation of the aspect provides that each image frame comprises one or more of a grayscale image, a color image, or a color image with associated depth information, wherein each image example represents a desired visual style of a corresponding target object, and wherein each region mask identifies spatial regions corresponding to one or more target objects in the image frames of the guiding video.

[0021] Optionally, in any of the preceding aspects, another implementation of the aspect provides that generating the hybrid inputs comprises selecting, for each target object and each image frame, a matching image example from the set of image examples based on viewpoint similarity; and generating the hybrid inputs for each target object and each image frame by masking a targetAtty. Docket No. 4502-84201 (6000729PCT02)object region in the image frame and concatenating the masked frame with the selected matching image example.

[0022] Optionally, in any of the preceding aspects, another implementation of the aspect provides that obtaining the camera views comprises estimating the camera views of the 3D scene using a structure-from-motion algorithm.

[0023] Optionally, in any of the preceding aspects, another implementation of the aspect provides receiving metadata associated with the guiding video, wherein the metadata includes the camera views for each image frame of the guiding video; and obtaining the camera views from the metadata.

[0024] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the multi-style diffusion model and the restoration model are trained based on a set of multi-style image pairs with associated masks.

[0025] A second aspect relates to a computing device, comprising: a memory or storage means configured to store instructions; and one or more processors or processing means coupled to the memory or the storage means and configured to execute the instructions to cause the computing device to perform the method in any of the disclosed embodiments.

[0026] A third aspect relates to a computer program product comprising computer-executable instructions stored on a non-transitory computer-readable storage medium, the computerexecutable instructions when executed by one or more processors of a computing device, cause the computing device to perform the method in any of the disclosed embodiments.

[0027] For the purpose of clarity, any one of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create a new embodiment within the scope of the present disclosure.

[0028] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims.Atty. Docket No. 4502-84201 (6000729PCT02)BRIEF DESCRIPTION OF THE DRAWINGS

[0029] For a more complete understanding of this disclosure, reference is now made to the following brief description, taken in connection with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0030] FIG. 1 is a schematic diagram of an overall workflow for generating a hybrid-style free-view video using GAI and GS according to an embodiment of the disclosure.

[0031] FIG. 2 is a schematic diagram of an overall workflow of semantic-aware 3D Gaussian splat module according to an embodiment of the disclosure.

[0032] FIG. 3 is a schematic diagram of an overall workflow of a semantic-aware multi-style 3D Gaussian splat module according to an embodiment of the disclosure.

[0033] FIG. 4 is a schematic diagram of a learning method for a multi-style diffusion module and a restoration module according to an embodiment of the disclosure.

[0034] FIG. 5 is a schematic diagram of a workflow of joint model tuning according to an embodiment of the disclosure.

[0035] FIG. 6 is a schematic diagram of an overall workflow for generating a hybrid-style free-view video using GS according to an embodiment of the disclosure.

[0036] FIG. 7 is a schematic diagram of an overall workflow of multi-style editing module according to an embodiment of the disclosure.

[0037] FIG. 8 is a method implemented by a computing device according to an embodiment of the disclosure.

[0038] FIG. 9 is a schematic diagram of a network apparatus according to an embodiment of the disclosure.DETAILED DESCRIPTION

[0039] It should be understood at the outset that although an illustrative implementation of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or in existence. The disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims along with their full scope of equivalents.Atty. Docket No. 4502-84201 (6000729PCT02)

[0040] Generating video content through GAI remains one of the most complex challenges in the field of content creation. In particular, controlled video generation that produces desired visual content presents significant technical difficulties. Existing text-to-video generation engines, such as Sora or CustomVideo, face the innate difficulty of using text prompts to accurately describe the desired generation target, such as the scene composition, scene motion, visual details, and related elements. In addition, how to achieve spatial-temporal consistency of both geometry and texture in the generated video is an open challenge. In addition, how to achieve spatial-temporal consistency of both geometry and texture in the generated video is an open challenge.

[0041] Disclosed herein are various systems and methods for generating hybrid-style videos based on a guiding video and image examples of desired styles of target objects in the generated video with associated masks. The guiding video serves as the basis to define the dynamic semantic development in the video, such as scene composition, scene content, scene movements, etc. For each target object in the guiding video with a desired presenting style in the generated video, the exemplar images with associated masks provides detailed visual description about the desired appearance of the object in the generated video. Different target objects can have different desired presenting styles, and the final generated video integrates these different styles into the guiding video while following the semantic and scene development defined by the guiding video.

[0042] Furthermore, the proposed system framework supports two types of applications. The first type has a training process where given the guiding video and image examples of desired styles of target objects in the generated video with associated masks, a 3D representation in the form of multi-style 3D Gaussians is trained to model the 3D scene of the target multi-style video, and novel video frames where target objects follows target styles of the target image examples can be generated from novel viewpoints that can be the same as or different from viewpoints of the guiding video. The second type is a zero-shot application. By using the learned 3D representation of the first type, given new image examples of new desired styles of target objects, novel video frames can be generated where target objects follow new target styles from novel viewpoints that can be the same as or different from viewpoints of the guiding video.

[0043] Great success has been achieved for Al Generated Content (AIGC) by using a wide range of image generative models, including generative adversarial networks (GAN), diffusion models, and auto-regressive (AR) models. The goal is to enable fast and accessible high-quality content creation. Various methods have been developed to allow for efficient manipulation of the generatedAtty. Docket No. 4502-84201 (6000729PCT02)content using different types of prompt inputs, such as using text descriptions and / or spatial / spatiotemporal compositions like sketches or segmentation maps. However, it is innately difficult to use prompt inputs to accurately describe detailed aspects of the image content, such as size, shape, color, location of various objects, scene composition, etc. It is difficult to control the generated output to have any specific desired attributes or visual details.

[0044] Image-to-image transformation has been largely used for transferring image styles, such as the StyleGAN method. Given an input image, a style generator generates a synthesized output image with a desired style. The style generator is trained to encode the style information from a set of example images, and the style encoded latent feature is injected into the generation model to combine with the embedding feature of the input, the style-transferred output usually preserves the high-level semantic attributes of the input image. However, it is in general difficult to control the properties of generated images, which requires both disentangling features or properties in images and adding controls over these properties to the style generator. As a result, existing GAN-based image-to-image style transfer methods are applied to two scenarios: transferring the global style of the image, e.g., changing from photos to oil paintings; or transferring specific features of an object category with well-defined high-level attributes like human faces, e.g., change skin color, makeup, or hair style.

[0045] Recent works use pretrained text-to-image diffusion models like Imagen to create variations of images, or to manipulate specific image regions. Text-guided image-to-image transformation gives large flexibility to generate novel image content. However, same as general text-to-image generation, due to the lack of description power of using text to define visual content, it is innately difficult to control the generated output to have any specific desired visual details.

[0046] Video generation is one of the most difficult problems among all GAI-based content generation problems. Controlled video generation that produces desired visual content is even harder. Existing text-to-video generation engines like Sora or CustomVideo face the innate difficulty of using text prompts to accurately describe the desired generation target, such as the scene composition, scene motion, visual details and so on.

[0047] Instead of using text descriptions, using an input video as a guiding example to accurately illustrate and define scene composition, semantic content, scene motion, etc. is much more efficient. Also, using image examples to illustrate and define the details of objects in the generated video enables the system to control the generation results.Atty. Docket No. 4502-84201 (6000729PCT02)

[0048] In addition, howto achieve spatial-temporal consistency of both geometry and texture in the generated video is an open challenge. It is innately difficult to use text descriptions to adequately describe the desired spatial-temporal consist properties the generated video, especially in terms of geometric accuracy. Using an input video as a guiding example, accurate 3D information of the video scene can be computed, which functions as a proxy to largely alleviate the problem of generating spatial-temporal consist content.

[0049] The disclosed embodiments provide techniques for generating a hybrid-style free-view video with controls over both the semantic scene content and desired styles of various objects. In an embodiment, the semantic scene content is controlled using a guiding video input, and the desired styles of specific objects are controlled by image examples of the objects. In an embodiment, GS is employed as an efficient scene representation and rendering technique to improve the quality and performance of free-view video synthesis, enabling real-time, photorealistic rendering from arbitrary viewpoints while preserving spatial-temporal consistency.

[0050] The disclosure is related to generating video content using GAI and GS. The disclosed approach introduces a system framework that takes as input a guiding video, one or multiple semantic masks indicating the target objects in one or multiple frames in the guiding video, and one or multiple exemplar object images corresponding to each semantic mask, and optionally one or multiple text descriptions corresponding to each semantic mask, and then generates as output a hybrid-style video where each masked object in the video is altered to follow the style of the corresponding exemplar object images in a spatial-temporal consistent way, and the remaining part of the generated video maintains the original style as the guiding video. The proposed system provides the capability for controlled generation of video content where the scene composition and semantic content is faithful to the guiding video, while spatial-temporal-consistently changing the object of interests in the video into different desired styles that are explicitly described by image examples (and optional text descriptions).

[0051] The disclosed techniques have a wide range of applications across industries requiring precise, high-fidelity video content generation. For example, in entertainment and media sector, the framework embodied by the disclosed techniques can be used for efficient post-production editing by restyling characters or objects without reshooting or reanimation. For example, in fashion and retail, the framework supports virtual try-on experiences by replacing garments or accessories in video with styled alternatives. In gaming, the framework allows dynamic asset customization andAtty. Docket No. 4502-84201 (6000729PCT02)cutscene generation with varied visual themes. Additionally, the disclosed techniques can be employed in augmented and virtual reality environments for seamless integration of personalized content into existing scenes. In marketing and social media, it supports the creation of engaging, brand-aligned video content tailored to target audiences. The ability to selectively stylize video elements while preserving scene integrity makes the framework a powerful tool for controlled, creative video manipulation.

[0052] FIG. 1 is a schematic diagram of an overall workflow 100 (a.k.a., framework) for generating a hybrid-style free-view video using GAI and GS according to an embodiment of the disclosure. In an embodiment, the overall workflow 100 is for application 1, which involves training a 3D representation for the 3D scene of the target multi-style video. In an embodiment, the overall workflow 100 is implemented by or on a personal computer (PC), a smart phone, a smart tablet, or some other computing device used to play games or consume entertainment.

[0053] As shown in FIG. 1, the overall framework 100 (a.k.a., system) is given an input comprising (1 ) a guiding video represented by a sequence of T image frames Xlt..., XT, (2) a set of image examples Y^, ••• Y^, ■■■, Y^, --- Y^k, and (3) a corresponding set of region masks M, ■■■ M?, ■■■, Mi, M? associated with the target objects represented in the image examples and image frames.

[0054] In an embodiment, each image frame X, may be represented in various formats, including but not limited to: 1 -channel (e.g., gray scale image), 3-channel (e.g., red green blue (RGB) color image), and / or 4-channel (e.g., red green blue-depth (RGBD) image with color and associated depth), etc. In an embodiment, camera parameters Ccamof the camera capturing the guiding video input Xi,..., XTis also provided to the system. For example, for pinhole camera, Ccamcan be the fx ()\camera intrinsic matrix Ccam— I 0 fyy0I, where fx, fyare focal length along the two axes of \0 0 1 / the imaging plane, x0, y0are principle point offsets of two axes of the imaging plane, and 5 is the axis skew.

[0055] In an embodiment, the guiding video Xr,..., XTthat captures three-dimensional (3D) scene from different camera views, the corresponding camera views v±,...,vTfor the frames Xi,..., XTare either provided directly (e.g., as metadata associated with the video) or can be computed in a computing view module 102, which employs 3D reconstruction techniques such asAtty. Docket No. 4502-84201 (6000729PCT02)structure-from-motion to recover camera poses based on scene geometry and image correspondences.

[0056] In an embodiment, in the set of image examples, each K, ••• Y^ rij > 1) illustrates the desired presenting style of a target y-th object in the guiding video. In an embodiment, in the region masks, each M, ■■■comprises a set of masks marking a spatial region of the target y-th object in the T frames of the guiding video. The system 100 generates a new video in which the target y-th object is rendered with the appearance style consistent with the corresponding image examples K, Y^ In an embodiment, the image examples K, ••• Y^ have the exemplary appearance of the / -th object looking from different views. In some alternative embodiments, the system may operate with only a single image example per object, corresponding to a single viewpoint. In an embodiment, viewpoints of the target objects in the Image Examples Y^, --- Y^ can be completely different from the viewpoints of the corresponding objects in Xt,..., XT, or partially the same.

[0057] In an embodiment, the system 100 is provided a set of text descriptions / -j1, ••• 1^, •••, / i, / £, each ItJis textual instruction for the y-th object in the / -th frame in the generated video, such as semantic instruction describing the desired properties or semantic characteristics like color, shape, resolution, semantic category, etc. In an embodiment, for each j-th object, can be all the same, i.e., the whole generated video can share a same rough textual instruction for the y-th object, or different, i.e., detailed specific instructions can be given for different frames.

[0058] In an embodiment, the system 100 further comprises an alignment and masking module 102 configured to process the guiding video Xlt..., XT, the image examples fj1, Y^, Y*, ■■■ Y*k, and the region masks M, •••M to produce a set of hybrid inputs’ HT> ’ Hi > In an embodiment, for the target y-th object and each guiding frame X with regions mask Mt], among the n7- image examples Y^, --- Y the alignment and masking module 102 identifies the best matching image example Y(J, whose view point of the target y-th object is the most similar to the view point of the target y-th object in(7. Once the best example is selected, a hybrid input HLJis generated by concatenating two components: (1) a masked guiding frame X, where the pixels in output of XtJ0 MtJcorresponding to the marked region of the target y-thAtty. Docket No. 4502-84201 (6000729PCT02)object in X are zeroed out to remove its original appearance, and (2) the matching image example YLJ, where MtJis the opposite operation. This combination ensures that each hybrid input captures the original scene context without the object, paired with a stylized reference for how the object should appear. This process is applied to every object across all frames, resulting in a comprehensive set of hybrid inputs, • • •, H, • • •, H, • • •, H. In an embodiment, example K can be either provided as a meta data during image capture or can be estimated through 3D reconstruction methods like structure-from-motion based on YJ, •••..

[0059] In an embodiment, the system 100 further comprises a multi-style diffusion module 104 configured to process the set of hybrid inputs Hl, •••, H?, •••, H to generate a set of multi-style latent features z, ■■■, zl, •••,z^,,z. Each z is an embedding latent feature in a learned class-guided multi-style joint embedding space for the object class of the target / -th object. In an embodiment, the text descriptions l, / ,, 11, ••• I are also fed into the multi-style diffusion module 104, so that the multi-style latent features z^, •••,z|, •••, z*,,z* also take into account the textual instruction provided by the l, --- l, ---, ll, --- 1. This embedding space is trained to capture the relationship between object class, appearance style, and viewpoint, thereby encoding complex style variations across different object instances and camera perspectives.

[0060] In an embodiment, the system 100 further comprises a restoration module 106 configured to generate multi-style outputs Xlt•••, XTbased on the multi-style latent features zj, •••, z, •••,z, •••,z and the hybrid inputs H, ■■■, H, ■■■, H^, ■■■, H. Each output image Xtcorresponds to the original input X, with the same resolution and channel dimensions, i.e., ) can have 1-channel (e.g., gray scale image), 3-channel (e.g., RGB color image), or 4-channel (e g. RGBD image with color and associated depth), etc. In an embodiment, the multi-style diffusion module 104 and the restoration module 106 are assumed to be pre-trained in FIG. 1.

[0061] In an embodiment, the system 100 further comprises a semantic-aware multi-style 3D Gaussian splat module 108 configured to compute a multi-style Gaussian model (φx(v, M)) based on the camera Views v1,...,vT, the camera parameters Ccam, the region masks Ml, ■■■Ml, ■■■, Mi, ■■■MT, and the multi-Style outputs, XT. In an embodiment, the semantic-aware multi-style 3D Gaussian splat module 108 learns a set of 3D Gaussians to represent the 3DAtty. Docket No. 4502-84201 (6000729PCT02)scene, each Gaussian associated with properties like color, opacity, and semantic masks for splatting rendering.

[0062] In an embodiment, the system 100 further comprises a multi-style rendering module 110 configured to compute a synthesized frame Xoutbased on a target camera view voutand using the multi-style Gaussian model (pxv,The synthesized frame Xoutis a synthesized image of the 3D scene projected to the imaging plane corresponding to vout. The rendering is done by positioning the 3D Gaussians on the camera planes and rendering color using point-based alpha-blending. In an embodiment, voutcan be one of vlt..., vT, or a novel view.

[0063] In an embodiment, the system 100 further comprises a semantic-aware 3D Gaussian splat module 112 configured to compute a GS model px(v, M) based on the guiding video input X1,..., XT, the camera parameters Ccam, the camera views vlt...,vT, and the region masks, • ■ • M,, M, • • • M. In an embodiment, the GS model (px(v, M) is fed into the semantic-aware 3D Gaussian splat module 112 to improve the stability of the multi-style Gaussian model <px(v, M).

[0064] Semantic-aware 3D Gaussian splat & Semantic-aware multi-style 3D Gaussian splat

[0065] In this disclosure, let 0 denote the set of 3D gaussians representing the 3D scene structure. In general Gaussian splatting, each Gaussian 0Zis defined by a set of parameters: Θi= {pi, si, qi, αi, ci}, where p_i ∈ ℝ3, s_i ∈ ℝ3, q_i ∈ ℝ4are the mean, scale, and rotation quaternion, αi∈ [0,1] is the opacity, and ci∈ ℝ3is the color. During rendering, to render an image with view vout, the 3D Gaussianis transformed into the image coordinates and projected onto the image plane of the target 2D image to obtain a 2D Gaussian Θ2Di,vbased on p^S, qt, and vout. The rendered color of a pixel x in the target 2D image is computed using point-based alpha-blending on each ray:= S"=i qaz0^Out(x) I]j=i(l - ay0^Out(x)), where N is the total number of Gaussians, and 0^out(x) is the Gaussian response of Q^out over the pixel.

[0066] In an embodiment, each Gaussian includes an additional parameter mi= { mi1, ···, mik+1}, i.e., Θi= {pi, si, qi, αi, ci, mi}, where mi7- is the semantic Gaussian mask for and the j-th semantic label, j =+ 1 where k is the total number of semantic objects corresponding to objects marked by the region masks M4, M?, ■■■, M*, ■■■ M?, and one additional semantic label is assigned to the background regions not marked by any masks.Atty. Docket No. 4502-84201 (6000729PCT02)

[0067] Semantic-aware 3D Gaussian splat

[0068] FIG. 2 is a schematic diagram of an overall workflow 200 (a.k.a., framework) of semantic-aware 3D Gaussian splat module according to an embodiment of the disclosure. In FIG. 2, the multi-style diffusion module is considered predefined and remains fixed. Additionally, the restoration module is also assumed to be available. In an embodiment, the restoration module can either remain fixed or be fine-tuned jointly with the semantic-aware multi-style 3D Gaussian splat module, and optionally with the semantic-aware 3D Gaussian splat module.

[0069] As shown in FIG. 2, the overall framework 200 (a.k.a., system) comprises a 3D Gaussian splat module 202 configured to compute an initial 3D GS Model(v) based on guiding video input Xlt..., XTand the camera Views vlt..., vT, and using Gaussian splatting methods.

[0070] In an embodiment, the system 200 further comprises an inverse semantic projection module 204 configured to un-project the 2D semantic labels back to the initial 3D GS Model φxinit(v), where each Gaussian= {pi, si, qi, αi, ci}, i = 1, ···, N based on the region Masks Mf, ■••, M*, ••• and the camera views v1,..., vT. In an embodiment, for each pixel xtG Xtwith semantic label lj masked by MtJ, the inverse semantic projection module 204 tries to locate the 3D Gaussian Θxthat accounts for xt, so that label lj can be associated with Θn. The inverse semantic projection module 204 can use different methods to do so.

[0071] In one embodiment, a depth map Dtis provided in Xt. For example, for the case the input Xtis a 4-channel RGBD image comprising both color and depth. A depth value dtis associated with the pixel xt. Then, the 3D Gaussian whose center location ptis closest to the depth dtcan be located as 0n. In another embodiment, when ground-truth depth information is not provided, the inverse semantic projection module 204 can make use of the corresponding 3D point x3Dof pixel xtin the 3D space. Using the camera parameters Ccamand the camera view rt, the pixel coordinate p(xt) = (u, v) and the 3D coordinates of the corresponding 3D point p(xfD) — (u3D, v3D, w3D) has a fixed relation: p(xt) = K- E - p(x3D), where K is the camera intrinsic matrix obtained from the camera parameters Ccam, and E is the camera extrinsic parameters obtained from the camera view vt. That is, pixel xtcorresponds to a beam of ray in the 3D space. In terms of semantic label lj of xt, it is projected to the surface points in the 3D space. During rendering, the opacity is accumulated from near to far. Let on(xt) denote the accumulated opacity of xtafter adding the / / -th Gaussian,on(xt) = Il}=i(l—“y072vt(xt)), n =Atty. Docket No. 4502-84201 (6000729PCT02)1,..., N, when on(xt) is greater than a cutoff threshold ad, it is assumed to reach a surface point, and this zz-th 3D Gaussian can be located as 0n. adcan be given as a hyperparameter. Θ2Di,vis the projected 2D Gaussian of the 3D Gaussian 0j based on Pt. St, qt,vt, and 0?®t(xt) is the Gaussian response of Θ2Di,vover pixel xt.

[0072] In an embodiment, for each semantic label lj, a total weight 17can be computed for Gaussian to indicate how possible Qt carries the semantic label If.Wij= ΣTt=1βij(Xt, Mtj),where (Xt, MtJ) is a weight computed based on each input Xtand mask MtJ. For example, = Σx∈Xon(xt) · Mtj(xt)whereor 0 if xthas semantic label lj or not. nxis the located id of the 3D Gaussian 0n%tfor pixel xt. In an embodiment, the Gaussian 0, is assigned to the semantic label lj if WtJis greater than a threshold Ws, and entry mi7= 1 in m;. Otherwise,= 0. If mi7- = 0 for all semantic labels lj,j = 1, •••, / <:,mik+i=1 and the Gaussianis considered to belong to background unknown category. The threshold Wscan be given as a hyperparameter.

[0073] In other embodiments, the inverse semantic projection module 204 can use a neural network to predict the semantic label of 3D gaussians. Let T denote the neural network parameters, and using the initial 3D GS Model φxinit(v) as input, Γ(φxinit(v)) outputs m1, ···, mN. The final Gaussians 0;=si, αi, ci, mi}, i = 1, ···, N form the GS model φx(v, M) This disclosure does not put any restrictions on methods or the network models the inverse semantic projection module 204 uses.

[0074] In an embodiment, the system 200 optionally comprises a GS update module 206 configured to iteratively update the initial GS Model φxinit(v), e.g., by performing splitting and pruning, etc., and then perform inverse semantic projection to update the GS Model φx(v, M), so that the original gaussian splat loss can be optimized by considering semantic label un-projection.

[0075] Semantic-aware multi-style 3D gaussian splat

[0076] FIG. 3 is a schematic diagram of an overall workflow 300 (a.k.a., framework) of a semantic-aware multi-style 3D Gaussian splat module according to an embodiment of the disclosure. In FIG. 3, the multi-style diffusion module is considered predefined and remains fixed. Additionally, the restoration module is also assumed to be available. In an embodiment, the restoration moduleAtty. Docket No. 4502-84201 (6000729PCT02)can either remain fixed or be fine-tuned jointly with the semantic-aware multi-style 3D Gaussian splat module, and optionally with the semantic-aware 3D Gaussian splat module.

[0077] As shown in FIG. 3, the semantic-aware multi-style 3D Gaussian splat module configured to compute a multi-style GS ModelM) based on the multi-style outputs Xltthe camera views vlt...,vT, the region masksMj-, •••, M*, M*, the camera parameters Ccam, and optionally the GS Model <px(v, W). Inanembodiment, each 3D Gaussian in the semantic-aware multi-style 3D Gaussian splat module is represented by—{pt, qt, aitct, mi], where is the semantic Gaussian mask forand the j-th semantic label, j = 1, •••, k + 1 where k is the total number of semantic objects corresponding to objects marked by the Region Masks M^, ••• M?,, M*, ■■■ M*, and one additional semantic label is assigned to the background regions not marked by any masks.

[0078] As shown in FIG. 3, the overall framework 300 (a.k.a., system) comprises a 3D Gaussian splat module 302. When the GS Model φx(v, M) is not given as an input, based on the multi-style outputs^,, XTand the camera views vlt...,vT, an initial 3D multi-style GS model <pxl lty) is first built through the 3D Gaussian splat module 302 using Gaussian splatting methods, where each GaussianΘ̃i= {p̃i, s̃i, q̃i, ãi, c̃i}, i =

[0079] In an embodiment, the system 300 further comprises an inverse semantic projection module 304. In an embodiment, based on the region masks M*,and the camera views v,...,vT, the inverse semantic projection module 304 is configured to un-project the 2D semantic labels back to the initial 3D multi-style GS Model <plt(v, and computes as output a multi-style GS Model <px(v, M), where each Gaussian 0;= {pt, s^, qbaitCt, fki], i = 1,, N. The process of the inverse semantic projection module 304 is the same as the process of the inverse semantic projection module 204, with the difference that the input initial 3D GS Model (pxl lty) is replaced with the initial 3D Multi-Style GS Model

[0080] In an embodiment, the system 300 further comprises an optional GS update module 306 configured to iteratively update the initial 3D multi-style GS Model(v), e.g., by performing splitting and pruning, etc., and then perform inverse semantic projection to update the multi-style GS Model φ̃x(v, M), so that the original Gaussian Splat loss can be optimized by considering semantic label un-projection.Atty. Docket No. 4502-84201 (6000729PCT02)

[0081] In an embodiment, when the GS Model <px(v, M) is given as an input, <px(v, M) can be used to initiateand to compute a semantic regularization loss L(φx(v, M), φ̃x(v, M)) to improve the cross-view consistency of the multi-style GS Model (px.v, M). In one embodiment, for each Gaussians 0j ={pi, Si, qiti = 1, •■•, N of φx(v, M), p^s^qi can be used to initialize the Gaussian 0;= {pt, st, cp, aitc, i = 1, •■•,7V, where at, ctcan be learned by 3D Gaussian splat module 302. Then, L px(, M),φx(v, M)) can be jointly optimized in the GS update module 306 to regularize the multi-style GS Model px(v, M). L(<px(v, M), <px(v, M)) measures the semantic difference between the semantic Gaussian Masks m1, ■■■, mNin φx(v, M) and the updated semantic Gaussian Masks m1, ■■■,mNin <px(v, M), e.g., L(<px(v, M),<px(v, M)) = mJ, where Lfm^mJ measures the difference between and mt, such as cross-entropy loss. This disclosure does not put restrictions on the type of loss, and how the GS update module optimizes various types of loss.

[0082] In an embodiment, using the multi-style GS model φx(v, M), given a target camera View vout, the multi-style rendering module renders the synthesized Xout. The rendered color of a pixel x G Xoutis computed using point-based alpha-blending on each ray:= ΣNi=1c̃iãiΘ̃2Di,v(x) Πi-1j=1(1 - ãjΘ̃2Dj,v(x)), (1) where N is the total number of Gaussians, 0^out(x) is the Gaussian response of Q ^out over the pixel,and 0^outis the 2D Gaussian obtained by projecting the 3D Gaussianonto the image plane of the target Camera View voutbased on pt, §i, cp, and vout.

[0083] Multi- Style Diffusion & Restoration

[0084] FIG. 4 is a schematic diagram of an overall workflow 400 (a.k.a., framework or system) of a learning method to obtain a multi-style diffusion module 402 and a restoration module 404 according to an embodiment of the disclosure. In an embodiment, multi-style diffusion module 402 configured to compute a set of multi-style latent features, • • •, z*, • • •, z, • • •, z based on the hybrid inputs Hl, ■■■, Hr, ■■■, H*, •■•. H* and optionally the text descriptions / *, ••• / ^, •••,, •■• / (as shown in FIG. 1). In an embodiment, the multi-style diffusion module 402 may comprise two main processing modules, a multi-style embedding module 406 and a reverse diffusion module 408. The restoration module 404 computes the multi-style output Xlt•■•, XTbased on the multi-style latent features z*, •••,z, •■■,z, •■•, z and the hybrid inputs H, ■••, H?, ■••, H in FIG. 1.Atty. Docket No. 4502-84201 (6000729PCT02)

[0085] In an embodiment, the multi-style diffusion module 402 and the restoration module 404 are trained based on a large set of multi-style image pairs with associated masks. Further, the system is given a training guiding image Input Xtr, a training target image Ytr, a training region mask Mtr, and a training text description Itr. The masked regions indicated by the region mask Mtrof the training target image Ytrand the training guiding image input Xtrare the same object with different appearance styles. Xtrand Ytrhas the same channel as the guiding video input Xlt, XTto which the multi-style diffusion module 402 will be applied. The text description Itrgives textual instructions of how the target style transfer looks like, e.g., change the polar bear to a brown bear. The masked objects in Xtrand Ytrare usually the same object (i.e., the same geometry and structure layout) where only the styles (e.g., color, texture, visual styles like oil painting versus photorealistic, simplified textures like sketches versus fine-grained texture, etc.) are different.

[0086] In an embodiment, the system 400 further comprises a masking module 410. Based on the training guiding image input Xtr, the training target image Ytr, and the raining region mask Mtr, the masking module 410 computes a training masked hybrid input Htr, which is a concatenated tensor of a masked guiding input, an unmasked guiding input Xtr® Mtr, an unmasked guiding input Xtr® Mtr(where pixels not marked by the mask are set to zeros), and the target input Ytr.

[0087] In an embodiment, the multi-style embedding module 406 configured to compute an embedded training joint latent ztrusing Htrand the text descriptions Itr. Latent ztris a latent feature representation in a multi-style embedding space for the object with the style of Xtrand the style of Ytr, and description ltrof this style transfer. Various neural networks can be used for the multi-style embedding module 406. For example, the image and text encoders of a vision-language model (VLM) like CLIP can be used to compute ztr, e.g., by encoding Ytr, Xtr® Mtrand Xtr® Mtrusing the image encoder and Itrusing the text encoder, and concatenate the encoded feature into ztr.

[0088] In an embodiment, the reverse diffusion module 408 generates a training multi-style latent feature ztrbased on the training joint latent ztrand a random noise n (e.g., a gaussian random noise), using a reverse diffusion method. After that, the restoration module 404 computes a training multi-style output Xtrbased on training multi-style latent feature ztrand the training masked hybrid input Htr. Various neural networks can be used for the reverse diffusion moduleAtty. Docket No. 4502-84201 (6000729PCT02)408 such as convolutional diffusion networks or diffusion transformers. Various neural networks can be used for the restoration module 404, such as the decoder. The disclosure does not put any restrictions on how to compute the network structure of the multi-style embedding module 406, the reverse diffusion module 408 or the restoration module 404, or how the latent feature ztris constructed, e.g., by concatenating or modulating different types of features.

[0089] In an embodiment, the system 400 further comprises a compute multi-style loss module 412 and a backpropagation module 414. In an embodiment, the compute multi-style loss module 412 is configured to compute a multi-style loss L(Xtr, Ytr, Mtr, X̃tr) using the training guiding image input Xtr, the training target image Ytr, the training region mask Mtrand the training multistyle output Xtrand the gradient of L Xtr, Ytr, Mtr, Xtr) is calculated and backpropagated to update the weight parameters of the multi-style diffusion module 402 and the restoration module 404 by a backpropagation module 414. There are many ways to compute the multi-style loss L(Xtr, Ytr, Mtr, Xtr. In the preferred embodiment, L Xtr, Ytr, Mtr, Xtr^ is a weighted combination of several losses including (1) the pixel-level distortion between the unmasked guiding input XtrMtrand the unmasked multi-style output XtrMtr, Lpixel(X̃tr⊗ M̄tr, X̃tr⊗ M̄tr), such as the LI or L2-norm between Xtr0 Mtrand Xtr0 Mtr, to enforce the regions not indicated by the targeted object mask in the generated Xtrto match the original style of Xtr(2) the pixel-level distortion between the masked target output Ytr0 Mtrand the masked multi-style output Xtr0 Mtr, Lpixel(Ytr⊗ Mtr, X̃tr⊗ Mtr), such as the LI or L2-norm between Xtr0 Mtrand Ytr0 Mtr, to enforce the target object in the generated Xtrto match the style of the target output Ytr. Other loss functions can also be added here, such as the GAN loss that promotes the naturality of the generated output Xtr.

[0090] In an embodiment, the update of the multi-style diffusion module and the restoration module can be performed together or at different frequencies. Also, the weight parameters of both modules can be updated all together or part by part. This disclosure does not put any restrictions on the methods for the model update.

[0091] In the preferred embodiment, the large training dataset to learn the multi-style diffusion module and the restoration module comprises multi-style image pairs with associated masks for various object categories, which includes the target object categories that the trained model will be applied to in the inference stage. Also, for each object category, the training multi-style imageAtty. Docket No. 4502-84201 (6000729PCT02)pairs with associated masks have various paired image styles, which includes the guiding and target styles of the object that the trained model will be applied to in the inference stage. Also, one set of weight parameters can be trained for the multi-style diffusion module for each object category, or for a group of object categories. Similarly, one set of weight parameters can be trained for the restoration module for each object category, or for a group of object categories.

[0092] Inference with the multi-style diffusion module & the restoration module

[0093] In an embodiment, during the inference stage, the multi-style diffusion module 402 and the restoration module 404 generate the multi-style outputs Xlt, XTbased on the hybrid inputs, H^,, W. Each output framed has multiple objects, each rendered in a desired target style indicated by the image examples Y,Y*, •■• Y^k. Furthermore, the / -th target object in corresponding to X7follows the style defined by Y^.-’Y^. In an embodiment, the desired styles are iteratively merged into the guiding video input Xlt..., XT.

[0094] Furthermore, letz1denote a previous restored version of the multi-style output X(after j-1 iterations. In the / -th iteration, using-1to replace the original guiding video frame X;, the hybrid input Ht]targeting at the / -th object can be computed as the concatenation of the masked replaced guiding frame-1® MLJ, the unmasked replaced guiding frame-1® MtJ, and the matching image example K whose view point of the target / -th object is the most similar to the view point of the target / -th object in X among the n,- image examples Fj7, ••• Y^ When the text description 1 for the j-th object is not given, a predefined text is assigned to / to be used as input to the multi-style embedding module. Then the multi-style embedding module computes an embedded Joint Latent z using HLJand / . After that, the reverse diffusion module 408 generates the multi-style latent feature z7based on the joint latent ztrand a random noise n (e.g., a gaussian random noise). The restoration module 404 finally computes the restored multi-style output X based on the multi-style latent feature z7and the masked hybrid input Ht].

[0095] FIG. 5 is a schematic diagram of a workflow 500 of joint model tuning according to an embodiment of the disclosure. In an embodiment, the multi-style GS Modelcan be tuned in the end-to-end fashion with the workflow described in FIG. 5 to improve the spatial-temporalAtty. Docket No. 4502-84201 (6000729PCT02)consistency of its synthesized images. Optionally, the multi-style diffusion module and the restoration module can be also finetuned in this process.

[0096] Given the multi-style outputs Xr,, XT, the camera ViewsvT, the region masks M*, ■■■My, ■■■, Mi, ■■■ MT, the camera parameters Ccam, and optionally the GS Model φx(v, M)~, the semantic-aware multi-style 3D Gaussian splat module computes the multi-style GS model <px(v, M using the process described by FIG. 3. Then using the multi-style GS model (px.v, M), based on the camera views v,..., vT, the multi-style rendering module renders the synthesized X, ■■■, XTbased on Equation (1). After that, the compute multi-style loss module computes a set of multi-style loss L(XhYi, M, X^, l = 1 — T,j = l, —,k, based on the synthesized Xlt, XT, the guiding video input Xi,, XT, the region masks M*, ■■■ M^,, Mi, ■■■MT, and the image examples Yi, Y, ■■■, Yi, Yk- The gradient of i XL, Y, M{,is calculated, which is fed into the GS update module to combine with L(<px(v, M),^x(v, M)) to iteratively update the multi-style GS model (φx(v, M) Optionally, the gradient of L(Xi, YtJ. M^X^) can also be backpropagated to update the model parameters of the multi-style diffusion module and the restoration module.

[0097] In an embodiment, L^X^ Y^, MtJ, X is a weighted combination of several losses including (1) the pixel-level distortion between the unmasked guiding input A; ® U;=1MtJand the unmasked synthesized XL® Uy=1M7, Lpixel(Xi ® Uy=1M7, Xt® Uy=1M7), such as the LI or L2-norm between Xt® Uy=1M7and Xt® Uy=1M7, to enforce the regions not indicated by any targeted object mask in the generated XLto match the original style of X where U() as the union operation; (2) the pixel-level distortion between the masked target output YtJ® M7and the masked synthesized output XL0 M., Lpixei(Yi MtJ, Xi 0 Mt], such as the LI or L2-norm between F;0 M and XL0 M{, to enforce the target object in the generated Xtto match the style of the target output YtJ. Other loss functions can also be added here, such as the GAN loss that promotes the naturality of the generated output Xtr.

[0098] FIG. 6 is a schematic diagram of an overall workflow 600 for generating a hybrid-style free-view video using GS according to an embodiment of the disclosure. FIG. 6 gives the overall workflow of the framework for application 2, which make uses of the 3D representation already learned for the 3D scene of the target multi-style video.Atty. Docket No. 4502-84201 (6000729PCT02)

[0099] As shown in FIG. 6, the overall framework 600 (a.k.a., system) comprises multi-style rendering module 602 configured to compute a synthesized output Xoutbased on a novel camera view v°ut, and by using the multi-style GS model px(v, M~) with Equation (1). At the same time, a set of image examples Y^, Y^, Yk, ••• Ykkare provided to the system. Each Y^, Y^ (nj > 1) consists of a set of image examples illustrating the desired presenting style of a target / -th object in the guiding input. Optionally, the system is also provided with a set of text descriptions I1, ■■■, Ik, each 77is a textual instruction for the / -th object in the desired output.

[0100] The system 600 further comprises a multi-style editing module 604 configured to compute the final output X° ttbased on the synthesized output Xout, the multi-style GS model <px(v, M), the image examples Y^, ■■■ Y^, Yk, ■■■ Ykk, the camera view vout, and optionally the text descriptions I1, ■■ ■, Ik. In the final output X°^t, the / -th target object follows the style given by image examples Y^, and the remaining region not marked by any target objects follows theoriginal style of X°ut.

[0101] In an embodiment, the object styles presented by image examples E, ••• Y^ can be the same as or different from the image examples in application type 1 (as shown in FIG. 1), where the multi-style GS model (px, M was obtained. Also, given two novel camera views v°ut, v2ut, the generated output X°^t l, X° it 2will maintain spatial-temporal consistency in terms of 3D scene structure.

[0102] FIG. 7 is a schematic diagram of an overall workflow 700 of multi-style editing module according to an embodiment of the disclosure. As shown in FIG. 7, the overall framework 700 (a.k.a., system) comprises an alignment module 702. The multi-style GS model (φx(v, M)) consists of N 3D Gaussians, each Gaussian Qt=Si, qt, at, ct, mt}, i = 1,, N. For each / -th target object, the 3D structure information D, e.g. depth information, can be computed from <px(v, M) based on the corresponding Gaussians whose semantic mask 7Hiy=l for Qt. For the target j-th object, among the rij Image Examples E, Y^ the best matching Image Example E7is found through the alignment module 702, whose viewpoint of the target j-th object is the most similar to the view point of the target / -th object in Xout.Atty. Docket No. 4502-84201 (6000729PCT02)

[0103] In an embodiment, in the alignment module 702, using the view point of K7, a corresponding aligned synthesized image Xall3nis computed using the multi-style GS model (Px(.v, M with Equation (1).

[0104] The system 700 further comprises a multi-style inversion module 704 configured to iteratively invert either the aligned xall9nor a latent feature zoutcomputed from the aligned xaligninto a corresponding Gaussian noise Xnoiseor znoise, by using denoising diffusion probabilistic models or denoising diffusion implicit models, guided by D and the best matching image example YJ. For example, the ControlNet model can be used as the multi-style inversion module.

[0105] The system 700 further comprises a reverse diffusion module 706. After that, using the Gaussian noise Xnoiseor znoise, the 3D structure information Dy, the best matching image example F7, and the synthesized output Xout, the reverse diffusion module 706 computes the final output ^edit The reverse diffusion module 706 uses the iterative reverse diffusion process of the corresponding inversion process in the multi-style inversion module 704, such as the reverse diffusion process of denoising diffusion probabilistic models or the reverse diffusion process of the denoising diffusion implicit models. The neural networks of the multi-style inversion and reverse diffusion is usually pre-trained, and finetuned in the iterative inversion and reverse diffusion process.

[0106] In an embodiment, optionally, when the text descriptions I1, •••, Ikare provided, it can be used by both the multi-style inversion module 704 and the reverse diffusion module 706 as additional controlling guidance, e.g., by converting into a controlling latent code via a VLM.

[0107] FIG. 8 is a method 800 implemented by a computing device according to an embodiment of the disclosure. In an embodiment, the computing device is a computer, a smart phone, a smart tablet, or other device configured to play games or display video content. In an embodiment, the method 800 is implemented during gaming or when video content is being consumed by a user.

[0108] In block 802, the computing device receives one or more of a guiding video, a set of image examples, a set of region masks associated with the image examples. In an embodiment, the guiding video comprises a sequence of image frames. In an embodiment, each image frame comprises one or more target objects.

[0109] In an embodiment, each image frame comprises one or more of a grayscale image, a color image, or a color image with associated depth information. In an embodiment, each image example represents a desired visual style of a corresponding target object. In an embodiment, eachAtty. Docket No. 4502-84201 (6000729PCT02)region mask identifies spatial regions corresponding to one or more target objects in the image frames of the guiding video.

[0110] In block 804, the computing device generates a set of hybrid inputs based on the guiding video, the set of image examples, and the set of region masks. In an embodiment, the computing device generates the hybrid inputs by selecting, for each target object and each image frame, a matching image example from the set of image examples based on viewpoint similarity; and generating the hybrid inputs for each target object and each image frame by masking a target object region in the image frame and concatenating the masked frame with the selected matching image example.[OHl] In block 806, the computing device generates a set of multi-style latent features based on the hybrid inputs using a multi-style diffusion model. In an embodiment, the computing device further receives one or more textual descriptions associated with the target objects; and generates the set of multi-style latent features based on the hybrid inputs and the textual descriptions using the multi-style diffusion model.

[0112] In block 808, the computing device generates, using a restoration model, a sequence of multi-style output frames based on the set of multi-style latent features and the set of hybrid inputs. In an embodiment, the restoration model reconstructs the multi-style output frames to match resolution and channel structure of the guiding video.

[0113] In block 810, the computing device obtains camera view and camera parameters from the guiding video. In an embodiment, the computing device obtains the camera views by estimating the camera views using a structure-from-motion algorithm. In an embodiment, the computing device receives metadata associated with the guiding video, wherein the metadata includes the camera views for each image frame of the guiding video; and obtains the camera views from the metadata.

[0114] In block 812, the computing device constructs, using a semantic-aware multi-style three-dimensional (3D) Gaussian splat (GS) model, a multi-style Gaussian model representing a 3D scene based on the multi-style output frames, the set of region masks, the camera parameters, and the camera views. In an embodiment, the multi-style Gaussian model comprises a set of 3D Gaussians representing the 3D scene, wherein each 3D Gaussian comprises a position vector, a scale vector, a rotation quaternion, an opacity value, a color vector, or a semantic mask vector for splatting-based rendering.Atty. Docket No. 4502-84201 (6000729PCT02)

[0115] In block 814, the computing device renders a synthesized image of the 3D scene from a target camera view using the multi-style Gaussian model. In an embodiment, rendering the synthesized image comprises performing point-based alpha by projecting each 3D Gaussian into a 2D Gaussian in the target camera view. In an embodiment, the target camera view is a novel view not present in the guiding video.

[0116] In block 816, the computing device generates a hybrid-style video from the synthesized image, wherein the hybrid-style video maintains an original style of the guiding video for non-target regions and applies visual styles from the image examples to the one or more target objects with spatial and temporal consistency. In an embodiment, the hybrid-style video is displayed on a display, screen, or monitor of the computing device for the benefit and enjoyment of the user. In an embodiment, the hybrid-style video is stored in a non-transitory computer-readable storage medium and / or transmitted to a remote device over a network for display or further processing.

[0117] FIG. 9 is a schematic diagram of a computing device 900 (e.g., a personal computer, smart phone, smart tablet, handheld gaming device, etc.) according to an embodiment of the disclosure. The computing device 900 is suitable for implementing the disclosed embodiments as described herein. The computing device 900 comprises ingress ports / ingress means 910 (a.k.a., upstream ports) and receiver units (Rx) / receiving means 920 for receiving data; a processor, logic unit, or central processing unit (CPU) / processing means 930 to process the data; transmitter units (Tx) / transmitting means 940 and egress ports / egress means 950 (a.k.a., downstream ports) for transmitting the data; and a memory / memory means 960 for storing the data. In an embodiment, the receiver units (Rx) / receiving means 920 comprise a discrete circuit, integrated circuit, chip set, package, hardware module, electronic device, or other structure capable of receiving signals. In an embodiment, the transmitter units (Tx) / transmitting means 940 comprise a discrete circuit, integrated circuit, chip set, package, hardware module, electronic device, or other structure capable of transmitting signals. The computing device 900 may also comprise optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the ingress ports / ingress means 910, the receiver units / receiving means 920, the transmitter units / transmitting means 940, and the egress ports / egress means 950 for egress or ingress of optical or electrical signals.

[0118] The processor / processing means 930 is implemented by hardware and software. The processor / processing means 930 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field-programmable gate arrays (FPGAs), application specific integratedAtty. Docket No. 4502-84201 (6000729PCT02)circuits (ASICs), and digital signal processors (DSPs). The processor / processing means 930 is in communication with the ingress ports / ingress means 910, receiver units / receiving means 920, transmitter units / transmitting means 940, egress ports / egress means 950, and memory / memory means 960. The processor / processing means 930 comprises a hybrid-style video generation module 970. The hybrid-style video generation module 970 is able to implement the methods disclosed herein. The inclusion of the hybrid-style video generation module 970 therefore provides a substantial improvement to the functionality of the computing device 900 and effects a transformation of the computing device 900 to a different state. Alternatively, the hybrid-style video generation 970 is implemented as instructions stored in the memory / memory means 960 and executed by the processor / processing means 930.

[0119] The computing device 900 may also include input and / or output (I / O) devices or I / O means 980 for communicating data to and from a user. The I / O devices or I / O means 980 may include output devices such as a display for displaying video data, speakers for outputting audio data, etc. The I / O devices or I / O means 980 may also include input devices, such as a keyboard, mouse, trackball, etc., and / or corresponding interfaces for interacting with such output devices. The memory / memory means 960 comprises one or more disks, tape drives, and solid-state drives and may be used as an over-flow data storage device, to store programs when such programs are selected for execution, and to store instructions and data that are read during program execution. The memory / memory means 960 may be volatile and / or non-volatile and may be read-only memory (ROM), random-access memory (RAM), ternary content-addressable memory (TCAM), and / or static RAM (SRAM).

[0120] Embodiments of the present disclosure provide at least the following technical advantages.

[0121] A. A novel system to generate a hybrid-style free-view video with controls over both the semantic scene content and desired styles of various objects. The semantic scene content is controlled by the guiding video input, and the desired styles of specific objects are controlled by image examples (and optionally text descriptions) of the objects. A 3D representation in the form of 3D Gaussians is learned upon the hybrid-style video frames. The learned 3D representation supports two types of free-view video generations: 1) generation of video frames from novel views with the hybrid style upon which the 3D representation was built; and 2) generation of video frames fromAtty. Docket No. 4502-84201 (6000729PCT02)novel views where specific objects can have new styles controlled by new image examples (and optionally text descriptions), not the same as the style for learning the 3D representation.

[0122] B. A novel function that transforms a guiding video input into a hybrid-style free-view video by using image examples of target styles of various objects (and optionally text descriptions) and corresponding semantic object masks in the guiding video input. The masked objects of the guiding video frames are transferred to the desired styles of the corresponding image examples (optionally assisted with text descriptions), and the remaining region retains the original style of the guiding video, resulting in a hybrid-style output video. The learned 3D Gaussian representation not only promotes spatial-temporal consistency for generated video, but also associates different Gaussians to different semantic objects, which not only supports free-view generation of hybridstyle video for the object styles upon which the 3D Gaussian representation was learned, but also supports free-view generation of hybrid- style video where specific objects can change to a different style illustrated by different corresponding image examples (and optionally text descriptions).

[0123] While several embodiments have been provided in the present disclosure, it may be understood that the disclosed systems and methods might be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered as illustrative and not restrictive, and the intention is not to be limited to the details given herein. For example, the various elements or components may be combined or integrated in another system or certain features may be omitted, or not implemented.

[0124] In addition, techniques, systems, subsystems, and methods described and illustrated in the various embodiments as discrete or separate may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of changes, substitutions, and alterations are ascertainable by one skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

Atty. Docket No. 4502-84201 (6000729PCT02)CLAIMSWhat is claimed is:

1. A method implemented by a computing device, comprising:receiving a guiding video, a set of image examples, and a set of region masks associated with the image examples, wherein the guiding video comprises a sequence of image frames, and wherein each image frame comprises one or more target objects;generating a set of hybrid inputs based on the guiding video, the set of image examples, and the set of region masks;generating a set of multi-style latent features based on the hybrid inputs using a multi-style diffusion model;generating, using a restoration model, a sequence of multi-style output frames based on the set of multi-style latent features and the set of hybrid inputs;obtaining camera parameters and camera views corresponding to the guiding video; constructing, using a semantic-aware multi-style three-dimensional (3D) Gaussian splat (GS) model, a multi-style Gaussian model representing a 3D scene based on the sequence of multi-style output frames, the set of region masks, the camera parameters, and the camera views;rendering a synthesized image of the 3D scene from a target camera view using the multistyle Gaussian model;generating a hybrid-style video from the synthesized image, wherein the hybrid-style video maintains an original style of the guiding video for non-target regions and applies visual styles from the set of image examples to the one or more target objects with spatial and temporal consistency; anddisplaying the hybrid-style video on a display device.

2. The method of claim 1, further comprising:receiving one or more textual descriptions associated with the target objects; and generating the set of multi-style latent features based on the hybrid inputs and the textual descriptions using the multi-style diffusion model.

3. The method of any of claims 1-2, wherein the multi-style Gaussian model comprises a set of 3D Gaussians representing the 3D scene.Atty. Docket No. 4502-84201 (6000729PCT02)4. The method of any of claims 1-3, wherein each 3D Gaussian comprises a position vector, a scale vector, a rotation quaternion, an opacity value, a color vector, or a semantic mask vector for splatting-based rendering.

5. The method of any of claims 1-4, wherein rendering the synthesized image comprises performing point-based alpha by projecting each 3D Gaussian into a two-dimensional (2D) Gaussian in the target camera view.

6. The method of any of claims 1-5, wherein the target camera view is a novel view not present in the guiding video.

7. The method of any of claims 1-6, further comprising generating a semantic-aware 3D GS model based on the guiding video to stabilize the multi-style Gaussian model.

8. The method of any of claims 1-7, wherein generating the semantic-aware 3D GS model comprises:receiving the guiding video, the set of region masks, and the camera views; generating an initial 3D GS model using a Gaussian splatting method; andassigning semantic labels to Gaussians in the initial 3D GS model using an inverse semantic projection module to generate the semantic-aware 3D GS model.

9. The method of any of claims 1-8, further comprising updating the semantic-aware 3D GS model by refining the Gaussians and reapplying an inverse semantic projection.

10. The method of any of claims 1-9, further comprising generating the semantic-aware multistyle 3D GS model by:receiving the multi-style output frames, the set of region masks, the camera views, and the camera parameters;generating an initial multi-style 3D GS model using Gaussian splatting; andAtty. Docket No. 4502-84201 (6000729PCT02)assigning semantic labels to multi-style Gaussians in the initial multi-style 3D GS model using an inverse semantic projection module to generate the semantic-aware multi-style 3D GS model.

11. The method of any of claims 1-10, further comprising generating the semantic-aware multistyle 3D GS model by:receiving the multi-style output frames, the set of region masks, the camera views, the camera parameters, and the semantic-aware 3D GS model, wherein the semantic-aware 3D GS model comprising Gaussians with first semantic masks;generating an initial multi-style 3D GS model using Gaussian splatting; andassigning semantic labels to multi-style Gaussians in the initial multi-style 3D GS model using an inverse semantic projection module to generate the semantic-aware multi-style 3D GS model, wherein the semantic-aware multi-style 3D GS model comprising Gaussians with second semantic masks.

12. The method of any of claims 1-11, further comprising:computing a semantic regularization loss based on differences between the first semantic masks and the second semantic masks;computing a multi-style loss for each synthesized image;updating the semantic-aware multi-style 3D GS model based on the semantic regularization loss and the multi-style loss; andbackpropagating the multi-style loss to update parameters of the multi-style diffusion model and the restoration model.

13. The method of any of claims 1-12, wherein the multi-style diffusion model comprises a multi-style embedding model and / or a reverse diffusion model, and wherein training the multi-style diffusion model comprises:generating a training masked hybrid input based on a training guiding image, a training target image, a training region mask, and a training text description;computing, using the multi-style embedding model, an embedded training joint latent feature based on the training masked hybrid input;Atty. Docket No. 4502-84201 (6000729PCT02)generating, using the reverse diffusion model, a training multi-style latent feature based on the embedded training joint latent feature and a random noise;processing, using the restoration model, the training multi-style latent feature and the training masked hybrid input to obtain a training multi-style output; andupdating the multi-style embedding model and / or reverse diffusion model based on the training multi-style output.

14. The method of claim 1, further comprising:receiving, for each of the one or more target objects in the 3D scene, the image examples illustrating a desired appearance style, a textual description of the desired appearance style, and the synthesized image; andcomputing, using a multi-style editing module, an edited output image by modifying the synthesized image to match the desired appearance style for each target object.

15. The method of claim 14, wherein generating the edited output image comprises:identifying, for each target object, a best matching image example from the image examples, wherein the best matching image example has a viewpoint most similar to that of the synthesized image;computing, using an alignment module, an aligned synthesized image based on the viewpoint;inverting, using a multi-style inversion module, the aligned synthesized image into a Gaussian noise representation based on 3D structure information of a target object and the best matching image example; andgenerating, using a reverse diffusion module, the edited output image based on the Gaussian noise representation, the best matching image example, the synthesized image, and the 3D structure information of the target object.

16. The method of any of claims 1-15, wherein each image frame comprises one or more of a grayscale image, a color image, or a color image with associated depth information, wherein each image example represents a desired visual style of a corresponding target object, and wherein eachAtty. Docket No. 4502-84201 (6000729PCT02)region mask identifies spatial regions corresponding to one or more target objects in the image frames of the guiding video.

17. The method of any of claims 1-16, wherein generating the hybrid inputs comprises:selecting, for each target object and each image frame, a matching image example from the set of image examples based on viewpoint similarity; andgenerating the hybrid inputs for each target object and each image frame by masking a target object region in the image frame and concatenating the masked frame with the selected matching image example.

18. The method of any of claims 1-17, further comprising obtaining the camera views of the 3D scene using a structure-from-motion algorithm.

19. The method of any of claims 1-18, further comprising:receiving metadata associated with the guiding video; andobtaining the camera views from the metadata.

20. The method of any of claims 1-19, wherein the multi-style diffusion model and the restoration model are trained based on a set of multi-style image pairs with associated masks.

21. A computing device, comprising:a memory configured to store instructions; andone or more processors coupled to the memory and configured to execute the instructions to cause the computing device to perform a method according to any of claims 1-20.

22. A computer program product comprising computer-executable instructions stored on a non-transitory computer-readable storage medium, the computer-executable instructions when executed by one or more processor of a computing device, cause the computing device to perform a method according to any of claims 1-20.Atty. Docket No. 4502-84201 (6000729PCT02)23. A computing device, comprising:a storage means configured to store instructions; andone or more processing means coupled to the storage means and configured to execute the instructions to cause the computing device to perform a method according to any of claims 1-20.

24. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause a computing device to perform operations according to any of claims 1-20.

Citation Information

Cited By

  • A Spatial Intelligent 3D Video Generation Method and System with Spatial Adaptive Noise Injection

    CN122138026A