Systems and method for enhancing aesthetics and interaction within a three-dimensional video environment

The system addresses accessibility and interactivity limitations in 3D video by using a server-client architecture for real-time rendering and interaction, enabling immersive experiences on consumer devices with AI-driven world generation and dynamic audio synchronization.

GB2643597APending Publication Date: 2026-02-25LANKESAR AMITH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
GB2024017720
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-12
Filing Date
2024-12-03
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

Existing 3D video technologies require specialized hardware and significant computational resources, limiting accessibility and interactivity, and lack seamless integration of animations, audio-visual synchronization, and interactive advertising in immersive environments.

Method used

A system utilizing a server-client architecture for encoding and streaming RGB/A and depth video data, enabling real-time rendering and interaction on consumer devices with AI-driven world generation, dynamic audio synchronization, and interactive advertisements.

Benefits of technology

Democratizes immersive 3D content access by providing interactive, scalable, and adaptable experiences on everyday devices, enhancing user engagement and interactivity without specialized hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The system includes a method of enhancing aesthetics and interaction in a three-dimensional video environment. The method includes rendering text and or graphics across three distinct spatial planes w
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention pertains to digital media technology and, more specifically, to the enhancement of three-dimensional (3D) video aesthetics. The invention applies advanced computational techniques, including artificial intelligence, to integrate visual and auditory effects that improve the interactivity and sensory impact of 3D video content across a wide range of display and audio systems, from personal devices to immersive platforms. BACKGROUND OF THE INVENTION

[0002] The field of immersive media consumption has undergone rapid evolution, with increasing demand for engaging and interactive experiences. Traditional two-dimensional (2D) video content remains dominant across a variety of platforms, including entertainment, education, advertising, and gaming. However, 2D content inherently lacks the depth, interactivity, and realism desired by today’s tech-savvy audiences. With the proliferation of smartphones, tablets, and other consumer devices equipped with high-resolution displays and advanced processing capabilities, users now expect a higher degree of immersion and interaction, which 2D content cannot adequately provide.

[0003] Recent advancements in augmented reality (AR) and virtual reality (VR) technologies have shown the potential for highly immersive experiences. However, these technologies often require specialized and costly hardware, such as headsets or AR glasses, which limits their accessibility to a broader audience. Despite improvements in VR headsets like the Oculus Quest and AR systems such as Microsoft’s HoloLens, these solutions remain niche due to factors such as high cost, limited portability, and complex setup requirements. The adoption of immersive technology thus remains constrained, creating a market gap for solutions that provide interactive and engaging experiences on commonly available devices without requiring specialized hardware.

[0004] The rise of 3D video content has further expanded the potential for immersive digital experiences. However, producing and consuming 3D video content faces significant technical and practical challenges. Existing methods for creating and displaying 3D content often rely on advanced tools such as depth cameras, sophisticated rendering software, and significant computational resources. For example, methods leveraging LiDAR sensors or multi-camera arrays are capable of creating detailed 3D reconstructions but are inaccessible to most consumers due to their high cost and complexity. Academic work such as “NeuVV: Neural Volumetric Videos with Immersive Rendering and Editing” (JIAKAI ZHANG ETAL. "NeuVV: Neural Volumetric Videos with Immersive Rendering and Editing." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11 February 2022, China) highlights advanced approaches to volumetric video rendering and editing but underscores the computational demands and limited accessibility of such techniques for consumer-level applications.

[0005] Interactive elements in 3D environments, such as objects, animations, and visual effects, are a critical component of immersive experiences. Current implementations often involve pre-programmed behaviors or require manual adjustments using specialized software. For example, game engines like Unity and Unreal provide tools for creating interactive environments but demand technical expertise and extensive time investment, which can be prohibitive for casual users or small content creators. Similarly, video synchronization with interactive elements remains challenging, as it requires precise alignment between animations, events, and audio-visual cues. Research such as" VOCA: Voice Operated Character Animation" (TIMO BOLKART ET AL. "VOCA: Voice Operated Character Animation." Proceedings of ACM SIGGRAPH 2020, 17 August 2020, Germany) demonstrates advancements in synchronizing audio input with animations for character-driven interactions but highlights the need for robust computational frameworks and expertise to implement such solutions effectively.

[0006] The integration of audio-driven animations into 3D environments has gained attention in both academic and commercial spheres. Techniques such as Fast Fourier Transform (FFT) are commonly employed to analyze audio signals and synchronize visual effects or animations. While effective, these methods often require significant programming expertise and computational resources, further limiting their accessibility to non-expert users. For example, VOCA's real-time audio-to-animation pipeline (Bolkart etal., 2020) illustrates the potential for seamless integration of speech-driven animations but relies on neural network-based methodologies that may not be practical for broader consumer use. Similarly, solutions like those explored in NeuVV (Zhang et al., 2022) provide immersive editing and rendering capabilities but require substantial computational power, limiting their usability for lightweight or real-time applications.

[0007] Additionally, the use of 3D advertising within immersive environments has emerged as a promising avenue for enhancing viewer engagement. Platforms like ARKit and ARCore have enabled developers to overlay digital content into real-world settings, but the seamless integration of advertisements into 3D video environments remains an underexplored area. Traditional advertisements are static and disruptive, failing to leverage the full potential of immersive technologies to deliver interactive and engaging content.

[0008] Another area of innovation is real-time 3D world generation, which involves creating dynamic environments that adapt to video content or user interactions. Current approaches, such as procedural generation algorithms in gaming, allow for the creation of expansive virtual worlds but are computationally intensive and challenging to implement in real-time applications. Research like “Procedural Modeling of Cities” (Parish and Muller, 2001) has laid the groundwork for procedural content generation, yet these techniques are often tailored for offline use and lack the flexibility to adapt dynamically to streaming video content.

[0009] Despite these advances, there remains a significant gap in systems and methods that democratize the creation, interaction, and consumption of 3D video environments. Specifically, existing technologies fail to address the need for:

[0010] Accessible Immersive Content: A system that enables consumers to experience 3D video environments on everyday devices without requiring specialized hardware or extensive computational resources.

[0011] Dynamic Interaction and Synchronization: Methods to seamlessly integrate animations, events, and object behaviors with 3D video playback, allowing for real-time interaction and precise temporal alignment.

[0012] Real-time Audio-Visual Synchronization: A mechanism to dynamically adapt animations and visual elements based on audio input, enhancing the coherence and immersiveness of the experience.

[0013] Interactive 3D Advertising: A framework for embedding advertisements as interactive 3D elements within immersive environments, enabling non-disruptive and engaging brand integrations.

[0014] Procedural 3D World Generation: A system capable of generating adaptive 3D environments in real-time based on video content, metadata, or user input, enhancing the contextual relevance and visual appeal of the experience.

[0015] The invention described herein addresses these unmet needs by introducing a system and method for creating and interacting with immersive 3D video environments. By leveraging advancements in audio-visual synchronization, real-time rendering, and Al-driven world generation, the system provides a comprehensive solution that is accessible, scalable, and adaptable. It eliminates the dependency on specialized hardware, reduces technical barriers, and delivers an interactive, immersive experience on widely available consumer devices. This innovation represents a significant step forward in the evolution of immersive media, bridging the gap between advanced technologies and practical applications for everyday users. SUMMARY OF THE INVENTION

[0016] The present invention relates to a system and method for enhancing user engagement and interactivity within a three-dimensional (3D) video environment, enabling seamless integration of immersive visual elements, dynamic audio synchronization, and interactive features. The invention overcomes the limitations of traditional 2D video content and complex 3D systems by providing an accessible, scalable, and interactive 3D environment that adapts to user inputs and contextual data in real-time.

[0017] The system utilizes a server-client architecture to process and deliver 3D video content enriched with depth information. A server encodes and streams RGB / A and depth video data, metadata, and auxiliary information to user devices using standardized transmission protocols like HLS, DASH, TCP, and UDP. The user device, which may include smartphones, tablets, augmented reality (AR) or virtual reality (VR) headsets, and other display-capable devices, decodes the stream and renders an interactive 3D video environment using a dedicated system processor, GPU, and shaders.

[0018] The invention dynamically integrates a variety of interactive elements and behaviors into the 3D video space. Users can control the depth and camera perspective within the 3D video environment, offering the ability to modify spatial perception by moving closer to or further from the rendered video. This allows users to personalize the level of immersion, adjusting between a compressed or expanded depth configuration. These camera adjustments enhance the perceived layering of objects and improve the overall realism and interactivity of the 3D experience.

[0019] One aspect of the invention is the synchronization of animations, events, and object behaviors with the 3D video playback sequence. Using a playhead mechanism, users can seamlessly navigate through time, observing past, current, and future object states within the 3D environment. This system ensures precise alignment of object behaviors such as movement, transitions, or triggered events with the video’s playback. The playhead can be manipulated through direct input, dragging, or external controls, enabling fine-grained or rapid adjustments to playback while maintaining synchronization of visual and audio elements.

[0020] Another innovation is the dynamic generation and rendering of a 3D world surrounding the video content. The invention analyzes video frames, metadata, and transcribed text to generate relevant 3D objects, textures, and structures. These elements are mapped onto existing or newly created 3D meshes, forming an immersive environment that evolves in real-time based on the video content. For example, a dialogue referencing specific locations or objects, such as “a boat at the docks,” can trigger the automatic generation of corresponding 3D elements, like a dock or a boat, which integrate seamlessly into the environment. The system ensures that all generated elements adhere to the art style, color palette, and thematic content of the 3D video.

[0021] The invention also features an advanced audio integration mechanism, which processes auditory inputs to synchronize object states and animations with sound. Through frequency spectrum analysis and channel filtering, the system extracts key audio characteristics such as rhythm, tempo, and amplitude and uses them to drive dynamic animations and events. For instance, visual effects or object movements can be synchronized to match the beat or intensity of the music, enhancing the audiovisual coherence of the 3D environment. The system also supports audio-driven animations, such as lip-synching for characters, which is achieved using blendshapes mapped to audio frequency bands.

[0022] A further feature is the inclusion of interactive advertisements within the 3D environment. The system renders advertisements as 3D objects, text, or characters that can be positioned dynamically around the video. These advertisements may appear during playback, paused states, or as standalone interactive elements. Users can engage with the advertisements through touch, gestures, or other input mechanisms, creating an integrated and engaging advertising experience.

[0023] The invention's modular design supports various applications, including entertainment, education, gaming, and professional use cases. By leveraging AI, programmatic algorithms, and optimized rendering techniques, the system ensures real-time adaptability while maintaining high-quality visuals. The invention democratizes access to immersive 3D content, enabling users to engage with rich interactive environments on commonly available devices without requiring specialized hardware.

[0024] This invention addresses the growing demand for immersive and interactive digital experiences, making it accessible to a broader audience. Its novel integration of real-time rendering, dynamic object generation, audio synchronization, and customizable depth perception establishes a versatile platform for creating, delivering, and interacting with 3D video content in innovative ways.

[0025] This summary is provided merely for purposes of summarizing some example embodiments, to provide a basic understanding of some aspects of the subject matter described herein. Accordingly, it will be appreciated that the above-described features are merely examples and should not be construed to narrow the scope or spirit of the subject matter described herein in any way. Other features, aspects, and advantages of the subject matter described herein will become apparent from the following detailed description and figures.

[0026] The abovementioned embodiments and further variations of the proposed invention are discussed further in the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The content presented here is depicted as illustrative examples rather than restrictive depictions in the accompanying figures. To enhance the simplicity and comprehensibility of these illustrations, the elements shown are not to be interpreted as being to exact scale. In certain instances, the size of some elements may be intentionally enlarged in comparison to others to facilitate a clearer understanding. Additionally, when it is deemed suitable, the same reference markers may be utilized across multiple diagrams to signify elements that are equivalent or have similar functions. The invention will now be described by the way of example and reference to the accompanying diagrams in which;

[0028] FIG. 1A shows a block diagram illustrating the general system configuration of the invention, where a server processes RGB / A and depth video data, and streams it to user devices via standard streaming protocols (HLS, DASH, HTTP / S, TCP, UDP). The user device incorporates a system processor, network system, user input controller, and a display for rendering 3D video content.

[0029] FIG. IB shows a detailed block diagram of the user device, including its components such as the system processor, storage system, network system, and I / O interface. The system processor integrates a 3D engine, CPU, GPU, and shader functionalities to process and render 3D video content. The network system supports connectivity through Wi-Fi and cellular networks (e.g., 4G, 5G, 6G) for seamless streaming.

[0030] FIG. IC shows examples of different frame layouts used for integrating RGB / A video frames with depth maps in both landscape and portrait orientations. Four distinct configurations are illustrated, demonstrating how depth and color information are aligned to create a cohesive 3D representation.

[0031] FIG. ID shows a flowchart of the color extraction program used to extract colors from specific regions of a video frame render texture. The program defines window regions, retrieves pixel data, and iteratively processes and stores extracted colors, which can later be applied to enhance visual elements in the 3D video environment.

[0032] FIG. 2A is a flowchart illustrating a method for dynamically adapting the color properties of particles in a particle system. The extracted colors are stored and gradually blended with the existing particle colors, ensuring smooth transitions and visual coherence.

[0033] FIG. 2B is an illustration of a particle system within a 3D video environment. The particles are dynamically adapted to enhance depth perception and complement the visual characteristics of the 3D video, using distinct patterns and color distributions to create a cohesive scene.

[0034] FIG. 3 A is a flowchart depicting a method for adapting the color properties of lights or mesh objects in a 3D environment. The system assigns colors to objects, gradually blending the new colors for smooth transitions, synchronized with the dynamics of the 3D video content.

[0035] FIG. 3B illustrates how light colors are distributed within a 3D environment, with symmetrical and dynamic configurations enhancing visual coherence. The colors adapt to the video’s dynamics to maintain a harmonious scene.

[0036] FIG. 3C illustrates how light colors affect projection distances in the 3D environment. Darker colors create shorter projections, enhancing focus, while lighter colors extend projections, emphasizing depth and broad illumination.

[0037] FIG. 4A illustrates a multiplayer gaming interface integrated within a 3D video environment. It depicts gameplay elements such as character movement, action buttons, and reward systems, fostering engagement and interaction with the 3D video content.

[0038] FIG. 4B illustrates a flowchart outlining the gameplay method within the 3D video environment. The process includes initializing the game, selecting characters, performing actions, earning points, and redeeming rewards, creating an interactive gaming experience.

[0039] FIG. 5A illustrates a 3D video environment overlaid with various user interface (UI) elements. The supertitle text is positioned at the top of the screen, while subtitle text appears at the bottom. Webcam feeds and interactive overlays enhance accessibility and engagement without obstructing the central content.

[0040] FIG. 5B illustrates a method for selecting and positioning 3D UI elements. It demonstrates selecting an object of interest, choosing UI element types like subtitles or 3D text, and assigning spatial coordinates. The method ensures UI elements are dynamically integrated within the 3D video environment.

[0041] FIG. 6A illustrates a curved 3D wireframe mesh with front and side views. This design creates a concave appearance, enhancing immersion by simulating depth in the viewer's perspective.

[0042] FIG. 6B illustrates a flat rectangular 3D wireframe mesh with front and side views. This configuration provides a traditional planar display, suitable for straightforward content presentation.

[0043] FIG. 6C illustrates a hybrid configuration combining a curved mesh in front and a flat mesh as the background. This arrangement enhances depth perception and visual layering.

[0044] FIG. 6D illustrates a 3D wireframe mesh in motion, with a top-down view illustrating rotational dynamics and side views showing spatial interactions. This emphasizes the flexibility of 3D meshes in adapting to user interactions or video content.

[0045] FIG. 7A illustrates a 3D video environment enhanced with reflections, using a central video display and a planar reflection layer. The reflection adds depth and realism to the scene.

[0046] FIG. 7B illustrates a method for generating reflections in a 3D video environment. It starts with a main camera feed, sets up a reflection camera, calculates a reflection matrix, and applies shaders for realism before rendering the final output. This method ensures reflections are dynamically synchronized with the video content.

[0047] FIG. 8 A illustrates a 3D video environment where animated characters are dynamically synchronized with voice audio input. Each character is situated within a 3D video environment, with features like shadows and floor reflections, enhancing realism and interactivity.

[0048] FIG. 8B illustrates various blendshapes assigned to a single animated character. These blendshapes represent incremental transitions in the character's facial expressions or body movements, driven by corresponding audio signals.

[0049] FIG. 8C illustrates a method for synchronizing animations with voice audio input. The process begins with voice audio acquisition, followed by frequency spectrum analysis to identify relevant characteristics. After filtering channels and sensitivities, blendshape ranges are determined and assigned. Animations are gradually transitioned to maintain smooth and realistic movements aligned with the audio input.

[0050] FIG. 9Aillustrates a 3D video environment where auditory data dynamically influences object behavior. Objects like particles and dynamic animations react to audio input, with elements such as ripples on the floor and oscillations reflecting the intensity and frequency of the sound, enhancing visual engagement and realism.

[0051] FIG. 9B illustrates a method for using audio input to alter object states and trigger events within a 3D environment. The process involves acquiring audio input, analyzing its frequency spectrum, filtering channels based on sensitivity thresholds, and extracting relevant data. Objects dynamically change their state in response to audio signals, with gradual transitions ensuring smooth visual effects, and additional events triggered by significant audio features.

[0052] FIG. 10A illustrates a 3D video playback sequence showcasing interactive playback controls and synchronized object behavior. The figure demonstrates a scenario where a panther, a human, and a ball dynamically interact within a 3D environment. The playback sequence represents specific playback positions, allowing users to navigate through the video. Adjusting the playhead updates object animations, such as the panther’s jump or the human kneeling, ensuring synchronization between user input and visual content.

[0053] FIG. 10B illustrates another instance of user interaction with the playback sequence, highlighting dynamic adjustments to object behavior. As the playhead is moved backward or forward, previously rendered animations are updated or rewound, creating a seamless transition for objects such as the panther and ball. The figure emphasizes the real-time adaptability of the 3D environment to user actions, ensuring immersive control.

[0054] FIG. 10C illustrates a flowchart representing the process of synchronizing 3D video playback with user interactions. The system detects user inputs, adjusts the playhead position accordingly, and synchronizes visual effects, animations, and object behaviors with the new playback sequence state. This ensures real-time playback updates and consistent object adaptations to enhance the user experience.

[0055] FIG. HAillustrates the generation of a 3D scene based on video content and metadata. The scene includes elements like a character speaking a line of dialogue, a boat, a crane, and a futuristic tower. Metadata tags such as “synthwave,” “docks,” and “retro” are processed to dynamically create objects, textures, and environmental features, ensuring contextual relevance between the video and the generated 3D environment.

[0056] FIG. 11B shows a workflow diagram for generating and rendering a dynamic 3D environment. The system initializes the video environment, processes video frames and metadata, generates 3D world elements, and maps them onto meshes. The enhanced environment is rendered in real-time or pre-generated for display. The process loops as the video progresses, adapting the environment to new frames or chapters seamlessly.

[0057] FIG. 12A illustrates a 3D video environment showcasing dynamic depth adjustments. The environment includes interactive elements such as cranes, structures, and background features, emphasizing depth dimensions and spatial layering. The figure highlights control elements like zoom-in, zoom-out, and depth sliders for modifying the camera’s perspective.

[0058] FIG. 12B illustrates a detailed configuration for manipulating the depth in the 3D environment. Users interact with control elements to adjust the camera’s position and perspective dynamically, enabling enhanced visualization of the environment. The figure shows how users can fine-tune depth length (Dz) and the camera distance (Cz) to achieve precise depth perception.

[0059] FIG. 12C illustrates a flow diagram for dynamic camera control in the 3D video environment. The process begins with environment initialization and camera control activation. It incorporates user input for forward and backward camera movement, which adjusts the depth dynamically, enhancing the immersive experience.

[0060] FIG. 13 illustrates the integration of a 3D interactive advertisement within the environment. A branded object, "T-Cola," is shown floating dynamically in 3D space, surrounded by interactive avatars and supporting elements. The advertisement uses the spatial depth and interactive components of the environment to engage users and enhance visual appeal.

[0061] FIG. 14A illustrates a 3D video environment presenting multiple-choice options integrated into three-dimensional space. The options, visually represented as buttons or interactable elements, allow the user to influence the video’s outcome. Interaction methods include gestures, visual highlights, or voice commands. The 3D space adds depth, with visual effects enhancing the immersive decision-making experience.

[0062] FIG. 14B illustrates a transition to a 2D interface overlay within the 3D video environment. Choices are displayed as labelled buttons for user interaction, allowing selection through various input methods such as button presses or voice commands. The visual representation of options and user feedback mechanisms ensures clarity while maintaining the integration with the overarching 3D video environment.

[0063] FIG. 14C illustrates the use of 3D objects as interactive elements within the 3D video environment. Users can interact with these objects through gestures, directional inputs, or voice commands to influence the video’s progression. The integration of these objects into the 3D space creates a dynamic, non-linear storytelling experience, enhancing interactivity and immersion.

[0064] FIG. 15A illustrates a 3D video environment using a chroma key or alpha cutoff technique to create transparent video layers. A 3D video mesh projects content across three planes, seamlessly integrating objects like a palm tree behind the video. Transparency effects allow selective visibility, emphasizing unlit rendering for consistent video output and dynamic blending with the 3D environment.

[0065] FIG. 15B illustrates a 3D video environment where a chroma key or alpha cutoff shader selects specific colors or patterns for transparency within the 3D video mesh. This transparency reveals objects, such as a palm tree, behind the video layer, seamlessly blending the video content with the 3D environment while maintaining unlit rendering for consistent visibility.

[0066] FIG. 15C illustrates a rectangular 3D video mesh where a chroma key or alpha cutoff shader applies transparency to selected areas of the video. This allows the background objects, such as a palm tree, to be visible through the video layer. The shader pattern ensures seamless blending of the transparent video with the 3D environment, enhancing visual integration.

[0067] FIG. 15D illustrates a 3D video environment where the transparent areas of the video layer, processed via chroma key or alpha cutoff, are fully removed, revealing objects like a palm tree seamlessly integrated into the scene. The video dynamically transitions between unlit and lit states, influenced by directional lighting and volumetric effects, emphasizing depth and environmental blending. DETAILED DESCRIPTION

[0068] The foilowing detailed description refers to the accompanying drawings, which illustrate specific embodiments of the invention. The descriptions provide clarity on the structure and operational functionality of the various system components, methods and user devices depicted in the drawings. The intention is to furnish comprehensive details that will enable those skilled in the pertinent technical field to practice the invention based on the representations and instructions herein. Reference numbers are consistently applied across multiple figures to denote identical or functionally similar elements, highlighting the cohesive nature of the system's design and operation.

[0069] In the foregoing sections, some features are grouped together in a single embodiment for streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the disclosed embodiments of the present disclosure must use more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the detailed description, with each claim standing on its own as a separate embodiment.

[0070] The specification may refer to “an”, “one” or “some” embodiment(s) in several locations. This does not necessarily imply that each such reference is to the same embodiment(s), or that the feature only applies to a single embodiment. Single feature of different embodiments may also be combined to provide other embodiments.

[0071] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well unless expressly stated otherwise. It will be further understood that the terms “includes”, “comprises”, “including” and / or “comprising” when used in this specification, specify the presence of stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations and arrangements of one or more of the associated listed items.

[0072] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It will be further understood that terms, such as those defined in commonly used dictionaries should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein. 1. Overview

[0073] As illustrated in FIG. lAand FIG. IB, the system and method for enhancing aesthetics and interaction within a three-dimensional (3D) video environment focus on creating an immersive multimedia experience by integrating advanced processing, rendering, and interaction capabilities.

[0074] FIG. 1A presents a high-level block diagram, starting with the Server(s) (101), responsible for processing and distributing 3D video content. The system processes RGB / A and depth video data, leveraging the Streaming System / Broadcaster (103) to transmit this content to User Devices (110) via transmission protocols such as HTTP / S, TCP, UDP, HLS, and DASH. The server ensures efficient streaming of real-time 3D video content and associated metadata, supporting dynamic interaction within the playback environment.

[0075] The User Device (110) represents a range of devices, including smartphones, tablets, virtual reality headsets, augmented reality glasses, laptops, personal computers, and other 3D display devices. Each device incorporates a System Processor (112), a Network System (113), and a Display (115) to receive, process, and render the 3D video content. The system also includes a User Input Controller (116) that enables interaction with the 3D environment through gestures, touch, or other inputs.

[0076] FIG. IB provides a detailed schematic of the User Device (110) components. The Storage System (111) holds video, audio, and 3D environment assets required for rendering and interaction. The System Processor (112) consists of a 3D engine that integrates CPU, GPU, and shaders for real-time video rendering and advanced depth perception. The Network System (113) facilitates connectivity through Wi-Fi and cellular networks (e.g., 5G, 6G), supporting various streaming protocols. An I / O Interface (114) connects the processor to the Display (115) and User Input Controller (116), ensuring seamless communication and user interaction. Together, these components enable real-time depth adaptation, 3D UI rendering, and interactive elements within the 3D video environment.

[0077] Through the integration of server-side processing, efficient streaming, and device-side rendering, this system enhances the visual depth, accessibility, and interactivity of 3D video content for a wide range of applications, including entertainment, education, and gaming. 2. A system for enhancing aesthetics and interaction within a three-dimensional (3D) video environment

[0078] The disclosed system is a comprehensive framework that integrates real-time audio analysis, visual synchronization, and spatial object manipulation to create immersive and dynamic 3D video experiences, allowing users to engage with interactive elements responsive to environmental or user-generated inputs.

[0079] FIG. 1A illustrates a system architecture designed to deliver and process RGB / A and depth video streams for interactive 3D video applications. The system integrates a server (101) and a user device (110) to enable seamless delivery, processing, and interaction with 3D video content.

[0080] The server (101) hosts a streaming system and broadcaster (103) that encodes and transmits RGB / A and depth video data using protocols such as HLS (HTTP Live Streaming), DASH (Dynamic Adaptive Streaming over HTTP), and other standard transport methods like HTTP / S, TCP, and UDP. This server-side architecture ensures that high-quality video streams are transmitted efficiently to the user device, supporting adaptive streaming and compatibility across different network environments.

[0081] At the user device (110), the network system (113) receives the RGB / A and depth video stream from the server. The system processor (112) processes the video data, integrating the depth information with the RGB / A data to create an immersive 3D video experience. User interaction is facilitated by a user input controller (116), which allows the system processor to respond dynamically to user commands, altering the 3D video output or interactive elements as required. The final processed 3D video is displayed on the display unit (115), completing the pipeline from server to user interaction.

[0082] This architecture enables efficient delivery and processing of 3D video content while incorporating interactive controls, providing a responsive and visually rich user experience.

[0083] FIG. IB illustrates the architecture of the user device (110) within the 3D video environment system. This figure expands on the internal components and their interrelation to enable the reception, processing, and interaction with 3D video content.

[0084] The network system (113) facilitates connectivity, receiving RGB / A and depth video streams from the server using various protocols (HLS, DASH, TCP, LDP, HTTP / S) and communication channels like Wi-Fi or advanced cellular networks (e.g., 5G, 6G). The data is passed through the I / O interface (114), which acts as the central communication hub for all internal subsystems of the device.

[0085] The incoming data is processed by the system processor (112), which comprises several modules: A 3D engine integrates the video data with the depth information to render immersive 3D environments.

[0086] The CPU and GPU handle computational and graphical processing tasks, ensuring smooth performance and high-quality visuals.

[0087] Video, audio, and shader modules enhance the audiovisual fidelity, applying effects, optimizing rendering, and synchronizing with the depth information.

[0088] To support these tasks, the storage system (111) provides access to locally stored assets such as video, audio, and 3D environment data. These assets can be used to supplement or enhance streamed content, enabling complex scenes or preloaded elements for a seamless experience.

[0089] User interaction is enabled through the user input controller (116), which sends commands via the I / O interface (114) to the system processor (112). This dynamic input allows for real-time control and interaction with the 3D video content, such as altering perspectives or engaging with virtual elements.

[0090] The processed video is then sent to the display (115), which renders the 3D video output for the user, completing the interactive loop. The integration of these components ensures efficient processing and delivery of an immersive 3D experience while maintaining flexibility for various devices, including smartphones, tablets, virtual and augmented reality headsets, and more.

[0091] FIGS. IC illustrate the treatment of video frames within the system, of a male and female 3D animated dancers. The RGB / A video frames (130, 132) are paired with depth maps (131, 133) to construct a comprehensive data set that a shader utilises to render the 3D video.

[0092] Element 134 - Frame Layout 1: The first example of a landscape frame layout where the RGB / A and depth map are placed adjacent to each other. A top-down approach where the RGB / A frame is on top.

[0093] Element 135 - Frame Layout 2: Another landscape layout configuration, which might illustrate an alternative way of integrating RGB / A and depth data. A top-down approach where the depth data frame is on top.

[0094] Element 136 - Frame Layout 3: A portrait equivalent frame layout, showing the RGB / A and depth map arranged in a manner suitable for portrait-mode displays. A left-right approach where the RGB / A frame is on the left.

[0095] Element 137 - Frame Layout 4: Yet another variation of the portrait frame layout, which might cater to different 3D rendering techniques or the needs of particular portrait-oriented display hardware.

[0096] Various frame layouts are supported, demonstrating the system's flexibility and its ability to optimise for both transmission efficiency and display fidelity.

[0097] FIG. 2B illustrates an implementation of a particle system within a three-dimensional (3D) video environment, where the particles dynamically adapt their visual properties to enhance depth perception and visual coherence. The central element (220) represents the 3D video, acting as the primary backdrop or scene within which the particle system operates. The particle system surrounds this video and spans across the environment, utilizing various rendering techniques and visual adaptations to integrate seamlessly into the 3D space.

[0098] The particles (221, 222, 223, 224, 225) are rendered with distinct patterns in the figure to illustrate how different colors are distributed among the particles. These patterns are a visual representation for illustrative purposes, emphasizing the variety of colors and dynamic changes that occur within the particle system. Each particle's color properties are determined and adapted in response to the color dynamics of the 3D video or environment. This color adaptation ensures that the particles visually complement the 3D video content, maintaining a coherent and immersive experience.

[0099] The system employs a combination of rendering techniques, such as billboard images, planar elements, two-dimensional (2D) meshes, or three-dimensional (3D) meshes, to create particles that effectively simulate environmental effects. These particles enhance the depth perception of the 3D video by creating layers and spatial context within the environment. To ensure seamless integration with the video, alpha blending and masking techniques are applied, allowing the particles to blend harmoniously with the background and foreground elements of the 3D video.

[0100] The real-time or programmed delay adaptation of particle colors further enhances the realism of the scene. This adaptation is controlled by programmatic algorithms, which dynamically adjust the particle colors based on the color properties of the 3D video frame or environment. The algorithms ensure visual coherence and continuity, making the transitions and interactions within the scene appear natural and fluid.

[0101] FIG. 2B demonstrates the interplay between the particle system and the 3D video environment, highlighting how particles with unique color properties can adapt dynamically to enhance visual depth, realism, and coherence. The use of patterns in this figure provides a conceptual representation of how colors are distributed and adapted among the particles to match the aesthetics and dynamics of the 3D video content.

[0102] FIG. 3B illustrates a system for dynamically adapting the color properties of lights in a 3D environment, where the central 3D video (348) acts as the focal point. The surrounding lights (340, 341, 342, 343, 344, 345, 346, 347, 349, 350) are represented with distinct patterns to illustrate the distribution and variation of colors. These patterns are symbolic and intended to represent the unique or paired color properties applied to each light source.

[0103] The system can adapt light colors based on the dynamics of the 3D video frame or the surrounding environment. The adaptation process allows for different configurations, such as random colors, unique colors for each light, or paired colors, as illustrated by the symmetry of light pairs like 340 and 346. This pairing demonstrates the flexibility of the system in assigning consistent or symmetrical color properties to enhance visual coherence. However, the system can also distribute completely random or fully unique colors to achieve varied aesthetic effects, depending on the desired outcome.

[0104] The color properties of these lights are applied to shaders and scripts that dynamically control their behavior, ensuring that the lights integrate seamlessly with the 3D video. The spatial distribution and adaptation of colors enhance the depth perception of the environment, making the 3D video appear more immersive. The interplay between light colors and their spatial arrangements is controlled programmatically to maintain visual continuity and harmony across the entire scene.

[0105] FIG. 3C builds upon the concept shown in FIG. 3B but focuses on how light colors affect projection distances. Lights with darker colors (360, 361, 362) are shown to have reduced projection distances, creating an effect where they appear closer to the central 3D video (348). Conversely, lighter colors are projected farther, emphasizing their intensity and creating a broader illumination effect. This behavior demonstrates how color intensity is not only a visual property but also influences the perceived spatial impact of each light.

[0106] The system leverages this dynamic adjustment of light projection to create depth and visual focus within the 3D environment. The real-time adaptation of light color properties ensures that the lighting system remains responsive to changes in the 3D video or environment. By controlling these properties through programmatic algorithms, the system ensures that the changes are smooth and consistent, maintaining visual coherence throughout the experience.

[0107] Together, FIG. 3B and FIG. 3C illustrate a sophisticated lighting adaptation system that enhances the visual dynamics of a 3D environment. The ability to modify light color, intensity, and projection distance in real-time allows for highly customizable and immersive experiences that respond seamlessly to the underlying video content.

[0108] FIG. 4A illustrates a system integrating a live multiplayer gaming component within a 3D video environment, enabling interactive gameplay during video playback. The figure demonstrates a dynamic interface where multiple users can engage in real-time activities, interact with elements within the 3D environment, and accrue points or rewards that enhance user engagement.

[0109] The user interface provides essential gameplay information and controls. The number of players in the multiplayer environment is displayed in the top-left comer (400), ensuring users are aware of their fellow participants. Adjacent to this, the in-game currency, such as gems, is indicated (401), providing players with an overview of their accrued resources. A score display (402) keeps track of individual player performance, showing points earned during gameplay. Together, these UI elements form an intuitive interface that informs players about their status and progress.

[0110] The central element (403) is the 3D video environment in which gameplay occurs. Within this space, players and non-player characters (NPCs) are rendered, interacting with both the environment and one another. The controlled player is marked with a triangle indicator above their character (404), making it clear which entity is under the player’s control. Other entities, such as animal characters or NPCs (405, 417), and other uncontrolled players or NPCs (406), populate the 3D environment, enriching the scene with diverse interactions. These characters move and interact on a horizontal plane (407), ensuring that all objects remain aligned and within the gameplay field.

[0111] The environment also includes interactive elements, such as collectable in-game currency (408), like gems, which players can pick up during gameplay to increase their resources. These items provide a direct incentive for exploration and engagement within the game.

[0112] The system features a robust control mechanism, represented by the D-PAD (411). Players can move their character in any direction, as indicated by directional arrows (409) within the D-PAD boundary (410). This control scheme enables fluid and precise navigation, allowing players to explore and interact with the 3D environment seamlessly. Additionally, action buttons (415,416) provide functionalities such as attacking, jumping, or interacting with objects, offering a deeper level of interaction and gameplay variety.

[0113] Players can also access and manage in-game items or digital merchandise through an inventory interface (412). Items such as a sword (414) and a shield (413) are displayed, indicating the equipment or assets currently in the player’s possession. These items can be used strategically during gameplay, adding an element of tactical decision-making.

[0114] FIG. 4A showcases an engaging multiplayer gaming system integrated within a 3D video environment. Through intuitive controls, interactive characters, and real-time rewards, the system fosters active user participation and enhances the overall experience. Players can interact with the environment, other characters, and in-game elements, creating a rich, immersive gaming experience synchronized with the 3D video content.

[0115] FIG. 5 A illustrates a system that integrates three-dimensional (3D) user interface (UI) elements within a 3D video environment to enhance interactivity, accessibility, and aesthetic quality. The central element (501) represents the 3D video content, serving as the backdrop for various interactive and dynamic UI components rendered in 3D space. These components are strategically designed to overlay or integrate into the 3D video environment, creating a seamless and immersive user experience.

[0116] A key feature shown in the figure is the use of supertitle text (502), which appears prominently above the video content. This supertitle provides an additional layer of information, such as contextual or descriptive content, and is dynamically positioned to remain visible and accessible within the 3D environment. Complementing this, subtitle text (512) is placed closer to the bottom of the video, offering text-based captions or dialogue synchronized with the video content. Both supertitles and subtitles can be auto-generated, Ai-generated, or pre-generated, providing flexibility in their creation and responsiveness to user needs.

[0117] Webcam feeds are also integrated into the 3D environment. A square-shaped webcam feed (503) and a square with circular borders (504) are displayed within the scene, showcasing different layout possibilities. These webcam feeds are interactive and can serve purposes such as user identification, live interaction, or feedback. Reflections of these webcam feeds are rendered on a reflective plane or mesh (513), as shown in (505) for the square feed and (506) for the bordered feed. These reflections enhance the realism of the 3D environment and add a layer of depth to the visual presentation.

[0118] The system further includes 3D text (508), which is rendered dynamically within the 3D space. This text provides additional interaction points or contextual information for the user. The reflection of the 3D text (509) on the reflective plane or mesh reinforces the spatial depth and immersive quality of the environment. Similarly, a transparent image (510) is positioned in 3D space, offering a visual element that can serve as a branding, logo, or decorative asset. The reflection of the transparent image (511) on the reflective plane adds to the visual complexity and aesthetic of the system.

[0119] An important feature highlighted is the 3D depth protrusion (507), which emphasizes the spatial positioning of the 3D video depth beyond the traditional x and y axes, extending into the z-plane.

[0120] The reflective plane or mesh (513) serves as a critical element in this system, adding realistic reflections for various UI components, such as the webcam feeds, 3D text, and transparent images. This reflective property ties the UI elements cohesively to the 3D environment, making them feel like a natural part of the scene.

[0121] FIG. 5 A demonstrates a sophisticated system for rendering 3D UI elements within a 3D video environment. The integration of text, webcam feeds, transparent images, and reflective properties, along with dynamic depth positioning, creates an engaging and visually compelling experience. This system is designed to respond to user interactions and adapt dynamically, ensuring an intuitive and aesthetically pleasing interface that enhances the overall functionality of the 3D video environment.

[0122] FIG. 6D illustrates a system for enhancing depth perception and visual immersion within a 3D video environment by integrating an additional 2D or 3D mesh positioned behind the primary 3D video content. The diagram provides both a front-top view and a side view to showcase the spatial arrangement of elements, including the primary video content, the additional mesh, and the camera setup.

[0123] The central element of the system is the primary 3D video (620), which serves as the focal point of the environment. Positioned behind this video is an additional mesh or plane (622), designed to create a layered visual effect that enhances the depth and richness of the 3D environment. The additional mesh can display video content that is either synchronized or asynchronous with the primary 3D video. This flexibility allows for dynamic visual effects, such as background animations that complement or contrast with the main video content. The back mesh can also take on various forms flat, curved, or irregular providing versatility in adapting to the design of the 3D scene.

[0124] The system employs a camera (625) that is positioned along a rotational path (626) to capture the scene from various angles, creating a dynamic and immersive perspective for the viewer. The camera's field of view (FOV) (624) encompasses both the primary 3D video and the additional mesh, ensuring that both layers are visible and contribute to the overall depth of the environment. The camera's rotation path allows it to track and center on the camera target (621), which is typically the primary 3D video (620), maintaining focus on the main content while integrating the surrounding elements.

[0125] The horizontal plane for perspective (623) establishes a spatial reference for the scene, ensuring that all elements, including the primary video, the back layer mesh, and the camera, align correctly within the 3D environment. This plane aids in maintaining visual consistency and realism, allowing the viewer to perceive the scene as a cohesive and immersive space.

[0126] Dynamic rendering techniques are applied to the back layer mesh (622) to adapt its visual properties based on the primary 3D video content, thus duplicating the 3D video content in real-time. This can be achieved from the same video feed. This adaptation can involve changes in color, texture, or animation to synchronize with the main video or create complementary effects. By leveraging this dynamic adaptability, the system reinforces the perception of depth and interaction within the 3D environment, making the experience more engaging and visually compelling.

[0127] FIG. 6D demonstrates a system for creating layered and immersive visual effects within a 3D video environment. The integration of a back layer mesh (622) behind the primary 3D video (620), along with the dynamic camera setup (625, 626) and field of view (624), enhances depth perception and visual interaction. The system’s flexibility in configuring the mesh and synchronizing it with the primary video content ensures a rich and customizable 3D experience.

[0128] FIG. 7A illustrates a 3D video environment where a primary 3D video layer (701) is reflected onto a planar surface, enhancing visual depth and the immersive quality of the system through the addition of reflective properties. This system integrates both visual and auditory elements into a cohesive environment, emphasizing depth, dimensionality, and interactivity.

[0129] The primary 3D video layer (701) serves as the central content within the environment, structured with depth across three spatial planes. This depth enables the video to extend into the environment, giving it a realistic sense of space and dimensionality. The 3D video layer interacts dynamically with the surrounding elements, ensuring that it seamlessly integrates into the immersive environment.

[0130] The reflection layer (702) is the focal element in this illustration, acting as a planar surface that reflects both the 3D video content (701) and additional visual elements, such as a sound wave visualization (704). This layer is equipped with a shader that dynamically configures the reflections to match the visual style of the environment, including options for blurring, darkening, lightening, or applying transparency effects. These configurable properties enhance the realism of the reflection and allow it to adapt to various aesthetic requirements.

[0131] The reflection of the 3D video (703) appears on the planar surface, preserving the multiplane depth of the original video (701). This reflection mirrors the position and dimensionality of the primary video, reinforcing the perception of depth and continuity in the 3D space. By dynamically adjusting to the movements and characteristics of the primary 3D video, the reflection creates a seamless and realistic visual experience.

[0132] Above the reflection layer, a 3D sound wave visualization (704) is displayed. This element represents the audio content of the 3D video, visually synchronized with the sound. By translating audio data into a 3D waveform, the visualization adds a dynamic and interactive component to the environment, making sound visible and reinforcing the connection between auditory and visual elements.

[0133] The reflection of the sound wave (705) appears on the planar surface below, aligning with the reflection of the 3D video. This reflected sound wave is slightly blurred to enhance the perception of depth and realism, while maintaining coherence with the surrounding elements. By mirroring both the video and the sound wave, the system achieves an extended visual layer that connects the audio and visual components into a unified and immersive experience.

[0134] In this system, the reflection layer (702) serves as a bridge between the 3D video (701) and the additional visual elements, capturing and mirroring their properties to enhance depth perception. The ability to apply various visual effects to the reflection layer ensures that it adapts seamlessly to the 3D environment, creating a dynamic and visually rich experience. By synchronizing the reflections of the video and the sound wave, the system reinforces the interplay between visual and auditory stimuli, making the environment engaging and immersive for the viewer.

[0135] FIG. 8 A demonstrates the spatial placement of multiple 3D characters around a primary 3D video (801). The 3D video (801) extends across the x, y, and z planes (803), serving as the central backdrop for the scene. Surrounding this video are several dynamically positioned 3D characters, each fulfilling a specific role within the environment.

[0136] A 3D character torso (802) is positioned adjacent to the left side of the video, while a second torso (804) is placed to the right. These adjacent placements allow the characters to interact contextually with the video content, such as narrating or responding to the visual and auditory elements of the 3D video. These characters are rendered with reflections (805 and 806) that are visible on a planar reflective surface below, ensuring seamless integration into the 3D environment.

[0137] In front of the video, a full-body 3D character (808) is displayed. This character is more prominently positioned to engage directly with the audience, serving as the focal point for interaction or storytelling. Like the adjacent characters, the full-body character’s reflection (809) is visible on the reflective surface, mirroring its movements and further reinforcing the realism of the scene.

[0138] Additionally, the 3D video’s reflection (807) appears on the same reflective surface, dynamically interacting with the characters and the environment. This interplay between reflections and animations creates a cohesive, visually engaging experience, where characters and video content appear integrated within a unified 3D space.

[0139] FIG. 8B provides an example of how the 3D characters in FIG. 8 A are equipped with lip-syncing capabilities using blendshapes. These blendshapes represent the range of mouth movements required for realistic lip-syncing. While four specific examples are shown, the system can support additional blendshapes for finer detail.

[0140] Blendshape 820 represents the character’s idle or rest position, with no active mouth movement. This state is used when there is no audio or when the character is pausing between speech segments.

[0141] Blendshape 821 depicts smaller mouth movements, typically corresponding to sounds like "u," "w," or "o." These subtle movements are essential for accurately representing nuanced speech.

[0142] Blendshape 822 shows wider mouth movements, commonly associated with sounds like "c," "d," "g," "k," "n," "r," "s," "th," "y," and "z." This intermediate range of movement aligns with more pronounced speech sounds.

[0143] Blendshape 823 illustrates larger mouth movements, corresponding to vowel sounds such as "a," "i," and "e." These exaggerated movements are crucial for conveying clear and impactful speech.

[0144] By leveraging these blendshapes, the system ensures that the characters' lip movements align precisely with the timing, pitch, and cadence of the audio. Whether the audio source is local, streamed, or cloud-based, the synchronization of lip-syncing enhances the realism and engagement of the characters within the 3D environment.

[0145] FIG. 9A illustrates a system integrating visual music elements rendered as dynamic two-dimensional (2D) or three-dimensional (3D) effects within a 3D video environment. These visual elements interact with and adapt in real-time to audio inputs, creating a cohesive and immersive audiovisual experience synchronized with the rhythm, tempo, and other musical characteristics of the audio.

[0146] At the core of the system is the 3D video (901), which serves as the central content displayed within the environment. Surrounding this 3D video are visual music elements, including a 3D fish (902) that emits particles and visual effects (VFX) based on the rhythm and characteristics of the music. These particles and VFX (903) are generated dynamically in response to real-time audio inputs, creating a synchronized visual representation of the music's energy and rhythm. The system ensures that these visual elements align with the tempo, beat, and other audio features to provide a unified and engaging experience.

[0147] Below the scene, the reflective surface adds an additional layer of realism and depth to the environment. The reflection of the 3D fish, particles, and VFX (904) appears on the planar reflection layer, mirroring the dynamic actions above. Similarly, the reflection of the 3D video (905) ensures that the video content is seamlessly integrated into the reflective environment.

[0148] Another critical component of the system is the soundwave spectrum (907), which visually represents the audio input in real-time. This soundwave reacts dynamically to the rhythm and tempo of the music, providing a direct visual correlation to the audio. Its movement and shape are synchronized with the musical characteristics, offering an additional layer of interactivity. The reflection of the soundwave spectrum (906) on the reflective surface reinforces the connection between the auditory and visual elements, creating a cohesive and immersive experience.

[0149] By combining these elements dynamic 3D visual effects, synchronized particle emissions, soundwave visualizations, and realistic reflections the system transforms the 3D video environment into an engaging and interactive audiovisual space. The seamless integration of visual music elements with real-time audio inputs ensures that the system adapts dynamically to changing musical characteristics, providing users with a captivating and immersive experience.

[0150] FIG. lOAand FIG. 10B demonstrate a system where interactive animations, events, and object behaviors within a 3D environment are synchronized with a 3D video and dynamically adjusted based on user-controlled playback. The system allows users to navigate the playback sequence, seamlessly transitioning between past, current, and future object states using a playhead, while the objects remain synchronized with the video and events in the environment. The "playback sequence" in this context refers to the progression of the 3D video and associated animations, events, and object behaviors overtime. Users can control the playback, moving to earlier, current, or later moments in the video, and the system ensures that the objects, animations, and events remain synchronized with the specific point in time being viewed or interacted with.

[0151] FIG. 10A (Current and Future Object States), the playback is currently positioned at 6 seconds on the playback sequence. The objects and animations displayed are synchronized with the 3D video content at this specific point in time, with their projected future states shown at 14 seconds.

[0152] The 3D video mesh (1001) serves as the central element, projecting the male character (1002) in a real-time animation sequence. At 6 seconds, the male character is walking forward, while in the future (14 seconds), he is depicted kneeling beside a ball (1003) as part of the event sequence. This transition reflects the character's movement and interaction with the environment over time.

[0153] In the 3D environment outside the video, a panther (1005) is synchronized to walk alongside the male character. This behavior is coordinated to match the playback sequence, ensuring it appears as if the panther is naturally moving with the character. At 6 seconds, the panther is walking, and at 14 seconds, it is shown in mid-air (1006), jumping over a ball (1004) that appears as part of the 3D environment. The ball itself, synchronized with the 3D environment, begins to fall at 6 seconds as part of an animation event triggered by the playback sequence. At 14 seconds, the ball has landed (1007) as the man kneels to examine it.

[0154] The boundary of the screen (1024) defines the visible area, ensuring that all interactions occur within the user’s view. The playhead (1013) on the seekbar (1017) represents the current playback position at 6 seconds, with its future position at 14 seconds marked as 1014. Moving the playhead forward along the seekbar triggers the animations and object states corresponding to future events. For example, dragging the playhead from 6 seconds to 14 seconds in a single motion would cause the panther to immediately jump, the ball to fall and land, and the man to kneel, all in perfect synchronization with the 3D video and audio content.

[0155] Additional controls are Skip back 10 seconds (1018), skip forward 10 seconds (1023) for rapid time navigation. As well as Rewind (1019) and fast forward (1022) for controlled backward or forward seeking. Play (1020) and pause (1021) functions are for managing playback state.

[0156] The current playback time (1010), displayed above the playhead, shows the exact moment in the content being played and dynamically updates when the playhead is moved. The current play time for the player interface (1011) synchronizes seamlessly with the content's playback sequence. A progress meter (1012) visually represents the proportion of the content played relative to its total duration, providing real-time feedback. The playhead time (1015) indicates the specific time at 14 seconds, highlighting the playhead’s position within the playback sequence. The total playback duration (1016) shows the complete length of the content, marking the endpoint for the playback sequence.

[0157] FIG. 10B (Past and Current Object States), the playback is positioned at 14 seconds, and the playback sequence now displays the current object states alongside their past states at 6 seconds. Past states are represented with dashed lines, while current states are shown with solid lines, clearly distinguishing between past and present events.

[0158] At 14 seconds, the male character (1003) is kneeling beside the ball, reflecting his current position. His past state at 6 seconds, when he was walking, is represented by 1002 with a dashed outline. The panther's current position (1006) shows it in mid-air as it leaps over the ball. Its past state at 6 seconds, where it was walking beside the man, is shown as 1005 with a dashed outline. Similarly, the ball (1007) is now stationary on the ground at 14 seconds, while its past falling state at 6 seconds is depicted as 1004 with a dashed outline.

[0159] The seekbar (1017) updates dynamically, with the playhead's current position at 14 seconds (1014) and its past position at 6 seconds marked as 1013. Dragging the playhead backward from 14 seconds to 6 seconds replays the animations, events, 3D video and audio in reverse, with the ball appearing to rise, the panther landing back on the ground, and the man transitioning from kneeling to walking. This backward navigation ensures that the system synchronizes all object states and events accurately with the adjusted playback sequence.

[0160] For Synchronization Between Animations and Time Events, the system ensures that all animations and object behaviors are perfectly synchronized with the playback sequence. Objects like the panther, ball, and man transition smoothly between their past, current, and future states as the playhead moves. This synchronization extends to interactions, such as the panther jumping over the ball or the man kneeling, ensuring that these events align with the correct playback sequence position. By representing past states with dashed outlines and current states with solid outlines, the system visually differentiates between what has already occurred and what is happening now, reinforcing the temporal coherence of the scene.

[0161] For Seekbar Behavior, the seekbar enables users to precisely control playback using the playhead. Dragging the playhead slowly allows for detailed observation of animations, creating a slow-motion effect that highlights subtle movements like the panther’s leap or the ball’s descent. Rapid movements of the playhead, such as skipping or fast-forwarding, trigger immediate synchronization of all animations and events, ensuring that objects adjust seamlessly to the new playback sequence position. For example, skipping forward from 6 seconds to 14 seconds causes the man to instantly kneel, the panther to jump, and the ball to settle on the ground, all without disrupting the visual flow.

[0162] FIG. 11A illustrates a method of dynamically generating a three-dimensional (3D) world surrounding a 3D video, enhancing the immersive qualities of the environment by incorporating elements inspired by the content of the video itself. These generated elements adapt in real-time or are pre-generated using programmatic algorithms, artificial intelligence (AI), or other 3D creation tools. The integration creates a seamless blend between the video and its surrounding 3D space, enriching the viewer's experience and engagement.

[0163] At the center of the environment is the 3D video (1100), which serves as the primary content. The video contains visual depth and storytelling elements, such as a male character (1104) and a fantasy teddy bear (1106), interacting as part of the narrative. Within the video, a retro synthwave sun (1103) and dynamic objects like a flying object (1105) further enhance visualisations, sense of motion and action within the scene.

[0164] The video content is augmented by textual and metadata inputs. For example, transcribed text (1101) in the form of a supertitle overlays the video, displaying the dialogue, "Cody: 'Let’s take a boat to the diner by the docks.'" This text provides context for the scene and hints at the interactive possibilities of the environment. Additionally, the video contains metadata tags (1102), such as "synthwave, adventure, action, fantasy, shipping docks, city, palm tree, retro, sun, island," which guide the generation of 3D elements outside the video.

[0165] Based on the metadata and transcribed text, 3D objects are programmatically generated and placed around the video. For instance:

[0166] A docking container crane (1107), retro palm tree (1108), and retro American diner (1109) are generated as part of the shipping dock and diner setting described in the transcribed text.

[0167] A boat (1110) is generated to align with the dialogue's reference to traveling by boat.

[0168] Retro city buildings (1111) and a city tower (1112) populate the background, complementing the scene's urban theme.

[0169] A retro island (1113) surrounds the environment, further immersing the viewer in the "docks and island" setting.

[0170] These objects are dynamically placed within the 3D environment, arranged in front of and around the video to create a coherent, immersive world. The generated objects adopt the same art style, color palette, and thematic elements as the video content to ensure seamless blending. This adaptability allows the environment to evolve in response to the video’s storyline, creating a synchronized audiovisual experience.

[0171] FIG. 12A and FIG. 12B depict the camera control mechanism within a three-dimensional (3D) video environment. These figures demonstrate how user interaction with camera controls adjusts the camera's position along the Z-plane relative to the 3D video, altering the perceived depth of the 3D video. The system enables a dynamically adjustable depth perception, enhancing the immersive quality of the 3D video.

[0172] In FIG. 12A, the camera (1201) is positioned closest to the 3D video (1200), presenting the environment with a normal depth configuration. The camera captures the primary elements of the scene, such as a docking container crane (1202) and a palm tree (1203), which protrude slightly into the 3D space, maintaining a natural appearance of depth. The outline of the back layer of the video (1204) serves as a reference boundary for the extent of the depth effect. The camera Z-plane distance (1208), denoted by the letter Cz, and the corresponding measurement arrow (1209) indicate the distance between the camera and the origin point of the 3D video. At this position, the depth of the 3D video Z-plane (1214), marked as Dz, remains at its default configuration, maintaining a consistent scale.

[0173] User interaction with camera controls allows precise adjustments to the camera's position. Amove camera forward button (1205), identified with a minus sign (-), decreases the distance between the camera and the 3D video. Conversely, a move camera backward button (1207), identified with a plus sign (+), increases the distance. A slider (1215) offers an additional means of adjusting the camera position, while the camera control icon (1206) visually indicates these functionalities. The field of view (1210) and the angular viewport (1212) restrict what is visible to ensure objects outside the viewport are not displayed, maintaining the spatial boundaries of the scene. The move camera icon and button (1216) in the up state along the slider (1215), further illustrates the functionality, where moving the button upward decreases the camera distance, and moving it downward increases the distance.

[0174] In FIG. 12B, the camera (1201) has been moved to its furthest position from the 3D video (1200), resulting in an increased depth effect. The docking container crane (1202) and palm tree (1203) now appear to protrude further into the 3D space, enhancing the visual perception of depth. The outline of the back layer of the video (1204) and 3D video (1200) appears smaller, showing the increased distance between the camera and the 3D video, however the depth is much larger. The camera Z-plane distance (1208), marked as Cz, and its measurement arrow (1209) show an extended length, indicating the increased distance from the origin point. Similarly, the 3D video Z-plane distance (1213) and its measurement arrow reflect the stretched depth configuration, denoted by Dz (1214), which adapts dynamically as the camera moves. The move camera icon and button (1216) in the down state along the slider (1215), further illustrates the functionality, where moving the button upward decreases the camera distance, and moving it downward increases the distance.

[0175] The move camera icon and buttons functionality ensure that users can interactively adjust the depth perception of the scene, either through buttons, sliders, or other control mechanisms, such as a joystick or hand gestures, depending on the implementation.

[0176] These figures highlight the system's ability to dynamically modify the perceived depth of the 3D video environment by adjusting the camera's position relative to the 3D video, offering users a customizable and immersive viewing experience.

[0177] FIG. 13 illustrates a three-dimensional (3D) advertising system integrated into a 3D video environment. This system allows advertisements to be rendered as interactive 3D elements within or around the 3D video space, dynamically engaging users while seamlessly blending with the immersive environment.

[0178] The central component of the figure is the 3D video mesh (1300), which acts as the primary medium for streaming video content. This video occupies the 3D space (1301), extending along the X, Y, and Z axes, providing depth and dimensionality to the video content. Positioned adjacent to this 3D video is a video feed of a woman (1302), displayed as part of the immersive environment, showcasing the capability to overlay live or pre-recorded streams within the 3D space.

[0179] The horizontal ground plane (1303) serves as the reflection layer, mirroring the objects and advertisements within the scene to create a cohesive and visually rich environment. Reflected elements, such as the transparent image (1305) and interactive advertisement, enhance the depth perception and realism of the scene.

[0180] A 3D character (1304) is prominently featured, interacting with an advertised 3D object, in this case, a soda can (1309). This character points toward the object, visually guiding the user’s attention. The soda can is surrounded by a 3D border (1307), indicating its interactivity. The border serves as a visual cue that users can engage with the object through touch, gesture, or other forms of interaction.

[0181] The soda can is branded with 3D text (1306) displaying a fictitious name, "T-Cola," rendered in 3D space. This branding showcases the system’s ability to embed advertisements dynamically and seamlessly into the environment. Below the object is a 3D call-to-action text (1308) that reads "TAP," inviting users to interact with the object. This element emphasizes the interactive nature of the advertisement, allowing for real-time engagement and a more immersive advertising experience.

[0182] Additionally, a transparent image (1305) is displayed within the 3D environment, further showcasing the versatility of the system to incorporate various media types, such as 2D overlays or semi-transparent branding elements.

[0183] The advertisement system dynamically integrates and positions these elements within the 3D environment, ensuring a seamless user experience. It supports rendering during active playback, paused states, or as a standalone component, enabling flexibility and adaptability in engaging viewers.

[0184] FIG. 14A, This figure illustrates a 3D video environment where multiple-choice options are presented as interactive elements within a three-dimensional space. The text "Cody: Where would you like to go?" (1401) is shown as part of the 3D video environment (1402), prompting user interaction. The character Cody (1403) interacts with a bear character (1404), who appears to be making a choice. Three options, "Cinema" (1405), "Arcade" (1406), and "Playground" (1407), are presented as interactable elements. Each option is represented visually with text (1415) and may include additional representations, such as an image of a cinema camera (1416), a joystick (1418), or a child playing (1419). These choices are integrated into the 3D space with visual effects such as reflections (1412) and highlight patterns (1414) to indicate selection. The hand in 3D space (1413) demonstrates gesture-based interaction. These elements can also be activated through voice commands, as represented by the sound icon (1417). The boundary of the screen (1411) defines the visible area of interaction.

[0185] FIG. 14B, This figure transitions to a 2D interface overlay, illustrating the same choices as UI elements within the 3D video environment (1402). The text "Cody: Where would you like to go?" (1401) remains visible, while the options "Cinema" (1408), "Arcade" (1409), and "Playground" (1410) are shown as labeled buttons with corresponding letters (A, B, and C) for alternative input methods, such as button presses. The options retain their additional visual representations, including a cinema camera (1416), joystick (1418), and child playing (1419). The sound icon (1417) highlights the possibility of selecting an option through voice commands. The curved 3D video environment provides depth, and the boundary of the screen (1411) defines the visible area for interaction.

[0186] FIG. 14C, This figure illustrates the integration of interactive 3D objects within the environment. The text "Cody: Where would you like to go?" (1401) is part of the 3D video environment (1402), where Cody (1403) prompts the bear (1404) to make a choice. Interactive elements "Cinema" (1422), "Arcade" (1423), and "Playground" (1424) are represented as 2D or 3D objects in the environment. These objects can be manipulated through user inputs, such as gestures represented by the hand in 3D space (1413) or directional control with a triangle (1420). The player character (1421) also serves as an interactable character for making a selection. Each choice retains its associated visual representations, such as a cinema camera (1416), joystick (1418), or child playing (1419). The boundary of the screen (1411) defines the visible area, ensuring that all interactions occur within the user’s view.

[0187] In all cases, interacting with any of the elements (1405, 1406, 1407, 1408, 1409, 1410, 1422, 1423, or 1424) triggers the system to load a new 3D scene or 3D video segment, creating a dynamic and non-linear storytelling experience. These elements are objects usually contain hidden 2D / 3D colliders in the form of a mesh, square, rectangle, circle, 2D or 3D geometric shape or polygon to trigger an event.

[0188] Upon selection of one of the interactable elements, the system dynamically transitions to the associated 3D scene or video segment. This transition could involve loading a new 3D scene, a new part of the current video, or an entirely different video, allowing for a non-linear progression of the narrative. This feature provides the user with the benefit of influencing the story outcome dynamically, rather than being limited to a linear experience.

[0189] The system supports multiple input methods, including buttons, sliders, joysticks, voice commands, hand gestures, or interactions with characters or objects, enabling seamless transitions and enhanced user engagement, processing inputs in real-time to ensure a seamless user experience. By combining diverse input options and dynamic transitions, this system transforms passive video consumption into an interactive and immersive activity, enabling users to engage deeply with the content.

[0190] FIG. 15A illustrates a three-dimensional (3D) video environment with a curved mesh outline (1501) representing the boundary of the video projection. A man (1502) is shown walking within the 3D video, with the projection extending outward across three planes into the surrounding 3D environment. The video is displayed in an unlit mode, meaning it does not react to environmental lighting sources, providing consistent visibility regardless of the scene's lighting. A 3D palm tree (1503) is positioned behind the 3D video, with its trunk partially obscured by the video layer, demonstrating that objects can coexist within the 3D space behind the video. The chroma key color or alpha channel of the video (1504) is uniform throughout, except for the man, indicating that this specific color or channel is targeted for removal using a cutoff shader or made transparent with an alpha shader. This setup exemplifies how chroma keying or transparency techniques can reveal objects behind the video layer while maintaining the integrity of the 3D scene.

[0191] FIG. 15B builds on the scenario shown in FIG. 15 A, again featuring the curved mesh outline (1501) and the man walking in the 3D video (1502). However, in this figure, the 3D palm tree (1503) positioned behind the video is now partially visible through the video layer, emphasizing how the transparency effect allows for elements within the scene to be seen through the video. The chroma key color or alpha channel (1505) is highlighted with a pattern, illustrating the mechanism by which the shader identifies the chroma key or alpha channel for removal. The shader applies this selection process uniformly across the video to render areas of the video transparent, enabling seamless integration of the 3D video with the surrounding environment.

[0192] FIG. 15C transitions to a rectangular mesh outline (1515) representing the boundary of the 3D video. Similar to the previous figures, the man (1502) walks outward from the video, projected into the 3D environment, and remains unlit for consistent visual representation. The 3D palm tree (1503) continues to demonstrate how objects behind the video are visible through the transparency effect, with the chroma key color or alpha channel (1505) patterned to show how the shader applies the cutoff or transparency process. This figure highlights the flexibility of the system in handling different mesh shapes, emphasizing that the transparency techniques work seamlessly regardless of the video’s mesh configuration.

[0193] FIG. 15D illustrates the final result of the transparency effect, where the mesh outline of the 3D video (1501 or 1515) is no longer visible. The chroma key colors or alpha channels (1505 and 1506) have been fully cut off or rendered transparent, seamlessly blending the 3D video with the surrounding environment. The man (1502) is now displayed in two sections— one unlit and one lit, divided by a clear line (1507) to demonstrate the effect of lighting in the environment. The unlit portion maintains consistent visibility independent of lighting conditions, while the lit portion (1508) reacts to a light source symbolized by an icon (1509). The directional arrow (1510) indicates the path of the light source, which is further accentuated by volumetric light rays (1511) extending into the scene. The 3D palm tree (1503) now fully blends into the scene, showcasing the system's ability to integrate transparent videos with their surroundings, making them visually cohesive and dynamic.

[0194] The combination of FIGS. 15A, 15B, 15C, and 15D demonstrates the versatility of the system in handling transparent 3D videos. The figures illustrate how the transparency effect works across different mesh shapes (curved, rectangular, or others) and lighting conditions. The setup is not limited to these configurations, as the mesh shape can be irregular, spherical, cubic, or other geometries, providing flexibility for various 3D video applications. WHAT IS CLAIMED IS: 1. A system for enhancing aesthetics and interaction within a three-dimensional (3D) video environment, comprising: rendering text and / or graphics across three distinct spatial planes surrounding and within a 3D video, wherein the spatial planes are configured to enhance visual depth and interaction with the 3D video; wherein the text and / or graphics are dynamically integrated into the 3D video environment to enrich visual depth and interaction, ensuring a cohesive aesthetic enhancement; wherein the system processes and streams video and auditory data through a system processor, a system network interface, and a system storage configured to store video and audio data; wherein the system processor transmits video and auditory data to a user device through the system network interface; and wherein the user device processes the received video and auditory data to modify and present an enhanced 3D visualization experience through a display and / or a user interface. 2. The system of claim 1, further comprising: a particle system spanning the three-dimensional (3D) environment, with particles rendered using one or more of the following: billboard images, planar elements, two-dimensional (2D) meshes, or three-dimensional (3D) meshes, to simulate environmental effects and enhance the depth perception of the 3D video; wherein alpha blending or masking techniques integrate particles seamlessly with the 3D video’s background and foreground, adapting their unique color properties in response to the 3D video frame or environment; and wherein the color adaptation is performed in real-time or with a programmed delay, controlled by programmatic algorithms to ensure visual coherence and continuity. 3. The system of claim 1, further comprising: adapting the color properties of lights or 2D / 3D meshes in response to the color dynamics of the 3D video frame or environment, applied through shaders or scripts; and wherein the color adaptation occurs in real-time or with a programmed delay, controlled by programmatic algorithms to ensure visual coherence and continuity throughout the experience. 4. The system of claim 1, further comprising: a live multiplayer gaming component within the 3D video environment, enabling users to participate in real-time events or mini-games during 3D video playback; allowing user interaction with or control of game elements, including user interface controls and characters such as animals, humanoids, or non-player characters (NPCs); and enabling players to earn points or rewards based on performance, redeemable for ingame currency, digital merchandise, or other items of value to enhance engagement with the 3D video content. 5. The system of claim 1, further comprising: rendering three-dimensional (3D) user interface (UI) elements in the 3D video environment to enhance interactivity and accessibility; the 3D UI elements include: 2D / 3D text in 3D space, captions, subtitles, and supertitles; text and audio content that is auto-generated, Ai-generated, or pre-generated; a webcam layer displayed within the 3D space, bordered in shapes such as square, circular, or square with circular edges; and images with transparent layers positioned within the 3D space, displayed on a 2D plane or mapped onto a 3D mesh. 6. The system of claim 1, further comprising: generating an additional 2D or 3D mesh positioned behind the primary 3D video content to create a layered visual effect that enhances depth perception and simulates elements emerging from the video; configuring the additional mesh in various forms, including flat, curved, or irregular shapes, tailored to match or contrast with the original 3D video environment; enabling the mesh to display video content that is either synchronized or asynchronous with the primary 3D video feed; and applying dynamic rendering techniques to adapt the mesh to the visual elements of the 3D video, creating an immersive experience by reinforcing the perception of depth and layered interaction within the 3D environment. 7. The system of claim 1, further comprising: a planar reflection layer within the 3D environment, configured with square, circular, or rectangular borders, and incorporating a shader to reflect the 3D video content and surrounding environment; customizing the reflection layer with visual effects, including blurred or non-blurred reflections, darkened or lightened tones, water reflection, and transparency or cut-off effects; and utilizing alpha blending, cut-off techniques, or both, to achieve a desired level of translucency or opaqueness, enhancing visual realism and depth perception in the 3D video environment. 8. The system of claim 1, further comprising: a real-time or pre-programmed 2D or 3D character dynamically positioned within the three-dimensional (3D) video environment, in front of or adjacent to the 3D video content, configured with lip-syncing capabilities that align mouth movements with audio or speech from the 3D video audio stream, local or cloud-based sources; synchronizing lip-syncing movements with the timing, pitch, and cadence of the audio to ensure seamless integration of character animation with the audio content, enhancing realism and viewer engagement; and dynamically adjusting the character’s position, orientation, and animation in real-time or according to pre-set programming to maintain cohesion with the 3D video scene. 9. The system of claim 1, further comprising: visual music elements rendered as two-dimensional (2D) or three-dimensional (3D) effects or characters surrounding the 3D video mesh, dynamically responding to or adapting in real-time based on audio inputs from streaming or cloud-based sources; and synchronizing the response or adaptation of these visual elements with the rhythm, tempo, or other musical characteristics of the audio to create a cohesive and immersive audiovisual experience within the 3D video environment. 10. The system of claim 1, further comprising: a mechanism to control animations, events, and visual effects of two-dimensional (2D) and three-dimensional (3D) objects positioned around and within a 3D video environment, enabling user interaction with a playhead to adjust its position in real-time by moving it forward or backward within the 3D video playback bar; dynamically synchronizing animations, events, and visual effects associated with 2D and 3D objects to specific points in the 3D video playback bar; and providing seamless adaptation of object behavior and visual properties based on the 3D video’s progress to maintain coherent synchronization throughout playback. 11. The system of claim 1, further comprising: 3D world generation surrounding the 3D video, with elements generated based on video content or individual frames, including transcribed text from speech or metadata; the generated elements being applied as references onto existing or newly generated meshes in real-time or pre-generated through 3D software, programmatic algorithms, or artificial intelligence (AI); adapting dynamically to match the art style, thematic content, or narrative elements of the video to ensure visual coherence and immersive integration; and responding to changes in the video content or playback, enabling real-time or event-driven updates to the environment. 12. The system of claim 1, further comprising: a camera control mechanism allowing user-controlled movement of the camera position within the three-dimensional (3D) video environment; movement of the camera position backward from the 3D video dynamically increasing the perceived depth effect by extending and enhancing spatial layers, creating a gradual depth expansion effect; movement of the camera position forward reducing the perceived depth effect, returning spatial layers to their original configuration; the camera control mechanism supporting multiple input methods, including buttons, sliders, joysticks, voice commands, hand gestures, and mouse inputs; and dynamic adaptation of visual and spatial properties, including lighting and shadow effects, to maintain coherence during camera movement. 13. The system of claim 1, further comprising: a three-dimensional (3D) advertising system integrated into the 3D video environment, with advertisements rendered as 3D objects, interactive elements, or visual overlays dynamically positioned within or around the 3D video space; advertisements appearing during active playback, paused states, or independently as interactive elements; advertisements taking various forms, including text, images, videos, 3D objects, 3D characters, and transparent overlays, dynamically mapped to generated or existing meshes; user interactions with advertisements enabled through touch, gesture recognition, gaze tracking, or voice commands, allowing real-time engagement; adaptive placement of advertisements using algorithms based on video content metadata, user behavior, or contextual data; and collection of interaction data for analysis to refine advertisement strategies and personalize content. 14. The system of claim 1, further comprising: a 3D video environment configured to present multiple choice options within a three-dimensional space, represented as interactable elements, each associated with a unique subsequent action that includes loading a new 3D scene, a new segment of the 3D video, or an entirely new 3D video; the interactable elements visually represented by text, images, 2D objects, or 3D objects, and selectable by the user through input methods such as buttons, sliders, joysticks, voice commands, hand gestures, mouse inputs, or interactions with characters or objects within the 3D environment; the system highlighting selected choices using visual feedback or 2D / 3D colliders, including reflections, highlights, or animations to indicate user interaction; and dynamically transitioning to the associated 3D scene or video segment upon selection, enabling a non-linear progression of the 3D video experience, while supporting simultaneous interaction through multiple input methods, processing user inputs in real-time to trigger the associated outcome. 15. The system of claim 1, further comprising: a three-dimensional (3D) video environment incorporating transparent video layers through a chroma key or alpha cutoff technique to selectively remove portions of the video layer; the transparent video layers rendered across three distinct planes, enabling seamless integration of foreground, midground, and background elements; the system revealing objects or elements positioned behind the transparent video layer and blending them into the 3D environment using real-time rendering techniques; the transparent video layer adapting to lighting conditions within the 3D environment through unlit rendering modes for consistent visibility or light-reactive modes to integrate shadows and highlights from dynamic light sources; a shader mechanism, such as a chroma key alpha, cutoff shader, or transparent alpha shader, analyzing pixel colors or alpha values and discarding those matching a user-defined chroma key color or alpha threshold; blending the transparent video with the 3D environment through properties like reflections, refractions, or transparency gradients to create a cohesive visual output; and dynamically transitioning the transparent video layer between transparency and opacity to reveal or hide elements within the 3D environment in real-time, supporting interactive storytelling or visual effects. 16. A method for enhancing aesthetics and interaction within a three-dimensional (3D) video environment, comprising: rendering text and / or graphics across three distinct spatial planes within and surrounding the 3D video, integrating these elements to enrich visual depth, interaction, and cohesion with the 3D video environment. 17. The method of claim 16, further comprising: implementing a particle system within the 3D video environment, with particles rendered as billboard images, planar elements, two-dimensional (2D) meshes, or three-dimensional (3D) meshes to simulate environmental effects and enhance depth perception; adjusting particle transparency or semi-transparency through alpha blending or masking techniques to seamlessly integrate particles with background and foreground elements, enhancing visual realism; and adapting particle color properties in response to the color dynamics of the 3D video frame or environment, performed in real-time or with a programmed delay, using programmatic algorithms to ensure visual coherence and continuity. 18. The method of claim 16, further comprising: adapting the color properties of lighting and simulated 2D / 3D meshes in response to the color dynamics of the 3D video frame or surrounding environment, executed in real-time or with a predefined delay and managed by programmatic algorithms to maintain visual coherence and continuity throughout the 3D video experience. 19. The method of claim 16, further comprising: enabling live multiplayer gaming within the 3D video environment, including directional character control and action buttons for interacting with animals, humanoids, or non-player characters (NPCs), allowing users to earn points or rewards redeemable for in-game currency, digital merchandise, or other items of value, enhancing engagement and interaction with the 3D video content. 20. The method of claim 16, further comprising: selecting and rendering three-dimensional (3D) user interface (UI) elements within the 3D video environment, including 3D text, subtitles, supertitles, webcam layers with customizable borders, and images with transparency or cut-off properties; positioning the UI elements in 3D space using specified coordinates (x, y, z) and adjusting their placement based on element type or function, such as subtitles at the bottom and supertitles at the top; and enabling user interaction for real-time customization and adjustment of the UI elements, enhancing interactivity and accessibility within the 3D video environment. 21. The method of claim 16, further comprising: rendering a primary video on a front-facing mesh configured as a flat or curved structure; generating an additional mesh positioned behind or in front of the primary video mesh, configured independently as flat, curved, or irregular to match or contrast with the primary video environment; synchronizing the additional mesh to display video content that complements or contrasts with the primary video, enhancing depth perception and creating a layered visual effect; adapting the video content on the meshes using dynamic techniques responsive to changes in the 3D video content; and combining the meshes to create the appearance of elements emerging from or receding into the video environment, enhancing depth perception and visual richness within the 3D scene. 22. The method of claim 16, further comprising: rendering a planar reflection layer within the 3D video environment to mirror the 3D video content and surrounding objects; configuring the reflection layer with border shapes, including square, circular, or rectangular, and applying shader effects to enhance reflection quality; displaying reflections with visual effects such as blurred or sharp details, adjusted tones, water-like appearances, and transparency or cut-off properties; and achieving reflections using alpha blending, cut-off techniques, or both, enabling realtime adjustments to translucency and opaqueness to enhance depth perception and realism in the 3D environment. 23. The method of claim 16, further comprising: acquiring voice or audio input data through a recording or input system; performing frequency spectrum analysis on the audio data using techniques such as Fast Fourier Transform (FFT), Short-Time Fourier Transform (STFT), Wavelet Transform, or equivalent methods to extract spectral characteristics; filtering the analyzed spectrum data by applying channel and sensitivity thresholds to isolate specific audio features; assigning a blendshape based on the identified audio feature or frequency range, determined by the spectral characteristics and the required dynamic range of motion for the animation; and transitioning the current blendshape to the assigned blendshape gradually, ensuring smooth and realistic animations synchronized with the audio content. 24. The method of claim 16, further comprising: processing and streaming video and auditory data through a system processor, a network interface, and a storage system that stores video and audio data; wherein the system processor transmits video and auditory data to a user device, enabling the user device to modify and present an enhanced 3D visualization experience based on the received video and auditory data. 25. The method of claim 16, further comprising: synchronizing animations, events, and visual effects of 2D and 3D objects in the 3D video environment based on user interactions with a playhead; detecting user interactions through input devices, touch interfaces, or programmatic controls; adjusting the playhead position to enable real-time control over object behaviors, including state changes, triggered animations, and event activations, relative to specific moments in the 3D video playback; and providing real-time visual feedback during playhead adjustments to ensure seamless updates to animations, events, and visual effects. 26. The method of claim 16, further comprising: generating a 3D world surrounding the 3D video based on the video content or individual frames, including transcribed speech text or metadata; mapping generated elements onto existing or new meshes in real-time or pre-generated using programmatic algorithms or artificial intelligence (AI); adapting the generated elements to match the art style, thematic context, or narrative flow of the video to ensure visual coherence; storing the generated 3D elements for reuse or adaptation across frames or chapters, enabling efficient and continuous enhancement of the environment; and supporting user interaction with the generated elements, including object manipulation or event triggering, to create an interactive and immersive experience. 27. The method of claim 16, further comprising: providing a camera control function that allows user interaction to adjust the camera position within the three-dimensional (3D) video environment; enabling backward camera movement to increase perceived depth by extending and enhancing spatial layering, creating an expanded depth experience; enabling forward camera movement to decrease perceived depth, returning spatial layers to their original configuration; supporting multiple input methods, including buttons, sliders, joysticks, voice commands, hand gestures, and mouse inputs, for camera control; applying dynamic updates to shader effects, including lighting, reflections, and depth-of-field adjustments, to maintain visual coherence during depth changes; and adjusting audio spatialization and scaling effects to synchronize with perceived depth, enhancing the immersive audiovisual experience. AMENDMENTS TO THE CLAIMS HAVE BEEN FILED AS FOLLOWS: 28 10 25

Claims

1. A computer implemented system for rendering a three-dimensional (3D) video environment, the system comprising a processor, memory and a display, configured to generate a composite 3D scene from three independently rendered planes in respective coordinate spaces and to present the composite 3D scene to a user, characterised in that(a) the three planes comprise:(i) a front-plane for user interface graphics rendered in a first planar coordinate space;(ii) a mid-plane containing decoded 3D video imagery rendered in a second coordinate space using depth or disparity associated with the 3D video; and(iii) a rear plane comprising procedural effects including particles or lighting rendered in a third coordinate space;(b) a compositing module combines the planes using depth aware per-pixel blending that, for each output pixel, applies a weight dependent on at least one of: local depth of the mid-plane, parallax between planes, and an occlusion state, thereby producing perceptual depth separation between the planes; and(c) a synchronisation module constrains the front and rear planes to the mid-plane timing and pose by applying per-plane parallax transforms and time base controlso that plane updates remain temporally coherent with the 3D video.

2. The system of claim 1, wherein the planes are rendered in orthogonal coordinate spaces and composited rear, mid and front with z-buffer preserving alpha such that the mid-plane occludes the rear-plane without altering a final alpha of the mid-plane.

3. The system of any preceding claim, wherein the compositing module performs temporal smoothing of the blending weights to reduce flicker during camera motion or scene cuts.

4. The system of any preceding claim, wherein the synchronisation module enforces independent frame rates for the planes and resamples the front and rear-planes to a mid-plane time base to avoid judder.28 10 255. The system of any preceding claim, further comprising a colour extraction program configured to sample a current mid-plane frame within one or more windows, generate a bounded palette of n representative colours, and publish the palette with optional weights.

6. The system of claim 5, wherein rear-plane particles and lighting consume the palette and gradually blend element colours toward assigned palette entries at a bounded rate to maintain visual stability.

7. The system of any preceding claim, wherein a particle system in the rear-plane anchors spawn position and size to mid-plane depth so that particles adhere to scene depth layers.

8. The system of any preceding claim, wherein a lip synchronised avatar is rendered on the front-plane and is time aligned to associated audio by a phoneme driven animation track referenced to mid-plane frame time stamps.

9. The system of any preceding claim, wherein an audio visualiser is rendered on the rear-plane with parameters driven by frequency band energy of the audio associated with the 3D video and constrained by mid-plane depth to avoid overlap with foreground subjects.

10. The system of any preceding claim, wherein a theme or colouration of front-plane user interface elements is derived from the palette of claim 5 with luminance / contrast clamped to maintain legibility over the mid-plane.

11. The system of any preceding claim, wherein the compositing module computes a per tile parallax for the front plane from local mid-plane depth to preserve readability while providing depth cues.

12. The system of any preceding claim, wherein advertisements are inserted on the front-plane into reserved regions whose colour and luminance are adapted using the palette of claim 5 to reduce local contrast steps at boundaries.

13. The system of any preceding claim, wherein the rear plane provides scene adaptive illumination whose hue and intensity follow the palette of claim 5 and whose direction follows a dominant motion vector of the mid plane.28 10 2514. The system of any preceding claim, wherein the mid-plane depth is obtained from stereo disparity, depth camera capture or neural depth estimation.

15. The system of any preceding claim, wherein the compositing module rejects or down weights rear-plane samples when mid-plane depth indicates foreground occlusion.

16. A computer implemented method of rendering within the 3D video environment of claim 1, performed by one or more processors, comprising:(a) rendering user interface graphics on a front-plane in a first coordinate space;(b) rendering decoded 3D video on a mid-plane in a second coordinate space using associated depth or disparity;(c) rendering particles or lighting on a rear-plane in a third coordinate space;(d) compositing the planes using adaptive depth aware blending based on mid-plane depth and parallax to obtain a composite 3D scene; and(e) outputting the composite 3D scene to a display.

17. The method of claim 16, further comprising extracting a bounded colour palette from the mid-plane and adapting rear-plane particles or lighting colours using a gradual blend toward palette entries.

18. The method of claim 16 or 17, wherein plane updates are time aligned by resampling the front and rear-planes to the mid-plane frame timing.

19. The method of any of claims 16-18, wherein front-plane parallax is computed per image tile from mid-plane depth to maintain readability while providing depth cues.

20. The method of any of claims 16-19, wherein a transparency workflow preserves the mid-plane alpha pipeline while compositing rear-plane effects behind mid-plane foreground subjects.

21. The method of any of claims 16-20, wherein user head pose or camera motion is applied as a stabilised parallax transform to the front and / or rear planes while the mid plane follows the decoded 3D video pose to reduce discomfort.28 10 2522. The method of any of claims 16-21, wherein reflections are generated by sampling the mid-plane frame to form a reflection map and compositing the map onto front-plane surfaces with depth aware attenuation.

23. The method of any of claims 16-22, wherein remote user avatars are composited on the front plane and synchronised to the mid-plane time base for multi user interaction.

24. The method of any of claims 16-23, wherein a global play-head controls decoding of the mid plane and triggers updates of the front and rear-planes at defined cadences.

25. The method of any of claims 16-24, wherein rear-plane geometry is procedurally generated using a depth histogram of the mid-plane to create depth coherent background structures.

26. A computer program which, when run on a computer, causes the computer to perform the method of any of claims 16-25.

27. A computer-readable storage medium storing instructions which, when executed by one or more processors, perform the method of any of claims 16-25.

28. The system of any preceding claim, wherein a zoom or scale transform is applied to the mid-plane based on user input or detected scene depth, and the front and rear-planes are reprojected with corresponding parallax offsets to maintain depth coherent composition.

Citation Information

Patent Citations

  • Methods, systems, and computer program product for managing and displaying webpages in a virtual three-dimensional space with a mixed reality system

    US20220292788A1

  • Devices, methods, and graphical user interfaces for interacting with media and three-dimensional environments

    WO2023049418A2