Systems and methods for digital synthesis
The compositing engine addresses CGI synchronization issues by managing data timing and synchronization in video production, enabling real-time correction of visual errors and reducing production costs through accurate CGI integration with live-action footage.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- FD IP & LICENSING LLC
- Filing Date
- 2024-04-17
- Publication Date
- 2026-05-27
AI Technical Summary
Existing digital compositing methods in video production face challenges with CGI synchronization issues, leading to visual problems such as unnatural movement, timing discrepancies, and lighting/shadow inconsistencies, which are costly and time-consuming to correct in post-production.
A compositing engine that manages data timing and synchronization by partitioning cache memory for video frames, scene-related information frames, and digital assets, using latency and network data to ensure accurate alignment and real-time or near-real-time compositing, thereby minimizing visual errors during filming.
Enables real-time correction of CGI synchronization issues, reducing production errors and post-production costs by ensuring seamless integration of CGI assets with live-action footage, allowing for more efficient and realistic visual effects.
Smart Images

Figure 2026516978000001 
Figure 2026516978000002 
Figure 2026516978000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the benefit of U.S. Patent Application No. 18 / 309,978, filed May 1, 2023, entitled "Systems and Methods for Digital Synthesis," which is hereby incorporated by reference in its entirety.
[0002] The present disclosure generally relates to video production, and more specifically, to systems and methods for digital synthesis.
Background Art
[0003] Synthesis is a process or technique that combines visual elements from separate sources into a single image, often creating the illusion that all of those elements are part of the same scene. Today, most, but not all, synthesis is achieved through digital image manipulation. All synthesis involves replacing a selected part of an image, often but not always, with other material from another image. In digital methods of synthesis, software commands specify a narrowly defined color as part of the image to be replaced. The software then replaces all pixels within the specified color range with pixels from another image that are aligned to appear as part of the original image. Visual effects (sometimes abbreviated as VFX) is a process in which images are created or manipulated outside the context of live action filmed in video production and video production. VFX involves the integration of live action video (which may include in - camera special effects), and other live action video or computer - generated imagery (CGI) (digital or optical, animals, or creatures) that looks like reality, but is dangerous, expensive, unrealistic, time - consuming, or impossible to capture on film.
Summary of the Invention
[0004] To provide a basic understanding, various details of this disclosure are summarized below. This summary is not intended to be a broad overview of the disclosure, nor is it intended to identify or describe the scope of any particular element of the disclosure. Rather, the main purpose of this summary is to present some of the concepts of the disclosure in a simplified form before the more detailed explanations presented below.
[0005] In the example, the computer implementation method may include identifying a video frame from among the video frames stored in the cache memory space based on video frame latency data. The video frame latency data may specify the number of the video frame that is stored in the cache memory space before the video frame is selected. The method may further include the steps of identifying a scene-related information frame from among the scene-related information frames based on the timecode of the video frame, and providing a composite video frame based on the video frame and the scene-related information frame.
[0006] In another example, a system for providing augmented video data may comprise a memory for storing machine-readable instructions and data, and one or more processors for accessing the memory and executing the machine-readable instructions. The machine-readable instructions may include a video frame retriever for identifying video frames among video frames based on video frame latency data. The video frame latency data may specify the number of a video frame that is stored in memory before the video frame is selected. The machine-readable instructions may further include a scene-related frame retriever for identifying scene-related information frames among scene-related information frames based on the timecode of the video frame, a digital asset retriever for retrieving digital asset data that characterizes the digital asset, and a digital compositor for providing augmented video data with the digital asset based on the digital asset data, scene-related information frames, and video frame.
[0007] In a further example, a computer implementation method may include a step of identifying video frames among video frames provided by a video camera that represent a scene in film production, based on video frame latency data. The video frame latency data may specify the number of a video frame that is stored in the memory of the computing platform before the video frame is selected. The method may further include a step of identifying scene-related information frames among scene-related information frames provided by a scene capture device, based on an evaluation of the timecode of the video frame against the timestamp of the scene-related information frame. The timestamp for the scene-related information frame may be generated based on a frame delta value representing the amount of time between each scene-related information frame provided by each scene capture device of the scene capture device. The method may further include a step of providing augmented video data with digital assets based on the scene-related information frames and video frames.
[0008] Any combination of the various embodiments and implementations disclosed herein may be used in further embodiments in a manner consistent with this disclosure. These and other embodiments and features can be understood from the following descriptions of the specific embodiments presented herein, in the accompanying drawings and claims. [Brief explanation of the drawing]
[0009] Embodiments will be described with reference to the attached drawings. In the drawings, similar reference numerals may indicate the same or functionally similar elements. The drawing in which an element first appears is generally indicated by the leftmost digit of the corresponding reference numeral.
[0010] [Figure 1] This is a block diagram of an example of a compositing engine that could be used for digital compositing during production.
[0011] [Figure 2] This is a block diagram of an example computing platform.
[0012] [Figure 3] This is a block diagram of an example system for visualizing CGI video scenes.
[0013] [Figure 4] This is an example of a method for providing composite video frames during production.
[0014] [Figure 5] This is an example of a method for providing augmented video data during production.
[0015] [Figure 6] Examples of computing environments that may be used to carry out the methods according to the aspects of this disclosure are provided. [Modes for carrying out the invention]
[0016] Next, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Similar elements in various drawings may be indicated by the same reference numerals for consistency. Furthermore, the following detailed description of embodiments of the present disclosure will mention a number of specific details in order to provide a more thorough understanding of the claimed subject matter. However, it will be apparent to those skilled in the art that embodiments disclosed herein may be practiced without these specific details. In other cases, well-known features are not described in detail to avoid unnecessarily complicating the description. Furthermore, it will be apparent to those skilled in the art that the scale of elements presented in the accompanying drawings may vary without departing from the scope of the present disclosure.
[0017] Examples of digital compositing during the filming (also known as production) of a scene are disclosed herein. When video data and other relevant scene information (necessary for digital compositing) arrive at the compositing device (also known as a compositor), in some cases the compositor may not use most of the relevant scene data. This can result in visual problems in the composite video data provided by the compositor during filming. Because the compositor may receive data at different times (as different devices are used to provide data over a network), in some cases the movement of digital assets may not be properly synchronized with the video footage. Multiple visual problems that can result from CGI synchronization issues may cause CGI assets to appear unconvincing or unrealistic during filming. Such visual problems may include, for example, unnatural movement, timing issues, lighting and shadow issues, and / or scale and viewpoint issues.
[0018] Visual issues can arise during on-set filming, for example, during complex VFX scenes, because the timing and choreography of the action need to be synchronized to ensure that the director's vision or expectations are met. Additionally, visual issues can also make filming more difficult because they are distracting and can transmit production errors downstream (e.g., post-production), requiring the production team to rapidly expend resources to correct these errors. For example, if a particular VFX shot turns out not to be as planned, or if there are mistakes in the original footage, the VFX team will need to spend resources (e.g., computing resources) in post-production to identify the errors / mistakes and correct these issues.
[0019] This specification discloses examples of digital compositing in which a data manager can be used to identify relevant scene-related information data and correct the timing of scene-related information data so that visual problems in composite video data (or augmented video data) can be mitigated or eliminated during on-location shooting. By implementing the systems and methods disclosed herein, real-time (or near real-time (e.g., with a delay of less than a second)) can be achieved from the outset (e.g., during pre-production). The systems and methods disclosed herein combine different shooting techniques with managing the data timing of various data source devices (or systems) communicating over a communication network so that filmmakers or productions (e.g., teams) can view the results (with a delay of less than a second) and use those results to notify and capture performance. Although the systems and methods disclosed herein are presented in relation to production, the examples herein should not be interpreted and / or limited to production only. The systems and methods disclosed herein may be used in any application where visual problems need to be corrected, which may include post-production.
[0020] Figure 1 is a block diagram of an example of a compositing engine 100 that may be used for digital compositing during production. The compositing engine 100 may run on a computing device such as a computing platform (e.g., computing platform 200 as shown in Figure 2, or computing platform 310 as shown in Figure 3) or on a computing system (e.g., computing system 600 as shown in Figure 6). The compositing engine 100 may include a data loader 102. The data loader 102 may partition a cache memory space 104 for storing input data 106 corresponding to data to be used by the digital compositor 108 of the compositing engine 100 to generate augmented video data 110. For example, the data loader 102 may partition the cache memory space 104 into multiple cache locations (or partitions) for storing video frames 112, scene-related information frames 114, and digital asset data 116 about the scene.
[0021] The data loader 102 may store the video frames 112 in a first set of partitions of the cache memory space 104, store the scene-related information frames 114 in a second set of partitions of the cache memory space 104, and store the digital asset data 116 in a third set of partitions of the cache memory space 104. In some examples, the first set of partitions may include even and odd video frame cache locations, and each video frame of the video frames 112 may be stored in one of the even and odd video frame cache locations based on the respective frame number of the video frame. In some examples, the video frames 112 are provided in real time during production, and thus when the scene is being shot, from a video camera (such as a digital video camera). In other examples, the video frames 112 are provided from a storage device, for example, during post-production. Thus, in some cases, the synthesis engine 100 may be used during post-production to provide the extended video data 110.
[0022] The scene-related information frames 114 may be provided over a network from different types of scene capture devices used to capture information related to the shooting of the scene being captured by the video camera. The video frames 112 are provided by the video camera. The network may be a wired and / or wireless network. Each scene-related information frame of the scene-related information frames 114 may include a timestamp generated based on the local clock of the respective scene capture device. The video frames may be provided by the video camera at a different frequency than that at which the scene-related information frames 114 are provided by the scene capture device.
[0023] As used herein, the term “frame” may represent or refer to a discrete unit of data. Therefore, in some examples, the term “frame” as used herein may refer to a video frame (e.g., one or more still images captured or generated by a video camera) or a data frame (e.g., a pose data frame). For example, each scene-related information frame 114 may contain camera pose data for a video camera providing a video frame 112. Pose data for a video camera typically refers to the camera’s position and orientation at a specific point in time, represented in three-dimensional (3D) space. It typically includes information such as the camera’s location, orientation, and field of view, and other relevant parameters such as focal length, aperture, and shutter speed. In some examples, when the cache memory space 104 is partitioned, the digital compositor 108 may initiate digital compositing, which may include applying VFX to at least one video frame (e.g., inserting a CGI asset).
[0024] For example, digital compositor 108 may issue a video frame request 118 for a video frame among video frames 112. The composition engine 100 includes a data manager 142 that includes a video frame retriever 120 that can process the video frame request 118 to provide a selected video frame 122. The video frame retriever 120 may provide the selected video frame 122 based on video frame latency data 124. The video frame latency data 124 may specify the number of video frames stored in the cache memory space 104 for digital composition so that visual problems (such as those described herein) present in the extended video data 110 are minimal or non-existent. Thus, the video frame latency data 124 may specify how quickly a composite video frame (which is part of the extended video data 110) is rendered on the display. For example, if the video frame latency data 124 indicates a three-frame latency, the video frame retriever 120 may provide the third stored frame as the selected video frame 122 to the digital compositor 108.
[0025] The selected video frame 122 may also be provided to the scene-related information frame retriever 126 of the data manager 142. The scene-related information frame retriever 126 may extract a timecode from the selected video frame 122. The scene-related information frame retriever 126 may use the timecode to identify the scene-related information frame of scene-related information frame 114. In some examples, the scene-related information frame is identified by evaluating the extracted timecode against the timestamp of scene-related information frame 114 and selecting the scene-related information frame that is temporally or geographically closest to the timecode of the selected video frame 122. Other techniques for identifying the scene-related information frame of scene-related information frame 114, such as those disclosed herein, may also be used. The scene-related information frame retriever 126 may provide a scene-related information frame as a selected scene-related information frame 128, as shown in Figure 1, which may be received by the digital compositor 108. The compositing engine 100 also includes a digital asset retriever 130 that can retrieve relevant digital asset data about a scene stored in the cache memory space 104 and provide this data to the digital compositor 108 as selected digital asset data 132. In some examples, the data manager 142 includes the digital asset retriever 130.
[0026] In some examples, the digital compositor 108 may receive a virtual environment of a scene and provide a 3D model 134 of the scene. The 3D model 134 may be used by the digital compositor 108 to generate augmented video data 110. In some examples, the 3D model 134 may be generated by a model generator 144 based on environmental data 146. The environmental data 146 may be provided by one or more environmental sensors (not shown in Figure 1). One or more environmental sensors may include, for example, infrared systems, light detection and ranging (LiDAR) systems, thermal imaging systems, ultrasonic systems, stereoscopic systems, RGB cameras, optical systems, or any device / sensor systems currently known in the art, and combinations thereof, that can measure / capture depth and distance data of objects in the environment. Thus, the environmental data 146 may characterize depth and distance data of objects in the environment (e.g., a scene).
[0027] In some cases, a QR code (registered trademark) (or other image mark) may be used to establish a common origin in the coordinate system of a real environment (e.g., a scene). Multiple QR codes (registered trademarks) may be used to define the original location of the coordinate system for the real environment. The QR codes (registered trademarks) may have different orientations to improve the X, Y, and Z information of the coordinate system. Information characterizing the coordinate system and the common origin may be provided as coordinate system data 148. For example, a device (e.g., a mobile phone, tablet, or other portable device) may be used to record the location / position / orientation of each QR code (registered trademark) in the coordinate system of the real environment and provide this data as coordinate system data 148. The model generator 144 receives the coordinate system data 148 and may register virtual markers in the 3D model 134 as starting locations for digital assets to be inserted into the 3D model 134. Thus, the model generator 144 may insert digital assets into the 3D model 134 at the virtual markers registered in the 3D model 134. In some examples, graphical instructions are provided in the 3D model 134 indicating where the digital assets are to be placed within the 3D model 134. Thus, the coordinate system data 148 can provide the transformation (or registration) of the digital assets into the actual environment.
[0028] The 3D model 134 may include a virtual camera representing a video camera used to film the scene, and digital assets (e.g., CGI assets) to be inserted into the scene. The digital compositor 108 may output a composite video frame, which may be provided as extended video data 110 based on the 3D model 134, as well as a selected video frame and a scene-related information frame, and in some cases, digital asset data provided from the cache memory space 104. Thus, each composite video frame may be provided based on the selected video frame 122, the selected scene-related information frame 128, and in some cases, the selected digital asset data 132.
[0029] In some examples, the 3D model 134 may be updated by the digital compositor 108 based on a selected scene-related information frame 128. For example, the digital compositor 108 may update the pose of the virtual camera in the 3D model 134 based on the closest pose data (reflected as the selected scene-related information frame 128) captured or measured for the video camera. In a further example, if the position of the video camera changes to a new camera position in the real-world environment (from the previous camera position), the 3D model 134 may be updated to reflect the new position of the video camera, and the position of any CGI assets in the 3D model 134 may be updated to match the new video camera position based on the selected scene-related information frame 128.
[0030] The digital compositor 108 may, for example, adjust the position, orientation, and / or scale of the CGI asset in the 3D model 134 so that the CGI asset appears in the exact location in the 3D model 134 relative to real-world footage (e.g., a real-world environment). The digital compositor 108 may provide augmented video data 110 using the updated 3D model. For example, the digital compositor 108 may overlay the CGI asset on a selected video frame so that the CGI asset appears in the updated position in the scene based on the updated 3D model 134, providing a composite video frame.
[0031] Since the scene-related information frame 114 is provided over a network (e.g., LAN, WAN, and / or the Internet), network latency may occur in the scene-related information frame 114. To compensate for the network latency effect on the scene-related information frame 114, the synthesis engine 100 may include a network analyzer 136. The network analyzer 136 can analyze the network through which the scene capture device provides the scene-related information frame 114 (e.g., network 253 shown in Figure 2, or network 335 shown in Figure 3) to determine the network latency of the network. In some examples, the data manager 142 includes a network analyzer 136. The network analyzer 136 may output network latency data 138 that characterizes the network latency of the network, which can be received by the scene-related frame retriever 126. The scene-related frame retriever 126 may update the timestamp of each scene-related information frame 114 based on the network latency of the network through which the scene-related information frame 114 is provided, based on the network latency of the network through which the scene-related information frame 114 is provided, based on the network latency data 138. Therefore, in some examples, the scene-related information frame identified by the scene-related frame retriever 126 can be obtained based on an evaluation of the adjusted timestamp (with respect to network latency) of the scene-related information frame 114 and the timecode of the video frame.
[0032] In additional or alternative examples, the timestamp of each scene-related information frame 114 may be updated based on a frame delta value 140. The frame delta value 140 may represent the amount of time between each scene-related information frame provided by each scene capture device. Since the internal clocks of each scene capture device are internally synchronized, the delta values between timestamps of adjacent scene-related information frames from the same scene capture device may be approximately the same over a period of time. The scene-related frame retriever 126 may use the frame delta value to adjust the timestamps of the scene-related information frames 114 provided by each scene capture device to compensate for the clock drift of each scene capture device.
[0033] By configuring the compositing engine 100 based on video frame latency data 124, and in some cases further based on network latency data 138 and / or frame delta value 140, the compositing engine 100 can provide extended shots of the scene at a given frame latency (e.g., 3 frames latency) sufficient to allow shots of the scene to be visualized during production with little or no visual problems. This is because the compositing engine 100 uses a data manager 142 to identify relevant scene-related information frames and correct the timing of these frames according to a timestamp adjustment scheme (such as those disclosed herein) to enable accurate alignment of non-frame data from various and dissimilar scene sources (e.g., scene capture devices) to video frames.
[0034] Figure 2 is a block diagram of an example computing platform 200 that may be used for digital synthesis. The computing platform 200 described herein may synthesize video frames and digital assets to provide a composite stream of video frames for rendering on a display. The computing platform 200 may include a processor unit 202. The processor unit 202 may be implemented on a single integrated circuit (IC) (e.g., die) or using multiple ICs. The processor unit 202 may be a general-purpose processor, a dedicated processor, an application-specific processor, an embedded processor, or the like. The processor unit 202 may include N processors, which may be collectively referred to as a CPU processing pool 206 in the example of Figure 2. In some cases, each processor may include multiple cores.
[0035] The processor unit 202 may include cache memory, which in the example in Figure 2 may be collectively referred to as the CPU memory space 208. The cache memory may include, for example, L1, L2, and / or L3 caches. The CPU memory space 208 may be a logical representation of the partitioned cache memory of the processor unit 202, as disclosed herein. In some examples, the CPU memory space 208 is distributed across multiple dies. For the sake of clarity and brevity, not all components of the processor unit 202 (e.g., functional blocks such as the memory controller, system interface, input / output (I / O) device controller, control unit, and arithmetic logic unit (ALU)) are shown in the example in Figure 2.
[0036] In some cases, the CPU memory space 208 includes a cache (or subset of caches) for each processor (or processor core). In further examples, the CPU memory space 208 may include a shared cache (e.g., an L3 cache) that can be shared between processors and / or cores. The CPU memory space 208 may be partitioned by the synthesis engine 216 for the real-time generation of augmented video data 212. In some examples, the augmented video data 212 corresponds to the augmented video data 110, as shown in Figure 1. The augmented video data 212 may include one or more synthesized video frames that can be rendered on the display 214.
[0037] In some cases, the real-time generation of the augmented video data 212 may correspond to the generation of video frames with a latency of less than or equal to three frames. One or more processors of the CPU processing pool 206 may read and execute program instructions (or parts thereof) representing the synthesis engine 216. The synthesis engine 216 may correspond to the synthesis engine 100 shown in Figure 1. Thus, in the example of Figure 2, the example of Figure 1 may be referenced. The synthesis engine 216 may run on a processor unit 202 to implement at least some of the functions disclosed herein. In some cases, the synthesis engine 216 represents an application (e.g., a software application) that may run on the computing platform 200.
[0038] The computing platform 200 may include a graphics processing unit (GPU) 218 for providing enhanced video data 212. For example, the GPU 218 may include a GPU processing pool 220 that can provide synthetic data that can be transformed to provide enhanced video data 212. For example, the GPU processing pool 220 may include multiple processing units. The GPU 218 may include a GPU memory space 222. For example, the GPU memory space 222 may include similar and / or different types of memory, such as local memory, shared memory, global memory (e.g., DRAM), texture memory, and / or constant memory. Thus, the GPU memory space 222 may include cache memory such as an L2 cache. For the purpose of clarity and brevity, not all components of the GPU 218 (e.g., interconnects, I / O interfaces, registers, ALU, texture units, load / store units, other fixed-function blocks, etc.) are shown in the example in Figure 2. As the CPU memory space 208, the GPU memory space 222 may be a logical representation of the cache memory of a partitionable GPU 218, as shown in the examples disclosed herein. In some examples, the CPU memory space 208 and / or the GPU memory space 222 may correspond to the cache memory space 104, as shown in Figure 1.
[0039] For example, the computing platform 200 may include multiple interfaces, such as a video interface 224 and an I / O interface 226. The video interface 224 may be used to receive video frames 228, which may be provided to the GPU 218 via the bus 230. In some examples, the video frames 228 are provided by a video capture device 232, as shown in Figure 1. The video frames 228 may represent images of a scene. The video frames 228 may also include metadata such as a timecode (or timestamp), frame rate, resolution, and / or other information describing the video frame 228 itself.
[0040] In some examples, the video capture device 232 may be implemented as a video capture card and therefore as part of the computing platform 200. In other examples, the video capture device 232 may be implemented as a standalone device capable of communicating with the computing platform 200 (e.g., using wired and / or wireless media). The video capture device 232 may receive video data (e.g., video data 336 shown in Figure 3) from a video camera (e.g., video camera 306) or from another video source (e.g., storing recorded video data in, for example, the cloud, on disk, or in another storage location). In some examples, the video camera is a digital video camera such as a digital motion picture camera (e.g., Arri Alexa developed by Arri). In other examples, the video camera may correspond to the camera of a portable device (e.g., a mobile phone, tablet, etc.). In some examples, video frame 228 may correspond to one of the video frames 112, as shown in Figure 1.
[0041] In some examples, the I / O interface 226 may be used to receive scene-related information frames 234 and digital asset data 236. Scene-related information frames 234 may include, for example, face tracking data, body tracking data, asset pose data, and / or camera pose data. In some examples, scene-related information frames 234 may correspond to one of the scene-related information frames 114, as shown in Figure 1. The type of scene-related information frame 234 may depend on the type of scene capture device used for the scene. For example, the I / O interface 226 may include a network interface that can be used to communicate with external devices, systems, and / or servers to receive scene-related information frames 234 and digital asset data 236. Scene-related information frames 234 and / or digital asset data 236 may be provided via network 253, as shown in Figure 1.
[0042] Network 253 may include one or more networks and / or the Internet. One or more networks may include, for example, a local area network (LAN) or a wide area network (WAN). In additional or alternative examples, I / O interface 226 may include a hard disk drive interface, a magnetic disk drive interface, and an optical drive interface, or an input device interface that can be used to enable scene-related information frames 234 and digital asset data 236 to be loaded (e.g., stored) into memory 238. Memory 238 may represent one or more memory devices, such as random access memory (RAM) devices, which may include static and / or dynamic RAM devices. One or more memory devices may include, for example, double data rate 2 (DDR2) devices, double data rate 3 (DDR3) devices, double data rate 4 (DDR4) devices, low power DDR3 (LPDDR3) devices, low power DDR4 (LPDDR4) devices, wide I / O 2 (WIO2) devices, high bandwidth memory (HBM) dynamic random access memory (DRAM) devices, HBM2 DRAM (HBM2 DRAM) devices, double data rate 5 (DDR5) devices, and low power DDR5 (LPDDR5) devices (e.g., mobile DDR). In some examples, one or more memory devices may include DDR SDRAM type devices.
[0043] The processor unit 202, GPU 218, video interface 224, I / O interface 226, and memory 238 may be connected to a system bus 230. The system bus 230 may represent a communication / transmission medium through which data can be provided between components, as shown in Figure 2. In some cases, the system bus 230 includes any number of buses, such as a backside bus, a frontside (system) bus, a peripheral component interconnect (PCI), or a PCIe bus. The system bus 230 may include corresponding circuits and / or devices that enable components to communicate and exchange data to provide augmented video data 212 based on video frames 228, scene-related information frames 234, and digital asset data 236.
[0044] In some examples, the processor unit 202 and / or the GPU 218 may convert video frames 228 to texture files, depending on the implementation and design of the computing platform 200. For example, video frames 228 may be decoded and decompressed (e.g., by the processor unit 202 or the GPU 218) and then converted to a texture file format that can be loaded and rendered by the GPU 218 along with VFX, according to examples disclosed herein. In examples where the processor unit 202 has sufficient computing power, the processor unit 202 may implement the conversion from video frames to texture files, and the GPU 218 may use the texture files for video frames 228 for image rendering. In some examples, the conversion from video frames to texture files is performed by both the processor unit 202 and the GPU 218, each handling different operations of the conversion process. For example, the processor unit 202 may be responsible for decoding and decompressing video frames 228, while the GPU 218 may be responsible for converting the decoded video frames to a texture file format and for image rendering.
[0045] Overall, the specific implementation of the video frame to texture file conversion process may depend on the system architecture, hardware capabilities, and software design of the computing platform 200. The compositing engine 216 may be used to implement digital compositing to provide a composite video frame based on video frames 228, scene-related information frames 234, and digital asset data 236, by controlling the components of the computing platform 200. For example, the compositing engine 216 may be implemented as machine-readable instructions that can be executed by the CPU processing pool 206 to control when data is provided to and from the GPU 218, and / or loaded / retrieved from and processed by the CPU and / or GPU memory spaces 208 and 222.
[0046] During CGI filming, it is desirable to know the location and / or behavior (movement) of CGI assets and other elements in relation to each other in the scene. That is, during production, the director or producer may want to know the position and / or location of CGI assets so that other elements (e.g., actors, props, etc.) or the CGI assets themselves can be adjusted to create a seamless integration of the two. For example, the director may want CGI assets to move and behave in a manner consistent with live-action footage. Knowledge of where CGI assets are located and how they behave is important for the director during filming (e.g., production) because it can minimize retakes and further reduce post-production time and costs. For example, knowledge of how CGI assets behave in a scene can help the director adjust the position or behavior of actors relative to the CGI assets. Generally, to help actors know where to look and how to interact with CGI assets, filmmakers often use on-set visual cues such as reference objects, markers, targets, and verbal cues, and in some cases, show actors a pre-visualization of the scene. Pre-visualization (or "pre-vis") is a technique used in filmmaking to create a rough, animated version of the final sequence, giving actors a general idea of what the CGI assets will look like, where they will be located, and how they will behave before the scene is shot.
[0047] Alternative previsualization techniques have been developed that allow directors to view assets in real time on a display during filming. For example, a director may use a previsualization system configured to provide a composite video feed with CGI assets incorporated while the scene is being filmed. An example of such a device / system is described in U.S. Patent Application No. 17 / 410,479, filed August 24, 2021, “Previsualization Devices and Systems for the Film Industry,” which is incorporated herein by reference in its entirety. A previsualization system allows a director to have a “good enough” shot of a scene with CGI assets, and as a result, production issues (e.g., where actors should actually look, where CGI assets should be positioned, how CGI assets should behave, etc.) can be corrected during filming, and therefore at the production stage. A previsualization system may implement digital compositing to provide a composite video frame with CGI assets incorporated.
[0048] During digital compositing, the movement of CGI assets is synchronized with the video footage in film or video, and as a result, when the composite video frames are played back on the display, the movement of the CGI assets is coordinated with the movement of the live-action footage. A pre-visualization system may be configured to provide CGI asset-live-action footage synchronization. For example, a pre-visualization system may implement compositing techniques / operations for combining different elements such as live-action footage, CGI assets, CGI asset movement, and other special effects into augmented video data (one or more composite video frames). The augmented video data may be rendered and visualized on the display by the user (e.g., the director).
[0049] In some cases, data for generating augmented video data may arrive at the pre-visualization system's compositing device (also known as a compositor) at different times, for example, during filming. Data may arrive at the compositing device at different times via a network (such as network 253 shown in Figure 2). Additionally, data may be generated by different scene capture devices based on internal or local clocks (which may drift over time). Furthermore, in some cases, the scene capture device may not be time-synchronized (or jam-synced) with the video camera (e.g., connected to a central timecode generator). In some scenarios, one or more scene capture devices may not support jamming. Therefore, when video frames and other scene data arrive at the compositor, the compositor may not use most of the relevant scene-related information data (frames) for digital compositing, which can result in visual problems in the composite video data. Because the compositor may receive data at different times, in some cases, the movement of CGI assets may not be properly synchronized with the video footage, which can result in several visual problems that may make CGI assets appear unconvincing or unrealistic during filming.
[0050] According to examples disclosed herein, the compositing engine 216 may be configured to mitigate or (in some cases) eliminate visual problems in the augmented video data 212. The compositing engine 216 may control which scene-related information frames 234 are used for compositing to provide augmented video data 212 with reduced or absent visual problems. Thus, in some examples, the compositing engine may employ a data manager 142 as shown in Figure 1.
[0051] For example, during filming (production), a video camera may provide video frames 228 to a computing platform 200 for digital compositing, which may be communicated to a GPU 218 via a bus 230 through a video interface 224. In some examples, video frames 228 are loaded into CPU memory space 208 and / or GPU memory space 222 for VFX processing (e.g., CGI insertion). In some examples, CPU memory space 208 may include even and odd video frame cache locations 238-240, a scene cache location 242, and a digital asset cache location 244.
[0052] The processor unit 202 or GPU 218 may load video frame 228 into one of the even or odd video frame cache locations 238-240 based on the timecode of video frame 228. For example, if the timecode for video frame 228 appears as 01:23:45:12, then since "12" is an even number, the processor unit 202 or GPU 218 may load video frame 228 into the even video frame cache location 238. The timecode for a video frame is a numerical representation of the specific time at which the frame appears in the video, and includes hours, minutes, seconds, and frames. A video frame 228 loaded into an even video frame cache location may be referred to as an even video frame, and therefore a video frame loaded into an odd video frame cache location may be referred to as an odd video frame.
[0053] In some examples, the GPU 218 (or in combination with the processor unit 202) may convert video frames 228 into a texture file format which may be referred to as texture frame data. The texture frame data may be stored in the GPU memory space 222. For example, the compositing engine 216 may cause the GPU processing pool 220 (or in combination with the CPU processing pool 206) to convert video frames 228 and provide texture frame data. The compositing engine 216 may cause threads to be assigned to processing elements of the GPU 218 and / or the processor unit 202 for the conversion of video frames 228. Once converted, the texture frame data may be stored in an even-numbered video frame cache location 238 (based on the timecode for video frames 228) and may be referred to as even-numbered texture frame data.
[0054] Subsequently received video frames may be processed in the same or similar manner as disclosed herein and stored in odd video frame cache locations 240 instead of even video frame cache locations 238. In some examples, when subsequent video frames are converted to texture frame data, the texture frame data may be stored in odd video frame cache locations 240 (based on the timecode for the subsequent video frame) and may be referred to as odd texture frame data. For example, if the timecode for the subsequent video frame is "13", since "13" is odd, the processor unit 202 or GPU 218 may load the subsequent video frame (or odd texture frame data) into odd video frame cache location 240.
[0055] In some cases, the compositing engine 216 may receive video frame latency data 217. The video frame latency data 217 may specify a number of video frames to be stored in the CPU or GPU memory space 208 or 222 (before digital compositing) for digital compositing such that there are minimal or no visual problems present in the augmented video data 212. Thus, the video frame latency data 217 may specify how quickly the composite video frames are rendered on the display 219. Therefore, the number of frames specified by the video frame latency data 217 may be such that there are few or no visual problems in the augmented video data 212. In an unrestricted example, the video frame latency data 217 indicates three video frames. Based on the video frame latency data 217, the compositing engine 216 may select or identify video frames (or texture frame data) for digital compositing. For example, if the video frame latency data 217 indicates three video frames, the synthesis engine 216 may select a third video frame (stored in memory space) for digital synthesis.
[0056] The time or amount of time required to generate texture frame data may vary from video frame to video frame due to the processing or loading requirements of the computing platform 200. The amount of processing required may vary based on the complexity of the frame (e.g., if it contains more complex visual content). Therefore, the video frame conversion time (the amount of time required to generate texture frame data) may vary or change from video frame to video frame. Additionally, odd and even video frames (or odd and even texture frame data) may be loaded (for further processing) into their corresponding cache memory locations before the scene-related information frame 234 reaches the computing platform 200. In some cases, by the time the third video frame arrives and is stored (in some cases converted to the corresponding texture file and then stored), the scene-related information frame 234 (generated at approximately the same time as video frame 228) has already reached the computing platform 200. The compositing engine 216, for example, selects a third video frame for digital compositing, thereby providing sufficient time for the scene-related information frame 234 to reach the computing platform 200.
[0057] In some examples, the synthesis engine 216 may partition the GPU memory space 222. For example, if the GPU 218 has sufficient computing power, the GPU memory space 222 may be partitioned in the same or similar manner as the CPU memory space 208. Thus, in some examples, the GPU memory space 222 may include even and odd video frame cache locations 246-248, a scene cache location 250, and a digital asset cache location 252. Video frames (or corresponding texture frame data) may be stored in the even and odd video frame cache locations 246-248 in the same or similar manner as disclosed herein. Thus, in some implementations, the write of even / odd video frames (or corresponding texture files, if converted by the GPU processing pool 220) may be removed, and the video frame (or texture file) data may instead be stored in the GPU memory space 222. In an example where the corresponding video frame needs to be stored on an external device, the GPU218 may store the corresponding video frame in the CPU memory space 208, and as a result, the corresponding video frame can be provided to the external device.
[0058] When the scene-related information frame 234 is received, the compositing engine 216 may cause the scene-related information frame 234 to be stored in memory 239, scene cache location 242, or scene cache location 250 (for later retrieval). Following the example in Figure 1, the compositing engine 216 may cause the processor unit 202 or GPU 218 to select a third video frame (or third texture frame data) from CPU memory space 208 or GPU memory space 222 in response to receiving the scene-related information frame 234. The compositing engine 216 may cause the scene-related information frame 234 to be loaded into either scene cache location 242 or scene cache location 250. The compositing engine 216 may also cause the digital asset data 236 to be loaded into either digital asset cache location 244 or digital asset cache location 252 for digital compositing.
[0059] In some examples, the compositing engine 216 may cause the processor unit 202 or GPU 218 to generate a virtual environment (3D model) 254 that can represent a scene, and therefore a real-world environment. The compositing engine 216 (or another system) may insert CGI assets into the virtual environment 254 based on digital asset data 236. In other examples, a different system may be used to create the virtual environment, which may be provided to the computing platform 200. The process of creating the 3D model 254 may include creating objects and elements in the scene, such as buildings, props, characters, and special effects. The 3D model 254 may include a virtual camera that may correspond to a video camera used to capture the scene. The virtual camera may be adjusted based on camera pose data. The location of the CGI assets is known in the 3D model 254, and since the virtual environment represents a scene, the CGI assets may be inserted into the scene during digital compositing based on digital asset data 236. In some examples, the 3D model 254 corresponds to the 3D model 134 as shown in Figure 1.
[0060] In some examples, a scene capture device providing scene-related information frames 234 may record or measure the related data at a different frame rate than the video camera provides the video frames. For example, a scene capture device may provide scene-related information frames at a higher frame rate (e.g., about 60 FPS) than a digital camera (e.g., about 30 FPS). Because the scene capture device provides scene-related information frames 234 at a different frequency than the video camera provides the video frames 228, the use of inaccurate scene-related information may result in visual problems in the augmented video data 212. According to the examples disclosed herein, the compositing engine 216 may identify most appropriate or accurate scene-related information frames for digital compositing (e.g., by the digital compositor 108, as shown in Figure 1). In each example, the scene-related information frame 234 provided by the corresponding scene capture device may include a timestamp that may indicate when the scene-related information frame 234 was generated (based on the local clock of the scene capture device).
[0061] For example, the compositing engine 216 may cause the processor unit 202 or GPU 218 to update the 3D model 254 based on the scene-related information frame 234. The compositing engine 216 causes the processor unit 202 or GPU 218 to evaluate the scene-related information frame 234 provided at different points in time by the corresponding scene capture device and identify the scene-related information frame 234 with the timestamp closest to the timecode of the selected video frame (e.g., the third video frame). In some examples, the compositing engine 216 may cause the processor unit 202 or GPU 218 to convert the timecode for the selected video frame to a timestamp. For example, the compositing engine 216 may cause the processor unit 202 or GPU 218 to convert the timecode based on the frame rate of the video camera and the number of frames in the selected video frame. For example, for the timecode of the 14th frame of a video frame provided (by the video camera) at a frame rate of 30 FPS, the calculation is as follows: timestamp = (14-1) / 30 = 0.367 seconds. By using the calculated timestamp for the selected video frame, the synthesis engine 216 can cause the processor unit 202 or GPU 218 to identify a scene-related information frame 234 with a timestamp closest to the timestamp of the selected video frame, which may be referred to as the closest scene-related information frame.
[0062] The synthesis engine 216 may cause the processor unit 202 or GPU 218 to update the 3D model 254 based on the nearest scene-related information frame. For example, the synthesis engine 216 may cause the processor unit 202 or GPU 218 to update the pose of the virtual camera in the 3D model 254 based on the nearest pose data captured or measured for a video camera. When each instance of the scene-related information frame 234 reaches the computing platform 200, the scene-related information frame 234 may be stored in memory 238 or the corresponding cache memory space (e.g., one of the CPU or GPU memory spaces 208 and 222).
[0063] The compositing engine 216 can cause the processor unit 202 or GPU 218 to identify the nearest scene-related information frame for generating a composite video frame. For example, if the position of a video camera in a real-world environment changes (from a previous camera position) to a new camera position, the 3D model 254 may be updated to reflect the new position of the video camera, and the position of any CGI asset in the 3D model 254 may be updated to match the new video camera position based on the nearest scene-related information frame. The compositing engine 216 can cause the processor unit 202 or GPU 218 to adjust, for example, the position, orientation, and / or scale of the CGI asset in the 3D model 254 so that the CGI asset appears in the correct location in the 3D model 254 relative to the real-world image (e.g., the real-world environment).
[0064] For example, if the timestamp for a selected video frame (or selected texture frame data) is closer to the first timestamp value of the first scene-related information frame (e.g., in terms of distance, time, percentage, etc.) than the second timestamp value of the second scene-related information frame, the synthesis engine 216 may use the scene-related information frame 234 (e.g., one of the first or second scene-related information frames) with the given timestamp value as the closest scene-related information frame. In some examples, the synthesis engine 216 may cause the processor unit 202 or GPU 218 to use interpolated scene-related data to update the 3D model 254. For example, if the timestamp for a selected video frame is a given distance, time value, or percentage from the first and second timestamp values, the synthesis engine 216 may cause the processor unit 202 or GPU 218 to interpolate scene-related data based on the first and second scene-related datasets. The interpolated scene-related data may be used as the closest scene-related information frame to update the 3D model 254.
[0065] The processor unit 202 or GPU 218 may use the updated 3D model for digital compositing. In some examples, the 3D model is updated as part of the digital compositing. For example, the processor unit 202 or GPU 218 may overlay a CGI asset onto a selected video frame so that the CGI asset appears in the scene at an updated position based on the updated 3D model 254. In some examples, the compositing engine 216 may cause the processor unit 202 or GPU 218 to convert texture file data for the selected video frame back into a video frame format for blending (e.g., overlaying) the CGI on the selected video frame. Regardless of the technique or method used for digital compositing, the processor unit 202 or GPU 218 may provide a composite video frame for the selected video frame for rendering on the display 219. When inserting a CGI asset into a video frame, the compositing engine 216 may use various techniques to ensure that the CGI asset matches the lighting, shadows, reflections, and other visual characteristics of the live-action footage. This may involve adjusting color grading, applying depth of field or motion blur effects, and adjusting the transparency or blend mode of CGI assets to ensure they look like a natural part of the scene.
[0066] Since the scene-related information frame 234 is provided over the network 253 (e.g., LAN, WAN, and / or the Internet), network latency may occur in the scene-related information frame 234. Network latency may add an additional timing delay to the scene-related information frame 234 so that the synthesis engine 216 can select an inaccurate scene-related information frame. To compensate for the network latency effect on the scene-related information frame 234, the synthesis engine 216 may communicate using the network interface (corresponding to one of the I / O interfaces 226) and send a ping to each scene capture device that provided the scene-related information frame 234. The synthesis engine 216 may determine the amount of time required for the ping data to travel from the computing platform 200 to the scene capture device and back. This amount of time may correspond to network latency. The synthesis engine 216 may adjust the timestamp for the scene-related information frame 234 (from a given scene capture device) based on the determined network latency when it reaches the computing platform 200 over the network 253. The adjusted timestamps can then be evaluated in the same or similar manner as disclosed herein to identify the nearest scene-related information frame for the selected video frame (or selected texture frame data) for digital synthesis.
[0067] In some cases, a scene capture device may be time-synchronized with a video camera. For example, a scene capture device and a video camera may be configured with a timecode generator. Each timecode generator may produce a unique timecode signal that can be synchronized with a master timecode source. The master timecode source may be a standalone device such as a timecode generator. Once the timecode generator is synchronized, the scene capture device may provide a scene-related information frame 234 with a timestamp (or timecode) synchronized with the master timecode source. Similarly, the video camera may provide a video frame 228 with a timecode synchronized with the master timecode source. However, since the clock of the scene capture device is synchronized with the clock of its timecode generator, and clocks are known to drift over time, the timestamp (or timecode) of each instance of the scene-related information frame 234 provided by the scene capture device will differ from the expected time (if there is no drift).
[0068] Additionally, network latency may affect the timestamp (or timecode) of the scene-related information frame 234. This occurs when this data travels from the scene capture device to the computing platform 200 via the network 253. Since the compositing engine 216 is configured based on the video frame latency data 217, the compositing engine 216 can identify or select a video frame (or texture frame data) where sufficient time is provided for the scene-related information frame 234 to reach the computing platform 200 from the scene capture device. According to the examples disclosed herein, the compositing engine 216 can identify the nearest scene-related information frame for the selected video frame (or selected texture frame data) in the same or similar manner as disclosed herein for digital compositing.
[0069] In some cases, the compositing engine 216 causes the processor unit 202 or GPU 218 to adjust the timestamps of scene-related information frames 234 based on frame delta values. The frame delta value may represent the amount of time between each scene-related information frame provided by the scene capture device. Since the internal clocks of the scene capture device are internally synchronized, the delta values between timestamps of adjacent scene-related information frames may be approximately the same over a period of time. The compositing engine 216 may use the frame delta value to cause the processor unit 202 or GPU 218 to adjust the timestamps of scene-related information frames 234 to compensate for drift in the scene capture device. If the compositing engine 216 determines that visual problems are increasing in the augmented video data 212, the compositing engine 216 may automatically adjust the video frame latency to reduce these visual effects, or, in some cases, the user may provide updated video frame latency data that identifies the new video frame latency for the computing platform 200. Therefore, in some cases, the user can vary the amount of visual issues in the augmented video data 212 based on video frame latency data 217 to the computing platform 200. For example, the compositing engine 216 may provide a graphical user interface (GUI), or another module or device may provide a GUI with GUI elements for setting the video frame latency for the computing platform 200. The video frame latency set by the user based on the GUI may be provided as video frame latency data 217. Therefore, by controlling which video frames (or associated texture files) are used for digital compositing based on the video frame latency data 217, sufficient time is provided, and as a result, scene-related information frames 234 can be received from different scene capture devices used in film production.In several cases, a 3-frame latency was determined to provide sufficient time for different types of scene-related information frames to reach the computing platform 200 through network 253. Furthermore, a 3-frame latency was determined to provide sufficiently near real-time composited video, useful in production for directing live-action footage that supports CGI asset-live-action video synchronization.
[0070] In some examples, the computing platform 200 is located on location, i.e., on the film set. In other examples, the computing platform 200 is located in a cloud environment. When the computing platform 200 is located in a cloud environment, the augmented video data 212 can be transmitted via the network to a display 219 located on the film set. In this example, the components of the computing platform 200 are shown as being implemented on the same system, but in other examples, different components may be distributed across different systems (in some cases, in a cloud computing environment) and communicate, for example, via a network.
[0071] In some cases, display 219 is a camera viewfinder, and therefore the augmented video data 212 can be rendered to a camera operator, who may be the director in some cases. In some cases, display 219 is one or more displays, including a camera viewfinder and display of a device, such as a tablet or mobile device. In some cases, one or more displays include displays located in the video village, which is an area on the set, where members of the film crew can observe the augmented and / or (raw) video footage as it is being filmed. Members of the film crew can view the augmented and / or raw video footage to correct any potential identified film production problems (e.g., an actor looking in the wrong direction).
[0072] Since more than one video camera is used during shooting, in some examples the compositing engine 216 may provide corresponding augmented video data (in terms of each digital video camera) in the same or similar manner as disclosed herein. In some examples each video camera has its own computing platform, similar to computing platform 200, as shown in Figure 2, in order to provide corresponding augmented video data. In some examples each video camera is jam-synced. For example, if a scene is being shot by more than one video camera, the video cameras may be electronically time-synchronized so that all digital video cameras may provide video frames at a given point in time with the same or similar timecode. For example all video cameras may provide a given video frame at approximately the same time, and therefore with the same or similar timecode. In some examples the timecode is Society of Motion Picture and Television Engineers (SMPTE) timecode.
[0073] In some examples, the computing platform 200 can be dynamically adjusted based on video frame latency data 217 so that when a new scene capture device is used (e.g., on set, during filming, etc.), new scene-related information frames can be synchronized with the video frames to provide enhanced video data 212 with little to no visual problems. Thus, when the frame rate (the rate at which the scene capture device provides data) changes (e.g., becomes higher or lower), the computing platform 200 can be configured to adapt to the new frame rate based on video frame latency data 217 provided by the user.
[0074] In some examples, once a video frame 228 is used by the computing platform 200, the synthesis engine 216 may cause the processor unit 202 or GPU 218 to remove the video frame 228 from the corresponding cache memory space. Thus, the synthesis engine 216 may expire video frames and associated data (e.g., metadata) that are no longer needed to free up space in the corresponding cache memory space for new data / information that may be processed according to the examples disclosed herein.
[0075] In some cases, for example, if the digital video camera is offline or if video frame 228 cannot be provided to the computing platform 200, the computing platform 200 may not receive video frame 228. In these cases, the computing platform 200 may create an artificial video frame with a timecode that is processed in the same or similar manner as disclosed herein and is therefore temporally aligned with the scene-related information frame 234. For example, if a director wants to adjust a CGI asset but does not need the scene, the computing platform 200 may be used to provide the CGI asset (within the shot) in the extended video data 212, and as a result, the director may use the computing platform 200 to make the appropriate changes.
[0076] Therefore, the computing platform 200 can provide extended shots of a scene with a given frame latency (e.g., 3 frames latency) sufficient to allow the shots of the scene to be visualized during film production with little or no visual problems. The compositing engine 216 uses a data management and timecode scheme that enables precise alignment of non-frame data from various and dissimilar data sources (e.g., scene capture devices) to video frames.
[0077] Figure 3 is a block diagram of an example system 300 for visualizing a CGI video scene. System 300 may be used to provide real-time augmented video data on display 302, such as during filming. The augmented video data may correspond to augmented video data 110 as shown in Figure 1 or augmented video data 212 as shown in Figure 2. Therefore, the example in Figure 3 may refer to the examples in Figures 1 and 2. In the example in Figure 3, display 302 shows a composite video frame of augmented video data (at a certain point in time) that includes an image of scene 304 captured by video camera 306, with a digital asset 308 (e.g., a CGI asset). A computing platform 310 may be used to provide augmented video data with one or more images incorporating the digital asset. The composite video frame on display 302 is illustrative, and in other examples, it may include any number of digital assets or may not include any digital assets at all (e.g., when the digital asset moves out of the scene or the field of view of video camera 306). Computing platform 310 may correspond to computing platform 200, as shown in Figure 2.
[0078] For example, a computing platform 310 may receive scene-related data 312, which may include one or more scene-related information frames, such as those disclosed herein with respect to Figures 1-2. For example, the scene-related data may include rig tracking data 314, which may be provided by a rig tracking system (or device) 316. The rig tracking system 316 may be used to track the rotation and position of a rig on which a video camera 306 is mounted. Examples of rigs may include tripods, handheld rigs, shoulder rigs, overhead rigs, POV rigs, camera dollies, sliders, rigs, camera cranes, camera stabilizers, Snorricams, etc. The rig tracking data 314 may characterize the rotation and position of the rig. The rig tracking data may also include timestamps indicating when the rotation and position of the rig were captured.
[0079] In some examples, scene-related data 312 may include body data 318 that can be provided by a body tracking system (or device) 320. The body tracking system 320 may be used to track the location of a person in scene 304 in a body tracking coordinate system. The body data 318 may specify the location information of the person in the body tracking coordinate system. In some examples, the body data 318 may specify rotation and / or joint information for each joint of the person. The body data 318 may also specify the position and orientation of the person in the body tracking coordinate system. The body tracking system 320 may provide body data 318 with a timestamp indicating the time when the body movement was captured. Examples of body tracking systems may include motion capture tools, video camera systems (e.g., one or more cameras, such as a mobile phone, used to capture body movement), etc. In an example where the body tracking system 320 is implemented as a video camera system, the body movement captured by the video camera system may be extracted from the image and provided as or as part of the body data 318. Furthermore, in an example where the body tracking system 320 is implemented as a motion capture tool, the motion capture tool may be time-synchronized with a video camera 306 or multiple video cameras 306 that capture scene 304, and the body data 318 may include a timecode indicating the time when the body movement was captured. In some examples, face data 322 may be used to animate the movement of a digital asset 308, for example, the body of the digital asset 308.
[0080] In some cases, scene-related data 312 includes face data 322 that may be provided by a face tracking system 324. The face data 322 may characterize the movement of a person's face. The face tracking system 324 may be implemented as a face motion capture system. In some cases, the face tracking system 324 includes a mobile device (e.g., a mobile phone) with software for processing images captured by the mobile device's camera to determine facial features and / or the movement of a person. The face tracking system 324 may provide face data 322 with a timestamp indicating the time when the movement of a person's face was captured. The face data 322 may be used to animate the facial expressions of a digital asset 308, for example, the facial expressions of the digital asset 308.
[0081] In some examples, scene-related data 312 includes prop data 326 which may be provided by a prop tracking system (or device) 328. The prop tracking system 328 may be used to track one or more props in scene 304. The prop tracking system 328 may provide prop data 326 with a timestamp indicating the time when the prop movement was captured. A prop can be any asset that you wish to track in space. This could include an initial pose for an animated character, such as a CGI asset, a sword and shield held in someone's hands as they move, or a piece of furniture that is stored away, all located within the scene.
[0082] In some examples, the computing platform 310 may receive digital asset data 330 that may correspond to digital asset data 236, as shown in Figure 2. For example, the digital asset data 330 may include CGI asset data and related information for animating one or more CGI assets (e.g., digital asset 308) in a scene such as scene 304. In some examples, the digital asset data 330 may include controller data 332 provided by the asset controller 334. In one example, the asset controller 334 is an input device, such as a controller. The asset controller 334 may be used to control the movement and facial expressions of the digital asset 308 in scene 304. For example, the asset controller 334 may output controller data 332 that characterizes the movement of the face and / or body implemented by the digital asset 308. The controller data 332 may include a timecode indicating the time when the movement of the face and / or body was captured. The controller data 332 may be used to animate the digital asset 308 (e.g., by the computing platform 310). In some examples, controller data 332 may indicate the start and end times for a CGI action (e.g., movement and / or rotation). In some examples, multiple asset controllers may be used. For example, a first asset controller may be used to manipulate the torso of digital asset 308, another asset controller for the arms of digital asset 308, a further asset controller for manipulating the movement of the fingers of digital asset 308, another asset controller device for manipulating the lips of digital asset 308, and so on. Controller data 332 from each of the asset controllers may be provided to the computing platform 310 for processing according to the examples disclosed herein. Thus, multiple different asset controllers may be used to control different parts of digital asset 308, enabling different movements / movements of the digital asset to be superimposed on a single character.
[0083] In the example in Figure 3, the video camera 306 provides video data 336 with one or more video frames containing images of scene 304. Each video frame may contain timecode. If multiple digital video cameras are used to capture scene 304, the digital video cameras may be time-synchronized as described herein. In some examples, if timecode is not present, timecode may be generated from the computing platform 310 clock. This still allows synchronization, but may be less accurate than when using a dedicated timecode device. In some examples, one or more video frames of the video data 336 may include video frame 112 as shown in Figure 1 and / or video frame 228 as shown in Figure 2.
[0084] Computing platform 310 may be implemented in the same or similar manner as computing platform 200 to process video data 336, scene-related data 312, and / or digital asset data 330 to provide augmented video data rendered on display 302 in the example in Figure 3. In some examples, scene-related data 312 includes scene-related information frame 114 shown in Figure 1 and / or scene-related information frame 234 shown in Figure 2. Computing platform 310 may provide augmented shots of a scene with a given frame latency (e.g., 3 frames latency) sufficient so that shots of a scene can be visualized during film production with little or no visual problems. By configuring computing platform 310 with data manager 142 as shown in Figure 1, computing platform 310 is enabled to accurately align non-frame data from various and dissimilar sources, such as rig tracking system 316, body tracking system 320, face tracking system 324, and asset controller 334, to video frames.
[0085] In some examples, rig data 314, body data 318, face data 322, prop data 326, controller data 332, and video data 336 are provided through a network 335, such as network 253, as shown in Figure 1. In other examples, some or all of the data 314, 318, 322, 326, 332, and 336 are provided through network 335. In some examples, camera 306 may be connected using a physical cable to a computing platform 310, which may be considered external or not part of network 335. In some examples, video camera 306 and systems 316, 320, 324, 328, and asset controller 334 may be connected to computing platform 310 using network 335, which may include wired and / or wireless connections. For example, to transmit high-resolution video data to the computing platform 310 at a high rate, a high-bandwidth cable may be used to connect the HDMI® output or SDI output of the video camera 306 to the computing platform 310. Examples of HDMI® cables include, but are not limited to, HDMI 2.1 cables with a bandwidth of 48 Gbps for carrying resolutions up to 10K at frame rates up to 120 fps (8K60 and 4K120). Examples of SDI cables include, but are not limited to, DIN 1.0 / 2.3 to BNC female adapter cables. An external or internal hardware capture card (such as an internal PCI Express Blackmagic Design DeckLink SDI Capture Card or an external unit Matrox MAXO2 Mini Max for Laptops, supporting various video inputs such as HDMI® HD 10-bit, Component HD / SD 10-bit, Y / C (S-Video) 10-bit, or Composite 10-bit) may be used to provide video data 336 to the interface of the computing platform 310 (for example, one of the I / O interfaces 226 as shown in Figure 2).
[0086] In some examples, system 300 may include a server 337 that can be connected to a network 335, and rig data 314, body data 318, face data 322, prop data 326, controller data 332, and video data 336 may be routed through server 337. For example, server 337 may communicate with each system 316, 328, 320, 324 and / or asset controller 334 and route the data as messages to the computing platform 310. Thus, in some examples, the server may use a publisher / subscriber paradigm where each system 316, 328, 320, 324 and / or asset controller 334 is the publisher and the computing platform is the subscriber. In a further example, server 337 may be implemented on a device such as a laptop.
[0087] In some cases, the computing platform 310 may be used for post-production. For example, pre-recorded video data 338 with video frames may be processed in the same or similar manner as video data 336 with other scene-related information (data / frames) to provide augmented video data. For example, the computing platform 310 may be used to re-animate live-action footage and to insert new digital assets (not used during production) into video footage (represented by the pre-recorded video data 338). In some cases, the computing platform 310 may be used to modify the actions and / or behavior of inserted digital assets (e.g., digital asset 308) that are first used during production.
[0088] For example, computing platform 310 may include a virtual environment builder 340 that can recreate or regenerate a 3D model 342 representing a scene captured by one or more video cameras that are the source of pre-recorded video data 338, based on environmental data 344 captured about the scene. Environmental data 146 may be provided by one or more environmental sensors, such as those disclosed herein. Input device 346 may be used to provide commands / actions to model-up data 348 to adjust one or more virtual features of the 3D model 342. One or more virtual features may include the position / orientation of each virtual camera representing the video cameras that are the source of the pre-recorded video data 338, the position / orientation of digital assets in the 3D model 342, and other types of virtual features. Computing platform 310 may include a digital compositor 350 that, in some cases, can be implemented similarly to the digital compositor 108 shown in Figure 1. The digital compositor 350 can provide augmented video data based on the 3D model 342 and pre-recorded video data 338. Thus, in some cases, the computing platform 310 can be used in post-production to allow reshooting of scenes by adjusting the 3D model 342, resulting in the modification or adjustment of VFX (if available in the video data 338), or in some cases, the augmented video data can be inserted into the video data 338.
[0089] In light of the structural and functional features described above, the method examples are better understood by referring to Figures 4-5. For the sake of brevity, the method examples in Figures 4-5 are shown and described as being performed sequentially, but it is understood and recognized that the examples are not limited by the order in which they are shown. Some actions may occur multiple times and / or simultaneously in other examples, in a different order than those shown and described herein. Furthermore, it is not necessary for all described actions to be performed in order to implement the method.
[0090] Figure 4 shows an example of method 400 for providing a composite video frame. Method 400 can be implemented by a composite engine 100 as shown in Figure 1, or a composite engine 216 as shown in Figure 2. Therefore, the example in Figure 4 may refer to the examples in Figures 1 to 3. Method 400 may begin in 402, in which a video frame (e.g., one of the video frames 112 as shown in Figure 1) among video frames stored in a cache memory space (e.g., cache memory space 104 as shown in Figure 1) is identified (e.g., by a video frame retriever 120 as shown in Figure 1) based on video frame latency data (e.g., video frame latency data 124 as shown in Figure 1). The video frame latency data may specify the number of the video frame stored in the cache memory space before the video frame is selected. In 404, a scene-related information frame (e.g., one of the scene-related information frames 114 as shown in Figure 1) among scene-related information frames may be identified (e.g., by a scene-related frame retriever 126 as shown in Figure 1) based on the timecode of the video frame. In 406, a composite video frame (for example, a part of the extended video data 110, as shown in Figure 1) can be provided based on the video frame and the scene-related information frame.
[0091] Figure 5 is an example of method 500 for providing augmented video data during scene pre-visualization. Method 500 can be implemented by the compositing engine 100, as shown in Figure 1, or by the compositing engine 216, as shown in Figure 2. Thus, the example in Figure 5 may refer to the examples in Figures 1 to 4. Method 400 may begin in 502 by identifying a video frame (e.g., one of the video frames 112, as shown in Figure 1) from among the video frames provided by a video camera (e.g., camera 306, as shown in Figure 1) representing a scene in film production (e.g., scene 304, as shown in Figure 1), based on video frame latency data (e.g., video frame latency data 124, as shown in Figure 1). The video frame latency data may specify the number of a video frame that is stored in the memory of the computing platform before the video frame is selected. In 504, a scene-related information frame (e.g., one of the scene-related information frames 114 as shown in Figure 1) among the scene-related information frames provided by a scene capture device (e.g., one of the systems 316, 320, 324, and / or 328, and / or asset controller 334 as shown in Figure 1) may be identified (e.g., by the scene-related frame retriever 126 as shown in Figure 1) based on an evaluation of the timecode of the video frame against the timestamp of the scene-related information frame. The timestamp for the scene-related information frame may be generated based on a frame delta value representing the amount of time between each scene-related information frame provided by each of the scene capture devices. In 506, extended video data (e.g., extended video data 110 as shown in Figure 1) with a digital asset (e.g., digital asset 308 as shown in Figure 1) based on the scene-related information frame and video frame may be provided (e.g., for rendering on an output device such as the display 219 as shown in Figure 2 or the display 302 as shown in Figure 3).
[0092] While this disclosure has described several exemplary embodiments, it will be understood by those skilled in the art that various modifications may be made without departing from the spirit and scope of the invention, and that equivalents may be substituted for those elements. Furthermore, it will be understood by those skilled in the art that many modifications are made to adapt specific equipment, situations, or materials to the embodiments of this disclosure without departing from their essential scope. Thus, the invention is not limited to the specific embodiments disclosed or the best mode assumed for carrying out this invention, and the invention is intended to include all embodiments belonging to the appended claims. Furthermore, references in the appended claims to an apparatus or system, or component of an apparatus or system, adapted, arranged, capable, configured, enabled, operable, or operating to perform a particular function, encompass that apparatus, system, or component insofar as it is adapted, arranged, capable, configured, enabled, operable, or operating in such a way, regardless of whether its or their particular function is activated, turned on, or unlocked or not.
[0093] In view of the structural and functional descriptions described above, those skilled in the art will understand that some of the embodiments can be embodied as methods, data processing systems, or computer program products. Accordingly, these parts of the embodiments may take the form of hardware embodiments as a whole, software embodiments as a whole, or embodiments combining software and hardware, such as those shown and described with respect to the computer system in Figure 6. Furthermore, some of the embodiments may be computer program products on a computer-readable storage medium having computer-readable program code on that medium. Any non-temporary, tangible storage medium processing structure may be used, including but not limited to static and dynamic storage devices, hard disks, optical storage devices, and magnetic storage devices, that exclude any medium that is not eligible for patent protection under § 101 of the United States Patent Act (such as the propagation of electrical or electromagnetic signals themselves). For example, and not to limit, computer-readable storage media may include, as needed, semiconductor-based circuits or devices or other ICs (e.g., field-programmable gate arrays (FPGAs) or ASICs), hard disks, HDDs, hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tapes, holographic storage media, solid-state drives (SSDs), RAM drives, secure digital cards, secure digital drives, or other suitable computer-readable storage media, or any combination of two or more of these. Computer-readable non-temporary storage media may be volatile, non-volatile, or a combination of volatile and non-volatile, as needed.
[0094] Furthermore, specific embodiments are described herein with reference to block diagrams of methods, systems, and computer program products. It will be understood that the blocks in the diagrams and combinations of blocks within the diagrams can be implemented by computer executable instructions. These computer executable instructions are provided to one or more processors of a general-purpose computer, a dedicated computer, or other programmable data processing device (or combination of devices and circuits) so that instructions executed through the processors can create a machine that implements the functions specified within the blocks or combinations of blocks.
[0095] Furthermore, these computer-executable instructions may be stored in computer-readable memory, and these instructions can instruct a computer or other programmable data processing device to function in a particular manner, so that the instructions stored in computer-readable memory produce a product containing instructions that implement functions specified in a block or multiple blocks of a flowchart. Also, computer program instructions can be loaded into a computer or other programmable data processing device to execute a series of operational steps on the computer or other programmable device, thereby producing a computer implementation process, so that the instructions executed on the computer or other programmable device provide steps for implementing functions specified in a block or multiple blocks of a flowchart.
[0096] In this regard, Figure 6 shows an example of a computing system 600 that may be employed to carry out one or more embodiments of the present disclosure. The computing system 600 may be implemented on one or more general-purpose network computer systems, embedded computer systems, routers, switches, server devices, client devices, various intermediate devices / nodes, or standalone computer systems. In addition, the computing system 600 may be implemented on various mobile clients, such as personal digital assistants (PDAs®), laptop computers, pagers, and the like, provided that they have sufficient processing power. In other examples, the computing system 600 may be implemented on dedicated hardware.
[0097] The computing system 600 includes a processing unit 602, memory 604, and a system bus 606 that connects various system components, including memory 604, to the processing unit 602. Dual microprocessors and other multiprocessor architectures can also be used as the processing unit 602. In some examples, the processor 602 corresponds to the processor unit 202 as shown in Figure 2. The system bus 606 can be one of several types of bus structures, including a memory bus or memory controller, peripheral bus, and local bus, using any of various bus architectures. Memory 604 includes read-only memory (ROM) 610 and random access memory (RAM) 612. A basic input / output system (BIOS) 614 may reside on the ROM 610, containing basic routines that help transmit information between elements within the computing system 600.
[0098] The computing system 600 may include a hard disk drive 616, a magnetic disk drive 618 for reading from or writing to, for example, a removable disk 620, and an optical disk drive 622 for reading from, for example, a CD-ROM disk 624 or reading from or writing to other optical media. The hard disk drive 616, the magnetic disk drive 618, and the optical disk drive 622 are connected to the system bus 606 by a hard disk drive interface 626, a magnetic disk drive interface 628, and an optical drive interface 630, respectively. The drives and associated computer-readable media provide non-volatile storage of data, data structures, and computer-executable instructions for the computing system 600. While the above description of computer-readable media refers to hard disks, removable magnetic disks, and CDs, other types of computer-readable media, such as various forms of magnetic cassettes, flash memory cards, digital video discs, and the like, may also be used in the operating environment. Furthermore, any such media may contain computer-executable instructions for implementing one or more parts of the embodiments shown and described herein. The computing system 600 may include a GPU interface 658 that can be used for interfacing with the GPU 660. In some examples, the GPU 660 may correspond to a GPU 218 as shown in Figure 1. In further examples, the computing system 600 includes the GPU 660.
[0099] A drive and RAM 610 may store an operating system 632, one or more application programs 634, other program modules 636, and multiple program modules including program data 638. For example, RAM 610 may include the synthesis engine 100 shown in Figure 1, or the synthesis engine 216 shown in Figure 1. In further examples, program data 632 may include video frames (e.g., video frame 112 shown in Figure 1), scene-related information frames (e.g., scene-related information frame 114 shown in Figure 1), digital asset data (e.g., digital asset data 116 shown in Figure 1), video frame latency data (e.g., video frame latency data 124 shown in Figure 1), and other data disclosed herein.
[0100] A user may input commands and information to the computing system 600 through one or more input devices 640, such as a pointing device (e.g., mouse, touchscreen), keyboard, microphone, joystick, gamepad, scanner, and the like. For example, one or more input devices 640 may be used to provide video frame latency data, as disclosed herein. These and other input devices are often connected to the processing unit 602 through corresponding port interfaces 642 coupled to the system bus, but may be connected by other interfaces such as parallel ports, serial ports, or Universal Serial Bus (USB). One or more output devices 644 (e.g., displays, monitors, printers, projectors, or other types of display devices) are also connected to the system bus 606 via an interface 646, such as a video adapter.
[0101] The computing system 600 may operate in a network environment using logical connections to one or more remote computers, such as remote computer 648. The remote computer 648 may be a workstation, computer system, router, peer device, or other common network node, and typically includes many or all of the elements described for the computing system 600. The logical connections schematically shown in 650 may include local area networks (LANs) and wide area networks (WANs). When used in a LAN network environment, the computing system 600 may be connected to the local network via a network interface or adapter 652. When used in a WAN network environment, the computing system 600 may include a modem or be connected to a communication server on the LAN. A modem, which may be internal or external, may be connected to the system bus 606 via an appropriate port interface. In a network environment, application programs 634 or program data 638 described for the computing system 600 or a portion thereof may be stored in a remote memory storage device 654.
[0102] While this disclosure has described several exemplary embodiments, it will be understood by those skilled in the art that various modifications may be made without departing from the spirit and scope of the invention, and that equivalents may be substituted for those elements. Furthermore, it will be understood by those skilled in the art that many modifications are made to adapt specific equipment, situations, or materials to the embodiments of this disclosure without departing from their essential scope. Thus, the invention is not limited to the specific embodiments disclosed or the best mode assumed for carrying out this invention, and the invention is intended to include all embodiments belonging to the appended claims. Furthermore, references in the appended claims to an apparatus or system, or component of an apparatus or system, adapted, arranged, capable, configured, enabled, operable, or operating to perform a particular function, encompass that apparatus, system, or component insofar as it is adapted, arranged, capable, configured, enabled, operable, or operating in such a way, regardless of whether its or their particular function is activated, turned on, or unlocked or not.
[0103] In this specification, the term “or” is intended to be inclusive, not exclusive. Unless otherwise specified, “X adopts A or B” is intended to mean any natural and inclusive permutation. That is, “X adopts A or B” is satisfied if X uses A; X uses B; or X uses both A and B. In this specification, the terms “example” and / or “exemplary” are used to indicate one or more features as examples, cases, or illustrations. The subject matter disclosed herein is not limited by such examples. In addition, any aspect, feature and / or design disclosed herein as “example” or “exemplary” is not necessarily intended to be construed as preferred or advantageous. Similarly, any aspect, feature and / or design disclosed herein as “example” or “exemplary” is not intended to exclude equivalent embodiments (e.g., features, structures, and / or methods) known to those skilled in the art.
[0104] Those skilled in the art will understand that it is impossible to describe every conceivable combination of the various features disclosed herein (e.g., components, products and / or methods), and will recognize that many further combinations and permutations of the various embodiments disclosed herein are possible and conceivable. Furthermore, in this specification, the terms “includes,” “has,” “possesses,” and / or similar are intended to be comprehensive in the same manner as the term “comprising,” when used as a transitional clause in a claim.
[0105] The examples described above are merely illustrations. Naturally, it is not possible to describe all possible combinations of components or methods, but those skilled in the art will recognize that many further combinations and permutations are possible. Therefore, this disclosure is intended to encompass all such changes, modifications, and variations that fall within the scope of this application, including the appended claims. Where this disclosure or a claim lists “a,” “an,” “first,” or “another” element or its equivalent, it should be interpreted as including one or more such elements, and not requiring or excluding two or more such elements. As used herein, the term “including” means including but not limited to, and the term “includes” means including but not limited to. The term “based on” means “based at least in part on.”
[0106] The terms "approximately" and "roughly" can be used to include a number that can vary without changing its fundamental function. When used with a range, "approximately" and "roughly" also disclose a range defined by the absolute values of the two endpoints; for example, "approximately 2 to approximately 4" also discloses the range "2 to 4". In general, the terms "approximately" and "roughly" can refer to plus or minus 5 to 10% of the indicated number.
Claims
1. In the stage of receiving scene-related information frames about the scene to be filmed, each scene-related information frame includes a timestamp and is provided by the device; In the step of receiving a video frame representing the aforementioned scene, the video frame is provided by a video camera; A step of selecting a video frame from among the video frames based on video frame latency data; A step of updating the timestamp of each of the scene-related information frames based on the timing synchronization factor; The step of selecting a scene-related information frame from among the scene-related information frames accompanied by the updated timestamp; and A step of generating an extended video frame with digital assets based on the video frame and the scene-related information frame. A method for providing this.
2. The method according to claim 1, wherein the video frame is provided by the video camera at a different frequency from the scene-related information frame.
3. The method according to claim 1, further comprising the step of extracting a timestamp of the selected video frame, wherein the scene-related information frame is selected based on the timestamp of the selected video frame.
4. The method according to claim 1, wherein the video frame is further selected based on a video frame request.
5. The method according to claim 1, wherein a digital compositor generates the extended video frame.
6. The method according to claim 1, wherein a plurality of video frames, including the aforementioned video frame, are stored in a cache memory space, and the video frame latency data specifies the number of the plurality of video frames stored in the cache memory space before the video frame is selected from the cache memory space.
7. The method according to claim 6, wherein the last video frame stored among the video frames in the cache memory space is selected based on the video frame latency data.
8. The method according to claim 6, wherein the first, second, and third video frames are stored in the cache memory space, and the third video frame is selected from the cache memory space.
9. The step of creating a virtual environment representing the scene, wherein the virtual environment includes computer-generated image (CGI) elements for the digital asset and a virtual camera for the video camera, and the scene-related information frame specifies the position and / or orientation of the video camera at a given point in time; A step of updating the virtual environment by updating the position of the virtual camera and the CGI elements in the virtual environment based on the selected scene-related information frame; and The step of generating the extended video frame in response to updating the virtual environment. The method according to claim 1, further comprising:
10. The method according to any one of claims 1 to 9, further comprising the step of determining the network latency of the network, wherein the scene-related information frame is provided over the network.
11. The method according to claim 10, wherein the timing synchronization factor specifies the network latency of the network through which the scene-related information frame is provided.
12. The method according to claim 11, wherein the timing synchronization factor includes a frame delta value representing the amount of time between the scene-related information frames.
13. The step of partitioning the cache memory space into partitions including a first partition and a second partition; The steps of storing the video frame in the first partition; and Steps to store the scene-related information frame in the second partition. The method according to any one of claims 1 to 9, further comprising:
14. The method according to claim 13, wherein the partition includes a third partition, and the method further comprises the step of storing the digital assets in the third partition.
15. The method according to claim 13, wherein the first partition includes even and odd video frame locations, and the video frames are stored in one of the even and odd video frame cache locations based on the respective frame numbers assigned to each of the video frames.
16. The method according to any one of claims 1 to 9, wherein the video frame latency data is provided based on user input.
17. Receive scene-related information frames about a scene, each scene-related information frame containing a timestamp; Receiving a video frame representing the aforementioned scene; Selecting video frames of the video frame based on video frame latency data; Updating the timestamp of each of the scene-related information frames based on the timing synchronization factor; Selecting a scene-related information frame from among the scene-related information frames accompanied by the updated timestamp; and To generate an extended video frame with digital assets based on the aforementioned video frame and the aforementioned scene-related information frame. One or more computing platforms configured to perform this task A system equipped with these features.
18. The system according to claim 17, wherein the timing synchronization factor specifies the network latency of the network through which the scene-related information frames are provided, and a frame delta value representing the amount of time between each scene-related information frame of the scene-related information frames provided by the device.
19. A computer program having machine-readable instructions that can be executed by a processor, wherein the machine-readable instructions are: The device provides a scene-related information frame about the scene being captured, each scene-related information frame containing a timestamp; The step of receiving a video frame representing the aforementioned scene; A step of selecting a video frame of the video frame based on video frame latency data; A step of updating the timestamp of each of the scene-related information frames based on the timing synchronization factor; A step of identifying a scene-related information frame among the scene-related information frames accompanied by the updated timestamp; and A step of providing an extended video frame with digital assets based on the video frame and the scene-related information frame. A computer program programmed to cause the processor to perform a method comprising the following.
20. The method further comprises the step of partitioning the cache memory space into partitions including a first partition, a second partition, and a third partition, wherein: The video frame is stored in the first partition. The aforementioned scene-related information frame is stored in the second partition. The aforementioned digital assets are stored in the third partition. The computer program according to claim 19.