Camera tracking via dynamic fluoroscopy

Visual odometry-based camera tracking in virtual production systems addresses inaccuracies by comparing captured and predicted images, enhancing precision and reducing the need for external markers.

JP2026004415APending Publication Date: 2026-01-14NANTSTUDIOS LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025163524
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-20
Filing Date
2025-09-30
Publication Date
2026-01-14

AI Technical Summary

Technical Problem

Existing camera tracking systems in virtual production environments are prone to errors and require additional equipment like markers, leading to inaccuracies in rendering backgrounds and characters on digital screens, especially in crowded stages.

Method used

A method that uses visual odometry to compare captured images with predicted images from a display panel, determining the physical camera's location by analyzing differences between the two, allowing for accurate tracking without external markers.

Benefits of technology

This approach reduces registration errors and improves tracking precision by directly measuring image distortions, enabling more accurate rendering of virtual scenes on display surfaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026004415000001_ABST
    Figure 2026004415000001_ABST
Patent Text Reader

Abstract

To synchronize the position change of a camera for photographing the live action of an actor with a video displayed on a background digital screen in virtual production.SOLUTION: The computer system identifying a first position of the physical camera corresponding to a first time period, rendering a first virtual scene for the first time period, projecting the first scene onto a display surface to determine a first rendered image for the first time period, and receiving a first camera image of the display surface from the camera during the first time period; Determining a first corrected position of the camera by comparing the first rendered image and the first camera image, predicting a second position of the camera corresponding to a second time period, rendering a second virtual scene for the second time period, and projecting the second virtual scene onto the display surface to determine a second rendered image for the second time period.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 257,930, entitled "CAMERA TRACKING VIA DYNAMIC PERSPECTIVES," filed October 20, 2021, which is incorporated by reference in its entirety for all purposes. [Background technology]

[0002] In one production technique, a digital screen can be used as a backdrop for the physical scene being filmed. Large digital screens allow arbitrary backgrounds and characters to be inserted by live action actors. The particular image displayed on the screen can change over time and in synchronization with the physical production camera. Images are displayed on the screen based on a determined position (e.g., by markers on the physical production camera and a separate device that determines the position of the physical production camera by using the markers). Such tracking can be problematic on a crowded stage and can be prone to errors. Summary of the Invention [Means for solving the problem]

[0003] Some embodiments may use a captured image of a display panel to determine the location of a physical camera. The display panel may display a rendered image based on the predicted location of the physical camera. If the display panel geometry and the rendered image are known, the location of the physical camera can be determined through a comparison of the rendered image with an image captured by the physical camera. If the physical camera is in the predicted location, the captured image should reflect the rendered image, and the difference between the two images can be used to calculate the actual location of the physical camera.

[0004] In some embodiments, a technique for performing camera tracking in a virtual production environment may include identifying a first position of a physical camera corresponding to a first time period or time point. The technique may include rendering a first virtual scene for the first time period. The technique may include projecting the first virtual scene onto a display surface to determine a first rendered image for the first time period. The projection of the first virtual scene corresponds to a first position of the physical camera. The technique may include receiving a first camera image on the display surface. The first image is acquired using the physical camera during the first time period. The technique may include determining a first corrected position of the physical camera by comparing the first rendered image and the first camera image. The technique may include predicting a second position of the physical camera corresponding to a second time period using the first corrected position. The technique may include rendering a second virtual scene for the second time period. The technique may include projecting the second virtual scene onto the display surface to determine a second rendered image for the second time period. The projection of the second virtual scene corresponds to a second position of the virtual camera.

[0005] In some embodiments, predicting the second position uses information from a predetermined choreography file of the final footage to be captured.

[0006] In some embodiments, the first position of the physical camera is identified by using an initial image from the physical camera of the display screen.

[0007] In some embodiments, identifying the first position of the physical camera includes storing a model that maps images to physical positions and inputting the initial image into the model.

[0008] These and other embodiments of the present disclosure are described in detail below. For example, other embodiments are directed to systems, devices, and computer-readable media related to the methods described herein.

[0009] A better understanding of the nature and advantages of some embodiments of the present disclosure may be obtained with reference to the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0010] [Figure 1] 1 illustrates an overhead perspective view of a virtual production studio in accordance with at least one embodiment. [Figure 2] 1 illustrates an overhead perspective view of a virtual scene rendered onto a display surface according to at least one embodiment. [Figure 3] 3 shows a simplified diagram 300 of a projection of a three-dimensional object onto a two-dimensional surface in accordance with at least one embodiment. [Figure 4] 1 illustrates an overhead perspective view of a virtual production studio with registration error in accordance with at least one embodiment. [Figure 5A] FIG. 1 illustrates an overhead perspective view of a camera with registration error along line of sight in accordance with at least one embodiment. [Figure 5B] 1 illustrates an overhead perspective view of a camera with registration error orthogonal to line of sight in accordance with at least one embodiment. [Figure 6] 1 illustrates an overhead perspective view of a virtual production studio with concatenated expected and actual camera positions, but with cameras at different orientations, according to at least one embodiment. [Figure 7] 1 illustrates an overhead perspective view of a virtual production studio with registration errors in accordance with at least one embodiment. [Figure 8] 1 illustrates a flowchart of a method for predicting a new camera position based on a previous camera position according to at least one embodiment. [Figure 9] 1 illustrates a flowchart of a technique for implementing camera tracking within a virtual manufacturing environment in accordance with at least one embodiment. [Figure 10] 1 illustrates a diagram of an exemplary computer system according to one embodiment of the present disclosure in accordance with at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Virtual production is a visual effects technology that allows live-action sequences to be filmed with physical cameras in front of a computer-generated environment. Filming during virtual production occurs in a large studio with a center stage surrounded (e.g., at 180 degrees, 270 degrees, 4π radians, etc.) by a display surface composed of an array of high-definition display panels. Actors and physical props are filmed within the center stage while a virtual scene (e.g., a computer-generated backdrop) is displayed on the bounding display surface. The display surface panels allow for realistic lighting and reflections of actors and objects within the center stage. For example, reflections of a virtual scene in a mirror can be captured during filming in a virtual production studio, while a similar effect with traditional filming techniques may require adding reflections in post-production.

[0012] Because the camera's position (angle and distance) may move, the rendered image needs to be updated based on such movement. Some embodiments may determine the camera's position so as to render an appropriate image of the camera's current position. Some embodiments may determine the position without requiring a separate system to track the camera's movement (e.g., by using markers attached to the camera). Instead, some embodiments may compare the image measured by the camera with a predicted image from a predicted position. The error determined from the comparison (e.g., the error from the predicted position at a particular time) provides information about where the camera is located.

[0013] In this way, the camera can be tracked with less equipment and more accurately than external tracking systems, since the largest visual distortion errors that cause the largest differences between images are measured with the highest precision. Such errors would not persist in embodiments that use visual odometry, since the estimated position is updated in a manner consistent with the errors between images.

[0014] I. Virtual Creation Virtual production is a filmmaking technique that combines live-action scenes with computer-generated imagery during filming, without the need to add the computer-generated imagery in post-production. During virtual production, a computer-generated three-dimensional environment is created and used to render and show a virtual scene on a display screen that borders center stage. The virtual scene is rendered based on an intended perspective (perhaps relative to a camera perspective or an actor perspective) so that the two-dimensional image on the display screen appears three-dimensional when viewed from the intended perspective.

[0015] To accurately convey the simulated environment bounded by a studio or stage, a virtual scene may be rendered relative to a single viewpoint. A single viewpoint is selected because the virtual scene simulates a three-dimensional environment, but the virtual scene is rendered on a two-dimensional display screen. To render a convincing three-dimensional environment on a two-dimensional screen, objects in three-dimensional space are rendered with their relative size and location determined by the viewpoint (e.g., three-dimensional (3D) projection, camera viewpoint, etc.). The virtual scene may be projected onto a two-dimensional surface corresponding to the screen to determine the image displayed. The exact image displayed may depend on the camera position. The virtual scene may be a computer-generated three-dimensional environment generated by using a game engine (e.g., Unreal Engine, UnReal, etc.). Projection is the operation of mapping (e.g., geometric projection) a 3D virtual scene onto a two-dimensional surface corresponding to the position of the display surface. Generally, objects appear smaller as their distance from the viewpoint increases. If the virtual scene is viewed from an unintended viewpoint, the rendered objects may appear distorted because the size, proportions, and position of the objects do not match how the objects appear from the unintended viewpoint.

[0016] Additionally, camera motion can cause image distortion because while the environment appears three-dimensional from a set viewpoint, there is no parallax (e.g., displacement of objects as seen from a changing viewpoint) in the two-dimensional image rendered on the display surface. To compensate for the lack of parallax, simulated parallax conveys depth by updating the computer-generated image based on the location and orientation of the physical camera. To properly simulate parallax, the physical camera is tracked to determine its orientation relative to the display surface.

[0017] In an illustrated example, a physical set is constructed within a central stage within a virtual production studio. Computer-generated images are provided to a rendering system within the studio, which is connected to a display surface including an array of light-emitting diode (LED) panels. In existing implementations, the studio may also include a camera tracking system (e.g., Optitrack) that provides an initial location for the physical camera. The location determined by the camera tracking system is used to render a first virtual scene shown on the display surface. After viewing the rendered scene, the director may make adjustments to the lighting and position of buildings within the environment. Once the director approves the scene, actors are placed within the stage and the scene is filmed. As the camera operator moves the camera during filming, the camera's location can be tracked via visual odometry by comparing captured images with the rendered virtual scene relative to the camera's estimated position (e.g., by using visual simultaneous localization and mapping (vSLAM) or techniques when a map of the environment is known in advance). Simultaneous localization and mapping (SLAM) can map an area while tracking the location of entities within the area. In some implementations, a map of the environment may be determined in advance, and new measurements may be compared to the existing 3D map. As will be described later, such embodiments may include using machine learning models trained within the physical environment and using the same content displayed on a screen (such content may be dynamic).

[0018] Visual odometry can track features through successive camera frames. Comparison of features between successive frames can be used to triangulate the location of features within a 3D environment. Successive camera frames can also be used to estimate the pose of the camera within the 3D environment. Approximation methods can be used and include particle filters, extended Kalman filters, covariance intersection methods, and GraphSLAM.

[0019] FIG. 1 shows an overhead perspective view of a virtual production studio 100 according to at least one embodiment. A physical camera 102 is in front of a curved display surface 104. In some implementations, the display surface 104 may border a central stage 106 and may include an array of display panels (e.g., light-emitting diode (LED) panels, organic light-emitting diode (OLED) panels, liquid crystal display (LCD) panels, etc.). Various display panel geometries are contemplated, and in some embodiments, the display surface 104 may occupy a small portion of the surface bordering the central stage 106. For example, the display surface 104 may be a simulated window on a wall. In other implementations, the display surface 104 may occupy up to all of the surface bordering the central stage 106. For example, the display surface 104 may include panels on the walls, ceiling, and floor bordering the central stage 106 (e.g., MegaPixel). In some embodiments, display surface 104 may include panels on the ceiling or floor. Physical cameras 102 may film both center stage 106 and display surface 104. Sets, actors, and practical effects may be positioned and filmed within center stage 106.

[0020] One or more physical cameras, such as physical camera 102, may capture a portion of display surface 104 defined by the camera's field of view 108. Rendering system 110 may be a computer system that generates a virtual scene by using a game engine (e.g., Unreal Engine, UnReal, etc.), and the virtual scene is shown on display surface 104. Rendering system 110 may be connected to display surface 104 by physical connection 112 or wirelessly to display surface 104. The virtual scene may be a dynamically variable three-dimensional environment as opposed to a static image or landscape, and rendering the environment may include generating the portion of the environment that is visible from the current position of the physical camera.

[0021] As will be described in more detail below, physical camera 102 may be communicatively coupled to rendering system 110. In this manner, rendering system 110 may analyze images from physical camera 102 to determine its current position, which may then be used to determine the image to be displayed on display surface 104.

[0022] FIG. 2 shows an overhead perspective view of a virtual production studio 200 according to at least one embodiment. A virtual point 202 is rendered on a curved display surface 204 that bounds a stage area 206. In some implementations, the display surface 204 may include an array or ensemble of display panels, which may include display panels on the ceiling or floor of the virtual production studio or stage. The virtual point 202 may be rendered as part of a virtual scene by generating an apparent virtual point 208 on the display surface 204. The apparent virtual point 208 may be generated such that when the display surface 204 is viewed from an estimated camera position 210, the apparent virtual point 208 may be perceived as a virtual scene including the virtual point 202. The virtual scene including the virtual point 202 may form a virtual environment that may include characters and scenery. In some implementations, the virtual scene may be rendered on the entire display surface or a portion of the display surface. Further illustration of projection is provided next.

[0023] 3 shows a simplified diagram 300 of projecting a three-dimensional object onto a two-dimensional surface (e.g., a screen) according to at least one embodiment. The three-dimensional object 302 (a cube) may be composed of virtual points in a virtual scene (such as virtual points 202 described above). The three-dimensional object may be projected onto a two-dimensional observation surface 304 as a two-dimensional object 306. The two-dimensional observation surface 304 may be a display surface in a virtual production studio, such as display surface 104 described above. Depending on the orientation of the two-dimensional object 306 relative to the two-dimensional observation surface 304, the portion of the three-dimensional object 302 that can be seen from the two-dimensional observation surface 304 is represented as the two-dimensional object 306. The two-dimensional object 306 may be represented as a collection of apparent virtual points.

[0024] Referring back to FIG. 2 , the estimated camera position 210 may be estimated in various ways as described herein. For example, the camera's position at each time instant may be specified to be in a particular location or to have a particular movement (choreography). Such expected positions may be used in conjunction with the difference between the predicted and actual images to determine an offset (also called error or difference) from the expected position. In another example, a machine learning model (e.g., a neural network) may be trained by using images taken at various positions for a given display image at a particular time instant. In this way, the machine learning model may map an image taken at a particular time instant to a certain position. The camera's trajectory may also be determined, thereby enabling prediction of the next position as a time step.

[0025] II. Camera Tracking Issues in Virtual Production The image shown on the display surface of the virtual production studio may be distorted unless it is viewed from an estimated camera position (e.g., an intended viewpoint). The estimated camera position may be a predicted physical camera position within the virtual production studio. Registration error is the difference between the estimated camera position and the actual position of the physical camera, and registration error may result in a distorted display panel image if the image is viewed from the actual position of the physical camera.

[0026] 4 illustrates a virtual production studio 400 having registration errors caused by differences between actual and desired line of sight, according to at least one embodiment. The virtual production studio 400 includes a display surface 402 bounded by a central stage area 404, similar to display surface 104 and central stage area 106 described above. Virtual points 406 that form part of the virtual scene are rendered as apparent virtual points 408. The virtual points 406 are rendered at a predicted camera position 410 from an intended viewpoint.

[0027] Apparent virtual point 408 is shown on display surface 402 along a line of sight 412 between predicted location 410 and virtual point 406. However, physical camera 414 is not located at predicted location 410. Actual line of sight 416 extends from physical camera 414 to apparent virtual point 408. Proper line of sight 418 extends from the physical camera to virtual point 406. Apparent virtual point 408 should have been rendered as proper virtual point 420 to accurately represent virtual point 406 from the perspective of physical camera 414. The difference between apparent virtual point 408 and proper virtual point 420 is the registration error.

[0028] To simulate a real-world environment, as a physical camera moves through the virtual production studio, objects in the computer-generated environment displayed on the display panel should move at different speeds based on the objects' virtual positions relative to the physical camera. For example, as the physical camera traverses along an axis perpendicular to the camera's line of sight, closer objects in the foreground image should move at a faster speed (i.e., parallax) compared to distant background images. In addition, as the physical camera moves along its line of sight, the relative sizes of the foreground and background images should change, with the foreground image growing at a faster rate than the background image.

[0029] Both the estimated and actual camera positions may be three-dimensional positions with six degrees of freedom, and registration errors may occur due to differences between the expected and actual positions along or perpendicular to the line of sight of the physical camera. In some implementations, the display surface is larger than the camera's field of view, and the physical camera may rotate while capturing the display screen. If the estimated and actual camera positions share a center of projection, rotational differences between the estimated and actual camera positions do not cause registration errors because the estimated and actual camera positions are articulated.

[0030] The magnitude of the registration error depends on the distance between a virtual point (e.g., a point in a computer-generated environment) and the apparent virtual point (e.g., a point on a display panel array that depicts the virtual point). Virtual points that are connected to an apparent virtual point will not appear distorted despite the registration error because there is no parallax since the apparent and actual positions of the virtual points are the same regardless of movement. However, as discussed above, the distance between a virtual point and an apparent virtual point can result in a registration error because the apparent position of the virtual point can vary depending on the viewpoint.

[0031] Orientation of registration errors can also contribute to image distortion, with errors orthogonal to the line of sight causing greater distortion than registration errors along the line of sight. External camera tracking techniques are prone to orthogonal registration errors because such systems often include ceiling-mounted tracking cameras or standard markers to determine camera position. External markers on the physical camera are used by the tracked camera to determine camera position by using motion capture techniques. Generally, errors in external camera tracking occur along the axis between the tracked camera and the physical camera because the picture used by the camera tracking system does not allow for easy determination of the depth of the physical camera relative to the tracked camera.

[0032] 5A illustrates a virtual production studio 500 having a discrepancy between the expected and actual positions of a physical camera along the physical camera's line of sight, according to at least one embodiment. A display surface 502 surrounds a central stage area 504, and the features described in connection with FIG. 5A are similar to the related features described above. An apparent virtual point 506 has been drawn along a line of sight 508 extending from an expected position 510 and a virtual point 512. A physical camera 514 is not located at the expected position 510 and is instead located closer to the apparent virtual point 506 along the line of sight 508.

[0033] In this case, line of sight 508 and proper line of sight are the same, and apparent virtual point 506 and proper virtual point 516 are closely spaced, but the object represented by proper virtual point 516 is larger than the object represented by apparent virtual point 506. The registration error is the distance between corresponding pixels of the object constructed by apparent virtual point 506 and corresponding pixels of the object constructed by proper virtual point 516, and the registration error can be small. The registration error can be measured in pixels, virtual distance in the virtual environment, real-world distance in a real-world studio set, or other units as may be convenient for the stakeholders or the underlying game engine.

[0034] Figure 5B illustrates a virtual production studio 501 with a discrepancy between the location of a physical camera and its expected position along an axis perpendicular to the line of sight, according to at least one embodiment. The features described in connection with Figure 5B are similar to the related features discussed above. A display surface 503 surrounds a central stage area 505, and an apparent virtual point 507 is displayed on the display surface 503 along a line of sight 509 extending from an expected position 511 to a virtual point 513. A physical camera 515 is positioned along an axis 517 perpendicular to the line of sight 509. The proper axis 519 of the physical camera 515 extends through the proper virtual point 521 to the virtual point 513, and a registration error is caused by the difference between the proper virtual point 521 and the apparent virtual point 507.

[0035] 6 shows an overhead perspective view of a virtual production studio 600 according to at least one embodiment, in which a physical camera is rotated about an axis extending from the floor to the physical camera relative to its expected location, causing rotational error. A display surface 602 surrounds a central stage area 604, and an apparent virtual point 606 is shown on the display surface 602 along an axis 612 between the virtual point 608 and its expected location 610. A physical camera 614 is positioned at the same location as the expected location 610, but the orientation of the physical camera 614 is rotated relative to the expected camera location. In some implementations, the virtual scene is rendered on a portion of the display surface that is larger than the field of view of the physical camera 614. In these implementations, the physical camera 614 can be rotated without registration error because the apparent virtual point 606 and the actual virtual point can be aligned if the center of projection is the same for the expected location 612 and the physical camera 614.

[0036] Such errors can be addressed by using visual odometry. Specifically, a captured image taken by a physical camera can be compared to a predicted image if the physical camera were in a predicted position. The error between the images provides a position offset relative to the predicted position, thereby providing a measure of the actual or corrected position. This corrected position can be used to predict the position at the next time step.

[0037] III. Determining Position by Using Odometry Visual odometry is the process of determining the location and orientation of a camera from a sequence of images. Visual odometry can be used to determine the three-dimensional motion of a camera within an environment (e.g., Egomotion). The location of a physical camera within an environment can be determined by mapping elements in two-dimensional images to locations in the three-dimensional environment.

[0038] A. Visual Odometry Visual odometry can be used to determine the position of a camera within a three-dimensional environment by tracking changes in images captured by the camera. Features (information about the content of the image) can be detected in successive camera images, and features from successive images can be compared to determine the camera's motion.

[0039] Features may include naturally occurring characteristics in an image, such as edges, corners, or texture. Tracking can be facilitated by easily identifiable features (called fiducials) added to the scene to aid in tracking. Fiducials, also known as landmarks or markers, include point fiducials, typically consisting of 4-5 pixel-wide points on a clear circular background. Alternatively, flat fiducials often consist of black and white grids. Fiducials are useful for feature detection because they can be easy to identify due to their high contrast with the background. Feature recognition can occur through object recognition techniques, such as those disclosed in U.S. Patent Application Publication Nos. 2021 / 0027084, 2019 / 0318195, or 2019 / 0272646, the entire contents of which are incorporated herein by reference.

[0040] The motion of features in successive images can be used to generate optical flow (e.g., the pattern of apparent motion of objects, surfaces, and edges in successive images caused by camera motion) by matching features between images. Optical flow (e.g., optical flow generated using the Lucas-Kanade method) can be used to estimate camera motion by using a Kalman filter to maintain a state estimation distribution. Camera motion can also be estimated by finding geometric properties of features that minimize a cost function based on the re-projection error between two adjacent images (e.g., by mathematical minimization or random sampling).

[0041] As an example, a change in the size of an object (i.e., becoming smaller or larger) may indicate movement toward or away from the screen, with the general shape remaining the same. Apparent movement to the left (e.g., away from normal to the screen) may be determined by objects on the left side of the screen becoming larger while objects become smaller in width. A similar change may be detected for movement to the right.

[0042] B. Mapping images to various locations Visual odometry can be used to track a virtual production camera by comparing images captured by the camera with simulated images (based on what would be seen if the camera were at a predicted position relative to the display surface). The system can determine a simulated (expected) image by using the predicted positions to project the virtual scene onto a screen (i.e., within a 3D model). The same predicted positions used to determine the rendered image on the display screen can be used to determine the predicted image on the display screen. As part of the image comparison, the corresponding virtual scene can be determined and compared to determine an offset or mapping to a position from the predicted position.

[0043] 1.2D-3D tracking Traditional visual odometry techniques can be used to determine a location in 3D space from a 2D camera image. The viewpoint (e.g., pose) can be determined using 3D points in the virtual environment and points on the 2D observation plane (e.g., camera image) according to the following equation:

number

[0044] projection point

number

number

number

number

[0045] Observed 2D feature points

number

number

number

number

[0046] Equation (1) also expresses the three-dimensional virtual point X i a two-dimensional point

number

number

[0047] To determine the location of the camera in the virtual production studio, a 3D point from the virtual scene (e.g., virtual point X i ) is a 2D apparent virtual point

number

number

number

number

[0048] The apparent 3D points can be projected onto a display surface (e.g., display surface 104) by using the following equation:

number

[0049] Known virtual point X i (e.g., virtual point 406) and estimated position (rendered_pose t ) is an apparent virtual point on the display surface (e.g., display surface 402).

number

number

number

[0050] 2.3D-2D tracking Apparent virtual points at 3D locations in a virtual production studio

number

number

number

[0051]

number

number

number

number

number

[0052] Observed 2D point

number

number

number

number

[0053] Observed 2D points

number

number

number

number

[0054] Although the tracking method has been described in the context of feature matching, it can also be described via image registration or global image alignment, where we search for a camera pose by repeatedly rendering simulated images and adjusting the camera pose so that the average difference between the simulated and observed images is minimized.

[0055] 3. Unknown scene geometry In some implementations, the camera tracking system may not have easy direct access to either the virtual geometry (e.g., points 406) or the rendering pose (e.g., rendering pose 410), and thus the tracking algorithm may be completely separated from the rendering system. For example, the tracking algorithm may be separated from the rendering system if the rendered content needs to be protected. The camera tracking system may not need access to the rendering system unless the camera tracking system itself is a separate product and is able to establish a communication protocol. Separating the tracking algorithm from the rendering system may also be desirable to reduce communication latency between the systems. For example, latency may be reduced if estimation algorithms are built into the production camera.

[0056] A tracking system without access to the rendering system can estimate the camera pose from the observed images by using a fully simultaneous localization and mapping (SLAM) strategy tailored to the virtual production environment. Here, we estimate the camera pose actual_pose by using a nonlinear optimization of the form t Not only that, but also the virtual point X i and rendered_pose t can also be optimized:

number

number

[0057] Optimization using equation (5) is similar to equations (3) and (4) in that it uses a model of the display surface to calculate intermediate apparent virtual points that are projected into the camera observation plane (image buffer). However, in some situations, we may solve for the entire system state simultaneously. To help stabilize the system, we may constrain the camera motion (both estimated and actual) by using smoothness_penality.

[0058] As in other SLAM methods, the camera state can be added to equation (5) over time, and the virtual points (X i ) map is incrementally constructed. i The initial uncertainties about the camera poses and motions may be large but may improve over time. We may maintain uncertainty estimates for individual camera poses and virtual points, constrain the poses and virtual points, and remove them from further optimization once their uncertainties fall below a threshold. We may also use loop closure strategies to prevent the creation of redundant virtual structures and reduce drift.

[0059] IV. Using Odometry to Determine the Projection of the Next Scene Visual odometry can determine the current position of the camera after a scene has been rendered and displayed on a display surface. To minimize errors caused by differences between the estimated and actual positions of the camera, techniques are provided for predicting the future position of the camera based on the current position of the camera. The predicted position can be used to render the scene displayed on the display surface.

[0060] A. Determining the error between the predicted position and the actual position 7 shows an overhead perspective view of a virtual production studio 700 with registration errors caused by discrepancies between estimated and actual camera positions, according to at least one embodiment. A display surface 702 bounded by a center stage 704 shows a three-dimensional perspective projection of the virtual scene. When viewed from an estimated camera position 706, the two-dimensional image on the display panel array 702 appears to be part of a three-dimensional computer-generated environment.

[0061] To convey depth, the environment is rendered such that the image simulates light rays traveling from the virtual scene to the estimated camera position 706. An apparent virtual point 708 is indistinguishable from a virtual point 710 when viewed from the estimated camera position 706. To accurately convey the path of light from the virtual point 710 to the estimated camera position 706, the image formed by the apparent virtual point 708 distorts objects (e.g., foreshortening) so that the two-dimensional display surface 702 appears to exhibit depth. However, when the apparent virtual point is viewed from the real camera position 712, image distortion is apparent because the simulated light rays do not converge at the real camera position 712.

[0062] The equations described in Section III, including equations (2) and (3), can be used to determine the actual camera position 712.

number

number

[0063] B. Projecting the next scene The actual camera position determined via visual odometry can be used to identify the estimated position of the next scene in the next time period. A scene can be generated with the estimated position as the intended viewpoint, and this scene can be projected onto the panel array in the next time period.

[0064] FIG. 8 illustrates a method 800 for predicting a future camera position from an actual camera position according to at least one embodiment. The prediction of the future camera pose can be optimized to reduce registration error, which can be the difference between a pixel in the predicted image and a pixel in the actual image. The parameters of the prediction function can be optimized by using training samples, for example, until a convergence criterion is met (e.g., within an error of one pixel), or if the pixel difference is less than a threshold number of pixels, or can be further updated during the current session. The threshold number of pixels can depend on the distance between the display surface and the physical camera. The threshold number of pixels can increase as the distance between the physical camera and the display surface increases. In some implementations, an alarm can be triggered if the pixel difference exceeds a threshold. The alarm can trigger a calibration process, such as resynchronization.

[0065] The product of the prediction function may be a floating-point number that is stored by the computer system as a binary number. The binary number may be an approximation of a floating-point number because while there may be an infinite number of possible floating-point integers, there may be a fixed amount of binary numbers that can be stored in a given number of bits. For example, the value of 1 / 3 cannot be stored exactly in binary. If a floating-point number cannot be stored exactly as a binary number, it may be rounded to a binary approximation. The difference between the floating-point number and the binary approximation is the rounding error. In some implementations, if the magnitude of the registration error is less than the magnitude of the rounding error, the error does not trigger an alarm or cause a calibration process such as resynchronization.

[0066] In block 810, the virtual point X i and estimated camera position rendered_pose t is the apparent virtual point on the display panel surface during the first period

number

number

[0067] In block 820, the apparent virtual point

number

number

number

number

[0068] In block 830, in some implementations, a set of corresponding virtual points and feature points may be established via the methods described in Section III.B.2 above. The feature tracker then calculates the corresponding set of virtual points X i and feature points

number

[0069] In some implementations, the feature tracker data may be obtained by an initialization procedure that may occur after the scene has been rendered and transmitted to the display surface, but before filming occurs. The initialization process involves locating the pose in the virtual production studio while the virtual scene is being displayed on the panel array. t This may include a camera mounted on a robotic arm that cycles through some or all of the features captured by the physical camera.

number

number

number

[0070] In some implementations, the actual camera pose can be learned by directly mapping the observed images to the camera pose. This can be achieved by attaching the camera to a robotic arm and cycling through some or all of the possible positions and contents on the display. Because the position of the camera mounted on the robotic arm is precisely known, a regressor such as a deep neural network can learn the direct mapping between the captured images and the camera pose.

[0071] In block 840, the actual camera position is used to predict the future camera position at the next rendering time. The future camera position can be predicted using sensor data provided by the virtual fabricated camera or an external sensor. For example, if the camera is moving in a straight line at a known speed, the future camera position can be calculated from the difference between the actual position and the rendering time. In some implementations, higher-order deviations of the position can be used to predict the camera position at the next time period from the sensor data.

[0072] In block 850, in some implementations, predicting the future camera position in block 840 may include predicting based on historical camera poses. Predicting the position may be based on historical camera poses, may be achieved by machine learning models, and is discussed in more detail below in section VC.

[0073] V. Predicting Location Rendering a scene based on a predicted rather than actual camera position can reduce registration errors caused by camera motion. Rendering a scene is not instantaneous, so if the camera is moving, a scene rendered based on the camera's current position will be outdated by the time the scene is sent to the panel array. Various techniques for predicting the physical camera's position can be used to reduce registration errors caused by physical camera motion.

[0074] A. Predicting position based on movement data The estimated position of the physical camera during the next time period can be predicted from the motion of the physical camera rather than the position of the physical camera during the previous time period. In some implementations, the motion of the physical camera can be determined by sensors such as an accelerometer, and the motion of the camera can also be determined by an external tracking system (e.g., Optitrack) or via the visual odometry techniques described herein. The sensor data can be combined with the motion and visual odometry history of the physical camera to predict the location of the physical camera. The future position of the physical camera during the next time period can be determined by extrapolating from the current position, velocity, and acceleration of the physical camera. The position can be calculated by the following formula:

number

[0075]

number

number

number

number

[0076] Additionally, other higher order derivatives of the position, such as jerk (third derivative), snap (fourth derivative), crackle (fifth derivative), and pop (sixth derivative), can also be used to make accurate predictions about the future location of the physical camera. The formula for determining the position from the first through sixth derivatives of the position is given as follows:

number

[0077] In addition to the initial position, final position, initial velocity, and initial acceleration discussed above, the above equation includes a vector of higher order derivatives of the motion.

number

number

number

number

[0078] Motion equations such as those discussed above typically assume six degrees of freedom with respect to three-dimensional space. A camera floating in the air can move in many directions from an initial starting point. However, the camera may be attached to a rig, and predictions of the physical camera's motion may take into account how the camera rig constrains the physical camera motion. By constraining the camera's possible motions, predictions regarding the camera's position may be made more quickly and accurately because only possible motions are considered in estimating the position. For example, the camera's possible positions may be limited by the maximum height of the camera rig because the camera may not be in a location where it would be difficult to reach.

[0079] In some implementations, the camera rig may be designed to limit degrees of freedom to help improve prediction accuracy by standardizing the rig's behavior so that camera motion is easier to predict. The camera rig may also constrain camera motion to eliminate motion patterns that cause registration errors. For example, the camera rig may constrain sudden changes in orientation.

[0080] Predicting positions through choreography Physical camera positions can also be determined using staged choreography. In scripted shoots, physical camera positions within the virtual production studio can be known prior to shooting. Choreography can be stored as choreography files, perhaps generated by a game engine, that are used to control camera movement via a robotic rig and to inform motion prediction. In some implementations, scenes can be pre-rendered prior to shooting to minimize latency, and the pre-rendered scenes are sent to panel arrays determined at times by the choreography without the need to render the virtual scene during shooting.

[0081] In some implementations, a choreography file may include a start point, an end point, and a desired path to follow. The path may include one or more positions along the desired path. In some implementations, there may be a start point, an end point, and timestamps of the positions along the desired path. These positions may be defined by a curve (e.g., time derivative) or by a position in a coordinate system (e.g., x, y, z Cartesian coordinates). In some implementations, the desired path may include instructions for a desired camera orientation (e.g., rotation) along the path. In some implementations, the choreography file may be stored binary, as an extensible markup language (XML) file, or as a JavaScript Object Notation (JSON) file.

[0082] In some implementations, the predicted camera position may be compared to a position determined in a choreography file. The difference between the predicted camera position and the position indicated by the choreography file may be communicated to the camera operator. The difference between the predicted position and the position indicated in the choreography file may be used to generate a camera operator metric. For example, the camera operator's score may indicate the number of times the distance between the predicted position and the position determined in the choreography file exceeds a threshold.

[0083] The physical camera may be mounted on a robotic arm and controlled by a camera control system rather than by a human operator. Choreography files may be generated using a human camera operator. The human operator may control a physical camera that is tracked by an external tracking system or via visual odometry. As the human operator moves the camera, their positions may be recorded and added to the choreography file. Once the choreography file is prepared, it may be provided to a computer-controlled rig and the camera path may be replicated by the camera control system.

[0084] Because robotically controlled movements can be precisely repeated, the robotically controlled physical camera can be optimized prior to filming to minimize registration errors. A calibration sequence can be filmed with physical camera movements controlled by a choreography file. Footage shot during the calibration sequence can be compared to the rendered images to determine any registration errors. In some situations, the scene can be re-rendered to minimize any registration errors observed during the calibration sequence.

[0085] The rendering system and the camera control system may also communicate to minimize registration errors. If during a calibration sequence the rendering system determines that the registration errors cannot be reduced by redrawing the scene, the rendering system may notify the camera control system that the predetermined choreography file is causing the registration errors. The rendering system may also indicate the time period within the predetermined choreography file where the registration errors occur. In response, the camera control system may notify the filmmaker that the predetermined choreography file is causing registration errors at the indicated time period so that the filmmaker can take corrective action. In some implementations, the camera control system may take corrective action to reduce the registration errors (e.g., the camera control system may convert a sudden change in direction into a smooth turn).

[0086] B. Predicting location using machine learning models The machine learning model can be trained to determine the camera's position for the next time period based on the camera's current state information. Training the machine learning model can include training the algorithm to classify input data (e.g., state information). In this case, classifying the input data can mean providing the camera's final position in a two-dimensional or three-dimensional coordinate system whose origin is the camera's initial location. During training, training data (e.g., data with known classifications) can be input to the algorithm. Output from the algorithm can be monitored, and if the data is not properly classified, the algorithm's weights can be changed until the output classification matches the label of the training data. Once the algorithm is trained, the model can be used to classify data with unknown classifications.

[0087] In one implementation, the model can be trained on training data that includes current state information such as starting position, starting velocity, and starting acceleration. The training data can be represented as an n-dimensional vector (which describes the input data) called a feature vector. The input to the model can also include the duration between the initial state and the final state as an input, and the training data can be labeled by the final position relative to the starting position. After training, the model should be able to receive state information as input, and the model should provide the final position relative to the starting position as an output.

[0088] In some implementations, the input data may also include labels for the camera operators controlling the virtual production cameras. The model may be trained to provide predicted estimated positions for individual camera operators based on their idiosyncrasies. Additionally, camera operator performance may be evaluated during filming. For example, camera operators may be given a score for how many registration errors occurred while controlling the camera.

[0089] In some implementations, a predetermined choreography file can be used as input to train an algorithm to generate a machine learning model according to the techniques described above. The trained model can be used to identify movements within the predetermined choreography file that may generate registration errors. The training data used to train the algorithm can include past predetermined choreography files with choreography divided into several estimated positions. The training data can also include labels for the estimated positions that indicate whether registration errors were observed. An unlabeled choreography file is input to the trained model to identify choreographed moves that may generate registration errors. Flags can be raised on the estimated positions so that the choreography can be adjusted prior to filming.

[0090] VI. Visual Odometry for Training Simulations The techniques described herein are not limited to filmmaking applications: visual odometry camera tracking techniques can be used in a variety of applications, including immersive training simulations.

[0091] In some implementations, the techniques in this disclosure can be used for training simulations. For example, in an airplane flight simulator, the display surface can include a display panel that mimics the airplane's windows. A camera mounted on a helmet or glasses can be used to determine the trainee's drawing pose (the trainee's viewpoint) through visual odometry. Because the virtual scene shown on the display panel changes with the trainee's position, training allows for immersive movements (e.g., leaning forward to check a blind spot in a vehicle) that are difficult with a static screen.

[0092] More than one trainee may be accommodated if the physical camera used for visual odometry is mounted in polarized glasses (each trainee's glasses with different polarizations). Several virtual scenes with different polarizations may be shown simultaneously on the display surface so that the trainee can view the virtual scene through the polarized glasses, rendered based on the trainee's viewpoint.

[0093] VII. Flowchart 9 is a flowchart of an example process 900 for implementing camera tracking within a virtual production environment according to at least one embodiment. The method may be performed by a computer system communicatively coupled to a movable physical camera.

[0094] In block 910, a first position of the physical camera corresponding to the first time period is identified. The first position of the physical camera may be identified by using an initial image from the physical camera on the display screen. In some implementations, the first position may be determined by using an external tracking system. In other implementations, the first position of the physical camera may be identified from a predetermined choreography file. In some implementations, the display screen may present a set of criteria prior to the first time period to assist in determining the first position. This time period may be the length between subsequent frames captured by the physical camera.

[0095] In some embodiments, the first position may be identified through special measurements by a tracking system or via visual odometry. In other examples, the first position may be identified by prediction (e.g., by using a machine learning model), or the first position may be identified because a location of the first position is defined. For example, the first position may be defined in a predetermined choreography file.

[0096] In block 920, a first time period of the virtual scene is rendered using an animation engine. The virtual scene can be a still image, such as an empty scene, or a moving image, such as an intersection crowded with pedestrians. The first time period can be related to the frame rate of the camera, and in some circumstances, the frame rate of the camera can be increased or decreased to compensate for artifacts captured during capture. The animation engine can be a game engine (e.g., Unreal Engine, Unity, or other game engine) designed for video game development.

[0097] In block 930, the virtual scene is projected onto a display surface to determine a first rendered image for a first time period. The projected virtual scene may correspond to a first position of the physical camera. The display surface may be an array of display panels, including light-emitting diode (LED) panels, organic light-emitting diode (OLED) panels, liquid crystal display (LCD) panels, etc.

[0098] In block 940, the first rendered image is stored in a frame buffer for display on a display surface. A frame buffer may be a fixed memory location in a physical storage medium or a virtual data buffer implemented in software. The rendered image may be displayed on a display surface in a virtual production studio, and in some implementations, the rendered image may be displayed on a portion or the entire display surface.

[0099] In block 950, a first camera image of the display surface is received. The first camera image was acquired using a physical camera during a first time period. In some implementations, the physical camera may include one or more cameras. For example, the physical camera may include a film camera for recording the scene and a positional camera for capturing the first camera image. In some implementations, the positional camera may be a 360-degree camera or an array of cameras. In implementations where the physical camera includes multiple cameras, the rendered scene may include fiducials rendered outside the field of view of the film camera and fiducials rendered within the field of view of the positional camera. For example, the virtual scene may be rendered on one half of the display surface that bounds the intended field of view of the film camera, while fiducials may be rendered on the other half of the display surface.

[0100] In block 960, a first corrected position of the physical camera is determined by comparing the first rendered image and the first camera image. The first rendered image and the first camera image may be compared by identifying features in both images and determining any features that the two images have in common. The relative movement of the features may be used to determine how the orientation and position of the camera has changed relative to the display surface.

[0101] In block 970, a second position of the physical camera corresponding to a second time period is predicted using the first corrected position. During the second time period, a second rendered image will be displayed on the display surface. In various embodiments, a pre-defined choreography file may be used to predict the second position. In some implementations, the second position may be predicted using a trained machine learning model, where the camera operator, choreography, camera position history, and sensor data may be inputs to the machine learning model.

[0102] In block 980, a second virtual scene for a second time period is rendered. The second virtual scene may be rendered using an animation engine. A second virtual scene may be rendered in which the intended viewpoint of the second rendered scene is a second position of the physical camera. In some implementations, two or more separate physical cameras may be used to capture the same scene. For example, cameras may be positioned to capture a scene with two camera lines of sight that are orthogonal to each other. A choreography file may determine which physical camera of the two or more physical cameras is used for a given time period. The location of the physical camera may be tracked even when the camera is not in use, and the scene may be rendered for various cameras as needed.

[0103] In block 990, a second virtual scene is projected onto the display surface to determine a second rendered image for a second time period, the projection of the second virtual scene corresponding to a second position of the virtual camera.

[0104] Process 900 may include additional implementations, such as any single implementation or any combination of the implementations described below and / or in connection with one or more other processes described elsewhere herein.

[0105] In some implementations, the process 900 includes identifying a first position of the physical camera by storing a model that maps images to physical locations and inputting an initial image into this model.

[0106] 9 shows example blocks of process 900, in some implementations, process 900 may include additional, fewer, different, or differently arranged blocks than those depicted in FIG 9. Additionally or alternatively, two or more of the blocks of process 900 may be performed in parallel.

[0107] VIII. Computer Systems Any of the computer systems described herein may utilize any suitable number of subsystems. An example of such a subsystem is shown in computer system 1010 of FIG. 10. In some embodiments, a computer system includes a single computer device, where the subsystems may be components of the computer device. In other embodiments, a computer system may include multiple computer devices (each of which is a subsystem) with internal components. Computer systems may include desktop and laptop computers, tablets, mobile phones, and other mobile devices.

[0108] The subsystems shown in Figure 10 are interconnected via a system bus 1075. Additional subsystems are shown, such as a printer 1074, a keyboard 1078, a storage device 1079, a monitor 1076 (e.g., a display screen such as an LED), and other devices coupled to a display adapter 1082. Peripherals and input / output (I / O) devices (coupled to I / O controller 1071) can be connected to the computer system by any number of means known in the art (e.g., USB, FireWire®), such as input / output (I / O) ports 1077. For example, I / O ports 1077 or external interface 1081 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect computer system 1010 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via system bus 1075 allows central processing unit 1073 to communicate with each subsystem and control the exchange of information between the subsystems as well as the execution of instructions from system memory 1072 or storage device 1079 (e.g., fixed disk such as a hard drive or optical disk). The system memory 1072 and / or storage device 1079 may embody computer-readable media. Another subsystem is a data collection device 1085 such as a camera, microphone, accelerometer, etc. Any of the data described herein may be output from one component to another and may be output to a user.

[0109] A computer system may include multiple identical components or subsystems connected together, for example, by an external interface 1081, by an internal interface, or through a removable storage device that can be connected and removed from one component to another. In some embodiments, computer systems, subsystems, or devices may communicate over a network. In such cases, one computer may be considered a client and another computer may be considered a server, each of which may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0110] Some aspects of some embodiments may be implemented in the form of control logic by using hardware circuitry (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or by using computer software stored in memory by a generally programmable processor in a modular or integrated manner; thus, a processor may include memory storing software instructions (which constitutes not only a hardware circuit but also an FPGA or ASIC with configuration instructions). As used herein, a processor may include a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or on dedicated hardware as well as networked. Based on the disclosure and teachings provided herein, those skilled in the art will know and appreciate other ways and / or methods to implement some embodiments of the present disclosure using hardware and combinations of hardware and software.

[0111] Any of the software components or functions described in this application may be implemented as software code executed by a processor using any suitable computer language, such as Java, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python, for example, using conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive or floppy disk, or optical media such as a compact disc (CD) or DVD (Digital Versatile Disc) or Blu-ray disc, flash memory, etc. The computer-readable medium may be any combination of such devices. Additionally, the order of operations may be rearranged. A process may be terminated when its operations are completed, but may have additional steps not included in the figures. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or to the main function.

[0112] Such programs may also be encoded and transmitted using carrier wave signals adapted for transmission over wired, optical, and / or wireless networks conforming to a variety of protocols, including the Internet. Thus, computer-readable media may be generated using data signals encoded with such programs. Computer-readable media encoded with program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer-readable medium may reside on or within a single computer product (e.g., a hard drive, CD, or entire computer system), and may be present on or within various computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing any of the results described herein to a user.

[0113] Any of the methods described herein may be performed in whole or in part with a computer system including one or more processors that may be configured to perform certain steps. Accordingly, some embodiments may be directed to a computer system configured to perform any of the method steps described herein (perhaps with various components performing each step or group of steps). Although presented as numbered steps, method steps herein may be performed simultaneously, at different times, or in different orders. In addition, some of these steps may be used with some of other steps from other methods. Also, all or part of a step may be optional. In addition, any of the method steps of any of the methods may be performed by a module, unit, circuit, or other means of a system for performing these steps.

[0114] The specific details of the particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the embodiments of the present disclosure, however, other embodiments of the present disclosure may be directed to specific embodiments relating to each individual aspect or to particular combinations of these individual aspects.

[0115] The foregoing description of some exemplary embodiments of the present disclosure has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure to the precise form described above, and thus many modifications and variations are possible in light of the above teachings.

[0116] Enumeration of articles or indefinite articles (a, an, or the) is intended to mean "one or more" unless specifically indicated otherwise. The use of "or" and "or" is intended to mean "inclusive or" rather than "exclusive or" unless specifically indicated otherwise. A reference to a "first" part does not necessarily require that a second part is also provided. Furthermore, a reference to a "first" or "second" part does not limit the referenced parts to a particular location unless otherwise stated. The term "based on" is intended to mean "based at least in part on."

[0117] All patents, patent applications, publications, and descriptions mentioned herein are incorporated by reference in their entirety for all purposes. None are admitted to be prior art. In the event of a conflict between the instant application provided herein and the reference application, the instant application shall control.

Claims

1. 1. A method for implementing camera tracking in a virtual fabricated environment, the method comprising, by a computer system communicatively coupled to a movable physical camera, performing: identifying a first position of the physical camera corresponding to a first time period; Rendering the first virtual scene for the first time period using an animation engine; projecting the first virtual scene onto a display surface to determine the first rendered image for the first time period, the projection of the first virtual scene corresponding to the first position of the physical camera; storing the first rendered image in a frame buffer for display on the display surface; receiving a first camera image of the display surface, the first camera image being acquired using the physical camera during the first period of time; determining a first corrected position of the physical camera by comparing the first rendered image with the first camera image; predicting, using the first corrected position, a second position of the physical camera corresponding to a second time period during which a second rendered image is displayed on a display surface; Rendering the second virtual scene for the second time period using an animation engine; and projecting the second virtual scene onto the display surface to determine a second rendered image for the second time period, the projection of the second virtual scene corresponding to the second position of the physical camera. The method includes:

2. Comparing the first rendered image and the first camera image includes: identifying a first pixel location of an object in the first rendered image and a second pixel location of the object in the first camera image; determining a difference between the first pixel location and the second pixel location; and The method of claim 1 , comprising applying the difference to the first position to obtain the first corrected position.

3. The method of claim 1 , wherein predicting the second position uses information from a predetermined choreography file for the final video to be captured.

4. The method of claim 1 , wherein the first position of the physical camera is identified by using an initial image from the physical camera of the display screen.

5. Identifying the first position of the physical camera includes: storing a model that maps images to physical locations; and The method of claim 4 , further comprising inputting the initial image into the model.

6. The method of claim 1 , wherein predicting the second position of the physical camera uses information from a history of positions of the physical camera.

7. Predicting the second position of the physical camera includes: storing the trained model that predicts future locations from current locations; and The method of claim 1 , further comprising inputting the first location into the model.

8. The method of claim 7 , wherein inputting the first position into the model includes inputting a set of motion data of the physical camera at the first position.

9. The method of claim 8 , wherein the set of motion data includes at least one of velocity, acceleration, jerk, snap, crackle, and pop.

10. When executed, it causes the computer system to: identifying a first position of the physical camera corresponding to a first time period; Rendering the first virtual scene for the first time period using an animation engine; projecting the first virtual scene onto a display surface to determine the first rendered image for the first time period, the projection of the first virtual scene corresponding to the first position of the physical camera; storing the first rendered image in a frame buffer for display on the display surface; receiving a first camera image of the display surface, the first camera image being acquired using the physical camera during the first time period; determining a first corrected position of the physical camera by comparing the first rendered image with the first camera image; predicting, using the first corrected position, a second position of the physical camera corresponding to a second time period during which a second rendered image is displayed on a display surface; Rendering the second virtual scene for the second time period using an animation engine; and projecting the second virtual scene onto the display surface to determine the second rendered image for the second time period, the projection of the second virtual scene corresponding to the second position of the physical camera.

10. A computer product comprising a non-transitory computer-readable medium storing a plurality of instructions for controlling the computer to perform the steps of:

11. Comparing the first rendered image and the first camera image includes: identifying a first pixel location of an object in the first rendered image and a second pixel location of the object in the first camera image; determining a difference between the first pixel location and the second pixel location; and The computer product of claim 10 , further comprising applying the difference to the first position to obtain the first corrected position.

12. 11. The computer product of claim 10, wherein predicting the second position uses information from a predetermined choreography file of the final footage to be captured.

13. 11. The computer product of claim 10, wherein the first position of the physical camera is identified by using an initial image from the physical camera of the display screen.

14. Identifying the first position of the physical camera includes: storing a model that maps images to physical locations; and The computer product of claim 13 , further comprising inputting the initial image into the model.

15. The computer product of claim 10 , wherein predicting the second position of the physical camera uses information from a history of positions of the physical camera.

16. Physical cameras; a non-transitory computer-readable medium; and 1. A system including one or more processors communicatively coupled to the non-transitory computer-readable medium, wherein the one or more processors: identifying a first position of the physical camera corresponding to a first time period; Rendering the first virtual scene for the first time period using an animation engine; projecting the first virtual scene onto a display surface to determine the first rendered image for the first time period, the projection of the first virtual scene corresponding to the first position of the physical camera; storing the first rendered image in a frame buffer for display on the display surface; receiving a first camera image of the display surface, the first camera image being acquired using the physical camera during the first period of time; determining a first corrected position of the physical camera by comparing the first rendered image with the first camera image; predicting, using the first corrected position, a second position of the physical camera corresponding to a second time period during which a second rendered image is displayed on a display surface; Rendering the second virtual scene for the second time period using an animation engine; and projecting the second virtual scene onto the display surface to determine the second rendered image for the second time period, the projection of the second virtual scene corresponding to the second position of the physical camera. A system configured to:

17. The system of claim 16 , wherein predicting the second position of the physical camera uses information from a history of positions of the physical camera.

18. Predicting the second position of the physical camera includes: storing the trained model that predicts future locations from current locations; and The system of claim 16 , further comprising inputting the first location into the model.

19. 20. The system of claim 18, wherein inputting the first position into the model includes inputting a set of motion data for the physical camera at the first position.

20. 20. The system of claim 19, wherein the set of motion data includes at least one of velocity, acceleration, jerk, snap, crackle, and pop.