Virtual production based on display assembly pose and pose error compensation

By processing motion capture data to determine and correct pose errors in display assemblies, the method enhances the accuracy of virtual model representation and content rendering, improving the quality of virtual production.

JP2025526219APending Publication Date: 2025-08-13NANT HOLDINGS IP LLC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024558351
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-10
Filing Date
2023-07-10
Publication Date
2025-08-13

AI Technical Summary

Technical Problem

Existing virtual production systems face challenges in accurately rendering content due to pose errors between the physical and virtual models of display assemblies, which can result in alignment issues and reduced quality of the production.

Method used

A method and system for rendering content based on a display assembly's pose, involving motion capture data processing to determine the physical pose, generating a transformation to a virtual pose, and updating a virtual model to reflect the current physical pose of displays, thereby correcting positional and rotational errors.

Benefits of technology

Improves the accuracy of virtual model representation and content rendering by aligning the physical and virtual poses, enhancing the quality of virtual production by accurately positioning content on displays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025526219000001_ABST
    Figure 2025526219000001_ABST
Patent Text Reader

Abstract

Systems, devices, and methods are disclosed for rendering content based on a pose of a display assembly. In one example, motion capture data for a display of a plurality of displays is received, the display moving from a first physical pose to a second physical pose. The motion capture data is processed to determine coordinates of the second physical pose. A transformation of the second physical pose of the display to a virtual pose of the display is generated. A virtual model of the plurality of displays is updated, the virtual model including the virtual pose of the display. Content is rendered on the updated virtual model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 390,252, entitled "VIRTUAL PRODUCTION BASED ON POSE ERROR CORRECTION," filed July 18, 2022, and U.S. Provisional Patent Application No. 63 / 458,412, entitled "VIRTUAL PRODUCTION BASED ON DISPLAY ASSEMBLY POSE," filed April 10, 2023, the contents of which are incorporated herein by reference in their entireties for all purposes.

[0002] The field of the invention relates to virtual production. [Background technology]

[0003] The background discussion includes information that may be useful in understanding the subject matter of the present invention. None of the information provided herein is admitted to be prior art or admitted prior art by the applicant, or to be relevant to the subject matter of the presently claimed invention, or that any publication specifically or implicitly mentioned is prior art or admitted prior art by the applicant.

[0004] A virtual production, e.g., a motion picture production, typically includes a virtual stage for presenting scene-related content, camera devices for generating motion picture data by capturing video of people, objects, and content, and motion capture systems for tracking the camera, people, and / or objects. The content may be dynamic (e.g., video content that changes over time) and / or its presentation may be adjusted based on the tracking.

[0005] All publications identified herein are incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference. In the event that a definition or use of a term in an incorporated reference contradicts or is contrary to the definition of that term provided herein, the definition of that term provided herein shall apply and the definition of that term in the reference shall not apply.

[0006] In some embodiments, numerical values used to describe and claim particular embodiments of the inventive subject matter, for example, expressing quantities or units of data, should be understood to be modified in some cases by the term "about." Accordingly, in some embodiments, the numerical parameters set forth in the written description and appended claims are approximations that may vary depending upon the desired properties sought to be obtained by a particular embodiment. In some embodiments, the numerical parameters should be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. Notwithstanding that the numerical ranges and parameters setting forth the broad scope of some embodiments of the inventive subject matter are approximations, the numerical values set forth in the specific examples are reported as precisely as practicable. The numerical values set forth in some embodiments of the inventive subject matter may contain certain errors necessarily resulting from the standard deviation found in their respective testing measurements.

[0007] Unless the context dictates to the contrary, all ranges set forth herein should be construed as inclusive of their endpoints, and open-ended ranges should be construed as including only commercially practical values. Similarly, all lists of values should be considered as inclusive of intermediate values unless the context dictates to the contrary.

[0008] As used throughout this specification and the claims that follow, the meanings of "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Also, as used herein, the meaning of "in" includes "in" and "on" unless the context clearly dictates otherwise.

[0009] Recitation of ranges of values herein is merely intended to serve as a shorthand method of individually referencing each separate value falling within the range. Unless otherwise indicated herein, each separate value is incorporated herein as if set forth individually herein. All methods described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples or exemplary language (e.g., "such as") provided herein with respect to specific embodiments is intended merely to better clarify the inventive subject matter and does not impose limitations on the scope of the inventive subject matter otherwise recited in the claims. No language herein should be construed as indicating any non-claimed element essential to the practice of the inventive subject matter.

[0010] Groupings of alternative elements or embodiments of the subject matter disclosed herein are not to be construed as limitations. Each group member may be referenced and claimed individually or in any combination with other members of the group or other elements found herein. One or more members of a group may be included in, or deleted from, a group for reasons of convenience and / or patentability. When such inclusion or deletion occurs, the specification shall be deemed to include the group as modified to fulfill the written description of all Markush groups used in the appended claims.

[0011] It should be understood that many of the basic technical features provided in the following specification are presented to enable a compact review of the disclosed inventive subject matter. Although some of the basic technical features described herein may appear unclear, in many cases, such features can be considered to be within the understanding of those skilled in the art. Therefore, the presentation of such background art should not be considered limiting. Summary of the Invention [Means for solving the problem]

[0012] Embodiments described herein include a method for rendering content based on a pose of a display assembly. The method includes at least one processor receiving motion capture data for a display of a plurality of displays, the display moving from a first physical pose to a second physical pose. The at least one processor processes the motion capture data to determine coordinates of the second physical pose. The at least one processor generates a transformation of the second physical pose of the display to a virtual pose of the display. The at least one processor updates a virtual model of the plurality of displays, the virtual model including the virtual pose of the display. The at least one processor renders content on the display based on the updated virtual model.

[0013] Embodiments may further include a system comprising one or more processors and one or more memories storing instructions that, when executed by the one or more processors, configure the system to receive motion capture data of a display of a plurality of displays moving from a first physical pose to a second physical pose. The system may further process the motion capture data to determine coordinates of the second physical pose. The system may further update a translation of the second physical pose of the display to a virtual pose of the display. The system may further update a virtual model of the plurality of displays including the virtual pose of the display. The system may further render content on the display based on the updated virtual model.

[0014] Embodiments may further include a non-transitory computer-readable medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations including receiving motion capture data for a display of the plurality of displays, the display moving from a first physical pose to a second physical pose. The at least one processor processes the motion capture data to determine coordinates of the second physical pose. The at least one processor generates a transformation of the second physical pose of the display to a virtual pose of the display. The at least one processor updates a virtual model of the plurality of displays, the virtual model including the virtual pose of the display. The at least one processor renders content on the display based on the updated virtual model.

[0015] Embodiments may further include a computer-implemented method including: determining, by at least one processor, a first pose of a first display of a plurality of displays included in a display assembly, the display assembly being configured to display content on the plurality of displays; and determining, by the at least one processor, a virtual model of the display assembly, the virtual model being stored in computer-readable memory and including a virtual representation of each one of the plurality of displays. Further, the method may include determining, by the at least one processor, a transformation based on the first pose of the first display and the first virtual representation of the first display in the virtual model; and rendering, by the at least one processor, the content on at least some of the plurality of displays in accordance with the transformation and the virtual model.

[0016] Further, embodiments may include a system comprising one or more processors and one or more memories storing instructions that, when executed by the one or more processors, configure the system to determine a first pose of a first display of a plurality of displays included in a display assembly, the display assembly configured to display content on the plurality of displays. The instructions may further configure the system to determine a virtual model of the display assembly, the virtual model being stored in one or more memories and including a virtual representation of each one of the plurality of displays. The instructions may further configure the system to determine a transformation based on the first pose of the first display and the first virtual representation of the first display in the virtual model, and to render content on at least some of the plurality of displays in accordance with the transformation and the virtual model.

[0017] Embodiments may additionally include one or more non-transitory computer-readable storage media storing instructions that, when executed on a system, cause the system to perform operations including determining a first pose of a first display of a plurality of displays included in a display assembly, the display assembly being configured to display content on the plurality of displays. The instructions can cause the system to determine a virtual model of the display assembly, the virtual model being stored in a computer-readable memory and including a virtual representation of each one of the plurality of displays. The instructions can cause the system to determine a transformation based on the first pose of the first display and the first virtual representation of the first display in the virtual model, and to render content on at least some of the plurality of displays according to the transformation and the virtual model. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 illustrates an example of a virtual production system, according to an embodiment of the present disclosure. [Figure 2] FIG. 1 illustrates an example of a display assembly and a virtual model thereof, according to an embodiment of the present disclosure. [Figure 3] FIG. 10 illustrates an example of a motion path of a physical marker on a display assembly, according to an embodiment of the present disclosure. [Figure 4] FIG. 1 illustrates an example of measured motion capture data that can be used to determine the physical pose of a display of a display assembly, according to an embodiment of the present disclosure. [Figure 5] FIG. 10 illustrates an example presentation path of a virtual marker on a display assembly, according to an embodiment of the present disclosure. [Figure 6] FIG. 1 illustrates an example of transforming a display assembly and determining an updated virtual model according to an embodiment of the present disclosure. [Figure 7] FIG. 1 illustrates an example of transform-based position error correction according to an embodiment of the present disclosure. [Figure 8] FIG. 1 illustrates an example flow for determining and using correction for pose error, according to an embodiment of the present disclosure. [Figure 9] FIG. 1 illustrates an example flow for rendering content based on posing error correction, according to an embodiment of the present disclosure. [Figure 10] FIG. 1 illustrates an example flow for processing motion capture data according to an embodiment of the present disclosure. [Figure 11] FIG. 1 illustrates an example flow for determining a transformation based on a field of view of a camera device, according to an embodiment of the present disclosure. [Figure 12] FIG. 1 illustrates an example of a virtual production system, according to an embodiment of the present disclosure. [Figure 13] FIG. 1 illustrates an example of a virtual production system, according to an embodiment of the present disclosure. [Figure 14] FIG. 10 illustrates a plot of change-point data for a virtual production, according to an embodiment of the present disclosure. [Figure 15] FIG. 1 illustrates an example of a display assembly, according to an embodiment of the present disclosure. [Figure 16] FIG. 1 illustrates an example of transforming a display assembly and determining an updated virtual model according to an embodiment of the present disclosure. [Figure 17] 1 is a table showing the difference between measured and calculated parameters. [Figure 18] FIG. 1 illustrates an example process flow for determining an updated virtual model based on display movement, according to an embodiment of the present disclosure. [Figure 19] FIG. 10 illustrates an example process flow for determining initialization of a motion capture system for a virtual production, according to an embodiment of the present disclosure. [Figure 20] FIG. 1 illustrates an example flow for determining a transformation based on detecting movement of a display of a display assembly, according to an embodiment of the present disclosure. [Figure 21]FIG. 1 illustrates exemplary components of a computer system according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0019] It should be noted that any reference to a computer should be read to include any appropriate combination of computing devices, including a server, interface, system, database, agent, peer, engine, controller, module, or other type of computing device acting individually or collectively. A computing device should be understood to comprise at least one processor configured to execute software instructions stored on a tangible, non-transitory computer-readable storage medium (e.g., a hard drive, FPGA, PLA, solid-state drive, RAM, flash, ROM, etc.). The software instructions or set of software instructions configure or program the computing device or its processor to provide roles, responsibilities, or other functions as described below with respect to the disclosed apparatus or system. Furthermore, the disclosed technology may be embodied as a computer program product including a non-transitory computer-readable medium storing software instructions or sets of software instructions that cause one or more processors to perform the disclosed steps associated with implementing a computer-based algorithm, process, method, or other instructions. In some embodiments, the various servers, systems, databases, or interfaces exchange data using standardized protocols or algorithms, possibly based on HTTP, HTTPS, TCP, UDP, FTP, SNMP, IP, AES, public-private key exchange, web services or RESTful APIs, known financial operations protocols, or other methods of electronic information exchange. Data exchange between devices may be performed over a packet-switched network, the Internet, a LAN, a WAN, a VPN, or other type of packet-switched network, a circuit-switched network, a cell-switched network, or other type of network (wired or wireless).

[0020] As used throughout this specification and the claims that follow, when describing a system, engine, server, agent, device, module, or other computing element as configured to perform or execute functions on data in memory, the meaning of "configured to" or "programmed to" is defined as one or more processors or cores of the computing element being programmed with a set of software instructions stored in the memory of the computing element to perform a set of functions on target data or data objects stored in the memory. It should be understood that the combination of software and hardware working in concert creates a dedicated set of physical, real-world structures that provide a utility to one or more users that would not exist outside the confines of the physical, real-world assets.

[0021] It should be appreciated that the disclosed techniques provide many advantageous technical effects, including improved modeling of a physical display assembly and improved rendering of content on at least a portion of a display of the physical display assembly. For example, by relying on pose measurements of the physical display assembly, the accuracy of a virtual model representing the physical display assembly is improved. Because the accuracy of the virtual model is improved, content rendering using the virtual model is also improved, thereby, for example, more accurately positioning the presentation of content on the display.

[0022] Embodiments of the present disclosure are directed, among other things, to rendering content based on a virtual model representing a real-time pose of a display assembly. In one example, the display assembly includes multiple configurable displays, each of which is positioned at a specific pose within the display assembly (e.g., a specific physical location and a specific rotation within the display assembly). The pose of each display may be configurable for a desired virtual production. For example, the displays have six degrees of freedom, such as the ability to mechanically rotate, move forward, move backward, and tilt. In comparison, a virtual model includes a virtual representation of each display, where the virtual representation of the displays in the virtual model exhibits a real-time virtual pose that corresponds to the real-time physical pose of the displays. By updating the virtual model to reflect the current pose of each display of the display assembly, the quality of the virtual production may be improved.

[0023] By updating the virtual model to reflect the current pose of each display of the display assembly, the quality of the virtual production can be improved.

[0024] A display of the display assembly can be moved from a first physical pose to a second physical pose. For example, the display can be connected to an actuator, such as a winch, to reposition the display from the first physical pose to the second physical pose. Virtual markers can be projected onto the display. Based on detection of the repositioning, a motion capture system can be initialized to collect motion data using the virtual markers. The movement of the display can be detected by one or more image capture devices using the virtual markers.

[0025] Initialization of the motion capture system can be based on either manual initialization or sensor-based responsive initialization. Manual initialization can be performed by a user. In the case of sensor-based initialization, one or more sensors can be directed toward the display assembly and continuously collect streaming data. An algorithm, such as a forgetting factor-based change-point detection algorithm, can process the streaming data to determine when a change in one or more parameters of the display assembly (e.g., display rotation) has occurred.

[0026] A forgetting factor-based change point detection algorithm may be used to detect change points from streaming data. Some streaming applications for change point detection may require selecting two or more parameters for change point detection. However, the selection of parameters may be based on the expected change size, and for streaming data where multiple change sizes may occur, the selection may not be optimal. Therefore, a forgetting factor-based change point detection algorithm may be used, which only requires the selection of a single parameter.

[0027] A second physical pose can be determined by the computer system based on the motion data captured by the motion capture system. In some cases, repositioning of the display can occur during filming of the virtual production. In such cases, the second pose can be determined after the display stops moving, or the current real-time physical pose of the display can be continuously determined as the display is moving.

[0028] To account for shifts in the position or orientation of the displays of the display assembly, a fitting model is used to generate a transformation, and the current physical pose (e.g., after the shift is input into the fitting model) is used to determine the current virtual pose. For example, the fitting model may perform an implementation of one or more fitting algorithms, such as a Levenberg-Marquardt nonlinear least-squares algorithm, a chi-squared test algorithm, a curve fitting algorithm, a weighted least-squares fitting algorithm, a polynomial regression algorithm, a Gauss-Newton algorithm, a shift-and-cut algorithm, a gradient algorithm, or a Nelder-Mead (simplex) search algorithm. Different techniques are possible for determining the physical pose. In one example technique, a motion capture system is used to generate motion capture data that tracks virtual markers presented on a moving display. The motion capture data is correlated with the moving display (e.g., a first portion of the motion capture data, the first motion capture data, is associated with the moving display of the display assembly). The current pose of the display (e.g., the pose after the display has moved) is derived from the portion of the motion capture data associated with it.

[0029] Furthermore, embodiments of the present disclosure are directed, among other things, to rendering content based on correcting pose errors (e.g., positional and / or rotational errors) of a virtual model representing a display assembly. In one example, a display assembly includes multiple displays, each disposed at a specific pose within the display assembly (e.g., a specific physical position and a specific rotation within the display assembly). In comparison, a virtual model includes a virtual representation of each display, where the virtual representations of the displays in the virtual model exhibit virtual poses corresponding to the physical poses. Due to different factors (e.g., installation tolerances, incorrect installation, operating temperature, heat, thermal expansion, etc.), a mismatch may exist between the virtual pose and the physical pose. This mismatch can cause quality issues when content is rendered based on the virtual model for presentation on the display assembly. By correcting the pose errors, the mismatch can be reduced or even eliminated, thereby mitigating the quality issues.

[0030] To correct for pose errors, a fitting model is used to generate a transformation, and the physical pose and the virtual pose are input into the fitting model. For example, the fitting model can perform an implementation of one or more fitting algorithms, such as a Levenberg-Marquardt nonlinear least-squares algorithm, a chi-squared test algorithm, a curve fitting algorithm, a weighted least-squares fitting algorithm, a polynomial regression algorithm, a Gauss-Newton algorithm, a shift-and-cut algorithm, a gradient algorithm, or a Nelder-Mead (simplex) search algorithm. Different techniques are possible for determining the physical pose. In one example technique, a motion capture system is used to generate motion capture data that tracks physical markers placed at locations on a display according to a predefined motion path. In another example technique, rather than tracking physical markers, the motion capture system is used to generate motion capture data that tracks virtual markers presented on a display according to a presentation path. In either example approach, the motion capture data is correlated with the display (e.g., a first portion of the motion capture data, the first motion capture data, is associated with a first display of the display assembly, a second portion of the motion capture data, the second motion capture data, is associated with a second display of the display assembly, etc.) The physical pose of the display is derived from the portion of the motion capture data associated therewith, for example, by determining the coordinates and rotations of markers (physical or virtual) and including such data in the physical pose.

[0031] Different techniques are also possible for generating the transformation. A first exemplary technique, referred to herein as global determination, uses the full set of physical poses and the full set of virtual poses as input to the fitting model. A second exemplary technique, referred to herein as local determination, instead uses a subset of the physical poses and a corresponding subset of the virtual poses. For example, the presentation of content on a display assembly may be based on one or more parameters (e.g., its pose) of a camera device that generates video data showing the content. A subset of the displays may be within the field of view of the camera device, while the remaining displays may be outside the field of view. In this illustration, only the physical poses of the displays included in the subset and their corresponding virtual poses are input to the fitting model. The physical poses and virtual poses of the remaining subset are excluded from the input. In this manner, the transformation is locally optimized by considering only relevant pose data (e.g., positional and / or rotational data of the displays within the field of view). In general, local determination techniques use a subset of displays, where the subset is localized based on one or more parameters of the camera device. The subset may be defined in the XY plane, for example, by including "x" by "y" displays, where "x" is the total number of displays along the horizontal axis ("x total "), and / or "y" is the total number of displays along the vertical axis ("y total "). For example, a subset may be smaller than a vertical stack or a column of a display (e.g., "x" equals one "1" and "y" equals one "y"). total "), horizontal strips or rows of a display (e.g., "x" is equal to one "x total" and "y" equals "1"), a diagonal strip of displays, a contiguous block of "x" by "y" displays, a non-contiguous block of "x" by "y" displays (e.g., a first display and a second display are part of a block, but the display between these two displays is not part of the block), etc. By using a local subset of displays, fine-grained adjustments can be made to the transformation such that it is optimized to reduce pose error in a particular dimension, which may correspond to particular camerawork or movement.

[0032] For clarity of explanation, various embodiments of the present disclosure are described in the context of a virtual production use case in which content is presented on a virtual stage based on a virtual model of the virtual stage. However, the embodiments are not so limited and equally apply to other use cases, such as virtual reality, augmented reality, mixed reality, content projection (e.g., in a home theater, movie theater, or building), performance stage (e.g., music concert), etc. In general, embodiments of the present disclosure enable improved content presentation on a display assembly, which presentation relies on a virtual model of the display assembly. In particular, embodiments disclose techniques for updating a virtual model for a virtual production based on movement of display elements. Additionally or alternatively, embodiments disclose techniques for reducing pose error between the actual physical pose of a display element of a display assembly and its corresponding virtual pose. The display element may be an actual display represented as a rigid body in the virtual model. Additionally or alternatively, the display element may be a subdivision (e.g., a section) of the display, which is also represented as a rigid body in the virtual model.

[0033] Also, for clarity of explanation, various embodiments of the present disclosure are described in the context of position, position error, and position error correction (in real and virtual spaces). However, these embodiments are not so limited and equivalently apply to rotation, rotation error, and rotation error correction (in real and virtual spaces). These embodiments also apply to pose, which is a combination of position and rotation. Position error and / or rotation error may exist (e.g., in pose), and position error correction and / or rotation error correction may be performed.

[0034] FIG. 1 illustrates an example of a virtual production system 100 according to an embodiment of the present disclosure. As illustrated, the virtual production system 100 includes, among other things, a display assembly 110, a camera device 120, motion capture devices 130A, 130B, and 130C (generally referred to by the numeral "130"), and a computer system 140. The display assembly 110 may be configured as a virtual stage that presents content 112 and defines a volume in which a camera device 120 is positioned (multiple such camera devices 120 are also possible). The presentation of the content 112 may be controlled by a game engine (e.g., the UNITY game engine, the UNREAL game engine, etc.) running on the computer system 140. For example, the content 112 itself and / or parameters of the presentation (e.g., panning, angling, tilt, etc.) may be controlled based on a number of factors. Among these factors is the pose of the camera device 120 within the volume. Motion capture device 130 can generate motion capture data that is processed to determine the poses of objects 150A, 150B, 150C (generally referenced with the numeral "150" and which may be actors, stage furniture, scene furniture, equipment, etc.) over time, as well as the pose of camera device 120. Such motion capture data can be processed and used to control the presentation of content 112 on display assembly 110.

[0035] In one example, the display assembly 110 includes multiple displays arranged to form a content presentation screen. An example of such an arrangement is further described in FIG. 2. The position and orientation of each display may be configurable so that a display or set of displays of the display assembly can be moved from a first pose to a second pose. The content presentation screen may have different shapes (e.g., curved surfaces, flat intersecting surfaces, etc.) and may encompass an area in which an object 150 may be placed, thereby defining a volume. In some cases, one or more displays are moved to cause the content presentation screen to change from having a first shape to having a second shape. The content presentation screen may be used to present interactive and dynamic scenes. Objects 150 can interact with such scenes, and content 112 may be updated based on the interaction (e.g., actor 150A can interact with a virtual object presented in content 112). Although display assembly 110 is shown as having a vertical position (e.g., configured as a wall), display assembly 110 (or a second display assembly) may additionally or alternatively be positioned in other positions (e.g., a horizontal position to define a ceiling or floor). Furthermore, the volume may include a sliding set of displays that can be moved in and out of the volume to define a particular shape. For example, the volume may be formed in the shape of a horseshoe, and the sliding set may be positioned to close the open portion of the horseshoe. In this way, a camera device 120 within the volume may be surrounded by a full 360 degrees of displays. Other shapes are possible, whereby the volume may be, for example, a hemisphere, a full sphere (i.e., 4π steradians), a cylinder, or a cube.

[0036] Camera device 120 may be a cinema camera mounted on a rig (e.g., a floor rig and / or a ceiling rig) that can be repositioned within the volume, or a movable rig (e.g., a tripod, a gimble, etc.). In this manner, camera device 120 may be configured to film a scene by generating video data (and optionally audio data) indicative of one or more of objects 150 and / or part or all of content 112 presented on display assembly 110, particularly from multiple viewpoints. The camera device 120 may have a high resolution (e.g., 4K, 6K, 8K, 12K, etc.) and may be available from, for example, BLACKMAGIC (e.g., URSA MINI PRO 12K, STUDIO CAMERA 4K PLUS, STUDIO CAMERA 4K PRO, URSA BROADCAST G2, etc.), ARRI (e.g., ALEXA MINI LF, ALEXA LF, ALEXA MINI, ALEXA SXT W, AMIRA, AMIRA LIVE, ARRI MULTICAM SYSTEM, etc. equipped with an ARRI SIGNATURE PRIME 35mm T1.8 lens, ARRI SIGNATURE PRIME 75mm T1.8 lens, etc.).

[0037] The motion capture device 130 may be a motion capture camera (e.g., an infrared camera) and / or other type of motion sensor (e.g., a depth sensor) that is part of a motion capture system configured to track motion within a volume. The motion capture system may be available, for example, from VICON (e.g., using VANTAGE, VERO, VUE, VIPER, VIPERX cameras, etc., and SHOGUN software, etc.) or OPTITRACK (e.g., using PRIME-X 41, PRIME, SLIM-X, SLIM, FLEX cameras, etc., and UNREAL PLUGIN, UNITY PLUGIN, MOTION BUILDER PLUGIN, OPTICAL MOTION CAPTURE SOFTWARE, MAYA PLUGIN software, etc.). The motion of an object may optionally be tracked using a motion tracker attached to the object. Tracking may involve locating an object within the volume by determining the object's position and rotation over time. The coordinate system of the motion capture system (e.g., a Cartesian coordinate system or any other coordinate system) may be defined relative to any origin within the volume.

[0038] Computer system 140 may be configured to process at least a portion of the motion capture data and, optionally, a portion of the video data. For example, a game engine may render content 112 using a virtual model of display assembly 110 and position data of camera device 120. Rendering may involve using a rendering engine to composite images and / or image frames (e.g., two-dimensional or three-dimensional), which are then presented on display assembly 110 as content 112. In addition to displaying content 112, one or more displays of display assembly 110 may be configured to generate virtual light for illuminating one or more production elements. The rendering of the content may be affected by the physical pose of each display. Thus, the virtual model may be updated to include the current virtual pose of each display so that the content has a desired effect.

[0039] FIG. 2 illustrates an example of a display assembly 110 and its virtual model 230, according to an embodiment of the present disclosure. The display assembly 110 may include an arrangement 210 of multiple displays. Each display 220 may be a plate with a specific shape and may be individually controllable to display content. The display assembly 110 may form a volume for virtual production. This volume may be approximately 2,230 m 2 The virtual production space may include a display assembly 110 forming a curved LED wall approximately 16m wide by 20m long, 270 degree oval (360 degree configuration is also possible).

[0040] In one example, display 220 is a light-emitting diode (LED) plate having a flat screen that displays content. The flat screen may have a square shape (as shown by the grid on display 220 in FIG. 2, where each cell in the grid represents a pixel) and a particular pixel resolution. Nevertheless, other curvatures and / or shapes of the screen and the underlying technology of display 220 (e.g., liquid crystal display—LCD) are possible. In a particular exemplary use case, the display is a BLACK PEARL 2 display available from ROE CREATIVE DISPLAY. In this exemplary use case, the screen is a flat panel with dimensions of 500×500×90 mm (height×width×depth), a resolution of 176×176 (horizontal×vertical), a pixel pitch of 2.84 mm, an LED surface-mounted diode (SMD) configuration, a magnesium frame with magnetic connectors, and a locking system. Although each display 220 has a solid, well-defined shape, installing a large number of displays 220 to form a display assembly 110 according to a desired shape along multiple degrees of freedom (e.g., six degrees of freedom to form a 270-degree curved wall 16 m wide by 20 m long) may result in pose offsets that need to be accounted for and corrected for in the virtual model of the display assembly.

[0041] Each display 220 may further be coupled to an actuator (e.g., a winch system, a motor, a robotic arm) for moving the display 220 from a first physical pose to a second physical pose. At any time during the virtual production, one or more displays may be moved to create a desired effect for the content displayed on the display assembly. For example, the shape of the display assembly may be changed for a desired filming of the virtual production. In other examples, a single display may be moved to create a desired filming effect. In each example, moving the display 220 may change the visual parameters of the content displayed on or around the display 220. Thus, to display the desired content with the desired effect, a computer system (e.g., computer system 140 of FIG. 1) may use a virtual model that includes a virtual pose that is an accurate representation of the current physical pose of each display.

[0042] Arrangement 210 may include stacking displays to form a particular shape of display assembly 110. For example, displays are placed adjacent to each other to form a desired curvature, height, and length of display assembly 110. Each display has a physical pose (i.e., an actual, real-world pose) in arrangement 210. The physical pose may be defined relative to a point on the display (e.g., top left corner, center, etc.) relative to an origin in the production volume (e.g., the origin of a coordinate system used by a motion capture system).

[0043] The virtual model 230 may include a three-dimensional object that represents the display assembly 110 as a rigid body within the game engine. The three-dimensional object may also represent each display as a rigid body by including a virtual representation thereof (e.g., as a three-dimensional sub-object). In this manner, the virtual model 230 may be a virtual representation not only of the display assembly 110 but also of the displays that form the display assembly. As part of the virtual representation, the virtual model 230 may indicate the virtual shape, virtual dimensions, and virtual pose (e.g., position and rotation) of each display within the three-dimensional object. In the illustration of FIG. 2, the virtual representation is shown as a curved mesh that mimics the display assembly 110, with each cell in the mesh representing one of the displays.

[0044] Virtual model 230 may match arrangement 210, and the virtual pose of the virtual representation of the display in virtual model 230 may match the actual physical pose of the display in arrangement 210. However, due to multiple factors (e.g., installation tolerances, incorrect installation, operating temperature, heat, thermal expansion, human interaction, natural forces, etc.), there may be a mismatch between the virtual and physical positions and / or between the virtual and physical rotations. Such a mismatch may potentially cause alignment errors, particularly from the camera's perspective, when rendering content on the display.

[0045] In some cases, the current physical pose of the display is determined after the display has stopped moving. In such cases, the resting physical pose of the display 220 may be determined after the display 220 has stopped moving. Once the resting physical pose is determined, a fitting model may be used to determine the current position and orientation of the display 220. A transformation function may be applied to update the virtual model to include a virtual pose for the display that represents the current resting physical pose of the display 220. In other examples, the current physical pose of the display 220 may be determined in real time as the display 220 is moving. In such cases, the current physical parameters of the display 220 may be continuously updated and input to the fitting model. The fitting model may continuously output a virtual model that includes the current virtual pose of the moving display 220. In such cases, the current physical parameters may cease to be used as input once the display 220 has stopped moving or after a short time interval (e.g., a few seconds) has elapsed after the display 220 has stopped moving.

[0046] 3 shows an example of a motion path 330 of a physical marker 320 on a display assembly 310 (e.g., display assembly 110) according to an embodiment of the present disclosure. In particular, the physical marker 320 may be positioned on different displays of the display assembly 310 at different times according to the motion path 330. A motion capture system may track the physical marker 320, and the resulting motion capture data may be processed to determine the physical position and / or physical rotation of the displays. The motion capture system may be the same as that used in the virtual production system 100, such as systems available from VICON or OPTITRACK.

[0047] The arrangement of displays in the display assembly 310 may be indexed by rows 350 and columns 340. The motion path 330 is shown in FIG. 3 as starting from the right of the bottom row of the display assembly 310 (e.g., the display with index (C,3)), moving horizontally left to the end of the bottom row, moving up one row, moving horizontally to the right, etc. Of course, other types of motion paths are possible (e.g., an "S"-like path starting from the top left, top right, or bottom left, or even a zigzag path). On each display, a physical marker 320 is placed at a location corresponding to a point used to define the physical pose of the object, such as the top left corner of the display (or any other point, such as the center, bottom right corner, etc.). The physical marker 320 remains at that location for a predefined time interval (e.g., 1 second, 2 seconds, etc.) before being moved to the next location according to the motion path 330. A time of 2 seconds has been found to be acceptable.

[0048] Depending on the motion capture system, the physical marker 320 may be a rigid body implementing motion tracking technology. For example, in the case of an infrared motion capture camera, the physical marker 320 may include one or more infrared-emitting (active or passive) points (each using a different infrared frequency). Generally, the greater the number of points, the more accurate the position estimation may be. In the above example of placing the physical marker 320 in the upper left corner of the display, the marker's upper left infrared-emitting red point may be placed at this location on the display and used as a reference point (e.g., the root of the rigid body) in the motion capture data to determine the physical position and / or physical rotation of the display. The motion capture system may be the same motion capture system used during virtual production involving the display assembly 310.

[0049] In one example, the physical marker 320 includes a single point detectable by an infrared motion capture camera. In this case, at least three infrared motion capture cameras may be required to detect the position of the physical marker 320. In particular, each of the three cameras would generate a two-dimensional image showing the marker's position in two dimensions. Because the position, orientation, and field of view of each camera are known, a three-dimensional vector on which the physical marker 320 is located can be determined from the set of three two-dimensional positions. In another example, the physical marker 320 includes multiple points detected by infrared motion capture cameras. In this case, a single infrared motion capture camera may be sufficient to detect the position of the physical marker 320. In particular, the relative positions of the points are known a priori, and this knowledge is used to process the images generated by the cameras. Of course, technologies other than infrared may also be used. For example, a two-dimensional visual marker may be used that encodes its dimensions, an image of which may be generated by an optical sensor operating in the human visible wavelength range. The pose of the visual marker may be determined by decoding the dimensions and applying geometric reconstruction to the image.

[0050] At some point (e.g., before production begins, during virtual production, etc.), the physical marker 320 may be moved by an operator (e.g., a human or a machine such as a robot or unmanned vehicle) according to the motion path 330. For example, the operator may first position the physical marker 320 (e.g., its upper left infrared emitting point) over the upper left corner of the (C,3) display for two seconds, then reposition the physical marker 320 (e.g., by aligning its upper left infrared emitting point) to the upper left corner of the (B,3) display for two seconds, etc.

[0051] The motion capture data may be processed according to the motion path 330 to determine the physical pose of the display. An example of processing is described further herein below.

[0052] 4 illustrates an example of measured motion capture data 410 that can be used to determine the physical pose of a display in a display assembly (e.g., display assembly 310) according to an embodiment of the present disclosure. In particular, a plot 400 is shown illustrating the x-coordinate (vertical axis) of a physical marker (e.g., physical marker 320) versus time (horizontal axis). In other words, the motion capture data 410 illustrated in FIG. 4 corresponds to the motion of the physical marker along the X-axis. This motion capture data 410 can be used to determine the pose of each display along the X-axis.

[0053] Motion capture data for a physical marker may be captured along other axes (e.g., Y and Z axes). For clarity of explanation, the X coordinate is described herein. However, embodiments equally apply to other coordinates of a physical marker to determine the position and rotation of the physical marker in three-dimensional space (including X, Y, Z coordinates and rotations). Embodiments also equally apply to non-Cartesian coordinate systems (e.g., if a polar coordinate system were used instead, ray and angle coordinates could be tracked and processed to determine the pose in three-dimensional space).

[0054] In one example, the motion capture data 410 is generated at a particular frame rate (e.g., 144 frames per second (FPS)) such that a single x-coordinate is available at the particular frame rate (e.g., approximately once every 7 milliseconds (ms)). Further, a predefined motion path can indicate the timing (e.g., approximately every 2 seconds) to statically place a physical marker at a location on the display, which timing can be related to an index of the display (e.g., referring back to FIG. 3 , from approximately 0 seconds to 2 seconds, the index is (C,3); from approximately 2 seconds to 4 seconds, the index is (B,3), etc.). Processing of the motion capture data 410 can determine a physical position on the display based on the frame rate and the predefined motion path.

[0055] In particular, a second x coordinate follows the first x coordinate (or two sets of subsequent x coordinates can be used to compare values and determine a change in x position). The value of the second x coordinate can be compared to the value of the first x coordinate. If the difference between the two values is less than a predefined threshold, this small difference indicates that the x position has remained substantially the same. If the difference between the two values is greater than a predefined threshold, this large difference indicates that the x position has changed. Different types of comparisons are available, such as comparing magnitudes, comparing changes in slope, etc.

[0056] This type of comparison-based determination is illustrated in FIG. 4 using the numeral 420. In particular, between time "t1" and time "t2," a change 420 is determined, and the change 420 is greater than a predefined threshold. Therefore, between time "t1" and "t2," the physical marker was relocated from a first location to a second location. A time window "TW0" between times "t0" and "t1" corresponds to the first location. This time window "TW0" has a time length of "t1-t0," which may typically be approximately 2 seconds. Between times "t1" and "t2," the relocation of the physical marker occurs as indicated by the change 420. The next time window "TW1," beginning at approximately "t2," when the change 420 is no longer observed, corresponds to the second location, and should be approximately 2 seconds long. This 2-second time window is provided herein for illustrative purposes. Different lengths of time may be used (e.g., about 1 second, about 5 seconds, etc.) and / or the different time windows need not have approximately the same lengths of time (e.g., one time window may be about 2 seconds while another time window may be about 5 seconds).

[0057] Considering the motion path, a first time window "TW0" corresponds to a first display index (C,3). Similarly, a second time window "TW1" corresponds to a second display index (B,3). Next, a portion of the motion capture data 410 having timing between times "t0" and "t1" (referred to herein for clarity as "first motion capture data") is processed to determine the x-position of the physical marker during the first time window "T0", and equivalently, the x-position of the display having the first display index (C,3). For example, all of the first motion capture data starting at time "t0" and ending at time "t1", a percentage (e.g., 60%) of the first motion capture data starting after time "t0" and ending before time "t1" (e.g., 25 ms after time "t0" and 30 ms before time "t1"), or a subset thereof, is statistically analyzed to determine x-position statistics (e.g., averages). Similar processing may be applied to the second motion capture data corresponding to the second time window "TW2," etc. Similar processing may also be applied to determine x and z positions and x, y and z rotations.

[0058] FIG. 5 illustrates an example of a presentation path 530 of a virtual marker 520 on a display assembly 510 (e.g., display assembly 110) according to an embodiment of the present disclosure. Unlike the use of physical markers 320, here the virtual markers 520 are used to determine the physical pose of the displays included in the display assembly 510. The presentation path 530 may be similar to the motion path 330, whereby the virtual markers 520 are presented on different displays according to a predefined sequence (e.g., an "S"-like sequence starting from the right of the bottom row, a zigzag sequence starting from the left of the top row, etc.). Optionally, the display index 522 of each display is also presented on the display in parallel with the presentation of the virtual marker 520 on the display. In this manner, the presentation path 530 need not be predefined and may be random, as long as the virtual markers 520 are presented on different displays over time.

[0059] In one example, the virtual marker 520 may be a multi-dimensional model (e.g., a two-dimensional model, a three-dimensional model, etc.) of a rigid body. The virtual marker 520 may be presented at a specific location on the display (e.g., the center as shown in the figure, but other locations such as the upper left corner are possible). Rather than physically moving between locations as with the physical marker 320, the presentation of the virtual marker 520 may remain on the display for a time window (e.g., 2 seconds) at a specific location, then stop on the display, and begin simultaneously or shortly thereafter on the next display (which may be, but need not be, an adjacent display). The presentation of the display index may also occur in parallel and thus follow the presentation path 530. The display index 522 is generally presented at a display location other than the display location of the virtual marker 520 (e.g., the lower right corner, as opposed to the virtual marker 520 being presented in the center).

[0060] Generally, virtual markers 520 do not use infrared technology unless each display is capable of emitting light in the infrared range. Instead, virtual markers 520 may include one or more virtual points that emit light in the human visible wavelength range, and a camera operating in that wavelength range may be used to capture one or more images of virtual marker 520 during presentation. The camera may be, but need not be, a motion capture camera. Similar to physical marker 320, virtual marker 520 may include at least three points, each of which may be a different color and / or shape, or even unique to a particular display (e.g., a barcode, QR code, unique shape, etc.), such that a single camera may be sufficient to generate an image of virtual marker 520, which may be processed to determine a corresponding presentation pose on the display. Alternatively, virtual marker 520 may include a single point, and three or more cameras may be used to generate images of virtual marker 520, which may be processed to determine a corresponding presentation pose on the display. Alternatively, the virtual marker 520 may be a virtual visual marker that encodes its dimensions, and a single camera may be sufficient to generate an image of this visual marker, which can be processed to determine a corresponding presentation pose on the display. Alternatively, the virtual marker 520 may have an asymmetric shape (e.g., a rectangular prism but not a square cube). Its presentation on the display may change its orientation (e.g., rotation, angle, etc.), and at least one image showing that change may be captured and processed, along with other images showing other changes, to determine the presentation pose of the virtual marker 520 on the display. Regardless of the technique used, the presentation pose of the virtual marker 520 on the display corresponds to the physical pose of the display.

[0061] Alternatively, rather than rearranging the presentation of the virtual markers 520 between displays, different virtual markers can be presented simultaneously or non-simultaneously on the displays. The virtual markers can have different shapes, and each shape can be associated with a display index. In this manner, processing the images can include recognizing the shape and then associating each virtual marker with a corresponding display. In another example, the virtual markers can have the same shape, and upon presentation of the virtual marker on a display, a display index is also displayed on the display. In this manner, processing the images can also include recognizing the display index. In an exemplary use case, virtual markers (of different shapes or the same shape with a display index) are presented simultaneously on a display assembly. One or more images are generated and processed to determine a pose of each of the virtual markers presented on the displays and associate the pose with the physical pose of the display.

[0062] Once an image of the virtual marker 520 is generated, the image can be processed to determine its presentation pose on each of the displays, and thus, equivalently, the physical pose of each display. Different processing techniques are possible. As an example, if a display index is not provided and instead a predefined presentation path is used, the processing described in connection with FIG. 4 can be applied. In particular, motion capture data is derived from the images by determining the presentation pose in each image, this data is processed to determine change over time, and a no-change time window is correlated with a particular display index that provides the timing indicated by the predefined presentation path. In another example, a display index is provided. Upon processing the images to detect the presentation pose of the virtual marker 520 therein, the display index indicated in the image can also be detected (e.g., by using optical character recognition and / or object detection machine learning models). Thus, a presentation pose can be associated with a display index and, accordingly, can represent the physical pose of the corresponding display.

[0063] A combination of techniques using physical and virtual markers is possible for collecting motion capture data (or, more generally, image data). As an example, a virtual marker is presented on a display as a placement instruction. An operator can then place a physical marker at the presented location, thereby covering the virtual marker. In another example, a virtual marker may or may not be presented. However, a display index is presented on the display. In this way, in addition to generating motion capture data corresponding to a physical marker placed at a location on the display, image data for capturing the display index can be generated in parallel. The motion capture data can be processed to determine the physical pose of the physical marker, and the image data can be processed to detect the display index. When the timing of the motion capture data and the timing of the image data match, the physical pose is associated with the display index.

[0064] FIG. 6 illustrates an example of determining a transform 610 and an updated virtual model 620 of a display assembly (display assembly 110) according to an embodiment of the present disclosure. In one example, virtual model 602 (e.g., virtual model 230) represents a display assembly. Transform 610 includes a set of functions (e.g., point-by-point rotation and / or translation along each dimensional axis, warping, twisting, bending, randomization, etc.) that can be applied to virtual model 602, resulting in updated virtual model 620. Updated virtual model 620 compensates for (e.g., reduces or even eliminates) pose errors between virtual model 602 and the display assembly (in other words, the pose errors of updated virtual model 620 are much smaller, if at all, than the pose errors of virtual model 602). An example of position error compensation is shown in the following figure.

[0065] To generate the transform 610, a physical pose 604 of a display included in the display assembly is determined and input along with the virtual model 602 (or more specifically, along with the corresponding virtual position) into a fitting model 630. The output of the fitting model 630 includes the parameters (e.g., coefficients) of the transform 610. The physical pose 604 may be derived based on motion capture data such as described in FIGS. 3-4, image data (which may include motion capture data) such as described in FIG. 6, and / or other positioning techniques.

[0066] The fitting model 630 may be a data fitting model that iteratively estimates the parameters of the transformation 610 to best fit the transformed virtual position to the physical position. Different types of data fitting models are possible, such as those based on an implementation of a Levenberg-Marquardt nonlinear least-squares algorithm, a chi-squared test algorithm, a curve fitting algorithm, a weighted least-squares fitting algorithm, a polynomial regression algorithm, a Gauss-Newton algorithm, a shift-and-cut algorithm, a gradient algorithm, a Nelder-Mead (simplex) search algorithm, or other types of fitting algorithms. Additionally or alternatively, multiple known virtual models and corresponding display assemblies can be used to train a machine learning model, such as a regression model or a convolutional neural network, to output the transformation parameters. Once trained, the virtual model 602 and the physical pose 604 can be input into the machine learning model, which outputs the parameters of the transformation 610.

[0067] As further explained in the next figure, the pose error may not be constant and may vary depending on the sub-area of the display assembly (e.g., the pose error of a display in the bottom left corner of the display assembly may be significantly different from the pose error of a display in the center of the display assembly, which in turn may be significantly different from the pose error of a display in the top right corner of the display assembly). To optimize the pose error variation, a local rather than a global determination technique is used. The global determination technique involves inputting the entire set of virtual poses of the virtual model 602 and the entire set of physical poses 604 into the fitting model 630. In this way, a single transformation is generated and used to correct the pose error for content rendering across the entire display assembly.

[0068] In comparison, the local determination technique involves dividing the display assembly into subareas. Each subarea includes a subset of displays. A transformation is generated for each subset and used to correct pose error for rendering a portion of the content, which is presented on the subset of displays. A first transformation associated with a first display subset may differ from a transformation for a second display subset. Generating the first transformation may include pose data (e.g., virtual and physical positions) for the first display subset and exclude pose data for the second display subset. In particular, to generate the first transformation, a subset of virtual poses and a corresponding subset of physical poses 604 are input into fitting model 630. The two subsets are associated with the first display subset. Another subset of virtual poses and another subset of physical poses 604 may be input into fitting model 630 to generate another transformation, and so on. As mentioned above, a subset may be defined in the XY plane, for example, by including "x" by "y" displays, where "x" is the total number of displays along the horizontal axis ("x total"), and / or "y" is the total number of displays along the vertical axis ("y total "). The subset may be selected based on many factors. For example, the subset corresponds to displays that are within the field of view of the camera. In another example, content is rendered in a particular way (e.g., with special effects) on a subset of displays, and the accuracy of the content presentation (including, e.g., how well the special effects are visually perceptible) depends on the positional error of these displays. In this case, the subset is one that is used in a local determination technique. In yet another example, a coarse estimate of the transform may be used, e.g., to reduce computational overhead or processing latency. In this example, every other display or other selection pattern (e.g., random selection distributed across the display assembly) may be used to define the subset. In a further example, a multi-granularity approach may be used, starting with a coarse, quick calculation of the transform, followed by a more targeted calculation (e.g., selection of displays within the field of view and / or selection of displays for particular special effect presentation).

[0069] In the context of virtual production, different transformations (each associated with a subarea of a display assembly, e.g., by generating transformations to correct for pose errors localized to the subarea) can be generated offline and used as needed during virtual production. Alternatively, each transformation can be generated in real time based on need. In particular, during virtual production, environmental factors can cause pose changes for a particular display, and transformations can be calculated in real time to correct for the resulting pose errors. Environmental factors can include, for example, an increase in temperature within or around a volume, or equipment / people accidentally bumping into the display assembly. In one example, the virtual production involves a camera device (e.g., camera device 120), whereby the content presented on the display assembly and / or its presentation is controlled, at least in part, based on the pose of the camera device. This pose (which can be tracked using a motion capture system) can indicate that the camera is at a distance from the display assembly and oriented in a particular direction, such that the resulting camera field of view includes the subarea of the display assembly. In this situation, a subset of displays included in the subarea can be determined. For example, a sub-area is defined as a projection of the field of view onto a display assembly, where the projection is determined based on the orientation of the camera device and the distance to the display assembly. Given a display index, the display belonging to the projection is identified, and the corresponding virtual and physical poses are obtained and input into fitting model 630 to generate a transformation in real time for use in rendering the content.

[0070] FIG. 7 illustrates an example of correcting position errors based on a transformation, according to an embodiment of the present disclosure. The transformation may be generated using any of the techniques described in connection with FIG. 6. In particular, FIG. 7 illustrates a plot 700 in which the horizontal axis corresponds to the x-coordinate, the vertical axis corresponds to the y-coordinate, and points in the plot correspond to two-dimensional positions of a display defined by (x, y) coordinates in the XY plane. Three types of two-dimensional positions ((x, y) coordinates) are shown: circles correspond to virtual positions from a virtual model of the display assembly; crosses correspond to physical positions of displays in the display assembly; and triangles correspond to generated corrected virtual positions. Each virtual position corresponds to a physical position and a corrected virtual position, and the corrected virtual positions are generated by applying a transformation to the virtual positions. Each triple of virtual position, physical position, and corrected virtual position corresponds to a single display. The bottom row of triples in plot 700 corresponds to a first row of the display at a first height (e.g., row "3" in FIG. 3), while the top row of triples in plot 700 corresponds to another row of the display at a second height (e.g., row "2" in FIG. 3). Conversely, each column of triples corresponds to a column of the display at a different height.

[0071] If there were no position error (e.g., zero translation, diagonal unity translation, etc.), each virtual location would coincide with the corresponding physical location (e.g., both locations would have the same (x,y) coordinates). However, as explained above, due to different factors, this is not the case as shown in Figure 7. The difference between a virtual location and its corresponding physical location is the position error (e.g., in the XY plane). As shown in Figure 7, the position error is reduced, and in some cases eliminated, causing the corrected virtual locations to more closely coincide with (and partially overlap) the corresponding physical locations. As an example, we performed measurements of a display assembly including a BLACK PEARL 2 display available from ROE CREATIVE DISPLAY. Each such display has a 500 x 500 (height x width) flat panel on the front. The display assembly forms a 270-degree curved wall measuring 16 m wide x 20 m long. Without using embodiments of the present disclosure, position errors along the X-axis, Y-axis, and Z-axis vary between approximately 150-250 mm, 20-50 mm, and 175-225 mm, respectively. By implementing the global determination techniques of the embodiments, these position errors can be significantly reduced to approximately 30-60 mm, 5-15 mm, and 40-56 mm, respectively.

[0072] Although not explicitly shown in Figure 7, the position error may not be uniform across different displays. As described herein above, local decision techniques that generate multiple transforms, each corresponding to a subset of the displays, can be used to optimize the correction of the position error to account for the non-uniform distribution.

[0073] 8-11 illustrate flows associated with pose determination and pose error correction in the context of content rendering that relies on a virtual model of a display assembly on which the content is displayed. The operations of the flows may be performed by a computer system (e.g., at least one processor, at least one computer-readable memory, etc.), such as computer system 140 of FIG. 1 . Some or all of the instructions for performing the operations may be implemented as hardware circuits and / or stored as computer-readable instructions on a non-transitory computer-readable medium of the computer system. When implemented, the instructions represent components including circuitry or code executable by a processor of the computer system. The use of such instructions configures the computer system to perform particular operations described herein. Each circuit or code, in combination with an associated processor, represents means for performing the respective operation. While the operations are shown in a particular order, it should be understood that no particular order is required and one or more operations may be omitted, skipped, performed in parallel, and / or reordered.

[0074] FIG. 8 shows an example flow for determining and using corrections for position errors according to an embodiment of the present disclosure. The flow may begin at operation 802, where a computer system collects motion capture data. For example, the motion capture data may be generated by a motion capture system and transmitted to the computer system. The motion capture data may track the pose of a physical marker on a display assembly over time, where the pose changes according to a predefined motion path. Additionally, or alternatively, the motion capture data may track the pose of a virtual marker presented on the display assembly. As described in connection with FIG. 5, in addition to or as an alternative to the motion capture data, image data may be generated when a virtual marker is used, and the image data may likewise be collected by the computer system.

[0075] In operation 804, the computer system processes the motion capture data to determine the physical pose of the display. The processing may depend on the collection technique. In one example using physical or virtual markers, changes in the motion capture data over time are used to determine a time window, and a predefined motion path is used to associate the time window with a display index. Motion capture data having timing within the time window is used to determine the physical pose (e.g., as a statistical measure such as an average applied to this data) of the display with the corresponding display index. In another example using virtual markers, the image data is processed to determine the pose of the virtual marker, and the timing of the image data is used to determine the corresponding display index according to a presentation path, and / or the display index is also presented and recognized directly from the image.

[0076] In operation 806, the computer system accesses a virtual model representing the display assembly. For example, the virtual model is loaded from the computer system's memory or retrieved from a remote data store.

[0077] In operation 808, the computer system generates a transformation. In one example, the virtual pose and physical pose of the virtual model are input into a fitting model, which then outputs parameters of the transformation, and the transformation is associated with the display assembly. In another example, a subset of the virtual poses and a corresponding subset of the physical poses are input into a fitting model, which then outputs a transformation, and the transformation is associated with a subarea of the display assembly.

[0078] In operation 810, the computer system renders the content by correcting pose errors in the virtual model based on the transformation. For example, an updated virtual model is generated by applying the transformation to the virtual model (or a portion thereof corresponding to a subset of the virtual positions). The updated virtual model is used by a game engine running on the computer system to render the content, and the rendered content is then displayed by a display assembly.

[0079] FIG. 9 illustrates an example of a flow for rendering content based on pose error correction according to an embodiment of the present disclosure. The operations of the flow may be implemented as suboperations of the flow of FIG. 8. In one example, the flow of FIG. 9 may begin at operation 902, in which a computer system determines a first pose (e.g., a first position and / or a first rotation) of a first display of multiple displays included in a display assembly. The display assembly is configured to display content on the multiple displays. The first pose may be a physical pose of the display, which may be determined based on processing of motion capture data and / or image data, as described in FIGS. 3-5. The physical pose may be defined, for example, by using a point (e.g., the upper left corner) of the display relative to the origin of a coordinate system of the motion capture system.

[0080] In operation 904, the computer system determines a virtual model of the display assembly. The virtual model includes a virtual representation of each one of the multiple displays. For example, the virtual representation of the display includes a multidimensional object that represents the display and its placement relative to the other displays, and indicates a virtual pose of the multidimensional object. The virtual pose can also be defined by using a point (e.g., the upper left corner) of the multidimensional object relative to the origin of a coordinate system.

[0081] In operation 906, the computer system determines a transformation based on the first pose of the first display and the first virtual representation of the first display in the virtual mode. For example, the first virtual pose and the first pose shown by the first virtual representation are input to a fitting model. Inputs to this model may include virtual poses and physical poses associated with other displays, depending on whether a global or local determination technique is used. The fitting model can then output parameters of a function (e.g., rotation, translation) that defines the transformation.

[0082] At operation 908, the computer system renders content based on the transformation and the virtual model on at least a portion of the display of the display assembly. For example, an updated virtual model is generated from the virtual model by translating each, some, or all of the virtual points of the virtual model and / or rotating a virtual object formed by multiple virtual points of the virtual model according to parameters of the transformation. A game engine can use the virtual model, along with other data such as the pose of a camera device, to render the content.

[0083] 10 illustrates an example of a flow for processing motion capture data according to an embodiment of the present disclosure. The flow corresponds to a motion capture data processing technique that relies on a predefined motion path followed to reposition physical markers between different positions. The operations of the flow may be implemented as suboperations of the flow of FIG. 8. In one example, the flow of FIG. 10 may begin at operation 1002, where a computer system receives motion capture data. The motion capture data may be generated by a motion capture system at a particular frame rate (e.g., 144 fps).

[0084] In operation 1004, the computer system determines a change in the motion capture data. The motion capture data may be multidimensional. The change may be determined for each dimension. For example, the motion capture data may include x-coordinates, and each x-coordinate is generated at a particular frame rate (e.g., approximately every 7 ms). Thus, two x-coordinate values (or the average of two ranges of x-coordinates) may be compared to determine a difference, which may be compared to a predefined distance threshold. If the difference is greater than the threshold, the change is detected, and the computer system determines that it corresponds to a transition of a physical marker from one location on one display to another location on another display.

[0085] In operation 1006, the computer system determines the timing associated with the change. For example, the timing may be available from a timestamp of the motion capture data and may coincide with the end or beginning of a time window (e.g., a 2-second time window) during which the physical marker is expected to be substantially statically positioned at a location on the display. Each time window may be represented by a predefined motion path.

[0086] In operation 1008, the computer system determines that the first motion capture data corresponds to a first display. For example, a change in the motion capture data is detected. Two consecutive changes correspond to the start and end of a time window. The timing of the changes is correlated with a display index for each predefined motion path. In this way, the portion between the start and end of the motion capture data becomes the first motion capture data. This first motion capture data can then be associated with the display index.

[0087] In operation 1010, the computer system determines a pose of the first display based on the first motion capture data. For example, a statistical metric may be applied (e.g., averaged) to the first motion capture data to determine its position. In certain circumstances, a subset of the first motion capture data (e.g., a percentage thereof, or a portion beginning a few milliseconds after the start of the time window and ending a few milliseconds before the end of the time window) may be subjected to the statistical metric to calculate the pose.

[0088] FIG. 11 illustrates an example of a flow for determining a transformation based on a field of view of a camera device according to an embodiment of the present disclosure. This flow corresponds to a local determination technique. The operations of the flow may be implemented as suboperations of the flow of FIG. 8. In one example, the flow of FIG. 11 may begin at operation 1102, in which a computer system determines a subset of displays within a field of view of a camera device. In one example, the content presented on the display assembly and / or the presentation of the content may be altered based on the pose of the camera. A motion capture system may output motion capture data tracking the camera device, and the computer system may process this data to determine the pose. Given the pose, a projection of the field of view of the camera onto the display assembly may be used to determine a subarea of the display assembly within the field of view. Displays belonging to this subarea are identified, and their display indexes are included in the subset.

[0089] In operation 1104, the computer system determines the physical pose of the displays that are in the field of view. The pose may be determined based on motion capture data of the physical and / or virtual trackers and / or image data of the virtual trackers, as described in Figures 3-5. In particular, the display indexes of the displays are used to look up and obtain the corresponding physical position data.

[0090] In operation 1106, the computer system determines a virtual pose of the virtual representation of the display in the virtual model. For example, the virtual representation (e.g., a multidimensional object representing the display) is also indexed with the same display index. In this manner, the display index is used to look up and obtain the corresponding virtual position data.

[0091] In operation 1108, the computer system generates a transformation, for example, the virtual pose and the physical pose are input into a fitting model, which then outputs the parameters of the transformation.

[0092] 12 illustrates an example of a virtual production system 1200, according to an embodiment of the present disclosure. As shown, the virtual production system 1200 includes a display assembly 1210, a camera device 1220, motion capture devices 1230A and 1230B (generally referenced by the numeral "1230"), and a computer system 1240. The display assembly 1210 may include perimeter walls 1250, a back wall 1260, and a roof 1270. The display assembly 1210 may be configured as a virtual stage that presents content and defines a volume in which the camera device 1220 is located.

[0093] The roof 1270 may include a display 1280 that can be moved from a first physical pose to a second physical pose. As shown, the display 1280 is positioned in a first physical pose such that the display is flush with the surface of the roof 1270. The display 1208 may be coupled to an actuator that can move the display 1280 from the first physical pose to the second physical pose. The actuator 1290 may include, for example, a winch system, a motor, a lever, a pulley system, or a robotic arm. In some cases, the display 1280 may be moved from the first physical pose to the second physical pose, for example, to create a desired effect on the display 1280. This may be performed by using the actuator 1290 to move the display 1280. Additionally or alternatively, one or more other displays that may be positioned on the roof 1270 and / or any other walls 1250-1260 may be moved to effect a pose change. The pose change of any display may be updated in a virtual model (e.g., virtual model 230).

[0094] 13 illustrates an example of a virtual production system 1300, according to an embodiment of the present disclosure. As shown, the virtual production system 1300 includes a display assembly 1310, a camera device 1320, motion capture devices 1330A and 1330B (generally referred to by the numeral "1330"), and a computer system 1340.

[0095] Figure 13 may differ from Figure 12 in that display 1350 (corresponding to display 1280) is positioned in a second physical pose. In Figure 12, display 1280 is positioned in a first physical pose such that display 1280 is flush with the surface of roof 1270. In Figure 13, display 1350 has been moved to a second physical pose such that display 1350 is rotated away from roof 1270. Computer system 1340 can still display content on display 1350. However, to create the desired visual effect, the virtual model may be updated to include a virtual pose of display 1350 that represents the current physical pose of display 1350.

[0096] As indicated above, the process for updating the virtual model may be initiated manually or based on a response to a sensor measurement. For manual initialization, the computer system 1340 may be in operative communication with a switch 1370. The switch 1370 may include one or more keys on a keyboard, a command prompt, a physical switch external to the computer system 1340, or other suitable switch 1370. A user may use the switch to initialize the display pose estimation process. A virtual marker 1380 may be displayed on the display 1350. The virtual marker 1380 may include a pattern, such as a synthetic marker (e.g., a CharUCo marker or an ArUco marker). As shown, the virtual marker 1380 spans the entire surface of the display 1350. However, in some cases, the virtual marker 1380 may be displayed on a portion of the display surface. The virtual marker may have a presentation pose that corresponds to the physical pose of the display 1350.

[0097] By using switch 1370, the motion capture system including motion capture devices 1330A and 1330B can begin capturing motion data using the virtual markers. In some cases, motion capture devices 1330A and 1330B can be configured to focus on the display at the coordinates of display 1350. In other examples, motion capture devices 1330A and 1330B can capture the entire display assembly 1310. Motion capture devices 1330A and 1330B can continuously capture frames of motion data while display 1350 is moving. Motion capture devices 1330A and 1330B can further transmit the motion data to computer system 1340.

[0098] The computer system 1340 can use a final frame (e.g., a frame captured after the display 1350 has stopped moving) to determine pose data for the display 1350. The pose data can include physical coordinates of the display 1350. The computer system 1340 can generate inputs based on the pose data and the virtual model to feed to the fitting model. The fitting model can use a transformation function to convert the current physical pose of the display 1350 into the current virtual pose of the display 1350. The computer system 1340 can further update the virtual model to include the current virtual pose of the display 1350. The computer system 1340 can use the updated model to display content on the display assembly 1310.

[0099] As an alternative to using switch 1370 to initiate the pose estimation process, a sensor-based response trigger may be used. Virtual production system 1300 may include one or more sensors 1390 configured to collect data (velocity, acceleration, six degrees of freedom) from display assembly 1310. Sensors 1390 may include, for example, proximity sensors, light sensors, pressure sensors, infrared sensors, ultrasonic sensors, or any other suitable sensors. Sensors 1390 may be in operative communication with computer system 1340 and provide collected data to computer system 1340.

[0100] The sensor-based data can be presented as a time series, and the computer system 1340 can detect change points in the sensor data using a change point detection algorithm. Change point detection is the process of detecting changes in the characteristics (e.g., physical coordinates of a display) represented by the time series. The computer system 1340 can identify boundaries between changes in the time series using a change point detection algorithm (e.g., forgetting factor-based change point detection algorithm, window-based segmentation, binary segmentation, bottom-up segmentation, pruning and extraction linear time, and exact segmentation dynamic programming). A plot of the first detected change point and the second detected change point is provided in FIG. 14.

[0101] In some cases, the first detected change point can represent the start of movement of the display 1350, and the second detected change point can represent the end of movement of the display 1350. Thus, the time interval between the first detected change point and the second detected change point can represent the time interval over which the display is moving.

[0102] Based on the detection of the first change point, a virtual marker can be displayed on the moving display 1350. The computer system 1340 can further initialize the motion capture devices 1330A and 1330B. Initializing the motion capture devices 1330A and 1330B can include determining the position and orientation of each motion capture device and configuring the device for capturing display motion (e.g., configuring the frame rate). The motion capture devices 1330A and 1330B can capture frames of the display 1350 as it moves. The motion capture devices 1330A and 1330B can continue capturing frames until the computer system 1340 detects a second change point, indicating that the display 1350 has stopped moving. The computer system 1340 can further send a signal to the motion capture devices 1330A and 1330B to stop collecting motion data for the display 1350. In some cases, motion capture devices 1330A and 1330B continue to capture one or two frames after display 1350 has stopped moving.

[0103] The computer system 1340 can use the collected sensor data to determine the physical pose (e.g., physical coordinates) of the display 1350 after the display 1350 stops moving. The computer system 1340 can analyze characteristics of the virtual markers to determine the presentation pose of the markers (e.g., coordinates of the virtual markers corresponding to physical coordinates of the display). The characteristics can include, for example, corners of a chessboard, centers of circles, and other image features. The computer system can further use an algorithm to determine the physical pose of the display 1350. For example, the computer system 1340 can use a point-and-point (PnP) pose calculation algorithm to evaluate the physical pose of the display in a desired coordinate system (e.g., the coordinate system of a camera in a motion capture system). The computer system can further generate inputs for a fitting model based on the virtual model (e.g., the virtual model generated before the display 1350 moved) and the physical pose. The fitting model can output an updated virtual model, including an updated virtual pose of the display 1350 representing the physical pose of the display after the display 1350 stops moving.

[0104] It should be appreciated that in some embodiments, rather than waiting for display 1350 to stop moving, computer system 1340 can begin determining the physical pose of display 1350 as it is moving. For example, in response to detecting a first change point, virtual markers can be displayed on display 1350, and the computer system can initialize motion capture devices 1330A and 1330B to begin capturing motion data using the virtual markers. Computer system 1340 can continuously generate input for the fitting model using a virtual model (e.g., a virtual model generated before display 1350 moved) and the current physical pose of display 1350 as it is moving. The fitting model can use the input to convert the virtual pose of display 1350 from before the display moved into a virtual pose of display 1350 as it is moving. For example, if the frame rate of motion capture devices 1330A and 1330B is 10 frames per second and the time interval over which the display is moving is 10 seconds, computer system 1340 can generate 100 updated virtual poses for 100 iterations of the updated virtual model.

[0105] FIG. 14 is a plot 1400 of change-point data for a virtual production, according to an embodiment of the present disclosure. The plot 1400 includes time on the X-axis and values on the Y-axis. The values may be based on sensor-based data. It should be understood that the collected sensor-based data may require preprocessing. For example, the sensor-based data may include outliers, missing values, or corrupted values. A computer system may preprocess the data before performing data analysis. The values on the Y-axis may represent the sensor-based data after cleaning.

[0106] As shown, the computer system uses a change point detection algorithm to detect a first change point 1410 at time 50 and a second change point 1420 at time 150. Data from time "0" to time "50" and data after time "150" can represent times when the display of the display assembly is not moving. Data from time "50" to time "150" can represent times when the display is moving. The computer system can use the change point detection algorithm to detect the first change point 1410. Based on the detection, virtual markers can be displayed on the moving display. The computer system can further initialize a motion capture system to capture motion data from the display using the virtual markers. In response to detecting the second change point 1420, the computer system can send a signal to the motion capture system to stop collecting motion data. The computer system can then determine the physical pose of the display based on the motion data. The computer system can further generate inputs for a fitting model using the virtual model generated prior to the display movement and the physical pose of the display. The fitting model can use the input to transform the virtual pose of the display into an updated virtual pose that reflects the physical pose of the display after the movement.

[0107] FIG. 15 illustrates an example display assembly 1500 according to an embodiment of the present disclosure. The display assembly 1500 is comprised of multiple displays. As illustrated, the display assembly is comprised of nine displays, conveniently identified by a grid system. As illustrated, a display 1510 positioned at position (A,1) has been moved from a first physical pose (e.g., the first pose indicated by the dashed lines) to a second physical pose in which the display 1510 protrudes away from the surfaces of the other displays in the display assembly 1500. In the first physical pose, the surface of the display 1510 is flush with the surfaces of the other displays in the display assembly 1500, but in the second physical pose, this is not the case. For example, an actuator can be used to move the display 1510 so that it protrudes away from the surfaces of the other displays in the display assembly 1500.

[0108] The virtual marker 1520 is displayed on part or all of the display 1510. In examples where the virtual marker 1520 is displayed on a portion of the display, the remainder of the display 1510 can display a portion of the scene. For example, the remainder of the display 1510 can display a portion of the scene that is displayed before the display of the virtual marker 1520. This portion can be predefined, for example, in a central portion or any other portion of the display 1510 (e.g., the upper left corner), and can have a predefined size (e.g., a particular number of pixels in width and height). The virtual marker can be displayed based on detecting that the display 1510 is moving. In one example, the virtual marker 1520 can represent a multi-dimensional model (e.g., a two-dimensional model, a three-dimensional model, etc.) of a rigid body. The virtual marker 1520 can be presented in a particular location on the display (e.g., the center as shown in the figure, but other locations such as the upper left corner are possible) according to a particular size. Generally, the virtual marker 1520 does not use infrared technology unless each display is capable of emitting light in the infrared range. Alternatively, the virtual marker 1520 may comprise one or more virtual points that emit light in the human visible wavelength range, and a camera operating in that wavelength range may be used to capture one or more images of the virtual marker 1520 as it is presented. The camera may be, but need not be, a motion capture camera. The virtual marker 1520 may comprise at least three points, each of which may be a different color and / or a different shape, or even possibly unique to a particular display (e.g., a barcode, a QR code, a unique shape, etc.), so that a single camera may be sufficient to generate an image of the virtual marker 1520, which may be processed to determine the corresponding physical pose of the display. Alternatively, the virtual marker 1520 may comprise a single point, and three or more cameras may be used to generate images of the virtual marker 1520, which may be processed to determine the corresponding physical pose of the display.Alternatively, the virtual marker 1520 may be a virtual visual marker that encodes its dimensions, and a single camera may be sufficient to generate an image of this visual marker, which can be processed to determine the corresponding physical pose of the display. Alternatively, the virtual marker 1520 may have an asymmetric shape (e.g., a rectangular prism but not a square cube). Of course, combinations of various techniques are possible. For example, the virtual marker 1520 may include at least three points, and multiple cameras may be used to capture its images so that pose determination may have greater accuracy. In another example, the virtual marker 1520 may be updated, and one or more cameras may be used to generate one or more images corresponding to each update. In particular, the virtual marker 1520 may initially be presented as including at least three points and then updated to present a single point that multiple cameras may capture.

[0109] The motion capture system can further use virtual markers 1520 to identify display 1510 and capture motion data representing the movement of display 1510. While FIG. 15 shows display 1510 moving, it should be understood that other displays in display assembly 1500 can also be configured to move. Furthermore, a set of displays can be configured to move. For example, displays located at positions (B,1), (B,2), (C,1), and (C,2) can be configured to move. In this case, a virtual marker can be displayed on one or more displays. When multiple displays are moved, the virtual markers can be the same or different. Using different virtual markers can facilitate tracking of positional changes of each display and can enable the use of the same camera system. Nevertheless, it is possible to use the same camera system for tracking even when the same virtual marker is presented on different moving displays. In this case, the initial pose of each display is known and tracking can be relative to the initial pose, so that each final pose is associated with the initial pose, which is then associated with the display.

[0110] 16 is a diagram 1600 illustrating an example of determining a transform 1610 of a display assembly and an updated virtual model 1620, according to an embodiment of the present disclosure. In one example, virtual model 1630 represents a display assembly before a display of the display assembly is moved. Transform 1610 includes a set of functions (e.g., point-by-point rotation and / or translation along each dimensional axis, warping, twisting, bending, randomization, etc.) that can be applied to virtual model 1630, resulting in updated virtual model 1620. Updated virtual model 1620 can represent the physical pose of each display of the display assembly.

[0111] To generate the transform 1610, a physical pose 1640 of a display included in the display assembly is determined and input into a fitting model 1650 along with the virtual model 1620 (or more specifically, along with the corresponding virtual position). As shown, the physical pose 1640 includes the physical pose of the display 1660 before the display 1660 is moved and the physical pose of the display 1660 after the display 1660 is moved. In some cases, the physical pose includes only the physical pose of the display 1660 after the display 1660 is moved. The output of the fitting model 1650 includes parameters (e.g., coefficients) of the transform 1610. The physical pose 1640 may be derived based on motion capture data, image data (which may include motion capture data), and / or other positioning techniques.

[0112] The fitting model 1650 may be a data fitting model that iteratively estimates the parameters of the transformation 1610 so as to best fit the transformed virtual positions to the physical positions. Different types of data fitting models are possible, such as those based on the Levenberg-Marquardt nonlinear least squares algorithm, a chi-squared test algorithm, a curve fitting algorithm, a weighted least squares fitting algorithm, a polynomial regression algorithm, a Gauss-Newton algorithm, a shift-and-cut algorithm, a gradient algorithm, a Nelder-Mead (simplex) search algorithm, or an implementation of other types of fitting algorithms. Additionally or alternatively, a number of known virtual models and corresponding display assemblies can be used to train a machine learning model, such as a regression model or a convolutional neural network, to output transformation parameters. Once trained, the virtual model 1630 and the physical pose 1640 can be input into the machine learning model, which outputs parameters for the transformation 1610.

[0113] FIG. 17 is a table 1700 showing the differences between measured and calculated parameters. Table 1700 includes a column of test data containing values obtained by actual measurements and a column of calculated (CV) data containing values obtained using the techniques described herein. As can be seen, the test data value for the distance between two poses of the display is 2395.117210 millimeters (mm), while the CV value for the distance between the two poses of the display is 2396.601829 mm. The difference between these two values is 1.484619 mm. In addition, the test data values for the relative rotation between the two poses of the display are −0.482724°, −41.687118°, and 1.1130060°. The calculated values for the relative rotation between the two poses of the display are 0.578781°, −41.718646°, and −0.956472°.

[0114] 18-20 illustrate flows related to updating a virtual model in the context of content rendering that relies on a virtual model of a display assembly on which the content is displayed. The operations of the flows may be performed by a computer system (e.g., at least one processor, at least one computer-readable memory, etc.), such as computer system 140 of FIG. 1. Some or all of the instructions for performing the operations may be implemented as hardware circuits and / or stored as computer-readable instructions on a non-transitory computer-readable medium of the computer system. As implemented, the instructions represent components including circuitry or code executable by a processor of the computer system. The use of such instructions configures the computer system to perform particular operations described herein. Each circuit or code, in combination with an associated processor, represents means for performing the respective operation. While the operations are shown in a particular order, it should be understood that no particular order is required and that one or more operations may be omitted, skipped, performed in parallel, and / or reordered.

[0115] FIG. 18 illustrates an example process flow 1800 for determining an updated virtual model based on movement of a display, according to an embodiment of the present disclosure. In operation 1802, a computer system may collect motion capture data of a display moving from a first physical pose to a second physical pose. For example, the motion capture data may be generated by a motion capture system associated with the virtual production and transmitted to the computer system. The motion capture data may track the pose of a virtual marker on the display moved from the first physical pose to the second physical pose. The computer system may use one or more characteristics of the virtual marker to determine a presentation pose of the virtual marker. The presentation pose may correspond to the physical pose of the display. For example, a production crew may use actuators to move the display from the first pose to the second pose to create a desired effect on content rendered on the display assembly.

[0116] In operation 1804, the computer system can process the motion capture data to determine the physical pose of the moved display. The processing can depend on the collection technique. In one example using virtual markers, the motion capture data can be used to determine the physical coordinates of the display in the coordinate system of the camera of the motion capture system. For example, the computer system can use a PnP pose algorithm to determine the physical pose of the moved display (physical coordinates in the coordinate system of the motion capture system).

[0117] In operation 1806, the computer system may generate a transformation of the second physical pose of the display to the virtual pose of the display. In one example, the virtual pose of the virtual model and the physical pose of the moved display are then used to generate inputs to a fitting model that outputs parameters of the transformation, the transformation being associated with the display assembly.

[0118] In operation 1808, the computer system may update the virtual model to include the virtual pose. For example, the computer system may update the virtual model determined before the movement of the display to constitute a virtual pose for the display. The updated virtual model may include a virtual pose that corresponds to the second physical pose of the moved display.

[0119] At operation 1810, the computer system may use the updated virtual model to render content on a display assembly. The updated virtual model may be used by a game engine running on the computer system to render content, which is then displayed by the display assembly.

[0120] 19 shows an example of a process flow 1900 for determining initialization of a motion capture system of a virtual production, according to an embodiment of the present disclosure. In one example, flow 1900 may begin at operation 1902, where a computer system receives input to initialize the motion capture system. The input may be a user-based manual input or a sensor-based input. A user-based manual input may be received based on a user operating a switch that generates a control command to initialize the motion capture system.

[0121] Alternatively, the input may be sensor-based. The virtual production may include one or more sensors that collect streaming data from the display assembly. The streaming data may be received by a computer system that can analyze the streaming data to determine whether an indication exists that one or more displays of the display assembly have been moved. The indication may be determined based on detecting change points in the streaming data. For example, the computer system may use a forgetting factor-based change point detection algorithm to detect first and second change points in the streaming data. The first change point may represent an indication that a display has started to move. The second change point may represent an indication that a display has stopped moving.

[0122] At operation 1904, the computer system may initialize the motion capture system, which may include determining the position and orientation of each motion capture device in the system (e.g., based on the coordinate system of the motion capture system) and configuring the devices for capturing display motion (e.g., configuring the frame rate).

[0123] In some embodiments, the motion capture system captures frames including the display after the display is moved. For example, if the input is a sensor-based input, the motion capture system can capture frames after the second point. In other embodiments, the motion capture system can capture images from the display from when the display starts moving to when the display stops moving.

[0124] 20 illustrates an example flow 2000 for determining a transformation based on detecting movement of a display of a display assembly, according to an embodiment of the present disclosure. In one example, the flow may begin at operation 2002, in which a computer system determines that a display of a display assembly has changed from having a first physical pose to having a second physical pose. The display may be moved, for example, by an actuator, based on the desires of a production crew for the virtual production. The computer system may determine the change to the second pose based on processing motion data received from a motion capture system.

[0125] In operation 2004, the computer system can determine a second physical pose (e.g., physical coordinates in the coordinate system of the motion capture system) of the moved display. The second physical pose can be determined based on the motion data of the virtual tracker and / or the image data of the virtual tracker. In particular, the display indexes of these displays are used to look up and obtain the corresponding physical position data.

[0126] In operation 2006, the computer system may determine a virtual pose that corresponds to the second physical pose of the display.

[0127] In operation 2008, the computer system may generate a transformation. For example, a virtual pose corresponding to the second physical pose and the virtual model determined before the display movement are input into a fitting model, which then outputs parameters for the transformation. The output parameters may then be used to update the virtual model.

[0128] Figure 12 illustrates exemplary components of a computer system 2100, according to an embodiment of the present disclosure. Computer system 2100 is an example of computer system 140 of Figure 1. Although the components of computer system 2100 are shown as belonging to the same computer system 2100, computer system 2100 may also be distributed (e.g., among multiple user devices).

[0129] The computer system 2100 includes at least a processor 2102, a memory 2104, a storage device 2106, input / output peripherals (I / O) 2108, communication peripherals 2110, and an interface bus 2112. The interface bus 2112 is configured to communicate, transmit, and transfer data, control, and commands between various components of the computer system 2100. The memory 2104 and the storage device 2106 include computer-readable storage media such as RAM, ROM, electrically erasable programmable read-only memory (EEPROM), hard drives, CD-ROMs, optical storage devices, magnetic storage devices, electronic non-volatile computer storage such as Flash memory, and other tangible storage media. Any such computer-readable storage media may be configured to store instructions or program code embodying aspects of the present disclosure. The memory 2104 and the storage device 2106 also include computer-readable signal media. The computer-readable signal media includes a propagated data signal having computer-readable program code embodied therein. Such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any combination thereof. Computer-readable signal media includes any computer-readable medium that can communicate, propagate, or transmit a program for use in connection with computer system 2100, other than computer-readable storage media.

[0130] Additionally, memory 2104 includes an operating system, programs, and applications. Processor 2102 is configured to execute stored instructions and includes, for example, a logic processing unit, microprocessor, digital signal processor, and other processors. Memory 2104 and / or processor 2102 may be virtualized and / or hosted within another computer system, for example, in a cloud network or data center. I / O peripherals 2108 include user interfaces such as keyboards, screens (e.g., touch screens), microphones, speakers, other input / output devices, and computing components such as graphics processing units, serial ports, parallel ports, universal serial buses, and other input / output peripherals. I / O peripherals 2108 are connected to processor 2102 via any of the ports coupled to interface bus 2112. Communications peripherals 2110 are configured to facilitate communication between computer system 2100 and other systems over a communications network and include, for example, network interface controllers, modems, wireless and wired interface cards, antennas, and other communications peripherals.

[0131] It should be apparent to those skilled in the art that many further modifications beyond those already described are possible without departing from the inventive concepts herein. Accordingly, the present subject matter should not be limited except in the spirit of the appended claims. Moreover, in interpreting both this specification and the claims, all terms should be interpreted in the broadest possible manner consistent with the context. In particular, the terms "comprises" and "comprising" should be interpreted as referring to elements, components, or steps in a non-exclusive manner, indicating that a referenced element, component, or step can be present in, utilized with, or combined with other elements, components, or steps not expressly referenced. When a specification or claim refers to at least one of something selected from a group consisting of A, B, C-N, the text should be interpreted as requiring only one element from that group, and not A+N, B+N, etc.

Claims

1. receiving, by at least one processor, motion capture data for a display of a plurality of displays, the display moving from a first physical pose to a second physical pose; processing, by at least one processor, the motion capture data to determine coordinates of the second physical pose; generating, by the at least one processor, a transformation of the second physical pose of the display into a virtual pose of the display; updating, by the at least one processor, a virtual model of the plurality of displays, the virtual model including the virtual pose of the displays; and rendering, by the at least one processor, content on the display based on the updated virtual model; and 11. A computer-implemented method comprising:

2. receiving input for initializing a motion capture system to generate the motion capture data, the motion capture system comprising a motion capture device; initializing the motion capture system based on the input, wherein initializing the motion capture system includes determining a position and orientation of the motion capture device; The computer-implemented method of claim 1 further comprising:

3. the input is a user-based input, and the computer-implemented method comprises: projecting a virtual marker onto the display based on receiving the user-based input; determining a presentation pose of the virtual marker, the presentation pose corresponding to the second physical pose; determining the coordinates of the second physical pose based on the presented pose; The computer-implemented method of claim 2 further comprising:

4. the input is a sensor-based input, and the computer-implemented method comprises: receiving streaming data from a sensor configured to collect data related to the display; detecting a first change point and a second change point from the streaming data; determining that the display is in the second physical pose based on detecting the second change point; processing the motion capture data to determine the coordinates of the second physical pose based on the determination; The computer-implemented method of claim 2 further comprising:

5. The computer-implemented method of claim 4 , wherein the first change-point and the second change-point are detected using a forgetting factor based change-point detection algorithm.

6. The computer-implemented method of claim 4 , wherein determining the coordinates of the second physical pose comprises using a Point-n-Point (PnP) pose computation algorithm.

7. 5. The computer-implemented method of claim 4, wherein determining the coordinates of the second physical pose comprises determining the coordinates in a coordinate system of a motion capture device used to capture the motion capture data.

8. The computer-implemented method comprises: starting to receive the motion capture data based on detecting the first change point; While the display is moving, continuing to receive the motion capture data; and continuously processing the motion capture data to determine current coordinates of the display; generating a current transform using the current coordinates of the display; updating a virtual model of the plurality of displays to include a current virtual pose of the displays, the current virtual pose being associated with the current coordinates of the displays; The computer-implemented method of claim 4 further comprising:

9. 1. A system comprising: one or more processors; when executed by the one or more processors, receiving motion capture data of a display of the plurality of displays moving from a first physical pose to a second physical pose; processing the motion capture data to determine coordinates of the second physical pose; updating a translation of the second physical pose of the display to a virtual pose of the display; updating a virtual model of the plurality of displays to include the virtual pose of the displays; Rendering content on the display based on the updated virtual model. one or more memories storing instructions for configuring the system to: A system comprising:

10. The instructions, when executed by the one or more processors, receiving an input for initializing a motion capture system to capture the motion capture data, the motion capture system comprising a motion capture device; initializing the motion capture system based on the input, wherein initializing the motion capture system includes determining a position and orientation of the motion capture device. The system of claim 9 , further configured to:

11. the input is a user-based input, and the instructions, when executed by the one or more processors, projecting a virtual marker onto the display based on receiving the user-based input; determining a presentation pose of the virtual marker, the presentation pose corresponding to the second physical pose; determining the coordinates of the second physical pose based on the presented pose; The system of claim 10 , further configured to:

12. 12. The system of claim 11, wherein the motion capture system comprises at least one motion capture device, the virtual markers include a plurality of virtual points projected onto the display, and the at least one motion capture device is configured to capture motion capture data based on tracking of the plurality of virtual points.

13. 12. The system of claim 11, wherein the motion capture system comprises a plurality of motion capture devices, the virtual marker comprises at least one virtual point, and the plurality of motion capture devices are configured to capture motion capture data based on tracking of the at least one virtual point.

14. The system of claim 11 , wherein the virtual marker covers a portion of a display, and a remainder of the display projects a portion of a scene.

15. The system of claim 11 , wherein the virtual marker covers the entire display.

16. The instructions, when executed by the one or more processors, starting to receive the motion capture data based on the detection of the first change point; While the display is moving, continuing to receive the motion capture data; continuously processing the motion capture data to determine current coordinates of the display; generating a current transformation using the current coordinates of the display; generating a current virtual model of the plurality of displays including a current virtual pose of the displays, the current virtual pose being associated with the current coordinates of the displays; The system of claim 12 , further configured to:

17. One or more non-transitory computer-readable storage media storing instructions that, when executed on a system, cause the system to: receiving motion capture data of a display of a plurality of displays moving from a first physical pose to a second physical pose; processing the motion capture data to determine coordinates of the second physical pose; generating a transformation of the second physical pose of the display into a virtual pose of the display; updating a virtual model of the plurality of displays to include the virtual pose of the displays; Rendering content on the display based on the updated virtual model; and One or more non-transitory computer-readable storage media for performing operations including:

18. The instructions, when executed by the one or more processors, receiving an input to initialize a motion capture system to capture the motion capture data, the motion capture system comprising a motion capture device; initializing the motion capture system based on the input, wherein initializing the motion capture system includes determining a position and orientation of the motion capture device; 20. The one or more non-transitory computer-readable storage media of claim 17, further configuring the system to perform operations including:

19. the input is a user-based input, and the instructions, when executed by the one or more processors, projecting a virtual marker onto the display based on receiving the user-based input; determining a presentation pose of the virtual marker, the presentation pose corresponding to the second physical pose; determining the coordinates of the second physical pose based on the presented pose; 20. The one or more non-transitory computer-readable storage media of claim 18, further configuring the system to perform operations including:

20. the input is a user-based input, and the instructions, when executed by the one or more processors, receiving streaming data from a sensor configured to collect data related to the display; detecting a first change point and a second change point from the streaming data; determining that the display is in the second physical pose based on detecting the second change point; processing the motion capture data to determine the coordinates of the second physical pose based on the determination; 20. The one or more non-transitory computer-readable storage media of claim 18, further configuring the system to perform operations including:

21. determining, by at least one processor, a first pose of a first display of a plurality of displays included in a display assembly, the display assembly being configured to display content on the plurality of displays; determining, by the at least one processor, a virtual model of the display assembly, the virtual model being stored in a computer readable memory and including a virtual representation of each one of the plurality of displays; determining, by the at least one processor, a transformation based on the first pose of the first display and a first virtual representation of the first display in the virtual mode; rendering, by the at least one processor, the content on at least some of the displays according to the transformation and the virtual model; 11. A computer-implemented method comprising:

22. Determining the first pose includes first: receiving, by at least one processor, first motion capture data of a physical marker placed at a first location on the first display, wherein the first pose is determined based on the first motion capture data; 22. The computer-implemented method of claim 21, comprising:

23. Determining the first pose comprises: receiving, by at least one processor, second motion capture data of the physical marker positioned at a second location on a second display according to a predefined motion path of the physical marker on the display assembly; determining, by at least one processor, a change between the first motion capture data and the second motion capture data; determining, by at least one processor, that the first motion capture data corresponds to the first location based on the change; determining, by at least one processor, a position and rotation of the physical marker at the first location based on the first motion capture data; generating, by at least one processor, the first pose of the first display by including the position and the rotation in the first pose; 23. The computer-implemented method of claim 22, further comprising:

24. Determining the first pose comprises: determining, by at least one processor, a timing of the first motion capture data; determining, by at least one processor, that the timing corresponds to the first location based on the predefined motion path; 24. The computer-implemented method of claim 23, further comprising:

25. Determining the first pose comprises: displaying, by at least one processor, a virtual marker on the display; receiving, by at least one processor, motion capture data of the virtual marker, wherein the first pose is determined based on the motion capture data; 22. The computer-implemented method of claim 21, comprising:

26. Determining the first pose comprises: displaying, by at least one processor, a display identifier on the display; determining, by at least one processor, that the motion capture data corresponds to the display based on the display identifier; 26. The computer-implemented method of claim 25, further comprising:

27. determining, by at least one processor, a second pose of a second display of the plurality of displays, the transformation being further determined based on the second pose and a second virtual representation of the second display in the virtual model; 22. The computer-implemented method of claim 21, further comprising:

28. determining, by at least one processor, a pose corresponding to each one of the plurality of displays, the transformation being further determined based on the pose and the virtual representation of the plurality of displays in the virtual model; 22. The computer-implemented method of claim 21, further comprising:

29. 29. The computer-implemented method of claim 28, wherein the virtual representations show virtual poses each corresponding to one of the plurality of displays, and the transformation is determined by at least fitting the poses of the plurality of displays to the virtual pose shown by the virtual model.

30. the plurality of displays includes a first subset of first displays within a field of view of a camera device, the content is further rendered based on the field of view, and the computer-implemented method further comprises: determining, by at least one processor, first poses each corresponding to one of the first displays; determining, by at least one processor, first virtual representations from the virtual model, each corresponding to one of the first displays, the transformation being further determined based on the first pose and the first virtual representation; 22. The computer-implemented method of claim 21, further comprising:

31. 31. The computer-implemented method of claim 30, wherein the plurality of displays includes a second subset of second displays that are outside the field of view of the camera device, and the transformation is determined independently of a second pose of the second displays and a second virtual representation of the second displays in the virtual model.

32. generating, by the at least one processor, an updated virtual model by correcting a pose error of the virtual model using the transformation, wherein the content is rendered based on the updated virtual model; 22. The computer-implemented method of claim 21, further comprising:

33. 1. A system comprising: one or more processors; when executed by the one or more processors, determining a first pose of a first display of a plurality of displays included in a display assembly, the display assembly configured to display content on the plurality of displays; determining a virtual model of the display assembly, the virtual model being stored in the one or more memories and including a virtual representation of each one of the plurality of displays; determining a transformation based on the first pose of the first display and a first virtual representation of the first display in the virtual mode; rendering the content on at least some of the displays according to the transformation and the virtual model; one or more memories storing instructions for configuring the system to: A system comprising:

34. said executing said instructions receiving motion capture data from a motion capture system configured to track motion of a physical marker according to a predefined motion path on the display assembly, the motion capture data including first motion capture data corresponding to the physical marker being placed at a first location on the first display, and determining the first pose based on the first motion capture data; 34. The system of claim 33, further configured to:

35. the motion capture data includes second motion capture data corresponding to the physical marker being positioned at a second location on a second display of the display assembly, and the execution of the instructions includes: determining a change between the first motion capture data and the second motion capture data; determining, based on the change, that the first motion capture data corresponds to the first location; determining a position and rotation of the physical marker at the first location based on the first motion capture data; generating the first pose of the first display by including the position and the rotation in the first pose; 35. The system of claim 34, further configured to:

36. said executing said instructions determining a timing of the first motion capture data; determining that the timing corresponds to the first location based on the predefined motion path; 36. The system of claim 35, further configured to:

37. One or more non-transitory computer-readable storage media storing instructions that, when executed on a system, cause the system to: determining a first pose of a first display of a plurality of displays included in a display assembly, the display assembly being configured to display content on the plurality of displays; determining a virtual model of the display assembly, the virtual model being stored in a computer readable memory and including a virtual representation of each one of the plurality of displays; determining a transformation based on the first pose of the first display and a first virtual representation of the first display in the virtual mode; rendering the content on at least some of the displays according to the transformation and the virtual model; One or more non-transitory computer-readable storage media that cause the computer to perform operations including:

38. The operation is determining a pose each corresponding to one of the plurality of displays, the transformation being further determined based on the pose and the virtual representation of the plurality of displays in the virtual model; 38. The one or more non-transitory computer-readable storage media of claim 37, further comprising:

39. the plurality of displays includes a first subset of first displays within a field of view of a camera device, the content is further rendered based on the field of view, and the action comprises: determining first poses each corresponding to one of the first displays; determining, from the virtual model, first virtual representations each corresponding to one of the first displays, the transformation being further determined based on the first pose and the first virtual representation; 38. The one or more non-transitory computer-readable storage media of claim 37, further comprising:

40. The operation is generating an updated virtual model by correcting pose errors of the virtual model using the transformation, wherein the content is rendered based on the updated virtual model; 38. The one or more non-transitory computer-readable storage media of claim 37, further comprising: