Determination of Camera Control Points for Virtual Production
By determining the offset between a marker and the control point using a camera model, the method addresses inaccuracies in estimating the pose of camera control points, improving the accuracy and quality of content rendering in virtual production systems.
Patent Information
- Application Number
- JP2024558353
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-18
- Filing Date
- 2023-05-23
- Publication Date
- 2025-07-30
AI Technical Summary
Existing virtual production systems face inaccuracies in estimating the pose of camera control points due to the physical separation of markers from the control points, leading to misalignment and quality issues in content rendering.
A method to determine the pose of a control point by using a camera model to calculate an offset between a marker attached to the camera device and the control point, allowing for accurate tracking and rendering of content based on the control point's pose.
Improves the accuracy of content rendering by accurately estimating and tracking the pose of the control point, reducing misalignment and enhancing the quality of the output.
Smart Images

Figure 2025524322000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 390,255, filed on July 18, 2022, entitled "DETERMINING A CAMERA CONTROL POINT FOR VIRTUAL PRODUCTION", the content of which is hereby incorporated by reference in its entirety for all purposes.
[0002] The field of the present invention relates to virtual production and camera control technology.
Background Art
[0003] The background description includes information that may be useful in understanding the subject matter of the present invention. None of the information provided in this specification is admitted to be prior art or prior art recognized by the applicant, or relevant to the subject matter of the currently claimed invention, or that any publication specifically or implicitly referred to is prior art or prior art recognized by the applicant.
[0004] Virtual production generally includes a virtual stage that presents content related to a scene, a camera device that generates movie data by capturing video of a person, an object, and the content, and a motion capture system that tracks the camera, the person, and / or the object. The content can be dynamic (e.g., video content that changes over time) and / or its presentation can be adjusted based on tracking.
[0005] All publications identified in this specification are hereby incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference. If the definition or use of a term in an incorporated reference is inconsistent with or contrary to the definition of that term provided herein, the definition of that term provided herein shall apply and the definition of that term in the reference shall not apply.
[0006] In some embodiments, numerical values used to describe and claim particular embodiments of the subject matter of the invention, for example, amounts of data or units, should be understood to be modified in some instances by the term "about". Accordingly, in some embodiments, the numerical parameters set forth in the written description and attached claims are approximations that may vary depending upon the desired properties sought to be obtained by a particular embodiment. In some embodiments, the numerical parameters should be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. Notwithstanding that the numerical ranges and parameters setting forth the broad scope of some embodiments of the subject matter of the invention are approximations, the numerical values set forth in the specific examples are reported as precisely as practicable. The numerical values presented in some embodiments of the subject matter of the invention may include certain errors necessarily resulting from the standard deviation found in the respective test measurements.
[0007] Unless the context indicates otherwise, all ranges defined herein should be construed to include their endpoints, and open-ended ranges should be construed to include only commercially practical values. Similarly, all lists of values should be considered to include intermediate values unless the context indicates otherwise.
[0008] As used throughout this specification and the claims that follow, the meanings of "a", "an", and "the" include plural references unless the context clearly dictates otherwise. Also, as used herein, the meaning of "in" includes "in" and "on" unless the context clearly dictates otherwise.
[0009] The recitation of a range of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each separate value is incorporated herein as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., "such as") provided herein with respect to particular embodiments is merely intended to better illuminate the subject matter of the invention and does not impose a limitation on the scope of the subject matter of the claims. No language in this specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
[0010] The grouping of alternative elements or embodiments of the subject matter of the invention disclosed herein should not be construed as a limitation. Members of each group may be referred to and claimed individually, or in any combination with other members of the group or other elements found herein. One or more members of a group may be included in, or deleted from, the group for reasons of convenience and / or patentability. When such inclusion or deletion occurs, the specification is considered to include the group as modified so as to satisfy the written description of all Markush groups used in the appended claims.
[0011] It should be understood that many of the basic technical features provided in the following specification are presented to enable a compact consideration of the subject matter of the disclosed invention. Some of the basic technical features described herein may appear unclear, but in many cases, such features can be considered within the understanding of those skilled in the art. Therefore, the presentation of such background art should not be considered limiting.
Summary of the Invention
Means for Solving the Problems
[0012] The subject matter of the present invention provides a system, apparatus, and method for rendering content based on control points of a camera device. Embodiments of the present disclosure are directed, inter alia, to determining the pose of a control point (e.g., its focus) of a camera device.
[0013] One embodiment of the present invention includes a method of determining a first pose of a marker in a real-world space via at least one processor, the marker being attached to a camera device. The method may further include determining, via at least one processor, an offset between the marker and a control point of the camera device based on the first pose and at least one calibration parameter of the camera device. Additionally, the method may include setting, via at least one processor, a second pose of a virtual camera in a virtual space in a computer-readable memory based on the offset, the virtual camera representing the camera device in the virtual space, the virtual space representing the real-world space. Further, the method may include rendering, via at least one processor, content on a display based on the second pose of the virtual camera, the content being captured by the camera device when the content is presented in the real-world space.
[0014] Various objects, features, aspects, and advantages of the subject matter of the present invention will become more apparent from the following detailed description of the preferred embodiments, along with the accompanying drawings in which like numerals represent like components.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
DETAILED DESCRIPTION OF THE INVENTION
[0016] It should be noted that any language referring to a computer should be read to include any suitable combination of computing devices, including servers, interfaces, systems, databases, agents, peers, engines, controllers, modules, or other types of computing devices operating individually or collectively. A computing device should be understood to comprise at least one processor configured to execute software instructions stored on a tangible non-transitory computer-readable storage medium (e.g., hard drive, FPGA, PLA, solid state drive, RAM, flash, ROM, etc.). The software instructions or set of software instructions configure or program the computing device or their processors to provide the roles, responsibilities, or other functions as described hereinafter with respect to the disclosed apparatus or system. Further, the disclosed technology may be embodied as a computer program product including a non-transitory computer-readable medium storing software instructions or a set of software instructions that cause one or more processors to execute the disclosed steps related to the implementation of computer-based algorithms, processes, methods, or other instructions. In some embodiments, various servers, systems, databases, or interfaces may use standardized protocols or algorithms based on, for example, HTTP, HTTPS, TCP, UDP, FTP, SNMP, IP, AES, public key-private key exchange, web services or RESTful APIs, known financial application protocols, or other electronic information exchange methods to exchange data. Data exchange between devices may be performed via a packet-switched network, the Internet, a LAN, a WAN, a VPN, or another type of packet-switched network, a circuit-switched network, a cell-switched network, or another type of network (wired or wireless).
[0017] When used throughout this specification and the following claims, when a system, engine, server, agent, device, module, or other computing element is described as being configured to perform or execute a function on data in memory, the meaning of "configured to" or "programmed to" is defined as one or more processors or cores of the computing element being programmed by a set of software instructions stored in the memory of the computing element to execute a set of functions on target data or data objects stored in the memory. It should be understood that the combination of software and hardware operating in concert creates a dedicated set of physical real-world structures that provide utility to one or more users that would not exist outside the scope of physical, tangible real-world assets.
[0018] It should be understood that the disclosed techniques provide many advantageous technical effects, including improving the calibration of a camera device and improving the use of control points for a given camera device. For example, content can be rendered on a set of displays based on the pose of a control point, where the pose can be the position and / or orientation of the control point in the real-world space. Calibration includes estimating an offset between the control point and at least one marker attached to the camera device. Given this offset, the pose of the control point can be tracked accurately over time by tracking the marker, thereby improving the quality of content rendering.
[0019] Embodiments of the present disclosure are particularly directed to determining the pose of a control point (e.g., its focus) of a camera device. In one example, at least one marker is attached to the camera device, and the pose of the marker (e.g., its position, orientation, etc.) in the real-world space can be tracked over time. Due to the physical characteristics of the camera, the marker cannot be attached to the same location as the control point of the camera device, so the pose of the marker may not accurately represent the pose of the control point. However, certain applications rely on the pose of the control point to generate an output (e.g., the presented content), and in particular, rely on the output rendered from the perspective of the control point. Thus, inaccuracies in estimating the pose of the control point can affect the accuracy of the output (e.g., causing quality issues related to the presentation of the content, misalignment between real and virtual items, etc.).
[0020] To improve the estimation of the pose of the control point, a camera model including, among other things, the external parameters, internal parameters, and / or distortion model of the camera device is used. Given the camera model and the pose of the marker, an offset between the pose of the marker and the pose of the control point can be determined, and the offset can include a rotational offset and / or a positional offset. Instead of setting the pose of the control point to the pose of the marker, the pose of the control point can be estimated by offsetting the pose of the marker by the corresponding rotational offset and / or positional offset, if applicable. As a result, the pose of the control point is estimated more accurately and tracked over time, thereby improving the accuracy of the output (e.g., reducing quality issues related to the presentation of the content).
[0021] The camera model can be determined according to a camera calibration procedure. In one example, the camera calibration procedure involves tracking the pose of an object in the real-world space, while the camera device generates images of the object in different poses (e.g., in its video stream, still images, etc.). Each pose is associated with at least one image. The pose of the camera device can be determined by initializing and using the camera model to project at least the pose of the object into the image plane and determining a projection error. In particular, the pose of the object (having, for example, 3D coordinates, etc.) is projected onto the image plane, and this projection is compared with the pixel coordinates of the object in the associated image to determine the corresponding projection error with respect to the real-world pose. The projection error is determined across different pose-image pairs, and a fitting model can be used to update the camera model by reducing or mitigating the projection error.
[0022] For illustration purposes, consider an example of a use case for virtual production. A camera device is used to generate video frames of a scene in the real-world space. The content is typically presented on a display assembly that operates as a background and forms part of the scene. To render the background content on the display assembly such that it appears natural from the focus viewpoint, the content itself and / or its presentation can be adjusted based on the pose of the focus of the camera device (an example of a control point). By using the above techniques, a motion capture system is used in a camera calibration procedure to generate a camera model, and markers are attached to and tracked by the camera device, thereby determining the offset between the pose of the markers and the pose of the focus. In the use case of virtual production, the virtual camera device represents the camera device in the virtual space, while the virtual space can represent the real-world space. The pose of the virtual camera in the virtual space is set to correspond to the pose of the focus. The game engine can render the content based on the pose of the virtual camera, and the rendered content is presented on the display assembly. As a further illustration, the camera model includes a distortion model. As part of the rendering, the game engine can distort a portion (or the entirety) of the content presented on a particular sub-area (or across the entire display assembly) of the display assembly based on the pose of the virtual camera and the distortion model.
[0023] For clarity of explanation, various embodiments of the present disclosure are described in the context of use cases for virtual production (e.g., production of movies using virtual backgrounds, production of augmented reality content, etc.). However, the embodiments are not so limited and are equivalently applicable to other use cases such as virtual reality, augmented reality, mixed reality, content projection (e.g., in a home theater, cinema, or building), performance stages (e.g., music concerts), etc. Generally, embodiments of the present disclosure enable improvement to the output of an application or system that depends on control points of a camera device to generate an output. A control point can be a physical point of the camera device, and the physical point has a corresponding virtual point in the camera model of the camera device, and the virtual point can be used when controlling the camera device itself or the manner of use of the camera device (e.g., for rendering content). The pose of a control point can extend along one or more dimensions and can cover multiple degrees of freedom (DOF) (e.g., at least 6 DOF) for operating the camera device. For example, the pose can indicate the position and orientation of the camera device along each axis of a three-dimensional system and / or can correspond to the movement of the camera over time (in which case the movement can be estimated based on speed, acceleration, state information, etc. to predict the current, past, and / or next pose). A control point can be converted to an actual dimensionless point but does not necessarily have to be. Instead, embodiments can approximate the control point as a dimensionless point and can be used to control a volume (e.g., an elliptical volume, etc.). Thus, embodiments enable optimization based on specific properties of the camera device. Also described herein are various embodiments by using focus as an example of a control point. However, the embodiments are not so limited and are equivalently applicable to other types of control points such as, for example, those corresponding to the center point (or principal point) of the image plane of the camera model or points along the focal length between the center point and the focus. Further, the camera device can also have two or more lenses (e.g., a stereo camera) and / or can be a camera array of multiple individual cameras.In such a situation, the control points can be associated with a plurality of lenses (which can be, for example, the geometric center of the focus of the lens) and / or a camera array (which can be, for example, the geometric center of the individual control points of an individual camera).
[0024] FIG. 1 shows an example of a virtual production system 100 according to an embodiment of the present disclosure. As shown, the virtual production system 100 includes, among other things, a display assembly 110, a camera device 120, motion capture devices 130A, 130B, 130C (generally referred to by the numeral "130"), and a computer system 140. The display assembly 110 can be configured as a virtual stage that presents the content 112 and defines the volume in which the camera device 120 is disposed (a plurality of such camera devices 120 are also possible). The presentation of the content 112 can be controlled by a game engine (such as the UNITY game engine, the UNREAL game engine, Blender, etc.) running on the computer system 140. For example, the content 112 itself and / or the presentation parameters (such as panning, angling, tilting, etc.) can be controlled based on a number of factors. Among these factors is the pose of the camera device 120 within the volume. The motion capture device 130 can generate motion capture data that is processed to determine the poses not only of the objects 150 (150A, 150B, 150C, generally referred to by the numeral "150" and which can be an actor, stage furniture, set furniture, equipment, etc.) but also of the camera device 120. Such motion capture data can be processed and used to control the content 112 presented on the display assembly 110.
[0025] In one example, the display assembly 110 includes a plurality of displays arranged to form a content presentation screen or “wall”. The content presentation screen has different shapes (e.g., a curved surface, a flat intersecting surface, etc.) and can surround an area where the object 150 can be placed, thereby defining a volume. The content presentation screen is shown in a stationary installation form, but in some embodiments, the screen or individual displays may be more dynamic and, in some cases, may move around the production volume. The content presentation screen can be used to present an interactive and dynamic scene. The object 150 can interact with such a scene, and the content 112 can be updated based on the interaction (e.g., the actor 150A can interact with the virtual object presented in the content 112). The display assembly 110 is shown as having a vertical position (e.g., set as a curved wall), but the display assembly 110 (or a second display assembly) can be additionally or alternatively placed in other positions (e.g., a horizontal position for defining a ceiling or a floor, different geometric shapes, 4π steradian coverage, etc.). Further, the volume can include a slide set of displays that can be moved in and out of the volume to define a specific shape. For example, the volume can be formed in a horseshoe shape, and the slide set can be arranged to close the open portion of the horseshoe shape. In this way, the camera device 120 within the volume can be surrounded by a 360-degree full display. Other shapes are possible, whereby the volume can be, for example, a hemisphere, a perfect sphere (i.e., 4π steradians), a cylinder, a cube. The display assembly 110 can form a volume for virtual production. This volume is about 2,230m 2may include a virtual production space, and the display assembly 110 forms a curved LED wall that is approximately 16 m wide × 20 m long and elliptical at 270 degrees (a 360-degree configuration is also possible). In a particular exemplary use case, each display of the display assembly 110 is a BLACK PEARL 2 display available from ROE CREATIVE DISPLAY. In this exemplary use case, the screen is a flat panel with an LED surface mount diode (SMD) configuration, a magnesium frame with a magnetic connector, and a locking system, having dimensions of 500×500×90 mm (height × width × depth), a resolution of 176×176 (horizontal × vertical), and a pixel pitch of 2.84 mm.
[0026] The camera device 120 can be a movie camera attached to a rig (e.g., a floor rig and / or a ceiling rig) that can be repositioned within the volume, or a movable rig (e.g., a tripod, a gimbal, etc.). In this way, the camera device 120 can be configured to capture a scene by generating video data (and optionally audio data) that shows one or more of the objects 150 and / or a part or all of the content 112 presented on the display assembly 110, particularly from multiple viewpoints. The camera device 120 can have a high resolution (e.g., 4K, 6K, 8K, 12K, etc.) and can be available, for example, from BLACKMAGIC (e.g., URSA MINI PRO 12K, STUDIO CAMERA 4K PLUS, STUDIO CAMERA 4K PRO, URSA BROADCAST G2, etc.), ARRI (e.g., ALEXA MINI LF, ALEXA LF, ALEXA MINI, ALEXA SXT W, AMIRA, AMIRA LIVE, ARRI MULTICAM SYSTEM, etc. with ARRI SIGNATURE PRIME 35mm T1.8 lenses, ARRI SIGNATURE PRIME 75mm T1.8 lenses, etc. attached).
[0027] Marker 122 can be attached to the camera device 120 (e.g., removably attached to its upper surface) and can be tracked by a motion capture system. Tracking may involve determining the pose of marker 122. This pose may correspond to the pose of camera device 120. Marker 122 may include a plurality of tracking points (each corresponding to an individual marker), and as a result, marker 122 is a marker assembly that constitutes a rigid body to which the tracking points are attached. In one example, the rigid body comprises a base having an arm extending therefrom, and each tracking point is connected to the end of the arm. Six tracking points may be provided, each of which may be an individual marker of VICON (e.g., Pearl hard marker) or OPTITRACK (e.g., M3 market, M4 marker, etc.).
[0028] The motion capture device 130 can be a motion capture camera (e.g., an infrared camera) and / or other types of motion sensors (e.g., depth sensors) that are part of a motion capture system configured to track motion within a volume. The motion capture system can be available, for example, from VICON (e.g., using VANTAGE, VERO, VUE, VIPER, VIPERX cameras, etc., and SHOGUN software, etc.) or OPTITRACK (e.g., using PRIME-X 41, PRIME, SLIM-X, SLIM, FLEX cameras, etc., and UNREAL PLUGIN, UNITY PLUGIN, MOTION BUILDER PLUGIN, OPTICAL MOTION CAPTURE SOFTWARE, MAYA PLUGIN software, etc.). The motion capture system uses infrared technology, but any trackable markers, such as radio frequency markers (using radio frequency technology) or vision markers (using, for example, computer vision technology), can be used alone or in combination. The motion of an object can optionally be tracked by using a motion tracker attached to the object (where marker 122 can be used to track the motion of camera device 120). Tracking can involve determining the position of an object within a volume by determining the position and rotation of the object over time. The coordinate system of the motion capture system (e.g., a Cartesian coordinate system or any other coordinate system) can be defined with respect to any origin within the volume.
[0029] Computer system 140 may be configured to process at least a portion of the motion capture data and optionally a portion of the video data. For example, a game engine may use a virtual model of the display assembly 110, a virtual camera that models the camera device 120, and the pose of the camera device 120 to render the content 112. In one example, the pose of the virtual camera is set to the pose of the camera device 120. As will be further described below, the pose of the camera device 120 may be determined by offsetting the pose of the marker 122. Rendering may involve synthesizing multi-dimensional images and / or image frames (e.g., 2D, 3D, etc.) and / or applying visual transformations (e.g., distortion, keystone, skew, and / or the reverse thereof) to present them as content 112 on the display assembly. Computer system 140 may also be configured for prediction processing. For example, given the motion history of the camera device 120, computer system 140 can perform extrapolation to predict the next set of poses (e.g., future motion) of the camera device 120. This prediction may also be based on data (e.g., choreography data) that indicates the sequence of events in the scene (which may be included as part of an executable script). Given the prediction, computer system 140 can calculate (or pre-calculate) the transformations required to render the content (e.g., based on the offset between the camera marker 122 and the focus of the camera device 120). By doing so, the latency between the actual pose and the adjustment of the content rendering can be reduced because computer system 140 does not need to first observe the actual pose to calculate a transformation that predicts the actual pose that will subsequently be required.
[0030] Figure 2 shows an example of tracking camera device 210 according to an embodiment of the present disclosure. Marker 212 is attached to camera device 210 and enables the pose of camera device 210 to be tracked over time by a motion capture system including motion capture device 220. Camera device 210 generates image data 214 and transmits it to computer system 230. Image data 214 represents a video stream captured by camera device 210. Computer system 230 also receives motion capture data 222 indicating the pose of camera device 210. A video stream synchronization process 232 is executed on computer system 230 (e.g., as a process such as a game engine, some other rendering engine, etc.) to synchronize the rendering of the content with the pose of camera device 210. When presented (e.g., on a display assembly, an individual display, an aggregate of displays, etc.), the content is captured within the video stream of camera device 210. Camera device 210, marker 212, motion capture device 220, and computer system 230 are each examples of camera device 120, marker , motion capture device 130, and computer system 140 of FIG. 1 respectively.
[0031] In one example, a motion capture system implements a particular motion capture technology, such as any one or a combination thereof of infrared (IR) technology, radio frequency (RF) technology, and / or computer vision (CV) technology (e.g., available from a library of programming functions such as the Open Source Computer Vision (OpenCV) library). The IR technology may rely on active markers and / or passive markers that emit (e.g., transmit and / or respond to) IR signals detectable by the motion capture system. The RF technology may rely on RF beacons and / or RF signal triangulation. The CV technology may rely on image processing and object detection by one or more machine learning (ML) models (e.g., available from a library of programming functions such as the Scikit-Learn library).
[0032] In one example, marker 212 is a rigid body that implements motion capture technology according to a motion capture system. For example, in the case of IR technology, marker 212 may include one or more infrared emitting (active or passive) points (each using a different infrared frequency). Generally, the more points there are, the more accurate the pose estimation can be. In one example, marker 212 includes a single point detectable by an infrared motion capture camera. In this case, at least three infrared motion capture cameras are required to detect the pose of marker 212. In particular, each of the three cameras will generate a two-dimensional image showing the position of the marker in two dimensions. Since the position, orientation, and field of view of each camera are known, the three-dimensional vector in which marker 212 is located can be determined from the set of three two-dimensional positions. In another example, marker 212 includes a plurality of points detected by an infrared motion capture camera. In this case, a single infrared motion capture camera may be sufficient to detect the pose of marker 212. In particular, the relative pose of the points is known a priori, and this knowledge is used in the processing of the image generated by the infrared motion capture camera. In virtual production, depending on the virtual production setup, a large number of motion capture cameras may be used. Referring back to the exemplary virtual production example of FIG. 1, more than 12 motion capture cameras, including 50 to 60 or more motion capture cameras, may be used.
[0033] Of course, technologies other than infrared can also be used. For example, in RF technology, marker 212 can be implemented as a set of RF beacons and / or as a set of active and / or passive RF identification (RFID) tags. The RF signals can be received by motion capture device 220 and processed to determine the range and direction, and the intersection points thereof can indicate the pose of marker 212. In the case of CV technology, an optical sensor operating in the visible wavelength range of humans can generate an image showing a two-dimensional visual marker, and a two-dimensional visual marker encoding its dimensions can be used. The pose of the two-dimensional visual marker can be determined by decoding the dimensions and applying geometric reconstruction to the image. Further, in certain situations, the CV technology may not rely on two-dimensional visual markers. Instead, the CV technology can detect a set of feature parts of camera device 210 and track the pose of this set. In this case, marker 212 is a set of feature parts rather than a rigid body removably attached to camera device 210. Such techniques can be used individually or in various combinations to reduce errors or increase accuracy.
[0034] Regardless of the underlying motion capture technique, the pose of marker 212 is tracked over time and represents the motion of marker 212. Motion can be tracked by defining a tracking point 213 on marker 212 (e.g., the center of a rigid body, the center of one of the light-emitting points of marker 212, etc.) and determining the position and rotation of tracking point 213 over time in the coordinate system of the motion capture system. The origin of the coordinate system can be a point in the real-world space (e.g., a point within the volume shown in FIG. 1). The motion capture data 222 can indicate the position and rotation in the coordinate system and can be generated based on a specific rate (e.g., 24 frames per second (FPS), 60 FPS, 144 FPS, etc.).
[0035] As shown in FIG. 2, the camera device 210 may include a focus 216. As further described in FIG. 8, the focus 216 may be an aperture point at a distance from the image plane of the camera device 210, and this distance is equal to the focal length of the camera device 210. Generally, the camera model of the camera device 210 includes external parameters that may indicate the focus 216. Further, in some embodiments, the focus may change when the operating parameters of the camera (e.g., focal length, lens change, etc.) change. The disclosed techniques are robust with respect to such changes.
[0036] There is an offset 240 between the tracking point 213 on the marker 212 and the focus 216 of the camera device 210. This offset 240 may include either a position offset (e.g., indicating the relative distance between the tracking point 213 and the focus 216) and / or a rotation offset (e.g., indicating the relative direction between the tracking point 213 and the focus 216), or a combination thereof.
[0037] The video stream synchronization process 232 may depend on the motion of the camera device 210 to synchronize content rendering with the video stream captured by the camera device 210, and the synchronization may follow an acceptable threshold. There are various options for determining the motion. In one option, the pose of the camera device 210 is set to be the same as the pose of the marker 212. In other words, the offset 240 is ignored, and as a result, the focus 216 is assumed to be the same as the tracking point 213. This option may introduce errors that affect the quality of the output of the video stream synchronization process 232, as further described in FIG. 3. In another option, the offset 240 is determined and used to transform the pose of the marker 212 to the pose of the camera device 210 (e.g., translate each pose of the tracking point 213 to the corresponding pose of the focus 216 based on the position offset, rotate based on the rotation offset, etc.). This option can reduce or even eliminate errors, thereby improving the quality of the output of the video stream synchronization process 232.
[0038] Figure 3 shows an example of tracking an error according to an embodiment of the present disclosure. The left side shows a left projection view 300 along the YZ plane. The right side shows an upper projection view 350 along the XZ plane. The camera device 310 has been moved from the first pose 302 to the second pose 304. The marker 312 is attached to the camera device 310 to determine the first pose 302 and the second pose 304. The camera device 310 also includes a focus 314, the pose of which can be used for content rendering synchronization. The camera device 310 and the marker 312 are examples of the camera device 210 and the marker 212 in FIG. 2, respectively.
[0039] As shown in the left projection view 300, when the camera device 310 is moved from the first pose 302 to the second pose 304, the marker 312 moves downward, while the focus 314 moves upward. Also, as shown in the upper projection view 350, when the camera device 310 is moved from the first pose 302 to the second pose 304, the marker 312 moves to the right, while the focus 314 moves to the left.
[0040] Therefore, when tracking the motion of the camera device 310 using the pose of the marker 312 without considering the offset between the marker 312 (e.g., the tracking point thereon) and the focus 314, there is a pose error. For example, referring back to two poses (the first pose 302 and the second pose 304), the camera device 310 is assumed to have moved down and to the right. However, its focus 314 has actually moved up and to the left. Thus, the video stream synchronization process may misjudge the bottom-right motion as the top-right motion instead, even when it should be used to synchronize the rendering of the content with the video stream of the content captured by the camera device 310. Such a difference can cause an unrealistic rendering of the background content on the screen and may disrupt the viewer's experience. By determining the poses of the markers 312 and converting them to the poses of the focus 314 based on the offset between the markers 312 and the focus 314, the pose error can be reduced and even eliminated.
[0041] There are problems regarding determining the offset between the focus of a camera device and a marker attached to the camera device. Usually, the focus is not marked on the camera device (since it is generally inside the camera device). Further, the positioning of the marker is not predefined and can be set flexibly. For example, markers of different types or configurations (having different light-emitting points, different sizes, etc.) can be used. Once a marker is selected, its attachment to the camera device can involve a quick-mount mechanism. Alternatively, when using CV technology to detect a set of features on the surface of the camera device for use as a marker, the set of features may not be predefined. It may be possible to use tools (such as a ruler, a protractor, etc.) to measure the offset between the focus and the marker. However, these measurements may not be accurate enough (especially considering that the focus is not marked on the camera device or otherwise identified). Instead, as further described herein below, the offset can be determined based on the camera model of the camera device. Using a camera calibration procedure, some or all of the external parameters, internal parameters, and distortion model of the camera model can be determined.
[0042] FIG. 4 shows an example of a camera calibration system 400 configured to determine the focal position of a camera device 410 according to an embodiment of the present disclosure. The camera calibration system 400 includes a motion capture system including a motion capture device 420, a computer system 430, and an object 440. The camera calibration system 400 executes a camera calibration procedure by processing at least in part motion capture data 442 generated by the motion capture system and image data 414 generated by the camera device 410. A marker (not shown) may be attached to the object 440 to track its motion by the motion capture system. The motion capture data 442 indicates the motion of the object 440 over time, while the image data 414 includes an image showing the pose of the object 440 (e.g., representing a video stream of the motion of the object 440). The processing may be performed by a camera calibration process 432 executed on the computer system 430 that receives the motion capture data 442 and the image data 414 and then calibrates the camera model of the camera device 410. The camera device 410, the motion capture device 420, and the computer system 430 are each examples of the camera device 210, the motion capture device 220, and the computer system 230 of FIG. 2.
[0043] A marker 412 may be attached to the camera device 410 to track the movement of the camera device 410 before, during, or after the start of the camera calibration procedure. The marker 412 is an example of the marker 212 of FIG. 2.
[0044] In one example, during the camera calibration procedure, the camera device 410 remains stationary (e.g., not moved). In this case, the marker 412 does not need to be attached until after the completion of the camera calibration procedure, and if attached, it may not be used. At the completion of the camera calibration procedure, the object 440 is deleted and is no longer used. The camera calibration process 432 can output the display of the focus 416 of the camera device 410 (e.g., the position and / or rotation of the focus 416 in a coordinate system, such as the coordinate system of the camera device 410 if the marker 412 is not used, or the coordinate system of the motion capture system if the marker 412) based on the calibrated camera model.
[0045] Thereafter, the camera device 410 can be moved, resulting in a change to the pose of its focus 416. The marker 412 is attached to the camera device 410 (if not already attached), and an offset with respect to the focus 416 can be estimated, such that the motion of the focus 416 (e.g., a sequence of poses) can be determined by tracking the motion of the marker 412. In particular, the motion capture system generates motion capture data 422 indicative of the pose of the marker 412 over time. The camera calibration process 432 receives the motion capture data 422 and determines the pose of the camera device 410 using the determined pose of the marker 412 (e.g., from each video frame captured in the motion capture data 422), and can determine the offset between the marker 412 (e.g., the tracking point thereon) and the focus 416. The offset can be output to a video stream synchronization process 434 running on a computer system. Based on the motion capture data 422 and the offset, the video stream synchronization process 434 determines the motion (e.g., pose) of the focus 416 by translating and / or rotating the motion (e.g., pose) of the marker 412, if applicable. Next, the video stream synchronization process 434 can synchronize content rendering with the video stream captured by the camera device 410 based on the motion of the focus 416.
[0046] The offset may only need to be determined once (e.g., based on the initial pose of marker 412, which may correspond to the pose of camera device 410 during the camera calibration procedure). However, it may also be possible to determine the offset each time the pose of marker 412 changes, or at some predefined rate. The predefined rate may be related to the FPS rate used by the motion capture system and / or the camera device. For example, the predefined rate may be once per second corresponding to 144 motion capture frames when a 144 FPS motion capture rate is used. Alternatively, the predefined rate may be, for example, once every 40 milliseconds corresponding to approximately 24 video frames when a 24 FPS video frame rate is used.
[0047] Referring back to system 100 of FIG. 1 and camera calibration system 400 of FIG. 4, the same motion capture system used in virtual production may also be used to calibrate a camera device (e.g., camera device 120 or 410). Before starting virtual production or shooting a scene, object 440 may be placed within the volume and the calibration procedure may be executed. Thereafter, object 440 may be removed and virtual production may be started and / or a scene may be shot by adjusting content 112 according to the motion of focus 416 using the camera device.
[0048] In FIG. 4, during the camera calibration procedure, while the camera device 410 remains stationary, the object 440 is being moved. However, embodiments of the present disclosure are not so limited. Instead, the camera device 410 may be additionally or alternatively movable. For example, the object 440 may remain stationary while the camera device is moved. In this case, instead of the motion capture data 442 of the object 440, the motion capture data 422 of the camera device 410 is used by the camera calibration process 432 to calibrate the camera model. In another example, the object 440 and the camera device 410 are moved during the camera calibration procedure, and the camera calibration process 432 uses both the motion capture data 422 and the motion capture data 442 in the calibration.
[0049] As described above in this specification, the offset can be determined only once, or can be repeatedly determined based on changes in the pose. The change with respect to the pose can be an actually observed change or a predicted change (e.g., a change predicted to occur with a likelihood exceeding a pre-defined likelihood threshold). By predicting the change, the offset can be pre-computed before the actually observed pose, thereby reducing the latency associated when the offset is determined and available for use by the video stream synchronization process 434. Further, in either case (e.g., whether the offset is calculated only once or repeatedly calculated), the video stream synchronization process 434 can calculate the transformation to use during the rendering of the content, and the transformation can be calculated based on the offset. To reduce the latency associated with the rendering of the content, the video stream synchronization process 434 can be configured for prediction processing. For example, the computer system 430 can store (e.g., in its memory) the motion history of the camera device 410. The video stream synchronization process 434 can perform extrapolation on the history data to predict the next set of poses (e.g., future motion) of the camera device 410. This prediction can also be based on scene data (e.g., choreography data), which can also be stored in the memory, indicating the sequence of events in the scene. Given the prediction, the video stream synchronization process 434 can predict the transformation required (e.g., based on the offset) to render the content. By doing so, the latency between the actual pose and the adjustment of the content rendering can be reduced. Additionally or alternatively to extrapolating based on the motion history, the computer system 140 can store (e.g., in the memory) the history of the transformation, and the video stream synchronization process 434 can predict the next set of transformations before the actual next set of poses is observed, and this prediction can involve extrapolation of the history of the transformation.
[0050] FIG. 5 shows an example of an object 510 that can be used in a camera calibration procedure according to an embodiment of the present disclosure. The object 510 is an example of the object 440 in FIG. 4. Generally, the object 510 includes a surface on which a pattern 520 of features appears, and this surface having the pattern 520 can be imaged by a camera device. The pattern 520 can be indexed (in FIG. 3, a row index 530 using numbers and a column index 540 using characters are shown, but other types of indexing are possible and may depend on the pattern 520). Markers are attached to the surface so that the pose of the object 510 can be tracked over time. In the example of FIG. 5, the markers include three separate markers (e.g., IR light emitting points, recognizable markers, etc.) 550A, 550B, and 550C arranged at different corners of the pattern 520 (however, different numbers of separate markers and / or different dispersions of such markers are also possible). A local coordinate system 560 (e.g., a Cartesian coordinate system, etc.) can be defined for the object 510, and the origin of this coordinate system can be set to a point on the surface (e.g., the origin is the root rigid body point). In the example of FIG. 5, the origin is set at the upper left corner of the pattern 520 and corresponds to the center of the individual marker 550A (however, another point such as the center of the pattern 520 can be used as the origin).
[0051] Object 510 may have different shapes and / or dimensions. Generally, object 510 is large enough to have a pattern 520 that can be accurately detected by the implementation of an image processing algorithm, and yet small enough so that object 510 can be easily moved. The surface of object 510 can be a (e.g., flat) plane or a (e.g., curved) non-plane. Pattern 520 can be a pattern such as a chessboard, but other patterns are also possible. In a chessboard pattern, the features can be a set of rectangular corners having a specific color (e.g., white and black), and can be SIFT features, SURF features, ORB features, etc. of a computer vision algorithm. Other configurations (e.g., shape, color, etc.) of the features are also possible. Generally, the features need to be detectable based on the implementation of an image processing algorithm (e.g., OpenCV algorithm). For accurate detection and processing, the size of the pattern (e.g., its cross-section) needs to be substantially the same as or equivalent to the size of a marker (e.g., the cross-section of an individual light-emitting point) attached to the calibrated camera device. Substantially the same or equivalent means that the difference between the two sizes is within a pre-defined size margin (e.g., ±10 percent) of each other. For example, while the cross-section of the marker is in the range of 200 - 300 millimeters (mm), the size of the pattern is 215.9 mm × 279.4 mm.
[0052] A virtual object (e.g., a three-dimensional object) can model object 510 and can include a virtual representation of pattern 520 and the features. For example, pattern 520 can be represented as a mesh, and each feature can be represented as a mesh point having a position defined as follows in a local coordinate system 560. That is, X ij = I × cell_size, Y ij = J × cell_size, and in the case of a flat surface, Z ijis equal to 0, where "I" corresponds to the column index 540 and "J" corresponds to the row index 530. The mesh (e.g., the virtual representation of the pattern 520) can be used to project the object 510 onto the image plane based on the detected pose of the object 510. Mesh points (e.g., the virtual representation of features) can be used to determine the projection error of the projection.
[0053] FIG. 6 shows an example of generating pose data and image data of an object 610 during a camera calibration procedure according to an embodiment of the present disclosure. As described in connection with FIG. 4, in addition to the object 610, the camera calibration procedure involves a motion capture system including a camera device 620 and a motion capture device 630. In the example of FIG. 6, during the camera calibration procedure, the object 610 is movable while the camera device 620 remains stationary. The motion capture system is used to generate motion capture data of the object 610 indicating the motion of the object 610 over time. The camera device 620 generates image data representing a video stream of the motion. The motion capture frame rate and the video frame rate can be different. For example, a 144 FPS motion capture frame rate and a 24 FPS video capture frame rate can be used. The pose indicated by the motion capture data can be associated with an image (e.g., a video frame) based on timing, as further shown in FIG. 7.
[0054] As shown in FIG. 6, the object 610 is placed in the first pose 602 for a specific amount of time (e.g., a half - second time window, or some other length of time window). While in the first pose 602, the first motion capture data 632 is generated by the motion capture system at the motion capture frame rate, indicating the first pose 602. Also, while in the first pose 602, the first image data 622 is generated by the camera device 620 at the view frame rate, indicating the object 610 in the first pose 602.
[0055] Next, object 610 is moved to the second pose 604. It may take a certain amount of time (e.g., a few milliseconds such as 50 milliseconds) for the object to be placed in the second pose 604. During this time, the motion capture data and the image data can still be generated. However, as described in connection with FIG. 7, these motion capture data and image data can be identified and discarded.
[0056] Object 610 remains in the second pose 604 for a certain amount of time (e.g., a further half - second time window, or a time window of some other length) before moving to the next pose. While in the second pose 604, the second motion capture data 634 is generated by the motion capture system at the motion capture frame rate and represents the second pose 604. Also, while in the second pose 604, the second image data 624 is generated by the camera device 620 at the video frame rate and represents the object 610 in the second pose 604.
[0057] The above process of moving object 610 to different poses every time window (e.g., every half - second) can be repeated to generate motion capture data representing the poses of object 610 over time, and image data representing a video stream of these poses. In one example, it may be sufficient to determine a minimum number of poses (e.g., about 4 poses) of object 610 to calibrate the camera device 620. However, a larger number of poses (e.g., in the range of 30 - 60 or more) can improve the accuracy of camera calibration.
[0058] The movement of the object 610 can be performed by an operator. The operator can be a human. Alternatively, the operator can be a robotic system. For example, the robotic system can include a movable body (e.g., a body attached to wheels) having a controller (e.g., a set of processors) and a memory, a robotic arm extending from the movable body and controllable by the controller, and an end effector having one end permanently attached to the robotic arm and the other end removably attached to the object 610 and also controllable by the controller. The memory can store instructions executed by the processor to control the robotic arm and the end effector. The memory can further store a motion path for controlling the robotic arm and the end effector, and this motion path indicates a sequence of poses where the object 610 should be placed.
[0059] As described above, the pose indicated by a part of the motion capture data can be associated with the image indicated by a part of the image data. The association can be based on timing (e.g., the timing of the pose matches the timing of the image). However, other types of associations can also be used. For example, a visual association is also possible. In particular, referring back to the robotic system, the motion path can include an identifier for each pose. The robotic system can also include a display presenting the identifier for each pose. When imaging the pose, the camera device 620 generates an image, and each image not only shows the object 610 in a specific pose but also shows the identifier of the pose presented on the display. Thus, each image is associated with the corresponding pose, and the pose is identified in the motion path.
[0060] FIG. 7 shows an example of pose data 710 and image data 720 associated with an object according to an embodiment of the present disclosure. The pose data 710 may include motion capture data generated by a motion capture system (as described in connection with FIGS. 4 and 6, for example). Generally, the pose data 710 indicates the pose of the object over time, and the sequence of poses represents the motion of the object. The pose can be defined as a position and rotation in a coordinate system (e.g., in the coordinate system of the motion capture system). Thus, a part of the pose data 710 corresponding to the pose can indicate the position and rotation by including x, y, z coordinates and x, y, z rotations (however, a non-orthogonal coordinate system can be used). The image data 720 may include images generated by a camera device, and the sequence of images indicates the motion of the object over time.
[0061] For clarity, a plot of the x coordinate having values available from the pose data 710 is shown in FIG. 7. The horizontal axis of the plot corresponds to time, and the vertical axis corresponds to the value. However, the y and z coordinates, as well as the x, y, and z rotations, can be plotted similarly. The pose data 710 changes over time. The magnitude or slope of the change can indicate the pose or the transition to the next pose. For example, two x coordinates (or two sets of average x coordinates) can be compared to determine the change in the x coordinate (e.g., the change in the slope of the straight line connecting the x coordinates). If the change is less than a threshold value, the x coordinate has not substantially changed, indicating that the object is still in the same pose. Otherwise, the x coordinate has substantially changed, indicating that the object has moved to the next pose.
[0062] Referring to the plot of FIG. 7, during a first time window "TW0" (e.g., a 500 ms time window) defined by a start time "t0" and an end time "t1", the corresponding x - coordinate 712 indicates that the object did not change significantly between times "t0" and "t1" and was in a first pose. During a transition time window (e.g., a time window of about 50 milliseconds) starting at time "t1" and ending at time "t2", a substantial change occurs, indicating a transition to the next second pose. During a second time window "TW1" (e.g., a 500 ms time window) defined by a start time "t2" and an end time "t3", the corresponding x - coordinate 716 indicates that the object did not change significantly between times "t2" and "t3" and was in the second pose. During a transition time window (e.g., a time window of about 50 milliseconds) starting at time "t3" and ending at time "t4", a substantial change occurs, indicating a transition to the next third pose.
[0063] The first pose is associated with the first time window "TW0". To determine the x - coordinate of the first pose, some or all of the x - coordinates 712 belonging to the first time window "TW0" can be used. In particular, a statistical metric (e.g., an averaging function) can be applied to the relevant x - coordinates 712 to calculate the x - coordinate of the first pose. This type of calculation can be repeated across different time windows as well as across the z - axis and y - axis to determine different poses.
[0064] The image data 720 includes images corresponding to the poses (e.g., a first image 722 and a second image 724). For example, the first image 722 corresponds to the first pose, while the second image 724 corresponds to the second pose. Of course, the number of images per pose can depend on the video frame rate and the length of time the object remained in that pose (e.g., the length of time of each time window "TW0", "TW1", etc.). If more than two images are available for a pose, a single image (or a subset of these images) can be randomly selected for further processing.
[0065] FIG. 8 shows an example of a camera model 810 determined based on a camera calibration procedure according to an embodiment of the present disclosure. The camera model 810 includes a focus 812 and an image plane 814 separated by a distance representing a focal length 816. The focus 812 can be the center of the aperture through which light rays from the object 850 pass and are projected as an inverted image onto the image plane 814.
[0066] The camera model 810 can characterize the corresponding camera device by including an external camera parameter 820, an internal camera parameter 822, and a distortion model 824. The external camera parameter 820 represents the location of the camera device in the real-world space. This location can correspond to the pose of the focus 812. The internal camera parameter 822 represents the optical center and focal length of the camera. For example, the internal camera parameter 822 includes the focal length 816 and the optical center, which is also known as the principal point, and can be represented as a camera matrix
Equation
[0067] The camera matrix does not take into account the distortion of the lens. To accurately represent the camera device, the distortion model 824 includes radial and tangential lens distortion. Radial distortion occurs when light rays bend more near the edge of the lens than at the optical center of the lens. The smaller the lens, the greater the distortion. The radial distortion coefficients model this type of distortion. The distorted point is represented as (x distorted ,y distorted ), where x distorted =x(1 + k1*r 2 + k2*r 4 + k3*r 6 ) and y distorted =y(1 + k1*r 2 + k2*r 4 + k3*r 6 ), where x and y represent the pixel positions without distortion (x and y are normalized image coordinates, which are calculated by translating the pixel coordinates to the optical center and dividing by the focal length in pixel units), k1, k2, and k3 represent the radial distortion coefficients of the lens, and r 2 =x 2 +y 2 .
[0068] Tangential distortion occurs when the lens and the image plane are not parallel. The tangential distortion coefficients model this type of distortion. The distorted point is represented as (x distorted ,y distorted ), where x distorted =x + [2*p1*x*y + p2*(r 2 + 2*x 2 )] and y distorted =y + [p1*(r 2 + 2*y 2 ) + 2*p2*x*y], where x and y represent the pixel positions without distortion, p1 and p2 are the tangential distortion coefficients of the lens, and r 2 is also equal to x 2 +y 2 .
[0069] As will be further described in the following figures, each of the external camera parameters 820, the internal camera parameters 822, and the distortion model 824 can be determined based on a camera calibration procedure. Alternatively, the internal camera parameters 822 and / or the distortion model 824 can be pre-defined (e.g., available from the camera device manufacturer or a third party and defined data for a device model or device serial number, etc.). If the internal camera parameters 822 or the distortion model 824 are available, the camera calibration procedure may determine only the external camera parameters 820 without updating the pre-defined internal camera parameters 822 and / or the distortion model 824 (if applicable). In this case, alternatively, the camera calibration procedure initializes the camera model 810 to use the pre-defined internal camera parameters 822 and / or the distortion model 824 (if applicable), and then, in addition to determining the external camera parameters 820, may also update the pre-defined internal camera parameters 822 and / or the distortion model 824.
[0070] Figures 9 - 11 illustrate a flow related to the determination and use of one or more poses of the focus of a camera device. The operations of the flow can be executed by a computer system such as computer system 140 or 430. Some or all of the instructions for executing the operations can be implemented as a hardware circuit and / or stored as computer-readable instructions on a non-transitory computer-readable medium of the computer system. When implemented, the instructions represent components including circuits or code executable by a processor of the computer system. By using such instructions, the computer system is configured to perform the specific operations described herein. Each circuit or code, in combination with the associated processor, represents a means for performing each operation. Although the operations are shown in a particular order, it should be understood that no particular order is required and one or more operations may be omitted, skipped, executed in parallel, and / or reordered.
[0071] FIG. 9 shows an example of a flow for rendering content based on the determination of the focus of a camera device according to an embodiment of the present disclosure. The flow may start at operation 902 where a computer system determines a first pose of a marker attached to the camera device. In one example, the computer system receives pose data over time, such as motion capture data generated by a motion capture system, which indicates the motion of the marker and is defined with respect to tracking points (e.g., its rigid body points) on the marker. A portion of the pose data may have a timestamp corresponding to the time when the camera device was in the first pose in the real-world space. That portion of the pose data may include x, y, and z coordinates and x, y, and z rotations indicating the position and rotation of the marker in a coordinate system (e.g., the coordinate system of the motion capture system).
[0072] At operation 904, the computer system determines an offset between the marker and the focus of the camera device based on the first pose and at least one calibration parameter of the camera device. In one example, the camera device is associated with a camera model that includes at least one calibration parameter, such as one or more external camera parameters and / or one or more internal camera parameters. The external camera parameters can specify the position of the focus in the coordinate system of the camera device. The internal camera parameters can indicate an image plane parallel to the plane containing the focus and can also indicate the focal length between the image plane and the focus. Based on the external camera parameters and the internal camera parameters, a position offset and a rotation offset are calculated from the position and rotation of the marker (e.g., its tracking points). The position offset and the rotation offset form an offset between the marker and the focus. The camera model may be pre-defined or may be determined partially or wholly based on a camera calibration procedure.
[0073] In operation 906, the computer system sets a second pose of the virtual camera based on the offset. The virtual camera can be a virtual representation of a camera device in a virtual space, while the virtual space is a virtual representation of the real-world space. In one example, the pose of the focus in the coordinate system of the real-world space (the coordinate system of the motion capture system in which the first pose is determined) is determined by translating the position of the marker by the position offset and rotating the rotation of the marker by the rotation offset. The pose of the focus can be represented as the position and rotation of the marker in the coordinate system. Given the mapping between the coordinate system of the real-world space and the coordinate system of the virtual space, the position and rotation of the focus are mapped to a virtual position and a virtual rotation in the virtual space. The virtual position and the virtual rotation define the second pose of the virtual camera.
[0074] In operation 908, the computer system renders content based on a second pose. For example, the computer system synthesizes an image and / or an image frame (2D or 3D) to generate the content. The synthesis can apply a transformation (e.g., skew, angling, distortion, etc.) based on the second pose such that when the content is presented (e.g., on a display assembly), the content itself and / or its presentation is adjusted taking into account the pose of the focus in the real-world space. In a further example, the camera model may include a distortion model. The rendering can distort the content based on the second pose by using radial distortion and / or tangential distortion (or vice versa) to give a special effect to the presentation of the content based on the pose of the focus in the real-world space. For example, the display assembly is modeled as a virtual display in a virtual space. The computer system determines a virtual display area (e.g., the entire virtual display or a portion thereof) based on the second pose (e.g., a virtual display area within the field of view of a virtual camera given the second pose). The computer system also determines the portion of the content that will be associated with the virtual display area (e.g., will be displayed in the display area of the display assembly corresponding to the virtual display area). The computer system then distorts the portion of the content based on the distortion model (e.g., by using radial distortion, tangential distortion, the reverse, etc.).
[0075] In operation 910, the computer system causes the content to be presented. For example, the computer system outputs the content to a display assembly, and the display assembly presents the content. When the content is presented, the camera device can generate a video stream showing the content.
[0076] The flow of FIG. 9 is described in relation to a single pose of the camera device. Some or all of the operations of the flow can be repeated for different poses of the camera device, and the sequence of poses corresponds to the motion of the camera device. For example, the offset can be determined once and reused, or can be calculated when the pose of the marker changes. In either case, the second pose of the focus can be updated based on the latest pose of the marker. The updated pose can be used to update the rendering of the content.
[0077] FIG. 10 shows an example of a flow for determining the offset between the focus of the camera device and a marker attached to the camera device according to an embodiment of the present disclosure. The offset can be determined simultaneously with the execution of the camera calibration procedure. The operation of the flow of FIG. 10 can be implemented as a sub-operation of the flow of FIG. 9. In one example, the flow of FIG. 10 can start at operation 1002 where the computer system determines the pose of the object. The sequence of poses represents the motion of the object over time. As an example, the computer system receives pose data indicating the pose, such as motion capture data generated by a motion capture system and tracking a marker attached to the object. Given a change in the pose data, the computer system can determine the pose.
[0078] At operation 1004, the computer system receives an image indicating the pose. For example, the image is generated by the camera device, represents a video stream showing the motion of the object, and is received as image data (e.g., a video frame) from the camera device.
[0079] In operation 1006, the computer system associates an image with a pose. For example, an image showing the pose of an object is associated with the pose. In one example, the association can be time-based. In particular, the timing of each image is determined (e.g., as a timestamp when the image was generated) and matched to the timing of the pose (e.g., within a time window in which the pose was calculated and associated with pose data). The match can correspond to the timestamp being between the start time and end time of the time window. If there is a match, an association between the image and the pose is established. Other types of associations can also be used additionally or alternatively. For example, referring back to a robot system that moves an object and displays an identifier for each pose, an image showing the pose can be processed to detect the identifier of the pose displayed by the robot system.
[0080] In operation 1008, the computer system calculates a camera model based on the image and the pose (and the association between them). For example, the image and the pose are used in a camera calibration procedure to determine, if applicable, external camera parameters, internal camera parameters, and / or a distortion model. In one example, the camera calibration procedure includes a plurality of steps.
[0081] In the first step, a subset of the images is randomly selected (e.g., at least 4, or more typically, in the range of 30 - 60 images, etc.). Visual features on the object (such as formed by pattern 520 in FIG. 5) are detected on each of the selected images as 2D coordinates on the image. Any type of 2D feature (e.g., SIFT, SURF, ORB, etc.) based on any recognizable pattern (e.g., a chessboard pattern) can be used.
[0082] In the second step, a correspondence relationship is set between the three-dimensional feature part (for example, three-dimensional coordinates in millimeters in the coordinate system of the motion capture system, which is called the global coordinate system in this specification) and the two-dimensional feature part (for example, two-dimensional image coordinates in pixels in the coordinate system of the image, which is called the image coordinate system in this specification). For example, when a chessboard pattern is used, the position of any chessboard mesh point (for example, any detectable feature part) can be described as follows in the local chessboard pattern coordinate system. That is, X ij =I×cell_size, Y ij =J×cell_size, in the case of a flat surface, Z ij =0, where "I" and "J" are the indexes of the feature part, and cell_size is the size of the feature part in millimeters.
[0083] Next, the camera calibration procedure can be executed for the extracted feature parts. The motion capture data of the recognizable pattern is used as the initial value of the external camera parameters. Each selected image has six external parameters, that is, three position coordinates and three Rodriguez angles. The internal camera parameters are the same for the selected images. Depending on the user settings, the number of internal parameters can include the focal length (in the X and Y directions) and the principal point (X and Y image coordinates), and as described above in this specification, the distortion model can include a plurality of distortion parameters ("k's" (radial distortion) and "p's" (tangential distortion)).
[0084] In an example of a distortion model with four distortion parameters (such as k8 to k11 in the case of a thin prism model), the following formula can be used. That is, x’’=x’*(1+k1r 2 +k2r 4 +k3r 6 ) / (1+k4r 2 +k5r 4 +k6r 6 )+2p1x’y’+p2*(r 2 +2x’ 2 )、y’’=y’*(1+k1r 2+k2r 4 +k3r 6 ) / (1 + k4r 2 +k5r 4 +k6r 6 ) + p1(r 2 +2y’ 2 ) + 2p2x’y’ and r 2 = x’ 2 +y’ 2 where x’ and y’ are the normalized projection coordinates (normalized image coordinates) of the 3D points of the object. The distortion model may also include the distortion of the thin prism model. In this case, the additional offset in the 3D point projection is X’’ += k8 * r 2 +k9 * r 4 and Y’’ += k10 * r 2 +k11 * r 4 which can be described as. <000038'7>
[0085] [[ID='32]]In another example of a distortion model that involves using angular parameters (also, for example, in the case of a thin prism model, τ x and τ y ), the following equation (referred to as the tilt matrix) can be used.
Equation
[0086] By using the pose associated with each of the selected images, the computer system receives the 3D coordinates of the features in the pose, where these coordinates are represented as P ij_global = R pattern * P ij_local + P pattern . P ij_global and P ij_local are the 3D positions of any feature in the corresponding global coordinate system and local coordinate system, R pattern is the rotation matrix of the pattern, and P pattern is the position of the pattern.
[0087] Note: In the translation, some tags like etc. are just preserved as they are because they seem to be some kind of specific identifiers in the original text format and there's no clear indication of how to translate them meaningfully. Also, the equations and notations are kept as close to the original as possible to maintain the integrity of the technical content.After determining the correspondence between 3D coordinates (in millimeters in the global coordinate system) and 2D coordinates (in pixels in the image coordinate system) for any image, the perspective-n-point (PnP) problem can be solved to estimate the pose of the camera device calibrated from the "n" correspondences from 3D to 2D feature parts. The pose of the camera device (Rk camera , the rotation of the camera device for the k-th image, and Pk camera , the position of the camera device for the k-th frame) is received for any image.
[0088] Then, the computer system can estimate the initial camera projection plane rotation (R camera ) and the initial camera focal position (P camera ) as a weighted average of the received pose data as follows. R camera = Σ n k=0 weight_k * Rk camera and P camera = Σ n k=0 weight_k * Pk camera . The weights can be calculated as the normalized reprojection error in the PnP solution for each image. The reprojection error can be calculated by determining the error (e.g., distance error) between the projection of the object of the pose (or its feature part) on the image plane of the camera device and the two-dimensional coordinates of the object (or its feature part) in the image indicating the pose and the projection.
[0089] Using camera calibration parameters and initial pose estimation, the internal and external parameters can be estimated by applying a fitting model to camera calibration parameters (including any of external camera parameters, internal camera parameters, and / or distortion parameters). Generally, the fitting model is a data fitting model that iteratively estimates the camera calibration parameters such that the reprojection error is reduced or minimized. For example, different types of data fitting models are possible, such as the Levenberg-Marquardt nonlinear least squares algorithm, chi-square test algorithm, curve fitting algorithm, weighted least squares fitting algorithm, polynomial regression algorithm, Gauss-Newton algorithm, shift cut algorithm, gradient algorithm, Nelder-Mead (simplex) search algorithm, etc. During the camera calibration procedure, if the camera device remains stationary, the same pose of the camera device is expected for all images, so the number of variables in the fitting model may depend on the internal parameters and distortion model of the camera and may not depend on the number of selected images. The reprojection error is used in the least squares curve fitting problem. Additionally, or alternatively, the fitting model may include a machine learning model such as a regression model or a convolutional neural network trained using a plurality of known camera models, poses of known objects, and training images showing such poses of the objects.
[0090] Furthermore, the fitting model can be called twice. The first time, a full set of selected images is used to estimate the camera model and is processed by the fitting model. Next, a subset of the selected images is used. The subset can be selected by choosing the images with the minimum reprojection error (for example, if 30 - 60 images are selected, a subset of 10 - 20, or some other number of images with the minimum reprojection error is selected). The subset can be processed again by the fitting model to further refine the parameters of the camera model.
[0091] In operation 1010, the computer system determines the pose of a marker attached to a camera. For example, the marker was attached before the start of the camera calibration procedure. The pose can be determined, for example, from pose data generated by a motion capture system. During the camera calibration procedure, the camera can remain stationary. In such a situation, during the camera calibration procedure, the pose data does not substantially vary. Thus, the position and rotation of the marker can be calculated in the global coordinate system as the average value of the pose data.
[0092] In operation 1012, the computer system determines a rotational offset and a position offset of the focus of the camera device based on the pose of the marker and the camera model. For example, the pose of the marker indicates a first rotation of the marker in the real-world space (e.g., in the global coordinate system, this first rotation is represented as a rotation matrix R marker about the X, Y, and Z axes). The camera model (e.g., one or more of its external camera parameters) can indicate a second rotation of the focus (e.g., this second rotation is represented as a rotation matrix R camera about the X, Y, and Z axes). The rotational offset can be determined based on the first rotation and the second rotation, such as by using the following equation. That is, R offset =R camera *R marker -1 , where R marker -1 is the inverse of the rotation matrix R marker of the marker.
[0093] The pose of the marker also indicates a first position of the marker in the real-world space (e.g., in the global coordinate system, this first position is represented as a position matrix T marker along the X, Y, and Z axes). The camera model (e.g., one or more of its external camera parameters) can indicate a second position of the focus (e.g., this second position is represented as a position matrix T camerarepresented as). The position offset can be determined based on the first rotation, the first position, and the second position, such as by using the formula, T offset =R marker -1 *(T camera -T marker ).
[0094] FIG. 11 shows an example of a flow for associating pose data with image data in a camera calibration procedure according to an embodiment of the present disclosure. The association is time-based. The operations of the flow of FIG. 11 can be implemented as sub-operations of the flow of FIG. 9 and / or the flow of FIG. 10. In one example, the flow of FIG. 11 can start at operation 1102 where the computer system receives motion capture data indicating the pose of an object over time. For example, a motion capture system can generate motion capture data at a particular frame rate (e.g., 144 FPS) and transmit this data to the computer system.
[0095] At operation 1104, the computer system receives image data representing a video stream of the pose of the object. For example, a camera device can generate image data as video frames at a particular frame rate (e.g., 24 FPS) and transmit this data to the computer system.
[0096] At operation 1106, the computer system determines the change over time of the motion capture data. This change can be with respect to the magnitude and / or inclination of either the x, y, z coordinates or the x, y, z rotations.
[0097] At operation 1108, the computer system determines the pose of the object based on the change. The change can be compared to a threshold defined for each of the x, y, z axes. If it is smaller for all three axes, a pose is detected and associated with a time window during which the change between them is smaller than the threshold. Otherwise, a transition between the current pose and the next pose is detected.
[0098] In operation 1110, the computer system associates an image with a pose based on the pose and the timing of the image. For example, each received image has a timestamp. Each pose is associated with a time window. If the timestamp of the image is within the time window of the pose, the image is associated with the pose, and the association indicates that the image depicts the pose.
[0099] FIG. 12 shows exemplary components of a computer system 1200 according to an embodiment of the present disclosure. The computer system 1200 is an example of the computer system 140 or 430. Although the components of the computer system 1200 are shown as belonging to the same computer system 1200, the computer system 1200 may also be distributed (e.g., among multiple user devices).
[0100] Computer system 1200 includes at least a processor 1202, a memory 1204, a storage device 1206, an input / output peripheral (I / O) 1208, a communication peripheral 1210, and an interface bus 1212. The interface bus 1212 is configured to communicate, transmit, and transfer data, control, and commands among the various components of computer system 1200. The memory 1204 and the storage device 1206 include computer-readable storage media such as RAM, ROM, electrically erasable programmable read-only memory (EEPROM), hard drives, CD-ROMs, optical storage devices, magnetic storage devices, electronic non-volatile computer storage such as Flash (registered trademark) memory, and other tangible storage media. Any such computer-readable storage media can be configured to store instructions or program code embodying aspects of the present disclosure. The memory 1204 and the storage device 1206 also include a computer-readable signal medium. The computer-readable signal medium includes a propagated data signal having computer-readable program code embodied therein. Such a propagated signal takes any of a variety of forms including, but not limited to, electromagnetic, optical, or any combination thereof. The computer-readable signal medium includes any computer-readable medium that can communicate, propagate, or transmit a program for use in connection with computer system 1200, which is not a computer-readable storage medium.
[0101] Furthermore, the memory 1204 includes an operating system, programs, and applications. The processor 1202 is configured to execute the stored instructions and includes, for example, an arithmetic logic unit, a microprocessor, a digital signal processor, and other processors. The memory 1204 and / or the processor 1202 may be virtualized and hosted, for example, within another computer system of a cloud network or a data center. The I / O peripheral devices 1208 include a keyboard, a screen (e.g., a touch screen), a microphone, a speaker, other input / output devices, and a user interface such as computing components including a graphics processing unit, a serial port, a parallel port, a universal serial bus, and other input / output peripheral devices. The I / O peripheral devices 1208 are connected to the processor 1202 via any of the ports coupled to the interface bus 1212. The communication peripheral devices 1210 are configured to facilitate communication between the computer system 1200 and other systems via a communication network and include, for example, a network interface controller, a modem, wireless and wired interface cards, an antenna, and other communication peripheral devices.
[0102] It should be apparent to those skilled in the art that many further modifications are possible without departing from the inventive concept described herein. Accordingly, the subject matter of the present invention should not be limited except in the spirit of the appended claims. Further, in interpreting both this specification and the claims, all terms should be construed in the broadest possible manner consistent with the context. In particular, the terms "comprises" and "comprising" should be construed as referring to elements, components, or steps in a non-exclusive manner, indicating that the recited elements, components, or steps may be present with, utilized with, or combined with other elements, components, or steps not explicitly recited. Where the specification or claims refer to at least one of something selected from a group consisting of A, B, C to N, the text should be construed as requiring only one element from that group, rather than A+N, B+N, etc.
Claims
1. Determining, via at least one processor, a first pose of a marker in a real-world space, wherein the marker is attached to a camera device; and Determining, via the at least one processor, an offset between the marker and a control point of the camera device based on the first pose and at least one calibration parameter of the camera device; and Setting, via the at least one processor, in a computer-readable memory, a second pose of a virtual camera in a virtual space based on the offset, wherein the virtual camera represents the camera device in the virtual space, and the virtual space represents the real-world space; and Rendering, via the at least one processor, content on a display based on the second pose of the virtual camera, wherein the content is captured by the camera device when the content is presented in the real-world space; A computer-implemented method comprising the above.
2. Determining the offset comprises Determining, based on the first pose, a first rotation of the marker in the real-world space; and Determining, based on the at least one calibration parameter, a second rotation of the control point; and Determining a rotation offset based on the first rotation and the second rotation; The computer-implemented method according to claim 1, comprising the above.
3. Determining the offset further comprises Determining, based on the first pose, a first position of the marker in the real-world space; and Determining, based on the at least one calibration parameter, a second position of the control point; and Determining a position offset based on the first rotation, the first position, and the second position; The computer-implemented method according to claim 2, further comprising the above.
4. The computer-implemented method according to claim 1, wherein the second pose is set as a translation and a rotation of the first pose based on the offset.
5. The at least one calibration parameter is determined from a camera model of the camera device, the camera model includes a distortion model, and rendering the content Based on the second pose, determining a virtual display area in the virtual space; determining a portion of the content associated with the virtual display area; distorting the portion of the content based on the distortion model; The computer-implemented method according to claim 1, comprising:
6. wherein the at least one calibration parameter is determined from a camera model of the camera device, the camera model is determined based on a calibration procedure, and the calibration procedure determines a plurality of poses of an object in the real-world space; receives a plurality of images generated by the camera device, the plurality of images representing a video stream showing the object in the plurality of poses; determines the camera model of the camera device based on the plurality of poses and the plurality of images, the camera model including external camera parameters, internal camera parameters, and a distortion model; The computer-implemented method according to claim 1, comprising:
7. wherein the at least one calibration parameter is determined based on a camera calibration procedure, and the camera calibration procedure determines a third pose of an object in the real-world space; receives an image generated by the camera device, the image showing the object while the object is in the third pose; determines the at least one calibration parameter based on the third pose and the image; The computer-implemented method according to claim 1, comprising:
8. The computer-implemented method according to claim 7, wherein the camera calibration procedure is performed while one of the camera device or the object is stationary and the other of the camera device or the object is movable.
9. The camera calibration procedure determines a plurality of poses of the object in the real-world space while the camera device remains in the same pose; receives a plurality of images generated by the camera device, the plurality of images representing a video stream showing the object in the plurality of poses; determines that the third pose is shown in the image; determining a projection error between the image and a projection of the object in the third pose onto the image plane of the camera model of the camera device, wherein the camera model includes the at least one calibration parameter; selecting a subset of the images based on the projection error, wherein the subset of the images is associated with a subset of the poses of the plurality of poses; updating the camera model based on the subset of the images and the subset of the poses; The computer-implemented method according to claim 7, further comprising.
10. The camera calibration procedure is determining a projection error between the image and a projection of the object in the third pose onto the image plane of the camera model of the camera device, wherein the camera model includes the at least one calibration parameter; updating the camera model based on a fitting model that reduces the projection error, wherein an output of the fitting model indicates one or more parameters of the camera model; The computer-implemented method according to claim 7, further comprising.
11. The camera calibration procedure is determining a first estimated pose of the camera device based on the third pose of the object and the image; determining a projection error between the image and a projection of the object in the third pose onto the image plane of the camera model of the camera device, wherein the camera model includes the at least one calibration parameter; associating an initial weight with the first estimated pose based on the projection error; determining a final estimated pose of the camera device based on the first estimated pose and the initial weight; updating the camera model based on a fitting model that uses the final estimated pose and the projection error; The computer-implemented method according to claim 7, further comprising.
12. The at least one calibration parameter is determined from a camera model of the camera device, the camera model is determined based on a camera calibration procedure, and the camera calibration procedure is Receiving motion capture data representing a plurality of poses of an object in the real-world space over time while the camera device remains in the same pose; Determining the timing of a third pose of the object based on a change in the motion capture data; Determining the at least one calibration parameter based on the third pose; The computer-implemented method according to claim 1, comprising:
13. The camera calibration procedure is: Receiving image data representing a plurality of images generated by the camera device over time while the camera device remains in the same pose, the plurality of images representing a video stream showing the object in the plurality of poses; Associating the third pose with an image among the plurality of images based on the timing of the third pose and the timing of the image, wherein the at least one calibration parameter is further determined at least partially based on the timing; The computer-implemented method according to claim 12, further comprising:
14. A system, comprising: One or more processors; One or more memories storing instructions, which when executed by the one or more processors: Determine a first pose of a marker in a real-world space, the marker being attached to a camera device; Determine an offset between the marker and a control point of the camera device based on the first pose and at least one calibration parameter of the camera device; In the one or more memories, set a second pose of a virtual camera in a virtual space based on the offset, the virtual camera representing the camera device in the virtual space, the virtual space representing the real-world space; Render content on a display based on the second pose of the virtual camera, the content being captured by the camera device when the content is presented in the real-world space; One or more memories configured to configure the system as such; A system comprising:
15. The at least one calibration parameter is determined from a camera model of the camera device, the camera model is determined based on a calibration procedure, and the calibration procedure is: determining a plurality of poses of an object in the real-world space; receiving a plurality of images generated by the camera device, the plurality of images representing a video stream of the object in the plurality of poses; determining the camera model based on the plurality of poses and the plurality of images; The system according to claim 14, comprising: **Claim 16** The system according to claim 15, wherein the object is arranged in the plurality of poses by a robot system. **Claim 17** The system according to claim 15, wherein the marker includes a marker unit, the object includes a pattern of feature portions, and the feature portions of the pattern of feature portions and the marker unit of the marker have substantially the same size. **Claim 18** One or more non-transitory computer-readable storage media storing instructions, which when executed on a system, cause the system to determine a first pose of a marker in a real-world space, the marker being attached to a camera device; determine an offset between the marker and a control point of the camera device based on the first pose and at least one calibration parameter of the camera device; set a second pose of a virtual camera in a virtual space based on the offset in a computer-readable memory, the virtual camera representing the camera device in the virtual space, the virtual space representing the real-world space; render content on a display based on the second pose of the virtual camera, the content being captured by the camera device when the content is presented in the real-world space; One or more non-transitory computer-readable storage media that cause an operation including the above to be executed. **Claim 19** Determining the offset includes determining a first rotation and a first position of the marker in the real-world space based on the first pose; determining a second rotation and a second position of the control point based on the at least one calibration parameter; Determine a rotational offset based on the first rotation and the second rotation, and determine a position offset based on the first rotation, the first position, and the second position; One or more non-transitory computer-readable storage media according to claim 18, comprising: **Claim 20** One or more non-transitory computer-readable storage media according to claim 18, wherein the second pose is set as a translation and a rotation of the first pose based on the offset.