Reference Frame for Motion Capture
The use of HMDs with retroreflectors and synchronized audio cues addresses the challenge of providing spatial references to remote actors, facilitating coordinated performance adjustments in geographically separated environments.
Patent Information
- Application Number
- JP2023531029
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-25
- Filing Date
- 2021-11-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-11-26
AI Technical Summary
Challenges arise in providing a physical reference to remote actors during motion capture due to geographical separation, which complicates the adjustment of performances in collaborative work environments such as movie production and computer simulations.
A method involving the use of a head-mounted display (HMD) with retroreflectors and light emitters to provide a reference frame to actors, synchronized motion capture videos, and synchronized audio cues to align performances across geographically distant locations.
Enables coordinated remote performance guidance by providing accurate spatial references and synchronized video feeds, allowing for seamless integration of multiple actors' motions into a single scene, enhancing the efficiency of collaborative projects like movie production and computer simulations.
Smart Images

Figure 0007711194000001 
Figure 0007711194000002 
Figure 0007711194000003
Abstract
Description
Technical Field
[0001] This application generally relates to technically innovative and non-conventional solutions that are necessarily rooted in computer technology and bring about specific technical improvements. In particular, this application relates to techniques for enabling coordinated remote performance guidance at multiple locations.
Background Art
[0002] Due to concerns regarding health and cost, people are increasingly performing collaborative work from remote locations. As understood herein, the production of movies and computer simulations (e.g., computer games) through collaborative work using remote actors can give rise to unique adjustment problems. This is because a director may need to direct a plurality of actors who may be in their own studio or a soundproof studio when producing a movie or for computer simulation-related activities such as motion capture (MoCap). For example, in the way the performance is adjusted, there are challenges in providing a physical reference to remote actors on individual stages. This principle provides techniques for addressing some of these adjustment challenges.
Summary of the Invention
[0003] Accordingly, this principle provides a method that includes providing a reference frame to at least a first actor at a first location while filming the first actor for motion capture (mocap) by at least partially presenting at least one reference image on a head-mounted display (HMD) worn by the first actor. The light reflected from the retroreflector can be from a light emitter. Additionally or alternatively, the method can include providing a retroreflector on a wall at the first location to reflect light towards the first actor. Additionally or alternatively, the method can include providing a visible marker on a floor at the first location.
[0004] In some examples, the method can include providing a reference frame to at least a second actor at a second location while photographing the second actor for mocap. The first location can be geographically distant from the second location, and the mocap from the first actor and the second actor can be presented on at least one director display that communicates with the first location and the second location during a WebEx.
[0005] A reference frame can be provided to the first actor using at least in part the audio reproduced at the first location. The plurality of light emitters may be provided on the HMD. The Mocap videos of the first actor and the second actor can be synchronized in time.
[0006] In another aspect, the device includes at least one computer storage, and the at least one computer storage includes instructions executable by at least one processor, not a transient signal, for receiving a motion capture (mocap) video of a first actor from a first camera at a first location. The instructions are executable to receive a mocap video of a second actor from a second camera at a second location, synchronize the mocap videos with each other, and merge the mocap videos into a single scene on at least one display at a third location that is geographically distant from the first location and the second location.
[0007] In another aspect, the apparatus includes at least one head-mounted display (HMD) assembly, and the at least one HMD assembly includes at least one processor configured with instructions and at least one display controlled by the processor. The HMD may also include a speaker. At least one projector is configured to project motion capture (mocap) reference light onto at least one surface visible to the wearer of the HMD assembly to provide a spatial reference to the wearer during mocap.
[0008] The details of the present application can be best understood with reference to the accompanying drawings in both its structure and operation, in which like reference numerals refer to like parts.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Modes for Carrying Out the Invention
[0010] Referring now to FIG. 1, the present disclosure generally relates to a computer ecosystem having aspects of a computer network that may include devices of consumer electronics (CE). The systems herein may include server components and client components connected via a network such that data can be exchanged between the client components and the server components. The client components may include one or more computing devices, including portable televisions (e.g., smart TVs, Internet-enabled TVs), portable computers such as laptop computers and tablet computers, and other mobile devices including smart phones and additional examples to be considered below. These client devices may operate in a variety of operating environments. For example, some of the client computers may use, by way of example, an operating system of Microsoft®, or Unix®, or an operating system manufactured by Apple Computer® or Google®. These operating environments may be used to execute one or more browsing programs, such as a browser created by Microsoft® or Google® or Mozilla®, or other browser programs that can access websites hosted by Internet servers discussed below.
[0011] The server and / or gateway may include one or more processors that execute instructions configuring the server to receive and transmit data over a network such as the Internet. Alternatively, the client and server can be connected through a local intranet or a virtual private network. The server or controller may be instantiated by a game console such as Sony PlayStation®, a personal computer, or the like.
[0012] Information can be exchanged between a client and a server through a network. For this purpose and for security, the server and / or the client may include a firewall, a load balancer, temporary storage, and a proxy, as well as other network infrastructure for reliability and security.
[0013] As used herein, an instruction refers to a computer-implemented step for processing information within a system. An instruction can be implemented in software, firmware, or hardware and can include any type of programmed step performed by a component of the system.
[0014] A processor can be a general-purpose single-chip processor or a general-purpose multi-chip processor that can execute logic through various lines such as address lines, data lines, and control lines, as well as registers and shift registers.
[0015] The software modules described herein by flowcharts and user interfaces can include various subroutines, procedures, etc. Without limiting the present disclosure, the logic defined to be executed by a particular module can be redistributed to other software modules and / or aggregated into a single module and / or made available in a shareable library. It should be understood that a flowchart format may be used, but the software may also be implemented as a state machine or other logical method.
[0016] The principles described herein can be implemented as hardware, software, firmware, or a combination thereof. Accordingly, the exemplary components, blocks, modules, circuits, and steps are described from the perspective of their functionality.
[0017] Furthermore, with respect to what was suggested above, the logic blocks, modules, and circuits described below can be implemented or executed by a general-purpose processor, a digital signal processor (DSP), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC), other programmable logic devices such as individual gates or transistor logic, individual hardware components, or any combination thereof, designed to perform the functions described herein. The processor can be implemented by a controller or state machine of a computing device, or a combination thereof.
[0018] The functions and methods described below, when implemented in software, although not limited thereto, can be described in a suitable language such as C# or C++, and stored in or transmitted via a computer-readable storage medium such as random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), other optical disk storage such as compact disc read-only memory (CD-ROM) or digital versatile disc (DVD), magnetic disk storage or other magnetic storage devices such as removable thumb drives. A computer-readable medium can be established by a connection. Such connections can include, by way of example, hardwire cables including optical fiber and coaxial wire, and digital subscriber line (DSL) and twisted pair wire.
[0019] The components included in one embodiment can be used in any suitable combination in other embodiments. For example, any of the various components described herein and / or shown in the figures may be combined, interchanged, or excluded from other embodiments.
[0020] "A system having at least one of A, B, and C" (similarly, "a system having at least one of A, B, or C" and "a system having at least one of A, B, C") includes, for example, a system having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together.
[0021] Here, specifically referring to FIG. 1, an exemplary system 10 is shown that may include one or more of the exemplary devices described above and further described below in accordance with this principle. Note that the computerized devices described in the figures herein can include some or all of the components described for the various devices of FIG. 1.
[0022] The first device among the exemplary devices included in system 10 is a device of a consumer electronics (CE) product, and this CE device is configured as an exemplary primary display device. In the illustrated embodiment, without limitation, it is an audio-video display device (AVDD) 12 such as an Internet-enabled TV with a TV tuner (equally, a set-top box for controlling a TV). AVDD 12 may be an Android (registered trademark)-based system. Alternatively, AVDD 12 may also be a computer-controlled Internet-enabled ("smart") phone, a tablet computer, a notebook computer, for example, a computer-controlled Internet-enabled clock, a computer-controlled Internet-enabled bracelet, other wearable computer-controlled devices such as other computer-controlled Internet-enabled devices, other computer-controlled devices, a computer-controlled Internet-enabled music player, a computer-controlled Internet-enabled headset, a computer-controlled Internet-enabled implantable device such as an implantable device for the skin, etc. In any case, it should be understood that AVDD 12 and / or other components described herein are configured to implement the present principles (e.g., communicate with other CE devices to implement the present principles, execute the logic described herein, and perform any other functions and / or operations described herein).
[0023] Therefore, to implement such a principle, AVDD12 can be established by some or all of the components shown in FIG. 1. For example, AVDD12 can include one or more displays 14, which may be implemented by a flat screen with high or ultra-high resolution, such as "4K" or higher resolution, and may or may not be touch-responsive to receive user input signals via touch on the display. Also, AVDD12 can include one or more speakers 16 for outputting audio according to this principle, and at least one additional input device 18, such as an audio receiver / microphone, for inputting audible commands to AVDD12 to control AVDD12. Further, an exemplary AVDD12 can include one or more network interfaces 20 for communicating through at least one network 22, such as the Internet, other wide area networks (WANs), local area networks (LANs), personal area networks (PANs), etc., under the control of one or more processors 24. Thus, the interface 20 can be, without limitation, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as a mesh network transceiver. The interface 20 can be a Bluetooth transceiver, Zigbee transceiver, IrDA transceiver, wireless USB transceiver, wired USB, wired LAN, Powerline, or MoCA, but is not limited thereto. It will be understood that the processor 24 controls AVDD12 to implement this principle, including other elements of AVDD12 described herein, such as controlling the display 14 to present images and receiving inputs therefrom. Further, note that the network interface 20 can be, for example, a wired or wireless modem or router, or other suitable interface, such as a wireless telephone transceiver or the aforementioned Wi-Fi transceiver.
[0024] In addition to the above, AVDD12 may also include one or more input ports 26, such as a high-definition multimedia interface (HDMI (registered trademark)) port or a USB port for physically connecting (e.g., using a wired connection) to another CE device, and / or a headphone port for connecting headphones to AVDD12 to provide audio to the user through the headphones. For example, the input port 26 may be connected wired or wirelessly to a cable or satellite source 26a of audio-video content. Thus, the source 26a may be, for example, a separate or integrated set-top box, or a satellite receiver. Alternatively, the source 26a may be a game console or a disc player.
[0025] AVDD12 may further include one or more computer memories 28, such as disk-based storage or solid-state storage, which are not temporary signals, and these storages may, in some cases, be embodied as a stand-alone device within the chassis of AVDD, or as a personal video recording device (PVR) or a video disc player either inside or outside the chassis of AVDD for playing AV programs, or as a removable memory medium. Also, in some embodiments, AVDD12 is configured to receive geographical location information from, for example, at least one satellite or cell phone tower and provide that information to the processor 24, and / or is configured to determine the altitude at which AVDD12 is disposed together with the processor 24, and may include receivers for position or location, such as a cell phone receiver, a GPS receiver, and / or an altimeter 30. However, it should be understood that other suitable position receivers, other than a cell phone receiver, a GPS receiver, and / or an altimeter, may be used to determine the position of AVDD12 in all three dimensions, for example, in accordance with this principle.
[0026] Continuing with the description of AVDD12, in one embodiment, AVDD12 may include one or more cameras 32, and the one or more cameras 32 may be, for example, digital cameras such as thermal imaging cameras, web cameras, and / or cameras integrated with AVDD12 and controllable by processor 24 to collect photos / images and / or videos in accordance with the present principle. Also, AVDD12 may include a Bluetooth® transceiver 34 and other near field communication (NFC) elements 36, which communicate with other devices using Bluetooth® and / or NFC technology, respectively. An exemplary NFC element may be a radio frequency identification (RFID) element.
[0027] Furthermore, AVDD12 may include one or more auxiliary sensors 38 (such as motion sensors like accelerometers, gyroscopes, cyclometers, or magnetic sensors, IR sensors for receiving infrared (IR) commands from a remote control, optical sensors, speed sensors and / or cadence sensors, gesture sensors (such as sensors for detecting gesture commands), etc.) that provide an input to processor 24. AVDD12 may include a wireless TV broadcast port 40 for receiving OTA (over-the-air) TV broadcasts that provide an input to processor 24. In addition to the above, it should be noted that AVDD12 may also include an infrared (IR) transmitter and / or IR receiver and / or IR transceiver 42 such as an infrared data association (IRDA) device. A battery (not shown) may be provided to power AVDD12.
[0028] Furthermore, in some embodiments, AVDD12 may include a graphics processing unit (GPU) 44 and / or a field programmable gate array (FPGA) 46. The GPU and / or FPGA may be utilized by AVDD12 for artificial intelligence processing, such as training a neural network and performing operations (e.g., inference) of the neural network, according to, for example, the present principles. Note, however, that the processor 24 can also be used for artificial intelligence processing, such as when the processor 24 can be a central processing unit (CPU).
[0029] Referring further to FIG. 1, in addition to AVDD12, the system 10 may include one or more other types of computer devices that may include some or all of the components shown in AVDD12. In one example, a first device 48 and a second device 50 are shown and may include components similar to some or all of the components of AVDD12. Fewer or more devices than shown may be used.
[0030] The system 10 may also include one or more servers 52. The server 52 may include at least one server processor 54, at least one computer memory 56, such as disk-based storage or solid-state storage, and at least one network interface 58 that enables communication with other devices in FIG. 1 through the network 22 under the control of the server processor 54 and, in fact, may facilitate communication between servers, controllers, and client devices according to the present principles. Note that the network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi® transceiver, or other suitable interface, such as, for example, a wireless telephone transceiver.
[0031] Thus, in some embodiments, server 52 may be an Internet server, may include a "cloud" function, may execute a "cloud" function, and enable the devices of system 10 to access a "cloud" environment via server 52 in an exemplary embodiment. Alternatively, server 52 may be implemented by a game console or other computer that is in the same room as, or near, the other devices shown in FIG. 1.
[0032] The following devices can incorporate some or all of the above elements.
[0033] "Geographically separated" refers to positions that are beyond visual and auditory range from each other, usually positions that are more than one mile apart from each other.
[0034] FIG. 2 illustrates exemplary logic consistent with this principle in an exemplary flowchart format. Basically, the projector is used for motion capture (mocap), tracks the mocap of the perspectives (POV) of multiple actors in multiple stages that are geographically separated, and is used to perform position tracking for fidelity.
[0035] Starting from block 200, the movements of each of the multiple actors are captured at their respective stages or other locations, for example, using a projector, by reflecting light from reflective tags or other markers held by the actors. In block 202, the video images of the mocap data of each actor are merged into a single scene by aligning the frames in time with each other in real-world time or the time of the video scene, for example, using timestamps added to the frames of the mocap of each actor. In block 204, the mocap of the multiple actors is integrated into a single scene, which is done in real-time or near real-time when the actors are being filmed, so in block 206, the director can give instructions to the actors by providing stage directions to the actors, as will be described in more detail below.
[0036] In practice, FIG. 3 shows a screen shot 300 of an exemplary stage display 302 that can be attached to a soundproof studio or other filming location taking into account the actor when the actor performs for mocap. A text stage direction 304 can be presented on the display to prompt the actor to perform a specific action, such as looking up to the left at the virtual position of the dragoon in the video. Audio prompts, such as beep sounds or voice instructions, can be emitted by one or more speakers 306 for the same effect, for example, to emit a beep sound from a speaker located at the upper left corner of the display. In this way, the mocap actor can look at the monitors lined up on the motion capture stage to confirm himself and the animation he is reacting to (e.g., so as not to collide with people or walls in the animation).
[0037] FIG. 4 shows an example of a distributed acting guidance environment 400 that exemplifies two remotely located studios or movie sets 402 where one or more respective actors 404 are performing for mocap purposes. One or more displays 406 and / or speakers 408 can be attached to the studio 402 as shown and can be instantiated, for example, by the display 302 shown in FIG. 3. The video feed of the actor 404 can be transmitted from the studio or set 402 to a director location 410 remote therefrom via a wired and / or wireless path of a wide area network (WAN) in, for example, a WebEx-style feed, where a person 412, such as a director or a quality control (QC) technician, can operate a director computer 414 that presents the video from each set or studio 402.
[0038] Thus, FIG. 4 shows that with multiple stages / studio, an actor's mocap video can be captured and these videos can be streamed into one virtual reality (VR) world, where an orchestrator computer 414 integrates each stream into a single video. This facilitates creating large scenes of multiple people based on actors who are geographically separated from each other, and in particular, helps to aggregate the body motion capture of multiple actors on their respective multiple stages.
[0039] The operator / leader (QC operator) who integrates the mocap videos can be located at a position 410 remote from the stage 402, for example, using a virtual private network (VPN) on the network. Using remote access software, the mocap videos can be moved to the QC computer 414, and the QC computer can provide feedback on stage production and other information such as whether the remote camera collided and whether another take is needed.
[0040] FIG. 5 shows a wall of a retroreflector 500 arranged within a grid that can be illuminated by one or more projectors 502 attached by one or more booms 504 to a head-mounted display (HMD) 506 of a mocap actor to illuminate a retroreflector applicable to the wall of a movie set 402, for example. This gives the mocap actor wearing the HMD 506 a real-world reference point when viewing the reflection of the projector from the retroreflector 500 by the HMD display 508. One or more speakers 510 can be provided on the HMD, and the output of the HMD, including the control of the projector 502, can be provided by one or more processors 512 accessing one or more transceivers 514.
[0041] The "walls" of the retroreflective material 500 and the projector 502 on the HMD 506 give different actors different reference frames with respect to the walls. Only the person wearing the HMD 506 can see the projection reflection from their own perspective. In this way, by using the walls of the retroreflective material 500, the virtual set of actors can check the necessary references without interfering with each other and can see the aligned reference reflections.
[0042] It should be understood that the HMD 506 may include one or more internal cameras 516 to track the wearer's head and eyes, better resolve what the actor would see in the virtual environment, feed that scene to the projector 502, and project an appropriate image onto the retroreflector 500.
[0043] Figure 6 shows further features of the retroreflector 500, where an HMD projector such as the projector 502 of Figure 5 projects various images from an existing virtual scene in which the actor's mocap is integrated. The actor's image 602 can be projected onto the retroreflective material 500 along with various visual or audible identifications 604 of the images for viewing the actor's image. In the example of Figure 6, an image 606 of a dragon is projected at a location within a virtual world that emulates the dragon, along with an image 608 of another character-based actor that uses mocap video from other actors.
[0044] Figure 7 shows additional exemplary logic in an exemplary flowchart format that is consistent with an embodiment in which the reference projection is presented on the display of the HMD 506. Starting at block 700, for an HMD with an internal visible retroreflector, at block 702, the head and eye pose can be tracked based on an image from the internal camera of the HMD. Proceeding to block 704, the scale and dimensions of the image projected onto the HMD 506, derived from an existing video of the virtual scene, can be changed based on the head / eye tracking at block 702. In this way, the actor can look up at the image of the head of a character emulated as being taller than the actor, or look down at the image of a character emulated as being shorter than the actor.
[0045] Alternatively, FIG. 8 shows a stage set 800 where the floor 802 has retroreflective markers 804 that allow the actor 806 to walk, and a projector on the HMD or elsewhere within the movie set projects an image onto the marker 804 to assist the actor 806 when navigating virtual objects within the movie set during mocap. It should be understood that the techniques described above provide pre-recorded animations or videos against which mocap actors can perform.
[0046] FIG. 9 shows additional exemplary logic where, in block 900, a computer game engine is used with a plugin computer program to stream game data to the plugin, and the plugin automatically sends the game data, which includes audio and video, to any of the projectors in this specification, and the projector presents a reference image to assist the performance of the mocap actor.
[0047] FIG. 10 shows a screenshot on an exemplary HMD 506. Similar to the outer wall of the retroreflector 500 shown in FIGS. 5 and 6, an internal projector within the HMD 506 can project the actor's image 1000 along with visible or audible identifications 1002 of various images onto the display of the HMD for the actor to view their image. In the example of FIG. 10, the image 1004 of a dragon is projected to a position within a virtual world emulating the dragon, along with the image 1006 of another character-based actor using mocap video from other actors. The display of the HMD may be a reflective surface such as a visor on the HMD being used for the projection, and head tracking is used to accurately size and scale the various images.
[0048] Regardless of whether it is projected onto a wall, floor, or HMD component, a character such as the aforementioned dragon is part of a reference video and can be displayed to the actor as being in the same position within the VR space on different stages that are geographically remote. Also, by sending the mocap feed of one actor to the display of another remotely located actor, both actors can receive the presence of the other actor on another stage with the same dragon and the actor they are facing. In this way, each actor is captured in the video, this video is sent to an unrealistic virtual representation, and virtual characters such as dragons are merged with the mocap video of real actors into a single scene. In this way, people on all stages can view the same integrated scene and make appropriate adjustments. Regardless of what (another actor or a pre-recorded character) the actor is supposed to react to, the screen or cage system gives an indication to the actor of where they need to look.
[0049] Therefore, the problem being solved is to give the actor a physical reference on stage. The audio being played can also be a cue to the actor.
[0050] In the above example, the reference image can be stabilized by converting the image based on the movement of the head of the mocap actor. Also, just before a character appears in the VR world, an audible alert such as a "beep sound" may be emitted before the time when the character appears. The network synchronization protocol is preferably implemented among the distributed computers of this specification to ensure that various videos are aligned by frames and the same scene.
[0051] Although the principle has been described with reference to several exemplary embodiments, these are not intended to be limiting, and it will be understood that various alternative configurations may be used to implement the subject matter claimed in this specification.
Claims
1. while photographing a first actor wearing a head-mounted display (HMD) for motion capture (mocap), providing a reference frame to the first actor at least partially at a first position by presenting, on the HMD, the head of at least one emulated character, the HMD includes one or more internal cameras for tracking the head and eyes of the first actor, wherein the scale and dimensions of the reference frame are changed based on the tracking of the head and eyes, a method.
2. providing a reference frame to at least a second actor at a second position while photographing the second actor for mocap, wherein the first position is geographically separated from the second position, the providing, presenting mocap from the first actor and the second actor on at least one director display communicating with the first position and the second position during a WebEx, The method according to claim 1, comprising:
3. The method according to claim 1, comprising providing a reference frame to the first actor using at least partially the audio reproduced at the first position.
4. The method according to claim 1, comprising providing a plurality of light emitters on the HMD.
5. The method according to claim 2, comprising synchronizing the mocap videos of the first actor and the second actor in time.
6. An apparatus comprising at least one head-mounted display (HMD) assembly, wherein the at least one HMD assembly includes: at least one processor configured by instructions, at least one display, at least one speaker, at least one projector configured to project a reference image for motion capture (mocap) onto at least one surface visible to the wearer of the HMD assembly to provide a spatial reference to the wearer on the HMD assembly during mocap, one or more internal cameras for tracking the head and eyes of the wearer, including wherein the scale and dimensions of the spatial reference are changed based on the tracking of the head and eyes, the apparatus.
7. The apparatus according to claim 6, wherein the projector is mounted on a boom of the HMD assembly.
8. The apparatus of claim 6, comprising at least one wireless transceiver on the HMD assembly for receiving commands from a remotely located performer computer. **Claim 9** The apparatus of claim 8, wherein the instructions are executable to present the commands on the display and / or to play audio on the speaker in response to the commands.
Citation Information
Patent Citations
Head-mounted video display device
JP2001177851A
Information presenting device
JP2008198196A
Mixed reality cinematography using remote activity stations
US20190102949A1