Motion capture reference frame

The method uses HMDs with retroreflectors and projectors to create spatial references for remote actors, synchronizing and merging motion capture videos, addressing coordination challenges in multi-location film and simulation productions.

JP2025122164APending Publication Date: 2025-08-20SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025088789
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-25
Filing Date
2025-05-28
Publication Date
2025-08-20

AI Technical Summary

Technical Problem

Collaborative production of films and computer simulations using remote actors faces challenges in coordinating physical references across geographically distant locations, particularly in motion capture (MoCap) scenarios.

Method used

A method is provided that includes using a head-mounted display (HMD) with a retroreflector and projector to create a frame of reference for actors, synchronized with audio cues, and merging motion capture videos from multiple locations into a single scene, facilitated by a computer network with synchronized video feeds and director displays.

Benefits of technology

Enables effective coordination and synchronization of remote actors' performances, allowing for integrated audio-visual productions by providing real-time spatial references and synchronized video merging, thus overcoming location-based coordination challenges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025122164000001_ABST
    Figure 2025122164000001_ABST
Patent Text Reader

Abstract

To provide a technique for facilitating the coordination of audio-video (AV) productions using multiple actors (404) at remote locations (402) such that the activities of the multiple remote actors (404) can be coordinated to produce an integrated AV product.SOLUTION: Motion capture (mocap) of multiple actors who are geographically separated from one another can be facilitated (200).SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates generally to technically inventive and unconventional solutions that are necessarily rooted in computer technology and result in specific technical improvements. In particular, this application relates to techniques for enabling multi-location collaborative remote performance instruction. [Background technology]

[0002] Due to health and cost considerations, people are increasingly collaborating from remote locations. As understood herein, collaborative production of films and computer simulations (e.g., computer games) using remote actors can pose unique coordination challenges. This is because a director must direct multiple actors, each of whom may be in their own studio or soundstage, when producing a film and for computer simulation-related activities such as motion capture (MoCap). For example, in ways in which performances are coordinated, challenges exist in providing physical references for remote actors on individual stages. The present principles provide techniques for addressing some of these coordination challenges. Summary of the Invention

[0003] Thus, the present principles provide a method that includes providing a frame of reference for at least a first actor at a first location while filming the first actor for motion capture (mocap) by at least partially presenting at least one reference image to a head-mounted display (HMD) worn by the first actor. The light reflected from the retroreflector can be from an illuminator. Additionally or alternatively, the method can include providing a retroreflector on a wall of the first location to reflect light toward the first actor. Additionally or alternatively, the method can include providing a visible marker on a floor of the first location.

[0004] In some examples, the method can include providing a frame of reference for at least a second actor at a second location while filming the second actor for mocap. The first location can be geographically remote from the second location, and the mocap from the first actor and the second actor can be presented on at least one director display in communication with the first and second locations during a web rehearsal (WebEx).

[0005] The first actor can be provided with a frame of reference at least in part using audio played at the first location. Multiple lights may be provided on the HMD. The Mocap videos of the first and second actors can be synchronized in time.

[0006] In another aspect, the device includes at least one computer storage, the at least one computer storage including instructions, other than the transient signal, executable by the at least one processor, for receiving motion capture (mocap) video of a first actor from a first camera at a first location, the instructions being executable to receive mocap video of a second actor from a second camera at a second location, synchronizing the mocap videos with each other, and merging the mocap videos into a single scene on at least one display at a third location geographically remote from the first and second locations.

[0007] In another aspect, the device includes at least one head-mounted display (HMD) assembly, the at least one HMD assembly including at least one processor configured with instructions and at least one display controlled by the processor. The HMD may also include a speaker. The at least one projector is configured to project a motion capture (mocap) reference light onto at least one surface visible to a wearer of the HMD assembly to provide a spatial reference for the wearer during mocap.

[0008] The details of the present application, both as to its structure and operation, can best be understood in reference to the accompanying drawings, in which like reference numerals refer to like parts, and in which: [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a block diagram of an exemplary system consistent with the present principles. [Figure 2] Exemplary logic consistent with the present principles is presented in exemplary flow chart form. [Figure 3] 1 shows a screenshot of an exemplary stage display. [Figure 4] 1 illustrates an exemplary distributed performance direction environment showing an example of two remote studios or film sets and a remote director's computer presenting video from each set or studio. [Figure 5] 1 shows a camera on a boom of a head-mounted display (HMD) for illuminating retroreflectors on the walls of a movie set. [Figure 6] 1 illustrates further features of the retroreflector. [Figure 7] Additional exemplary logic consistent with the present principles is illustrated in exemplary flow chart form. [Figure 8] 1 shows markers on the floor of a movie set to assist actors within the set. [Figure 9] Additional exemplary logic consistent with the present principles is illustrated in exemplary flow chart form. [Figure 10] 1 shows a screenshot on an exemplary HMD. DETAILED DESCRIPTION OF THE INVENTION

[0010] Referring now to FIG. 1 , the present disclosure generally relates to a computer ecosystem having aspects of a computer network that may include consumer electronics (CE) devices. The system herein may include a server component and a client component connected via a network such that data may be exchanged between the client component and the server component. The client component may include one or more computing devices, including portable computers such as portable televisions (e.g., smart TVs, Internet-enabled TVs), laptop computers, and tablet computers, as well as smartphones and other mobile devices, including additional examples discussed below. These client devices may operate in a variety of operating environments. For example, some client computers may use, by way of example, Microsoft® operating systems, or Unix® operating systems, or operating systems manufactured by Apple Computer® or Google®. These operating environments may be used to run one or more browsing programs, such as browsers created by Microsoft®, Google®, or Mozilla®, or other browser programs capable of accessing websites hosted by Internet servers, as discussed below.

[0011] The server and / or gateway may include one or more processors that execute instructions that configure the server to receive and transmit data over a network such as the Internet. Alternatively, the client and server may be connected through a local intranet or a virtual private network. The server or controller may be instantiated by a game console such as a Sony PlayStation®, a personal computer, or the like.

[0012] Information may be exchanged between the client and the server over a network. For this purpose and for security, the server and / or client may include firewalls, load balancers, temporary storage, and proxies, as well as other network infrastructure for reliability and security.

[0013] As used herein, instructions refer to computer-implemented steps for processing information in a system. Instructions can be implemented in software, firmware, or hardware and can include any type of programmed step performed by a component of the system.

[0014] The processor may be a general purpose single-chip processor or a general purpose multi-chip processor capable of implementing logic through various lines such as address lines, data lines, and control lines, as well as registers and shift registers.

[0015] The software modules described herein by flowcharts and user interfaces may include various subroutines, procedures, etc. Without limiting the disclosure, the logic specified to be performed by a particular module may be redistributed among other software modules and / or aggregated together in a single module and / or made available in a shareable library. While a flowchart format may be used, it should be understood that the software may also be implemented as a state machine or other logical method.

[0016] The principles described herein may be implemented as hardware, software, firmware, or a combination thereof. Accordingly, the illustrative components, blocks, modules, circuits, and steps are described in terms of their functionality.

[0017] Further, as alluded to above, the logic blocks, modules, and circuits described below may be implemented or performed by a general purpose processor, digital signal processor (DSP), field programmable gate array (FPGA), or other programmable logic device, such as an application specific integrated circuit (ASIC), discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described herein. A processor may be implemented by a controller or state machine, or combination, of a computing device.

[0018] The functions and methods described below, when implemented in software, can be written in a suitable language, such as, but not limited to, C# or C++, and can be stored on or transmitted via a computer-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), other optical disk storage, such as compact disk read-only memory (CD-ROM) or digital versatile disk (DVD), magnetic disk storage, or other magnetic storage devices, including removable thumb drives, etc. The computer-readable medium can be established by a connection. Such connections can include, by way of example, hardwire cables, including optical fiber and coaxial wire, as well as digital subscriber line (DSL) and twisted pair wire.

[0019] Components included in one embodiment may be used in other embodiments in any suitable combination. For example, any of the various components described herein and / or illustrated in the figures may be combined, interchanged, or omitted from other embodiments.

[0020] "A system having at least one of A, B, and C" (and similarly "a system having at least one of A, B, or C" and "a system having at least one of A, B, C") includes systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.

[0021] 1, there is shown an exemplary system 10 in accordance with the present principles that may include one or more of the exemplary devices described above and further below. It should be noted that the computerized devices depicted in the figures herein may include some or all of the components described for the various devices in FIG.

[0022] The first of the exemplary devices included in system 10 is a consumer electronics (CE) device, which is configured as an exemplary primary display device and, in the illustrated embodiment, is an audio-video display device (AVDD) 12, such as, without limitation, an Internet-enabled TV with a TV tuner (equivalently, a set-top box that controls the TV). AVDD 12 may also be an Android®-based system. Alternatively, AVDD 12 may also be a computer-controlled Internet-enabled (“smart”) phone, a tablet computer, a notebook computer, a wearable computer-controlled device such as, for example, a computer-controlled Internet-enabled watch, a computer-controlled Internet-enabled bracelet, or other computer-controlled Internet-enabled device, other computer-controlled device, a computer-controlled Internet-enabled music player, computer-controlled Internet-enabled headphones, a computer-controlled Internet-enabled implantable device such as an implantable skin device, or the like. In any event, it should be understood that AVDD 12 and / or other components described herein are configured to implement the present principles (e.g., communicate with other CE devices to implement the present principles, execute the logic described herein, and perform any other functions and / or operations described herein).

[0023] Accordingly, to implement these principles, AVDD 12 can be established by some or all of the components shown in FIG. 1. For example, AVDD 12 can include one or more displays 14, which may be implemented by flat screens with high or ultra-high resolution (4K) or higher resolution, and which may or may not be touch-enabled for receiving user input signals via touching the display. AVDD 12 can also include one or more speakers 16 for outputting audio in accordance with the present principles and at least one additional input device 18, such as an audio receiver / microphone, for inputting audible commands to AVDD 12 and controlling AVDD 12. Furthermore, an exemplary AVDD 12 can include one or more network interfaces 20 for communicating over at least one network 22, such as the Internet, another wide area network (WAN), a local area network (LAN), or a personal area network (PAN), under the control of one or more processors 24. Thus, interface 20 may be, without limitation, a Wi-Fi® transceiver, which is an example of a wireless computer network interface, such as, without limitation, a mesh network transceiver. Interface 20 may also be, without limitation, a Bluetooth® transceiver, a Zigbee® transceiver, an IrDA transceiver, a wireless USB transceiver, a wired USB, a wired LAN, Powerline, or MoCA. It will be understood that processor 24 controls AVDD 12 to implement the present principles, including other elements of AVDD 12 described herein, such as, for example, controlling display 14 to present images and receiving input therefrom. Furthermore, it should be noted that network interface 20 may be, for example, a wired or wireless modem or router, or other suitable interface, such as, for example, a wireless telephony transceiver or the Wi-Fi® transceiver described above.

[0024] In addition to the above, AVDD 12 may also include one or more input ports 26, such as a High-Definition Multimedia Interface (HDMI®) port or a USB port for physically connecting to another CE device (e.g., using a wired connection) and / or a headphone port for connecting headphones to AVDD 12 to provide audio from AVDD 12 to a user through the headphones. For example, input port 26 may be connected wired or wirelessly to a cable or satellite source 26a of audio-video content. Thus, source 26a may be, for example, a separate or integrated set-top box, or a satellite receiver. Alternatively, source 26a may be a game console or disc player.

[0025] AVDD 12 may further include one or more computer memories 28, such as non-transitory disk-based or solid-state storage, which may in some cases be embodied within the AVDD chassis as a standalone device, or as a personal video recording device (PVR) or video disc player, or as removable memory media, either internal or external to the AVDD chassis for playing AV programs. In some embodiments, AVDD 12 may also include a position or location receiver, such as, but not limited to, a cellular telephone receiver, a GPS receiver, and / or an altimeter 30, configured to receive geographic location information from at least one satellite or cellular telephone tower and provide that information to processor 24 and / or to determine the altitude at which AVDD 12 is located in conjunction with processor 24. However, it should be understood that another suitable position receiver other than a cellular telephone receiver, a GPS receiver, and / or an altimeter may be used in accordance with the present principles, for example, to determine the position of AVDD 12 in all three dimensions.

[0026] Continuing with the description of AVDD 12, in one embodiment, AVDD 12 may include one or more cameras 32, which may be, for example, a thermal imaging camera, a digital camera such as a webcam, and / or a camera integrated into AVDD 12 and controllable by processor 24 to collect pictures / images and / or video in accordance with the present principles. AVDD 12 may also include a Bluetooth transceiver 34 and other near field communication (NFC) elements 36, which communicate with other devices using Bluetooth and / or NFC technology, respectively. An exemplary NFC element may be a radio frequency identification (RFID) element.

[0027] Furthermore, AVDD 12 may include one or more auxiliary sensors 38 (e.g., motion sensors such as an accelerometer, gyroscope, cyclometer, or the like, or magnetic sensors, infrared (IR) sensors for receiving IR commands from a remote control, optical sensors, speed and / or cadence sensors, gesture sensors (e.g., sensors for detecting gesture commands), etc.) that provide input to processor 24. AVDD 12 may include an over-the-air TV broadcast port 40 for receiving over-the-air (OTA) TV broadcasts that provide input to processor 24. In addition to the above, it should be noted that AVDD 12 may also include an infrared (IR) transmitter and / or IR receiver and / or IR transceiver 42, such as an infrared data association (IRDA) device. A battery (not shown) may be included to power AVDD 12.

[0028] Additionally, in some embodiments, AVDD 12 may include a graphics processing unit (GPU) 44 and / or a field programmable gate array (FPGA) 46. The GPU and / or FPGA may be utilized by AVDD 12 for artificial intelligence processing, such as training neural networks and performing neural network operations (e.g., inference), in accordance with present principles. Note, however, that processor 24 may also be used for artificial intelligence processing, such as in the case where processor 24 may be a central processing unit (CPU).

[0029] 1, in addition to AVDD 12, system 10 may include one or more other types of computing devices that may include some or all of the components shown in AVDD 12. In one example, first device 48 and second device 50 are shown and may include components similar to some or all of the components of AVDD 12. Fewer or more devices than shown may be used.

[0030] System 10 may also include one or more servers 52. Server 52 may include at least one server processor 54, at least one computer memory 56, such as disk-based or solid-state storage, and at least one network interface 58 that, under the control of server processor 54, enables communication with other devices of FIG. 1 over network 22 and, indeed, may facilitate communication between servers, controllers, and client devices in accordance with the present principles. Note that network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi® transceiver, or other suitable interface, such as, for example, a wireless telephony transceiver.

[0031] Thus, in some embodiments, server 52 may be an Internet server or may include or perform "cloud" functionality, allowing devices of system 10 to access the "cloud" environment via server 52 in exemplary embodiments. Alternatively, server 52 may be implemented by a game console or other computer located in or near the same room as the other devices shown in FIG. 1.

[0032] The devices described below may incorporate some or all of the elements described above.

[0033] "Geographically separated" refers to locations beyond sight and hearing of one another, typically more than one mile from one another.

[0034] 2 illustrates, in exemplary flow chart form, exemplary logic consistent with the present principles. Essentially, projectors are used for motion capture (mocap) to track the points of view (POV) of multiple actors on multiple geographically separated stages and to perform positional tracking for fidelity.

[0035] Beginning in block 200, the movements of each of multiple actors are captured at their respective stage or other locations, e.g., using a projector to reflect light from reflective tags or other markers carried by the actors. Video images of each actor's mocap data are merged into a single scene in block 202, e.g., by using timestamps added to each actor's mocap frames to time-align the frames with one another in real-world time or the time of the video scene. The mocaps of the multiple actors are combined into a single scene in block 204, which occurs in real time or near real time as the actors are being filmed, so that in block 206 a director can direct the actors by providing them with stage directions, as described in more detail below.

[0036] In practice, Figure 3 shows a screenshot 300 of an exemplary stage display 302 that can be mounted in a soundstage or other filming location to accommodate actors as they perform for mocap. Textual stage directions 304 can be presented on the display to prompt actors to take specific actions, such as looking up and to the left for the virtual location of the dragoon in the video. Audio prompts, such as beeps or voice instructions, can be emitted by one or more speakers 306 for the same effect, e.g., a beep from a speaker located in the upper left corner of the display. In this way, mocap actors can view monitors aligned with the motion capture stage to see themselves and the animation they are reacting to (e.g., to avoid bumping into people or walls in the animation).

[0037] 4 shows an example of a distributed acting coaching environment 400 illustrating two remote studios or film sets 402 with one or more respective actors 404 performing for mocap purposes. One or more displays 406 and / or speakers 408 can be attached to the studios 402 as shown, and can be instantiated, for example, by the display 302 shown in FIG. 3. Video feeds of the actors 404 can be transmitted from the studios or sets 402 over wired and / or wireless paths of a wide area network (WAN), for example, during a Web rehearsal (WebEx) style feed, to a remote director's location 410, where a person 412, such as a director or quality control (QC) technician, can operate a director's computer 414 presenting the video from each set or studio 402.

[0038] 4 thus shows that multiple stages / studios can be used to capture mocap video of actors, which can then be streamed into a single virtual reality (VR) world and merged into a single video at director computer 414. This facilitates creating large multi-person scenes based on actors who are geographically distant from each other, and is particularly useful for aggregating body motion capture on multiple stages for each of multiple actors.

[0039] The operator / leader (QC operator) integrating the mocap video can be at a location 410 remote from the stage 402, for example using a virtual private network (VPN) over the network. Using remote access software, the mocap video can be moved to a QC computer 414, which can provide feedback on the stage direction and other information such as whether a remote camera has collided or whether another take is needed.

[0040] 5 shows a wall of retroreflectors 500 arranged in a grid that can be illuminated by one or more projectors 502 attached by one or more booms 504 to a mocap actor's head-mounted display (HMD) 506 to illuminate the retroreflectors that can be applied, for example, to the walls of a movie set 402. This gives a mocap actor wearing the HMD 506 a real-world reference point when he or she views the projector's reflection from the retroreflector 500 through the HMD display 508. One or more speakers 510 can be provided on the HMD, and the output of the HMD, including control of the projector 502, can be provided by one or more processors 512 accessing one or more transceivers 514.

[0041] The "wall" of retroreflective material 500 and the projector 502 on the HMD 506 give different actors different frames of reference relative to the wall. Only the person wearing the HMD 506 can see the projected reflection from their perspective. In this way, using the wall of retroreflective material 500, actors on the virtual set can see the reference they need without obstructing each other and can see the reference reflection to which they are aligned.

[0042] It should be appreciated that the HMD 506 may include one or more internal cameras 516 to track the wearer's head and eyes to better resolve what the actor would see in the virtual environment and feed that scene to the projector 502 to project an appropriate image onto the retroreflector 500.

[0043] Figure 6 illustrates a further feature of retroreflector 500, in which an HMD projector, such as projector 502 in Figure 5, projects various images from an existing virtual scene into which actor mocap is integrated. An actor image 602 may be projected onto retroreflector 500 along with a visual or audible identification 604 of the various images to view the actor's image. In the example of Figure 6, an image of a dragon 606 is projected at a location within the virtual world emulating a dragon, along with an image of another character-based actor 608 using mocap video from another actor.

[0044] 7 illustrates, in exemplary flowchart form, additional exemplary logic consistent with an embodiment in which a reference projection is presented on the display of the HMD 506. Beginning at block 700, in the case of an HMD with an internally visible retroreflector, head and eye poses can be tracked based on images from the HMD's internal camera at block 702. Proceeding to block 704, the scale and dimensions of the image projected onto the HMD 506, derived from existing video of the virtual scene, can be altered based on the head / eye tracking at block 702. In this manner, the actor can look up toward an image of the emulated character's head that is taller than the actor, or down toward an image of the emulated character that is shorter than the actor.

[0045] 8 shows a stage set 800 in which a floor 802 has retro-reflective markers 804 on which an actor 806 can walk, and a projector on an HMD or elsewhere in the stage set 800 projects images onto the markers 804 to assist the actor 806 in navigating virtual objects in the movie set during mocap. It should be understood that the techniques described above can provide pre-recorded animations or videos against which mocap actors can perform.

[0046] FIG. 9 illustrates, in block 900, additional exemplary logic for using a computer game engine in conjunction with a plug-in computer program to stream game data to the plug-in, which automatically transmits the game data, including audio and video, to any of the projectors herein, which present reference images to assist the mocap actors in their performance.

[0047] Figure 10 shows a screenshot on an exemplary HMD 506. Similar to the exterior wall of the retroreflector 500 shown in Figures 5 and 6, an internal projector within the HMD 506 can project images 1000 of actors along with visual or audible identification 1002 of the various images onto the HMD's display for viewing the actor's images. In the example of Figure 10, an image 1004 of a dragon is projected onto a location within the virtual world emulating a dragon, along with an image 1006 of another character-based actor using mocap video from another actor. The HMD's display may be a reflective surface, such as a visor on the HMD, used for projection, and head tracking is used to ensure accurate dimensions and scale of the various images.

[0048] Whether projected onto a wall, floor, or HMD component, a character, such as the aforementioned dragon, can be part of a reference video and appear to actors in the same location in VR space on a different, geographically remote stage. Also, by transmitting one actor's mocap feed to another remote actor's display, both actors can receive the presence of the other actor on a different stage, complete with the same dragon and facing actor. In this way, each actor can be captured on video, this video can be transmitted to an unreal virtual representation, and the virtual character, such as a dragon, can be merged with the mocap video of the real actor into a single scene. In this way, people on all stages can see the same integrated scene and make appropriate adjustments. Regardless of what the actor is supposed to react to (another actor or a pre-recorded character), a screen or cage system gives the actor an indication of where they need to look.

[0049] Thus, the problem being solved is to give actors a physical reference on stage. The audio being played can also be a cue for the actors.

[0050] In the above example, the reference image can be stabilized by transforming the image based on the mocap actor's head movement. Also, an audible alert, such as a "beep," may be emitted just before a character appears in the VR world. A network synchronization protocol can desirably be implemented between the distributed computers herein to ensure that various videos are aligned by frame and by the same scene.

[0051] While the present principles have been described with reference to certain exemplary embodiments, it will be understood that these are not intended to be limiting and that various alternative configurations may be used to implement the subject matter claimed herein.

Claims

1. providing a frame of reference for at least a first actor at a first location while filming the first actor wearing a head-mounted display (HMD) for motion capture (mocap), at least in part by providing a retroreflector on a wall of the first location to reflect light toward the first actor; The HMD includes a projector that illuminates the retroreflector to provide the first actor with a real-world reference point when the first actor views the projector's reflection from the retroreflector through the HMD, and the retroreflector provides different actors with different frames of reference relative to the wall so that actors on a virtual set can see the references they need without obstructing each other.

2. 1. A method comprising: while filming a first actor wearing a head-mounted display (HMD) for motion capture (mocap), providing a frame of reference for at least a first actor at a first location, at least in part, by providing visible markers on a floor at the first location and a projector that projects images onto the visible markers to assist the first actor in navigating virtual objects during motion capture of the first actor.

Citation Information

Patent Citations

  • Head-mounted video display device

    JP2001177851A

  • Information presenting device

    JP2008198196A

  • Mixed reality cinematography using remote activity stations

    US20190102949A1