Information processing device, information processing method, and system
The system addresses the challenge of presenting images around a subject by controlling imaging and display timing, enabling 3D model generation and interactive performances with enhanced entertainment value.
Patent Information
- Application Number
- JP2023510290
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-31
- Filing Date
- 2022-01-13
- Publication Date
- 2025-12-03
- Estimated Expiration
- 2042-01-13
AI Technical Summary
Existing methods for extracting a subject area from a captured image, such as using a green screen, struggle to present images other than the screen around the subject, limiting the entertainment value and realism in applications like remote live performances.
An information processing system that controls the timing of imaging and displaying images around a subject to allow simultaneous capture of 3D model images and presentation of audience images, using multiple imaging units and display areas, with staggered timing to facilitate 3D model generation and interactive performance.
Enables the generation of free viewpoint 3D images for remote audiences and interactive actions by performers, enhancing entertainment value and realism in live performances.
Smart Images

Figure 0007779310000001 
Figure 0007779310000002 
Figure 0007779310000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, an information processing method, and a system. [Background technology]
[0002] Conventionally, a green screen or a blue screen has been used to facilitate extraction of a person area (silhouette image of a subject) from a captured image. Regarding extraction of a subject area from a captured image, for example, Patent Document 1 listed below discloses a technology for generating a three-dimensional model of a subject using N RGB images acquired from N RGB cameras provided at positions surrounding the subject, and M pieces of active depth information indicating distances to the subject acquired from M active sensors similarly provided at positions surrounding the subject. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2019 / 107180 Summary of the Invention [Problem to be solved by the invention]
[0004] However, while using a green screen or the like can make it easier to extract the area of the subject, it has been difficult to present an image other than a green screen or the like around the subject.
[0005] Therefore, the present disclosure proposes an information processing device, an information processing method, and a system that can enhance entertainment value by simultaneously capturing an image of a subject and displaying an image around the subject. [Means for solving the problem]
[0006] According to the present disclosure, an information processing device is proposed that includes a control unit that controls imaging by multiple imaging units to obtain three-dimensional information of a subject and displays images obtained from the outside in one or more display areas located around the subject, and the control unit controls to make the timing of imaging different from the timing of displaying the images obtained from the outside in the display areas.
[0007] According to the present disclosure, an information processing method is proposed, which includes a processor controlling imaging by multiple imaging units to obtain three-dimensional information of a subject, and displaying images obtained from outside in one or more display areas located around the subject, and controlling the timing of performing the imaging to differ from the timing of displaying the images obtained from outside in the display areas.
[0008] According to the present disclosure, a system is proposed that includes a plurality of imaging devices arranged around a subject to acquire three-dimensional information of the subject, one or more display areas arranged around the subject, and an information processing device having a control unit that controls imaging by the plurality of imaging devices and display control for displaying images acquired from outside in the one or more display areas, wherein the control unit controls to differentiate the timing of capturing the images from the timing of displaying the images acquired from outside in the display areas. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a diagram illustrating an overview of an information processing system according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a diagram illustrating the arrangement of a display area (display) and an imaging unit (camera) in a studio where information for generating a 3D model of a performer is acquired according to this embodiment. [Figure 3] FIG. 1 is a diagram illustrating the arrangement of a display area (screen) and an imaging unit (camera) in a studio where information for generating a 3D model of a performer is acquired according to this embodiment. [Figure 4]2 is a block diagram mainly showing a specific configuration example of a display processing unit of the performer information input / output system according to the present embodiment. FIG. [Figure 5] 10 is a diagram showing the relationship between the emphasis degree of brightness / color correction and the length of a non-display period according to the present embodiment. FIG. [Figure 6] 10A and 10B are diagrams illustrating an example of timing control of display ON / OFF and imaging ON / OFF according to the present embodiment. [Figure 7] 2 is a block diagram mainly showing a specific example of the configuration of the video acquisition unit of the performer information input / output system according to the present embodiment. FIG. [Figure 8] 1 is a block diagram mainly showing a specific example of the configuration of a performer information generation unit of the performer information input / output system according to this embodiment. [Figure 9] 10A and 10B are diagrams showing an example of performer gaze expression processing for a 2D performer video according to this embodiment. [Figure 10] 1 is a diagram illustrating the matching between a concert venue on the audience side and a studio on the performer side according to this embodiment. FIG. [Figure 11] 10A and 10B are diagrams illustrating specific examples of performer gaze expressions when a performer selects a specific concert venue according to this embodiment. [Figure 12] 10A and 10B are diagrams illustrating other specific examples of performer gaze expressions when a performer selects a specific concert venue according to this embodiment. [Figure 13] 10A and 10B are diagrams illustrating an example of a performer's gaze expression when a performer designates a specific audience avatar according to this embodiment. [Figure 14] 10 is a flowchart showing an example of the flow of display and image capture processing in the performer information input / output system according to the present embodiment. [Figure 15] FIG. 2 is a diagram illustrating an example of the configuration of an information processing system according to a first modified example of the present embodiment. [Figure 16] 10A and 10B are diagrams illustrating an example of timing control of display ON / OFF and imaging ON / OFF according to a first modified example of the present embodiment. [Figure 17]10A and 10B are diagrams illustrating an example of presentation of a virtual 2D image in a concert hall according to a first modified example of the present embodiment. [Figure 18] FIG. 10 is a diagram illustrating an example of the configuration of an information processing system according to a second modified example of the present embodiment. [Figure 19] FIG. 10 is a diagram showing an example of timing control of display ON / OFF, imaging ON / OFF, and illumination ON / OFF according to a second modified example of the present embodiment. [Figure 20] 10A to 10C are diagrams illustrating a virtual 2D image lighting effect reflection process and a performer image lighting effect reflection process according to a second modified example of the present embodiment. [Figure 21] 1 is a block diagram illustrating an example of the hardware configuration of an information processing device that realizes an audience information output system, a performer information input / output system, or a performer video display system according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0010] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0011] The explanation will be given in the following order: 1. Overview of information processing system according to one embodiment of the present disclosure 2.Configuration example 2-1. Audience information output system 1 2-2. Performer Information Input / Output System 2 2-3. Performer video display system 3 3. Operation processing 4. Variations 4-1. First modified example 4-2. Second modified example 5. Hardware Configuration 6. Supplementary Information
[0012] <<1. Overview of Information Processing System According to One Embodiment of the Present Disclosure>> 1 is a diagram illustrating an overview of an information processing system according to an embodiment of the present disclosure. As shown in FIG. 1, the information processing system according to this embodiment includes an audience information output system 1, a performer information input / output system 2, and a performer video display system 3.
[0013] In this embodiment, we will explain an example of a live concert (also called a remote live concert) in which video of performers filmed in a studio is provided to a remote audience in real time. A performer is a person who performs. A performer is also an example of a subject. A remote location means a location different from where the performer is located. Video of the performers filmed in the studio is acquired by a performer information input / output system 2, transmitted to a performer video display system 3 via a network 42, and presented to the audience by the performer video display system 3.
[0014] Examples of spectators include spectators at a concert venue that can accommodate a large number of people, such as a stadium, arena, or hall (first spectator example), spectators watching performer videos streamed to their own display devices (television devices, PCs (personal computers), smartphones, tablet devices, projectors (projection devices), etc.) using a telecommunications system (second spectator example), and spectators participating as avatars in a live performance held in a virtual space (third spectator example). Note that virtual space includes VR (Virtual Reality) space. The spectator examples described above are merely examples, and this embodiment is not limited to the first to third spectator examples.
[0015] In this embodiment, audience video is acquired by audience information output system 1, transmitted to performer information input / output system 2 via network 41, and provided to the performers by performer information input / output system 2. This allows the performers to perform a remote live show while watching the audience's situation.
[0016] Here, in the performer information input / output system 2, a 3D model of the subject is generated based on dozens of captured images obtained by simultaneously capturing images from various directions using dozens of cameras arranged to surround the subject, and performer information is acquired using a technology (for example, Volumetric Capture technology) that generates high-quality 3D images of the subject viewed from any direction. body Since a 3D image of the subject is generated from the 3D model, it is possible to generate an image from a viewpoint (virtual viewpoint) where there is no camera, allowing for more flexible viewpoint manipulation by the broadcaster and the audience. In this embodiment, a performer is used as an example of the subject, but the present disclosure is not limited to this, and the subject is not limited to humans. The subject includes a wide range of objects to be captured, such as animals, insects, automobiles, airplanes, robots, and plants. The performer information input / output system 2 according to this embodiment transmits 3D images generated from the 3D model of the performer to the performer image display system 3 as the performer's image.
[0017] (Identifying issues) When generating a 3D model of a subject from a captured image, it is necessary to extract the subject area (silhouette image of the subject) from the captured image. extraction However, it is difficult to present an image other than a green screen or blue screen around the subject.
[0018] For example, when performing a remote live performance like the one mentioned above, if the audience's video could be displayed to the performers in real time, the performers could perform interactive actions while watching the audience's situation, which would improve the entertainment value. It would also be desirable to provide a more realistic live feeling to the performers who are recording in a studio to generate 3D models.
[0019] Therefore, in an embodiment of the present disclosure, the timing control unit 24 of the performer information input / output system 2 controls the timing of capturing images of the performer from around the performer and the timing of displaying the audience image around the performer (time-sharing control at a high rate), thereby making it possible to capture images for generating a 3D model of the performer and for the performer to view the audience image at the same time.
[0020] 2 and 3 are diagrams illustrating the arrangement of display areas (display regions) 233 (e.g., displays or screens) and imaging units 251 (cameras) in a studio that acquires information for generating a 3D model of a performer according to this embodiment. As shown in FIG. 2, for example, m cameras are arranged as imaging units 251 around performer A (e.g., in a circular shape), and displays (e.g., LED displays) are arranged as display areas 233A to fill the gaps. In another example, as shown in FIG. 3, a projector screen (display area 233B) may be arranged in a location that would previously have been a green screen, and a short-focus rear projector (projector 234) may be installed behind it. Note that the screen is assumed to be, for example, a colored screen (e.g., a green screen). In the example shown in FIGS. 2 and 3, n display areas 233 and m imaging units 251 are arranged in a circular shape, but they may also be rectangular or have other shapes, and the imaging units 251 and display areas 233 may be arranged in different shapes. Furthermore, the imaging units 251 and display areas 233 are not limited to being arranged in a single row around performer A, but may be arranged in multiple rows in the vertical direction.
[0021] In this way, by placing the display area 233 and the imaging unit 251 around performer A and staggering the timing of imaging and the timing of displaying the audience video, it becomes possible to simultaneously capture images for generating a 3D model and display the audience video. That is, the timing control unit 24 of the performer information input / output system 2 controls the timing so that the display is turned off when imaging is taking place and the imaging is turned off when displaying is taking place. As a result, when the display is turned off, the LEDs of the LED display are turned off and the background becomes black, or the original color of the screen (e.g., green) becomes the background, making it possible to obtain an image that makes it easy to extract the performer's area.
[0022] As explained above, when performing a remote live performance, free viewpoint 3D images of the performers can be generated and provided to the audience, and audience images can be presented to the performers, who can then take interactive actions on the audience images while observing the audience's situation, thereby further improving the entertainment value.
[0023] This system is not limited to live concerts, but can also be widely applied to games, telecommunications, and other situations where interactive actions are performed via video. While this system does not address audio, it can be processed separately in practice, with the performer's audio transmitted to the audience and vice versa. For example, the performer's audio can be encoded along with the performer's video and sent to the performer video display system 3, which can then output the audio along with the performer's video from the performer video display system 3 (which also has an audio output function).
[0024] The information processing system according to an embodiment of the present disclosure has been outlined above. Next, the specific configuration of each device included in the information processing system according to the present embodiment will be described with reference to the drawings.
[0025] <<2. Configuration Example>> <2-1. Audience Information Output System 1> As shown in Fig. 1, audience information output system 1 has audience information acquisition unit 10 and transmission unit 20. The audience information output system 1 may be composed of multiple information processing devices, or may be a single information processing device. It is envisioned that audience information output system 1 will be applied to a device (or a system composed of multiple devices) that processes the acquisition of audience footage at each concert venue, or to a display terminal (information processing device) used by each audience member.
[0026] (Audience information acquisition unit 10) The audience information acquisition unit 10 acquires images of the audience (live-action images), and if the audience is an avatar, movement information of each part of the avatar, and audience attribute information (for example, the audience's shooting conditions (camera information, etc.), gender, age, region, venue information, fan club membership information, enthusiasm level analyzed online, etc.).
[0027] Audience example 1 (concert venue) If the audience is at a concert venue that can accommodate a large number of people, such as a stadium, arena, or hall, the audience information acquisition unit 10 captures a wide range of audience seats and generates wide-field-of-view video as the audience video. Specifically, for example, the audience information acquisition unit 10 may generate wide-field-of-view video by stitching (joining) video captured by multiple monocular cameras (multiple video data captured of different areas), or may use equipment dedicated to wide-field shooting, such as a 360-degree all-around camera. Furthermore, the audience information acquisition unit 10 may process the wide-field-of-view video into a format compatible with various formats (e.g., equirectangular format or cube map format) (data format conversion process) before outputting the wide-field-of-view video.
[0028] Audience example 2 (viewing on each individual display device) In the case where the audience is watching at home on their own display terminals using a telecommunications system, the audience information acquisition unit 10 photographs the audience with a monocular camera mounted on a PC or smartphone and outputs the photograph as the audience video.
[0029] Audience example 3 (avatar in virtual space) In the case where spectators participate as avatars in a live performance (where 3D images of performers are displayed) held in a virtual space generated by 3DCG or the like, the spectator information acquisition unit 10 acquires movement information of the spectator's avatar (3DCG character). The movement information is information indicating the movement of each part of the avatar (information for moving it). As a display device for viewing the video in the virtual space, for example, a non-transparent HMD (Head Mounted Display) that covers the entire field of view is assumed. The spectator information acquisition unit 10 can predict the movement of corresponding parts based on signals acquired from various sensors (such as an audio pickup unit, an RGB camera, an eye tracking sensor, and an IMU (Inertial Measurement Unit) sensor) provided in the HMD, and output the predicted movement information (motion capture data) of the avatar. Techniques such as machine learning may be used to predict the movement of corresponding parts. Specifically, the audience information acquisition unit 10 generates (avatar) mouth movements based on speech audio signals acquired from an audio recording unit, generates (avatar) facial expressions based on image signals acquired from an RGB camera, generates (avatar) pupil movements based on near-infrared LED signals acquired from an eye tracking sensor, and generates (avatar) eye movements based on acceleration sensor and gyro sensor signals acquired by an IMU sensor. of) Translational and rotational movements of the head are generated. Note that the display device and various sensors for viewing the video of the virtual space described above are merely examples, and the present embodiment is not limited to these. Various sensors may be attached to the spectators' hands and feet, or various sensors may be installed around the spectators. Furthermore, the spectators may use remote controllers to control the movements of their own avatars.
[0030] (Transmitter 20) The transmitting unit 20 transmits the audience video or the movement information of the audience avatars together with the audience attribute information to the performer information input / output system 2 via the network 41.
[0031] Specifically, the transmitter 20 may function as both an encoder and a multiplexer. For example, the encoder encodes the audience video or audience avatar motion information and the audience attribute information. The multiplexer then multiplexes the encoded streams (the audience video encoded stream or avatar motion stream, and the audience attribute information stream) and transmits the data to the performer information input / output system 2.
[0032] Video compression processing (e.g., AVC (H.264), HEVC (H.265), etc.) may be applied to the encoding of audience video. Also, specialized encoding for the avatar rig configuration (such as bones) of the avatar (3DCG, etc.) may be applied to the avatar movement information. Also, dedicated encoding processing may be applied to the encoding of audience attribute information.
[0033] <2-2. Performer Information Input / Output System 2> As shown in FIG. 1, performer information input / output system 2 includes receiving unit 21, distributed display data generating unit 22, display processing unit 23, timing control unit 24, video acquisition unit 25, performer information generating unit 26, and transmitting unit 27. Performer information input / output system 2 may be composed of multiple information processing devices, or may be a single information processing device. Furthermore, distributed display data generating unit 22, display processing unit 23, timing control unit 24, video acquisition unit 25, and performer information generating unit 26 are listed as functions of the control unit of performer information input / output system 2. Furthermore, receiving unit 21 and transmitting unit 27 are listed as functions of the communication unit of performer information input / output system 2.
[0034] The display processing unit 23 can also perform display processing on a display area realized by a display device (a display or projector). The video acquisition unit 25 also includes the acquisition of video signals by a camera. In the following description of the configuration of the performer information input / output system 2, reference will also be made to the block diagrams shown in Figures 4, 7, and 8 as appropriate. Figure 4 is a block diagram mainly showing a specific example of the configuration of the display processing unit 23 of the performer information input / output system 2 according to this embodiment. Figure 7 is a block diagram mainly showing a specific example of the configuration of the video acquisition unit 25 of the performer information input / output system 2 according to this embodiment. Figure 8 is a block diagram mainly showing a specific example of the configuration of the performer information generation unit 26 of the performer information input / output system 2 according to this embodiment.
[0035] (2-2-1. Receiving unit 21) The receiving unit 21 receives audience video (or avatar movement information) and audience attribute information from the audience information output system 1, and outputs them to the distributed display data generating unit 22.
[0036] Specifically, for example, the receiving unit 21 functions as a demultiplexing unit and a decoding unit. The demultiplexing unit separates the data received from the audience information output system 1 into an audience video encoded stream or avatar motion stream, and an audience attribute information stream, and outputs them to the decoding unit. The decoding unit then performs decoding processing using the corresponding decoders. Specifically, the decoding unit decodes the input audience video encoded stream and outputs it as audience video information. Alternatively, the decoding unit decodes the input avatar motion stream and outputs it as avatar movement information. The decoding unit also decodes the audience attribute information stream and outputs it as audience attribute information.
[0037] (2-2-2. Distribution display data generation unit 22) Based on the audience video information input from the receiving unit 21, the distributed display data generating unit 22 generates distributed display data (video signals) to be distributed and displayed in multiple display areas 233 (see FIGS. 2 and 3) arranged around the performers, and outputs the data to the display processing unit 23. Here, when audience videos are sent from multiple concert venues, the distributed display data generating unit 22 may output audience videos from the concert venue selected by the performer. Furthermore, when audience videos of each individual are sent via a telecommunications system, the distributed display data generating unit 22 may output audience videos of attributes selected by the performer (e.g., age group, gender, specific membership number, etc.). Furthermore, when movement information of avatars of audience members participating in a live performance taking place in a virtual space is sent, the distributed display data generating unit 22 may control the movement of each avatar according to the movement information. Furthermore, the distributed display data generating unit 22 may generate video (a field of view including the audience avatars) from the performer's viewpoint in the virtual space (e.g., on the stage in the virtual space) and output it as the audience video.
[0038] A detailed description will be given below with reference to FIG. 4. As shown in FIG. 4, the distributed display data generation unit 22 receives as input audience video information 510 and audience attribute information 520 decoded by the receiving unit 21, as well as pre-generated studio attribute information 530 (e.g., the type, size, and number of display devices, the relative positional relationship between the performer and the display device, ambient brightness, etc.). The performer interaction information 540 is generated based on the performer's operations and gestures, and includes information on the venue selected by the performer and audience attributes. The performer interaction information 540 is generated in the performer information input / output system 2 by, for example, analyzing the performer's speech, analyzing gestures (e.g., pointing) from captured images, operating buttons by the performer (e.g., a switch on the microphone held by the performer), or operations by staff on the distributor's side, and is input to the distributed display data generation unit 22. By providing a switch on the microphone held by the performer, the performer can perform operations seamlessly even during a live performance. In addition, a part of the dance choreography performed by the performer may be recognized using an image or a sensor and reflected in the performer interaction information.
[0039] The distributed display data generation unit 22 takes into consideration the performer interaction information and determines the display format (data selection, position, size, orientation, etc.) for the downstream display processing unit 23. Furthermore, the distributed display data generation unit 22 processes audience video and the like according to the determined display format, and outputs the processed video signal to the display processing unit 23 as distributed display data (data to be distributed and displayed in multiple display areas 233). Below, the function of the distributed display data generation unit 22 will be specifically explained for each of the first to third audience examples.
[0040] First audience case (concert venue) When the audience is at a concert venue that can accommodate a large number of people, such as a stadium, arena, or hall, the distributed display data generation unit 22 can function as both an audience venue selection unit and a data generation unit. In the first audience example, live streaming to multiple different concert venues is assumed. In this case, a performer performing a live performance in a studio can also communicate with a specific concert venue (for example, by shouting or speaking to a specific concert venue). When the performer selects a specific concert venue, the audience venue selection unit selects that concert venue, and the data generation unit appropriately processes the audience video at that concert venue. The processed data (video signal) is then output to the display processing unit 23, which displays it in the display area 233.
[0041] More specifically, the audience venue selection unit selects audience video information and accompanying audience attribute information for the concert venue selected by the performer from among multiple different concert venues based on performer interaction information (including identification information of the venue selected by the performer), and outputs this to the subsequent data generation unit.
[0042] The data generation unit processes the selected audience image information so that the image of the audience from the performer's viewpoint is displayed at life size, taking into consideration the display conditions for the performer indicated in the studio attribute information 530 (e.g., type, size, number of display areas, relative positional relationship between the performer and the display area, ambient brightness, etc.) and the audience shooting conditions indicated in the selected audience attribute information (e.g., camera installation position, FOV (Field of View), etc.), and outputs the processed image information as distributed display data.
[0043] When no specific concert venue is selected, the audience venue selection unit may randomly and periodically select one or more concert venues, or may select all concert venues, thereby causing audience videos from one or more concert venues to be randomly and periodically switched and displayed in display area 233, or audience videos from all concert venues to be displayed in display area 233.
[0044] Second example of audience (viewing on their own display device) In the case of audience members viewing at home on their own display terminals using a telecommunications system, the distributed display data generation unit 22 can function as both an audience grouping analysis and selection unit and a data generation unit. In the second audience example, a live broadcast to audience members at home using a telecommunications system is assumed. In this case, a performer performing a live performance in a studio can communicate (call out or speak to specific audience groups) with a specific audience group (e.g., a group of women, a group of children, a group of adults, a group of residents in a specific area, a group of excited fans, a group of people wearing glasses, etc.). The audience grouping analysis and selection unit selects audience video of audience members belonging to the audience group specified (selected) by the performer, and the selected audience video is processed appropriately by the data generation unit. The processed data (video signal) is then output to the display processing unit 23, which displays it in the display area 233.
[0045] More specifically, the audience grouping analysis and selection unit selects the audience video information and accompanying audience attribute information of the audience group designated (selected) by the performer from among the grouped audience groups based on the performer interaction information (including the identification information of the audience group designated by the performer), and outputs this to the downstream data generation unit. The audience grouping analysis and selection unit may perform grouping based on pre-registered audience information (which may also be included in the audience attribute information), or may perform grouping based on information obtained by analyzing individual audience videos (such as age, gender, and facial expression obtained by face recognition technology, or excitement level obtained by analyzing head movement). Information that changes over time (for example portion In the case of grouping based on performance level, facial expression, etc., grouping may be performed at any time, or may be performed when performer interaction information is input.
[0046] The data generation unit takes into consideration the display conditions for the performers indicated in the studio attribute information 530 (e.g., type, size, number of display areas, relative positional relationship between the performers and display areas, ambient brightness, etc.) and the shooting conditions for the audience indicated in the selected audience attribute information (e.g., camera installation position, FOV (Field of View), etc.), and processes the selected audience video information so that it is displayed in a tiled manner at a size that allows each audience member's face to be visible, and outputs it as distributed display data.
[0047] When no specific audience group is designated (selected), the audience grouping analysis and selection unit may randomly and periodically select one or more audience groups, or may select all audience members. This allows audience videos of one or more audience groups to be randomly and periodically switched and displayed in the display area 233, or audience videos of all audience members to be displayed in the display area 233.
[0048] Third example of audience (avatars in virtual space) In the case where the audience is participating as avatars in a live performance (where 3D images of the performers are displayed) held in a virtual space generated using 3DCG or the like, the distributed display data generation unit 22 can function as both a performer viewpoint movement unit and a data generation unit. In a third audience example, a live concert can be held in which 3D images (volumetric images) of the performers are displayed in real time in the virtual space. The audience, for example, wears an HMD that obscures their field of vision and watches the video of the live concert taking place in the virtual space (from the audience's viewpoint in the virtual space (e.g., the viewpoint of the audience's avatar or a viewpoint from which the audience's avatar is visible)). In addition, in the studio where the performer is performing, a video from the performer's viewpoint in the virtual space (e.g., a view of the audience seats where the avatar is located as seen from the stage in the virtual space) is displayed in the display area 233 around the performer, allowing the performer to perform the live performance while observing the situation of the audience. In this case, the performer can approach a specific avatar and communicate with it. The performer viewpoint movement unit identifies the avatar designated (selected) by the performer, and the data generation unit renders the image of the identified avatar, thereby generating an audience image in which the performer's viewpoint appears to be closer to the avatar in the virtual space. The generated data (image signal) is then output to the display processing unit 23, which displays it in the display area 233.
[0049] More specifically, the performer viewpoint movement unit identifies the avatar designated (selected) by the performer based on the performer interaction information (including identification information of the avatar the performer wants to approach), selects information about that avatar (information for displaying the avatar, such as movement information and 3DCG) and accompanying audience attribute information, and outputs this to the subsequent data generation unit.
[0050] The data generation unit takes into consideration the display conditions for the performers indicated in the studio attribute information 530 (e.g., type, size, number of display areas, relative positional relationship between the performers and display areas, ambient brightness, etc.) and the rendering conditions for the specified avatar indicated in the selected audience attribute information (e.g., position, orientation, size of the avatar in the virtual space, texture material information, lighting, etc.), generates an image that makes the specified avatar appear to be close to the performers in the virtual space, and outputs it as distributed display data.
[0051] (2-2-3. Display processing unit 23) The display processing unit 23 separates the distribution display data (video signal) output from the distribution display data generating unit 22 and performs processing to display the data in a plurality of display areas 233. This will be specifically described below with reference to FIG.
[0052] 4, display processing unit 23 has a video signal separation unit 231, multiple video processing units 232, and multiple display areas 233. Video signal separation unit 231 separates the distributed display data (video signal) output from distributed display data generation unit 22 for each display area, and outputs the separated data to multiple video processing units 232 that each control display for each display area. Each video processing unit 232 appropriately corrects the received data (separated data) and then controls display in the corresponding display area 233.
[0053] Multiple display areas 2 3As described above, in the first audience example, the audience video constructed by 3 is a video of an audience at a concert hall that accommodates a large number of people. In the second audience example, it may be a video in which, for example, a video chat screen of a telecommunications system (video of the audience captured by a PC camera) is arranged in a tiled pattern. In the third audience example 3, it is a video from the performer's perspective in a virtual space. The performer's perspective video may be a field of view video (including the audience's avatars) from the position of the face (eyes) of a 3D video (a live-action 3D video generated from a 3D model of the performer; a volumetric image) placed as a performer avatar in the virtual space. The performer's perspective may also be a viewpoint slightly away from the performer avatar (3D video) (for example, behind the performer avatar) from which both the performer avatar and the audience avatars are visible.
[0054] Here, each video processing unit 232 according to this embodiment displays the audience video at a timing based on the display timing information 551 input from the timing control unit 24. The display ON timing indicated by the display timing information 551 is picture This is shifted (different) from the timing of imaging ON indicated by imaging timing information 552 output to the acquisition unit 25. For this reason, in this embodiment, it is possible to turn off the display at the timing of imaging ON, and it becomes possible for the imaging acquisition unit 25 to acquire captured images suitable for generating a 3D model of the performer. Note that, since the multiple video processing units 232 control the display timing (control the display rate) according to the same display timing information 551, the display timing in all display areas 233 (display areas 233-1 to 233-n) can be synchronized (all displays are turned ON / OFF at the same timing).
[0055] Display area 233 may be display area 233A realized by the display shown in Fig. 2, or may be display area 233B realized by the screen shown in Fig. 3. In the case of a screen, display on display area 233B may be performed by projector 234.
[0056] The video signal separator 231 and the video processors 232 may be implemented by an information processing device that is connected to and communicates with each display or projector. Alternatively, the receiver 21, the distributed display data generator 22, the video signal separator 231, the video processors 232, and the timing controller 24 may be implemented by an information processing device that is connected to and communicates with a large number of displays or projectors.
[0057] A more detailed explanation will be given below.
[0058] Video signal separator 231 As a method of separating data (video signals), for example, when an audience image is displayed by connecting multiple individual LED displays (display areas 233A-1 to 233A-n) or multiple screens for projector irradiation (display areas 233B-1 to 233B-n), the video signal separation unit 231 distributes the video signal to correspond to each display or each screen according to the arrangement of each display or each screen.
[0059] In addition, if the display or screen is not separated (if a single display or screen is used), the video signal separation unit 231 may configure the audience image to correspond to multiple display areas set within one display area on the display or screen.
[0060] Video Processing Unit 232 The video processing unit 232 can function as, for example, a brightness correction unit 2320a, a color correction unit 2320b, and a display rate control unit 2320c. Note that the correction described here is an example, and the present embodiment is not limited to this. Also, correction does not necessarily have to be performed.
[0061] For example, the video processing unit 232 appropriately performs brightness correction of the video by the brightness correction unit 2320a and color correction of the video by the color correction unit 2320b on the separated data (video signal separated according to the display area) input from the video signal separation unit 231, depending on the magnitude of the display rate specified by the display timing information 551 separately input from the timing control unit 24.
[0062] Specifically, if the display area 233 is an LED display, during periods when no image is displayed (non-display periods), the LEDs are turned off and a black screen is displayed. Because humans perceive visual information by integrating it over time, the longer the non-display period, the darker the image appears. Therefore, as shown on the left side of FIG. 5, the brightness correction unit 2320a corrects the brightness of the separated data so that the longer the non-display period, the greater the emphasis of the display's brightness correction. On the other hand, when projecting an image onto a colored screen (e.g., a green screen), if the period when no image is projected (achieved, for example, by attaching an LCD shutter to the projector) becomes long, humans perceive visual information by integrating it over time, causing the image to appear greenish. Therefore, as shown on the right side of FIG. 5, the color correction unit 2320b corrects the color of the separated data so that the longer the non-display period, the greater the emphasis of the projector's color correction. Either brightness correction or color correction may be performed, or both, depending on the type of display area 233, etc.
[0063] The actual correction strength may be adjusted by displaying a test signal at a predetermined display rate, manually adjusting the brightness and color visually, and setting the correction parameters in advance. Alternatively, the display area 233 may be photographed with a separate camera, and the brightness correction unit 2320a and color correction unit 2320b may perform automatic correction using the photographed image.
[0064] Then, the display rate control unit 2320c controls the display of the corrected video in the corresponding display area 233 so that the display rate is determined by the display timing information 551. Specifically, the display rate control unit 2320c controls the turning on and off of the LEDs in the case of an LED display, and controls the opening and closing of a liquid crystal shutter provided in the projector in the case of a projector.
[0065] In addition, if the performer's direction (viewing direction) is determined in advance (for example, if the front direction is determined), the display processing unit 23 does not need to display images in all display areas, and display areas in positions that are in the performer's blind spot may be left unlit (display OFF) to save power.
[0066] (2-2-4. Timing control unit 24) The timing control unit 24 generates display timing information 551 and outputs it to the display processing unit 23, and on the other hand, generates imaging timing information 552 and controls outputting it to the video acquisition unit 25. Specifically, the timing control unit 24 generates and outputs timing information that shifts (differentiates) the timing of display ON and imaging ON.
[0067] Fig. 6 is a diagram showing an example of timing control of display ON / OFF and imaging ON / OFF according to this embodiment. In this embodiment, as shown in Fig. 6, control is realized in which the display is turned OFF when imaging is ON, and the display is turned ON when imaging is OFF. As a result, when the display is OFF, as described above, it becomes possible to capture an image using the imaging unit (camera) that acquires information for generating a 3D model of the performer, with the subject being a black screen or green screen as the background.
[0068] More specifically, the timing control unit 24 generates display timing information (display synchronization signal) that turns on the display of the audience video when the imaging for generating the 3D model of the performer is off, and outputs this to the display processing unit 23, and on the other hand, generates imaging timing information (imaging synchronization signal) that turns off the display of the audience video when the imaging for generating the 3D model of the performer is on, and outputs this to the video acquisition unit 25.
[0069] It is desirable that the frequency that is turned on at the display timing be set to a value equal to or higher than the critical fusion frequency (approximately 30 to 40 Hz) so that flicker is not perceived. In other words, timing control unit 24 controls the display of the audience video at a display rate (high-speed rate) that at least satisfies the critical fusion frequency.
[0070] Furthermore, since a time lag may actually occur in each device such as a display or a camera, in order to provide a transition time for switching from ON to OFF (or from OFF to ON), for example, the image capturing rate control unit 2510a (see FIG. 7) of the video acquisition unit 25 may adjust the shutter speed of the camera (image capturing unit) and set the exposure time so that the period is shorter than the ON period of the image capturing timing shown in FIG. 6. Similarly, the display rate control unit 2320c also sets the lighting of the LED (or the opening time of the liquid crystal shutter of the projector) so that the period is shorter than the ON period of the display timing.
[0071] (2-2-5. Video acquisition unit 25) The video acquisition unit 25 has the function of acquiring video (captured images) for generating a 3D model of the performer. As shown in FIGS. 2 and 3 , the video acquisition unit 25 simultaneously captures images of the performer from various angles (by controlling the shutters) according to image capture timing information 552 input from the timing control unit 24 using multiple (e.g., dozens of) cameras (image capture units 251) arranged around the performer, thereby acquiring multiple captured images. The video acquisition unit 25 also integrates the multiple captured images and outputs them as multi-viewpoint data to the performer information generation unit 26, which generates a 3D model of the performer. The camera (image capture unit 251) may include various devices for sensing depth information. In this case, the multi-viewpoint data may include not only RGB signals but also depth signals and their underlying sensing signals (e.g., infrared signals).
[0072] A more detailed description will be given below with reference to Fig. 7. Fig. 7 is a block diagram mainly showing a specific example of the configuration of the video acquisition unit 25 of the performer information input / output system 2 according to this embodiment.
[0073] 7, video acquisition unit 25 includes a plurality of imaging units (cameras) 251 and multi-viewpoint data generation unit 252. For example, multi-viewpoint data generation unit 252 and performer information generation unit 26 may be realized by an information processing device that is communicatively connected to a plurality of imaging units 251 (cameras). Alternatively, timing control unit 24, multi-viewpoint data generation unit 252, performer information generation unit 26, and transmission unit 27 may be realized by an information processing device that is communicatively connected to a plurality of imaging units 251 (cameras).
[0074] As shown in FIG. 7 , each imaging unit 251 has the functions of an imaging rate control unit 2510a, an imaging signal acquisition unit 2510b, and a signal correction unit 2510c. The imaging rate control unit 2510a outputs information such as shutter speed and aperture value to the downstream imaging signal acquisition unit 2510b in accordance with the imaging rate indicated by imaging timing information 552 input from the timing control unit 24. The imaging signal acquisition unit 2510b captures an image of a subject (performer) using various camera parameters such as shutter speed and aperture value, acquires a captured image (imaging signal), and outputs the captured image to the downstream signal correction unit 2510c. The signal correction unit 2510c performs various signal correction processes such as noise reduction, resolution conversion, and dynamic range conversion, and outputs the corrected captured image to the multi-viewpoint data generation unit 252. Note that the correction contents are not limited to those described above, and it is not necessary to perform all of the corrections described here.
[0075] Furthermore, the captured image output to the multi-viewpoint data generation unit 252 may be only an RGB signal captured by an RGB camera, or may be a signal including a depth signal acquired by various depth sensors and the sensing signal that is the source of the depth signal (e.g., an infrared signal).
[0076] The multi-viewpoint data generator 252 integrates the input captured images from each viewpoint (for example, several tens of captured images), and outputs the integrated images as multi-viewpoint data 560 to the performer information generator 26.
[0077] (2-2-6. Performer information generation unit 26) Performer information generation unit 26 generates a 3D model of the performer based on multi-viewpoint data 560 input from video acquisition unit 25, generates performer video (e.g., live-action 3D video of the performer) from the 3D model, and outputs it to transmission unit 27. Performer information generation unit 26 also generates performer gaze information indicating which spectator (or spectator avatar) displayed in multiple display areas 233 the performer is looking at, from the three-dimensional position and orientation of the performer (e.g., six patterns of movement: up and down gaze movement, left and right gaze movement, head tilt movement, forward and backward body movement, left and right body movement, and up and down body movement) detected from multi-viewpoint data 560 and display area arrangement information 570 indicating the arrangement of the display areas, and outputs this information to transmission unit 27.
[0078] The performer information generating unit 26 will be described in detail with reference to Fig. 8. Fig. 8 is a block diagram mainly showing a specific example of the configuration of the performer information generating unit 26 of the performer information input / output system 2 according to this embodiment.
[0079] As shown in FIG. 8, the performer information generating unit 26 functions as a preprocessing unit 263, a performer video generating unit 261, and a performer gaze information generating unit 262.
[0080] The preprocessing unit 263 performs processes such as calibration and subject silhouette extraction (foreground / background separation), and outputs the preprocessed multi-viewpoint data to the performer video generation unit 261 and performer gaze information generation unit 262 at the subsequent stages.
[0081] The performer image generation unit 261 generates a 3D model (3D modeling data) of the performer based on the preprocessed multi-viewpoint data, and from the 3D model, it can generate 2D performer image (free viewpoint image) rendered from a certain viewpoint, or data (data including 3D modeling data and texture data) for rendering 3D performer image intended for viewing on a 3D display such as a stereoscopic hologram, 3D display, or HMD.
[0082] About Modeling The performer video generation unit 261 has the function of a modeling unit that generates a 3D model. The modeling unit generates 3D modeling data (3D model) based on the preprocessed multi-view data. 3D modeling methods may include, but are not limited to, a Shape from Silhouette method (SFS method) such as Visual Hull (visual volume intersection method) or a Multi-View Stereo method (MVS method). The data format of the 3D modeling data may be any representation format, such as Point Cloud, voxel, or mesh.
[0083] -Generation of 2D performer video (free viewpoint video) The performer video generation unit 261 also has the functionality of a 2D video generation unit and can generate 2D performer video (free viewpoint video) from a 3D model (3D modeling data). In such an example, it is assumed that the audience will view the performer video on a 2D display. For example, in the first audience example described above, it is conceivable that 2D performer video is presented on a large screen or large display at a concert venue. In the second audience example, it is also conceivable that each audience member will use a telecommunications system to view 2D performer video on a 2D display at home, etc. In the third audience example, it is conceivable that the performer video is displayed in 2D in a virtual space (for example, displayed on a virtual screen).
[0084] The 2D image generation unit generates 2D images (free viewpoint images) rendered from a certain viewpoint from the 3D modeling data and the RGB data contained in the preprocessed multi-viewpoint data, and outputs them as performer image display information. Note that the viewpoint set by the 2D image generation unit may be determined by staff on the distributor's side (such as the video production director), or may be determined by information interactively specified by the audience (information separately sent from the audience).
[0085] -Generating data for rendering 3D performer images The performer image generation unit 261 also has the functionality of a 3D image display data generation unit and can generate data for rendering 3D performer images from a 3D model (3D modeling data). In such an example, it is assumed that the audience will view performer images in 3D display, such as a 3D hologram, 3D display, or HMD. For example, in the first audience example described above, a 3D hologram of the performer could be presented as the performer image in a concert venue. In the second audience example, it is also conceivable that each audience member could use a telecommunications system to view performer images displayed in 3D on a 3D display at home, etc. In the third audience example, it is also conceivable that the performer images could be displayed in 3D in a virtual space.
[0086] The 3D image display data generation unit generates 3D texture data corresponding to the 3D modeling data from the RGB data included in the preprocessed multi-viewpoint data. Next, the 3D image display data generation unit outputs 3D image display data (volumetric data) in which the 3D texture data and the 3D modeling data are multiplexed to the transmission unit 27 as performer image display information. Note that the 3D texture data may be generated in a format that takes into account viewpoint-dependent rendering, or may include data that takes into account the texture of the subject's surface.
[0087] 8 extracts the performer's three-dimensional position and orientation (for example, six patterns of movement, such as up-and-down gaze movement, left-and-right gaze movement, head tilt movement, forward-and-back body movement, left-and-right body movement, and up-and-down body movement) from the preprocessed multi-view data and estimates the performer's gaze direction. The gaze direction may be detected using the results of detecting the performer's head orientation (in this case, the performer's head orientation may be determined by analyzing the preprocessed multi-view data, or may be detected from an IMU (inertial measurement unit) device worn by the performer).
[0088] Next, the performer gaze information generation unit 262 combines the performer's line of sight with display area arrangement information 570 indicating the arrangement of the multiple display areas to generate performer gaze information indicating which spectator (or spectator avatar) displayed in the multiple display areas 233 the performer is looking at, and outputs this to the transmission unit 27. Note that the performer gaze information generation unit 262 may also generate performer gaze information (such as information about the concert venue selected by the performer) from performer interaction information.
[0089] In this way, performer information generating unit 26 outputs performer video display information 580 (2D video or 3D video display data) and performer gaze information 590 to transmitting unit 27 as performer information.
[0090] (2-2-7. Transmitter 27) The transmitting unit 27 transmits the performer information (performer video display information 580 and performer gaze information 590) to the performer video display system 3 via the network 42. The transmitting unit 27 may encode the data in a data format appropriate for the receiving side, and then transmit the encoded data to the performer video display system 3.
[0091] For example, the transmission unit 27 functions as a performer video encoding unit, a performer gaze information encoding unit, and a data multiplexing unit. The performer video encoding unit encodes the performer video (2D video or 3D video display data) using a predetermined codec and outputs it as a performer video encoded stream. The performer gaze information encoding unit encodes the performer gaze information using a predetermined codec and outputs it as a performer gaze information encoded stream. Note that the codec for the 3D video display data may be a Point Cloud-based V-PCC codec standardized by MPEG, or another method combining mesh data encoding may be used.
[0092] The data multiplexing unit multiplexes the performer video coded stream and the performer gaze information coded stream, and the multiplexed data is output from the transmitting unit 27.
[0093] <2-3. Performer video display system 3> As shown in Figure 1, performer video display system 3 has a receiving unit 31, a display control unit 32, and a display unit 30. Performer video display system 3 may be composed of multiple information processing devices, or may be a single information processing device. Performer video display system 3 is expected to be applied to a device (or a system composed of multiple devices) that processes video display at each concert venue, or to a display terminal (information processing device) used by each audience member.
[0094] Furthermore, the display control unit 32 is considered to be a function of the control unit of the performer video display system 3. Furthermore, the receiving unit 31 is considered to be a function of the communication unit of the performer video display system 3. Furthermore, the display unit 30 is realized by a 2D display (such as a PC, smartphone, or tablet device), a 3D display (such as an HMD), or a 3D hologram presentation device.
[0095] (Receiving unit 31) The receiving unit 31 outputs the performer information received from the performer information input / output system 2 to the display control unit 32. More specifically, the receiving unit 31 separates the multiplexed data (multiplexed performer information) received from the performer information input / output system 2 into a performer video coded stream and a performer gaze information coded stream by demultiplexing. Next, the receiving unit 31 decodes the performer video coded stream and the performer gaze information coded stream using a predetermined decoder, and outputs the performer video (2D video or 3D video display data) and the performer gaze information to the display control unit 32.
[0096] (Display control unit 32) The display control unit 32 processes the 2D image or generates a 3D image as needed based on the performer image (data for displaying 2D or 3D image) and performer gaze information output from the receiving unit 31, and controls the display of the 2D or 3D performer image on the display unit 30.
[0097] The display control unit 32 according to this embodiment can provide a more realistic live performance to the audience by adding special expressions (productions) to the audience that the performer is gazing at based on the performer gaze information. This will be explained in detail below.
[0098] 2D performer video The performer video generation unit 321 refers to the performer gaze information, and if the performer is gazing at the audience (the audience watching the performer video presented by the performer video display system 3), it processes the performer video appropriately to generate a performer video that makes it clear that the audience is gazing at the performer. Note that the performer gaze information may be transmitted from the performer information input / output system 2 only to the audience the performer is gazing at.
[0099] FIG. 9 shows an example of performer gaze expression processing applied to a 2D performer video according to this embodiment. The upper part of FIG. 9 shows image 310 before gaze expression processing, and the lower part of FIG. 9 shows images 311a to 311c after gaze expression processing. For example, in image 311a, a frame is added around the image to indicate that the performer is gazing. In image 311b, the performer's face is zoomed in to indicate that the performer is gazing. In image 311c, an arrow or the like is added to emphasize that the performer is looking forward (directly at the camera), indicating that the performer is gazing. Note that performer information generation unit 26 of performer information input / output system 2 may generate performer video rendered from a viewpoint that makes the performer appear to be facing forward for the audience member the performer is gazing at.
[0100] In this way, the audience can see that the performer is gazing at the audience while performing live. Note that the processing pattern for gaze expression according to this embodiment is not limited to the example shown in FIG.
[0101] 3D performer footage In this embodiment, it is also assumed that the audience will view 3D performer images using a stereoscopic hologram, a 3D display, an HMD, etc. The performer image generation unit 321 can render 3D performer images using the 3D texture data and 3D modeling data included in the decoded 3D image display data.
[0102] Assuming a first audience example in which live streaming is being performed for multiple different concert venues, the performer gaze information would be information indicating the concert venue the performer is gazing at (the venue selected by the performer to engage in specific communication). In the first audience example, as shown in FIG. 10, the display may be controlled so that the audience's view of the concert venue and the performer's view of the studio are consistent with each other (e.g., the relative positions and size of the performers and audience). In the example shown in FIG. 10, a 3D performer video (stereoscopic hologram) 312 is displayed on the stage at the concert venue, and audience groups B1 to B3 are positioned on three sides around the stage. Each of the audience groups B1 to B3 is photographed by a monocular camera, and a wide-field audience video stitched together from the three audience videos is transmitted to the performer information input / output system 2. To correspond to the positional relationship between the performers and the audience at the concert venue, performer information input / output system 2 distributes wide-field audience images to display areas 233-1 to 233-3 located on three sides of performer A in the studio, as shown on the right in Figure 10, and displays audience images of audience groups B1 to B3, respectively. This ensures that the way both sides see the same thing.
[0103] When such display control is performed, a specific example of a performer's gaze expression when a specific concert venue is selected by the performer will be described with reference to Fig. 11. In the example shown in Fig. 11, it is assumed that the performer selects venue D (concert venue D) by, for example, saying "venue D!", operating a switch attached to the microphone he is holding, or pointing at the display area where venue D is displayed.
[0104] In this case, as shown in the upper part of FIG. 11, in the performer's studio, the distributed display data generation unit 22 of the performer information input / output system 2 displays the information of the audience group B1 of the concert venue D on the display areas 233-1 to 233-3 based on the performer interaction information (information generated from the performer's speech, switch operation, pointing, etc.). D ~B3 D The relative positions of the performers and audience in the studio are controlled to match those in concert venue D.
[0105] Meanwhile, as shown in the lower part of FIG. 11 , at multiple different concert venues (e.g., concert venue C, concert venue D), 3D performer video 312 is displayed on the center stage of the concert venue using, for example, a stereoscopic hologram. At each venue, audience members are positioned on three sides surrounding the center stage. If concert venue D is selected (if the gaze information indicates it as the concert venue to be gazed at), a circular effect image (which may be a 3D image) is displayed at the feet of 3D performer video 312 on the center stage at concert venue D, as shown on the right side of the lower part of FIG. 11 . This makes it possible for the audience at concert venue D to clearly see that they are being gazed at by the performer. Note that the method of expressing performer gaze is not limited to the example shown in FIG. 11 ; images of other shapes may be displayed at the feet of 3D performer video 312, or dramatic 3DCG may be displayed around the performer. Furthermore, at concert venue D, gaze expression may be achieved using effects other than images, such as flashing lights, fireworks, confetti, and sound effects.
[0106] While Fig. 11 describes performer gaze expressions when 3D performer images are presented at a concert venue using 3D holograms, this embodiment is not limited to this, and various performer gaze expressions can also be made when 3D performer images are presented using a large-screen display (or screen) such as that shown in Fig. 12. In the example shown in Fig. 12, as shown on the right side of the lower half of Fig. 12, the fact that concert venue D has been selected can be expressed by displaying a frame image, lighting up the area around the performer, or displaying a dramatic image around the performer on large-screen display (or screen) 30D of concert venue D that the performer is gazing at.
[0107] Audience members at a concert venue who are selected by a performer can intuitively understand visually (or audibly) that they have been selected by the performer through the various performer-gazing expressions described above, and can feel an interaction with the performer, providing an experience that is close to that of an actual live performance.
[0108] The performer gaze expression in the case of the third audience example will also be described with reference to FIG. 13. FIG. 13 is a diagram illustrating an example of a performer gaze expression when a performer designates a specific audience avatar according to this embodiment. The third audience example is a case in which audience members participate in a live concert held in a virtual space as audience avatars (audience avatars). In the example shown in FIG. 13, it is assumed that a performer designates audience avatar T when a performer avatar 313 (a 3D image of the performer) performing a live performance and audience avatars participating in the live concert are arranged in the virtual space. In this case, as shown on the right side of FIG. 14, the performer avatar 313 approaches audience avatar T in the virtual space. In this state, a rendering image from the audience avatar's viewpoint is generated and displayed on the display device (e.g., HMD) of each audience member corresponding to audience avatar T. This allows audience members to experience a simulated live performance in which performers approach the audience and perform.
[0109] The above describes in detail each component of the information processing system according to this embodiment. In this embodiment, a 3D model is generated from captured images of a performer captured in a studio, and 2D or 3D video of the performer from any viewpoint is live-streamed from the 3D model to a remote audience. In this case, displaying the audience video on the background (surroundings) of the performer, which was previously a green screen, and capturing the performer for generating the 3D model can both be achieved by performing time-sharing control at a high rate to offset the timing. Furthermore, by communicating the visibility of the audience video of the performer to the remote audience, it is possible to provide an experience closer to an actual live performance, where the performer can feel like they are interacting with the performer.
[0110] <<3. Operation Processing>> FIG. 14 is a flowchart showing an example of the flow of the display and image capture process in the performer information input / output system 2 according to this embodiment.
[0111] As shown in FIG. 14, first, the receiving unit 21 of the performer information input / output system 2 receives the audience video from the audience information output system 1 (step S103).
[0112] Next, the distribution display data generation unit 22 selects an audience venue / audience group / audience avatar based on the performer interaction information (step S106), and generates distribution display data based on the video of the selected audience venue / audience group or the movement information of the audience avatar (step S109).
[0113] Next, the display processing unit 23 performs control to simultaneously display the distributed display data in multiple display areas 233 arranged around the performer in accordance with the display timing information input from the timing control unit 24 (step S112). Note that the display is performed when imaging is turned off.
[0114] Meanwhile, the video acquisition unit 25 controls the multiple imaging units 251 arranged around the performer to simultaneously capture images in accordance with the imaging timing information input from the timing control unit 24 (step S115). Note that the imaging is performed when the display is turned off. This allows the video acquisition unit 25 to obtain an image that makes it easy to extract the subject's silhouette.
[0115] Next, the video acquisition unit 25 extracts silhouette images of the performers from each of the plurality of captured images to acquire multi-viewpoint data (step S118).
[0116] Next, the performer information generation unit 26 generates a 3D model of the performer based on the multi-view data, generates 2D or 3D performer video from the 3D model (step S121), and also generates performer gaze information based on the multi-view data (step S124).
[0117] Then, the transmitting unit 27 transmits the performer video and the performer gaze information to the audience side (performer video display system 3) (step S127).
[0118] The above is a description of the operation processing according to this embodiment. Note that the flow of the operation processing shown in Fig. 14 is an example, and the present disclosure is not limited to this.
[0119] <<4. Modifications>> Next, a modified example of the information processing system according to this embodiment will be described with reference to FIGS.
[0120] <4-1. First modified example> In the first modified example, a function is added to generate a virtual 2D image in which an audience is positioned around the performers. The virtual 2D image is obtained by timing the display of the audience image in each display area 233 in the studio and the capture of the performers by each imaging unit 251 to occur simultaneously. In other words, when the performers are captured by each imaging unit 251, the audience image is displayed in each display area 233 arranged around the performers (including the background), thereby obtaining a captured image (multi-view data for virtual 2D image) in which the audience image is reflected around the performers.
[0121] Fig. 15 is a diagram showing an example of the configuration of an information processing system according to the first modified example. The system shown in Fig. 15 is configured by adding a virtual 2D video generation unit 280, a transmission unit 281, a network 43, and a virtual 2D video display system 4 to the system shown in Fig. 1.
[0122] (Timing control) The timing control unit 24a according to the first variant generates timing information including control to shift (make different) the timing of display ON and imaging ON, and control to synchronize (make the same) the timing of display ON and imaging ON, and outputs the information to the display processing unit 23 and the video acquisition unit 25.
[0123] FIG. 16 is a diagram illustrating an example of timing control of display ON / OFF and imaging ON / OFF according to the first modified example. As shown in FIG. 16, timing is controlled so that the period of the display timing is doubled relative to the imaging timing. That is, the timing control unit 24a generates timing information that realizes the control of turning the display OFF when imaging is ON, the control of turning the display ON when imaging is ON, and the control of turning the display ON when imaging is OFF, as shown in FIG. 16. This allows captured images for generating a 3D model of the performer to be acquired when the display is OFF, and captured images for virtual 2D images to be acquired when the display of the audience video is ON. Furthermore, a timing for turning off imaging when the display of the audience video is ON is also provided. In this modified example, the display timing information is also input to the video acquisition unit 25 from the timing control unit 24a. When generating multi-viewpoint data, the video acquisition unit 25 refers to the display timing information and can acquire captured images captured when the display was ON as captured images for virtual 2D images.
[0124] In the example shown in FIG. 16, the display-on period is longer than the imaging-on period, but it is desirable to shorten the imaging timing period of the camera and capture images at a high rate so that the display-on period satisfies the condition of being equal to or higher than the critical fusion frequency (approximately 30 to 40 Hz).
[0125] (Generation of virtual 2D images) The virtual 2D video generation unit 280 acquires multi-viewpoint data for virtual 2D video that integrates multiple captured images acquired when imaging is ON and the audience video display is ON from the video acquisition unit 25. On the other hand, multiple captured images acquired when imaging is ON and the audience video display is OFF are output from the video acquisition unit 25 to the performer information generation unit 26 as multi-viewpoint data for generating a 3D model, as in the above-described embodiment.
[0126] The virtual 2D video generation unit 280 selects a 2D video (captured image) from a certain viewpoint from the multi-viewpoint data for virtual 2D video. The selection may be performed by a distributor's staff member (such as a video production director) taking into consideration the performer's position and the appearance of the audience image, or the selection may be performed automatically using image analysis technology. The virtual 2D video generation unit 280 also processes the selected captured image to reflect the director's intentions. Possible processes include trimming (cropping) and scaling. The virtual 2D video generation unit 280 outputs the processed video signal as a virtual 2D video to the transmission unit 281. The transmission unit 281 encodes the virtual 2D video using a predetermined codec, for example, and transmits the encoded virtual 2D video data to the virtual 2D video display system 4 via the network 43.
[0127] (Presenting virtual 2D images) As shown in FIG. 15, virtual 2D video display system 4 has a receiving unit 401, a display control unit 402, and a display unit 403. Receiving unit 401 decodes the virtual 2D video encoded data using a predetermined decoder and outputs the virtual 2D video to display control unit 402. Display control unit 402 displays the virtual 2D video on a large screen 431 (an example of display unit 403) in a concert venue, for example, as shown in FIG. 17. Display control unit 402 displays the virtual 2D video toward the audience in the concert venue. Note that at the concert venue, 3D performer video (stereoscopic hologram) 312 may be separately displayed on the stage by performer video display system 3.
[0128] While viewing the 3D performer video (stereoscopic hologram) 312, the audience can also see the performers performing while watching the audience video on a large screen 431 in a third-person perspective. Because the audience knows that the performers are performing for the audience, they can experience a remote live performance with a greater sense of unity.
[0129] 17 are merely examples, and the present embodiment is not limited to these. The display unit 403 may be, for example, a large display.
[0130] In addition, in this modification, it is assumed that the virtual 2D image is generated by shooting the performers together with the audience image displayed in the display area 233, but it is not limited to the audience image. For example, the virtual 2D image may be generated by shooting the performers together with CG image (staged image) that changes with music. Furthermore, the generated virtual 2D image can also be used as a live video at a later date.
[0131] <4-2. Second modified example> In the second modified example, a function for reflecting lighting effects on performer images, etc. is added. FIG. 18 is a diagram showing an example configuration of an information processing system according to the second modified example. The system shown in FIG. 18 is configured by adding a lighting device 29, a virtual 2D image lighting effect reflecting unit 290, and a performer image lighting effect reflecting unit 291 to the system shown in FIG. 15. The lighting device 29 is used for staging, and one or more lighting devices 29 are installed in the studio. There are no particular restrictions on the installation location of the lighting device 29.
[0132] (Lighting timing control) The timing control unit 24b according to the second modified example generates timing information including control to shift (make different) the timing of display ON and imaging ON, control to synchronize (make the same) the timing of display ON and imaging ON, and further control to synchronize (make the same) the timing of imaging ON and lighting ON, and outputs the information to the display processing unit 23, the video acquisition unit 25, and the lighting device 29.
[0133] Fig. 19 is a diagram showing an example of timing control of display ON / OFF, imaging ON / OFF, and lighting ON / OFF according to the second modified example. The timing control unit 24b generates timing information for realizing control of turning off the display and lighting when imaging is ON, control of turning on the display and turning off the lighting when imaging is ON, and control of turning off the display and turning on the lighting when imaging is ON, as shown in Fig. 19.
[0134] This allows captured images for generating a 3D model to be acquired when the display and lighting are off, captured images for virtual 2D images to be acquired when the display of the audience video is on and the lighting is off, and further allows captured images for lighting effects to be acquired when the display of the audience video is off and the lighting is on.
[0135] 19, the display OFF period is as long as two repeated ON / OFF cycles of imaging, and the display ON period is as long as one ON / OFF cycle of imaging. In this case as well, it is desirable to shorten the period of the camera's imaging timing and capture images at a high rate so that the display ON period satisfies the condition of being equal to or higher than the critical fusion frequency (approximately 30 to 40 Hz).
[0136] The timing control shown in FIG. 19 is an example, and the timing control section 24b is not limited to this as long as it generates timing information that generates at least the above three combinations of ON / OFF control.
[0137] (Generation of multi-view data) Video acquisition unit 25 combines multiple captured images acquired by multiple imaging units 251 to generate multi-viewpoint data. In this modification, display timing information and lighting timing information are also input to video acquisition unit 25 from timing control unit 24b, and video acquisition unit 25 refers to the display timing information and lighting timing information when generating multi-viewpoint data. Images captured when the display was turned on may be acquired as multi-viewpoint data for virtual 2D video, and images captured when the lighting was turned on may be acquired as multi-viewpoint data for lighting effects. Images captured when both the display and lighting were turned off may be acquired as multi-viewpoint data for generating a 3D model (multi-viewpoint data for performer video).
[0138] (Reflecting lighting effects) Virtual 2D image lighting effect reflection unit 290 performs alignment processing such as motion correction on the multi-viewpoint data for lighting effects based on the multi-viewpoint data for virtual 2D images and the multi-viewpoint data for lighting effects output from image acquisition unit 25, and reflects the results in the multi-viewpoint data for virtual 2D images. Virtual 2D image lighting effect reflection unit 290 outputs the multi-viewpoint data for virtual 2D images after reflection to virtual 2D image generation unit 280.
[0139] Furthermore, performer image lighting effect reflection unit 291 performs alignment processing such as motion correction on the multi-viewpoint data for lighting effects based on the multi-viewpoint data for 3D model generation and the multi-viewpoint data for lighting effects output from image acquisition unit 25, and processes the result to be reflected in the multi-viewpoint data for 3D model generation. Performer image lighting effect reflection unit 291 outputs the reflected multi-viewpoint data for 3D model generation to performer information generation unit 26.
[0140] Both the virtual 2D video lighting effect reflecting unit 290 and the performer video lighting effect reflecting unit 291 perform frame interpolation processing and lighting effect reflecting texture generation processing.
[0141] 20 is a diagram illustrating the virtual 2D video lighting effect reflection process and the performer video lighting effect reflection process according to the second modified example. For example, performer video lighting effect reflection unit 291 performs frame interpolation processing and generates a lighting effect reflection texture for each of the top two rows of data shown in FIG. 20 (performer video multi-viewpoint data and lighting effect multi-viewpoint data).
[0142] Specifically, first, in the performer video multi-viewpoint data and lighting effect multi-viewpoint data shown in FIG. 20, the frames at the time points indicated by the dotted lines are generated by interpolating using data from existing frames in the past and future. For example, frame 562-1 ab is generated from a real past frame 562a and a real future frame 562b. ab is generated from an existing past frame 562a and an existing future frame 562b. For the frame interpolation process, a technique for predicting and automatically generating intermediate frames using machine learning, for example, may be used.
[0143] Next, the performer video lighting effect reflecting unit 291 reflects the lighting effect from the frame of the multi-viewpoint data for lighting effects on the frame of the performer video multi-viewpoint data (multi-viewpoint data for generating a 3D model). The performer video lighting effect reflecting unit 291 reflects the lighting effect from the frame of the multi-viewpoint data for lighting effects (for example, frame 562-1) corresponding to the same time as the frame of the target performer video multi-viewpoint data at a certain time (for example, frame 561-L indicated by diagonal lines). ab ) or the closest real frame (e.g., frame 562a) of the multi-viewpoint data for lighting effects is used as reference data. Specifically, the performer video lighting effect reflection unit 291 searches the reference data for data similar to frame 561-L of the multi-viewpoint data for live-action video, and replaces it as the data reflecting the lighting effects. This process is a so-called template matching technique, and is generally performed for each local region. Furthermore, the replacement may involve deformation, and various geometric transformation processes such as affine transformation can be applied. Furthermore, various indices indicating the similarity of images (e.g., sum of absolute difference (SAD), sum of squared difference (SSD), normalized cross-correlation (NCC), zero-means normalized cross-correlation (ZNCC)) can be used as cost functions during the search.
[0144] On the other hand, the virtual 2D image lighting effect reflection unit 290 performs frame interpolation processing and generates a lighting effect reflection texture on the bottom two rows of data shown in Figure 20 (multi-viewpoint data for lighting effects and multi-viewpoint data for virtual 2D images) in the same manner as above.
[0145] Through the above processing, textures that reflect lighting effects such as reflections, skin gloss, and shine can be generated and reflected in performer videos and virtual 2D images, providing audiences with images that reflect the lighting effects of a real live performance.
[0146] <<5. Hardware configuration example>> Next, with reference to FIG. 21 , an example hardware configuration of an information processing device according to an embodiment of the present disclosure will be described. The processing by the above-described audience information output system 1, performer information input / output system 2, and performer video display system 3 can be realized by one or more information processing devices. FIG. 21 is a block diagram showing an example hardware configuration of an information processing device 900 that realizes the audience information output system 1, performer information input / output system 2, or performer video display system 3 according to an embodiment of the present disclosure. Note that the information processing device 900 does not necessarily need to have all of the hardware configuration shown in FIG. 21 . Furthermore, some of the hardware configuration shown in FIG. 21 may not be present in the audience information output system 1, performer information input / output system 2, or performer video display system 3.
[0147] 21 , the information processing device 900 includes a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 903, and a RAM (Random Access Memory) 905. The information processing device 900 may also include a host bus 907, a bridge 909, an external bus 911, an interface 913, an input device 915, an output device 917, a storage device 919, a drive 921, a connection port 923, and a communication device 925. Instead of or in addition to the CPU 901, the information processing device 900 may include a processing circuit such as a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), or an ASIC (Application Specific Integrated Circuit).
[0148] The CPU 901 functions as an arithmetic processing unit and control unit, and controls all or part of the operations within the information processing device 900 in accordance with various programs recorded in the ROM 903, the RAM 905, the storage device 919, or the removable recording medium 927. The ROM 903 stores programs and calculation parameters used by the CPU 901. The RAM 905 temporarily stores programs used in the execution of the CPU 901 and parameters that change as appropriate during the execution. The CPU 901, the ROM 903, and the RAM 905 are interconnected by a host bus 907 constituted by an internal bus such as a CPU bus. Furthermore, the host bus 907 is connected to an external bus 911 such as a PCI (Peripheral Component Interconnect / Interface) bus via a bridge 909.
[0149] The input device 915 is a device operated by a user, such as a button. The input device 915 may include a mouse, a keyboard, a touch panel, a switch, a lever, or the like. The input device 915 may also include a microphone that detects the user's voice. The input device 915 may be, for example, a remote control device that uses infrared or other radio waves, or an externally connected device 929 such as a mobile phone that supports operation of the information processing device 900. The input device 915 includes an input control circuit that generates an input signal based on information input by the user and outputs the signal to the CPU 901. The user operates the input device 915 to input various data to the information processing device 900 and to instruct processing operations.
[0150] The input device 915 may also include an imaging device and a sensor. The imaging device is a device that captures real space and generates a captured image using an imaging element such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor) and various components such as a lens for controlling the formation of a subject image on the imaging element. The imaging device may capture a still image or a moving image.
[0151] The sensors include various types of sensors such as a distance measurement sensor, an acceleration sensor, a gyro sensor, a geomagnetic sensor, a vibration sensor, an optical sensor, and a sound sensor. The sensors acquire information about the state of the information processing device 900 itself, such as the attitude of the housing of the information processing device 900, and information about the surrounding environment of the information processing device 900, such as the brightness and noise around the information processing device 900. The sensors may also include a GPS (Global Positioning System) sensor that receives a GPS signal and measures the latitude, longitude, and altitude of the device.
[0152] The output device 917 is configured with a device capable of visually or audibly notifying the user of acquired information. The output device 917 may be, for example, a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display, or an audio output device such as a speaker or headphones. The output device 917 may also include a PDP (Plasma Display Panel), a projector, a hologram, a printer, or the like. The output device 917 outputs the results obtained by processing by the information processing device 900 as video such as text or images, or as sound such as voice or audio. The output device 917 may also include a lighting device that brightens the surroundings.
[0153] The storage device 919 is a data storage device configured as an example of a storage unit of the information processing device 900. The storage device 919 is configured, for example, by a magnetic storage device such as a hard disk drive (HDD), a semiconductor storage device, an optical storage device, or a magneto-optical storage device. The storage device 919 stores programs and various data executed by the CPU 901, as well as various data acquired from the outside.
[0154] The drive 921 is a reader / writer for a removable recording medium 927 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, and is built into or externally attached to the information processing device 900. The drive 921 reads information recorded on the attached removable recording medium 927 and outputs the information to the RAM 905. The drive 921 also writes information to the attached removable recording medium 927.
[0155] The connection port 923 is a port for directly connecting a device to the information processing device 900. The connection port 923 may be, for example, a USB (Universal Serial Bus) port, an IEEE 1394 port, or a SCSI (Small Computer System Interface) port. The connection port 923 may also be an RS-232C port, an optical audio terminal, or an HDMI (registered trademark) (High-Definition Multimedia Interface) port. By connecting an external device 929 to the connection port 923, various types of data can be exchanged between the information processing device 900 and the external device 929.
[0156] The communication device 925 is, for example, a communication interface configured with a communication device for connecting to the network 931. The communication device 925 may be, for example, a communication card for a wired or wireless local area network (LAN), Bluetooth (registered trademark), Wi-Fi (registered trademark), or WUSB (Wireless USB). The communication device 925 may also be a router for optical communication, a router for an asymmetric digital subscriber line (ADSL), or a modem for various types of communication. The communication device 925 transmits and receives signals, for example, between the Internet and other communication devices using a predetermined protocol such as TCP / IP. The network 931 connected to the communication device 925 is a network connected by wire or wirelessly, for example, the Internet, a home LAN, infrared communication, radio wave communication, or satellite communication.
[0157] <<6. Supplementary Information>> Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the present technology is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical ideas described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0158] For example, an example of an audience member may be an audience member watching a performer's performance (such as a live concert) using AR (Augmented Reality) or MR (Mixed Reality).
[0159] Furthermore, although it has been described that the timing control unit 24 outputs imaging timing information and display timing information, the present disclosure is not limited to this. For example, the display processing unit 23 can perform display ON / OFF control at a predetermined timing and instruct the video acquisition unit 25 to perform imaging ON / OFF control at a corresponding predetermined timing. Conversely, for example, the video acquisition unit 25 can perform imaging ON / OFF control at a predetermined timing and instruct the display processing unit 23 to perform display ON / OFF control at a corresponding predetermined timing.
[0160] Furthermore, at the timing when imaging is turned on to acquire the captured image for generating a 3D model, the display is turned off, but control (display ON control) may be performed to display an image in solid green or solid blue used as a green screen or blue screen.
[0161] In addition, in the second variant, an example has been described in which the function of generating virtual 2D images shown in the first variant is further added to the system with a function of reflecting lighting effects, but the present disclosure is not limited to this, and only the function of reflecting lighting effects shown in the second variant may be added to the system described with reference to Figure 1.
[0162] It is also possible to create one or more computer programs for causing hardware such as a CPU, ROM, and RAM built into the information processing device 900 to perform the functions of the audience information output system 1, the performer information input / output system 2, or the performer video display system 3. A computer-readable storage medium storing the one or more computer programs is also provided.
[0163] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0164] The present technology can also be configured as follows. (1) a control unit that controls imaging by a plurality of imaging units to acquire three-dimensional information of a subject, and controls display of an image acquired from an external source in one or more display areas located around the subject, The control unit performs control to differentiate a timing of capturing the image from a timing of displaying the externally acquired image in the display area. (2) The information processing device described in (1), wherein the image acquired from outside is an audience image capturing an audience watching a two-dimensional or three-dimensional performer image generated based on three-dimensional information of the performer, who is the subject. (3) The information processing device described in (1), wherein the image acquired from outside is an image of a virtual space that includes in its field of view an audience avatar viewing a two-dimensional or three-dimensional performer image generated based on three-dimensional information of the performer, who is the subject, in the virtual space. (4) The control unit The information processing device described in (2) or (3) extracts the area of the performer from multiple captured images obtained simultaneously by the multiple imaging units located around the subject, generates a three-dimensional model of the performer, and generates a video of the performer from a free viewpoint from the three-dimensional model. (5) The information processing device described in any one of (2) to (4), wherein the control unit selects a specific spectator or a specific spectator avatar in accordance with the performer's instructions, and controls the display of spectator footage of the selected spectator or spectator avatar in the display area as an image obtained from the outside. (6) The information processing device according to any one of (1) to (5), wherein the control unit generates display timing information that instructs the control unit to control the image not to be displayed when the image is captured and the control unit to control the image to be displayed when the image is not captured. (7) The information processing device described in any one of (1) to (6), wherein the control unit generates imaging timing information that instructs the control unit to control so that the imaging is not performed at the timing when the image is displayed, and to control so that the imaging is performed at the timing when the image is not displayed. (8) The information processing device according to any one of (1) to (7), wherein the control unit controls the display of the externally acquired image at a display rate that satisfies at least a critical fusion frequency. (9) The information processing device described in any one of (1) to (8), wherein the control unit performs the control of making the timing different and the control of making the timing of taking the image and the timing of displaying the image obtained from the outside in the display area the same. (10) The information processing device described in (9) above, wherein the control unit controls the transmission to the audience of an image of the performer, who is the subject, including the image displayed in the display area as a background, obtained by capturing the image at the time the image is displayed. (11) The control unit a first imaging control for performing the imaging at a timing when the image is not displayed and when illumination of the subject is not performed; a second imaging control for performing the imaging at a timing when the image is not displayed and at a timing when the subject is illuminated; The information processing device according to any one of (1) to (10) above, which executes the above. (12) The processor: A device that controls imaging by a plurality of imaging units to acquire three-dimensional information of a subject, and controls display of an image acquired from an external device in one or more display areas located around the subject; performing control to make the timing of capturing the image different from the timing of displaying the externally acquired image in the display area; An information processing method, including: (13) a plurality of imaging devices arranged around a subject to acquire three-dimensional information of the subject; one or more display areas arranged around the subject; an information processing device having a control unit that controls image capture by the plurality of image capture devices and displays an image acquired from an external device in the one or more display areas; Equipped with The control unit performs control to differentiate a timing of capturing the image from a timing of displaying the externally acquired image in the display area. [Explanation of symbols]
[0165] 1. Spectator information output system 2. Performer information input / output system 21 Receiving unit 22 Distribution display data generation unit 23 Display processing section 24 Timing control section 25 Video acquisition unit 26 Performer information generation unit 27 Transmitter 3 Performer video display system 900 Information Processing Equipment
Claims
1. a control unit that controls imaging by a plurality of imaging units to acquire three-dimensional information of a subject, and controls display of an image acquired from an external device in one or more display areas located around the subject, the control unit performs control to make a timing of capturing the image different from a timing of displaying the image acquired from the outside in the display area; The externally acquired image is an audience image capturing an audience watching a two-dimensional or three-dimensional performer image generated based on three-dimensional information of the performer, who is the subject of the image. Information processing device.
2. The information processing device described in claim 1, wherein the image acquired from outside is an image of a virtual space that includes, in its field of view, an audience avatar viewing a two-dimensional or three-dimensional performer image generated in a virtual space based on three-dimensional information about the performer who is the subject.
3. The control unit The information processing device described in claim 1 extracts the area of the performer from multiple captured images obtained simultaneously by the multiple imaging units located around the subject, generates a three-dimensional model of the performer, and generates a video of the performer from a free viewpoint from the three-dimensional model.
4. The information processing device described in claim 1, wherein the control unit selects a specific spectator or a specific spectator avatar in accordance with instructions from the performer, and controls the display of an spectator image of the selected spectator or spectator avatar in the display area as an image acquired from the outside.
5. The information processing device according to claim 1 , wherein the control unit generates display timing information instructing to control not to display the image at the timing when the image is captured and to control to display the image at the timing when the image is not captured.
6. The information processing device according to claim 1 , wherein the control unit generates imaging timing information instructing control to not capture the image at a timing when the image is displayed and to capture the image at a timing when the image is not displayed.
7. The information processing device according to claim 1 , wherein the control unit controls the display of the externally acquired image at a display rate that satisfies at least a critical fusion frequency.
8. The information processing device according to claim 1 , wherein the control unit performs the control for making the timings different and the control for making the timing of capturing the image and the timing of displaying the externally acquired image in the display area the same.
9. The information processing device described in claim 8, wherein the control unit controls the transmission to the audience of an image of the performer, who is the subject, including the image displayed in the display area as a background, obtained by capturing the image at the timing when the image is displayed.
10. The control unit a first imaging control for performing the imaging at a timing when the image is not displayed and when illumination of the subject is not performed; a second imaging control for capturing the image at a timing when the image is not displayed and at a timing when the subject is illuminated; The information processing apparatus according to claim 1 , wherein the information processing apparatus executes the following:
11. The processor: A device that controls imaging by a plurality of imaging units to acquire three-dimensional information of a subject, and controls display of an image acquired from an external device in one or more display areas located around the subject; performing control to make the timing of capturing the image different from the timing of displaying the externally acquired image in the display area; Including, The externally acquired image is an audience image capturing an audience watching a two-dimensional or three-dimensional performer image generated based on three-dimensional information of the performer, who is the subject of the image. Information processing methods.
12. a plurality of imaging devices arranged around a subject to acquire three-dimensional information of the subject; one or more display areas arranged around the subject; an information processing device having a control unit that controls image capture by the plurality of image capture devices and displays an image acquired from an external device in the one or more display areas; Equipped with the control unit performs control to make a timing of capturing the image different from a timing of displaying the image acquired from the outside in the display area; The externally acquired image is an audience image capturing an audience watching a two-dimensional or three-dimensional performer image generated based on three-dimensional information of the performer, who is the subject of the image. system.
Citation Information
Patent Citations
Remote communication method in which 3-dimensional display device with shooting function is used and three-dimensional display with camera which is used in the same method
JP2006078597A
Image processing unit, image data communication system, and image processing program
JP2009187387A
Game system and program
JP2018075259A
Asynchronous camera / projector system for video segmentation
US20080084508A1
Apparatus and method for capturing image in electronic device
US20150222880A1