Information processing device and program
The information processing apparatus addresses the issue of subjects being unaware of virtual camera positions by generating a 3D model and providing virtual camera information, improving performance engagement.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2022-03-08
- Publication Date
- 2026-05-15
AI Technical Summary
Subjects performing actions like singing or dancing cannot effectively engage with a virtual camera due to a lack of awareness of its position, as conventional methods do not inform them of the virtual camera's installation location.
An information processing apparatus that acquires multiple real images from surrounding cameras, generates a 3D model of the subject, and presents virtual camera information to the subject through a display system, allowing them to be aware of the virtual camera's position and orientation.
Enables subjects to perform actions like singing or dancing while being cognizant of the virtual camera's position, enhancing their engagement and performance quality.
Smart Images

Figure 0007859444000001 
Figure 0007859444000002 
Figure 0007859444000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus and a program, and more particularly to an information processing apparatus and a program capable of informing a subject (performer) of the position of a virtual camera that is observing the subject.
Background Art
[0002] Conventionally, a method has been proposed in which information obtained by sensing a real 3D space, for example, multi-viewpoint images of a subject captured from different viewpoints, is used to generate a 3D object in a viewing space and a video (volumetric video) that appears as if the object exists in the viewing space (for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in Patent Document 1, since the subject cannot know the installation position of the virtual camera, there is a problem that when performing a performance such as singing or dancing, the subject cannot perform a performance while being aware of the position of the virtual camera.
[0005] The present disclosure proposes an information processing apparatus and a program capable of informing a subject of the position of a virtual camera that is observing the subject.
Means for Solving the Problems
[0006] To solve the above problems, one embodiment of the information processing apparatus according to the present disclosure is an information processing apparatus comprising: a first acquisition unit that acquires a plurality of real images captured by a plurality of first imaging devices arranged around a subject; a generation unit that generates a 3D model of the subject from the plurality of real images; and a presentation unit that presents information relating to a virtual viewpoint when rendering the 3D model into an image in a form corresponding to a viewing device to the subject. [Brief explanation of the drawing]
[0007] [Figure 1] This is a system configuration diagram showing an overview of the video processing system of the first embodiment. [Figure 2] This diagram shows an overview of the process for generating a 3D model of the subject. [Figure 3] This diagram shows the data required to represent a 3D model. [Figure 4] This diagram shows the schematic configuration of the imaging display device installed in the studio. [Figure 5] This figure shows an example of timing control for turning the display panel ON / OFF and the camera ON / OFF. [Figure 6] This figure shows an example of virtual camera information displayed on the display panel. [Figure 7] The first figure shows a specific example of virtual camera presentation information. [Figure 8] The second figure shows a specific example of virtual camera presentation information. [Figure 9] This figure shows typical variations of virtual camera presentation information. [Figure 10] This figure shows an example of virtual camera information indicating that the virtual camera is set to a location where there is no display panel. [Figure 11] This figure shows an example of virtual camera display information showing the camera's movement. [Figure 12] This figure shows an example of virtual camera display information when the setting positions of multiple virtual cameras overlap. [Figure 13]It is a functional block diagram showing an example of the functional configuration of the video processing system according to the first embodiment. [Figure 14] It is a diagram showing an example of the input / output information of the virtual camera information generation unit. [Figure 15] It is a flowchart showing an example of the processing flow performed by the video processing system according to the first embodiment. [Figure 16] It is a flowchart showing an example of the processing flow of the virtual camera information generation process in FIG. 15. [Figure 17] It is a flowchart showing an example of the processing flow of the virtual camera presentation information generation process in FIG. 15. [Figure 18] It is a flowchart showing an example of the processing flow of the virtual camera group display type determination process in FIG. 17. [Figure 19] It is a flowchart showing an example of the processing flow of the virtual camera group priority determination process in FIG. 17. [Figure 20] It is a flowchart showing an example of the processing flow of the virtual camera group presentation information generation process in FIG. 17. [Figure 21] It is a flowchart showing an example of the processing flow of the virtual camera presentation information generation process (normal) in FIG. 20. [Figure 22] It is a flowchart showing an example of the processing flow of the virtual camera presentation information generation process (position correction) in FIG. 20. [Figure 23] It is a flowchart showing an example of the processing flow of the virtual camera group presentation information generation process (normal) in FIG. 20. [Figure 24] It is a flowchart showing an example of the processing flow of the virtual camera group presentation information generation process (position correction) in FIG. 20. [Figure 25] It is a flowchart showing an example of the processing flow of the camera work display process in FIG. 20. [Figure 26] It is a flowchart showing an example of the processing flow of the virtual camera group voice generation process in FIG. 17. [Figure 27] It is a flowchart showing an example of the processing flow of the virtual camera presentation information output process in FIG. 15. [Figure 28] It is a flowchart showing an example of the flow of Volumetric video generation processing in FIG. 15. [Figure 29] It is a flowchart showing an example of the flow of superimposition processing of Volumetric video and background video in FIG. 15. [Figure 30] It is a system configuration diagram showing an overview of the video processing system of the second embodiment. [Figure 31] It is a functional block diagram showing an example of the functional configuration of the video processing system of the second embodiment. [Figure 32] It is a system configuration diagram showing an overview of the video processing system of the third embodiment. [Figure 33] It is a functional block diagram showing an example of the functional configuration of the video processing system of the third embodiment. [Figure 34] It is a diagram showing a method for a user to set camera work information using a viewing device. [Figure 35] It is a diagram showing a method for a user to set an operator video, an operator voice, and an operator message using a viewing device. [Figure 36] It is a diagram showing an example of virtual camera group presentation information according to the number of viewing users. [Figure 37] It is a diagram showing an example of virtual camera group presentation information when a viewing user changes the viewing position. [Figure 38] It is a diagram showing an example of a function for a viewing user and a performer to communicate with each other. [Figure 39] It is a flowchart showing an example of the flow of processing performed by the video processing system of the third embodiment. [Figure 40] It is a flowchart showing an example of the flow of communication video / audio generation processing in FIG. 39.
Embodiments for Carrying Out the Invention
[0008] Embodiments of this disclosure will be described in detail below with reference to the drawings. In each of the following embodiments, the same parts will be denoted by the same reference numerals to avoid redundant descriptions.
[0009] Furthermore, this disclosure will be explained in the order of the items shown below. 1. First Embodiment 1-1. Schematic Configuration of the Image Processing System of the First Embodiment 1-2. Explanation of Prerequisites - 3D Model Generation 1-3. Explanation of Prerequisites - Data Structure of 3D Models 1-4. Schematic Configuration of the Image Display Device 1-5. Explanation of information displayed by the virtual camera 1-6. Variations of information displayed by virtual cameras 1-7. Functional configuration of the video processing system of the first embodiment 1-8. Overall flow of processing performed by the video processing system of the first embodiment 1-9. Flowchart of virtual camera information generation process 1-10. Flowchart of the process for generating virtual camera display information 1-10-1. Flowchart for determining the display type of virtual camera group 1-10-2. Flowchart for determining virtual camera group priority 1-10-3. Flowchart for generating virtual camera group presentation information 1-10-4. Flowchart of virtual camera group audio generation process 1-11. Flowchart of virtual camera display information output processing 1-12. Flowchart of Volumetric Video Generation Process 1-13. Flowchart for superimposing volumetric video and background video 1-14. Effects of the First Embodiment 2. Second Embodiment 2-1. Schematic Configuration of the Video Processing System of the Second Embodiment 2-2. Functional Configuration of the Video Processing System in the Second Embodiment 2-3. Operation of the video processing system of the second embodiment 2-4. Effects of the second embodiment 3. Third Embodiment 3-1. Schematic Configuration of the Video Processing System of the Third Embodiment 3-2. Functional Configuration of the Video Processing System of the Third Embodiment 3-3. How to obtain virtual camera information 3-4. Format of information presented by virtual camera groups 3-5. Processing flow of the video processing system of the third embodiment 3-6. Effects of the Third Embodiment
[0010] (1. First Embodiment) [1-1. Schematic Configuration of the Image Processing System in the First Embodiment] First, the first embodiment of the video processing system 10a of this disclosure will be described using Figure 1. Figure 1 is a system configuration diagram showing an overview of the video processing system of the first embodiment.
[0011] The video processing system 10a comprises a volumetric studio 14a and a video processing device 12a. It is preferable to install the video processing device 12a in the volumetric studio 14a in order to process the video captured in the volumetric studio 14a with minimal latency.
[0012] Volumetric Studio 14a is a studio where images of the subject 22 are taken in order to generate a 3D model 22M of the subject 22. An image display device 13 is installed in Volumetric Studio 14a.
[0013] The imaging display device 13 captures images of the subject 22 using multiple cameras 16 arranged on the inner wall surface 15 of the Volumetric studio 14a so as to surround the subject 22. The imaging display device 13 also displays information related to the virtual viewpoint when rendering the 3D model 22M of the subject 22 into an image in a format appropriate to the user's viewing device on a display panel 17 arranged on the inner wall surface 15 of the Volumetric studio 14a so as to surround the subject 22. This information related to the virtual viewpoint includes, for example, information indicating the position and observation direction of the virtual camera.
[0014] The video processing device 12a generates a 3D model 22M of the subject 22 based on the actual camera image I acquired from the camera 16. The video processing device 12a also generates information related to a virtual viewpoint (virtual camera presentation information 20) when rendering the 3D model 22M of the subject 22 into an image in a format appropriate to the user's viewing device. The video processing device 12a then outputs the generated virtual camera presentation information 20 to the display panel 17. The video processing device 12a also generates a volumetric image 24 by rendering an image of the 3D model 22M of the subject 22 viewed from a set virtual viewpoint in a format appropriate to the viewing device. Specifically, if the user's viewing device is a 2D display such as a tablet or smartphone, the video processing device 12a renders the 3D model 22M of the subject 22 into a 2D image. If the user's viewing device is a viewing device capable of displaying 3D information, such as an HMD (Head Mount Display), the video processing device 12a renders the 3D model 22M of the subject 22 into a 3D image.
[0015] Furthermore, the video processing device 12a superimposes the generated volumetric video 24 onto the acquired background video 26a to generate video observed from a set virtual viewpoint. The generated video is distributed, for example, to the user's viewing environment and displayed on the user's viewing device. Note that the video processing device 12a is an example of an information processing device in this disclosure.
[0016] [1-2. Explanation of Prerequisites - 3D Model Generation] Next, using Figure 2, we will explain the process flow for generating a 3D model of the subject, which is a prerequisite for this embodiment. Figure 2 is a diagram illustrating the overview of the process for generating a 3D model of the subject.
[0017] As shown in Figure 2, the 3D model 22M of the subject 22 is generated through a process that involves imaging the subject 22 with multiple cameras 16 (16a, 16b, 16c) and generating a 3D model 22M containing 3D information of the subject 22 through 3D modeling.
[0018] Specifically, as shown in Figure 2, the multiple cameras 16 are positioned outside the subject 22, facing the subject 22, so as to surround the subject 22. Figure 2 shows an example with three cameras, where cameras 16a, 16b, and 16c are positioned around the subject 22. In Figure 2, the subject 22 is a person, but the subject 22 is not limited to a person. Also, the number of cameras 16 is not limited to three, and more cameras may be used.
[0019] 3D modeling is performed using multiple viewpoint images (actual camera footage I) that are synchronously and volumetrically captured (hereinafter referred to as Volumetric) by three cameras 16a, 16b, and 16c from different viewpoints. A 3D model 22M of the subject 22 is generated for each video frame from the three cameras 16a, 16b, and 16c.
[0020] The 3D model 22M is a model that contains 3D information of the object 22. The 3D model 22M contains shape information representing the surface shape of the object 22, for example, in the form of mesh data called a polygon mesh, which is represented by the connections between vertices. In addition, the 3D model 22M contains texture information representing the surface state of the object 22, corresponding to each polygon mesh. Note that the format of the information contained in the 3D model 22M is not limited to these, and other formats of information may also be used.
[0021] When reconstructing the 3D model 22M, texture mapping is performed, which involves applying textures that represent the color, pattern, and texture of a mesh according to its position. To improve the realism of the 3D model 22M, it is desirable to apply a View-Dependent (VD) texture that corresponds to the viewpoint position. This allows for higher-quality virtual images because the texture changes according to the viewpoint position when the 3D model 22M is imaged from an arbitrary virtual viewpoint. However, since this increases the bandwidth required for transmission, it is also acceptable to apply a View-Independent (VI) texture to the 3D model 22M that does not depend on the viewpoint position.
[0022] The volumetric video 24, which includes the retrieved 3D model 22M, is superimposed on the background video 26a and transmitted to a playback device, such as a mobile terminal 80, for playback. Rendering of the 3D model 22M and playback of the volumetric video 24 containing the 3D model 22M results in a 3D image being displayed on the user's mobile terminal 80.
[0023] [1-3. Explanation of Prerequisites - Data Structure of 3D Models] Next, we will explain the contents of the data required to represent the 3D model 22M using Figure 3. Figure 3 is a diagram showing the contents of the data required to represent the 3D model.
[0024] The 3D model 22M of the subject 22 is represented by mesh information M that shows the shape of the subject 22 and texture information T that shows the surface texture (color, pattern, etc.) of the subject 22.
[0025] Mesh information M represents the shape of the 3D model 22M by using several parts on the surface of the 3D model 22M as vertices and connecting these vertices (polygon mesh). Alternatively, instead of mesh information M, depth information Dp (not shown) representing the distance from the viewpoint position where the subject 22 is observed to the surface of the subject 22 may be used. The depth information Dp of the subject 22 is calculated, for example, based on the parallax for the same area of the subject 22 detected from images captured by an adjacent imaging device. Alternatively, instead of an imaging device, a sensor equipped with a ranging mechanism (e.g., a TOF (Time Of Flight) camera) or an infrared (IR) camera may be installed to obtain the distance to the subject 22.
[0026] In this embodiment, two types of data are used as texture information T. One is texture information Ta (VI) that is independent of the viewpoint position from which the 3D model 22M is observed. Texture information Ta is data that stores the texture of the surface of the 3D model 22M in the form of an unfolded diagram, such as the UV texture map shown in Figure 3. That is, texture information Ta is data that is independent of the viewpoint position. For example, if the 3D model 22M is a person wearing clothes, a UV texture map including the pattern of the clothes and the person's skin and hair is prepared as texture information Ta. The 3D model 22M can then be drawn by applying the texture information Ta corresponding to the mesh information M to the surface of the mesh information M representing the 3D model 22M (VI rendering). In this case, even if the observation position of the 3D model 22M changes, the same texture information Ta is applied to the mesh representing the same area. Thus, VI rendering using texture information Ta is performed by applying the texture information Ta of the clothing worn by the 3D model 22M to all the meshes representing parts of the clothing. As a result, the data size is generally small and the computational load of the rendering process is light. However, since the applied texture information Ta is uniform and the texture does not change even if the observation position is changed, the texture quality is generally low.
[0027] Another type of texture information T is VD (Variable Directional) texture information Tb, which depends on the viewpoint position from which the 3D model 22M is observed. Texture information Tb is represented by a set of images taken from multiple viewpoints of the subject 22. In other words, texture information Tb is data corresponding to the viewpoint position. Specifically, if the subject 22 is observed by N cameras, texture information Tb is represented by N images simultaneously captured by each camera. When rendering texture information Tb to an arbitrary mesh of the 3D model 22M, all regions corresponding to the relevant mesh are detected from the N images. Then, the textures captured in each of the detected regions are weighted and applied to the corresponding mesh. Thus, VD rendering using texture information Tb generally results in a large data size and a heavy computational load for rendering. However, because the applied texture information Tb changes according to the observation position, the texture quality is generally high.
[0028] [1-4. Schematic Configuration of the Imaging Display Device] Next, the schematic configuration of the imaging display device included in the image processing system 10a of the first embodiment will be described using Figures 4 and 5. Figure 4 is a schematic diagram of the imaging display device installed in a studio. Figure 5 is a diagram showing an example of timing control for the ON / OFF of the display panel and the ON / OFF of the camera.
[0029] In the Volumetric Studio 14a, multiple cameras 16 (16a, 16b, 16c…) are arranged around the subject 22, surrounding it. Multiple display panels 17 (17a, 17b, 17c…) are arranged to fill the gaps between adjacent cameras 16. The display panels 17 can be, for example, LED panels, liquid crystal panels, or organic EL panels. The multiple cameras 16 and multiple display panels 17 constitute the image display device 13a. In Figure 4, the cameras 16 and display panels 17 are arranged in a single row around the subject 22, but the cameras 16 and display panels 17 may be arranged in multiple rows in the vertical direction of the Volumetric Studio 14a.
[0030] In the imaging display device 13a, multiple cameras 16 simultaneously capture images of the subject 22 in order to generate a 3D model 22M of the subject 22. That is, the imaging timing of the multiple cameras 16 is synchronized.
[0031] Furthermore, in the imaging display device 13a, virtual camera presentation information 20 is displayed on multiple display panels 17. Details regarding the virtual camera presentation information 20 will be described later (see Figure 7).
[0032] Furthermore, the timing of the camera 16's image capture and the display panel 17's display timing are controlled so that they do not overlap. Details will be described later (see Figure 5).
[0033] The configuration of the imaging display device 13 is not limited to the imaging display device 13a. The imaging display device 13b shown in Figure 4 includes a projector 28 (28a, 28b, 28c…) and a transmissive screen 18 (18a, 18b, 18c…) onto which the image information projected by the projector 28 is projected, instead of a display panel 17 (17a, 17b, 17c…).
[0034] The projector 28 projects virtual camera information 20 from the back side of the transparent screen 18.
[0035] Furthermore, the imaging display device 13c shown in Figure 4 includes a projector 29 (29a, 29b, 29c…) and a reflective screen 19 (19a, 19b, 19c…) onto which the image information projected by the projector 29 is projected, instead of a display panel 17 (17a, 17b, 17c…).
[0036] The projector 28 projects virtual camera information 20 from the front side of the reflective screen 19.
[0037] Furthermore, as the simplest implementation of this disclosure, although not shown in the figures, instead of the display panel 17, a projection device such as a laser pointer capable of projecting a laser beam all around may be used to present the position of the virtual viewpoint as a bright spot.
[0038] The imaging of the subject 22 by the camera 16 and the display of virtual camera information 20 on the display panel 17 (or projectors 28, 29) are controlled based on the timing chart shown in Figure 5.
[0039] Specifically, the imaging display device 13 alternately performs the imaging operation of the camera 16 and the presentation of visual information to the display panel 17 (or projectors 28, 29) in a time-dependent manner. That is, when the camera 16 is imaging the subject 22, the presentation of visual information to the display panel 17 (or projectors 28, 29) (display of virtual camera presentation information 20) is not performed. On the other hand, when the virtual camera presentation information 20 is presented to the display panel 17 (or projectors 28, 29), the camera 16 does not image the subject 22. This prevents the virtual camera presentation information 20 from appearing in the background when the camera 16 is imaging the subject 22.
[0040] In Figure 5, the time spent by the camera 16 capturing images and the time spent presenting visual information (virtual camera-presented information 20) on the display panel 17 (or projectors 28, 29) are depicted as being approximately equal. The ratio of these times is set so that the movement of the subject 22 can be reliably captured, and the subject 22 can fully perceive the virtual camera-presented information 20.
[0041] The image processing device 12a performs a process to separate the subject 22 from the image that includes the captured subject 22. Therefore, while this process is being performed, virtual camera information 20 may be displayed on the display panel 17 (or projectors 28, 29). In addition, to reliably and easily separate the subject 22, imaging may be performed using an IR camera and IR light.
[0042] [1-5. Explanation of information displayed by the virtual camera] Next, specific examples of the virtual camera information 20 will be explained using Figures 6, 7, and 8. Figure 6 is a diagram showing an example of virtual camera information displayed on the display panel. Figure 7 is the first diagram showing a specific example of virtual camera information. Figure 8 is the second diagram showing a specific example of virtual camera information.
[0043] As shown in Figure 6, multiple display panels 17 are arranged on the inner wall surface 15 of the Volumetric studio 14a in the vertical direction along the H axis and the horizontal direction along the θ axis. Cameras 16 are installed adjacent to four of the display panels 17.
[0044] The image processing device 12a shown in Figure 1 displays a frame 21 at a position corresponding to the virtual viewpoint. Within the frame 21, for example, virtual camera presentation information 20, as shown in Figure 7, is displayed. The frame 21 is, for example, rectangular and is set at the position specified by the image processing device 12a: top-left vertex (θo, ho), width Wa, and height Ha. The virtual camera presentation information 20 is then displayed within the set frame 21.
[0045] As shown in Figure 6, the frame 21 that is set may overlap with multiple display panels 17. Also, since the number of virtual viewpoints set by the video processing device 12a is not limited to one, generally, multiple frame 21 are set on the inner wall surface 15 of the Volumetric studio 14a.
[0046] In this way, the frame 21 is configured to display, for example, the virtual camera presentation information 20 shown in Figure 7.
[0047] The virtual camera presentation information 20a(20) shown in Figure 7 includes a camera icon 30, a tally lamp 31, a cameraman icon 32, and a camera name 33 within the frame 21. The virtual camera presentation information 20a(20) is information that informs the subject 22 of the position of the virtual viewpoint set by the image processing device 12a. Note that the virtual camera presentation information 20 is an example of information related to the virtual viewpoint in this disclosure.
[0048] The camera icon 30 is an icon that simulates a virtual camera placed at the position of a virtual viewpoint set by the video processing device 12a. The camera icon 30 is displayed in a form that simulates the distance between the subject 22 and the virtual viewpoint, and the direction of the line of sight in the virtual viewpoint. In addition, the camera icon 30 is displayed in a form that makes it appear as if the camera is looking at the subject 22 from the other side of the inner wall surface 15 of the Volumetric studio 14a.
[0049] The tally lamp 31 indicates the operating status of the virtual camera placed at the virtual viewpoint. For example, when the virtual camera is capturing and broadcasting (On Air state), the tally lamp 31 lights up red. When the virtual camera is only capturing images, the tally lamp 31 lights up green.
[0050] The cameraman icon 32 is an icon uniquely associated with the operator controlling the virtual viewpoint, and any pre-set icon is displayed. By checking the cameraman icon 32, the subject 22 can recognize who the operator setting the position of the virtual viewpoint is. The size of the cameraman icon 32 may be changed according to the distance between the subject 22 and the virtual viewpoint. For example, the closer the subject 22 is to the virtual viewpoint, the larger the cameraman icon 32 may be displayed. The cameraman icon 32 may also be an image of the operator themselves.
[0051] Camera name 33 is an identifier uniquely associated with the virtual camera, and displays a pre-configured arbitrary name.
[0052] The virtual camera information 20 changes its form according to the state of the set virtual viewpoint. The virtual camera information 20b(20) shown in Figure 7 displays information related to a different virtual viewpoint than the virtual camera information 20a. More specifically, the virtual camera information 20b(20) is information from a different virtual camera than the virtual camera information 20a(20). Also, the line of sight direction in the virtual viewpoint is different from that of the virtual camera information 20a.
[0053] Furthermore, the camera icon 30 and cameraman icon 32 displayed in the virtual camera information 20b are drawn larger than the camera icon 30 and cameraman icon 32 in the virtual camera information 20a. This indicates that the position of the virtual viewpoint shown in the virtual camera information 20b is closer to the subject 22 than the position of the virtual viewpoint shown in the virtual camera information 20a.
[0054] Although not shown in Figure 7, the closer the virtual viewpoint is to the subject 22, the larger the size of the frame 21 may be.
[0055] The virtual camera presentation information 20c(20) shown in Figure 8 is an example in which an operator controlling the virtual viewpoint by the video processing device 12a displays a message to the subject 22. That is, the virtual camera presentation information 20c(20) includes message information 37.
[0056] [1-6. Variations of information displayed by virtual cameras] Next, we will explain the variations of the virtual camera presentation information 20 using Figures 9 to 12. Figure 9 is a diagram showing typical variations of the virtual camera presentation information.
[0057] In Figure 9, the virtual camera presentation information 20d(20) indicates that the virtual camera is facing the direction of the subject 22.
[0058] When virtual camera information 20d(20) is displayed and another virtual camera approaches, the video processing device 12a displays virtual camera information 20e(20). Virtual camera information 20e(20) indicates that "Camera 1" and "Camera 2" are in close proximity to each other. Virtual camera information 20 displayed in this grouped state is specifically called virtual camera group information 200.
[0059] Furthermore, when the virtual camera approaches the subject 22 while the virtual camera information 20d(20) is displayed, the virtual camera information 20f(20) is displayed. The virtual camera information 20f(20) indicates that the virtual camera has approached the subject 22 by drawing the camera icon 30 larger. At this time, the frame 21 may also be drawn larger. Although not shown in Figure 9, when the virtual camera moves away from the subject 22, the camera icon 30 is drawn smaller.
[0060] Virtual camera information 20g(20) is information that is presented when the orientation of the virtual camera changes from the state in which virtual camera information 20d(20) was presented. In Figure 9, virtual camera information 20g(20) indicates that the virtual camera has turned to the right.
[0061] The virtual camera information 20h(20) indicates that the virtual camera placed at the virtual viewpoint has actually started taking pictures. In this case, the display mode of the tally lamp 31 is changed to indicate that it is taking pictures.
[0062] Figure 10 shows an example of virtual camera presentation information indicating that the virtual camera is set to a location where there is no display panel.
[0063] The virtual viewpoint (virtual camera) can be placed at any position surrounding the subject 22. Therefore, the virtual camera can be placed in locations where the display panel 17 cannot be installed or is difficult to install, such as the ceiling or floor of the Volumetric studio 14a. In such cases, the video processing device 12a displays a camera position indicator icon 34 in the virtual camera display information 20 to indicate that the virtual camera is outside the installation location of the display panel 17.
[0064] The virtual camera presentation information 20i(20) shown in Figure 10 includes a camera position indicator icon 34a(34). The camera position indicator icon 34a(34) indicates that the virtual camera is set on the ceiling of the inner wall surface 15 of the Volumetric studio 14a.
[0065] Furthermore, the virtual camera presentation information 20j(20) includes a camera position indicator icon 34b(34). The camera position indicator icon 34b(34) indicates that the virtual camera is set on the floor surface of the inner wall surface 15 of the Volumetric studio 14a.
[0066] The virtual camera information 20k(20) shown in Figure 10 includes a camera position indicator icon 34c(34). The camera position indicator icon 34c(34) is a modified version of the camera position indicator icon 34a(34). The camera position indicator icon 34c(34) indicates approximately where the virtual camera is set on the ceiling. The rectangular area contained within the camera position indicator icon 34c(34) indicates the virtual camera's setting position. If the virtual camera is set at the top (ceiling) on the side where the virtual camera information 20k(20) is displayed, the rectangular area contained within the camera position indicator icon 34c(34) is displayed at the bottom of the icon. On the other hand, if the virtual camera is set at the top (ceiling) on the back side of the side where the virtual camera information 20k(20) is displayed, the rectangular area contained within the camera position indicator icon 34c(34) is displayed at the top of the icon. Furthermore, if the virtual camera is positioned directly above the subject 22, the rectangular area contained within the camera position indicator icon 34c(34) will be displayed in the center of the camera position indicator icon 34c(34).
[0067] Furthermore, the virtual camera information 20l(20) includes a camera position indicator icon 34d(34). The camera position indicator icon 34d(34) is a modified version of the camera position indicator icon 34b(34). The camera position indicator icon 34d(34) indicates approximately where the virtual camera is set on the floor. The rectangular area contained within the camera position indicator icon 34d(34) indicates the virtual camera's setting position. If the virtual camera is set at the bottom (floor) on the side where the virtual camera information 20l(20) is displayed, the rectangular area contained within the camera position indicator icon 34d(34) is displayed at the top of the camera position indicator icon 34d(34). On the other hand, if the virtual camera is set at the bottom (floor) on the back side of the side where the virtual camera information 20l(20) is displayed, the rectangular area contained within the camera position indicator icon 34d(34) is displayed at the bottom of the camera position indicator icon 34c(34). Furthermore, if the virtual camera is positioned directly below the subject 22, the rectangular area contained within the camera position indicator icon 34d(34) will be displayed in the center of the camera position indicator icon 34d(34).
[0068] Figure 11 shows an example where virtual camera display information shows the camera work of the virtual camera.
[0069] The virtual camera presentation information 20m(20) displayed within the frame 21 shown in Figure 11 includes camera work information 35, which indicates the movement trajectory of the virtual camera generated by the video processing device 12a, and camera work 36. The camera work information 35 indicates the name of the camera work.
[0070] The camera work 36 is an arrow indicating the actual direction of movement of the virtual camera. By representing the movement of the virtual camera with an arrow, the subject 22 can predict the movement of the virtual camera and perform accordingly. As shown in Figure 11, the direction of the camera work may be emphasized by displaying the front of the arrow indicating camera work 36 in a darker color and gradually fading the back of the arrow indicating camera work 36.
[0071] Furthermore, if the virtual camera's movement speed is slow, the current position of the virtual camera may be superimposed on the camera work 36 and displayed sequentially, as shown in Figure 11. However, if the virtual camera's movement speed is fast, the position of the virtual camera may be displayed at the endpoint of the camera work 36.
[0072] Figure 12 shows an example of virtual camera presentation information when the setting positions of multiple virtual cameras overlap.
[0073] The video processing device 12a sets up multiple virtual cameras on the inner wall surface 15 of the Volumetric studio 14a. Each of the set up virtual cameras can move freely. Therefore, the positions of the multiple virtual cameras may be close together.
[0074] Figure 12 shows how the two virtual cameras, as time t progresses, move towards each other, then pass each other and move away from one another.
[0075] In this case, initially, virtual camera presentation information 20n1(20) and virtual camera presentation information 20n2(20) corresponding to each virtual camera are displayed. Then, when the positions of two virtual cameras approach each other, virtual camera presentation information 20n3(20), i.e., virtual camera group presentation information 200, is displayed in one frame 21. The virtual camera group presentation information 200 includes the virtual camera presentation information 20 of multiple virtual cameras located in close proximity within a divided frame 21.
[0076] After the two virtual cameras pass each other, the virtual camera information 20n1(20) and virtual camera information 20n2(20) corresponding to each virtual camera are displayed again.
[0077] [1-7. Functional Configuration of the Video Processing System in the First Embodiment] Next, the functional configuration of the video processing system 10a will be explained using Figures 13 and 14. Figure 13 is a functional block diagram showing an example of the functional configuration of the video processing system of the first embodiment. Figure 14 is a diagram showing an example of input and output information of the virtual camera information generation unit.
[0078] As shown in Figure 13, the video processing system 10a comprises a video processing device 12a and a camera 16 and a display panel 17 that constitute the image capture and display device 13. The video processing system 10a also includes peripheral devices: a remote control 54, an intercom 55, a microphone 56, a speaker 57, and a viewing device 53a. The functions of the camera 16 and the display panel 17 are as described above, so their explanation will be omitted.
[0079] The video processing device 12a comprises a controller 40, a virtual camera information generation unit 41, a virtual camera presentation information generation unit 42, a UI unit 43, a studio video display unit 44, an audio output unit 45, a volumetric video shooting unit 46, a volumetric video generation unit 47, a master audio output unit 48, an audio recording unit 49, a CG background generation unit 50, a volumetric video / CG overlay / audio MUX unit 51, and a distribution unit 52. These functional units are realized by the CPU of the video processing device 12a, which has a computer configuration, executing an unillustrated control program that controls the operation of the video processing device 12a. Alternatively, all or some of the functions of the video processing device 12a may be realized by hardware.
[0080] The controller 40 generates information related to the virtual camera. The controller 40 is an information input device equipped with, for example, an operating device such as a joystick or selection buttons, and sets the position of the virtual viewpoint, camera work information, etc., according to the user's operating instructions. The video processing device 12a can set multiple virtual viewpoints by having multiple controllers 40.
[0081] The controller 40 also includes a camera and a microphone (not shown). The camera on the controller 40 captures images of the operator controlling the virtual viewpoint. The microphone on the controller 40 acquires the speech (voice) of the operator controlling the virtual viewpoint.
[0082] The controller 40 further includes operating devices such as selection buttons for selecting and sending messages from the operator controlling the virtual viewpoint.
[0083] The virtual camera information generation unit 41 acquires information related to the virtual viewpoint and information related to the operator from the controller 40. The information related to the virtual viewpoint includes, for example, the virtual camera position information Fa, camera work information Fb, and camera information Ff shown in Figure 14. The information related to the operator includes, for example, the operator video Fc, operator voice Fd, and operator message Fe shown in Figure 14. The virtual camera information generation unit 41 is an example of the second acquisition unit in this disclosure.
[0084] The virtual camera position information Fa includes the virtual camera's position coordinates, orientation, and field of view. The virtual camera position information Fa is set by operating an operation device such as a joystick provided by the controller 40.
[0085] Camera work information Fb is information relating to the movement trajectory of the virtual camera. Specifically, camera work information Fb includes the camera work start position, camera work end position, trajectory between the start and end positions, virtual camera movement speed, camera work name, etc. Camera work information Fb is set by operating an operation device such as a selection button provided on the controller 40.
[0086] Camera information Ff includes information related to the virtual viewpoint, such as camera number, camera name, camera status, camera icon / image, and camera priority.
[0087] The operator video Fc is video footage of the operator controlling the virtual viewpoint. The video processing device 12a may display the operator video Fc in the virtual camera presentation information 20 instead of the cameraman icon 32 (see Figure 7).
[0088] Operator voice Fd is an audio message that the operator controlling the virtual viewpoint communicates to the subject 22.
[0089] Operator message Fe is a text message that the operator controlling the virtual viewpoint communicates to the subject 22. Operator message Fe is set by operating an operation device such as a selection button on the controller 40.
[0090] The virtual camera information generation unit 41 generates virtual camera information F (see Figure 14) which is a compilation of the various acquired information for each virtual camera. The virtual camera information generation unit 41 then sends the generated virtual camera information F to the virtual camera presentation information generation unit 42. Regarding camera work information Fb, the virtual camera information generation unit 41 manages the playback status of the camera work internally, and if camera work playback is in progress, it updates the virtual camera position information sequentially.
[0091] The virtual camera presentation information generation unit 42 generates virtual camera presentation information 20 to be displayed on the display panel 17. More specifically, the virtual camera presentation information generation unit 42 generates information related to the virtual viewpoint when rendering the 3D model 22M of the subject 22 into an image in a form appropriate to the user's viewing device. More specifically, based on the virtual camera position coordinates and camera information contained in the virtual camera information F, the virtual camera presentation information 20 is generated by, as necessary, changing the display color of the tally lamp 31, combining multiple virtual camera presentation information 20, generating a camera position display icon 34 indicating that the virtual camera is on the ceiling or floor, etc. Furthermore, the virtual camera presentation information generation unit 42 generates audio output to be output from the audio output unit 45.
[0092] The UI unit 43 allows the subject 22 or the director to change the settings of various parameters used by the video processing device 12a using a remote control 54. By operating the UI unit 43, the subject 22 selects a specific operator to control the virtual viewpoint and engages in voice conversation with the selected operator. Note that the UI unit 43 is an example of a selection unit in this disclosure.
[0093] The studio video display unit 44 displays the virtual camera presentation information 20 received from the virtual camera presentation information generation unit 42 at corresponding positions on multiple display panels 17. Note that the studio video display unit 44 is an example of a presentation unit in this disclosure.
[0094] The audio output unit 45 outputs audio data received from the virtual camera information F to the intercom 55. This transmits various instructions from the operator controlling the virtual viewpoint to the subject 22.
[0095] The Volumetric video capture unit 46 is positioned around the subject 22 and captures real images of the subject 22 from multiple directions simultaneously using multiple externally synchronized cameras 16. The Volumetric video capture unit 46 also sends the real camera images I obtained by the capture to the Volumetric video generation unit 47 as Volumetric camera video data, which includes a frame number and identification information that identifies the camera 16 that captured the image. The Volumetric video capture unit 46 is an example of the first acquisition unit in this disclosure.
[0096] The Volumetric video generation unit 47 receives Volumetric camera video data from the Volumetric video shooting unit 46 and performs Volumetric video generation processing. The Volumetric video generation unit 47 holds calibration data that has been performed by internal calibration to correct distortion of the camera 16 and external calibration to determine the relative position of each camera 16, and corrects the captured actual camera video I using this calibration data. Then, based on the Volumetric camera video data acquired by the Volumetric video shooting unit 46, the Volumetric video generation unit 47 performs modeling processing of the subject 22, i.e., generates a 3D model 22M. After that, the Volumetric video generation unit 47 renders a Volumetric video of the 3D model 22M of the subject 22 as seen from a virtual viewpoint, based on the acquired virtual camera position information. The Volumetric video generation unit 47 sends the rendered Volumetric video, frame number, and virtual camera information F to the Volumetric video / CG superimposition / audio MUX unit 51. Note that the Volumetric video generation unit 47 is an example of a generation unit in this disclosure.
[0097] The master audio output unit 48 outputs the music played by the subject 22 during singing or dance performances through the speaker 57. The master audio output unit 48 also sends the audio data of the music to the audio recording unit 49.
[0098] The audio recording unit 49 generates audio data by mixing the audio data from the master audio output unit 48 with the audio data input from the microphone 56 (for example, the singing data of the subject 22), and sends it to the Volumetric video / CG superimposition / audio MUX unit 51.
[0099] The CG background generation unit 50 generates background CG data with frame numbers based on pre-prepared background CG data. The CG background generation unit 50 then sends the generated background CG data to the Volumetric video / CG overlay / audio MUX unit 51.
[0100] The Volumetric Video / CG Overlay / Audio MUX Unit 51 generates, for example, a 2D image as seen from a virtual viewpoint by rendering and overlaying the acquired Volumetric video data and background CG data based on the virtual camera position information contained in the Volumetric video data. The Volumetric Video / CG Overlay / Audio MUX Unit 51 then sends the distributed content, which is the generated 2D image and audio information multiplexed (MUXed), to the distribution unit 52. If the user's viewing device 53a is a device capable of displaying 3D information, the Volumetric Video / CG Overlay / Audio MUX Unit 51 generates a 3D image by rendering the 3D model 22M of the subject 22 into a 3D image.
[0101] The distribution unit 52 distributes the content received from the Volumetric video / CG superimposition / audio MUX unit 51 to the viewing device 53a.
[0102] The remote control 54, provided as a peripheral device, is used to change the settings of various parameters used by the video processing unit 12a.
[0103] The intercom 55 is worn by the subject 22 to hear audio from the operator controlling the virtual viewpoint.
[0104] Microphone 56 records the singing and conversations of subject 22.
[0105] Speaker 57 outputs music or other sounds that the subject 22 listens to during shooting.
[0106] The viewing device 53a is a device used by the user to view content delivered from 12a. Examples of viewing devices 53a include tablet devices and smartphones.
[0107] [1-8. Overall flow of processing performed by the video processing system of the first embodiment] Figure 15 will be used to explain the overall flow of processing performed by the video processing system 10a.
[0108] The virtual camera information generation unit 41 performs a virtual camera information generation process to generate virtual camera information F (step S11). Details of the virtual camera information generation process will be described later (see Figure 16).
[0109] The virtual camera presentation information generation unit 42 performs a virtual camera presentation information generation process to generate virtual camera presentation information 20 (step S12). Details of the virtual camera presentation information generation process will be described later (see Figure 17).
[0110] The in-studio video display unit 44 generates an image that displays the virtual camera display information 20 at the corresponding position on the display panel 17, and performs virtual camera display information output processing to output the generated image to the display panel 17 (step S13). Details of the virtual camera display information output processing will be described later (see Figures 17 and 27).
[0111] The volumetric video generation unit 47 performs volumetric video generation processing to generate volumetric video based on volumetric camera video data received from the volumetric video capture unit 46 (step S14). The flow of the volumetric video generation processing will be described later (see Figure 28).
[0112] The Volumetric Video / CG Overlay / Audio MUX Unit 51 performs the overlay processing of the Volumetric video and background video (step S15). The flow of the Volumetric video and background video overlay processing will be described later (see Figure 29).
[0113] The distribution unit 52 performs distribution processing to distribute the content received from the Volumetric video / CG superimposition / audio MUX unit 51 to the viewing device 53a (step S16).
[0114] [1-9. Flow of virtual camera information generation process] The flow of the virtual camera information generation process will be explained using Figure 16. Figure 16 is a flowchart showing an example of the virtual camera information generation process in Figure 15.
[0115] The virtual camera information generation unit 41 obtains virtual camera position information Fa and camera work information Fb from the controller 40 (step S21).
[0116] The virtual camera information generation unit 41 updates the camera work queue based on the camera work information Fb (step S22).
[0117] The virtual camera information generation unit 41 determines whether there is a camera work currently being played in the camera work queue (step S23). If it is determined that there is a camera work currently being played (step S23: Yes), the process proceeds to step S24. On the other hand, if it is not determined that there is a camera work currently being played (step S23: No), the process proceeds to step S26.
[0118] In step S23, if it is determined that there is camera work currently being played back, the virtual camera information generation unit 41 updates the virtual camera position information Fa based on the frame number and camera work information Fb of the camera work currently being played back (step S24).
[0119] Next, the virtual camera information generation unit 41 generates virtual camera information F and sets the camera work name and playback frame number based on the current camera work (step S25). After that, it returns to the main routine (Figure 15).
[0120] On the other hand, if it is determined in step S23 that there is no camera work currently being played back, the virtual camera information generation unit 41 maintains the position of the virtual camera at that time by clearing the camera work name and playback frame number (step S26). After that, it returns to the main routine (Figure 15).
[0121] [1-10. Flow of virtual camera display information generation process] The flow of the virtual camera information generation process will be explained using Figure 17. Figure 17 is a flowchart showing an example of the virtual camera presentation information generation process in Figure 15.
[0122] The virtual camera information generation unit 42 acquires all virtual camera information F for the current frame number (step S31).
[0123] The virtual camera presentation information generation unit 42 generates virtual camera presentation information 20 (step S32).
[0124] The virtual camera presentation information generation unit 42 generates virtual camera group presentation information 200 by grouping nearby cameras based on the generated virtual camera presentation information 20 (step S33).
[0125] The virtual camera display information generation unit 42 performs virtual camera group display type determination processing based on the virtual camera group display information 200 (step S34). Details of the virtual camera group display type determination processing will be described later (see Figure 18).
[0126] The virtual camera information generation unit 42 performs a virtual camera group priority determination process to sort the virtual camera information F included in the same group based on the camera state and camera priority (step S35). Details of the virtual camera group priority determination process will be described later (see Figure 19).
[0127] The virtual camera presentation information generation unit 42 performs a virtual camera group presentation information generation process to generate virtual camera group presentation information 200 (step S36). Details of the virtual camera group presentation information generation process will be described later (see Figure 20).
[0128] The virtual camera presentation information generation unit 42 performs virtual camera group audio generation processing to generate audio output to be presented to the subject 22 (step S37). Details of the virtual camera group audio generation processing will be described later (see Figure 26). After that, the process returns to the main routine (Figure 15).
[0129] [1-10-1. Flow of virtual camera group display type determination process] Using Figure 18, we will explain the flow of the virtual camera group display type determination process shown in step S34 of Figure 17. Figure 18 is a flowchart showing an example of the flow of the virtual camera group display type determination process in Figure 17.
[0130] The virtual camera display information generation unit 42 determines whether the number of virtual cameras is 2 or more and whether the maximum number of divisions when grouping virtual cameras is 2 or more (step S41). If the conditions are met (step S41: Yes), the process proceeds to step S42. On the other hand, if the conditions are not met (step S41: No), the process proceeds to step S43.
[0131] If it is determined in step S41 that the conditions are met, the virtual camera presentation information generation unit 42 determines whether the number of virtual cameras is 4 or more and whether the maximum number of divisions when grouping virtual cameras is 4 or more (step S42). If the conditions are met (step S42: Yes), the number of virtual cameras and the maximum number of divisions when grouping virtual cameras are increased, and the same determination as in steps S41 and S42 is continued. On the other hand, if the conditions are not met (step S42: No), the process proceeds to step S45.
[0132] The same determination process as in steps S41 and S42 continues, and if it is determined that the conditions are met, the virtual camera presentation information generation unit 42 determines whether the number of virtual cameras is 7 or more and whether the maximum number of divisions when grouping virtual cameras is 7 or more (step S44). If the conditions are met (step S44: Yes), the process proceeds to step S47. On the other hand, if the conditions are not met (step S44: No), the process proceeds to step S46.
[0133] If it is determined in step S41 that the condition is not met (step S41: No), the virtual camera display information generation unit 42 sets the virtual camera display type to 1, that is, the number of virtual camera display divisions to 1 (step S43). After that, the process returns to the flowchart in Figure 17.
[0134] If it is determined in step S42 that the condition is not met (step S42: No), the virtual camera display information generation unit 42 sets the virtual camera display type to 2, that is, the number of virtual camera display divisions to 2 (step S43). After that, the process returns to the flowchart in Figure 17.
[0135] If it is determined in step S44 that the condition is not met (step S44: Yes), the virtual camera display information generation unit 42 sets the virtual camera display type to 64, that is, the number of virtual camera display divisions to 64 (step S43). After that, the process returns to the flowchart in Figure 17.
[0136] If it is determined in step S44 that the condition is not met (step S44: No), the virtual camera display information generation unit 42 sets the virtual camera display type to 49, that is, the number of virtual camera display divisions to 49 (step S43). After that, the process returns to the flowchart in Figure 17.
[0137] [1-10-2. Flow of virtual camera group priority determination process] Using Figure 19, we will explain the flow of the virtual camera group priority determination process shown in step S35 of Figure 17. Figure 19 is a flowchart showing an example of the flow of the virtual camera group priority determination process in Figure 17.
[0138] The virtual camera information generation unit 42 sorts the virtual camera information F included in the same group according to the camera state and camera priority (step S51). Then, it returns to the flowchart in Figure 17.
[0139] [1-10-3. Flowchart for generating virtual camera group presentation information] Using Figure 20, we will explain the flow of the virtual camera group presentation information generation process shown in step S36 of Figure 17. Figure 20 is a flowchart showing an example of the flow of the virtual camera group presentation information generation process in Figure 17.
[0140] The virtual camera presentation information generation unit 42 determines whether there is only one virtual camera in the group (step S61). If it is determined that there is only one virtual camera in the group (step S61: Yes), the process proceeds to step S62. On the other hand, if it is not determined that there is only one virtual camera in the group (step S61: No), the process proceeds to step S68.
[0141] In step S61, if it is determined that there is only one virtual camera in the group, the virtual camera presentation information generation unit 42 determines whether the image frame is in a position where it can be displayed (step S62). If it is determined that the image frame is in a position where it can be displayed (step S62: Yes), the process proceeds to step S63. On the other hand, if it is not determined that the image frame is in a position where it can be displayed (step S62: No), the process proceeds to step S64.
[0142] In step S62, if it is determined that the image frame is in a position where it can be displayed, the virtual camera presentation information generation unit 42 generates normal virtual camera presentation information 20 (step S63). Then, the process proceeds to step S65. The detailed flow of the process performed in step S63 will be described later (see Figure 21).
[0143] If, in step S62, it is determined that the image frame is not in a position where it can be displayed, the virtual camera presentation information generation unit 42 generates position-corrected virtual camera presentation information 20 (step S64). Then, the process proceeds to step S65. The detailed flow of the process performed in step S64 will be described later (see Figure 22).
[0144] Following step S63 or step S64, the virtual camera presentation information generation unit 42 determines whether camera work is being played back (step S65). If it is determined that camera work is being played back (step S65: Yes), the process proceeds to step S66. On the other hand, if it is not determined that camera work is being played back, the process returns to the flowchart in Figure 17.
[0145] In step S65, if it is determined that camera work is being played back, the virtual camera presentation information generation unit 42 determines whether the camera work display setting is turned on (step S66). If it is determined that the camera work display setting is turned on (step S66: Yes), the process proceeds to step S67. On the other hand, if it is not determined that the camera work display setting is turned on (step S66: No), the process returns to the flowchart in Figure 17.
[0146] In step S66, if it is determined that the camera work display setting is turned on, the virtual camera presentation information generation unit 42 performs the camera work display process (step S67). After that, the process returns to the flowchart in Figure 17. The detailed flow of the process performed in step S67 will be described later (see Figure 25).
[0147] Returning to step S61, if it is not determined in step S61 that there is only one virtual camera in the group, the virtual camera presentation information generation unit 42 determines whether the frame is in a position where it can be displayed (step S68). If it is determined that the frame is in a position where it can be displayed (step S68: Yes), the process proceeds to step S69. On the other hand, if it is not determined that the frame is in a position where it can be displayed (step S68: No), the process proceeds to step S70.
[0148] In step S68, if it is determined that the image frame is in a position where it can be displayed, the virtual camera presentation information generation unit 42 generates normal virtual camera group presentation information 200 (step S69). After that, the process returns to the flowchart in Figure 17. The detailed flow of the process performed in step S68 will be described later (see Figure 23).
[0149] In step S68, if it is determined that the image frame is not in a position where it can be displayed, the virtual camera presentation information generation unit 42 generates position-corrected virtual camera group presentation information 200 (step S70). After that, the process returns to the flowchart in Figure 17. The detailed flow of the process performed in step S70 will be described later (see Figure 24).
[0150] Next, we will explain the flow of the normal virtual camera presentation information generation process 20 using Figure 21. Figure 21 is a flowchart showing an example of the normal virtual camera presentation information generation process in Figure 20.
[0151] The virtual camera presentation information generation unit 42 determines whether the display mode of the virtual camera presentation information 20 is normal (step S71). If it is determined that the display mode of the virtual camera presentation information 20 is normal (step S71: Yes), the process proceeds to step S72. On the other hand, if it is not determined that the display mode of the virtual camera presentation information 20 is normal (step S71: No), the process proceeds to step S73.
[0152] In step S71, if it is determined that the display mode of the virtual camera presentation information 20 is normal, the virtual camera presentation information generation unit 42 generates the virtual camera presentation information 20 based on the virtual camera information F (step S72). Then, the process returns to the flowchart in Figure 20. Note that the virtual camera presentation information 20p1(20) shown in Figure 21 is an example of the virtual camera presentation information generated in step S72.
[0153] On the other hand, if the display mode of the virtual camera presentation information 20 is not determined to be normal in step S71, the virtual camera presentation information generation unit 42 generates virtual camera presentation information 20 in which particles 38 resembling a virtual camera are drawn (step S73). After that, the process returns to the flowchart in Figure 20. Note that the virtual camera presentation information 20p2(20) shown in Figure 21 is an example of virtual camera presentation information generated in step S73.
[0154] Next, we will explain the process flow for generating the position-corrected virtual camera presentation information 20 using Figure 22. Figure 22 is a flowchart showing an example of the virtual camera presentation information generation process (position correction) flow in Figure 20.
[0155] The virtual camera presentation information generation unit 42 determines whether the display mode of the virtual camera presentation information 20 is normal (step S81). If it is determined that the display mode of the virtual camera presentation information 20 is normal (step S81: Yes), the process proceeds to step S82. On the other hand, if it is not determined that the display mode of the virtual camera presentation information 20 is normal (step S81: No), the process proceeds to step S83.
[0156] In step S81, if it is determined that the display mode of the virtual camera presentation information 20 is normal, the virtual camera presentation information generation unit 42 updates the field of view information based on the virtual camera information F and generates the virtual camera presentation information 20 (step S82). After that, the process returns to the flowchart in Figure 20. The virtual camera presentation information 20q1(20) and 20q2(20) shown in Figure 22 are examples of virtual camera presentation information generated in step S82 and displayed on the inner wall surface 15.
[0157] On the other hand, if the display mode of the virtual camera presentation information 20 is not determined to be normal in step S81, the virtual camera presentation information generation unit 42 generates virtual camera presentation information 20 in which particles resembling a virtual camera are drawn (step S83). After that, the process returns to the flowchart in Figure 20. Note that the virtual camera presentation information 20q3(20) and 20q4(20) shown in Figure 22 are examples of virtual camera presentation information generated in step S83 and displayed on the inner wall surface 15.
[0158] Next, we will explain the normal flow of generating virtual camera group presentation information 200 using Figure 23. Figure 23 is a flowchart showing an example of the normal flow of the virtual camera group presentation information generation process in Figure 20.
[0159] The virtual camera presentation information generation unit 42 determines whether the display mode of the virtual camera presentation information 20 is normal (step S91). If it is determined that the display mode of the virtual camera presentation information 20 is normal (step S91: Yes), the process proceeds to step S92. On the other hand, if it is not determined that the display mode of the virtual camera presentation information 20 is normal (step S91: No), the process proceeds to step S96.
[0160] In step S91, if it is determined that the display mode of the virtual camera presentation information 20 is normal, the virtual camera presentation information generation unit 42 determines whether there is remaining space in the divided display frame of the frame 21 (step S92). If it is determined that there is remaining space in the divided display frame of the frame 21 (step S92: Yes), the process proceeds to step S93. On the other hand, if it is not determined that there is remaining space in the divided display frame of the frame 21 (step S92: No), the process returns to the flowchart in Figure 20.
[0161] In step S92, if it is determined that there is remaining space in the divided display frame of the picture frame 21, the virtual camera presentation information generation unit 42 determines whether there is a virtual camera to be displayed (step S93). If it is determined that there is a virtual camera to be displayed (step S93: Yes), the process proceeds to step S94. On the other hand, if it is not determined that there is a virtual camera to be displayed (step S93: No), the process returns to the flowchart in Figure 20.
[0162] In step S93, if it is determined that there is a virtual camera to be displayed, the virtual camera presentation information generation unit 42 performs the normal virtual camera presentation information generation process 20 by executing the flowchart in Figure 21 (step S94).
[0163] Then, the virtual camera presentation information generation unit 42 draws the virtual camera presentation information 20 generated in step S94 onto the divided display frame (step S95). After that, the process returns to step S92 and the above-described process is repeated. Note that the virtual camera presentation information 200a (200) shown in Figure 23 is an example of the information generated in step S95.
[0164] On the other hand, if the display mode of the virtual camera presentation information 20 is not determined to be normal in step S91, the virtual camera presentation information generation unit 42 generates virtual camera group presentation information 200 in which particles 38 resembling a virtual camera are drawn (step S96). After that, the process returns to the flowchart in Figure 20. Note that the virtual camera presentation information 200b(200) shown in Figure 23 is an example of the information generated in step S96.
[0165] Next, we will explain the process flow for generating position-corrected virtual camera group presentation information 200 using Figure 24. Figure 24 is a flowchart showing an example of the virtual camera group presentation information generation process (position correction) in Figure 20.
[0166] The virtual camera presentation information generation unit 42 determines whether the display mode of the virtual camera presentation information 20 is normal (step S101). If it is determined that the display mode of the virtual camera presentation information 20 is normal (step S101: Yes), the process proceeds to step S102. On the other hand, if it is not determined that the display mode of the virtual camera presentation information 20 is normal (step S101: No), the process proceeds to step S107.
[0167] In step S101, if it is determined that the display mode of the virtual camera presentation information 20 is normal, the virtual camera presentation information generation unit 42 determines whether there is remaining space in the divided display frame of the frame 21 (step S102). If it is determined that there is remaining space in the divided display frame of the frame 21 (step S102: Yes), the process proceeds to step S103. On the other hand, if it is not determined that there is remaining space in the divided display frame of the frame 21 (step S102: No), the process proceeds to step S106.
[0168] In step S102, if it is determined that there is remaining space in the divided display frame of the picture frame 21, the virtual camera presentation information generation unit 42 determines whether there is a virtual camera to be displayed (step S103). If it is determined that there is a virtual camera to be displayed (step S103: Yes), the process proceeds to step S104. On the other hand, if it is not determined that there is a virtual camera to be displayed (step S103: No), the process proceeds to step S106.
[0169] In step S103, if it is determined that there is a virtual camera to be displayed, the virtual camera presentation information generation unit 42 performs the process of generating position-corrected virtual camera presentation information 20 by executing the flowchart in Figure 22 (step S104).
[0170] Then, the virtual camera presentation information generation unit 42 draws the virtual camera presentation information 20 generated in step S104 onto the divided display frame (step S105). After that, the process returns to step S102 and the above-described process is repeated. Note that the virtual camera presentation information 200c (200) shown in Figure 24 is an example of virtual camera group presentation information generated in step S105.
[0171] If, in step S102, it is determined that there is no remaining space in the divided display frame of the picture frame 21, or if, in step S103, it is determined that there is a virtual camera to be displayed, the virtual camera presentation information generation unit 42 corrects the position of the divided display frame and displays it (step S106). After that, the process returns to the flowchart in Figure 20.
[0172] Furthermore, if the display mode of the virtual camera presentation information 20 is not determined to be normal in step S101, the virtual camera presentation information generation unit 42 generates virtual camera group presentation information 200 in which particles resembling virtual cameras are drawn (step S107). After that, the process returns to the flowchart in Figure 20. Note that the virtual camera presentation information 200d(200) shown in Figure 24 is an example of the virtual camera group presentation information generated in step S107.
[0173] Next, we will explain the flow of the camera work display process using Figure 25. Figure 25 is a flowchart showing an example of the camera work display process in Figure 20.
[0174] The virtual camera presentation information generation unit 42 obtains frame information, camera work name, and camera work frame number from the generated virtual camera presentation information 20 (step S111). The frame information includes information such as the display position and frame size of the frame.
[0175] Next, the virtual camera presentation information generation unit 42 generates camera work presentation information based on the frame information, camera work name, and camera work frame number (step S112). The camera work presentation information is, for example, the camera work information 35 shown in Figure 11.
[0176] Then, the virtual camera presentation information generation unit 42 superimposes the camera work presentation information onto the virtual camera presentation information 20 (step S113). After that, the process returns to the flowchart in Figure 20.
[0177] [1-10-4. Flow of virtual camera group audio generation process] Using Figure 26, we will explain the flow of the virtual camera group audio generation process shown in step S37 of Figure 17. Figure 26 is a flowchart showing an example of the flow of the virtual camera group audio generation process in Figure 17.
[0178] The virtual camera presentation information generation unit 42 determines whether the virtual camera audio output mode is ALL, that is, a mode in which the audio data of all virtual camera information F is mixed and output (step S121). If it is determined that the virtual camera audio output mode is ALL (step S121: Yes), the process proceeds to step S122. On the other hand, if it is not determined that the virtual camera audio output mode is ALL (step S121: No), the process proceeds to step S123.
[0179] In step S121, if it is determined that the virtual camera audio output mode is ALL, the virtual camera presentation information generation unit 42 generates audio output data by mixing all the audio frame data (audio data corresponding to video frame data) of the virtual camera information F with the audio output unit 45 (step S122). Then, the process returns to Figure 17.
[0180] On the other hand, if it is not determined in step S121 that the virtual camera audio output mode is ALL, the virtual camera presentation information generation unit 42 determines whether the virtual camera audio output mode is On Air camera, that is, a mode that outputs audio data held by the virtual camera information F of the virtual camera that is performing imaging and distribution (step S123). If it is determined that the virtual camera audio output mode is On Air camera (step S123: Yes), the process proceeds to step S124. On the other hand, if it is not determined that the virtual camera audio output mode is On Air camera (step S123: No), the process proceeds to step S125.
[0181] In step S123, if it is determined that the virtual camera audio output mode is On Air, the virtual camera presentation information generation unit 42 generates audio output data from the audio frame data of the virtual camera information F whose camera state is On Air (step S124). Then, the process returns to Figure 17.
[0182] On the other hand, if it is not determined in step S123 that the virtual camera audio output mode is an On Air camera, the virtual camera presentation information generation unit 42 determines whether the virtual camera audio output mode is a Target camera mode, that is, a mode that outputs audio data possessed by a specified specific virtual camera information F (step S125). If it is determined that the virtual camera audio output mode is a Target camera (step S125: Yes), the process proceeds to step S126. On the other hand, if it is not determined that the virtual camera audio output mode is a Target camera (step S125: No), the process proceeds to step S127.
[0183] In step S125, if it is determined that the virtual camera audio output mode is the Target camera, the virtual camera presentation information generation unit 42 generates audio output data from the audio frame data of the virtual camera information F corresponding to the specified camera number (step S126). Then, the process returns to Figure 17.
[0184] On the other hand, if in step S125 the virtual camera audio output mode is not determined to be the Target camera, the virtual camera presentation information generation unit 42 generates silent audio output data (step S127). Then, the process returns to Figure 17.
[0185] [1-11. Flow of virtual camera display information output processing] Using Figure 27, we will explain the flow of the virtual camera presentation information output process shown in step S13 of Figure 15. Figure 27 is a flowchart showing an example of the flow of the virtual camera presentation information output process in Figure 15.
[0186] The studio video display unit 44 acquires virtual camera presentation information 20 from the virtual camera presentation information generation unit 42 (step S131). Alternatively, the studio video display unit 44 may acquire virtual camera group presentation information 200 from the virtual camera presentation information generation unit 42.
[0187] The in-studio video display unit 44 generates an image to be displayed on the inner wall surface 15 from the virtual camera presentation information 20 (step S132).
[0188] The in-studio video display unit 44 outputs the video generated in step S132 to each display panel 17 (step S133). If the video is projected from projectors 28 and 29, the in-studio video display unit 44 outputs the video generated in step S132 to each projector 28 and 29. Then, the process returns to Figure 17.
[0189] [1-12. Flow of Volumetric Video Generation Process] Figure 28 will be used to explain the flow of the volumetric video generation process shown in step S14 of Figure 15. Figure 28 is a flowchart showing an example of the flow of the volumetric video generation process in Figure 15.
[0190] The volumetric video generation unit 47 acquires video data (actual camera video I) captured by the camera 16 from the volumetric video shooting unit 46 (step S141).
[0191] The volumetric video generation unit 47 performs a modeling process to generate a 3D model 22M of the subject 22 based on the video data acquired in step S141 (step S142).
[0192] The volumetric video generation unit 47 obtains virtual camera position information Fa from the virtual camera presentation information generation unit 42 (step S143).
[0193] The volumetric video generation unit 47 performs rendering processing of a volumetric video of the 3D model 22M as viewed from a virtual viewpoint, based on the virtual camera position information Fa (step S144).
[0194] The volumetric video generation unit 47 calculates the depth, or distance, from the virtual viewpoint to the 3D model 22M based on the virtual camera position information Fa (step S145).
[0195] The volumetric video generation unit 47 outputs volumetric video data (RGB-D) to the volumetric video / CG superimposition / audio MUX unit 51 (step S146). The volumetric video data includes color information (RGB) and distance information (D). After that, the process returns to the main routine (Figure 15).
[0196] [1-13. Flow of superimposing volumetric video and background video] Using Figure 29, we will explain the process of superimposing volumetric video and background video as shown in step S15 of Figure 15. Figure 29 is a flowchart showing an example of the process of superimposing volumetric video and background video as shown in Figure 15.
[0197] The Volumetric video / CG superimposition / audio MUX unit 51 acquires Volumetric video data from the Volumetric video generation unit 47 (step S151).
[0198] The Volumetric video / CG superimposition / audio MUX unit 51 acquires background CG data from the CG background generation unit 50 (step S152).
[0199] The Volumetric Video / CG Overlay / Audio MUX Unit 51 renders background CG data in 3D (step S153).
[0200] The Volumetric Video / CG Overlay / Audio MUX Unit 51 overlays Volumetric video onto a 3D space on which background CG data has been drawn (step S154).
[0201] The Volumetric Video / CG Overlay / Audio MUX Unit 51 generates a 2D image of the 3D space created in step S154, viewed from a virtual viewpoint (step S155). If the user's viewing device 53a is capable of displaying 3D images, the Volumetric Video / CG Overlay / Audio MUX Unit 51 generates a 3D image.
[0202] The Volumetric Video / CG Overlay / Audio MUX Unit 51 outputs the 2D video (or 3D video) generated in step S155 to the distribution unit 52 (step S156). After that, it returns to the main routine (Figure 15).
[0203] Although not shown in the flowchart in Figure 29, the Volumetric Video / CG Overlay / Audio MUX Unit 51 also performs the process of multiplexing (MUXing) the generated 2D video (or 3D video) and audio information.
[0204] [1-14. Effects of the First Embodiment] As described above, the video processing device 12a (information processing device) of the first embodiment includes a Volumetric video shooting unit 46 (first acquisition unit) that acquires multiple real images (real camera images I) captured by multiple cameras 16 (first imaging devices) arranged around the subject 22, a Volumetric video generation unit 47 (generation unit) that generates a 3D model 22M of the subject 22 from the multiple real images, and a studio video display unit 44 (presentation unit) that presents information related to a virtual viewpoint when rendering the 3D model 22M into an image in a form corresponding to the viewing device 53a to the subject 22.
[0205] This allows Volumetric Studio 14a to recreate a situation where a cameraman is directly shooting with an actual camera. Therefore, the subject 22 can perform in a way that is aware of the virtual camera, thereby enhancing the sense of realism in the streamed content.
[0206] Furthermore, the video processing device 12a (information processing device) of the first embodiment further includes a virtual camera information generation unit 41 (second acquisition unit) that acquires information related to a virtual viewpoint.
[0207] This makes it possible to reliably and easily acquire information related to virtual cameras.
[0208] Furthermore, in the video processing device 12a (information processing device) of the first embodiment, the in-studio video display unit 44 (presentation unit) presents the position of the virtual viewpoint to the subject 22 (for example, virtual camera presentation information 20a, 20b).
[0209] This allows us to recreate the situation as if a photographer were actually taking pictures with a real camera.
[0210] Furthermore, in the video processing device 12a (information processing device) of the first embodiment, the in-studio video display unit 44 (presentation unit) displays information indicating that a virtual viewpoint exists at the virtual viewpoint location.
[0211] This allows subject 22 to intuitively understand the position of the virtual camera.
[0212] Furthermore, in the video processing device 12a (information processing device) of the first embodiment, the in-studio video display unit 44 (presentation unit) presents information indicating the position of a virtual viewpoint to the subject 22 (for example, virtual camera presentation information 20i, 20j, 20k, 20l).
[0213] This makes it possible to display the position of a virtual viewpoint even in studios where display panels 17 and projectors 28 and 29 cannot be installed.
[0214] Furthermore, in the video processing device 12a (information processing device) of the first embodiment, the in-studio video display unit 44 (presentation unit) presents the distance between the virtual viewpoint and the subject 22 to the subject 22 (for example, virtual camera presentation information 20f).
[0215] This allows subject 22 to intuitively understand the distance between itself and the virtual camera.
[0216] Furthermore, in the video processing device 12a (information processing device) of the first embodiment, the in-studio video display unit 44 (presentation unit) presents the observation direction from a virtual viewpoint to the subject 22 (for example, virtual camera presentation information 20g).
[0217] This allows subject 22 to intuitively understand the orientation of the virtual camera.
[0218] Furthermore, in the video processing device 12a (information processing device) of the first embodiment, the in-studio video display unit 44 (presentation unit) presents the direction of movement of the virtual viewpoint to the subject 22 (for example, virtual camera presentation information 20m).
[0219] This allows for unique volumetric camera work that is not possible with actual cameras, while simultaneously communicating the position of the virtual camera to the subject 22.
[0220] Furthermore, in the video processing device 12a (information processing device) of the first embodiment, the in-studio video display unit 44 (presentation unit) presents the operating status of a virtual camera placed at a virtual viewpoint to the subject 22 (for example, virtual camera presentation information 20h).
[0221] This allows subject 22 to intuitively understand the operating status of the virtual camera.
[0222] Furthermore, in the video processing device 12a (information processing device) of the first embodiment, the in-studio video display unit 44 (presentation unit) presents a message from the operator controlling the virtual viewpoint to the subject 22 (for example, virtual camera presentation information 20c).
[0223] This allows the subject 22 to perform while communicating with the operator controlling the virtual viewpoint.
[0224] Furthermore, in the video processing device 12a (information processing device) of the first embodiment, the in-studio video display unit 44 (presentation unit) synthesizes information related to multiple virtual viewpoints and presents it to the subject 22 when the positions of multiple virtual viewpoints approach each other (for example, virtual camera presentation information 20n3).
[0225] (2. Second Embodiment) [2-1. Schematic Configuration of the Video Processing System in the Second Embodiment] Next, a second embodiment of the video processing system 10b of this disclosure will be described using Figure 30. Figure 30 is a system configuration diagram showing an overview of the video processing system of the second embodiment.
[0226] The video processing system 10b has almost the same functions as the video processing system 10a described above, but differs in that it captures background data on which volumetric video data is superimposed using a real camera, and the position of the real camera capturing the background data is set as a virtual viewpoint. Below, the schematic configuration of the video processing system 10b will be explained using Figure 30. Note that the components common to the video processing system 10a will not be explained.
[0227] The video processing system 10b comprises a volumetric studio 14a, a 2D shooting studio 14b, and a video processing device 12b.
[0228] The 2D shooting studio 14b is a different studio from the Volumetric studio 14a. The 2D shooting studio 14b is equipped with multiple real cameras 60. Each real camera 60 can have its position, observation direction, field of view, etc., changed by operation by a cameraman or by an external control signal. In addition, an arbitrary background is drawn on the wall of the 2D shooting studio 14b, or an arbitrary background is projected by a projector or the like. Furthermore, the interior of the 2D shooting studio 14b is equipped with multiple lighting devices whose lighting state can be arbitrarily controlled. In the 2D shooting studio 14b, the 2D real image J captured by the real cameras 60 is input to the image processing device 12b. Note that the real cameras 60 are an example of the second imaging device in this disclosure.
[0229] The video processing device 12b generates a 3D model 22M of the subject 22 based on the actual camera image I acquired from the camera 16. The video processing device 12a also assumes that the actual camera 60 is in a virtual viewpoint and renders the 3D model 22M of the subject 22 as seen from that virtual viewpoint into an image in a format appropriate to the user's viewing device 53a. Furthermore, the video processing device 12a generates virtual camera presentation information 20 related to the virtual viewpoint based on information related to the actual camera 60 and outputs it to the display panel 17.
[0230] Furthermore, the video processing device 12b acquires a 2D real video J from the real camera 60. The video processing device 12b also superimposes a volumetric video 24 based on a 3D model 22M onto the acquired 2D real video J as a background video 26b. The generated video is then distributed, for example, to the user's viewing environment. Note that the video processing device 12b is an example of an information processing device in this disclosure.
[0231] [2-2. Functional Configuration of the Video Processing System in the Second Embodiment] Next, the functional configuration of the video processing system 10b will be explained using Figure 31. Figure 31 is a functional block diagram showing an example of the functional configuration of the video processing system of the second embodiment.
[0232] As shown in Figure 31, the video processing system 10b comprises a video processing device 12b, a camera 16 and a display panel 17 that constitute the imaging display device 13, and a real camera 60. The video processing system 10b also includes peripheral devices such as a remote control 54, an intercom 55, a microphone 56, a speaker 57, and a viewing device 53a.
[0233] The video processing device 12b comprises a virtual camera presentation information generation unit 42, a UI unit 43, a studio video display unit 44, an audio output unit 45, a volumetric video shooting unit 46, a volumetric video generation unit 47, a master audio output unit 48, an audio recording unit 49, a distribution unit 52, a virtual camera information acquisition unit 62, a virtual camera information transmission unit 63, a 2D video shooting unit 64, a virtual camera information receiving unit 65, a volumetric video / audio transmission unit 66, a volumetric video / audio receiving unit 67, and a volumetric video / 2D video superposition / audio MUX unit 68. These functional units are realized by the CPU of the video processing device 12b, which has a computer configuration, executing an unillustrated control program that controls the operation of the video processing device 12b. Alternatively, all or some of the functions of the video processing device 12b may be realized by hardware.
[0234] Of the functional components mentioned above, those listed to the left of the dotted line L1 in Figure 31 are installed in the Volumetric studio 14a. Those listed to the right of the dotted line L1 are installed in the 2D shooting studio 14b. Below, we will describe the functions of each functional component, but only those that are different from the video processing system 10a.
[0235] The virtual camera information acquisition unit 62 acquires information related to the actual camera 60 (second imaging device) on the 2D shooting studio 14b side. The information related to the actual camera 60 is the virtual camera information F when the actual camera 60 is considered as a virtual camera. The contents of the virtual camera information F are as described in the first embodiment. Note that the virtual camera information acquisition unit 62 is an example of the second acquisition unit in this disclosure.
[0236] The virtual camera information transmission unit 63 transmits the virtual camera information F acquired by the virtual camera information acquisition unit 62 to the Volumetric studio 14a.
[0237] The virtual camera information receiving unit 65 receives virtual camera information F from the 2D shooting studio 14b.
[0238] The 2D video shooting unit 64 generates a background 2D video from the 2D real video J captured by the real camera 60.
[0239] The Volumetric video / audio transmission unit 66 transmits the Volumetric video and audio data generated in the Volumetric studio 14a to the 2D shooting studio 14b.
[0240] The Volumetric video / audio receiver 67 receives Volumetric video and audio data from the Volumetric studio 14a.
[0241] The Volumetric Video / 2D Video Overlay / Audio MUX Unit 68 renders a 3D model 22M of the subject 22 into an image in a form appropriate to the user's viewing device 53a, and overlays it onto an image captured by a real camera 60 (second imaging device) located in a different location from the subject 22. The Volumetric Video / 2D Video Overlay / Audio MUX Unit 68 also multiplexes (MUX) the overlaid image with audio data. Note that the Volumetric Video / 2D Video Overlay / Audio MUX Unit 68 is an example of an overlay unit in this disclosure.
[0242] Furthermore, the video processing system 10b does not have the controller 40 (see Figure 13) that the video processing system 10a has. This is because, in the video processing system 10b, the actual camera 60 itself generates information related to the virtual camera. Specifically, the actual camera 60 has a gyro sensor and an accelerometer. The actual camera 60 detects its own shooting direction and movement direction by detecting the output of the gyro sensor and accelerometer.
[0243] Furthermore, the 2D shooting studio 14b where the actual camera 60 is placed is equipped with a position detection sensor (not shown) that measures the position of the actual camera 60 within the 2D shooting studio 14b. The position detection sensor consists of multiple base stations installed in the 2D shooting studio 14b that transmit IR signals with different light emission patterns, and an IR sensor installed on the actual camera 60 that detects the IR signals from the base stations. The IR sensor detects its own position in the 2D shooting studio 14b based on the intensity of the multiple IR signals it has detected. The actual camera 60 may also detect its own position and orientation in the 2D shooting studio 14b based on the image it has captured. In this way, the actual camera 60 generates information related to the virtual camera based on the information acquired by the various sensors.
[0244] Furthermore, the actual camera 60 includes an operating device such as a selection button for instructing the selection and start of camera work information, and a display device for displaying the camera work information options.
[0245] In Figure 31, the following components of the video processing device 12b are installed in the 2D shooting studio 14b where the actual camera 60 is located: the virtual camera information acquisition unit 62, the virtual camera information transmission unit 63, the 2D video shooting unit 64, the volumetric video / audio receiving unit 67, the volumetric video / 2D video superposition / audio MUX unit 68, and the distribution unit 52. The other functional components of the video processing device 12b are installed in the volumetric studio 14a.
[0246] [2-3. Operation of the Video Processing System of the Second Embodiment] The flow of processing performed by the video processing system 10b is the same as the flow of processing performed by the aforementioned video processing system 10a. Therefore, a detailed description of the processing flow is omitted.
[0247] In the video processing system 10a, the background CG video needed to have 3D information. However, in the video processing system 10b, virtual camera information F corresponding to the movement of the real camera 60 is generated for each frame. Then, the video processing device 12b generates a Volumetric video corresponding to the virtual camera information F and superimposes it on the background 2D video based on the 2D real video J captured by the real camera 60. Therefore, it is not necessary to prepare 3D background data (background CG video) as in the video processing system 10a.
[0248] Also, the video processing system 10b has characteristics different from those of virtual production, which is known as a system that generates a video as if it were taken at a target location. That is, in well-known virtual production, 3D CG is drawn on the background in accordance with the movement of the real camera, and a subject standing in front of it is photographed. In contrast, in the video processing system 10b, a Volumetric video of the subject 22 performing in accordance with the movement of the real camera 60 that photographs the actual background prepared in the 2D shooting studio 14b is generated. Therefore, the positioning of the subject and the background is opposite to that of well-known virtual production. Therefore, by using the video processing system 10b, the application range of current virtual production can be expanded.
[0249] [2-4. Effects of the Second Embodiment] As described above, the video processing apparatus 12b (information processing apparatus) of the second embodiment renders the 3D model 22M of the subject 22 into an image in a form corresponding to the viewing device 53a, and superimposes it on an image captured by a real camera 60 (second imaging device) located at a location different from the subject 22. The Volumetric video·2D video superimposition / audio MUX unit 68 (superimposition unit) is further provided. The virtual camera information acquisition unit 62 (second acquisition unit) regards the real camera 60 as a virtual camera placed at a virtual viewpoint, and acquires information related to the virtual viewpoint from the real camera 60.
[0250] As a result, when the real camera 60 installed at a distant location is regarded as a virtual camera, in the Volumetric studio 14a, it is possible to reproduce a situation as if a cameraman is directly shooting with an actual camera. Therefore, the subject 22 can perform a performance while being conscious of the virtual camera, so that the sense of presence of the distribution content can be further enhanced.
[0251] (3. Third Embodiment) [3-1. Schematic Configuration of the Video Processing System of the Third Embodiment] Next, the video processing system 10c, which is the third embodiment of the present disclosure, will be described with reference to FIG. 32. FIG. 32 is a system configuration diagram showing the outline of the video processing system of the third embodiment.
[0252] The video processing system 10c has substantially the same functions as the above-described video processing systems 10a and 10b. However, while the video processing systems 10a and 10b distributed the generated distribution content to the user's viewing device 53a unidirectionally, the video processing system 10c is different in that the user can interactively control the position of the virtual viewpoint using the viewing device 53b. Hereinafter, the schematic configuration of the video processing system 10c will be described with reference to FIG. 32. Note that descriptions of components common to the video processing systems 10a and 10b will be omitted.
[0253] The video processing system 10c comprises a volumetric studio 14a, a video processing device 12c, and a viewing device 53b. The video processing device 12c may be installed in the volumetric studio 14a.
[0254] The video processing device 12c generates a 3D model 22M of the subject 22 based on the actual camera image I acquired from the camera 16. The video processing device 12c also acquires virtual camera information F from the user's viewing device 53b. The video processing device 12c then renders the 3D model 22M of the subject 22 as seen from a virtual viewpoint based on the virtual camera information F into an image in a format appropriate to the user's viewing device 53b. The video processing device 12c also generates virtual camera presentation information 20 related to the virtual viewpoint and outputs it to the display panel 17. Here, the information related to the virtual viewpoint is information related to the viewpoint from which each of the multiple viewing users views the image rendered by the video processing device 12c on their own viewing device 53b.
[0255] Furthermore, the video processing device 12c superimposes a volumetric video 24 based on the generated 3D model 22M onto the acquired background video 26a to generate a video observed from a set virtual viewpoint. The video processing device 12c then distributes the generated video to the user's viewing device 53b. Note that the video processing device 12c is an example of an information processing device in this disclosure.
[0256] [3-2. Functional configuration of the video processing system of the third embodiment] Next, the functional configuration of the video processing system 10c will be explained using Figure 33. Figure 33 is a functional block diagram showing an example of the functional configuration of the video processing system of the third embodiment.
[0257] As shown in Figure 33, the video processing system 10c comprises a video processing device 12c, a viewing device 53b, and a camera 16 and a display panel 17 that constitute the image capture display device 13. The video processing system 10c also includes peripheral devices: a remote control 54, an intercom 55, a microphone 56, and a speaker 57.
[0258] The video processing device 12c comprises a virtual camera presentation information generation unit 42, a UI unit 43, a studio video display unit 44, an audio output unit 45, a volumetric video shooting unit 46, a volumetric video generation unit 47, a master audio output unit 48, an audio recording unit 49, a CG background generation unit 50, a volumetric video / CG superimposition / audio MUX unit 51, a distribution unit 52, a virtual camera information acquisition unit 62, a virtual camera information transmission unit 63, a virtual camera information reception unit 65, a distribution reception unit 70, a volumetric video output unit 71, and an audio output unit 72. These functional units are realized by the CPU of the video processing device 12c, which has a computer configuration, executing an unillustrated control program that controls the operation of the video processing device 12c. Alternatively, all or some of the functions of the video processing device 12c may be realized by hardware.
[0259] Of the functional components described above, those listed to the left of the dotted line L2 in Figure 33 are installed in the Volumetric studio 14a. The functional components listed to the right of the dotted line L2 are installed in the user environment holding the viewing device, and preferably are built into the viewing device 53b. Below, only the functional components that are different from the video processing systems 10a and 10b will be described, focusing on the functions of each functional component.
[0260] The virtual camera information acquisition unit 62 acquires virtual camera information F from the viewing device 53b, which includes virtual camera location information and user video, messages, etc.
[0261] The virtual camera information transmission unit 63 transmits the virtual camera information F acquired by the virtual camera information acquisition unit 62 to the Volumetric studio 14a.
[0262] The virtual camera information receiving unit 65 receives virtual camera information F from the virtual camera information transmitting unit 63.
[0263] The distribution receiving unit 70 receives the distribution content transmitted from the Volumetric studio 14a. The content received by the distribution receiving unit 70 is different from the content viewed by the user; it is simply a multiplexed combination of Volumetric video, background CG, and audio data.
[0264] The Volumetric video output unit 71 decodes the Volumetric video and background CG from the multiplexed signal received by the distribution receiving unit 70. The Volumetric video output unit 71 also renders a Volumetric video of the 3D model 22M of the subject 22 as seen from an observation position based on the virtual camera position information Fa. The Volumetric video output unit 71 then superimposes the rendered Volumetric video onto the background CG data. Finally, the Volumetric video output unit 71 outputs the video with the superimposed background CG data to the viewing device 53b.
[0265] The audio output unit 72 decodes audio data from the multiplexed signal received by the distribution receiving unit 70. The audio output unit 72 then outputs the decoded audio data to the listening device 53.
[0266] The Volumetric Video / CG Overlay / Audio MUX Unit 51 multiplexes (MUX) Volumetric video, background CG, and audio data. However, unlike the Volumetric Video / CG Overlay / Audio MUX Unit 51 (see Figure 13) of the video processing device 12a, the overlay of Volumetric video and background CG is performed by the Volumetric Video Output Unit 71, so only signal multiplexing (MUX) is performed here.
[0267] Note that the viewing device 53b has the functions of the controller 40 in the video processing device 12a. The viewing device 53b is, for example, a mobile terminal such as a smartphone or a tablet terminal, an HMD, a spatial reproduction display capable of naked-eye stereoscopy, or a combination of a display and a game controller. Note that the viewing device 53b has at least a function of specifying a position and a direction, a function of selecting and determining menu contents, and a function of communicating with the video processing device 12c.
[0268] By having these functions, the viewing device 53b sets the position and direction necessary for setting a virtual viewpoint, similarly to the controller 40. That is, the viewing device 53b itself serves as a virtual camera. Also, the viewing device 53b selects and determines the camera work of the virtual viewpoint (virtual camera). Further, the viewing device 53b selects and determines a message for the subject 22.
[0269] [3-3. Method for obtaining virtual camera information] A method for obtaining virtual camera information F from a mobile terminal 80, which is an example of the viewing device 53b, will be described using FIGS. 34 and 35. FIG. 34 is a diagram showing a method for a user to set camera work information using a viewing device. FIG. 35 is a diagram showing a method for a user to set an operator video, an operator voice, and an operator message using a viewing device.
[0270] In the mobile terminal 80, which is an example of the viewing device 53b, when a camera work setting menu is selected from a main menu (not shown) displayed when an application using the video processing system 10c is launched, a camera work selection button 74 shown in FIG. 34 is displayed on the display screen of the viewing device 53b. Note that the display screen of the mobile terminal 80 also has the function of a touch panel, and the GUI (Graphical User Interface) displayed on the display screen can be controlled using a finger.
[0271] The camera work selection button 74 is the button to press when starting the camera work settings.
[0272] When the camera work selection button 74 is pressed, the camera work selection window 75 is displayed on the mobile terminal 80's screen. The camera work selection window 75 displays a list of pre-set camera works. Additionally, the camera work start button 76 is displayed superimposed on any camera work selected in the camera work selection window 75.
[0273] The user of the mobile terminal 80 places the camera work start button 76 over the type of camera work they wish to set. By pressing the camera work start button 76, the camera work setting is completed. The set camera work is sent to the virtual camera information acquisition unit 62 as camera work information Fb.
[0274] Although not shown in Figure 34, the camera work settings menu also allows you to set the start and end positions of the camera work, as well as the camera work speed.
[0275] Furthermore, when an application using the video processing system 10c is launched on the mobile terminal 80, selecting the operator message setting menu from the main menu (not shown) displays the message selection button 77 shown in Figure 35 on the display screen of the mobile terminal 80.
[0276] Message selection button 77 is the button to press when starting the selection of an operator message.
[0277] When the message selection button 77 is pressed, the message selection window 78 is displayed on the mobile terminal 80's screen. The message selection window 78 displays a list of pre-set messages. Additionally, the message send button 79 is displayed superimposed on any message displayed in the message selection window 78.
[0278] The user of the mobile terminal 80 places the message send button 79 over the message they wish to set. By pressing the message send button 79, the setting of the operator message Fe is completed. The set operator message Fe is sent to the virtual camera information acquisition unit 62.
[0279] In addition to the preset messages, images and audio of the operator acquired using the IN camera 81 and microphone 82 built into the viewing device 53b may also be set as operator messages Fe.
[0280] Furthermore, the mobile terminal 80 detects virtual camera position information Fa, which indicates its own shooting direction and direction of movement, by detecting the output of the gyro sensor and accelerometer. This is the same method used by the actual camera 60 to detect virtual camera position information Fa in the second embodiment, so further explanation is omitted.
[0281] [3-4. Format of information presented by virtual camera groups] Figures 36, 37, and 38 illustrate the form of the virtual camera group information 200 presented by the video processing system 10c. Figure 36 shows an example of virtual camera group information according to the number of viewing users. Figure 37 shows an example of virtual camera group information when a viewing user changes their observation position. Figure 38 shows an example of a function that enables communication between viewing users and performers.
[0282] The video processing system 10c allows multiple users to freely set the position of their virtual viewpoint using their respective viewing devices 53b. Consequently, situations arise where the virtual viewpoints of many users are close together. Figure 36 shows an example of virtual camera group presentation information 200 presented in such a case.
[0283] The horizontal axis in Figure 36 shows the number of viewers from a specific location. The left side indicates fewer viewers, and the right side indicates more viewers.
[0284] For example, the virtual camera group information 200e, 200f, and 200g divide a single frame 21 and display human-shaped icons (corresponding to the cameraman icon 32 in Figure 7) in each divided area to indicate the presence of viewing users. This display format allows us to show how many users are viewing from the location where the virtual camera group information 200 is displayed. One human-shaped icon may represent one viewing user, or one human-shaped icon may be associated with a predetermined number of users. In this way, the virtual camera group information 200e, 200f, and 200g indicate the density of users viewing from a specific location. In the virtual camera group information 200g, the large size of one human-shaped icon indicates that several users are viewing from a position close to the subject 22. Furthermore, as will be described later, the human-shaped icons may also be enlarged according to other criteria (see Figure 38).
[0285] The number of people (10026) displayed at the top of the virtual camera group information 200e, 200f, and 200g indicates the current total number of viewing users. Alternatively, instead of displaying the current total number of viewing users, the number of viewing users viewing from the direction in which the virtual camera group information 200 is displayed may be displayed.
[0286] Furthermore, the method of displaying the number of viewers is not limited to this; a presentation format that allows viewers to intuitively understand the density of viewers may also be used, such as particle displays like virtual camera group information 200h, 200i, 200j.
[0287] Figure 37 shows an example of how the virtual camera group presentation information 200 changes when the viewer changes their virtual viewpoint.
[0288] Figure 37 shows the state at time t0 where virtual camera group information 200k and 200l are presented. Figure 37 also shows the state at time t1 where one or more viewing users U, displayed in virtual camera group information 200k, have changed their virtual viewpoint position. Furthermore, Figure 37 shows that at time t2, the virtual viewpoint position of viewing user U has reached the position where virtual camera group information 200l is presented.
[0289] At this time t1, the virtual camera group information 200k is changed to virtual camera group information 200m, in which the human-shaped icon corresponding to viewer user U is removed. Then, virtual camera information 20r corresponding to viewer user U is newly presented.
[0290] Furthermore, at time t2, the virtual camera presentation information 20r corresponding to viewer user U is deleted. Then, the virtual camera group presentation information 200l is changed to virtual camera group presentation information 200n, to which a human-shaped icon corresponding to viewer user U is added.
[0291] Furthermore, the virtual camera information 20r(20) corresponding to the viewing user U may be displayed in a simplified form, as shown in the lower part of Figure 37, as the virtual camera information 20s(20).
[0292] Figure 38 shows an example in which a viewer communicates with the subject 22 in the video processing system 10c.
[0293] The virtual camera group display information 200p(200) shown in Figure 38 is an example in which, when an operator message is sent from a specific viewing user, message information 37 is displayed in the split display frame for the corresponding user in the virtual camera display information 20r(20).
[0294] Furthermore, if subject 22 wishes to communicate with a specific viewer, subject 22 turns on the cursor display by providing operation information for its remote control 54 to the UI unit 43. When the cursor display is turned on, the cursor 90 is displayed superimposed on the virtual camera group presentation information 200q(200), as shown in Figure 38. Subject 22 moves the position of the displayed cursor 90 to the position of the viewer it wishes to communicate with by operating the remote control 54, and selects that viewer. Alternatively, it specifies the Target camera number with which it wishes to communicate.
[0295] Furthermore, subject 22 turns on communication mode. By turning on communication mode, the split display frame of the selected viewer is enlarged, and the virtual camera group presentation information 200r(200) shown in Figure 38 is displayed. Note that communication mode may always be left ON as the default setting of the video processing system 10c. In this case, when subject 22 selects a viewer with the cursor 90, it becomes possible to communicate with that viewer immediately. In this way, subject 22 can select any viewer through the operation of the UI unit 43, which is an example of a selection unit, and communicate with the selected viewer.
[0296] The virtual camera group information 200r(200) displays an enlarged image of the user, allowing the subject 22 to make eye contact with the selected user. At the same time, the subject 22 can hear the user's message through the intercom 55. This communication function can also be similarly implemented in the aforementioned video processing systems 10a and 10b.
[0297] The specific viewing users referred to here are, for example, high-priority users such as paid users or premium users. In other words, the viewing device 53b (virtual camera) of a high-priority user has a high camera priority in the camera information Ff (see Figure 14) described in the first embodiment. And high-priority users can communicate with the subject 22 with priority.
[0298] Furthermore, the video processing device 12c may also enable another viewer to view the communication between the subject 22 and a specific viewer, for example, as shown in Figure 38, the virtual camera group presentation information 200r(200) being displayed behind the subject 22.
[0299] [3-5. Processing flow of the video processing system of the third embodiment] Figures 39 and 40 illustrate the processing flow performed by the video processing system 10c. Figure 39 is a flowchart illustrating an example of the processing flow performed by the video processing system of the third embodiment. Figure 40 is a flowchart illustrating an example of the communication video / audio generation processing flow in Figure 39.
[0300] The virtual camera presentation information generation unit 42 performs the virtual camera presentation information generation process (step S161). The flow of the virtual camera presentation information generation process is as shown in Figure 17.
[0301] The UI unit 43 determines whether the cursor display is ON (step S162). If it is determined that the cursor display is ON (step S162: Yes), the process proceeds to step S164. On the other hand, if it is determined that the cursor display is not ON (step S162: No), the process proceeds to step S163.
[0302] In step S162, if it is determined that the cursor display is ON, the UI unit 43 generates an image of the cursor 90 (step S164). Then, the process proceeds to step S163.
[0303] On the other hand, if it is not determined in step S162 that the cursor display is ON, or after step S164 is executed, the UI unit 43 determines whether the communication mode is ON (step S163). If it is determined that the communication mode is ON (step S163: Yes), the process proceeds to step S166. On the other hand, if it is not determined that the communication mode is ON (step S163: No), the process proceeds to step S165.
[0304] In step S163, if it is determined that the communication mode is ON, the virtual camera presentation information generation unit 42 performs communication video / audio generation processing (step S166). Then, the process proceeds to step S165. Details of the video / audio generation processing are shown in Figure 40.
[0305] On the other hand, if it is not determined in step S163 that the communication mode is ON, or after step S166 is executed, the virtual camera presentation information generation unit 42 superimposes the communication video / audio and the cursor 90 video onto the virtual camera video / audio (step S165).
[0306] Next, the virtual camera presentation information generation unit 42 outputs virtual camera presentation information 20 (or virtual camera group presentation information 200) to the studio video display unit 44 and the audio output unit 45 (step S167). After that, the virtual camera presentation information generation unit 42 completes the process shown in Figure 39.
[0307] Next, Figure 40 will be used to explain the details of the video / audio generation process performed in step S166.
[0308] The virtual camera presentation information generation unit 42 acquires virtual camera presentation information 20 (or virtual camera group presentation information 200) corresponding to the virtual camera number of the communication target (step S171).
[0309] The virtual camera presentation information generation unit 42 generates communication video / audio from frame information, video frame data, audio frame data, and messages (step S172). After that, it returns to the main routine (Figure 39).
[0310] [3-6. Effects of the Third Embodiment] As described above, in the video processing device 12c (information processing device) of the third embodiment, the information relating to the virtual viewpoint is information relating to the viewpoint when each of the multiple viewing users views the rendered image on the viewing device 53b.
[0311] This allows for the delivery of images tailored to each viewer's perspective to multiple users.
[0312] Furthermore, in the video processing device 12c (information processing device) of the third embodiment, the in-studio video display unit 44 (presentation unit) presents information related to multiple virtual viewpoints to the subject 22 by arranging it within the divided picture frame 21.
[0313] This allows the subject 22 to grasp the approximate number of viewers viewing from a specific direction.
[0314] Furthermore, the video processing device 12c (information processing device) of the third embodiment is further equipped with a UI unit 43 (selection unit) that acquires operation information of the subject 22 and selects a viewing device 53b (virtual camera) placed in a virtual viewpoint, and the subject 22 communicates with the operator of the viewing device 53b selected by the UI unit 43.
[0315] This allows the subject 22 to communicate with any viewer.
[0316] The effects described herein are merely illustrative and not limiting, and other effects may also occur. Furthermore, the embodiments of this disclosure are not limited to those described above, and various modifications are possible without departing from the spirit of this disclosure.
[0317] For example, this disclosure can also take the following form:
[0318] (1) A first acquisition unit that acquires multiple real images captured by multiple first imaging devices arranged around the subject, A generation unit that generates a 3D model of the subject from the aforementioned multiple real images, A display unit presents information related to a virtual viewpoint when rendering the 3D model into an image in a format appropriate to the viewing device to the subject, An information processing device equipped with the following features. (2) The system further includes a second acquisition unit for acquiring information related to the virtual viewpoint. The information processing device described in (1) above. (3) The system further includes a superposition section that renders the aforementioned 3D model into an image in a form appropriate to the viewing device and superimposes it onto an image captured by a second imaging device located in a different location from the subject. The second acquisition unit considers the second imaging device as a virtual camera placed at a virtual viewpoint and acquires information related to the virtual viewpoint from the second imaging device. The information processing device described in (2) above. (4) The information relating to the virtual viewpoint is information relating to the viewpoint of each of the multiple viewing users when viewing the rendered image on their viewing device. The information processing device described in any one of (1) to (3) above. (5) The aforementioned display unit is, The position of the virtual viewpoint is presented to the subject. The information processing device described in any one of (1) to (4) above. (6) The aforementioned display unit is, Information indicating that a virtual viewpoint exists at the location of the virtual viewpoint is presented. The information processing device described in any one of (1) to (5) above. (7) The aforementioned display unit is, Information indicating the location of the aforementioned virtual viewpoint is presented to the subject. The information processing device described in any one of (1) to (6) above. (8) The aforementioned display unit is, The distance between the virtual viewpoint and the subject is presented to the subject. The information processing device described in any one of (1) to (7) above. (9) The aforementioned display unit is, The observation direction from the aforementioned virtual viewpoint is presented to the subject. The information processing device described in any one of (1) to (8) above. (10) The aforementioned display unit is, The direction of movement of the virtual viewpoint is presented to the subject. The information processing device described in any one of (1) to (9) above. (11) The aforementioned display unit is, The operating state of the virtual camera placed at the virtual viewpoint is presented to the subject. The information processing device described in any one of (1) to (10) above. (12) The aforementioned display unit is, The operator controlling the virtual viewpoint presents a message to the subject. The information processing device described in any one of (1) to (11) above. (13) The aforementioned display unit is, When the positions of multiple virtual viewpoints are close together, the information related to those multiple virtual viewpoints is synthesized and presented to the subject. The information processing device described in any one of (1) to (12) above. (14) The aforementioned display unit is, The information relating to the multiple virtual viewpoints is arranged within the divided frames and presented to the subject. The information processing device described in (13) above. (15) The system further includes a selection unit that acquires the operation information of the subject and selects a virtual camera placed in the virtual viewpoint, The subject communicates with the operator of the virtual camera selected by the selection unit. The information processing device described in any one of (1) to (14) above. (16) Computers, A first acquisition unit that acquires multiple real images captured by multiple first imaging devices arranged around the subject, A generation unit that generates a 3D model of the subject from the aforementioned multiple real images, A display unit presents information related to a virtual viewpoint when rendering the 3D model into an image in a format appropriate to the viewing device to the subject, A program that makes something work. [Explanation of Symbols]
[0319] 10a, 10b, 10c... Video processing system, 12a, 12b, 12c... Video processing device (information processing device), 13, 13a, 13b, 13c... Image display device, 14a... Volumetric studio, 14b... 2D shooting studio, 15... Inner wall surface, 16, 16a, 16b, 16c... Camera (first imaging device), 17... Display panel, 18... Transmissive screen, 19... Reflective screen, 20... Virtual camera presentation information (information related to virtual viewpoint), 21... Image frame, 22...Subject, 22M...3D model, 24...Volumetric video, 26a,26b...Background video, 28,29...Projector, 30...Camera icon, 31...Tally lamp, 32...Cameraman icon, 33...Camera name, 34...Camera position display icon, 35...Camera work information, 36...Camera work, 37...Message information, 38...Particle, 41...Virtual camera information generation unit (second acquisition unit), 43...UI unit (selection unit), 44...Studio Internal video display unit (presentation unit), 46…Volumetric video shooting unit (first acquisition unit), 47…Volumetric video generation unit (generation unit), 51…Volumetric video / CG superimposition / audio MUX unit, 53a, 53b…Viewing device, 60…Actual camera (second imaging device), 62…Virtual camera information acquisition unit (second acquisition unit), 74…Camera work selection button, 75…Camera work selection window, 76…Camera work start button, 77…Message selection button, 78…Message selection window, 79…Message send button, 80…Mobile terminal, 90…Cursor, 200…Virtual camera group presentation information, F…Virtual camera information, Fa…Virtual camera position information, Fb…Camera work information, Fc…Operator video, Fd…Operator voice, Fe…Operator message, Ff…Camera information, I…Actual camera video, J…2D actual video, M…Mesh information, Ta, Tb…Texture information, U…Viewing user
Claims
1. A first acquisition unit that acquires a plurality of real images captured by a plurality of first imaging devices arranged around a subject, A generation unit that generates a 3D model of the subject from the aforementioned multiple real images, A display unit presents information related to a virtual viewpoint when rendering the 3D model into an image in a format appropriate to the viewing device to the subject, A second acquisition unit that acquires information related to the virtual viewpoint, The system includes a superposition unit that renders the 3D model into an image in a form appropriate to the viewing device and superimposes it onto an image captured by a second imaging device located in a different location from the subject, The second acquisition unit considers the second imaging device as a virtual camera placed at a virtual viewpoint and acquires information related to the virtual viewpoint from the second imaging device. Information processing device.
2. A first acquisition unit that acquires a plurality of real images captured by a plurality of first imaging devices arranged around a subject, A generation unit that generates a 3D model of the subject from the aforementioned multiple real images, The system includes a display unit that presents information related to a virtual viewpoint when rendering the 3D model into an image in a format appropriate to the viewing device to the subject, The aforementioned display unit is, The operating state of the virtual camera placed at the virtual viewpoint is presented to the subject. Information processing device.
3. A first acquisition unit that acquires a plurality of real images captured by a plurality of first imaging devices arranged around a subject, A generation unit that generates a 3D model of the subject from the aforementioned multiple real images, The system includes a display unit that presents information related to a virtual viewpoint when rendering the 3D model into an image in a format appropriate to the viewing device to the subject, The aforementioned display unit is, When the positions of multiple virtual viewpoints approach each other, the information relating to those multiple virtual viewpoints is synthesized and presented to the subject. The aforementioned display unit is, The information relating to the multiple virtual viewpoints is arranged within the divided frame and presented to the subject. Information processing device.
4. A first acquisition unit that acquires a plurality of real images captured by a plurality of first imaging devices arranged around a subject, A generation unit that generates a 3D model of the subject from the aforementioned multiple real images, A display unit presents information related to a virtual viewpoint when rendering the 3D model into an image in a format appropriate to the viewing device to the subject, The system includes a selection unit that acquires operation information of the subject and selects a virtual camera placed in the virtual viewpoint, The subject communicates with the operator of the virtual camera selected by the selection unit. Information processing device.
5. The system further includes a second acquisition unit for acquiring information related to the virtual viewpoint. The information processing apparatus according to any one of claims 2 to 4.
6. The device further comprises a superposition section that renders the 3D model into an image in a form appropriate to the viewing device and superimposes it onto an image captured by a second imaging device located in a different location from the subject, The second acquisition unit considers the second imaging device as a virtual camera placed at a virtual viewpoint and acquires information related to the virtual viewpoint from the second imaging device. The information processing apparatus according to claim 5.
7. The display portion is, The operating state of the virtual camera placed at the virtual viewpoint is presented to the subject. The information processing apparatus according to claim 1, 3, 4, 5, or 6.
8. The aforementioned display unit is, When the positions of multiple virtual viewpoints are close together, the information related to those multiple virtual viewpoints is synthesized and presented to the subject. The information processing apparatus according to claim 1, 2, 4, 5, 6, or 7.
9. The display portion is, The information relating to the multiple virtual viewpoints is arranged within the divided frames and presented to the subject. The information processing apparatus according to claim 8.
10. Further comprising a selection unit that acquires operation information of the subject and selects a virtual camera placed in the virtual viewpoint, The subject communicates with the operator of the virtual camera selected by the selection unit. The information processing apparatus according to claim 1, 2, 3, 5, 6, 7, 8, or 9.
11. The information relating to the virtual viewpoint is information relating to the viewpoint of each of the multiple viewing users when viewing the rendered image on their viewing device. The information processing apparatus according to any one of claims 1 to 10.
12. The aforementioned display unit is, The position of the virtual viewpoint is presented to the subject. The information processing apparatus according to any one of claims 1 to 11.
13. The aforementioned display unit is, Information indicating that a virtual viewpoint exists at the location of the virtual viewpoint is presented. The information processing apparatus according to any one of claims 1 to 12.
14. The aforementioned display unit is, Information indicating the location of the aforementioned virtual viewpoint is presented to the subject. The information processing apparatus according to any one of claims 1 to 13.
15. The aforementioned display unit is, The distance between the virtual viewpoint and the subject is presented to the subject. The information processing apparatus according to any one of claims 1 to 14.
16. The aforementioned display unit is, The observation direction from the aforementioned virtual viewpoint is presented to the subject. The information processing apparatus according to any one of claims 1 to 15.
17. The aforementioned display unit is, The direction of movement of the virtual viewpoint is presented to the subject. The information processing apparatus according to any one of claims 1 to 16.
18. The aforementioned display unit is, The operator controlling the virtual viewpoint presents a message to the subject. The information processing apparatus according to any one of claims 1 to 17.
19. Computers, A first acquisition unit that acquires multiple real images captured by multiple first imaging devices arranged around the subject, A generation unit that generates a 3D model of the subject from the aforementioned multiple real images, A display unit presents information related to a virtual viewpoint when rendering the 3D model into an image in a format appropriate to the viewing device to the subject, A second acquisition unit that acquires information related to the virtual viewpoint, The 3D model is rendered into an image in a form appropriate to the viewing device and functions as an overlay section that superimposes it onto an image captured by a second imaging device located in a different location from the subject. The second acquisition unit considers the second imaging device as a virtual camera placed at a virtual viewpoint and acquires information related to the virtual viewpoint from the second imaging device. program.
20. A computer, A first acquisition unit that acquires multiple real images captured by multiple first imaging devices arranged around the subject, A generation unit that generates a 3D model of the subject from the aforementioned multiple real images, The 3D model is rendered into an image in a format appropriate to the viewing device, and the information related to the virtual viewpoint is presented to the subject as a presentation unit. The aforementioned display unit is, The operating state of the virtual camera placed at the virtual viewpoint is presented to the subject. program.
21. A computer, A first acquisition unit that acquires multiple real images captured by multiple first imaging devices arranged around the subject, A generation unit that generates a 3D model of the subject from the aforementioned multiple real images, The 3D model is rendered into an image in a format appropriate to the viewing device, and the information related to the virtual viewpoint is presented to the subject as a presentation unit. The aforementioned display unit is, When the positions of multiple virtual viewpoints approach each other, the information relating to those multiple virtual viewpoints is synthesized and presented to the subject. The aforementioned display unit is, The information relating to the multiple virtual viewpoints is arranged within the divided frame and presented to the subject. program.
22. A computer, A first acquisition unit that acquires multiple real images captured by multiple first imaging devices arranged around the subject, A generation unit that generates a 3D model of the subject from the aforementioned multiple real images, A display unit presents information related to a virtual viewpoint when rendering the 3D model into an image in a format appropriate to the viewing device to the subject, It acquires the operation information of the subject and functions as a selection unit that selects a virtual camera placed in the virtual viewpoint. The subject communicates with the operator of the virtual camera selected by the selection unit. program.