Information processing apparatus, information processing method, video distribution method, and information processing system
By distributing volumetric processing across multiple computers and employing high-performance computing and low-volume transmission, the problem of generating a wide range of demonstrations suitable for real-time video content was solved, thereby improving real-time performance and development efficiency.
Patent Information
- Application Number
- CN202180049906.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-06
- Filing Date
- 2021-07-15
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2041-07-15
AI Technical Summary
Existing technologies struggle to generate a wide range of presentations suitable for real-time video content such as live music performances, sporting events, lectures, and academic courses. Furthermore, they are difficult to guarantee real-time performance and development efficiency when processing large-volume videos on personal computers with limited resources.
By distributing the processing in volumetric technology to multiple computers, employing high-performance computing technology and low-volume data transmission methods, real-time volumetric video is generated and distributed, and wide-range demonstrations are achieved by combining the collaborative work of multiple computers.
It enables the generation of volumetric videos with a wide range of presentations from real-time video content, ensuring real-time performance, improving system development efficiency, and reducing latency.
Smart Images

Figure CN116075860B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to information processing apparatus, information processing method, video distribution method, and information processing system. Background Technology
[0002] A volumetric technique (also known as volumetric capture) has been proposed that uses multiple cameras arranged around a subject (object) to reconstruct the three-dimensional shape of the subject (object) from within and redraw the shape from a free viewpoint. By arranging the cameras to capture the back and top of the subject using this volumetric technique, video (volumetric video) can be generated that allows viewers to view the subject from all directions.
[0003] Citation List
[0004] Patent documents
[0005] Patent Document 1: WO 2019 / 021375A
[0006] summary
[0007] Technical issues
[0008] In a typical scenario of watching videos, users will see a video generated by overlaying 3D objects, created using volumetric techniques, onto a pre-created background object and rendering the combined object. However, the problem is that simply overlaying 3D objects onto a background object cannot achieve a presentation suitable for every type of video content, such as live music performances, sporting events, lectures, and academic courses.
[0009] Therefore, this disclosure provides information processing apparatus, information processing method, video distribution method, and information processing system that enable the generation of videos with a wide range of presentations from three-dimensional objects generated by volumetric techniques, etc.
[0010] Solution to the problem
[0011] To address the aforementioned problems, an information processing apparatus according to an embodiment of the present disclosure includes a first generation unit that generates a video based on a three-dimensional model of the subject generated by using multiple captured images obtained by imaging the subject, and a two-dimensional image, wherein the video simultaneously contains the subject generated from the three-dimensional model and the two-dimensional image. Attached Figure Description
[0012] Figure 1 This is a block diagram illustrating a schematic configuration of an information processing system according to one embodiment of the present disclosure.
[0013] Figure 2 This is a diagram illustrating an example of an imaging apparatus according to this embodiment.
[0014] Figure 3 This is a flowchart illustrating an example of the process performed by the information processing system according to this embodiment.
[0015] Figure 4 This is a block diagram illustrating an example of the hardware configuration of an information processing system according to this embodiment.
[0016] Figure 5 This is a block diagram illustrating a further detailed configuration example of the rendering unit according to this embodiment.
[0017] Figure 6 This is a diagram illustrating an example of an intermediate rendered video according to this embodiment.
[0018] Figure 7 This is a diagram illustrating an example of virtual viewpoint video (RGB) according to this embodiment.
[0019] Figure 8 This is a diagram illustrating an example of a virtual viewpoint video (depth) according to this embodiment.
[0020] Figure 9 This is a diagram showing an example of a real camera image according to this embodiment.
[0021] Figure 10 This is a view showing an example of an auxiliary video according to this embodiment.
[0022] Figure 11 This is a block diagram illustrating a further detailed configuration example of the initial virtual viewpoint video generation unit according to this embodiment.
[0023] Figure 12 This is a block diagram illustrating a further detailed configuration example of the final virtual viewpoint video generation unit according to this embodiment.
[0024] Figure 13 This is a diagram illustrating an example of the processing performed by the image quality enhancement unit according to this embodiment.
[0025] Figure 14 This is a diagram illustrating an example of video content distributed to a user in this embodiment.
[0026] Figure 15 This is a view showing another example of a volumetric video distributed to a user in this embodiment.
[0027] Figure 16 This is a system configuration diagram illustrating a specific example of an information processing system according to this embodiment. Detailed Implementation
[0028] Embodiments of this disclosure will now be described in detail with reference to the accompanying drawings. In each of the following embodiments, the same parts are denoted by the same reference numerals, and repeated descriptions thereof will be omitted.
[0029] The contents of this disclosure will be described in the following order.
[0030] 0. Introduction
[0031] 1. One implementation method
[0032] 1-1. Functional Configuration of the Information Processing System
[0033] 1-2. Processing flow performed by the information processing system
[0034] 1-3. Hardware Configuration of the Information Processing System
[0035] 1-4. Further details of this embodiment
[0036] 1-4-1. Further detailed configuration examples of rendering units
[0037] 1-4-2. Specific example of intermediate video rendering
[0038] 1-4-3. Further detailed configuration example of the initial virtual viewpoint video generation unit 131
[0039] 1-4-4. Further detailed configuration example of the final virtual viewpoint video generation unit 134
[0040] 1-5. Examples of demonstration according to this embodiment
[0041] 1-6. Specific examples of information processing systems
[0042] 1-7. Summary
[0043] 0. Introduction
[0044] Volumetric techniques are those that use multiple cameras arranged around a subject (object) to reconstruct the internal 3D shape of the subject (object) and redraw the shape from a free viewpoint. By arranging the cameras to capture the back and top, viewers can view the subject from all directions. Because the various types of processing in such volumetric techniques, such as capture, modeling, and rendering, require significant computational cost and long processing times, they are typically performed offline. However, leveraging the ability to perform various types of processing in volumetric techniques online in real time allows for the immediate generation of volumetric video from captured 3D objects and the distribution of the generated video to users. This leads to the need for real-time processing of various types of volumetric techniques in use cases where live performance is crucial, such as live music performances, sporting events, lectures, and academic courses. Incidentally, volumetric video can be, for example, video generated using 3D objects created by volumetric techniques.
[0045] For example, real-time execution of various types of processing in volumetric technologies can be achieved by leveraging high-performance computing (HPC) techniques used for large-scale processing in supercomputers or data centers.
[0046] Furthermore, as mentioned above, in the typical case of watching volumetric videos, people watch videos generated by overlaying 3D objects created using volumetric techniques onto a pre-created background object. However, it is not always possible to achieve a presentation suitable for every type of video, such as live music performances, sporting events, lectures, and academic courses, simply by overlaying 3D objects onto a background object.
[0047] To address this issue, the following implementation method enables the generation of videos with a wide range of demonstrations from 3D objects generated by volumetric techniques.
[0048] Furthermore, when the generation of volumetric video is performed by the limited resources of a personal computer (hereinafter referred to as a PC), there is a possibility that processing cannot keep up with the captured video data, leading to compromised real-time performance. In particular, combining various types of additional processing, such as overlaying another 3D object, image quality enhancement, and various effects, into volumetric techniques will increase the overall processing load, making it difficult to ensure real-time performance.
[0049] Furthermore, since each PC has its own suitable processing tasks depending on the specifications, aggregating various types of processing for volumetric techniques on a single PC can reduce development efficiency. For example, there may be a situation where, on the one hand, a PC running Linux (registered trademark) can perform low-latency distribution processing, in which each process is distributed to multiple graphics processing units (GPUs) (hereinafter referred to as GPU distribution processing), while on the other hand, the PC has few necessary processing libraries, resulting in poor development efficiency.
[0050] Therefore, in the following implementation, various types of processing in volumetric technology are distributed to multiple computers, thereby enabling the rapid generation of volumetric videos. For example, volumetric videos that ensure real-time performance can be generated. Furthermore, distributing processing to multiple computers increases the flexibility of the system's development environment, allowing the construction of systems that suppress the deterioration of development efficiency.
[0051] However, even when processing data distributed to multiple computers, there is an issue of increased latency unless a transmission method with low throughput is used. Therefore, in the following embodiments, a data transmission method between computers with low throughput will also be described by way of example.
[0052] 1. One implementation method
[0053] 1-1. Functional Configuration of the Information Processing System
[0054] First, refer to Figure 1 An overview of an information processing system according to one embodiment of this disclosure is provided. Figure 1 This is a block diagram illustrating a schematic configuration of the information processing system according to this embodiment.
[0055] like Figure 1 As shown, the information processing system 10 includes a data acquisition unit 11, a 3D model generation unit 12, a rendering unit 13, a sending unit 14, a receiving unit 15, and a display unit 16. Note that the display unit 16 is not required to be included in the information processing system 10.
[0056] (Data Acquisition Unit 11)
[0057] The data acquisition unit 11 acquires image data (hereinafter referred to as real camera images) used to generate a three-dimensional model of the subject 90 as the imaging target. (Note that the image data in this specification may also include video data such as moving images). For example, as Figure 2As shown, multiple viewpoint images captured by multiple real cameras 70a, 70b, 70c, 70d, 70e... (hereinafter, real cameras 70a, 70b, 70c, 70d, 70e... are also collectively referred to as real cameras 70) arranged around the subject 90 are acquired as real camera images. In this case, the multiple viewpoint images are preferably images captured simultaneously by the multiple real cameras 70. Furthermore, for example, the data acquisition unit 11 can acquire multiple real camera images from different viewpoints by moving one real camera 70 and imaging the subject 90 from multiple viewpoints. However, the invention is not limited to this, and the data acquisition unit 11 can acquire a single real camera image of the subject 90. In this case, the 3D model generation unit 12, which will be described below, can generate a three-dimensional model of the subject 90 based on a single real camera image, for example, using machine learning.
[0058] Note that the data acquisition unit 11 can perform calibration based on real camera images and acquire the internal and external parameters of each real camera 70. Furthermore, the data acquisition unit 11 can acquire, for example, multiple depth information lines indicating distances from multiple viewpoints to the subject 90.
[0059] (3D model generation unit 12)
[0060] The 3D model generation unit 12 generates a 3D model containing 3D information of the subject 90 based on a real camera image used to generate a 3D model of the subject 90. For example, the 3D model generation unit 12 can generate a 3D model of the subject 90 by sculpting the 3D shape of the subject 90 based on images from multiple viewpoints (e.g., contour images from multiple viewpoints) using a technique called a visual shell. In this case, the 3D model generation unit 12 can also perform a high-accuracy transformation on the 3D model generated using the visual shell by using multiple depth information lines indicating the distances from viewpoints at multiple locations to the subject 90. Furthermore, as described above, the 3D model generation unit 12 can generate a 3D model of the subject 90 from a single real camera image of the subject 90.
[0061] The 3D model generated by the 3D model generation unit 12 can also be defined as a motion image of the 3D model, since the model is generated in a time series on a frame-by-frame basis. Furthermore, the 3D model is generated using real camera images acquired by the real camera 70, and therefore can be defined as a real-time 3D model. The 3D model can be formed as shape information indicating the surface shape of the subject 90 expressed in the form of 3D shape mesh data, referred to as a polygon mesh, which is expressed by connections between vertices. The 3D shape mesh data includes, for example, the 3D coordinates of the vertices of the mesh and index information indicating which vertices are to be combined to form a triangular mesh. Note that the method of representing the 3D model is not limited to this, and the 3D model can be described by a technique called point cloud representation, which expresses the model through positional information formed by points.
[0062] Three-dimensional shape mesh data can be associated with information about color and pattern as texture (also known as a texture image). Texture association includes view-independent texture methods where color does not change in any viewing direction, and view-dependent texture methods where color changes depending on the viewing direction. In this embodiment, one or both of these methods may be used, and other texture methods may also be used.
[0063] (Rendering Unit 13)
[0064] For example, rendering unit 13 uses the three-dimensional shape mesh data of the projected three-dimensional model to draw the camera viewpoint (corresponding to the virtual viewpoint described below), and performs texture mapping that applies textures representing the color or pattern of the mesh to the shape of the projected mesh, thereby generating a volumetric video of the three-dimensional model. Since the viewpoint at this time is a freely set viewpoint independent of the camera position during imaging, the viewpoint is also referred to as a virtual viewpoint in this embodiment.
[0065] Texture mapping includes methods such as view-dependent methods (VD methods) that consider the user's viewing perspective and view-independent methods (VI methods) that do not consider the user's viewing perspective. The VD method changes the texture to be applied to the 3D model according to the position of the viewing perspective, and therefore has the advantage of successfully achieving higher quality rendering compared to the VI method. On the other hand, the VI method does not consider the viewpoint position during viewing, and therefore has the advantage of requiring less processing power compared to the VD method. Note that viewpoint data for viewing can be obtained such that the user's viewing position (region of interest) is detected by the user-side display device (also referred to as the user terminal) and then input from the user terminal to the rendering unit 13.
[0066] Furthermore, rendering unit 13 allows for the use of, for example, a billboard rendering method, where objects are rendered to maintain the vertical pose of the subject relative to the viewpoint. For instance, when rendering multiple subjects, the billboard rendering method can be used for subjects that are less of the viewer's attention, while another rendering method can be used for the other subjects.
[0067] In addition, the rendering unit 13 appropriately applies various types of processing to the generated volumetric video, such as the addition of shadows, image quality enhancement, and effects, to generate video content to be ultimately distributed to users.
[0068] (Sending Unit 14)
[0069] The sending unit 14 sends (distributes) the video stream of the video content output from the rendering unit 13 to one or more user terminals, including the receiving unit 15 and the display unit 16, via a predetermined network. The predetermined network can be various networks, such as the Internet, local area network (LAN) (including Wi-Fi, etc.), wide area network (WAN), and mobile communication network (including LTE, fourth-generation mobile communication system, fifth-generation mobile communication system, etc.).
[0070] (Receiver Unit 15)
[0071] The receiving unit 15 is installed in the aforementioned user terminal and receives video content sent (distributed) from the sending unit 14 via a predetermined network.
[0072] (Display Unit 16)
[0073] Display unit 16 is disposed in the aforementioned user terminal and displays video content received by receiving unit 15 to the user. The user terminal can be, for example, an electronic device capable of viewing moving image content, such as a head-mounted display, spatial display, television, PC, smartphone, mobile phone, or tablet terminal. Furthermore, display unit 16 can be a 2D monitor or a 3D monitor.
[0074] Note that the following is shown Figure 1 The information processing system 10 in the system directs a series of processes, starting from the data acquisition unit 11 that acquires captured images as material to be generated, up to the control of the display unit 16 of the user terminal used by the user for viewing. However, not all functional blocks are necessary to implement this embodiment, and this embodiment can be implemented for each functional block or a combination of multiple functional blocks. For example, Figure 1The configuration shown is an exemplary case where the rendering unit 13 is located on the server side. The configuration is not limited to this case, and the rendering unit 13 can be located on the user terminal side including the display unit 16. Furthermore, when the 3D model generation unit 12 and the rendering unit 13 are located in different servers (information processing devices) connected via a network, it is permissible to equip the server with the 3D model generation unit 12 with an encoding unit for compressing transmitted data, and to equip the server with the rendering unit 13 with a decoding unit for decoding compressed data.
[0075] When information processing system 10 is implemented, a single operator can perform all processing, or different operators can perform processing for each functional block or each process (step) described below. For example, Company X generates 3D content through data acquisition unit 11 and 3D model generation unit 12. There are also cases where multiple operators jointly process the content to allow distribution of the 3D content through Company Y's platform, and then the information processing device in Company Z performs operations such as rendering and display control of the 3D content.
[0076] Furthermore, each of the aforementioned functional blocks can be implemented in the cloud. For example, rendering unit 13 can be implemented on the user terminal side or on the server side. In this case, information is exchanged between the user terminal and the server.
[0077] Figure 1 The information processing system 10 includes a data acquisition unit 11, a 3D model generation unit 12, a rendering unit 13, a sending unit 14, a receiving unit 15, and a display unit 16. Alternatively, the information processing system 10 of this specification can be flexibly defined such that a system comprising two or more functional blocks is referred to as an information processing system, or for example, two of the data acquisition unit 11, the 3D model generation unit 12, the rendering unit 13, and the sending unit 14 can be collectively referred to as the information processing system 10, excluding the receiving unit 15 or the display unit 16.
[0078] 1-2. Processing flow performed by the information processing system
[0079] Next, we will refer to Figure 3 Describe the process flow performed by the information processing system 10. Figure 3 This is a flowchart illustrating an example of the process performed by the information processing system 10.
[0080] like Figure 3 As shown, this operation begins when the data acquisition unit 11 acquires real camera images of the subject captured by multiple real cameras 70 (step S11).
[0081] Next, the 3D model generation unit 12 generates a 3D model with the 3D information (3D shape mesh data and texture) of the subject based on the real camera image obtained in step S11 (step S12).
[0082] Next, the rendering unit 13 performs rendering of the 3D model generated in step S12 based on the 3D shape mesh data and texture to generate a video content to be presented to the user (step S13).
[0083] Next, the sending unit 14 sends (distributes) the video content generated in step S13 to the user terminal (step S14).
[0084] Next, the receiving unit 15 of the user terminal receives the video content sent from the sending unit 14 (step S15). Subsequently, the display unit 16 of the user terminal displays the video content received in step S15 to the user (step S16). After this, the information processing system 10 ends this operation.
[0085] 1-3. Hardware Configuration of the Information Processing System
[0086] Next, we will refer to Figure 4 Describe the hardware configuration of the information processing system 10. Figure 4 This is a hardware block diagram illustrating an example of the hardware configuration of the information processing system 10.
[0087] exist Figure 4 In this configuration, CPU 21, ROM 22, and RAM 23 are interconnected via bus 24. Bus 24 is also connected to input / output interface 25. Input / output interface 25 is connected to input unit 26, output unit 27, storage unit 28, communication unit 29, and driver 20.
[0088] Input unit 26 includes, for example, a keyboard, mouse, microphone, touch panel, input terminal, etc. Output unit 27 includes, for example, a display, speaker, output terminal, etc. Storage unit 28 includes, for example, a hard disk, RAM disk, non-volatile memory, etc. Communication unit 29 includes, for example, a network interface, etc. Driver 20 drives removable media, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.
[0089] In the computer configured as described above, for example, the CPU 21 loads the program stored in the storage unit 28 into the RAM 23 via the input / output interface 25 and the bus 24 and executes the program, thereby performing the series of processes described above. The RAM 23 also appropriately stores data, etc., required by the CPU 21 to perform various types of processing.
[0090] A computer-executed program can be used, for example, by recording the program on a removable medium, which serves as the encapsulation medium. In this case, by attaching the removable medium to the drive, the program can be installed in the storage unit 28 via an input / output interface.
[0091] Note that the program can also be provided via wired or wireless transmission media such as a local area network, the Internet, or digital satellite broadcasting. In this case, the program can be received by the communication unit 29 and installed in the storage unit 28.
[0092] 1-4. Further details of this embodiment
[0093] The above has described a schematic configuration example of the information processing system 10. Below, a more detailed configuration example based on the aforementioned information processing system 10 will be described in detail with reference to the accompanying drawings.
[0094] 1-4-1. Further detailed configuration examples of rendering units
[0095] Figure 5 This is a block diagram illustrating a further detailed configuration example of the rendering unit 13 according to this embodiment. (See diagram for details.) Figure 5 As shown, the rendering unit 13 includes an initial virtual viewpoint video generation unit 131, a sending unit 132, a receiving unit 133, and a final virtual viewpoint video generation unit 134. Although Figure 4 A final virtual viewpoint video generation unit 134 is shown, but the rendering unit 13 may include multiple final virtual viewpoint video generation units 134.
[0096] The rendering unit 13 receives input a three-dimensional model (three-dimensional shape mesh data and texture) generated by the 3D model generation unit 12.
[0097] In addition, the rendering unit 13 also receives input about the virtual viewpoint during the rendering of the 3D model (hereinafter referred to as virtual viewpoint information). The virtual viewpoint information may include information indicating the position of the virtual viewpoint (hereinafter referred to as virtual viewpoint position) and information about the 3D rotation matrix relative to the reference position of the virtual viewpoint (hereinafter referred to as virtual viewpoint rotation information).
[0098] In addition, the rendering unit 13 may also receive input of real camera images acquired by any one or more real cameras 70 (hereinafter, N viewpoints (N is an integer of 0 or greater)), the viewpoint position of each of the N viewpoints of the real camera 70 (hereinafter referred to as the real camera viewpoint position), and information of a three-dimensional rotation matrix indicating the reference position of the real camera 70 relative to the viewpoint position of each of the N viewpoints (hereinafter referred to as real camera viewpoint rotation information). The real camera viewpoint position and the real camera viewpoint rotation information may be information included in the internal and external parameters obtained, for example, through calibration performed by the data acquisition unit 11, and this information is collectively referred to as real camera viewpoint information in the following description.
[0099] (Initial Virtual Viewpoint Video Generation Unit 131)
[0100] The initial virtual viewpoint video generation unit 131 renders a 3D model from a virtual viewpoint based on input 3D shape mesh data, texture, and virtual viewpoint information, thereby generating a virtual viewpoint video. Subsequently, the initial virtual viewpoint video generation unit 131 uses the generated virtual viewpoint video to generate an intermediate rendered video to be sent to the final virtual viewpoint video generation unit 134. As described below, the intermediate rendered video can be a video obtained by tiling virtual viewpoint video (RGB) and virtual viewpoint video (Depth) together to aggregate them into a single image, and can be a video in which images such as auxiliary video or real camera images are further included in a single image. Details of the intermediate rendered video will be described below. For example, the initial virtual viewpoint video generation unit 131 may correspond to an example of the second generation unit in the claims.
[0101] (Transmitting unit 132 and receiving unit 133)
[0102] The transmitting unit 132 and the receiving unit 133 are configured to transmit the intermediate rendered video generated by the initial virtual viewpoint video generation unit 131 to one or more final virtual viewpoint video generation units 134. The transmission of the intermediate rendered video via the transmitting unit 132 and the receiving unit 133 can be performed through a predetermined network such as a LAN, WAN, the Internet, or a mobile communication network, or through a predetermined interface such as a High-Definition Multimedia Interface (HDMI) or a Universal Serial Bus (USB). However, the transmission method is not limited to these, and various communication means can be used.
[0103] (Final Virtual Viewpoint Video Generation Unit 134)
[0104] Based on the intermediate rendered video input from the initial virtual viewpoint video generation unit 131 via the sending unit 132 and the receiving unit 133, and based on the real camera viewpoint information shared by one or more final virtual viewpoint video generation units 134, the final virtual viewpoint video generation unit 134 performs processes not performed in the initial virtual viewpoint video generation unit 131 to generate video content to be finally presented to the user. For example, the final virtual viewpoint video generation unit 134 performs processes such as overlaying background objects or another 3D object onto the volumetric video generated from the 3D model, and enhancing the image quality of the volumetric video. Furthermore, the final virtual viewpoint video generation unit 134 may also perform the arrangement of the real camera image 33 relative to the generated volumetric video, effects processing, etc. For example, the final virtual viewpoint video generation unit 134 may correspond to an example of the first generation unit in the claims.
[0105] 1-4-2. Specific example of intermediate video rendering
[0106] Here, a specific example of intermediate-rendered video will be described. Figure 6 This is a diagram illustrating an example of an intermediate rendered video according to this embodiment. (See diagram for example.) Figure 6 As shown, the intermediate rendered video 30 is configured such that virtual viewpoint video (RGB) 31, virtual viewpoint video (depth) 32, real camera image 33, and auxiliary video 34 are combined (described herein as tiled) to form a single image data. For example, the real camera image 33 and auxiliary video 34 may each correspond to an example of a captured image in the claims. However, the real camera image 33 and auxiliary video 34 in this embodiment, as well as the multiple captured images in the claims, are not limited to real camera images acquired by the real camera 70, but can be various types of video content, such as movies, music videos, promotional videos thereof, videos broadcast on television, videos distributed on the Internet, and videos shared in video conferences, regardless of whether the video is captured online or created offline. Furthermore, for example, the intermediate rendered video 30 may correspond to an example of a packaged image in the claims.
[0107] (Virtual viewpoint video (RGB) 31)
[0108] Figure 7 This is a diagram illustrating an example of virtual viewpoint video (RGB) according to this embodiment. (See diagram for example.) Figure 7As shown, the virtual viewpoint video (RGB) 31 can be, for example, a volumetric video at the current point generated by rendering a 3D model (3D shape mesh data and texture) from a virtual viewpoint based on virtual viewpoint information. In this case, the virtual viewpoint video (RGB) 31 can retain the texture information obtained when viewing the 3D model from the virtual viewpoint. For example, the virtual viewpoint video (RGB) 31 can correspond to an example of the first texture image in the claim.
[0109] (Virtual Viewpoint Video (Depth) 32)
[0110] Figure 8 This is a diagram illustrating an example of virtual viewpoint video (depth) according to this embodiment. (See diagram for example.) Figure 8 As shown, the virtual viewpoint video (depth) 32 can be image data indicating the depth of each pixel in the virtual viewpoint video (RGB) 31 from the virtual viewpoint, and can be a depth image generated by calculating the distance (depth information) from the virtual viewpoint position to each point in the 3D shape mesh data. In this case, the virtual viewpoint video (depth) 32 can maintain the depth information from the virtual viewpoint to the 3D model obtained when viewing the 3D model from the virtual viewpoint.
[0111] (Real camera image 33)
[0112] Figure 9 This is a diagram illustrating an example of a real camera image according to this embodiment. (e.g.) Figure 9 As shown, for example, real camera image 33 can be a real camera image captured by any of the cameras selected from real cameras 70. Note that the selection of a camera can be random, a selection made on the content creator's side, or a selection made by the content viewer (user).
[0113] (Supporting Video 34)
[0114] Figure 10 This is a diagram illustrating an example of an auxiliary video according to this embodiment. (See diagram for example.) Figure 10 As shown, for example, supplementary video 34 is image data used to enhance the image quality of the volumetric video to be ultimately provided to the user, and may, for example, include an image of subject 90 captured from the same viewpoint as the virtual viewpoint but different from the virtual viewpoint video (RGB) 31. For example, supplementary video 34 may be a real camera image captured by any of the cameras selected from real cameras 70. Similar to real camera image 33, the selection of any real camera may be random, may be a selection made on the content creator's side, or may be a selection made by the content viewer (user). However, supplementary video 34 may be omitted. Note, for example, that supplementary video 34 may correspond to an example of the second texture image in the claim.
[0115] (Shared Virtual Viewpoint)
[0116] In this embodiment, the initial virtual viewpoint video generation unit 131 and one or more final virtual viewpoint video generation units 134 share virtual viewpoint information (virtual viewpoint position and virtual viewpoint rotation information) to enable processing to be performed based on the same virtual viewpoint. This allows processing based on the same virtual viewpoint to be performed in each of the initial virtual viewpoint video generation unit 131 and / or the multiple final virtual viewpoint video generation units 134, enabling operations such as distributed execution by multiple computers of each process in the initial virtual viewpoint video generation unit 131 and / or each process in the final virtual viewpoint video generation unit 134, and distributed arrangement of each of the multiple final virtual viewpoint video generation units 134 in multiple computers, to generate the final volumetric video (corresponding to video content) to be provided to the user, respectively.
[0117] 1-4-3. Further detailed configuration example of the initial virtual viewpoint video generation unit 131
[0118] Next, a more detailed configuration example of the initial virtual viewpoint video generation unit 131 will be described. Figure 11 This is a block diagram illustrating a further detailed configuration example of the initial virtual viewpoint video generation unit according to this embodiment. (As shown...) Figure 11 As shown, the initial virtual viewpoint video generation unit 131 includes a virtual viewpoint video (RGB) generation unit 1312, an auxiliary video generation unit 1313, a virtual viewpoint video (depth) generation unit 1314, and an intermediate rendering video generation unit 1315.
[0119] The virtual viewpoint video (RGB) generation unit 1312 receives inputs of 3D shape mesh data, texture, virtual viewpoint information, and real camera viewpoint information. The auxiliary video generation unit 1313 receives inputs of 3D shape mesh data, texture, virtual viewpoint information, real camera viewpoint information, and real camera images (N viewpoints). Note that the real camera images (N viewpoints) are also input to the intermediate rendering video generation unit 1315.
[0120] (Virtual viewpoint video (RGB) generation unit 1312)
[0121] As described above, as an operation of the rendering unit 13, the virtual viewpoint video (RGB) generation unit 1312 projects three-dimensional shape mesh data from the virtual viewpoint and performs texture mapping that applies textures representing the colors or patterns of the mesh to the projected mesh shape. The virtual viewpoint video (RGB) 31 thus generated is input to the virtual viewpoint video (depth) generation unit 1314 and the intermediate rendering video generation unit 1315, respectively.
[0122] (Auxiliary video generation unit 1313)
[0123] The auxiliary video generation unit 1313 generates auxiliary video 34 for image quality enhancement performed by the final virtual viewpoint video generation unit 134 described below. Note that real camera images acquired by any one or more of the real cameras 70 can be used as auxiliary video 34.
[0124] (Virtual viewpoint video (depth) generation unit 1314)
[0125] The virtual viewpoint video (depth) generation unit 1314 generates a virtual viewpoint video (depth) 31 as a depth image based on the depth information of each point (corresponding to a pixel) on the 3D model determined when the virtual viewpoint video (RGB) generation unit 1312 generates the virtual viewpoint video (RGB) 32. The generated virtual viewpoint video (depth) 32 is input to the intermediate rendering video generation unit 1315.
[0126] Note that while the virtual viewpoint video (RGB) generation unit 1312 determines depth information in absolute values (mm, etc.), the virtual viewpoint video (depth) generation unit 1314 can generate virtual viewpoint video (depth) 32 by quantizing the depth information of each pixel. For example, when the bit depth of other image data such as virtual viewpoint video (RGB) 31 is 8 bits, the virtual viewpoint video (depth) generation unit 1314 can quantize the depth information (mm, etc.) of each pixel into 256 grayscale depth information values from "0" to "255". This eliminates the need to increase the bit depth of intermediate rendered video, thus suppressing the increase in the amount of data to be sent.
[0127] (Intermediate Rendering Video Generation Unit 1315)
[0128] The intermediate rendering video generation unit 1315 generates an intermediate rendering video 30 by tiling the virtual viewpoint video (RGB) 31 input from the virtual viewpoint video (RGB) generation unit 1312, the virtual viewpoint video (depth) 32 input from the virtual viewpoint video (depth) generation unit 1314, the directly input real camera images (N viewpoints) 33, and the auxiliary video 34 input from the auxiliary video generation unit 1313, according to a predetermined arrangement. Figure 6The generated intermediate rendered video 30 is output to the sending unit 132 in the rendering unit 13.
[0129] 1-4-4. Further detailed configuration example of the final virtual viewpoint video generation unit 134
[0130] The following section will describe a more detailed configuration example of the final virtual viewpoint video generation unit 134. Figure 12 This is a block diagram illustrating a further detailed configuration example of the final virtual viewpoint video generation unit according to this embodiment. (As shown...) Figure 12 As shown, the final virtual viewpoint video generation unit 134 includes an object generation unit 1342, a camera image update unit 1343, a virtual viewpoint video generation unit 1344, a shadow generation unit 1345, an image quality enhancement unit 1346, and an effect processing unit 1347.
[0131] The object generation unit 1342 receives input of the object (content) to be overlaid on the 3D model and virtual viewpoint information. The camera image update unit 1343 receives input of the intermediate rendered video 30 and virtual viewpoint information. Note that the intermediate rendered video 30 is also input to the virtual viewpoint video generation unit 1344, the shadow generation unit 1345, and the image quality enhancement unit 1346. Furthermore, the virtual viewpoint information can also be input to the virtual viewpoint video generation unit 1344, the shadow generation unit 1345, the image quality enhancement unit 1346, and the effects processing unit 1347.
[0132] (Object generation unit 1342)
[0133] The object generation unit 1342 generates 3D objects, background objects, etc. (hereinafter referred to as additional objects) to be superimposed on the volumetric video to be provided to the user by using input objects (content). The generated additional objects are input to the virtual viewpoint video generation unit 1344.
[0134] (Camera image update unit 1343)
[0135] The camera image update unit 1343 extracts a real camera image 33 from the intermediate rendered video 30. Note that the real camera image 33 incorporated into the intermediate rendered video 30 can be a real camera image synchronized with the virtual viewpoint video (RGB), i.e., temporally corresponding to the virtual viewpoint video (RGB). In other words, the real camera image 33 can be a two-dimensional image that includes a subject 90 temporally corresponding to the subject 90 used to generate the 3D model. The extracted real camera image 33 is input to the virtual viewpoint video generation unit 1344.
[0136] (Virtual viewpoint video generation unit 1344)
[0137] The virtual viewpoint video generation unit 1344 extracts virtual viewpoint video (RGB) 31 and virtual viewpoint video (depth) 32 from the intermediate rendered video 30, and uses the extracted virtual viewpoint video (RGB) 31 and virtual viewpoint video (depth) 32 to reconstruct the 3D model. Furthermore, the virtual viewpoint video generation unit 1344 overlays additional objects input from the object generation unit 1342 onto the reconstructed 3D model as needed. That is, the virtual viewpoint video generation unit 1344 arranges the 3D model and additional objects in the same virtual space. At this time, for example, the virtual viewpoint video generation unit 1344 can adjust the positional relationship between the 3D model and the additional objects based on the depth information of the additional objects relative to the virtual viewpoint and based on the virtual viewpoint video (depth) 32.
[0138] The virtual viewpoint video generation unit 1344 then generates a volumetric video of the 3D model (and attached objects) by using a 3D model (and attached objects) reconstructed by virtual viewpoint rendering based on virtual viewpoint information (virtual viewpoint position and virtual viewpoint rotation information).
[0139] Additionally, the virtual viewpoint video generation unit 1344 generates video content by overlaying a real camera image 33 input from the camera image update unit 1343 onto a predetermined area of the generated volumetric video. The video content thus generated is, for example, a video in which a subject generated from a 3D model and an object based on the real camera image 33 coexist, and may correspond to an example of the video in the claims.
[0140] However, content generation is not limited to this, and the virtual viewpoint video generation unit 1344 can directly generate video content by arranging the real camera image 33 on a plane within a virtual space where a 3D model (and additional objects) are arranged, and by rendering the 3D model (and additional objects) and the real camera image 33 using a virtual viewpoint. The combined model of the generated 3D model (and additional objects) and the real camera image 33 can then correspond to an example of the combined 3D model in the claims.
[0141] The video content generated in this way is input into the shadow generation unit 1345.
[0142] (Shadow Generation Unit 1345)
[0143] The shadow generation unit 1345 applies shadows to objects included in the volumetric video of the input video content. This enhances the realism of the volumetric video. However, the shadow generation unit 1345 can be omitted or disabled if shadows are not applied. For example, the shadow generation unit 1345 applies shadows to the images of the 3D model and the additional objects based on the virtual viewpoint video (depth) 32 included in the intermediate rendered video 30, the depth information of the additional objects generated when the object generation unit 1342 generates additional objects, etc. The video content to which shadows are added to the objects is then input to the image quality enhancement unit 1346.
[0144] Note that the shadow generation unit 1345 also receives input about the position and type (hue, etc.) of the light source used to generate the shadow (hereinafter referred to as light source information). For example, the light source information may be included in the virtual viewpoint information, or it may be input separately to the shadow generation unit 1345.
[0145] (Image quality enhancement unit 1346)
[0146] The image quality enhancement unit 1346 performs image quality enhancement processing on the input video content. For example, the image quality enhancement unit 1346 performs processes such as contour blurring, alpha blending, and noise removal on images of 3D models and additional objects included in the volumetric video of specific video content, thereby suppressing temporal fluctuations that occur in the volumetric video. At this time, the image quality enhancement unit 1346 can use the auxiliary video 34 included in the intermediate rendered video 30 for the above processing.
[0147] The image quality enhancement unit 1346 removes the green screen color during rendering so that the green screen color is not left in the outline of the 3D model. In this specification, the video before removal is referred to as auxiliary video 34. For example, as... Figure 13 As shown, the image quality enhancement unit 1346 can calculate the difference between the auxiliary video 34 and the volumetric video 40 generated by the virtual viewpoint video generation unit 1344 to extract the contour 41 of the image of the 3D model (and additional objects) included in the volumetric video 40. Then, based on the extracted contour 41, the volumetric video can be subjected to blurring, alpha blending, noise removal, etc. At this time, the image quality enhancement unit 1346 can adjust the scale, tilt, color, etc. of the volumetric video 40 and / or the auxiliary video 34 to achieve a near-perfect match between the image of the 3D model included in the volumetric video 40 and the image of the subject 90 included in the auxiliary video 34.
[0148] (Effects Processing Unit 1347)
[0149] The effects processing unit 1347 can perform various effects processing on the video content. For example, the effects processing unit 1347 performs effects processing including applying effects such as lighting effects to images of 3D models or attached objects, applying effects such as mosaic effects to real camera images 33, and applying petal falling effects to the background. This allows for an expanded scope of presentation. Note that the effects processing unit 1347 can be omitted or disabled if no effects are applied. Subsequently, the effects processing unit 1347 outputs the video content with applied effects to the sending unit 14 as the final video content to be distributed to the user, as needed. Therefore, the final video content sent from the sending unit 14 is received by the receiving unit of the user terminal and displayed on the display unit 16 of the user terminal.
[0150] 1-5. Examples of demonstration according to this embodiment
[0151] Figure 14 This is a diagram illustrating an example of video content distributed to a user in this embodiment. Figure 14 In the video content 50A shown, a real camera image 33, synchronized with the volumetric video including the image 52 (corresponding in time to the volumetric video including the image 52), is superimposed on the volumetric video. Furthermore, in the video content 50A according to this embodiment, an auxiliary video 34, synchronized with the volumetric video (corresponding in time to the volumetric video), is also superimposed. The auxiliary video 34 has undergone an effect 51 as a lighting effect, which allows the displayed subject 90 to appear as if it is emitting light and shadow.
[0152] Figure 15 This is a view illustrating another example of a volumetric video to be distributed to users in this embodiment. Figure 15 In the video content 50B shown, the image 53 of the background object is superimposed on the image 52 of the 3D model. That is, a volumetric video is generated while the background object is positioned within a virtual space containing the 3D model. Similarly, in... Figure 15 In the process, a real camera image 33, synchronized with the volumetric video (corresponding to the volumetric video in time), is superimposed on the volumetric video.
[0153] In this way, this implementation allows another video to be overlaid on a volumetric video, thereby expanding the scope of the presentation in the video content provided to the user. Note that the video to be overlaid on the volumetric video is not limited to real camera images 33 (including real camera images of auxiliary video 34), but can be various separately prepared image data, such as a scene from a promotional video, animation, or movie.
[0154] Furthermore, according to this embodiment, various effects can be applied to volumetric videos and image data overlaid on them. This allows for a further expansion of the scope of demonstrations within the video content to be provided to users.
[0155] Furthermore, according to this embodiment, it is possible to perform the rendering and display of another three-dimensional object, background object, etc., superimposed on the volumetric video. This allows for a further expansion of the scope of the video content to be presented to the user.
[0156] Note that image data and additional objects overlaid on volumetric video can be configured to switch automatically or switch at any time based on user action.
[0157] 1-6. Specific examples of information processing systems
[0158] Next, a specific example of the information processing system 10 according to this embodiment will be described. Figure 16 This is a system configuration diagram illustrating a specific example of an information processing system according to this embodiment. For example... Figure 16 As shown, the information processing system 10 includes a cloud server 83, a publishing server 85, a distribution server 87, and a user terminal 89. The user terminal 89 does not need to be included in the information processing system 10.
[0159] For example, by utilizing one or more real cameras 70 installed in the volumetric studio 80 (see reference). Figure 2 The subject 90 is captured to obtain a real camera image 81 for generating volumetric video. At this time, audio data based on the sound signal recorded along with the real camera image 81 can be generated by an audio mixer 82. The volumetric studio 80 can be indoors or outdoors, as long as it is an environment capable of imaging the subject 90 in one or more directions.
[0160] Real camera images 81 acquired by real camera 70 and audio data generated by audio mixer 82 are input to cloud server 83 set on a predetermined network. Cloud server 83 includes, for example... Figure 1 The configuration shown includes a portion of the data acquisition unit 11, the 3D model generation unit 12, and the rendering unit 13—namely, the initial virtual viewpoint video generation unit 131 and the sending unit 132.
[0161] Using the 3D model generation unit 12, the cloud server 83 generates a 3D model of the subject 90 from the real camera image 81 acquired by the data acquisition unit 11. Then, the cloud server 83 generates a virtual viewpoint video (RGB) 31 and a virtual viewpoint video (depth) 32 from the 3D model in the initial virtual viewpoint video generation unit 131 of the rendering unit 13, and generates an intermediate rendered video 30 from the generated virtual viewpoint video (RGB) 31 and virtual viewpoint video (depth) 32 and from the input real camera image 81 (corresponding to real camera image 33 and auxiliary video 34).
[0162] Note that cloud server 83 is just an example, and it can be implemented by using various information processing devices set up on the network, such as fog servers and edge servers.
[0163] The generated intermediate rendered video 30 is sent from the sending unit 132 of the rendering unit 13 to the publishing server 85 via a predetermined network. The publishing server 85 includes, for example, the receiving unit 133 in the rendering unit 13 and the final virtual viewpoint video generation unit 134.
[0164] The publishing server 85 extracts virtual viewpoint video (RGB) 31, virtual viewpoint video (depth), real camera image 33, and auxiliary video 34 from the intermediate rendered video 30, and generates the final video content 50 to be presented to the user based on these videos. The specific operation can be similar to the operation of the final virtual viewpoint video generation unit 134 described above.
[0165] The video content 50 generated in this manner is converted into a distribution stream 88 corresponding to each user terminal 89 in a distribution server 87 equipped with a sending unit 14, and then distributed from the distribution server 87 to the user terminal 89 via a predetermined network. In response, the user terminal 89 receives the distributed distribution stream 88 through the receiving unit 15 and recovers the video content from the distribution stream 88. The user terminal 89 then displays the recovered video content 50 to the user on the display unit 16.
[0166] Summary of 1-7
[0167] As described above, according to this embodiment, real camera images 33 (including perspective videos, videos with effects, etc.) can be displayed in sync with and superimposed on the volumetric video 40. Furthermore, background objects can be displayed in the presentation in contrast to the volumetric video 40. This allows for a significant expansion of the presentation scope when distributing volumetric videos in real time.
[0168] Furthermore, in this embodiment, during the processing performed by the rendering unit 13, the three-dimensional data is converted into intermediate rendered video 30 as two-dimensional image data. This enables the transmission of data used to generate volumetric video via HDMI capture or codec compression, allowing for lower communication overhead during transmission and shorter data transmission time. Therefore, real-time performance in the distribution of video content including volumetric video can be enhanced.
[0169] Furthermore, by sharing virtual viewpoint information among multiple computers and distributing the intermediate rendered video 30 to each computer, processing related to volumetric video generation (e.g., virtual viewpoint video generation, shadow generation, image quality enhancement, effects processing, etc.) can be performed using distributed processing across multiple computers. This reduces the load on each computer. Additionally, processing related to volumetric video generation can be flexibly allocated to computers suitable for each process, further reducing the load on each computer. For example, job allocation can be performed such that the generation of virtual viewpoint video (RGB) 31 and virtual viewpoint video (depth) 32 is performed by a computer equipped with a Linux system capable of performing GPU-based distributed processing, while subsequent processing such as shadow application, image quality enhancement, and effects processing is performed by a computer equipped with an operating system (OS) capable of using libraries suitable for effortless implementation of the processing. Therefore, the overall processing time can be reduced by efficiently executing each process, enhancing the real-time performance of video content distribution, including volumetric video.
[0170] For example, the program for implementing each unit (each process) described in the above embodiments can be executed on a specific device. In this case, there will be no problem as long as the device has the necessary functional blocks and can obtain the necessary information.
[0171] Furthermore, for example, each step of a flowchart can be executed by a single device, or it can be shared and executed by multiple devices. Additionally, when a step includes multiple processes, these processes can be executed by a single device, or they can be shared and executed by multiple devices. In other words, multiple processes included in a step can also be executed as processes of multiple steps. Conversely, processes described as multiple steps can be executed together as a single step.
[0172] Furthermore, for example, with respect to a program executed by a computer, the processing of the steps describing the program can be performed sequentially in the order described in this specification, or it can be performed individually or in parallel at necessary timing points, such as when called. That is, as long as there is no contradiction, the processing of each step can be performed in an order different from the above-described order. In addition, the processing of the steps describing the program can be performed in parallel with the processing of another program, or it can be performed in combination with the processing of another program.
[0173] Furthermore, for example, multiple technologies related to this technology can be implemented independently, provided there are no contradictions. Naturally, any of the multiple technologies can be combined. For example, some or all of the technologies described in any embodiment can be combined with some or all of the technologies described in other embodiments. Additionally, some or all of the aforementioned technologies can be combined with other technologies not described above.
[0174] The embodiments of this disclosure have been described above. However, the technical scope of this disclosure is not limited to the embodiments described above, and various modifications can be made without departing from the scope of this disclosure. Furthermore, combinations of components and appropriate modifications across different embodiments are permitted.
[0175] The effects described in the various embodiments of this specification are merely illustrative, and therefore, other effects may exist, not limited to those illustrated.
[0176] Note that this technology can also have the following configurations. (1)
[0178] An information processing apparatus includes a first generation unit that generates a video based on a three-dimensional model of the subject generated by using multiple captured images obtained through imaging the subject and a two-dimensional image, wherein the video simultaneously contains the subject generated from the three-dimensional model and the two-dimensional image. (2)
[0180] According to the information processing device described in (1),
[0181] Wherein, the two-dimensional image is a two-dimensional image of at least one of the plurality of captured images used to generate a three-dimensional model of the subject, and
[0182] The first generation unit generates the video, in which a subject generated from the three-dimensional model and a subject based on the two-dimensional image corresponding to the subject exist simultaneously. (3)
[0184] According to the information processing device described in (2),
[0185] The first generation unit generates the video based on a three-dimensional model of the subject and a two-dimensional image of the subject that corresponds in time to the subject used to generate the three-dimensional model. (4)
[0187] The information processing apparatus according to any one of (1) to (3) further includes
[0188] The second generation unit generates a packaged image in which a texture image obtained by converting the three-dimensional model of the subject into two-dimensional texture information based on a virtual viewpoint and a depth image obtained by converting the depth information from the virtual viewpoint to the three-dimensional model of the subject into a two-dimensional image are packaged in one frame. The virtual viewpoint is set in a virtual space where the three-dimensional model is arranged. (5)
[0190] According to the information processing device described in (4),
[0191] The packaged image also includes at least one of the plurality of captured images. (6)
[0193] According to the information processing device described in (5),
[0194] The texture image and the captured image included in the packaged image are temporally corresponding images. (7)
[0196] The information processing apparatus according to any one of (4) to (6),
[0197] The packaged image includes the following texture images: a first texture image, which is obtained by converting the three-dimensional model of the subject into the two-dimensional texture information based on the virtual viewpoint, wherein the virtual viewpoint is set in the virtual space where the three-dimensional model is arranged; and a second texture image, which includes the subject from the same viewpoint as the virtual viewpoint and is an image different from the first texture image. (8)
[0199] The information processing apparatus according to any one of (4) to (7) further includes
[0200] The sending unit sends the packaged image; and
[0201] The receiving unit receives the packaged image from the sending unit.
[0202] The first generation unit reconstructs the 3D model based on the packaged image received by the receiving unit, renders the 3D model from a virtual viewpoint set in a virtual space where the 3D model is arranged, and generates a 2D image of a subject generated from the 3D model through the reconstruction and rendering operations, and uses the 2D image to generate the video. (9)
[0204] The information processing apparatus according to any one of (4) to (8) further includes
[0205] Multiple first generation units,
[0206] Each of the first generation units generates the video by using a 3D model reconstructed from the packaged images obtained from the second generation unit. (10)
[0208] The information processing apparatus according to any one of (1) to (9) further includes
[0209] A shadow generation unit that applies shadows to areas of the subject included in the video. (11)
[0211] The information processing apparatus according to any one of (1) to (10) further includes
[0212] An image quality enhancement unit enhances the image quality of the video by using at least one of the plurality of captured images. (12)
[0214] The information processing apparatus according to any one of (1) to (11) further includes
[0215] An effects processing unit performs effects processing on the video. (13)
[0217] The information processing apparatus according to any one of (1) to (12) further includes
[0218] A sending unit that sends the video generated by the first generating unit to one or more user terminals via a predetermined network. (14)
[0220] The information processing apparatus according to any one of (1) to (13),
[0221] The first generation unit sets a three-dimensional model of the subject generated using multiple captured images obtained by imaging the subject in a three-dimensional space, and arranges a two-dimensional image based on at least one of the multiple captured images in the three-dimensional space. Through this arrangement, the first generation unit generates a combined three-dimensional model including the three-dimensional model and the two-dimensional image. (15)
[0223] According to the information processing device described in (14),
[0224] The first generation unit generates the video by rendering the combined 3D model based on a virtual viewpoint set in a virtual space where the 3D model is arranged. (16)
[0226] An information processing method includes: generating a video using a computer based on a three-dimensional model of the subject generated by using multiple captured images obtained through imaging the subject, and based on a two-dimensional image, wherein the video simultaneously contains the subject generated from the three-dimensional model and the two-dimensional image. (17)
[0228] A video distribution method, comprising:
[0229] A three-dimensional model of the subject is generated by using multiple captured images obtained through imaging the subject;
[0230] Based on the 3D model and 2D image of the subject, generate a video in which both the subject generated from the 3D model and the 2D image are simultaneously present; and
[0231] The video is distributed to user terminals via a predetermined network. (18)
[0233] An information processing system, comprising:
[0234] An imaging device that images a subject to generate multiple captured images of the subject;
[0235] An information processing apparatus that generates a video containing both a subject generated from the three-dimensional model and a two-dimensional image of the subject generated using the plurality of captured images; and
[0236] A user terminal displays the video generated by the information processing device to a user.
[0237] List of reference numerals
[0238] 10. Information Processing System
[0239] 11 Data Acquisition Unit
[0240] 12 3D model generation units
[0241] 13 Rendering Units
[0242] 14 Transmitting Unit
[0243] 15 Receiving Unit
[0244] 16 display units
[0245] 20 drives
[0246] 21 CPU
[0247] 22 ROM
[0248] 23 RAM
[0249] 24 BUS
[0250] 25 Input / Output Interfaces
[0251] 26 Input Units
[0252] 27 Output Unit
[0253] 28 storage units
[0254] 29 Communication Unit
[0255] 30 Intermediate rendering videos
[0256] 31 Virtual Viewpoint Video (RGB)
[0257] 32 Virtual Viewpoint Videos (In-Depth)
[0258] 33,81 Real camera images
[0259] 34 Supporting Videos
[0260] 40-page video
[0261] 41. Outline
[0262] Video content of 50, 50A, and 50B
[0263] 51 Effects
[0264] 52 Images of 3D models
[0265] 53 Images of background objects
[0266] 70, 70a, 70b, 70c, 70d, 70e, ... real cameras
[0267] 80 Volume Studio
[0268] 82 Audio Mixer
[0269] 83 Cloud Servers
[0270] 85 Release Server
[0271] 87 Distribution Server
[0272] 88 Distribute Stream
[0273] 89 User Terminal
[0274] 90 subjects
[0275] 131 Initial Virtual Viewpoint Video Generation Unit
[0276] 132 Transmitting Unit
[0277] 133 Receiving Unit
[0278] 134 Final Virtual Viewpoint Video Generation Unit
[0279] 1312 Virtual Viewpoint Video (RGB) Generation Unit
[0280] 1313 Auxiliary Video Generation Unit
[0281] 1314 Virtual Viewpoint Video (Depth) Generation Unit
[0282] 1315 Intermediate Rendering Video Generation Unit
[0283] 1342 Object Generation Unit
[0284] 1343 Camera Image Update Unit
[0285] 1344 Virtual Viewpoint Video Generation Unit
[0286] 1345 Shadow Generation Unit
[0287] 1346 Image Quality Enhancement Units
[0288] 1347 Effects Processing Unit
Claims
1. An information processing apparatus, comprising a first generation unit, the first generation unit performing video generation based on a three-dimensional model of the subject generated by using multiple captured images obtained through imaging the subject, and a two-dimensional image, wherein the video simultaneously contains the subject generated from the three-dimensional model and the two-dimensional image. in, The information processing device further includes: a second generation unit, which generates a packaged image, wherein a texture image obtained by converting the three-dimensional model of the subject into two-dimensional texture information based on a virtual viewpoint and a depth image obtained by converting the depth information from the virtual viewpoint to the three-dimensional model of the subject into a two-dimensional image are packaged in one frame, and the virtual viewpoint is set in a virtual space in which the three-dimensional model is arranged.
2. The information processing device according to claim 1, in, The two-dimensional image is a two-dimensional image obtained by using at least one of the multiple captured images from the captured video used to generate a three-dimensional model of the subject, and The first generation unit generates the video, in which a subject generated from the three-dimensional model and a subject based on the two-dimensional image corresponding to the subject exist simultaneously.
3. The information processing device according to claim 2, in, The first generation unit generates the video based on a three-dimensional model of the subject and a two-dimensional image of the subject that corresponds in time to the subject used to generate the three-dimensional model.
4. The information processing device according to claim 1, in, The packaged image also includes at least one of the plurality of captured images.
5. The information processing apparatus according to claim 4, in, The texture image and the captured image included in the packaged image are temporally corresponding images.
6. The information processing apparatus according to claim 1, in, The packaged images include: a first texture image, which is obtained by converting the three-dimensional model of the subject into the two-dimensional texture information based on the virtual viewpoint, wherein the virtual viewpoint is set in the virtual space where the three-dimensional model is arranged; and a second texture image, which includes the subject from the same viewpoint as the virtual viewpoint and is an image different from the first texture image.
7. The information processing apparatus according to claim 1, further comprising: The sending unit sends the packaged image; and The receiving unit receives the packaged image from the sending unit. in, The first generation unit reconstructs the three-dimensional model based on the packaged image received by the receiving unit, renders the three-dimensional model with a virtual viewpoint set in a virtual space where the three-dimensional model is arranged, and through the reconstruction and rendering operations, the first generation unit generates a two-dimensional image including the subject generated from the three-dimensional model, and uses the two-dimensional image to generate the video.
8. The information processing apparatus according to claim 1, further comprising: Multiple first generation units, in, Each of the first generation units generates the video by using a 3D model reconstructed from the packaged images obtained from the second generation unit.
9. The information processing apparatus according to claim 1, further comprising: A shadow generation unit that applies shadows to areas of the subject included in the video.
10. The information processing apparatus according to claim 1, further comprising: An image quality enhancement unit enhances the image quality of the video by using at least one of the plurality of captured images.
11. The information processing apparatus according to claim 1, further comprising: An effects processing unit performs effects processing on the video.
12. The information processing apparatus according to claim 1, further comprising: A sending unit that sends the video generated by the first generating unit to one or more user terminals via a predetermined network.
13. The information processing apparatus according to claim 1, in, The first generation unit sets a three-dimensional model of the subject generated using multiple captured images obtained by imaging the subject in a three-dimensional space, and arranges a two-dimensional image based on at least one of the multiple captured images in the three-dimensional space. Through this arrangement, the first generation unit generates a combined three-dimensional model including the three-dimensional model and the two-dimensional image.
14. The information processing apparatus according to claim 13, in, The first generation unit generates the video by rendering the combined 3D model based on a virtual viewpoint set in a virtual space where the 3D model is arranged.
15. An information processing method, comprising: A video is generated using a computer based on a three-dimensional model of the subject, created by using multiple captured images obtained through imaging the subject, and a two-dimensional image. The video simultaneously contains both the subject generated from the three-dimensional model and the two-dimensional image. The information processing method further includes: using the computer to generate a packaged image, wherein a texture image obtained by converting the three-dimensional model of the subject into two-dimensional texture information based on a virtual viewpoint and a depth image obtained by converting the depth information from the virtual viewpoint to the three-dimensional model of the subject into a two-dimensional image are packaged in one frame, and the virtual viewpoint is set in a virtual space in which the three-dimensional model is arranged.
16. A video distribution method, comprising: A three-dimensional model of the subject is generated by using multiple captured images obtained through imaging the subject; Based on the three-dimensional model and two-dimensional image of the subject, a video is generated in which the subject generated from the three-dimensional model and the two-dimensional image are simultaneously present; as well as The video is distributed to user terminals via a predetermined network. In this process, a packaged image is generated, in which a texture image obtained by converting the three-dimensional model of the subject into two-dimensional texture information based on a virtual viewpoint and a depth image obtained by converting the depth information from the virtual viewpoint to the three-dimensional model of the subject into a two-dimensional image are packaged in one frame, and the virtual viewpoint is set in a virtual space where the three-dimensional model is arranged.
17. An information processing system, comprising: An imaging device that images a subject to generate multiple captured images of the subject; An information processing device that generates a video in which both the subject generated from the three-dimensional model and the two-dimensional image are present simultaneously, based on a three-dimensional model of the subject generated by using the plurality of captured images; as well as A user terminal displays the video generated by the information processing device to the user. In this process, a packaged image is generated, in which a texture image obtained by converting the three-dimensional model of the subject into two-dimensional texture information based on a virtual viewpoint and a depth image obtained by converting the depth information from the virtual viewpoint to the three-dimensional model of the subject into a two-dimensional image are packaged in one frame, and the virtual viewpoint is set in a virtual space where the three-dimensional model is arranged.
18. A computer-readable medium having a program recorded thereon, the program causing the computer to perform a method when executed by the computer, the method comprising: A video is generated based on a three-dimensional model of the subject, created by using multiple captured images obtained through imaging the subject, and a two-dimensional image, wherein the video simultaneously contains the subject generated from the three-dimensional model and the two-dimensional image. The method further includes generating a packaged image, wherein a texture image obtained by converting the three-dimensional model of the subject into two-dimensional texture information based on a virtual viewpoint and a depth image obtained by converting the depth information from the virtual viewpoint to the three-dimensional model of the subject into a two-dimensional image are packaged in one frame, and the virtual viewpoint is set in a virtual space where the three-dimensional model is arranged.
19. A computer-readable medium having a program recorded thereon, the program causing the computer to perform a method when executed by the computer, the method comprising: A three-dimensional model of the subject is generated by using multiple captured images obtained through imaging the subject; Based on the three-dimensional model and two-dimensional image of the subject, a video is generated in which the subject generated from the three-dimensional model and the two-dimensional image are simultaneously present; as well as The video is distributed to user terminals via a predetermined network. In this process, a packaged image is generated, in which a texture image obtained by converting the three-dimensional model of the subject into two-dimensional texture information based on a virtual viewpoint and a depth image obtained by converting the depth information from the virtual viewpoint to the three-dimensional model of the subject into a two-dimensional image are packaged in one frame, and the virtual viewpoint is set in a virtual space where the three-dimensional model is arranged.
Citation Information
Patent Citations
Video generation program, video generation method, and video generation device
WO2019021375A1
Three-dimensional model distribution method, three-dimensional model receiving method, three-dimensional model distribution device, and three-dimensional model receiving device
CN110114803A
Image processing device and method
CN110998669A
Information processing apparatus, information processing method and storing unit
US20190182470A1