Information processing systems, information processing methods, and computer programs
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2022-07-22
- Publication Date
- 2026-08-03
AI Technical Summary
【0008】 本開示によれば、仮想視点画像に3次元仮想オブジェクトを適切なタイミングで表示することができる。
Smart Images

Figure 0007898975000001 
Figure 0007898975000002 
Figure 0007898975000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing system, an information processing method, a computer program, etc. for generating a virtual viewpoint image.
Background Art
[0002] Techniques for generating a virtual viewpoint image from a specified virtual viewpoint using a plurality of images obtained by imaging with a plurality of imaging devices have attracted attention. Patent Document 1 describes a method for generating a virtual viewpoint image using the three-dimensional shape of a subject estimated from a plurality of captured images obtained by arranging a plurality of imaging devices at different positions and photographing the subject.
[0003] On the other hand, a synthetic image is generated by synthesizing a three-dimensional virtual object such as background CG or an effect according to the movement of the subject into a captured image. When creating such a synthetic image, since it is necessary to display the three-dimensional virtual object in accordance with the progress of shooting and the movement of the subject, the control of the three-dimensional virtual object may sometimes have to be manual input. Therefore, the shooting time and the display time of the three-dimensional virtual object are not recorded in association with each other.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, when generating a virtual viewpoint image after shooting is completed, it is necessary to input the three-dimensional virtual object again, and the timing of the input operation does not match the shooting time, and there are cases where the three-dimensional virtual object is displayed with an image at a different timing from the image at the shooting time.
[0006] This disclosure aims to display 3D virtual objects in a virtual viewpoint image at appropriate timings. [Means for solving the problem]
[0007] One of the disclosures of The information processing system of this embodiment is A first display means for displaying a composite image obtained by combining a 3D virtual object generated without relying on a second captured image for generating a virtual viewpoint image, and a first captured image captured by an imaging device, The aforementioned 3D virtual objects and , corresponding to the display time of the composite image, and associated with the second captured image. A storage means for storing 3D virtual object information that shows the correspondence with timecode, Second image The timecode associated with it, the position of the virtual viewpoint, The aforementioned Direction of view from a virtual viewpoint and 、 A means for acquiring viewpoint information that indicates, Based on the three-dimensional virtual object information and the viewpoint information, the time code indicated in the viewpoint information corresponds The aforementioned A generation means for generating a virtual viewpoint image including a 3D virtual object, It has. [Effects of the Invention]
[0008] According to this disclosure, a 3D virtual object can be displayed in a virtual viewpoint image at an appropriate time. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows an example of the device configuration of the information processing system according to Embodiment 1. [Figure 2] This figure shows the hardware configuration of the information processing system according to Embodiment 1. [Figure 3] This diagram shows the processing flow performed by the verification image generation device of Embodiment 1. [Figure 4] This figure shows a table in which the 3D virtual object of Embodiment 1 is recorded. [Figure 5] This figure shows the processing flow performed by the virtual viewpoint image generation device of Embodiment 1. [Figure 6] This figure shows the recording process for the first embodiment. [Figure 7] This figure shows how to edit a 3D virtual object stored in the 3D virtual object storage unit of Embodiment 2. [Modes for carrying out the invention]
[0010] Embodiments of this disclosure will be described below with reference to the drawings. However, this disclosure is not limited to the embodiments described below. In each drawing, the same reference numeral is used for the same member or element, and redundant descriptions are omitted or simplified.
[0011] <Embodiment 1> Figure 1 shows an example of the device configuration of an information processing system that generates virtual viewpoint images according to this embodiment. This system is composed of, for example, an imaging unit 101, a synchronization unit 102, a 3D shape estimation unit 103, a storage unit 104, a 3D virtual object storage unit 105, a confirmation imaging unit 115, a virtual viewpoint image generation device 116, and a confirmation image generation device 117. The virtual viewpoint image generation device 116 is a device that includes a viewpoint indication unit 106, a subject image generation unit 107, a background image generation unit 108, an image synthesis unit 109, and an output image display unit 110, and is, for example, a tablet terminal, a smartphone, or an image generation device with a joystick. The confirmation image generation device 117 is a device that includes a 3D virtual object manipulation unit 111, a confirmation background image generation unit 112, a chroma key synthesis unit 113, and a confirmation display unit 114, and is, for example, a tablet terminal, a smartphone, or a laptop computer.
[0012] This system may consist of one electronic device or multiple electronic devices.
[0013] An information processing system is a system that generates a virtual viewpoint image representing a scene from a specified virtual viewpoint based on a plurality of images captured by a plurality of imaging devices and the specified virtual viewpoint. The virtual viewpoint image in the present embodiment is also called a free viewpoint image, but is not limited to an image corresponding to a viewpoint freely (arbitrarily) specified by the user. For example, an image corresponding to a viewpoint selected by the user from a plurality of candidates is also included in the virtual viewpoint image. In the present embodiment, the case where the virtual viewpoint is specified by a user operation will be mainly described, but the virtual viewpoint may be automatically specified based on the result of image analysis or the like. In the present embodiment, the case where the virtual viewpoint image is a moving image will be mainly described, but the virtual viewpoint image may be a still image.
[0014] The viewpoint information used for generating the virtual viewpoint image is information indicating the position and orientation (line-of-sight direction) of the virtual viewpoint. Specifically, the viewpoint information is a parameter set including parameters representing the three-dimensional position of the virtual viewpoint and parameters representing the orientation of the virtual viewpoint in the pan, tilt, and roll directions. Note that the content of the viewpoint information is not limited to the above. For example, the parameter set as the viewpoint information may include a parameter representing the size (angle of view) of the field of view of the virtual viewpoint. Also, the viewpoint information may have a plurality of parameter sets. For example, the viewpoint information may have a plurality of parameter sets respectively corresponding to a plurality of frames constituting a moving image of the virtual viewpoint image, and may be information indicating the position and orientation of the virtual viewpoint at each of a plurality of consecutive time points.
[0015] The information processing system has a plurality of imaging devices that image an imaging area from a plurality of directions. The imaging area is, for example, an arena where competitions such as soccer or karate are held, or a stage where concerts or plays are performed. The plurality of imaging devices are installed at different positions so as to surround such an imaging area and perform imaging synchronously. Note that the plurality of imaging devices do not necessarily need to be installed over the entire circumference of the imaging area, and may be installed only in a part of the periphery of the imaging area depending on restrictions on the installation location or the like. Also, in the present embodiment, the description will be made on the premise that the present system is used in a studio that shoots a virtual viewpoint image configured with 360-degree green screen. Also, imaging devices having different functions such as a telephoto camera and a wide-angle camera may be installed.
[0016] The virtual viewpoint image is generated, for example, by the following method. First, a plurality of images (plurality of viewpoint images) are acquired by imaging from different directions by a plurality of imaging devices. Next, a foreground image obtained by extracting a foreground area corresponding to a predetermined object such as a person or a ball and a background image obtained by extracting a background area other than the foreground area are acquired from the plurality of viewpoint images. Also, a foreground model representing the three-dimensional shape of a predetermined object and texture data for coloring the foreground model are generated based on the foreground image, and texture data for coloring a background model representing the three-dimensional shape of the background such as an arena is generated based on the background image. Then, the texture data is mapped to the foreground model and the background model, and rendering is performed according to the virtual viewpoint indicated by the viewpoint information, whereby a virtual viewpoint image is generated. However, the method for generating the virtual viewpoint image is not limited to this, and various methods can be used, such as a method for generating a virtual viewpoint image by projective transformation of an imaging image without using a three-dimensional model.
[0017] A foreground image is an image extracted from an image captured by an imaging device, specifically the area of an object (foreground region). An object extracted as the foreground region refers to a dynamic object (a moving body) that is in motion (its absolute position and shape may change) when images are taken from the same direction over time. Examples of such objects include athletes and referees on the field during a sporting event, a ball in a ball game, or singers, musicians, performers, and presenters in a concert or entertainment event.
[0018] A background image is an image of an area (background region) that is at least different from the foreground object. Specifically, a background image is an image obtained by removing the foreground object from the captured image. The background refers to an object that remains stationary or nearly stationary when images are taken from the same direction over time. Such objects include, for example, a stage for a concert, a stadium for an event such as a sports competition, structures such as goals used in ball games, and a field. However, the background is at least an area different from the foreground object, and the object being captured may include other objects besides the foreground object and background.
[0019] A virtual camera is a virtual camera distinct from the multiple imaging devices actually installed around the imaging area, and is a concept used to conveniently explain the virtual viewpoint involved in the generation of virtual viewpoint images. In other words, a virtual viewpoint image can be considered an image captured from a virtual viewpoint set in a virtual space associated with the imaging area. The position and orientation of the virtual viewpoint in the said image can be represented as the position and orientation of the virtual camera. In other words, a virtual viewpoint image can be said to be an image that simulates the image captured by a camera, assuming that the camera exists at the position of the virtual viewpoint set in space. In this embodiment, the content of the changes in the virtual viewpoint over time is referred to as the virtual camera path. However, it is not essential to use the concept of a virtual camera to realize the configuration of this embodiment. That is, it is sufficient that at least information representing a specific position in space and information representing an orientation are set, and a virtual viewpoint image is generated according to the set information.
[0020] Each imaging unit 101 is a camera with its own independent housing capable of capturing images from a single viewpoint. However, it is not limited to this, and two or more imaging devices may be configured within the same housing. For example, multiple imaging devices may be installed, each being a single camera equipped with multiple lens groups and multiple sensors capable of capturing images from multiple viewpoints. Alternatively, the imaging unit 101 may have no lenses but have an image sensor. In this case, interchangeable lenses may be used with the imaging unit 101 to capture images of the subject onto the image sensor.
[0021] The synchronization unit 102 controls the timing at which the imaging unit 101 and the confirmation imaging unit 115 capture images. In other words, it controls the timing at which the imaging unit 101 and the confirmation imaging unit 115 capture images.
[0022] The 3D shape estimation unit 103 generates 3D model data of the subject using multiple images captured by the imaging unit 101. Specifically, the 3D shape estimation unit 103 generates 3D model data represented by a known representation method. The 3D model data may be point cloud data composed of points, mesh data composed of polygons, or voxel data composed of voxels.
[0023] The storage unit 104 stores the following material data as material used to generate a virtual viewpoint image. In this embodiment, the data for generating a virtual viewpoint image, which includes captured images taken by the imaging device and data generated based on said captured images, is referred to as material data. The material data generated based on the captured images includes, for example, foreground image data extracted from the captured images, 3D model data representing the shape of objects in the virtual space, and texture data for coloring the 3D model. In this embodiment, the 3D model data is assumed to be 3D model data generated by the 3D shape estimation unit 103, but it may also be a pre-created 3D model. The type of material data is not limited as long as it is data for generating a virtual viewpoint image. For example, camera parameters representing the imaging conditions of the imaging device that acquires the captured image may be included in the material data. Furthermore, although the above describes an example of material data when using a method to generate a virtual viewpoint image by generating a 3D model, it is not limited to this. When generating a virtual viewpoint image using an image-based rendering method that does not use a 3D model, the data required to generate the virtual viewpoint image may differ from the example of material data described above. Thus, the material data may differ depending on the method for generating the virtual viewpoint image. The source data is stored in association with the shooting time generated by the imaging unit 101.
[0024] The 3D virtual object storage unit 105 stores pre-created 3D virtual objects. These 3D virtual objects include background CG (background models) and effects, and their positions and orientations in the virtual space are predetermined. Furthermore, 3D virtual objects include objects whose shape and color change over time. Specifically, the storage includes a 3D virtual object (background model) representing the background that is always displayed as the background image, and a 3D virtual object (effect) representing an effect that is displayed temporarily. The time at which the 3D virtual objects are displayed in a series of shots is specified in the processing flow shown in Figure 3, which will be described later. In this embodiment, it is assumed that the time at which the 3D virtual objects are displayed is not specified before the processing flow shown in Figure 3 is executed. Also, in this embodiment, it is assumed that the user can change the 3D virtual objects, including the background model and effects, at any time during shooting. This allows the user to input instructions to switch the 3D virtual objects through the 3D virtual object operation unit 111, and have these instructions reflected in the captured image. However, it is not limited to this; 3D virtual objects that are background CG may also be stored in the storage unit 104. In that case, the background is a background model that does not change due to user operation.
[0025] The viewpoint instruction unit 106 is a viewpoint operation unit that operates a virtual viewpoint, which is a physical user interface such as a joystick or touch panel. Based on the input from the viewpoint operation unit, it generates virtual viewpoint information and outputs the generated virtual viewpoint information to the subject image generation unit 107 and the background image generation unit 108. In this embodiment, the virtual viewpoint information consists of information corresponding to external camera parameters such as the position and orientation of the virtual viewpoint, information corresponding to internal camera parameters such as focal length and angle of view, and time information that specifies the time of capture of the captured image captured by the imaging unit 101. The captured image captured by the imaging unit 101 is an image for generating a virtual viewpoint image, and the time of capture corresponds to a time code. Note that it is not limited to a time code, and may also be the number of frames at the time of capture.
[0026] The subject image generation unit 107 acquires data for the corresponding shooting time from the storage unit 104 based on the time information included in the input virtual viewpoint information. The subject image generation unit 107 places a 3D model of the subject in the virtual space from the acquired data and generates a subject image that depicts the subject from the input virtual viewpoint, and outputs it to the image synthesis unit 109. If the position and line of sight of the virtual viewpoint coincide with that of the imaging unit 101, the foreground image extracted from the image captured by the imaging unit 101 may be output to the image synthesis unit 109. In this case, the subject image output will be an image with transparency in parts other than the subject.
[0027] The background image generation unit 108 acquires data for the time of capture from the 3D virtual object storage unit 105 based on the time information included in the input virtual viewpoint information. From the acquired data, it places 3D virtual objects in the virtual space and outputs a background image, which is an image of the virtual space viewed from the virtual viewpoint specified by the viewpoint instruction unit 106, to the image synthesis unit 109.
[0028] The image synthesis unit 109 synthesizes the input subject image and background image. Since the background portion of the subject image is transparent, this transparency is used in the synthesis process, and the result is output to the output image display unit 9. Alternatively, distance information from a virtual viewpoint may be added to both the subject image and the background image, and this information may be used to render the image closer to the virtual viewpoint. Since both the subject image and the background image are images viewed from a virtual viewpoint, the synthesized image produced by the image synthesis unit 109 is a composite image created by combining two virtual viewpoint images.
[0029] The output image display unit 110 displays the composite image created by the image synthesis unit 109.
[0030] The 3D virtual object operation unit 111 is an operation unit that has, for example, a touch panel or a keyboard. By receiving input to the operation unit, it outputs an instruction to combine the 3D virtual object stored in the 3D virtual object storage unit 105 with the image acquired from the confirmation imaging unit 115. Specifically, it specifies a particular 3D virtual object from the 3D virtual object storage unit 105 and sends an instruction to the confirmation background image generation unit 108 to combine it with the image acquired from the confirmation imaging unit 115. At this time, multiple 3D virtual objects may be specified simultaneously. It may also be specified the position in which the 3D virtual object will be combined. In this embodiment, the 3D virtual object stored in the 3D virtual object storage unit is assumed to have predetermined position information in virtual space, but the stored position information may be updated in response to an operation to specify the position of the 3D virtual object.
[0031] The verification background image generation unit 112 receives input from the 3D virtual object manipulation unit 111 and receives the corresponding 3D virtual object from the 3D virtual object storage unit 105. Next, it places the received 3D virtual object in the virtual space and generates a verification background image taken from the position and orientation of the verification imaging unit 115 converted to the virtual space. The generated verification background image is transmitted to the chroma key compositing unit 113.
[0032] The chroma key compositing unit 113 performs chroma key compositing using the captured image received from the confirmation imaging unit 115 and the confirmation background image received from the confirmation background image generation unit 112. The generated composite image is output to the confirmation display unit 114.
[0033] The confirmation display unit 114 displays the composite image received from the chroma key compositing unit 113. If there is no input from the 3D virtual object manipulation unit, it displays the captured image received from the confirmation imaging unit 115. Note that the captured image does not have to be received directly from the confirmation imaging unit 115, but may be received via the chroma key compositing unit 113.
[0034] The confirmation imaging unit 115 is a camera capable of capturing images from a single viewpoint. In this embodiment, it is a different camera from the imaging unit 101, but it is not limited to this. One of the multiple imaging units 101 may be designated as the confirmation imaging unit 115.
[0035] The virtual viewpoint image generation device 116 is a device having a viewpoint operation unit that operates a virtual viewpoint, which is a physical user interface such as a joystick or a jog dial, and generates and displays a virtual viewpoint image based on the virtual viewpoint information output from the viewpoint operation unit.
[0036] The confirmation image generation device 117 is a device that composites a 3D virtual object specified by the user onto the captured image received from the confirmation imaging unit 115 and displays it. For example, it is a device for a laptop computer or tablet terminal where a user supervising the shooting can view the captured image being taken by the confirmation imaging unit 115 and give instructions to composite the 3D virtual object at the appropriate timing. In this embodiment, the confirmation image is treated as an image for the user to give instructions to composite the 3D virtual object, but it is not limited to this. The confirmation image can also be used as an image for distribution.
[0037] The virtual viewpoint image generation device 116 and the confirmation image generation device 117 are not limited to the above configuration. For example, the virtual viewpoint image generation device 116 may have a storage unit 104, and the confirmation image generation device 117 does not need to have a 3D virtual object manipulation unit 111. The virtual viewpoint image generation device 116 may not have a viewpoint indication unit 106 and may receive virtual viewpoint information from another device. Also, the virtual viewpoint image generation device 116 and the confirmation image generation device 117 may each have a 3D virtual object storage unit 105.
[0038] Figure 2 shows the hardware configuration of the virtual viewpoint image generation device 116. The virtual viewpoint image generation device 116 includes a CPU 211, ROM 212, RAM 213, auxiliary storage device 214, display unit 215, operation unit 216, communication I / F 217, and bus 218. The confirmation image generation device 117 has a similar hardware configuration, so its description is omitted.
[0039] The CPU 211 controls the entire virtual viewpoint image generation device 116 using computer programs and data stored in the ROM 212 and RAM 213, thereby realizing each function of the virtual viewpoint image generation device 116 shown in Figure 1. The virtual viewpoint image generation device 116 may have one or more dedicated hardware components separate from the CPU 211, and at least a portion of the processing performed by the CPU 211 may be executed by the dedicated hardware. Examples of dedicated hardware include ASICs (Application-Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), and DSPs (Digital Signal Processors). The ROM 212 stores programs and other data that do not require modification. The RAM 213 temporarily stores programs and data supplied from the auxiliary storage device 214, as well as data supplied externally via the communication interface 217. The auxiliary storage device 214 is, for example, a hard disk drive and stores various types of data such as image data and audio data.
[0040] The display unit 215 is composed of, for example, a liquid crystal display or LEDs, and displays a GUI (Graphical User Interface) for the user to operate the virtual viewpoint image generation device 116. The operation unit 216 is composed of, for example, a keyboard, mouse, joystick, touch panel, etc., and receives various instructions from the user and inputs them to the CPU 211. The CPU 211 operates as a display control unit that controls the display unit 215 and an operation control unit that controls the operation unit 216. The communication interface 217 is used for communication between the virtual viewpoint image generation device 116 and external devices. For example, if the virtual viewpoint image generation device 116 is connected to an external device by wire, a communication cable is connected to the communication interface 217. If the virtual viewpoint image generation device 116 has a function to communicate wirelessly with an external device, the communication interface 217 is equipped with an antenna. The bus 218 connects the various parts of the virtual viewpoint image generation device 116 and transmits information.
[0041] In this embodiment, the display unit 215 and the operation unit 216 are assumed to be located inside the virtual viewpoint image generation device 116, but at least one of the display unit 215 and the operation unit 216 may exist as a separate device outside the virtual viewpoint image generation device 116.
[0042] Figure 3 shows the processing flow performed by the confirmation image generation device 117. This processing flow is executed every frame until the capture is complete.
[0043] In step S301, the confirmation image generation device 117 determines whether it has received input information from the 3D virtual object operation unit 111. The input information is an input specifying a particular 3D virtual object from a plurality of 3D virtual objects pre-stored in the 3D virtual object storage unit 105, and an input displaying that 3D virtual object on the captured image. In this embodiment, the particular 3D virtual object is described as an effect whose shape and color change over time, but it is not limited to this. If the particular 3D virtual object is a background model, the timing of the background model switching is received as input information. Also, if the particular 3D virtual object is a background model or effect whose shape and color change due to user operation, the timing of the shape and color change is received as input information. In this embodiment, the 3D virtual object operation unit 111 is operated by the user and the input information is received. If no input information has been received, the process proceeds to step S302. If input information has been received, the process proceeds to step S303.
[0044] In step S302, the confirmation image generation device 117 displays the captured image received from the confirmation imaging unit 115. Then, the process proceeds to step S308.
[0045] In step S303, the verification image generation device 117 receives a 3D model of a specific 3D virtual object from the 3D virtual object storage unit 105 based on the input information received in step S301, and places it in the virtual space.
[0046] In step S304, the confirmation image generation device 117 receives position and orientation information of the confirmation imaging unit 115 from the confirmation imaging unit 115. Based on the received information, it converts the position and orientation of the confirmation imaging unit 115 into a position and orientation in virtual space, and generates an image of a 3D virtual object viewed from the corresponding position and orientation in the virtual space of the confirmation imaging unit 114.
[0047] In step S305, the confirmation image generation device 117 transmits the image created in S304 to the chroma key compositing unit 113, where it is combined with the captured image acquired from the confirmation imaging unit 115. Specifically, in this embodiment, it is assumed that images captured in a studio with a green screen are to be combined, so the green portion of the captured image acquired from the confirmation imaging unit 115 is combined with the image created in S304 to create a composite image. The composite image is then transmitted to the confirmation display unit 112.
[0048] In step S306, the confirmation image generation device 117 displays the composite image created in S305.
[0049] In step S307, the verification image generation device 117 stores in the 3D virtual object storage unit 105 the specified 3D virtual object and the corresponding shooting time in S301. The information stored in the 3D virtual object storage unit 105 in association with the specified 3D virtual object is not limited to this, and may also be the time when the input information was received from the 3D virtual object operation unit 111. Alternatively, it may be the time when the 3D virtual object operation unit 111 received operation information, or the time when the composite image was created by the chroma key compositing unit 113. It may also be the time when instructions were given by the user, or the display time when the composite image was displayed. It may also be the time code associated with the captured image captured by the verification imaging unit 115. The 3D virtual object information showing the correspondence between the specified 3D virtual object and the above-mentioned time or time code is stored in the 3D virtual object storage unit 105. The process then proceeds to S308.
[0050] In step S308, the confirmation image generation device 117 determines whether or not to terminate the shooting. If the shooting is to be terminated, the above process is terminated. If the shooting is not to be terminated, the process proceeds to S301. This loop process is executed for all frames until the shooting is completed.
[0051] This process associates the 3D virtual object with the time of capture and stores it in the 3D virtual object storage unit 105. Furthermore, when displaying the background image for verification, the verification imaging unit 115 captures the subject in the studio, and the chroma key compositing unit 113 performs chroma key compositing on this captured image and the verification background image, displaying it on the verification display unit 114.
[0052] Figure 6 shows the recording process for a shoot using the above-described process. In this embodiment, a display device 601 is placed in the studio. The image displayed on the display device 601 is the same image as the image displayed on the confirmation display unit 114. In this way, the subject (performer) can check in real time what kind of 3D virtual object they are currently being composited with and what kind of effects are being applied.
[0053] Figure 4 shows an example of data stored in the 3D virtual object storage unit 105, where a 3D virtual object and the time of capture are associated as a result of the process described in Figure 3. In this embodiment, a 3D virtual object (background model) that represents the background always displayed as a background image and a 3D virtual object that represents an effect that is displayed temporarily are stored. The background model is stored in association with the time of capture by the user's operation on the 3D virtual object operation unit 111. In other words, the time at which the user switches to another background model is stored. The effects are also stored in association with the time of capture by the user's operation on the 3D virtual object operation unit 111. Note that multiple effects may be associated with the same time of capture. Furthermore, each effect has a display time, and is displayed from the time of capture associated by the user's operation until the display time. In the example in Figure 4, a switch from background model 1 to background model 2 occurs at 12:11:00, effect 1 is displayed for 0.5 seconds from 11:30:00, and effect 2 is displayed for 3 seconds from 13:20:03. The effects here include explosion effects, as well as visual representations of correct / incorrect answers in quizzes. The position and orientation of the effects in the virtual space are predetermined (not illustrated).
[0054] Figure 5 shows the processing flow performed by the virtual viewpoint image generation device 116. This processing flow is performed after the 3D virtual object and the imaging time have been stored in the 3D virtual object storage unit 105 in association with each other, according to the processing flow described in Figure 3.
[0055] In step S501, the virtual viewpoint image generation device 116 searches the 3D virtual object storage unit for a 3D virtual object corresponding to the time of capture, based on the virtual viewpoint information received from the viewpoint indication unit 106.
[0056] In step S502, the virtual viewpoint image generation device 116 receives a 3D virtual object corresponding to the shooting time from the 3D virtual object storage unit 105 and places it in the virtual space. Next, in the virtual space, it generates a virtual viewpoint image (background image) that represents the view from the virtual viewpoint received from the viewpoint instruction unit 106. The generated background image is transmitted to the image synthesis unit 109.
[0057] In step S503, the virtual viewpoint image generation device 116 places a foreground model representing the subject in the virtual space from the material data received from the storage unit 104, and generates a virtual viewpoint image (subject image) as seen from the virtual viewpoint received from the viewpoint indication unit 106. The generated subject image is transmitted to the image synthesis unit 109.
[0058] In step S504, the virtual viewpoint image generation device 116 combines the subject image received from the subject image generation unit 107 and the background image received from the background image generation unit 108 to generate a virtual viewpoint image. The generated virtual viewpoint image is transmitted to the output image display unit 110.
[0059] In step S505, the virtual viewpoint image generation device 116 displays the virtual viewpoint image received from the image synthesis unit 109.
[0060] In step S506, the virtual viewpoint image generation device 116 determines whether to terminate the editing process for generating the virtual viewpoint image. If the process for generating the virtual viewpoint image is completed, this process is terminated. If not, the process returns to S501. Steps S501 to S505 are repeated until the editing process is completed.
[0061] The process described in Figure 5 above allows the 3D virtual object, which has been associated with the imaging time by the process described in Figure 3, to be displayed in the virtual viewpoint image.
[0062] As described above, a 3D virtual object input during shooting can be displayed in the virtual viewpoint image at the same time as the shooting.
[0063] <Embodiment 2> In this embodiment, a second 3D virtual object operation unit (not shown) equipped with a display unit is connected to the background image generation unit 108, and information regarding the 3D virtual object is displayed on the second 3D virtual object operation unit 700. The second 3D virtual object operation unit is an operation unit for editing data stored in the 3D virtual object storage unit to generate a virtual viewpoint image.
[0064] Figure 7 shows a graphical user interface for manipulating 3D virtual objects in the second background manipulation section. Specifically, there is a timeline 701 extending chronologically, as shown at the bottom of Figure 7, and within it is a seek bar 702 that indicates the time (frame) currently being played back by the editing process. This timeline 701 is divided into columns according to the classification of the 3D virtual objects, and the actual shooting time associated with the 3D virtual object may be displayed as a keyframe 703. In this case, the classification of the 3D virtual objects refers to classifications such as background models and effect patterns.
[0065] Furthermore, as shown in Figure 7, checkboxes 704 may be provided for each column of the timeline 701, allowing users to select whether or not to apply each category during playback. Alternatively, the shooting time displayed as a keyframe 703 may be sought, and the shooting time stored in the 3D virtual object storage unit 105 may be edited later. A user interface may also be provided that allows enabling / disabling or deleting each keyframe.
[0066] In this embodiment, the subject image generation unit 107 is described as making the background portion of the subject image transparent, but this is not necessarily the only option. For example, the background portion of the subject image may also be made a single color, such as green, and then the image synthesis unit 109 may use chroma keying to synthesize it.
[0067] In this embodiment, the subject image generation unit 107, which generates the subject image, and the background image generation unit 108, which generates the background image, are represented as separate components and combined by the image synthesis unit 109. However, the embodiment is not limited to this configuration. The subject image generation unit 107 and the background image generation unit 108 may be implemented in a single image generation unit. In that case, the image synthesis unit 109 would not be necessary.
[0068] In this embodiment, the storage unit 104 and the 3D virtual object storage unit 105 were described as separate components, but they may be a single storage unit.
[0069] In this embodiment, the output destination for the virtual viewpoint image is the output image display unit 110, but it does not necessarily have to be a display device. For example, the virtual viewpoint image may be output to an image recording device or an image distribution device.
[0070] While it is stated that the 3D virtual objects used in the background image generation unit 108 and the verification background image generation unit 112 are the same 3D virtual objects stored in the 3D virtual object storage unit 105, this is not necessarily the only option. Since the verification image is solely for verification purposes, it is sufficient if the object placement and effect display can be confirmed, and a simple 3D virtual object background may be used.
[0071] In this embodiment, the image captured by the confirmation shooting unit 114 and the confirmation background image are combined in the chroma key compositing unit 113 and displayed in the confirmation display unit 114. However, the system is not necessarily limited to this configuration. For example, the confirmation background image generated by the confirmation background image generation unit 112 may be displayed directly on the confirmation display unit 114.
[0072] In this embodiment, the configuration has been described in which the user manipulates a 3D virtual object using a 3D virtual object manipulation unit 111, but the system is not necessarily limited to this. For example, instead of the 3D virtual object manipulation unit 111, a 3D virtual object manipulation information input unit (not shown) that accepts external signal input may be used. In this case, for example, it may be connected to sound equipment, and the sound equipment outputs SE (sound effect), and some signal is input to the 3D virtual object manipulation input unit to realize background manipulation in accordance with the SE. Although sound equipment has been used here, the equipment to be connected is not limited to other equipment such as lighting equipment in this disclosure.
[0073] Furthermore, in this embodiment, some or all of the control may be performed by supplying a computer program that realizes the functions of the above-described embodiment to an information processing system, etc., via a network or various storage media. The computer (or CPU, MPU, etc.) in the information processing system, etc., may then read and execute the program. In that case, the program and the storage media storing the program constitute the present disclosure.
[0074] Furthermore, the disclosure of this embodiment includes the following configuration, method, and program.
[0075] (Configuration 1) A storage means for storing 3D virtual object information that shows the correspondence between a 3D virtual object superimposed on an image captured by an imaging device and a time code associated with the image, An acquisition means for acquiring a time code associated with an image for generating a virtual viewpoint image, the position of the virtual viewpoint, and viewpoint information indicating the direction of the line of sight from the virtual viewpoint. A system characterized by having a generation means that generates a virtual viewpoint image including a 3D virtual object corresponding to a time code shown in the viewpoint information, based on the 3D virtual object information and the viewpoint information.
[0076] (Configuration 2) Furthermore, it has an input means for inputting the display of the three-dimensional virtual object in the image captured by the imaging device, The system according to configuration 1, characterized in that the three-dimensional virtual object information is acquired based on operation information input by the user to an input device.
[0077] (Configuration 3) The system according to Configuration 1 or 2, characterized in that the three-dimensional virtual object information includes information indicating the display time.
[0078] (Configuration 4) The system according to any one of Configurations 1 to 3, characterized in that the three-dimensional virtual object information has position information indicating a position in virtual space.
[0079] (Configuration 5) The system according to Configuration 4, further comprising a modification means for modifying at least one of the three-dimensional virtual object information and the position information of the three-dimensional virtual object stored by the storage means.
[0080] (Configuration 6) The system according to any one of Configurations 1 to 5, further comprising a display means for displaying the image captured by the imaging device and the three-dimensional virtual object.
[0081] (Configuration 7) The system according to Configuration 6, characterized in that the display means displays a three-dimensional virtual object stored by the storage means.
[0082] (Configuration 8) The system according to Configuration 6 or 7, characterized in that the display means displays a seek bar for changing the time code associated with the image which is in a relationship with the three-dimensional virtual object stored by the storage means.
[0083] (Configuration 9) The system according to any one of Configurations 6 to 8, characterized in that the display means has a user interface for setting whether or not to display the three-dimensional virtual object in the virtual viewpoint image.
[0084] (Configuration 10) The system according to any one of Configurations 1 to 9, wherein the three-dimensional virtual object is a CG or effect for background purposes.
[0085] (Configuration 11) The system according to any one of Configurations 1 to 10, characterized in that the three-dimensional virtual object is modified in any one of its shape and color based on operation information input by the user to an input device.
[0086] (Configuration 12) The system according to Configurations 1 to 11, characterized in that the virtual viewpoint image is generated based on a plurality of images captured over time by a plurality of imaging devices.
[0087] (Configuration 13) Acquisition means for acquiring viewpoint information including a 3D virtual object to be composited into an image captured by an imaging device, 3D virtual object information associated with the time of acquisition when an input for displaying the 3D virtual object was received, the time of generating a virtual viewpoint image, the position of the virtual viewpoint, and the direction of line of sight from the virtual viewpoint. The apparatus is characterized by having a generation means that generates a virtual viewpoint image including the 3D virtual object, which is associated with the shooting time corresponding to the time the virtual viewpoint image is generated, based on the 3D virtual object information and the viewpoint information.
[0088] (Method) A storage step of storing 3D virtual object information that shows the correspondence between a 3D virtual object synthesized in an image captured by an imaging device and a time code associated with the image, An acquisition process to obtain a time code associated with an image for generating a virtual viewpoint image, viewpoint information indicating the position of the virtual viewpoint and the direction of line of sight from the virtual viewpoint, A method characterized by comprising a generation step of generating a virtual viewpoint image including a 3D virtual object corresponding to a time code indicated in the viewpoint information, based on the 3D virtual object information and the viewpoint information.
[0089] (Program) A program for causing a computer to function as one of the means of the system described in any one of the configurations 1 to 12. [Explanation of symbols]
[0090] 101 Imaging Unit 105 Viewpoint Indicator 107 Background Image Generation Unit 110 3D Virtual Object Manipulation Unit 113 3D Virtual Object Storage Unit 114 Confirmation imaging unit
Claims
1. A first display means for displaying a composite image obtained by combining a three-dimensional virtual object generated independently of a second captured image for generating a virtual viewpoint image and a first captured image captured by an imaging device, A storage means for storing three-dimensional virtual object information that shows the correspondence between the three-dimensional virtual object and a time code associated with the second captured image that corresponds to the display time of the composite image, Acquisition means for acquiring viewpoint information indicating a time code associated with the second captured image, the position of a virtual viewpoint, and the direction of line of sight from the virtual viewpoint, A generation means that generates a virtual viewpoint image including the three-dimensional virtual object corresponding to the time code shown in the viewpoint information, based on the three-dimensional virtual object information and the viewpoint information, An information processing system characterized by having the following features.
2. Furthermore, it has an input means for receiving input for displaying the composite image, The information processing system according to claim 1, characterized in that the three-dimensional virtual object information is acquired based on the input.
3. The information processing system according to claim 1, characterized in that the three-dimensional virtual object information includes information indicating the display time.
4. The information processing system according to claim 1, characterized in that the three-dimensional virtual object information has position information indicating a position in a virtual space.
5. Furthermore, the information processing system according to claim 4 is characterized by having a modification means for modifying at least one of the three-dimensional virtual object information and the position information of the three-dimensional virtual object stored by the storage means.
6. The information processing system according to claim 1, characterized in that the first captured image and the second captured image are different captured images.
7. The information processing system according to claim 1, further characterized in that the virtual viewpoint image is displayed on a second display means different from the first display means.
8. The information processing system according to claim 7, characterized in that the second display means displays a seek bar for changing the time code associated with the second captured image which is in a relationship with the three-dimensional virtual object in the three-dimensional virtual object information stored by the storage means.
9. The information processing system according to claim 7, characterized in that the second display means displays a user interface for setting whether or not to display the three-dimensional virtual object in the virtual viewpoint image.
10. The information processing system according to claim 1, characterized in that the three-dimensional virtual object is a CG or effect for the background.
11. The information processing system according to claim 2, characterized in that the three-dimensional virtual object is modified in at least one of its shape and color based on the input.
12. The information processing system according to claim 1, characterized in that the virtual viewpoint image is generated based on a plurality of second captured images captured over time by a plurality of imaging devices.
13. A first display step of displaying a composite image obtained by combining a three-dimensional virtual object generated without relying on a second captured image for generating a virtual viewpoint image and a first captured image captured by an imaging device, A storage step for storing three-dimensional virtual object information that shows the correspondence between the three-dimensional virtual object and the time code associated with the second captured image, which corresponds to the display time of the composite image. An acquisition step to acquire viewpoint information indicating a time code associated with the second captured image, the position of a virtual viewpoint, and the direction of line of sight from the virtual viewpoint, A generation step of generating a virtual viewpoint image including the three-dimensional virtual object corresponding to the time code shown in the viewpoint information, based on the three-dimensional virtual object information and the viewpoint information, An information processing method characterized by having the following features.
14. A program for causing a computer to function as one of the means of the information processing system described in any one of claims 1 to 12.