Image processing device, image processing method and program
The image processing device efficiently generates virtual viewpoint images by predicting subject positions and selecting optimal imaging devices, addressing delays and inefficiencies in existing technologies, especially with multiple subjects, and ensuring real-time image generation.
Patent Information
- Application Number
- JP2022078716
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-12
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-05-12
AI Technical Summary
Existing technologies face delays in generating virtual viewpoint images due to the need to switch and read captured images from multiple imaging devices when the virtual viewpoint changes, especially when multiple subjects are present, leading to inefficiencies in processing time.
An image processing device that includes an information acquisition means for virtual viewpoint information, a model acquisition means for three-dimensional modeling, a viewpoint prediction means, a model prediction means for subject positioning, a determination means for selecting imaging devices, and an image generation means to generate virtual viewpoint images based on predicted positions and parameters.
Enables the generation of virtual viewpoint images in real-time even with multiple subjects, reducing processing delays and optimizing communication bandwidth and processing load by selecting appropriate imaging devices for image capture.
Smart Images

Figure 0007775140000001 
Figure 0007775140000002 
Figure 0007775140000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technology for generating a virtual viewpoint image based on captured images acquired by a plurality of imaging devices. [Background technology]
[0002] In recent years, attention has been focused on a technology that synchronizes multiple imaging devices installed at different locations to capture multiple viewpoints, and then uses the captured images from the multiple viewpoints to generate images that appear to be captured not only from the installation locations of the imaging devices but also from any virtual viewpoint. The generation of virtual viewpoint images is achieved by consolidating the captured images from multiple viewpoints captured by the multiple imaging devices in an image processing device such as a server, and then having the image processing device perform processing such as rendering based on the desired virtual viewpoint. This technology for generating virtual viewpoint images can produce video content with powerful viewpoints from captured images of, for example, dance or acting. For example, a user viewing content can set a virtual viewpoint of their own choosing, allowing the user to freely move the viewpoint, providing a greater sense of realism to the user than conventional captured images that do not generate virtual viewpoint images.
[0003] Here, the placement positions of the multiple imaging devices correspond to respective positions in the virtual space. When generating a virtual viewpoint image for a virtual viewpoint at a position different from the placement positions of the imaging devices, images captured by an imaging device placed close to the virtual viewpoint are used. In other words, the imaging device that acquires the images required to generate the virtual viewpoint image differs depending on the position of the virtual viewpoint. Therefore, for example, when the virtual viewpoint is moved, the imaging device that acquires the imaging device used to generate the virtual viewpoint image also changes one after another. In this case, the captured images required to generate the virtual viewpoint image in response to the movement of the virtual viewpoint are sequentially switched and read out from the captured images for each imaging device collected in the server database, which takes time to generate the virtual viewpoint image, resulting in delays.
[0004] Patent Document 1 discloses a technology for calculating a predicted virtual viewpoint based on a virtual viewpoint related to a virtual viewpoint image, obtaining images necessary for generating a virtual viewpoint image according to the predicted virtual viewpoint from a storage that stores images captured by a plurality of imaging devices, and generating the virtual viewpoint image from the images. The technology disclosed in Patent Document 1 makes it possible to reduce the time required to generate a virtual viewpoint image. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 2019-79468 Summary of the Invention [Problem to be solved by the invention]
[0006] However, in the technology described in Patent Document 1, the images required to generate a virtual viewpoint image are determined based solely on a prediction of the virtual viewpoint, and therefore, there may be cases where a virtual viewpoint image cannot be generated, for example, when there are multiple subjects.
[0007] Therefore, an object of the present disclosure is to make it possible to generate a virtual viewpoint image even when multiple subjects are present. [Means for solving the problem]
[0008] The image processing device of the present disclosure is characterized by having an information acquisition means for acquiring virtual viewpoint information indicating the position and direction of a virtual viewpoint; a model acquisition means for acquiring a three-dimensional model of a subject generated based on captured images captured by a plurality of imaging devices; a viewpoint prediction means for predicting a virtual viewpoint in a second frame subsequent to a first frame based on a virtual viewpoint of a frame preceding a first frame in the virtual viewpoint image; a model prediction means for predicting a position of a three-dimensional model of a subject in a second frame based on a position of a three-dimensional model of the subject corresponding to a frame preceding the first frame; a determination means for determining an imaging device from among the plurality of imaging devices for acquiring captured images to be used when generating the second frame based on the predicted virtual viewpoint, the predicted position of the three-dimensional model, and shooting parameters of the plurality of imaging devices; and an image generation means for generating a virtual viewpoint image based on the captured image corresponding to the second frame acquired by the determined imaging device, the three-dimensional model corresponding to the second frame acquired by the model acquisition means, and virtual viewpoint information corresponding to the second frame acquired by the information acquisition means. [Effects of the Invention]
[0009] According to the present disclosure, a virtual viewpoint image can be generated even when multiple subjects are present. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram illustrating a schematic configuration of an image processing system. [Figure 2] FIG. 1 is a diagram illustrating an example of installation of a plurality of imaging devices. [Figure 3] FIG. 2 illustrates an example of a hardware configuration. [Figure 4] 1 is a diagram illustrating a functional configuration of an image generating apparatus according to a first embodiment. [Figure 5] 4 is a flowchart of image processing according to the first embodiment. [Figure 6] FIG. 1 is a conceptual diagram of a virtual space according to a first embodiment. [Figure 7]This is a conceptual diagram of the virtual space after one frame. [Figure 8] FIG. 10 is a diagram illustrating a functional configuration of an image generating apparatus according to a second embodiment. [Figure 9] 10 is a flowchart of image processing according to the second embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of priorities of imaging devices. [Figure 11] FIG. 10 is a diagram illustrating a functional configuration of an image generating apparatus according to a third embodiment. [Figure 12] 10 is a flowchart of image processing according to the third embodiment. [Figure 13] FIG. 10 is a conceptual diagram of a virtual space according to a third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. The following embodiments do not limit the present disclosure, and not all of the combinations of features described in the present embodiments are necessarily essential to the solutions of the present disclosure. The configurations of the embodiments may be modified or changed as appropriate depending on the specifications of the device to which the present disclosure is applied and various conditions (such as usage conditions and usage environment). In addition, the embodiments may be configured by appropriately combining parts of the embodiments described below. In the following embodiments, the same configurations and processes are described with the same reference symbols.
[0012] FIG. 1 is a diagram showing a schematic configuration of an image processing system 100 according to this embodiment. The image processing system 100 includes a plurality of imaging devices 110, an image generating device 120, and a terminal device 130. The imaging devices 110 and the image generating devices 120 are connected to each other via a communication cable such as a LAN (Local Area Network) cable. In this embodiment, the communication cable is a LAN cable, but the communication cable is not limited to this example. Furthermore, the connection between the devices is not limited to via a communication cable, and may be a wireless connection.
[0013] The imaging devices 110 are digital cameras that are installed in plurality to surround a specific imaging area at a predetermined imaging location in real space and are capable of capturing, for example, still images and videos. In the following description, unless a distinction between still images and videos is required, the still images and videos captured by the imaging devices 110 will be collectively referred to as captured images. In this embodiment, the imaging device 110 is a digital camera that outputs a video made up of images for each frame that are consecutive on a time axis.
[0014] FIG. 2 is a diagram showing a schematic installation example of multiple imaging devices 110. As shown in FIG. 2, each imaging device 110 is installed to surround a specific imaging area in a predetermined imaging location such as a photography studio, and each captures an image within that imaging area. If subjects such as people are present within the imaging area, the captured image of that imaging area will show those subjects in the foreground, and the part of the photography studio corresponding to the imaging area will be shown in the background. In this embodiment, an example is given in which multiple people are captured as subjects, such as a dance scene in a photography studio. The captured image data acquired by each imaging device 110 is transmitted to the image generation device 120. In the following description, image data handled within the imaging device 110, image generation device 120, terminal device 130, etc. will be simply referred to as "image" unless a separate explanation is required.
[0015] The image generating device 120 is an example application of the image processing device according to this embodiment. The image generating device 120 stores a plurality of captured images transmitted from a plurality of imaging devices 110. Information corresponding to an operation instruction from a user's terminal device 130 is also input to the image generating device 120. In this embodiment, the information corresponding to the operation instruction from the user's terminal device 130 includes at least virtual viewpoint information and playback time information, which will be described later. As will be described in detail later, when the virtual viewpoint information and playback time information are input from the terminal device 130, the image generating device 120 generates a virtual viewpoint image based on the stored captured images and the virtual viewpoint information and playback time information input from the terminal device 130. In this embodiment, the user of the terminal device 130 is assumed to be a video creator who creates content including virtual viewpoint images, a viewer who receives the content, or the like; hereinafter, these will be referred to as a user without distinction.
[0016] Here, virtual viewpoint information is information indicating the direction, etc., represented by the three-dimensional position and angle, of a virtual viewpoint (hereinafter referred to as the virtual viewpoint) in a virtual space constructed from captured images. The virtual viewpoint information has a predetermined position, such as the center of a photography studio, as the origin position, and includes at least the following: a relative position with respect to the origin position, i.e., position information for the front / back, left / right, and up / down relative to the origin position; and direction information, i.e., angle information about the direction from the origin position, i.e., the front / back, left / right, and up / down axes. Because the virtual viewpoint is thus represented by the three-dimensional position and angle, etc., in the following description, the virtual viewpoint including the three-dimensional position and angle, etc., will be referred to as the "virtual viewpoint position." Furthermore, playback time information is time information from the recording start time of the captured image. When a user specifies a playback time via the terminal device 130, the image generating device 120 generates a virtual viewpoint image from the playback time onward.
[0017] The image generating device 120 is, for example, a server device, and has a database function and an image processing function as described below. The database of the image generating device 120 stores captured images sent from the multiple image capturing devices 110, each associated with an identifier that identifies the corresponding image capturing device 110. In this embodiment, the database stores images of the interior of a photography studio captured by the multiple image capturing devices 110. The database stores, as background images, images captured by each image capturing device 110 of the photography studio when no subject, such as a person performing a dance, is present. The database also stores, as foreground images, object images of specific subjects separated by image processing from images captured by each image capturing device 110 of the photography studio when a subject, such as a person, is present. The subject to be separated as an object image from a captured image may not only be a person, but also an object with a predetermined image pattern, such as a prop.
[0018] In this embodiment, the virtual viewpoint image generated by the image generating device 120 in association with the virtual viewpoint information and playback time information is generated based on a background image and an object image of the subject managed in a database. For example, model-based rendering (MBR) is used as a method for generating the virtual viewpoint image. Note that MBR is a method for generating a virtual viewpoint image based on a three-dimensional model generated based on multiple captured images of a subject captured from multiple directions. Specifically, MBR is a technology for generating an image of the scene as seen from a virtual viewpoint using a three-dimensional model (three-dimensional shape) of a target scene obtained by a three-dimensional shape reconstruction method such as volume intersection or multi-view stereo (MVS). Rendering methods other than MBR may also be used to generate the virtual viewpoint image. The virtual viewpoint image generated by the image generating device 120 is transmitted to the terminal device 130 via a LAN cable or the like.
[0019] The terminal device 130 is, for example, a PC (Personal Computer) or a tablet terminal. In this embodiment, a controller 131 is connected to the terminal device 130. The controller 131 includes at least one of a mouse, a keyboard, a six-axis controller, a touch panel, and the like, and is operated by a user. The terminal device 130 displays a virtual viewpoint image received from the image generation device 120 on a display unit 132. The terminal device 130 converts a user operation input from the controller 131 into playback time information and information on an instruction to move the virtual viewpoint position (an instruction regarding the amount of movement and the direction of movement), and transmits the information to the image generation device 120. The instruction to move the playback time and the virtual viewpoint position is not limited to continuous movement of the playback time and the virtual viewpoint position. For example, the virtual viewpoint position can be moved to a predetermined virtual viewpoint position, such as a position in front of, behind, or looking down on the subject in the virtual space. It is also possible to set the playback time and virtual viewpoint position in advance, in which case it becomes possible to instantly move to the preset playback time or virtual viewpoint position in response to a command from the user.
[0020] FIG. 3 is a diagram showing an example of the hardware configuration of the image generating device 120. As shown in FIG. As shown in FIG. 3, the image generating device 120 includes a CPU 301, a ROM 302, a RAM 303, a HDD 304, a display unit 305, an input unit 306, and a communication unit 307. The CPU 301 reads control programs stored in the ROM 302 and executes various control processes. The RAM 303 is used as a temporary storage area, such as the CPU 301's main memory and work area. The HDD 304 stores various programs, including the image processing program according to this embodiment, and various data, including image data. Note that the image processing program according to this embodiment may be stored in the ROM 302. The display unit 305 displays captured images, generated virtual viewpoint images, and various other information. The input unit 306 includes a keyboard and a mouse and accepts various user instructions. The communication unit 307 performs communication with external devices, such as the image capturing device 110, via a network. Note that an example of the network is Ethernet (registered trademark). As another example, the communication unit 307 may communicate with external devices wirelessly. In this embodiment, the functions and processes of the image generating device 120, which will be described later, are realized by the CPU 301 reading and executing an image processing program stored in the HDD 304 or the ROM 302. The hardware configuration of the terminal device 130 is similar to the hardware configuration shown in Fig. 3, and therefore illustration and description thereof will be omitted.
[0021] First Embodiment FIG. 4 is a functional block diagram showing the functional configuration of the image generating device 120 according to the first embodiment. The image input unit 401 converts transmission signals input from each image capture device 110 via a LAN cable into captured image data and outputs the captured image data to the separation unit 402 . If the captured image input from the image input unit 401 is an image of a scene in which a subject does not exist, that is, an image captured before the start of a dance performance, the separation unit 402 outputs the captured image as a background image to the data storage unit 403. If the captured image input from the image input unit 401 is an image of a scene in which a subject exists, that is, an image captured of a scene in which a dance performance or the like is being performed, the separation unit 402 extracts the object of the subject from the captured image. Then, the separation unit 402 outputs the object image of the subject extracted from the captured image to the data storage unit 403 as a foreground image.
[0022] The data storage unit 403 is a database, and stores the background image and foreground image input from the separation unit 402. Then, the data storage unit 403 outputs the foreground image to a three-dimensional shape generation unit 405 (hereinafter referred to as 3D shape generation unit 405). The data storage unit 403 also outputs the foreground image and background image to a virtual viewpoint image generation unit 411. As will be described in detail later, the virtual viewpoint image generation unit 411 uses the foreground image and background image when generating a virtual viewpoint image.
[0023] The parameter storage unit 404 stores in advance the shooting parameters of each of the image capture devices 110 installed to surround a specific shooting area in the photography studio shown in FIG. 2. The shooting parameters are parameter information including the installation position and shooting direction of each of the image capture devices 110, and shooting setting information of each of the image capture devices 110, such as the focal length and exposure time. Each of the image capture devices 110 is installed at a predetermined position. In this embodiment, the shooting parameters of each of the image capture devices 110 will be referred to as "camera parameters" hereinafter. The parameter storage unit 404 outputs the camera parameters of each of the image capture devices 110 to the 3D shape generation unit 405, the selection unit 410, and the virtual viewpoint image generation unit 411.
[0024] The 3D shape generation unit 405 is a three-dimensional model generation unit that generates a three-dimensional model of a subject based on multiple captured images captured by multiple image capture devices arranged at different positions in real space and the camera parameters of each of the multiple image capture devices. In this embodiment, the 3D shape generation unit 405 estimates a three-dimensional model of the subject based on a foreground image read from the data storage unit 403 and camera parameters input from the parameter storage unit 404. The three-dimensional model of the subject is a three-dimensional shape, hereinafter referred to as a 3D shape. The 3D shape generation unit 405 generates 3D shape information of the subject using a three-dimensional shape reconstruction method such as volume intersection. The 3D shape generation unit 405 then outputs the 3D shape information to the 3D position prediction unit 406 and the virtual viewpoint image generation unit 411.
[0025] The 3D position prediction unit 406 is a model prediction unit that predicts the position of a 3D model in a second frame, which is a frame that follows the first frame on the time axis, based on a 3D model generated in a frame preceding the first frame among consecutive frames on the time axis. Here, for example, the first frame is the current frame, and the second frame is the frame following the current frame. In this embodiment, the 3D position prediction unit 406 predicts the 3D shape and its position in the next frame based on 3D shape information for multiple frames preceding the current frame over a predetermined period input from the 3D shape generation unit 405. In other words, it generates a predicted position of the subject in the next frame. More specifically, the 3D position prediction unit 406 calculates the amount of movement change of the 3D shape between two frames preceding the current frame, and further calculates the movement speed from the amount of movement change of the 3D shape. The 3D position prediction unit 406 then estimates the 3D shape and its predicted position in the next frame based on the movement speed of the 3D shape. Hereinafter, the estimated 3D shape and predicted position will be referred to as the predicted 3D shape position. The 3D position prediction unit 406 outputs the information on the predicted 3D shape position to the selection unit 410.
[0026] The user input unit 407 converts the transmission signal transmitted from the terminal device 130 via the LAN cable into user input data. If the user input data is playback time information and virtual viewpoint information, the user input unit 407 outputs the playback time information and virtual viewpoint information to the information setting unit 408.
[0027] The information setting unit 408 is an information acquisition unit that acquires virtual viewpoint information that indicates the position and direction of the virtual viewpoint. In this embodiment, the information setting unit 408 updates the current position and direction of the virtual viewpoint in the virtual space, and the playback time, based on the playback time information and virtual viewpoint information received from the user input unit 407. Thereafter, the information setting unit 408 outputs the playback time information and virtual viewpoint information to the viewpoint position prediction unit 409 and the virtual viewpoint image generation unit 411. It is assumed that the origin of the virtual space is set in advance, for example, to the center of the photography studio.
[0028] The viewpoint position prediction unit 409 predicts the position and direction of the virtual viewpoint in a second frame, which is subsequent to the first frame on the time axis, based on the position and direction of the virtual viewpoint in a frame preceding the first frame among consecutive frames on the time axis. That is, assuming that the first frame is the current frame and the second frame is the next frame, the viewpoint position prediction unit 409 predicts the position and direction of the virtual viewpoint in the next frame based on virtual viewpoint information for multiple frames over a predetermined period preceding the current frame, acquired from the information setting unit 408. Hereinafter, the position and direction of the virtual viewpoint predicted by the viewpoint position prediction unit 409 are collectively referred to as the predicted virtual viewpoint position. In this embodiment, the viewpoint position prediction unit 409 calculates the amount of movement change of a specific virtual viewpoint between two frames preceding the current frame, and further calculates the movement speed of the specific virtual viewpoint from the amount of movement change. The viewpoint position prediction unit 409 then estimates a predicted virtual viewpoint position, which represents the position and direction of the virtual viewpoint in the next frame, based on the movement speed of the virtual viewpoint. The viewpoint position prediction unit 409 outputs information about the predicted virtual viewpoint position to the selection unit 410.
[0029] The selection unit 410 determines an imaging device that will acquire a captured image to be used when generating a virtual viewpoint image for the second frame, based on the predicted virtual viewpoint position by the viewpoint position prediction unit 409, the predicted 3D shape position predicted by the 3D position prediction unit 406, and the camera parameters. That is, when the first frame is the current frame and the second frame is the next frame, the selection unit 410 selects an imaging device that captured an image required to render the subject at the next frame time, based on the predicted 3D shape position, the predicted virtual viewpoint position, and the camera parameters. The selection unit 410 then outputs the identifier of the determined imaging device, etc., to the virtual viewpoint image generation unit 411 as imaging device selection information.
[0030] In this embodiment, the selection unit 410 determines the visibility of a 3D shape when the 3D shape prediction position is captured from the predicted virtual viewpoint position, and selects an imaging device that is close to the predicted virtual viewpoint position from among the imaging devices determined to be visible. That is, the selection unit 410 selects an imaging device that is close to the predicted virtual viewpoint position from among the imaging devices from which the predicted 3D shape position is visible when viewed from the predicted virtual viewpoint position. The selection unit 410 then determines the identifier of the selected imaging device. As a result, the virtual viewpoint image generation unit 411 acquires captured images captured by the imaging device identified by the identifier. Note that when selecting an imaging device that is close to the predicted virtual viewpoint position, at least one of the multiple imaging devices used for capturing images is selected. For example, a predetermined number of two or more imaging devices may be selected as imaging devices that are close to the predicted virtual viewpoint position. In this case, the virtual viewpoint image generation unit 411 acquires an image by combining the pixels of images captured by the predetermined number of imaging devices.
[0031] The virtual viewpoint image generation unit 411 generates a virtual viewpoint image of the second frame based on the captured image and camera parameters of the imaging device determined by the selection unit 410, the 3D model generated by the 3D shape generation unit 405, and the virtual viewpoint information from the information setting unit 408. That is, the virtual viewpoint image generation unit 411 generates a virtual viewpoint image by performing rendering processing based on the virtual viewpoint information, imaging device selection information, the captured image read from the data storage unit 403 according to the imaging device selection information, and the 3D shape information. For example, the virtual viewpoint image generation unit 411 renders (colors) the 3D shape of the subject as seen from the virtual viewpoint position using color information of the image captured by the imaging device at the time corresponding to the playback time. Furthermore, when a subject based on the 3D shape is visible from the virtual viewpoint and the installation position of the imaging device is within a range in which the 3D shape is visible from the virtual viewpoint position, the color of the foreground image extracted from the captured image of the imaging device is used as the color of the 3D shape. The virtual viewpoint image generating unit 411 then synthesizes the image of the subject based on the virtual viewpoint position with the background image to generate a virtual viewpoint image. The virtual viewpoint image generated by the rendering process in the virtual viewpoint image generating unit 411 is sent to the image output unit 412.
[0032] The image output unit 412 converts the virtual viewpoint image input from the virtual viewpoint image generation unit 411 into a transmission signal that can be transmitted to the terminal device 130, and outputs the signal to the terminal device 130.
[0033] Next, the operation of the image generating device 120 will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the flow of image processing in the image generating device 120 according to the first embodiment.
[0034] In step S501, the image input unit 401 determines whether or not imaging has started in the multiple imaging devices and whether or not a captured image has been input from each of the imaging devices. If a captured image has not yet been input from any of the imaging devices, the image input unit 401 waits for input, whereas if a captured image has been input from each of the imaging devices, the image input unit 401 outputs each captured image to the separation unit 402. Then, the processing of the image generation device 120 proceeds to step S502.
[0035] In step S502, if the captured image is an image of a scene in which a subject does not exist, the separation unit 402 outputs the captured image as a background image to the data storage unit 403. If the captured image is an image of a scene in which a subject exists, the separation unit 402 extracts the object of the subject from the captured image and outputs the object image to the data storage unit 403 as a foreground image. As a result, in the next step S 503 , data storage unit 403 stores the foreground image and background image sent from separation unit 402 .
[0036] Next, in step S504, the 3D shape generation unit 405 generates 3D shape information of the subject based on the camera parameters received from the parameter holding unit 404 and the foreground image read from the data storage unit 403. The 3D shape generation unit 405 generates the 3D shape information of the subject using a three-dimensional shape reconstruction method such as the volume intersection method, as described above. Here, the 3D shape information of the subject is made up of a group of multiple points, each of which includes position information.
[0037] Next, in step S505, the information setting unit 408 determines whether a virtual camera path including playback time information and virtual viewpoint information has been input via the user input unit 407. The virtual camera path is virtual viewpoint information that represents the position and direction (posture) of each frame at the virtual viewpoint position, and is a set (sequence) of virtual camera parameters (called virtual camera parameters) at the virtual viewpoint position for each frame. For example, when a frame rate is set to 60 frames per second, one second's worth of information is a sequence of virtual camera parameters at the positions and directions of 60 virtual viewpoints. If a virtual camera path has not been input, the information setting unit 408 waits for input. On the other hand, if a virtual camera path has been input, the information setting unit 408 outputs the virtual camera path to the viewpoint position prediction unit 409.
[0038] Next, in step S506, the viewpoint position prediction unit 409 predicts the virtual viewpoint position of the next frame. For example, when the frame at the current playback time is set as the current frame, the viewpoint position prediction unit 409 calculates the movement speed of the virtual viewpoint based on the amount of movement change of the virtual viewpoint between the two frames before the current frame. Furthermore, the viewpoint position prediction unit 409 determines the predicted virtual viewpoint position of the next frame based on the movement speed. Note that the viewpoint position prediction unit 409 may calculate the movement speed of the virtual viewpoint, and further calculate acceleration from the movement speed of the virtual viewpoint, and calculate the predicted virtual viewpoint position using information about the acceleration.
[0039] Next, in step S507, the 3D position prediction unit 406 predicts the 3D shape position of the next frame based on the 3D shape information for a predetermined period input from the 3D shape generation unit 405, i.e., generates a predicted position of the object in the next frame. For example, if the frame at the current playback time is the current frame, the 3D position prediction unit 406 calculates the amount of movement change of the 3D shape between the two frames before the current frame, and further calculates the movement speed of the 3D shape information from the amount of movement change. Then, the 3D position prediction unit 406 determines the predicted position of the 3D shape in the next frame based on the movement speed. Note that the 3D position prediction unit 406 may calculate the movement speed of the 3D shape, further calculate the acceleration based on the movement speed, and calculate the position of the 3D shape using the acceleration information.
[0040] Next, in step S508, the selection unit 410 determines the imaging device that captured the image required for rendering the subject in the next frame time based on the 3D shape predicted position, the virtual viewpoint predicted position, and the camera parameters. Then, the selection unit 410 outputs imaging device selection information such as the identifier of the selected imaging device to the virtual viewpoint image generation unit 411.
[0041] Next, in step S509 , the virtual viewpoint image generating unit 411 starts receiving the captured image in the next frame based on the image capturing device selection information input from the selecting unit 410 .
[0042] Next, in step S510, the virtual viewpoint image generation unit 411 determines whether or not virtual viewpoint information for the next frame has been input, that is, whether or not a virtual camera path for the next frame has been input, from the information setting unit 408. If virtual viewpoint information for the next frame has not been input, the virtual viewpoint image generation unit 411 enters a waiting state, and if virtual viewpoint information for the next frame has been input, the process proceeds to step S511.
[0043] In step S511, the virtual viewpoint image generation unit 411 generates a virtual viewpoint image, which is an image of a viewpoint seen from the virtual viewpoint position of the next frame. That is, based on the imaging device selection information obtained in step S508, the virtual viewpoint image generation unit 411 performs rendering processing based on the captured image of the next frame read out from the data storage unit 403 in step S509 and the 3D shape information from the 3D shape generation unit 405. Then, the virtual viewpoint image generation unit 411 outputs the virtual viewpoint image generated by the rendering processing to the image output unit 412.
[0044] 6(a) and 6(b) are conceptual diagrams showing the positional relationship between the subject shape predicted in the virtual space and the predicted virtual viewpoint position. Note that in the examples of Fig. 6(a) and Fig. 6(b), for the sake of simplicity of illustration and explanation, only six of the eight image capture devices 110 shown in Fig. 2, namely, image capture devices 601 to 606, are depicted.
[0045] FIG. 6(a) is a diagram showing an example in which image processing according to this embodiment is not performed. FIG. 6(a) shows image capture devices 601 to 606 actually arranged in a corresponding manner in a virtual space, subjects 1411 and 1412 corresponding to the virtual space, a virtual viewpoint position 1421, and a predicted virtual viewpoint position 1422. For example, if it is assumed that the subject 1411 is captured at the predicted virtual viewpoint position 1422, images captured by the image capture devices 601 and 602 will be used based on the predicted virtual viewpoint position. However, if image processing according to this embodiment is not performed, the subject 1411 will overlap with the subject 1412 as seen from the image capture device 601, making it impossible to color the subject 1411 using the image captured by the image capture device 601.
[0046] For this reason, the image generating device 120 according to this embodiment determines the position of the imaging device that captures the image required to generate the virtual viewpoint image based on the virtual viewpoint predicted position, the 3D shape predicted position, which is the predicted subject position, and the camera parameters. As a result, the image generating device 120 according to this embodiment can generate a colored virtual viewpoint image even when multiple subjects are present. Furthermore, the image generating device 120 according to this embodiment can reduce the time required to generate the virtual viewpoint image by predicting the virtual viewpoint and 3D shape.
[0047] FIG. 6(b) is a diagram showing an example of image processing according to this embodiment performed in the image generating device 120. In FIG. 6(b), image capturing devices 601 to 606 are image capturing devices arranged in a virtual space, similar to the example in FIG. 6(a). A virtual viewpoint position 622 indicates the position and direction of the virtual viewpoint according to the playback time information and virtual viewpoint information input from the user input unit 407. Meanwhile, a virtual viewpoint position 621 indicates the position and direction of the virtual viewpoint in the previous frame, and a predicted virtual viewpoint position 623 indicates the predicted virtual viewpoint position indicating the position and direction of the virtual viewpoint predicted in the next frame. Furthermore, in FIG. 6(b), a predicted 3D shape position 612 indicates the predicted 3D shape position predicted in the next frame for a 3D shape 611 of an object associated in the virtual space. A predicted 3D shape position 614 indicates the predicted 3D shape position predicted in the next frame for a 3D shape 613 of the object. 6(b), for example, when capturing images of a 3D shape predicted position 612 and a 3D shape predicted position 614 from a predicted virtual viewpoint position 623, rendering at the 3D shape predicted position 614 is possible by using images captured by the image capturing devices 601 and 602. However, the 3D shape predicted position 612 overlaps with the 3D shape predicted position 614 as seen from the image capturing device 601, reducing visibility. For this reason, the image captured by the image capturing device 601 is not used for the 3D shape predicted position 612, and rendering at the 3D shape predicted position 612 is performed using images captured by the image capturing devices 602 and 606. That is, the image generating device 120 will use images captured by the image capturing devices 601, 602, and 606 for rendering after the next frame time.
[0048] FIG. 7 is a conceptual diagram of the positional relationship between the shape of a subject that has actually moved after the time of the next frame in virtual space and the virtual viewpoint position. In FIG. 7, image capture devices 601 to 606 are image capture devices arranged in a corresponding manner in virtual space, similar to the example in FIG. 6. FIG. 7 also shows 3D shapes 701 and 702 of the subject that have been associated with the virtual space, and a virtual viewpoint position 711 based on virtual viewpoint information input by a user. The virtual viewpoint position 711 does not necessarily coincide with the predicted virtual viewpoint position 623 based on the prediction described in FIG. 6(b). In contrast, the image capture device selected based on the virtual viewpoint prediction and the subject prediction coincides with the image capture device determined based on the virtual viewpoint position based on the user input at the time of the next frame and the actual subject position. Therefore, even if the subject or the like actually moves, rendering can be performed for the subject shape.
[0049] According to the first embodiment, when there are multiple subjects, an image capture device to be used for rendering in the next frame is selected based on the 3D shape predicted position, the virtual viewpoint predicted position, and the camera parameters, as described above. As a result, according to the first embodiment, a virtual viewpoint image can be generated even when there are multiple subjects, and the delay time from user input to display of the virtual viewpoint image can be shortened, enabling real-time display. Furthermore, according to this embodiment, since images captured by the selected image capture device are used, the amount of image data used is reduced, enabling a reduction in communication bandwidth and processing load.
[0050] <Second embodiment> Below, as a second embodiment, we will explain an example in which a priority is set for the imaging device that acquires the captured image based on the 3D shape predicted position, the virtual viewpoint predicted position, and the camera parameters, and the captured image is acquired from the imaging device based on that priority.
[0051] 8 is a diagram showing the functional configuration of an image generating device 800 according to the second embodiment. The image generating device 800 according to the second embodiment has a priority determination unit 801 instead of the selection unit 410 of the image generating device 120 according to the first embodiment shown in FIG. 4. The priority determination unit 801 receives input of information on the predicted 3D shape position from the 3D position prediction unit 406, information on the predicted virtual viewpoint position from the viewpoint position prediction unit 409, and camera parameters from the parameter storage unit 404. Note that, since the other functional units other than the priority determination unit 801 are generally similar to the corresponding functional units in the first embodiment described above, a description of those units will be omitted, and only the parts that are different from the first embodiment will be described below.
[0052] Based on the 3D shape predicted position, the virtual viewpoint predicted position, and the camera parameters, the priority determination unit 801 increases the priority (priority order) of the imaging device that captured the image required to render the subject for the next frame time, and decreases the priority of the other imaging devices. For example, the priority determination unit 801 determines the visibility of the 3D shape when the 3D shape predicted position is captured from the virtual viewpoint predicted position. Then, for each imaging device for which the 3D shape predicted position is determined to be visible, the priority determination unit 801 increases the priority of the imaging device the closer it is to the virtual viewpoint predicted position, and decreases the priority of the imaging device the farther it is from the virtual viewpoint predicted position. Note that, taking into consideration that the virtual viewpoint position may be moved to a predetermined virtual viewpoint position, the priority of the imaging device closer to the predetermined virtual viewpoint position may be increased. Then, the priority determination unit 801 outputs priority information, which associates the priority determined for each imaging device with the identifier of each imaging device, to the virtual viewpoint image generation unit 411. As a result, the virtual viewpoint image generation unit 411 acquires the captured images of each imaging device based on the priority.
[0053] Fig. 9 is a flowchart of image processing in an image generating device 800 according to the second embodiment. Note that steps S501 to S507 and steps S510 to S511 are the same as the corresponding steps in the flowchart shown in Fig. 5, and therefore their description will be omitted. In the flowchart of Fig. 9, after processing in step S507, processing proceeds to step S901, and further after processing in step S902, processing proceeds to step S510.
[0054] In step S901, the priority determination unit 801 sets priorities for the image capture devices based on the predicted 3D shape position, the predicted virtual viewpoint position, and the camera parameters. That is, the priority determination unit 801 increases the priority of the image capture device that captured an image necessary for rendering the subject in the next frame time, decreases the priority of the other image capture devices, and outputs priority information associated with the identifiers of the image capture devices to the virtual viewpoint image generation unit 411.
[0055] Next, in step S902, the virtual viewpoint image generation unit 411 starts receiving captured images of the next frame in descending order of priority from the image capture devices assigned with higher priorities based on the priority information input from the priority determination unit 801. Note that priorities may be assigned to all image capture devices, and a priority range that is actually desired to be acquired may be specified within the priority order. In this case, the virtual viewpoint image generation unit 411 may acquire captured images in descending order of priority from the image capture devices assigned with priorities within the priority range. Also, for example, priorities do not necessarily need to be assigned to all image capture devices. In this case, the virtual viewpoint image generation unit 411 may acquire captured images in descending order of priority from the image capture devices assigned with priorities.
[0056] FIG. 10 is a diagram showing an example in which priorities are assigned to image capture devices based on the 3D shape predicted position and the virtual viewpoint predicted position, and the priorities are assigned corresponding to the identifiers of the image capture devices and then sorted in order of priority. FIG. 10 shows an example of priorities assigned to image capture devices 601 to 606, using the positional relationship between the subject and the virtual viewpoint shown in FIG. 6(b) as an example. Note that in the example of FIG. 10, the reference numerals (601 to 606) assigned to each image capture device are used as the identifiers of the image capture devices. In the case of the 3D shape predicted position and the virtual viewpoint predicted position shown in FIG. 6(b), image capture device 602 is determined to have a priority of "1" because it is close to the virtual viewpoint predicted position 623, and image capture device 601 is determined to have a priority of "2" because it is the next closest after image capture device 602. Furthermore, image capture device 606 is determined to be in a position necessary for rendering 3D shape predicted position 612, and therefore is determined to have a priority of "3." Also, taking into consideration that the virtual viewpoint position may be moved to a predetermined virtual viewpoint position set in advance, if the priority of the image capture device 604 in the vicinity of the predetermined virtual viewpoint position is also set high, the priority is determined to be "4." On the other hand, the image capture device 603, which is in a position that is unlikely to move in the next frame time, is set to a priority of "5," and the image capture device 605, which is in an even less likely position, is set to a priority of "6."
[0057] As described above, according to the second embodiment, by acquiring captured images from the imaging device in order of priority, the captured image to be used for the virtual viewpoint position of the next frame time becomes available first, thereby shortening the delay time until the virtual viewpoint image is generated. Furthermore, in the second embodiment, for example, if there is spare capacity in the transmission bandwidth of the captured images, captured images with lower priorities may also be acquired sequentially. In this case, it is possible not only to accommodate movement to a predetermined virtual viewpoint position, but also to accommodate the unlikely event that the predicted virtual viewpoint position and the actual virtual viewpoint position differ.
[0058] <Third embodiment> Hereinafter, as a third embodiment, an example will be described in which the number of imaging devices that acquire captured images is changed based on the predicted 3D shape position and the moving speed of the virtual viewpoint when generating the predicted virtual viewpoint position. Fig. 11 is a diagram showing an example of the functional configuration of an image processing apparatus according to the third embodiment. An image generating apparatus 1100 according to the third embodiment has a number determination unit 1101 instead of the selection unit 410 of the image generating apparatus 120 according to the first embodiment shown in Fig. 4. The number determination unit 1101 receives input of information on the predicted 3D shape position from the 3D position prediction unit 406, information on the predicted virtual viewpoint position and the moving speed of the virtual viewpoint from the viewpoint position prediction unit 409, and camera parameters from the parameter storage unit 404. Note that, since the other functional units other than the number determination unit 1101 are the same as the corresponding functional units in the first embodiment described above, their description will be omitted, and only the parts different from the first embodiment will be described below.
[0059] The number determination unit 1101 determines the number of imaging devices that have captured images required for rendering the subject in the next frame time based on the 3D shape predicted position, the virtual viewpoint predicted position, the moving speed of the virtual viewpoint, and the camera parameters. Then, the number determination unit 1101 outputs the identifiers of each of the determined number of imaging devices to the virtual viewpoint image generation unit 411.
[0060] In the third embodiment, the viewpoint position prediction unit 409 also calculates the moving speed of the virtual viewpoint based on the virtual viewpoint positions in the two frames preceding the current frame, as described above, when calculating the predicted position of the virtual viewpoint. Here, if the moving speed of the virtual viewpoint is higher than the assumed speed (e.g., 3 m / s) of a person, who is the main subject, running, the predicted virtual viewpoint position may be located further away than the virtual viewpoint position that should be correctly predicted. Conversely, it is possible that the predicted virtual viewpoint position may stop at a position closer to the virtual viewpoint position that should be correctly predicted. In other words, a discrepancy may occur between the predicted virtual viewpoint position acquired by the viewpoint position prediction unit 409 and the virtual viewpoint position that should be correctly predicted. If the difference between the predicted virtual viewpoint position and the virtual viewpoint position that should be correctly predicted becomes large, when an imaging device is selected based on a prediction such as that in the first embodiment, the selected imaging device may differ from the imaging device that should be used at the virtual viewpoint position.
[0061] Therefore, in the third embodiment, the number determination unit 1101 determines the number of imaging devices to acquire captured images based on the predicted 3D shape position, camera parameters, predicted virtual viewpoint position, and moving speed of the virtual viewpoint. That is, the faster the moving speed of the virtual viewpoint is above a predetermined set speed, the more imaging devices the number determination unit 1101 will acquire captured images. As an example of the predetermined set speed, a speed (e.g., 3 m / s) that simulates the movement of a person, the main subject, when he or she starts running can be cited. Also, for example, if the number of imaging devices to acquire captured images used to generate a virtual viewpoint image is set to a predetermined number (e.g., 3), the number determination unit 1101 increases the number of imaging devices to a number greater than the predetermined number when the moving speed of the virtual viewpoint is faster than the set speed.
[0062] Furthermore, for example, even when the moving speed of the virtual viewpoint is equal to or less than the set speed, the predicted virtual viewpoint position may be farther away than the virtual viewpoint position that should be correctly predicted, or may stop at a position closer to the virtual viewpoint position that should be correctly predicted. However, the slower the moving speed of the virtual viewpoint, the smaller the difference between the predicted virtual viewpoint position and the virtual viewpoint position that should be correctly predicted is expected to be. In other words, the slower the moving speed of the virtual viewpoint, the smaller the difference between the number of imaging devices to be used at the virtual viewpoint position and the predetermined number is expected to be. Therefore, when the moving speed of the virtual viewpoint is equal to or less than the predetermined set speed, the number determination unit 1101 sets the number of imaging devices that will acquire captured images to the predetermined number. Note that in this embodiment, an example has been given in which the number of imaging devices that will acquire captured images is set to a predetermined number. However, the number determination unit 1101 may change the number of imaging devices that will acquire captured images to a smaller number as the moving speed of the virtual viewpoint becomes slower.
[0063] As described above, in the third embodiment, by changing the number of image capturing devices that acquire captured images according to the moving speed of the virtual viewpoint, it is possible to deal with fluctuations in the difference between the predicted virtual viewpoint position and the virtual viewpoint position that should be correctly predicted. Note that, although the example in this embodiment uses the moving speed of the virtual viewpoint, if the predicted virtual viewpoint position is calculated by calculating acceleration based on the moving speed of the virtual viewpoint, the number of image capturing devices may be determined based on the moving acceleration of the virtual viewpoint.
[0064] In the third embodiment, the number determination unit 1101 determines the visibility of the 3D shape when the 3D shape predicted position is captured by each of the image capturing devices of the number determined as described above. Furthermore, the number determination unit 1101 selects an image capturing device that is close to the virtual viewpoint predicted position from among the image capturing devices for which the 3D shape predicted position is determined to be visible, and determines the identifier of the selected image capturing device. As a result, the virtual viewpoint image generation unit 411 acquires an image captured by the image capturing device identified by the identifier and generates a virtual viewpoint image.
[0065] Fig. 12 is a flowchart of image processing in the image generating device 1100 according to the third embodiment. Note that steps S501 to S507 and steps S510 to S511 are the same as the corresponding steps in the flowchart shown in Fig. 5, and therefore their description will be omitted. In the flowchart of Fig. 12, after processing step S507, the process proceeds to processing step S1201, and after processing step S1202, the process proceeds to processing step S510.
[0066] In step S1201, the number determination unit 1101 determines the number of imaging devices required for rendering the subject in the next frame time, based on the 3D shape predicted position, the virtual viewpoint predicted position, the moving speed of the virtual viewpoint, and the camera parameters. Furthermore, the number determination unit 1101 determines the visibility of the 3D shape when each imaging device captures the 3D shape predicted position, and outputs the identifier of the imaging device selected based on the determination result to the virtual viewpoint image generation unit 411.
[0067] Next, in step S1202, the virtual viewpoint image generation unit 411 starts receiving, as the captured image of the next frame, the captured images of the imaging devices corresponding to the identifiers input from the number determination unit 1101. As a result, the virtual viewpoint image generation unit 411 generates a virtual viewpoint image based on the captured images of those imaging devices.
[0068] 13(a) and 13(b) are conceptual diagrams showing the 3D shape of a subject predicted in a virtual space and the predicted positional relationship with the virtual viewpoint. 13(a) and 13(b), similar to the example of FIG. 6(b), image capture devices 601 to 606 are image capture devices arranged in a corresponding manner in a virtual space. Furthermore, virtual viewpoint position 622 indicates a virtual viewpoint position corresponding to playback time information and virtual viewpoint information input by a user, and virtual viewpoint position 621 indicates a virtual viewpoint position in the previous frame. Furthermore, 3D shape predicted position 612 indicates a 3D shape predicted position in the next frame for 3D shape 611 of an object corresponding to the virtual space, and 3D shape predicted position 614 indicates a 3D shape predicted position in the next frame for 3D shape 613 of an object.
[0069] Here, Fig. 13(a) shows an example in which the moving speed of the virtual viewpoint is slower than a predetermined set speed, and predicted virtual viewpoint position 1301 indicates the predicted virtual viewpoint position in the next frame. On the other hand, Fig. 13(b) shows an example in which the moving speed of the virtual viewpoint is faster than a predetermined set speed, and predicted virtual viewpoint position 1302 indicates the predicted virtual viewpoint position in the next frame. In other words, the predicted virtual viewpoint position 1301 in Fig. 13(a) where the moving speed of the virtual viewpoint is slow is significantly different from the predicted virtual viewpoint position 1302 in Fig. 13(b) where the moving speed is fast.
[0070] In the example of Figure 13(a), since the moving speed of the virtual viewpoint is equal to or less than a predetermined set speed, the number of imaging devices determined corresponding to the predicted virtual viewpoint position is set to a predetermined number (e.g., three). On the other hand, in the example of FIG. 13(b), because the moving speed of the virtual viewpoint is faster than the set speed, the moving range of the predicted virtual viewpoint position expands as shown by the trajectory indicated by the arrow from the virtual viewpoint position 622 to the predicted virtual viewpoint position 1302. In this case, the number of imaging devices required to obtain captured images that may be necessary for generating a virtual viewpoint image increases. To accommodate this, the number of imaging devices is determined to be, for example, a number greater than the predetermined number (e.g., four). Note that in this embodiment, the number of imaging devices that acquire captured images is determined based on the moving speed of the virtual viewpoint. However, the number of imaging devices to acquire captured images may be increased or decreased based on not only the moving speed but also the installation locations and the number of imaging devices. For example, the number determination unit 1101 may determine a larger number of imaging devices that acquire captured images as the number of imaging devices that capture the same capturing range increases. Conversely, the number determination unit 1101 may determine a smaller number of imaging devices that acquire captured images as the number of imaging devices that capture the same capturing range decreases.
[0071] By determining the number of imaging devices according to the moving speed of the virtual viewpoint as described above, for example, when capturing images at the 3D shape predicted position 612 and the 3D shape predicted position 614 at the virtual viewpoint predicted position 1301 in FIG. 13( a), the number of imaging devices is three. That is, in this example, the three imaging devices 601, 602, and 606 are determined as the imaging devices for acquiring captured images to be used in the next frame. The number determination unit 1101 also determines the visibility of the 3D shape when the 3D shape predicted position is captured by the three imaging devices 601, 602, and 606. In the example of FIG. 13( a), the imaging devices 601 and 602 closest to the virtual viewpoint predicted position 1301 are selected, and the captured images acquired by them are used to render the virtual viewpoint image corresponding to the 3D shape predicted position 614. Similarly, the imaging devices 602 and 606 closest to the virtual viewpoint predicted position 1301 are selected, and the captured images acquired by them are used to render the virtual viewpoint image corresponding to the 3D shape predicted position 612.
[0072] For example, at the predicted virtual viewpoint position 1302 in FIG. 13B, when capturing images at the predicted 3D shape position 612 and the predicted 3D shape position 614, the number of imaging devices is four, as described above. That is, in this example, the four imaging devices 601, 602, 603, and 604 are determined as the imaging devices for acquiring captured images to be used in the next frame. The number determination unit 1101 also determines the visibility of the 3D shape when the predicted 3D shape position is captured by the four imaging devices 601, 602, 603, and 604. In the example of FIG. 13B, the imaging devices 602 and 603 closest to the predicted virtual viewpoint position 1302 are selected, and the captured images acquired by these devices are used to render a virtual viewpoint image corresponding to the predicted 3D shape position 614. Similarly, the imaging devices 602 and 603 closest to the predicted virtual viewpoint position 1302 are selected, and the captured images acquired by these devices are used to render a virtual viewpoint image corresponding to the predicted 3D shape position 612.
[0073] As described above, in the third embodiment, the imaging devices that will capture images to be used for rendering in the next frame and the number of imaging devices are determined based on the predicted 3D shape position, camera parameters, and the predicted virtual viewpoint position and moving speed of the virtual viewpoint. As a result, according to the third embodiment, it is possible to cover a predicted position range that may change depending on the moving speed of the virtual viewpoint.
[0074] In the first to third embodiments described above, the virtual viewpoint position is specified by a user operation, but this is not limited to being specified by a user operation, and a virtual viewpoint image may be generated using a virtual viewpoint position prepared in advance.
[0075] The present disclosure can also be realized by supplying a program that realizes one or more functions of each of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. The above-described embodiments are merely examples of specific implementations of the present disclosure, and should not be construed as limiting the technical scope of the present disclosure. In other words, the present disclosure can be implemented in various forms without departing from its technical idea or main features.
[0076] The disclosure of the present embodiment includes the following configurations, methods, programs, and systems. (Configuration 1) information acquisition means for acquiring virtual viewpoint information indicating the position and direction of the virtual viewpoint; a model acquisition means for acquiring a three-dimensional model of a subject generated based on images captured by a plurality of imaging devices; a viewpoint prediction means for predicting a virtual viewpoint in a second frame subsequent to a first frame based on a virtual viewpoint in a frame preceding a first frame in a virtual viewpoint image; a model prediction means for predicting a position of a three-dimensional model of the subject in the second frame based on a position of a three-dimensional model of the subject corresponding to a frame preceding the first frame; a determination means for determining, from among the plurality of image capture devices, an image capture device that will acquire a captured image used in generating the second frame, based on the predicted virtual viewpoint, the predicted position of the three-dimensional model, and imaging parameters of the plurality of image capture devices; an image generation means for generating a virtual viewpoint image based on a captured image corresponding to the second frame acquired by the determined imaging device, a three-dimensional model corresponding to the second frame acquired by the model acquisition means, and virtual viewpoint information corresponding to the second frame acquired by the information acquisition means; 1. An image processing device comprising: (Configuration 2) the first frame is a frame corresponding to a playback time designated by a user, 2. The image processing device according to configuration 1, wherein the viewpoint prediction means predicts a virtual viewpoint in the second frame based on virtual viewpoints in at least two frames before the first frame. (Configuration 3) the first frame is a frame corresponding to a playback time designated by a user, 3. The image processing device according to claim 1, wherein the model prediction means predicts a position of the three-dimensional model of the subject in the second frame based on the three-dimensional model of the subject in the at least two frames prior to the first frame. (Configuration 4) When the information acquisition means acquires virtual viewpoint information indicating the position and direction of the virtual viewpoint specified by the user, the determining means determines, based on virtual viewpoint information designated by the user, an imaging device that will acquire an image to be used when generating the virtual viewpoint image, from among the plurality of imaging devices; The image processing device according to any one of configurations 1 to 3, wherein the image generation means generates the virtual viewpoint image based on a captured image corresponding to the second frame captured by the determined imaging device and imaging parameters of the imaging device, virtual viewpoint information specified by the user, and a three-dimensional model corresponding to the second frame captured by the model acquisition means. (Configuration 5) The image processing device according to any one of configurations 1 to 4, wherein the determination means determines whether the position of the predicted three-dimensional model is visible from the predicted virtual viewpoint, and based on the result of the determination, determines an imaging device to acquire an image to be used when generating the virtual viewpoint image of the second frame. (Configuration 6) The image processing device according to configuration 5, wherein the determination means determines at least one image capturing device in the vicinity of the predicted virtual viewpoint, where the position of the predicted three-dimensional model is determined to be visible, as the image capturing device that will acquire the captured image to be used when generating the virtual viewpoint image. (Configuration 7) The image processing device according to configuration 6, characterized in that when a predetermined number of two or more imaging devices are determined as imaging devices near the predicted virtual viewpoint, the image generation means uses an image obtained by synthesizing each of the images captured by the predetermined number of imaging devices when generating the virtual viewpoint image. (Configuration 8) the determining means determines priorities of the image capturing devices for acquiring captured images used in generating the virtual viewpoint image of the second frame based on the predicted virtual viewpoint, the predicted position of the three-dimensional model, and the imaging parameters of the plurality of image capturing devices; 8. The image processing device according to any one of configurations 1 to 7, wherein the image generation means uses the captured images of the imaging device and the shooting parameters in an order according to the priority when generating the virtual viewpoint image. (Configuration 9) The image processing device according to configuration 8, wherein the determination means increases the priority of each of the image capturing devices determined as an image capturing device for acquiring a captured image to be used when generating a virtual viewpoint image of the second frame, and decreases the priority of each of the image capturing devices not determined as an image capturing device. (Configuration 10) The image processing device according to configuration 8 or 9, wherein the determining means determines whether the position of the predicted three-dimensional model is visible from the predicted virtual viewpoint, and increases the priority of the imaging devices that are determined to be visible in order of proximity to the predicted virtual viewpoint. (Configuration 11) 11. The image processing device according to any one of configurations 8 to 10, wherein the determining means assigns a higher priority to an image capturing device that is closer to a predetermined virtual viewpoint that has been set in advance. (Configuration 12) 12. The image processing device according to any one of configurations 8 to 11, wherein the image generation means, when generating the virtual viewpoint image, uses the captured image and the shooting parameters of the image capture device having a priority within a priority range that is specified in advance for the priority. (Configuration 13) the determining means determines the number of image capturing devices that will acquire captured images used when generating the virtual viewpoint image of the second frame, based on the predicted virtual viewpoint, the predicted position of the three-dimensional model, and the imaging parameters of the plurality of image capturing devices; 13. The image processing device according to any one of configurations 1 to 12, wherein the image generation means uses the captured images and the shooting parameters of the determined number of the imaging devices when generating the virtual viewpoint image. (Configuration 14) 14. The image processing device according to claim 13, wherein the determining means obtains the predicted moving speed of the virtual viewpoint, and determines the number of the virtual viewpoints to be larger as the moving speed becomes faster than a predetermined set speed. (Configuration 15) The image processing device according to configuration 13 or 14, wherein the determination means acquires the predicted moving speed of the virtual viewpoint, and if the moving speed is equal to or less than a predetermined set speed, sets the number of devices to a predetermined number. (Configuration 16) The image processing device described in any one of configurations 13 to 15, characterized in that the determination means determines whether the position of the predicted three-dimensional model is visible from the predicted virtual viewpoint, and based on the result of the determination, selects an imaging device from the determined number of imaging devices to obtain an image to be used when generating the virtual viewpoint image of the second frame. (Configuration 17) 17. The image processing device according to any one of configurations 13 to 16, wherein the determining means increases the number of imaging devices to be determined as the number of imaging devices that capture the same shooting range increases. (Method 1) An image processing method executed by an image processing device, an information acquisition step of acquiring virtual viewpoint information indicating the position and direction of the virtual viewpoint; a model acquisition step of acquiring a three-dimensional model of a subject generated based on images captured by a plurality of imaging devices; a viewpoint prediction step of predicting a virtual viewpoint in a second frame after a first frame based on a virtual viewpoint in a frame before the first frame in a virtual viewpoint image; a model prediction step of predicting a position of a three-dimensional model of the subject in the second frame based on a position of a three-dimensional model of the subject corresponding to a frame preceding the first frame; a determination step of determining, from among the plurality of image capture devices, an image capture device that will acquire a captured image to be used when generating the second frame, based on the predicted virtual viewpoint, the predicted position of the three-dimensional model, and imaging parameters of the plurality of image capture devices; an image generation step of generating a virtual viewpoint image based on a captured image corresponding to the second frame captured by the determined imaging device, a three-dimensional model corresponding to the second frame acquired in the model acquisition step, and virtual viewpoint information corresponding to the second frame acquired in the information acquisition step; An image processing method comprising: (Program 1) A program that causes a computer to function as the image processing device according to any one of configurations 1 to 17. (System 1) A plurality of imaging devices arranged in real space; an image processing device according to any one of configurations 1 to 17; An image processing system comprising: [Explanation of symbols]
[0077] 120: Image generating device, 406: 3D position predicting unit, 408: Information setting unit, 410: Viewpoint position predicting unit, 410: Selection unit, 411: Virtual viewpoint image generating unit
Claims
1. information acquisition means for acquiring virtual viewpoint information indicating the position and direction of the virtual viewpoint; a model acquisition means for acquiring a three-dimensional model of a subject generated based on images captured by a plurality of imaging devices; a viewpoint prediction means for predicting a virtual viewpoint in a second frame subsequent to a first frame based on a virtual viewpoint in a frame preceding a first frame in a virtual viewpoint image; a model prediction means for predicting a position of a three-dimensional model of the subject in the second frame based on a position of a three-dimensional model of the subject corresponding to a frame preceding the first frame; a determination means for determining, from among the plurality of image capture devices, an image capture device that will acquire an image used in generating the second frame, based on the predicted virtual viewpoint, the predicted position of the three-dimensional model, and imaging parameters of the plurality of image capture devices; an image generation means for generating a virtual viewpoint image based on a captured image corresponding to the second frame acquired by the determined imaging device, a three-dimensional model corresponding to the second frame acquired by the model acquisition means, and virtual viewpoint information corresponding to the second frame acquired by the information acquisition means; 1. An image processing device comprising:
2. the first frame is a frame corresponding to a playback time designated by a user, 2. The image processing apparatus according to claim 1, wherein said viewpoint prediction means predicts the virtual viewpoint in the second frame based on virtual viewpoints in at least two frames prior to the first frame.
3. the first frame is a frame corresponding to a playback time designated by a user, 2. The image processing device according to claim 1, wherein the model prediction means predicts the position of the three-dimensional model of the subject in the second frame based on the three-dimensional model of the subject in at least two frames before the first frame.
4. When the information acquisition means acquires virtual viewpoint information indicating the position and direction of the virtual viewpoint specified by the user, the determining means determines, based on virtual viewpoint information designated by the user, an imaging device that will acquire an image to be used when generating the virtual viewpoint image, from among the plurality of imaging devices; The image processing device according to claim 1, characterized in that the image generation means generates the virtual viewpoint image based on a captured image corresponding to the second frame acquired by the determined imaging device and shooting parameters of the imaging device, virtual viewpoint information specified by the user, and a three-dimensional model corresponding to the second frame acquired by the model acquisition means.
5. 5. The image processing device according to claim 1, wherein the determination means determines whether the position of the predicted three-dimensional model is visible from the predicted virtual viewpoint, and based on the result of the determination, determines an imaging device to acquire an image to be used when generating the virtual viewpoint image of the second frame.
6. The image processing device according to claim 5, wherein the determination means determines at least one image capturing device in the vicinity of the predicted virtual viewpoint, from which the position of the predicted three-dimensional model is determined to be visible, as the image capturing device that will acquire the captured image to be used when generating the virtual viewpoint image.
7. The image processing device according to claim 6, characterized in that, when a predetermined number of two or more imaging devices are determined as imaging devices near the predicted virtual viewpoint, the image generation means uses an image obtained by synthesizing each of the images captured by the predetermined number of imaging devices when generating the virtual viewpoint image.
8. the determining means determines priorities of the image capturing devices for acquiring captured images used in generating the virtual viewpoint image of the second frame based on the predicted virtual viewpoint, the predicted position of the three-dimensional model, and the imaging parameters of the plurality of image capturing devices; 5. The image processing device according to claim 1, wherein the image generating means uses the captured images and the imaging parameters of the imaging device in the order of priority when generating the virtual viewpoint image.
9. The image processing device according to claim 8, characterized in that the determination means increases the priority of each of the imaging devices determined as the imaging device for acquiring the captured image to be used when generating the virtual viewpoint image of the second frame, and decreases the priority of each of the imaging devices not determined.
10. The image processing device according to claim 8, wherein the determining means determines whether the position of the predicted three-dimensional model is visible from the predicted virtual viewpoint, and increases the priority of the imaging device that is determined to be visible in order of proximity to the predicted virtual viewpoint.
11. 9. The image processing apparatus according to claim 8, wherein the determining means assigns a higher priority to an image capturing apparatus that is closer to a predetermined virtual viewpoint that has been set in advance.
12. 9. The image processing device according to claim 8, wherein the image generating means uses the captured image and the shooting parameters of the image capturing device having a priority within a priority range that is pre-specified for the priority when generating the virtual viewpoint image.
13. the determining means determines the number of image capturing devices that will acquire captured images used when generating the virtual viewpoint image of the second frame, based on the predicted virtual viewpoint, the predicted position of the three-dimensional model, and the imaging parameters of the plurality of image capturing devices; 5. The image processing device according to claim 1, wherein the image generating means uses the captured images of the determined number of the image capturing devices and the imaging parameters when generating the virtual viewpoint image.
14. 14. The image processing apparatus according to claim 13, wherein the determining means acquires the predicted moving speed of the virtual viewpoint, and determines the number of the virtual viewpoints to be larger as the moving speed becomes faster than a predetermined set speed.
15. 14. The image processing device according to claim 13, wherein the determining means acquires the predicted moving speed of the virtual viewpoint, and when the moving speed is equal to or less than a predetermined set speed, sets the number of the virtual viewpoints to a predetermined number.
16. The image processing device according to claim 13, wherein the determination means determines whether the position of the predicted three-dimensional model is visible from the predicted virtual viewpoint, and based on the result of the determination, selects an imaging device from the determined number of imaging devices to obtain an image to be used when generating the virtual viewpoint image of the second frame.
17. 14. The image processing device according to claim 13, wherein the determining unit increases the number of image capturing devices to be determined as the number of image capturing devices that capture the same image capturing range increases.
18. An image processing method executed by an image processing device, an information acquisition step of acquiring virtual viewpoint information indicating the position and direction of the virtual viewpoint; a model acquisition step of acquiring a three-dimensional model of a subject generated based on images captured by a plurality of imaging devices; a viewpoint prediction step of predicting a virtual viewpoint in a second frame subsequent to a first frame in a virtual viewpoint image based on a virtual viewpoint in a frame preceding the first frame; a model prediction step of predicting a position of a three-dimensional model of the subject in the second frame based on a position of a three-dimensional model of the subject corresponding to a frame preceding the first frame; a determination step of determining, from among the plurality of image capture devices, an image capture device that will acquire a captured image used in generating the second frame, based on the predicted virtual viewpoint, the predicted position of the three-dimensional model, and imaging parameters of the plurality of image capture devices; an image generation step of generating a virtual viewpoint image based on a captured image corresponding to the second frame captured by the determined imaging device, a three-dimensional model corresponding to the second frame acquired in the model acquisition step, and virtual viewpoint information corresponding to the second frame acquired in the information acquisition step; An image processing method comprising:
19. Computer, information acquisition means for acquiring virtual viewpoint information indicating the position and direction of the virtual viewpoint; a model acquisition means for acquiring a three-dimensional model of a subject generated based on images captured by a plurality of imaging devices; a viewpoint prediction means for predicting a virtual viewpoint in a second frame subsequent to a first frame based on a virtual viewpoint in a frame preceding a first frame in a virtual viewpoint image; a model prediction means for predicting a position of a three-dimensional model of the subject in the second frame based on a position of a three-dimensional model of the subject corresponding to a frame preceding the first frame; a determination means for determining, from among the plurality of image capture devices, an image capture device that will acquire an image used in generating the second frame, based on the predicted virtual viewpoint, the predicted position of the three-dimensional model, and imaging parameters of the plurality of image capture devices; an image generation means for generating a virtual viewpoint image based on a captured image corresponding to the second frame acquired by the determined imaging device, a three-dimensional model corresponding to the second frame acquired by the model acquisition means, and virtual viewpoint information corresponding to the second frame acquired by the information acquisition means; A program that causes the image processing device to function as an image processing device having the above.
20. A plurality of imaging devices arranged in real space; The image processing device according to claim 1 ; An image processing system comprising:
Citation Information
Patent Citations
Image processor, image processing method, and program
JP2018036955A
Image processing system, control method therefor, and program
JP2019079468A
Method, system and apparatus for capture of image data for free viewpoint video
JP2020095717A
System and method for determining a virtual camera path
JP2022510658A
Information processing device, information processing method, and program
WO2021002116A1