Image processing device, control method thereof, and program

The image processing apparatus addresses abrupt video transitions by generating a virtual viewpoint video during a transition period, ensuring smooth switching between camera views and enhancing viewer experience.

JP7716232B2Active Publication Date: 2025-07-31CANON KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021089463
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-27
Publication Date
2025-07-31
Estimated Expiration
2041-05-27

AI Technical Summary

Technical Problem

Existing video switching technologies cause unnatural changes and discomfort to viewers due to abrupt transitions between multiple camera views, especially in music event photography and live streaming, where videos from different cameras are switched without smooth transitions.

Method used

An image processing apparatus that generates a virtual viewpoint video by acquiring information from multiple cameras, setting a transition period for switching, and creating a virtual viewpoint video during this period based on the viewpoints of both cameras, allowing for a smooth transition between the videos.

Benefits of technology

The solution reduces the sense of discomfort by smoothly transitioning between different camera views, providing a more natural and seamless viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007716232000001
    Figure 0007716232000001
  • Figure 0007716232000002
    Figure 0007716232000002
  • Figure 0007716232000003
    Figure 0007716232000003
Patent Text Reader

Abstract

To reduce an unnatural change in a video when two videos are output by being switched.SOLUTION: An image processing apparatus obtains, as information on a first video and a second video at least one of which is a captured video obtained by an image capturing apparatus, first viewpoint information for obtaining the first video and second viewpoint information for obtaining the second video at a time corresponding to the time of the first video, sets, in switching a video to be output from the first video to the second video, a period from the end of output of the first video to the start of output of the second video, generates information on a virtual viewpoint in the period, based on the first viewpoint information in the set period and the second viewpoint information in the set period, generates a virtual viewpoint video of the period based on the information on the virtual viewpoint in the period, and sequentially outputs the first video, the virtual viewpoint video of the set period, and the second video by switching them.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, a control method thereof, and a program.

Background Art

[0002] Recently, a technique of installing a plurality of cameras at different positions to perform synchronous shooting from multiple viewpoints and generating a virtual viewpoint video using the multi-viewpoint video obtained by the shooting has attracted attention. For example, Patent Document 1 discloses a technique of arranging a plurality of cameras so as to surround a subject and generating an image of an arbitrary viewpoint using the images of the subject taken by these plurality of cameras. According to the technique of generating a virtual viewpoint video from such a multi-viewpoint video, for example, highlight scenes of soccer or basketball can be viewed from various angles, so that a higher sense of presence can be given to viewers as compared with ordinary videos. In addition, in the shooting of music events, live distribution, music videos, etc., videos of an artist taken from various angles can be created.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In music event photography, live streaming, and the filming of music videos, etc., it is common to switch and use multiple videos obtained simultaneously from multiple cameras. For example, the first camera shoots a so-called "pull video" from a long-shot video including the surroundings of the subject to a bust shot of the subject. Also, for example, the second camera shoots a so-called "push video" from a bust shot video of the subject to a close-up shot. And by switching and using the videos shot by these first and second cameras, videos corresponding to various subject sizes can be generated. At this time, for example, it is conceivable to use the first camera as a virtual viewpoint (referred to as a virtual camera in this specification) for generating the above-described virtual viewpoint video, and the second camera as an actual camera (referred to as an actual camera in this specification) that shoots an image not used in the virtual viewpoint video.

[0005] Generally, in a video switching device that switches two videos and outputs one video, since the video instantaneously switches to another video, the video changes greatly at the time of switching. For this reason, viewers may feel a sense of discomfort. As a method for reducing the viewers' sense of discomfort at the time of video switching, it is known to add video effects such as fade-in and fade-out during video switching. However, at the time of switching, the video by the first camera and the video by the second camera are still used, and it is impossible to avoid the occurrence of unnatural video changes caused by video switching.

[0006] According to one aspect of the present invention, a technique for reducing unnatural changes in a video when switching and outputting two videos is provided.

Means for Solving the Problem

[0007] An image processing apparatus according to one aspect of the present invention has the following configuration. That is, An acquisition means for acquiring information related to a first video and a second video, at least one of which is a captured video obtained by an imaging device, the acquisition means acquiring information on a first viewpoint for obtaining the first video and information on a second viewpoint for obtaining the second video at a time corresponding to the time of the first video. Setting means for setting a period from the end of the output of the first video to the start of the output of the second video when switching the output video from the first video to the second video. First generation means for generating information on a virtual viewpoint during the period based on the information on the first viewpoint during the period and the information on the second viewpoint during the period. Second generation means for generating a virtual viewpoint video for the period based on the information on the virtual viewpoint during the period. Output means for switching and outputting in the order of the first video, the virtual viewpoint video for the period, and the second video. having and The first generation means generates a virtual viewpoint for the period based on the information of the first viewpoint, the information of the second viewpoint, and the ratio of the elapsed time from the start of the period to the total time of the period. An image processing apparatus according to another aspect of the present invention has the following configuration. That is, an acquisition means for acquiring information related to a first video and a second video, at least one of which is an imaging video obtained by an imaging device, the acquisition means acquiring information of a first viewpoint for obtaining the first video and information of a second viewpoint for obtaining the second video at a time corresponding to the time of the first video; a setting means for setting a period from the end of the output of the first video to the start of the output of the second video when switching the output video from the first video to the second video; a first generation means for generating information of a virtual viewpoint in the period based on the information of the first viewpoint in the period and the information of the second viewpoint in the period; a second generation means for generating a virtual viewpoint video for the period based on the information of the virtual viewpoint in the period; an output means for switching and outputting the first video, the virtual viewpoint video for the period, and the second video in this order; a setting means for setting a ratio according to a user operation received during the period; and has the first generation means generates a virtual viewpoint for the period based on the information of the first viewpoint, the information of the second viewpoint, and the ratio set by the setting means. An image processing apparatus according to still another aspect of the present invention has the following configuration. That is, an acquisition means for acquiring information related to a first video and a second video, at least one of which is an imaging video obtained by an imaging device, the acquisition means acquiring information of a first viewpoint for obtaining the first video and information of a second viewpoint for obtaining the second video at a time corresponding to the time of the first video; a setting means for setting a period from the end of the output of the first video to the start of the output of the second video when switching the output video from the first video to the second video; a first generation means for generating information of a virtual viewpoint in the period based on the information of the first viewpoint in the period and the information of the second viewpoint in the period; a second generation means for generating a virtual viewpoint video for the period based on the information of the virtual viewpoint in the period; an output means for switching and outputting the first video, the virtual viewpoint video for the period, and the second video in this order; having The first generation means generates a virtual viewpoint at each time in the period based on the information of the first viewpoint and the information of the second viewpoint at each time.

Advantages of the Invention

[0008] According to the present invention, unnatural changes in video when switching between two videos are reduced.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 10

Modes for Carrying Out the Invention

[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.

[0011] <First Embodiment> Hereinafter, an image processing apparatus that switches the output video from the video of the first viewpoint to the video of the second viewpoint will be described. In the first embodiment, the first viewpoint is set as the viewpoint of a virtual imaging device for generating a virtual viewpoint video from a plurality of images captured by a plurality of imaging devices, and the second viewpoint is set as the viewpoint of a physical imaging device that captures the video. That is, the video of the first viewpoint is a virtual viewpoint video, and the video of the second viewpoint is a video captured by a real camera (hereinafter, real camera video). Hereinafter, in an image processing system that generates a virtual viewpoint video, an example of generating a new virtual viewpoint video that smoothly connects these two videos when switching from the virtual viewpoint video to the real camera video will be described.

[0012] FIG. 1 is a block diagram showing a configuration example of an image processing system that generates a virtual viewpoint video according to the first embodiment. The camera group 101 is composed of a plurality of imaging devices (hereinafter referred to as cameras) that acquire multi-viewpoint images of the shooting range in order to generate a virtual viewpoint video. Each of the plurality of cameras includes an imaging element inside and a lens in front of it. The plurality of cameras are installed and fixed around the shooting range facing the shooting range. The camera control unit 102 controls each camera of the camera group 101. The camera control unit 102 is provided for each camera of the camera group 101 and is connected to each camera of the camera group 101 by a camera control cable and a camera image output cable. Also, between the plurality of camera control units 102, they are connected, for example, in a daisy chain via a local network cable or the like, and the images of the camera group 101 are transmitted to the image processing device 103 connected to the subsequent stage. Note that the network configuration for connecting the plurality of camera control units 102 is not limited to a daisy chain, and a star-type network configuration in which each camera control is connected to the image processing device may be used.

[0013] The image processing device 103 has a function of generating and outputting a virtual viewpoint video, which is a video from a virtual viewpoint, based on the images (multi-viewpoint images) acquired by the camera group 101. Hereinafter, the functional configuration of the image processing device 103 will be described.

[0014] The image acquisition unit 104 acquires the captured images (multi-viewpoint images) acquired by the camera group 101 from the camera control unit 102. The image acquisition unit 104 acquires in advance, as a background image, a captured image obtained by the camera group 101 capturing a shooting area that does not include the shooting target (foreground), and stores it in the background image storage unit 105. The separation unit 106 separates the shooting target (foreground) included in the captured image from the captured image of the shooting area. The separation unit 106 performs separation by, for example, background difference. More specifically, the separation unit 106 compares the background image acquired in advance and stored in the background image storage unit 105 with the captured image, and separates the foreground and the background by identifying the difference as the foreground that is the shooting target. The separation unit 106 stores the image including the separated foreground (hereinafter referred to as the foreground image) in the foreground image storage unit 107. Note that the method for separating the foreground and the background used by the separation unit 106 is not limited to the above-described separation method using background difference, and well-known separation methods such as a separation method using a distance image can be used.

[0015] The foreground image storage unit 107 stores a plurality of foreground images (a plurality of foreground images acquired by a plurality of cameras (that is, a plurality of viewpoints)) separated by the separation unit 106 from the captured images of the camera group 101 installed around the shooting area. The 3D model generation unit 108 acquires a foreground image from the foreground image storage unit 107 and generates a 3D model of the foreground. The 3D model generation unit 108 generates a 3D model of the foreground using, for example, the volume intersection method from the foreground images acquired from a plurality of viewpoints. The generated 3D model of the foreground and its position information are stored in the 3D model storage unit 109.

[0016] The virtual camera information generation unit 110 generates virtual camera information according to user operations received from a user interface such as a joystick and various input units, which indicate the position of a virtual viewpoint, the direction of a line of sight, and the like. The virtual camera information includes information on the position, orientation (line-of-sight direction), field of view angle (focal length), and time information of the virtual viewpoint of the virtual viewpoint video (hereinafter also referred to as a virtual camera). That is, the function of the virtual camera information generation unit 110 is to generate information for each time of the virtual viewpoint necessary for generating a virtual viewpoint video according to the operation of the virtual camera by an operator using an input unit such as a joystick.

[0017] The virtual viewpoint video generation unit 111 generates a virtual viewpoint video based on the time, the position, the orientation, and the field of view angle of the virtual camera represented by the virtual camera information generated by the virtual camera information generation unit 110 or the virtual camera information automatic generation unit 117 described later. For example, in order to generate a virtual viewpoint video, the virtual viewpoint video generation unit 111 acquires a foreground image at the time from the foreground image storage unit 107 and a foreground 3D model at the time from the 3D model storage unit 109, and generates a foreground image corresponding to the position, the orientation, and the field of view angle of the virtual camera. In addition, the virtual viewpoint video generation unit 111 acquires the background image stored in the background image storage unit 105, acquires a prepared background 3D model, and generates a background image corresponding to the position, the orientation, and the field of view angle of the virtual camera. The virtual viewpoint video generation unit 111 synthesizes the generated foreground image and background image and outputs them as a virtual viewpoint video. The virtual viewpoint video is provided to the video switching unit 115 and becomes one of the video candidates output as the final video.

[0018] The real camera 112 is a camera that can photograph the shooting range of the virtual camera independently of the camera group 101. The real camera 112 is not used to acquire the images necessary for the virtual viewpoint video, but is used to photograph the subject in close-up. In the present embodiment, in order to distinguish the camera group 101 that acquires the images necessary for the virtual viewpoint video and the virtual camera that is virtually arranged at a position as if it is acquiring the virtual viewpoint video although it does not actually exist, the name of "real camera" is used. The imaging video obtained by the real camera 112 is provided to the video switching unit 115 described later and becomes one of the video candidates to be output as the final video.

[0019] The real camera information acquisition unit 113 acquires information including the position, posture (line-of-sight direction), and angle of view (focal length) of the real camera 112. The real camera information acquisition unit 113 estimates the position and posture of the real camera 112, for example, from the position where the marker arranged in the range where the real camera 112 moves is reflected in the image photographed by the real camera 112. However, it is not limited to this. For example, an image of the marker may be obtained by connecting a camera that photographs the marker for position estimation separately from the real camera to the real camera 112. Further, without arranging the marker, a characteristic location whose position is known may be specified from the image photographed by the real camera 112, and the position and posture of the real camera 112 may be estimated.

[0020] The video determination unit 114 selects and determines an output video from among a plurality of output video candidates. The video determination unit 114 includes input units such as a switch for selecting video output and a fader for adjusting volume and the like. Also, various video effects (transitions) when switching videos can be added and switched. For example, it can be determined to output a virtual viewpoint video, switch from a virtual viewpoint video to a real camera video, or determine to add video effects such as fade-in and fade-out when switching. The video determination unit 114 transmits channel information specifying the selected video and information indicating the video effects to be executed when switching to the video switching unit 115. The video switching unit 115 selects a video from among the video candidates based on the information from the video determination unit 114 and outputs it to the video output unit 116. The video output unit 116 outputs the video supplied from the video switching unit 115 to the outside.

[0021] When the virtual camera information automatic generation unit 117 switches the output video from the video of the virtual camera to the video of the real camera, it automatically generates virtual camera information for obtaining a virtual viewpoint video that connects the videos before and after the switch. The virtual camera information generated by the virtual camera information automatic generation unit 117 is one of the video effects when switching videos. When the positions, postures (directions of the line of sight), and field angles (focal lengths (zoom values)) of the virtual camera and the real camera are different, new virtual camera information is automatically generated from the virtual camera information and the real camera information to smooth the change in the video when switching videos.

[0022] Next, the hardware configuration of the image processing apparatus 103 that realizes the functional configuration as described above will be described with reference to FIG. 10. The image processing apparatus 103 includes a CPU (Central Processing Unit) 1001, a ROM (Read Only Memory) 1002, a RAM (Random Access Memory) 1003, an auxiliary storage device 1004, a display unit 1005, an operation unit 1006, a communication I / F 1007, and a bus 1018.

[0023] The CPU 1001 controls the entire image processing apparatus 103 by using computer programs and data stored in the ROM 1002 and the RAM 1003, thereby realizing each function of the image processing apparatus 103 shown in FIG. 1. Note that the image processing apparatus 103 may have one or more dedicated hardware different from the CPU 1001, and at least a part of the processing by the CPU 1001 may be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a DSP (Digital Signal Processor). The ROM 1002 stores programs and the like that do not require modification. The RAM 1003 temporarily stores programs and data supplied from the auxiliary storage device 1004, and data supplied from the outside via the communication I / F 1007. The auxiliary storage device 1004 is composed of, for example, a hard disk drive or the like, and stores various data such as image data and audio data.

[0024] The display unit 1005 is composed of, for example, a liquid crystal display, an LED, or the like, and displays a GUI (Graphical User Interface) or the like for a user to operate the image processing apparatus 103. The operation unit 1006 is composed of, for example, a keyboard, a mouse, a joystick, a touch panel, or the like, and receives an operation by the user and inputs various instructions to the CPU 1001. The communication I / F 1007 is used for communication with a device outside the image processing apparatus 103. For example, when the image processing apparatus 103 is connected to an external device by wire, a communication cable is connected to the communication I / F 1007. When the image processing apparatus 103 has a function of wireless communication with an external device, the communication I / F 1007 includes an antenna. The bus 1018 connects each part of the image processing apparatus 103 and transmits information.

[0025] In this embodiment, it is assumed that the display unit 1005 and the operation unit 1006 are inside the image processing apparatus 103. However, at least one of the display unit 1005 and the operation unit 1006 may exist outside the image processing apparatus 103 as another device. In this case, the CPU 1001 may operate as a display control unit that controls the display unit 1005 and an operation control unit that controls the operation unit 1006.

[0026] Next, with reference to FIG. 2, the processing for switching the video of the virtual camera and the real camera by the image processing apparatus 103 having the above configuration will be described. FIG. 2 is a flowchart showing the output video determination processing by the image processing apparatus of the first embodiment. Note that in FIG. 2, the processing of storing the image acquired by the camera group 101 in the foreground image storage unit 107 and the processing of storing the foreground image separated by the separation unit 106 in the 3D model storage unit 109 are omitted.

[0027] In step S201, the virtual viewpoint video generation unit 111 acquires the virtual camera information generated by the virtual camera information generation unit 110. In step S202, the virtual viewpoint video generation unit 111 generates a virtual viewpoint video based on the acquired virtual camera information. In step S203, the video switching unit 115 acquires the switching information of the output video from the video determination unit 114. The switching information indicates, for example, the channel of the output video after switching determined by the video determination unit 114, the switching time, and the like. In step S204, the video switching unit 115 determines whether to stop the output video based on the switching information acquired in step S203. If it is determined to stop the output video (YES in step S204), in step S205, the video switching unit 115 stops the output of the video. If it is determined not to stop the output video (NO in S204), the process proceeds to step S206.

[0028] In step S206, the video switching unit 115 determines whether to switch the output video based on the switching information acquired in step S203. If it is determined not to switch the output video (NO in step S206), in step S207, the video switching unit 115 continues to output the video without switching the output video. Then, the process returns to step S201. On the other hand, if it is determined to switch the output video (YES in step S206), the process proceeds to step S208.

[0029] In step S208, the video switching unit 115 determines whether to automatically generate virtual camera information when switching the output video. If it is determined not to automatically generate virtual camera information (NO in step S208), in step S209, the video switching unit 115 immediately switches the video output to the video switching unit 116 based on the switching information to the switched video indicated by the switching information. For example, the virtual viewpoint video generated by the virtual viewpoint video generation unit 111 from the virtual viewpoint generated by the virtual camera information generation unit 110 is switched to the real camera video captured by the real camera 112. Then, the process returns to step S201. On the other hand, if it is determined to automatically generate virtual camera information (YES in step S208), the process proceeds to step S210.

[0030] The switching information from the video determination unit 114 is also provided to the virtual camera information automatic generation unit 117. In step S210, the virtual camera information automatic generation unit 117 acquires switching conditions from the switching information received from the video determination unit 114. The switching conditions include, for example, information on a transition period indicating a period (start time and end time) for automatically generating virtual camera information. The virtual camera information automatic generation unit 117 acquires virtual camera information necessary for generating a virtual viewpoint from the virtual camera information generation unit 110 and real camera information from the real camera information acquisition unit 113. In step S211, the virtual camera information automatic generation unit 117 generates information on a new virtual viewpoint (virtual camera information) when switching the video based on the virtual camera information, the real camera information, and the switching conditions. In step S212, the virtual viewpoint video generation unit 111 generates a virtual viewpoint video based on the new virtual viewpoint newly generated by the virtual camera information automatic generation unit 117. After outputting the virtual viewpoint video obtained from this new virtual viewpoint, the video switching unit 115 starts outputting the selected video (in this example, the real camera video). Then, the process returns to step S201.

[0031] The relationship among the virtual viewpoint video, the real camera video, and the output video at each time point when switching the output video from the virtual camera to the real camera will be described below with reference to FIG. 3. FIG. 3 is a diagram showing the time line of the video switching process in the first embodiment. In FIG. 3, the first virtual viewpoint video 301 is a virtual viewpoint video generated by the virtual viewpoint video generation unit 111 based on the virtual camera information (also referred to as the first virtual camera information) generated by the virtual camera information generation unit 110. The real camera video 302 is a video captured and output by the real camera 112. The second virtual viewpoint video 303 is a virtual viewpoint video generated by the virtual viewpoint video generation unit 111 based on the virtual camera information (also referred to as the second virtual camera information) generated by the virtual camera information automatic generation unit 117. The output video 304 is a video selected and output by the video switching unit 115 from among the first virtual viewpoint video 301, the real camera video 302, and the second virtual viewpoint video 303, which are candidate videos. Note that the horizontal axis represents time.

[0032] The virtual viewpoint video generation unit 111 generates and outputs a first virtual viewpoint video 301 according to the virtual camera information generated by the virtual camera information generation unit 110 in response to the virtual camera operation by the operator. The real camera 112 also outputs a real camera video 302 that it has captured. Note that the position, orientation, zoom, etc. of the real camera 112 are being operated by the cameraman. At time t0, the video determination unit 114 outputs switching information 310 to the video switching unit 115, indicating that the video is to be switched from the first virtual viewpoint video 301 to the real camera video 302 over a period of t7 - t2 seconds using the second virtual viewpoint video 303 after t2 - t0 seconds. In the example of FIG. 3, the period from the time t2 when the output of the first virtual viewpoint video ends to the time t7 when the output of the real camera video 302 starts is set as the transition period.

[0033] The switching information 310 received by the video switching unit 115 instructs to switch the output video from the first virtual viewpoint video 301 to the real camera video 302 and to use the second virtual viewpoint video 303 as the switching condition. Note that the second virtual viewpoint video 303 is a virtual viewpoint image generated by the virtual viewpoint video generation unit 111 based on the virtual camera information generated by the virtual camera information automatic generation unit 117. Also, the switching condition is set such that the time period from t2 to t7 is the transition period (the period during which the second virtual viewpoint video is output) for video switching.

[0034] When the switching information 310 including the above switching conditions is output from the video determination unit 114, it is determined as YES in steps S206 and S208 of FIG. 2. When the virtual camera information automatic generation unit 117 receives this switching condition, it generates a new virtual viewpoint (also referred to as the second virtual viewpoint) for creating the second virtual viewpoint video 303 for switching from the first virtual viewpoint video 301 to the real camera video 302 from time t2 to time t7. More specifically, first, the virtual camera information automatic generation unit 117 obtains virtual camera information from the virtual camera information generation unit 110 and real camera information from the real camera information acquisition unit 113 in order to create virtual viewpoint information from time t2 to time t7. The virtual camera information includes information on the position, line-of-sight direction, and field angle of the virtual viewpoint used by the virtual viewpoint video generation unit 111 to generate the first virtual viewpoint video 301. The real camera information includes information on the position, posture, and field angle of the real camera 112 that is capturing the real camera video 302. The video switching unit 115 selects the first virtual viewpoint video 301 and outputs it to the video output unit 116 until time t2. At time t2, the video switching unit 115 switches the video output to the video output unit 116 from the first virtual viewpoint video 301 to the second virtual viewpoint video 303. Further, at time t7, the video switching unit 115 switches the video output to the video output unit 116 from the second virtual viewpoint video 303 to the real camera video 302. The video output unit 116 outputs the video sent from the video switching unit 115.

[0035] An example of the automatic generation process of virtual camera information by the virtual camera information automatic generation unit 117 will be described in detail with reference to FIGS. 4A to 4C. FIGS. 4A to 4C are examples of the automatic generation process of virtual camera information in the first embodiment. FIG. 4A shows the positions and postures of the virtual camera for generating the first virtual viewpoint video 301, the virtual camera for generating the second virtual viewpoint video 303, and the real camera 112 that captures the real camera video 302 at each time from time t0 to t10. In the following, the positions of the virtual camera and the real camera will be described, but other camera information (posture, zoom state, etc.) can be calculated in the same way. Note that the time line from t0 to t10 corresponds to the time line shown in FIG. 3.

[0036] In FIGS. 4A to 4C, the first virtual camera information 401 indicates the position indicated by the position information of the first virtual camera generated by the virtual camera information generation unit 110 with a black dashed arrow. The first virtual camera moves moment by moment along the black dashed arrow in the direction of the arrow between t0 and t10. The real camera information 403 indicates the position indicated by the position information of the real camera 112 acquired by the real camera information acquisition unit 113 with a white dashed arrow. The real camera 112 moves moment by moment in the direction of the arrow along the white dashed arrow between t0 and t10. The virtual camera information automatic generation unit 117 generates the second virtual camera information 402 so as to approach the real camera information at each time, starting from the virtual camera information at time t2. The movement of the second virtual camera according to the second virtual camera information 402 is indicated by a black solid arrow in FIGS. 4A to 4C.

[0037] Hereinafter, with reference to FIGS. 4B and 4C, a method for the virtual camera information automatic generation unit 117 to generate the position of the second virtual camera 2 from the positions of the first virtual camera that moves moment by moment and the real camera 112 will be described. Hereinafter, an example of generating the information of the second virtual viewpoint based on the information of the first virtual camera and the real camera, the elapsed time from the start of the transition period, and the ratio of the total time of the transition period will be described.

[0038] 4a in FIG. 4B shows the positions of the first virtual camera, the real camera 112, and the second virtual camera at time t2. At time t2, the position of the second virtual camera is the same as the position of the first virtual camera. 4b shows the positions of the first virtual camera, the real camera 112, and the second virtual camera at time t3. The position of the second virtual camera at time t3 is determined based on the ratio of the elapsed time (t3 - t2) to the total time of the transition period (t7 - t2). More specifically, on the line segment connecting the position of the first virtual camera at time t2 and the position of the real camera 112 at time t3, the position advanced from the first virtual camera towards the real camera 112 by the ratio of (t3 - t2) / (t7 - t2) is the position of the second virtual camera at time t3. In other words, the position of the second virtual camera during the transition period is generated by weighted-averaging the position of the first virtual viewpoint and the position of the real camera 112 based on the ratio. 4c shows the positions of the first virtual camera, the real camera 112, and the second virtual camera at time t4. The position of the second virtual camera at time t4 is generated in the same way as at time t3. That is, on the line segment connecting the position of the first virtual camera at time t2 and the position of the real camera 112 at time t4, the position advanced from the first virtual camera towards the real camera 112 by the ratio of (t4 - t2) / (t7 - t2) is the position of the second virtual camera at time t4.

[0039] 4d in FIG. 4C shows the positions of the first virtual camera, the actual camera 112, and the second virtual camera at time t5. The position of the second virtual camera at time t5 is also generated in the same way as at time t3. That is, on the line segment connecting the position of the first virtual camera at time t2 and the position of the actual camera 112 at time t5, the position advanced from the first virtual camera in the direction of the actual camera 112 by the ratio of (t5 - t2) / (t7 - t2) becomes the position of the second virtual camera at time t5. 4e in FIG. C shows the positions of the first virtual camera, the actual camera 112, and the second virtual camera at time t6. The position of the second virtual camera at time t6 is also generated in the same manner as above. That is, it is the position advanced from the first virtual camera in the direction of the actual camera 112 by the ratio of (t6 - t2) / (t7 - t2) on the line segment connecting the position of the first virtual camera at time t2 and the position of the actual camera 112 at time t6. 4f in FIG. 4C shows the positions of the first virtual camera, the actual camera 112, and the second virtual camera at time t7. The position of the second virtual camera at time t7 is the position advanced from the first virtual camera in the direction of the actual camera 112 by the ratio of (t7 - t2) / (t7 - t2) on the line segment connecting the position of the first virtual camera at time t2 and the position of the actual camera 112 at time t7. That is, at time t7, which is the end time of the transition period, the position of the second virtual camera and the position of the actual camera 112 become the same.

[0040] As described above, according to the first embodiment, when switching from the virtual viewpoint video by the first virtual camera to the real camera video by the real camera 112, a transition period from time t2 to time t7 is provided. And in this transition period, the information of the second virtual camera that moves from the position of the first virtual camera to the position of the real camera 112 is generated based on the information of the first virtual camera and the information of the real camera in the transition period. Therefore, when switching from the video of the first virtual camera to the video of the real camera 112, even if the positions of the first virtual camera and the real camera are far apart, it is possible to automatically generate the information of the virtual camera that interpolates between them during the transition period. As a result, it is possible to provide a video without a sense of incongruity when switching from the video of the virtual camera to the video of the real camera. Although the process of switching from the virtual camera video to the real camera video has been described, the same process as above can be applied when switching from the real camera video to the virtual camera video. In that case, the position of the second virtual camera at the first time of the transition period is the same as the position of the real camera 112, and the position of the second virtual camera is gradually brought closer to the position of the first virtual camera.

[0041] Note that in FIGS. 4A to 4C, the position of the second virtual camera during the transition period gradually approaches the position of the real camera without depending on the position of the first virtual camera except at the start of the transition period, but it is not limited to this. For example, the second virtual camera information 402 may be automatically generated using the method shown in FIG. 5.

[0042] FIG. 5 shows another example of the virtual camera path generation method for the virtual viewpoint video in the first embodiment. Similar to FIGS. 4A to 4C, FIG. 5 shows the positions of the first virtual camera, the second virtual camera, and the real camera 112 at each time from time t0 to t10. In this example, a method of generating the information of the second virtual camera using the information of the first virtual camera and the real camera 112 at the same time will be described to generate the second virtual camera information 402. Similar to the method described in FIGS. 4A to 4C, at time t2, the position of the first virtual camera and the position of the second virtual camera are the same.

[0043] The position of the second virtual camera at time t3 is a position advanced from the first virtual camera in the direction of the real camera 112 by a ratio of (t3 - t2) / (t7 - t2) on the line segment connecting the positions of the first virtual camera and the real camera 112 at time t3. Similarly, the position of the second virtual camera at time t4 is a position advanced from the first virtual camera in the direction of the real camera 112 by a ratio of (t4 - t2) / (t7 - t2) on the line segment connecting the positions of the first virtual camera and the real camera 112 at time t4. Similarly, the position of the second virtual camera at time t5 is a position advanced from the first virtual camera in the direction of the real camera 112 by a ratio of (t5 - t2) / (t7 - t2) on the line segment connecting the positions of the first virtual camera and the real camera 112 at time t5. Similarly, the position of the first virtual camera at time t6 is a position advanced from the first virtual camera in the direction of the real camera 112 by a ratio of (t6 - t2) / (t7 - t2) on the line segment connecting the positions of the first virtual camera and the real camera 112 at time t6. Similarly, the position of the second virtual camera at time t7 is a position advanced from the first virtual camera in the direction of the real camera 112 by a ratio of (t7 - t2) / (t7 - t2) on the line segment connecting the positions of the first virtual camera and the real camera 112 at time t7. As described with reference to FIG. 4C (4f), at time t7 which is the end time of the transition period, the position of the second virtual camera and the position of the real camera 112 become the same.

[0044] As described above, in the method shown in FIG. 5, the virtual camera position when switching from the virtual viewpoint video by the first virtual camera to the real camera video by the real camera 112 is calculated based on the positions of the first virtual camera and the real camera 112 at the same time. According to this method, when switching from the virtual camera video to the real camera video or from the real camera video to the virtual camera video, the position of the second virtual camera is always calculated from the positions of the first virtual camera and the real camera 112 at the same time. Therefore, even if the second virtual camera changes its direction to move from the position of the first virtual camera to the position of the real camera 112 and then towards the virtual camera position 1 from the real camera position during the movement, there is no sense of discomfort and the switching can be performed without discomfort.

[0045] In addition, in the method of automatically generating the above two pieces of virtual camera information, the start time and end time for switching the video are specified, but it is not limited to this. Instead, the start time for switching and the time required for switching (the length of the transition period) may be specified. This makes it easy to specify in advance the time required for switching or to unify the switching times when generating the same video.

[0046] In addition, in the method of automatically generating the above two pieces of virtual camera information, the movement of the second virtual camera when switching the video is determined based on the ratio of the elapsed time to the movement period, but it is not limited to this. For example, instead of the ratio of the elapsed time to the movement period described above, a ratio specified by a user operation (hereinafter referred to as a transition ratio) may be used at each time in the movement period. For example, an input unit having a fader that can specify the video before switching and the video after switching to the video determination unit 114 and can specify the transition ratio may be provided, and the position of the second virtual viewpoint may be generated according to a user operation on the input unit.

[0047] Fig. 6 shows an example of the input unit 600 that can specify the transition ratio. The user operation by the input unit 600 is output to the video determination unit 114. The input unit 600 has a pre-switch button switch 601 and a post-switch button switch 602, and each is provided with button switches from channel 1 to channel 4. A fader 603 is provided so as to straddle between the pre-switch button switch 601 and the post-switch button switch 602. The fader 603 moves according to a user operation and indicates the transition ratio when switching the video according to its position. In the present embodiment, the virtual viewpoint video by the first virtual camera is assigned to channel 1, and the real camera video by the real camera 112 is assigned to channel 2.

[0048] In Fig. 6(a), the fader 603 is at the uppermost position. In this case, the video of the channel specified by the pre-switch button switch 601 is output. The pre-switch button switch 601 of channel 1 is lit, indicating that the video of channel 1 (the first virtual viewpoint video 301) is selected as the video output from the video switching unit 115. On the other hand, channel 2 is selected on the post-switch button switch 602, and channel 2 is lit. This indicates that channel 2 (the actual camera video 302) is selected as the video to be output after switching. When the fader 603 is moved from the uppermost position downward, the output video switches from the virtual viewpoint video by the first virtual camera to the virtual viewpoint video 2 by the second virtual camera. And the position of the second virtual camera is generated by the method described above with reference to Figs. 4A to 4C or Fig. 5 based on the transition ratio corresponding to the position of the fader 603. Note that the transition ratio can be set based on, for example, the distance from the uppermost position to the lowermost position of the fader 603 and the distance from the uppermost position to the current position of the fader 603.

[0049] In the example of Fig. 6(b), the fader 603 is at the 2 / 5 position between the uppermost and lowermost positions. In this case, on the line segment connecting the position of the first virtual camera and the position of the actual camera 112 at that time, the position advanced 2 / 5 of the line segment from the first virtual camera toward the actual camera 112 is the position of the second virtual camera (similar to 4c in Fig. 4B). Note that the time when the movement of the fader 603 starts from the state of Fig. 6(a) is the start time of the above-described transition period, and the time when the fader 603 reaches the lowermost position as shown in Fig. 6(c) is the end time of the transition period. That is, when the fader 603 reaches the lowermost position, the video switches from the video of the second virtual camera to the video of the actual camera 112, and the video switching is completed.

[0050] As described above, by operating the fader 603, it becomes possible to specify the transition ratio used by the virtual camera information automatic generation unit 117 to generate virtual camera information when switching videos. Therefore, it is possible to easily operate the switching time and the speed at which the virtual camera approaches the state of the actual camera.

[0051] As described above, the switching from the virtual viewpoint video to the actual camera video has been described, but it is not limited to this, and the above processing can also be applied to the switching from the actual camera video to the virtual viewpoint video. That is, one of the first viewpoint for obtaining the video before switching and the second viewpoint for obtaining the video after switching is the viewpoint of a virtual imaging device for generating the virtual viewpoint video, and the other may be the viewpoint of a physical imaging device for shooting the video. In that case, the video switches from the actual camera video to the virtual viewpoint video by the second virtual camera, and further switches to the virtual viewpoint video by the first virtual camera. The virtual viewpoint video is generated as if it has switched from the virtual viewpoint camera information 2 to the virtual viewpoint camera information 1. Also, in the switching between two virtual viewpoint videos by two virtual viewpoints and the switching between two self-camera videos by two actual cameras, the virtual viewpoint video from the second virtual camera generated by the virtual camera information automatic generation unit 117 can be used.

[0052] As described above, according to the first embodiment, in the switching from the first video obtained by the first viewpoint to the second video obtained by the second viewpoint, a new virtual camera is generated so as to complement the between the first viewpoint and the second viewpoint. Then, by using the virtual viewpoint video by the new virtual viewpoint between the first video and the second video, it is possible to realize a switching as if the first video and the second video after switching were shot by one viewpoint (camera). Also, by smoothly switching between the virtual viewpoint video and the video of the actual camera, a more dynamic video expression that cannot be shot with the actual camera becomes possible.

[0053] <Second Embodiment> In the first embodiment, the process of generating the information of the virtual viewpoint (the second virtual camera) based on the information of the first virtual camera and the information of the real camera was described. The information of the virtual viewpoint includes the position, the posture (the direction of the line of sight), the focal length (the zoom value), etc. In the process of the first embodiment, these were generated by the same process without particularly distinguishing them. In the second embodiment, among the information of the virtual viewpoint, the position information and the posture information are generated by independent processes. Note that the same reference numerals are given to the configurations equivalent to those of the first embodiment, and the detailed description thereof is omitted.

[0054] As described above, in the first embodiment, the position information of the second virtual camera was generated so as to move between the position information of the first virtual camera and the position information of the real camera 112, and it was assumed that the posture of the second virtual camera could also be generated by the same method. However, in the method of the first embodiment, there is a problem that depending on the posture or the focal length of the second virtual camera, the subject to be photographed may not be included in the shooting range of the second virtual camera. In the second embodiment, in order to solve such a problem, the position of the second virtual camera, the posture of the second virtual camera, and the information of the focal length are independently controlled.

[0055] FIG. 7 is a block diagram showing a configuration example of the image processing system according to the second embodiment. It has a configuration in which a subject identification unit 701 is added to the configuration of the first embodiment (FIG. 1). The subject identification unit 701 identifies the subject being photographed by the virtual camera or the real camera 112. That is, the subject identification unit 701 identifies the subject moving in the video of the virtual camera or the real camera 112 based on the camera information from the virtual camera information generation unit 110, the real camera information acquisition unit 113, the virtual camera information automatic generation unit 117, and the information from the 3D model storage unit 109. In addition, the image acquisition unit 104 also provides the video acquired from the camera control unit 102 to the video switching unit 115. As a result, the video switching unit 115 can also use the video of the camera group 101 used for the virtual viewpoint video as the video output.

[0056] FIG. 8 is a flowchart showing the output video determination process according to the second embodiment. The same step numbers are assigned to the processes equivalent to those in the first embodiment (FIG. 2). In step S801, the virtual camera information automatic generation unit 117 refers to the switching information and determines whether the transition ratios of the position and orientation of the second virtual camera are different during the transition period from the first virtual viewpoint video 301 to the real camera video 302. If it is determined that the transition ratios are not different (NO in step S801), the process proceeds to step S211. On the other hand, if it is determined that the transition ratios are different (YES in step S801), the process proceeds to step S802.

[0057] In step S802, the virtual camera information automatic generation unit 117 generates information on the position, orientation, and field of view angle of the second virtual camera when switching from the virtual camera video to the real camera video based on the information of the first virtual camera, the information of the real camera 112, and the switching conditions. The virtual camera information automatic generation unit 117 acquires the position transition period for switching from the position of the first virtual camera included in the switching conditions to the position of the real camera 112 and the orientation transition period for switching from the orientation of the first virtual camera to the orientation of the real camera 112. In the switching conditions, for example, the position transition period and the orientation transition period are set independently of each other and are indicated by the start time and the end time, respectively. The virtual camera information automatic generation unit 117 calculates the position and orientation of the second virtual camera at each time. Note that, similar to the first embodiment, an input unit 600 provided with a fader 603 for specifying the switching ratio may be used. In that case, a fader 603 is provided individually for each condition to be controlled independently.

[0058] Also, the posture of the second virtual camera may be calculated at a transition ratio different from the position transition ratio so as to preferentially project the subject included in the output video after the switching. FIGS. 9A to 9B show an example of a process of generating virtual camera information so as to preferentially project the subject included in the output video after the switching in step S802. The positions and postures of the first virtual camera, the second virtual camera, and the actual camera 112 at each time are as shown in FIG. 4A. In the first virtual camera, the subject 901 exists in its shooting range as the mainly shot subject, and in the actual camera 112, the subject 902 exists in its shooting range as the mainly shot subject. During the position transition period (from time t2 to t7), the position of the second virtual camera transitions from the position of the first virtual camera to the position of the actual camera 112 in the same manner as in the first embodiment. On the other hand, the posture and focal length (zoom value) of the second virtual camera are steeply changed to have the same angle of view as the actual camera 112 between time t2 and time t4, which is the posture transition period. After that, between time t4 and time t7, the posture and focal length of the second virtual camera are set to have the same angle of view as the actual camera 112. Note that the same angle of view refers to the posture and angle of view set so that the same subject appears at approximately the same position in the videos obtained from each viewpoint. Or, it refers to the posture and angle of view set so that the same subject appears at approximately the same size in the videos obtained from each viewpoint. Or, it refers to the posture and angle of view set so that the position and size of the same subject appearing in the videos obtained from each viewpoint are approximately the same.

[0059] Based on the information of the position, orientation, and focal length of the first virtual camera from the virtual camera information generation unit 110 and the position of the foreground from the 3D model storage unit 109 by the object identification unit 701, it is possible to confirm at which position in the virtual viewpoint video acquired by the first virtual camera the foreground exists. Similarly, based on the information of the position, orientation, and focal length of the real camera 112 and the position of the foreground from the 3D model storage unit 109, it is possible to confirm at which position in the real camera video captured by the real camera 112 the foreground exists. The virtual camera information automatic generation unit 117 of the present embodiment calculates the orientation of the second virtual camera as if it were capturing a video with the same field of view as the video after switching, that is, the video of the real camera 112, during the transition period when the virtual viewpoint video by the second virtual camera is output.

[0060] In FIG. 9A, 9a shows the position 911 and orientation 912 of the first virtual camera at time t2, and the position 931 and orientation 932 of the real camera 112 at time t2. At time t2, the position and orientation of the second virtual camera are the same as the position 931 and orientation 932 of the first virtual camera. 9b in FIG. 9A shows the position 913 and orientation 914 of the first virtual camera at time t3, the position 933 and orientation 934 of the real camera 112, and the position 951 and orientation 954 of the second virtual camera. The orientation 954 of the second virtual camera at time t3 is determined based on the orientation 912 (orientation 952) of the first virtual camera at time t2 and the orientation 953 that allows the second virtual camera to obtain the same field of view as the real camera 112 at time t3. That is, the orientation 954 of the second virtual camera at time t3 is an orientation that is inclined from orientation 952 to orientation 953 by a ratio of (t3 - t2) / (t4 - t2) between orientation 952 and orientation 954.

[0061] In FIG. 9B, 9c shows the position 915 and orientation 916 of the first virtual camera at time t4, the position 935 and orientation 936 of the actual camera 112, and the position 955 and orientation 956 of the second virtual camera. Similar to the case of time t3, the orientation 956 of the second virtual camera at time t4 is determined based on the orientation 912 of the first virtual camera at time t2 and the orientation that allows the second virtual camera to obtain the same field of view as the actual camera 112 at time t4. However, at time t4, since (t4 - t2) / (t4 - t2) = 1, the orientation 956 that can obtain the same field of view as the actual camera 112 is determined as the orientation of the second virtual camera at time t4.

[0062] 9d in FIG. 9B shows the position 917 and orientation 918 of the first virtual camera at time t5, the position 937 and orientation 938 of the actual camera 112, and the position 957 and orientation 958 of the second virtual camera. The orientation 958 of the second virtual camera at time t5 is determined so as to obtain the same field of view as the actual camera 112 at time t5. Similarly, 9e in FIG. 9C shows the position 919 and orientation 920 of the first virtual camera at time t6, the position 939 and orientation 940 of the actual camera 112, and the position 959 and orientation 960 of the second virtual camera. The orientation 960 of the second virtual camera at time t6 is determined so as to obtain the same field of view as the actual camera 112 at time t6. 9f in FIG. 9C shows the position 921 and orientation 922 of the first virtual camera at time t7, and the position 941 and orientation 942 of the actual camera 112. At time t7, the position and orientation of the second virtual camera are the same as the position 941 and orientation 942 of the actual camera 112.

[0063] <Other Embodiments> In each of the above-described embodiments, the real camera 112 has been described as a camera brought to the periphery of the shooting range of the virtual viewpoint video, which is different from the camera group 101 that generates the virtual viewpoint video. However, the present invention is not limited to this. For example, as in the second embodiment, if the video of some or all of the cameras in the camera group 101 is sent to the video switching unit 115 and can be selected as the output video, the real camera 112 may be any one of the cameras in the camera group 101. Thereby, even when switching from the virtual viewpoint video to the real camera video by one real camera among the camera group 101 for generating the virtual viewpoint video, it becomes possible to easily generate a new virtual viewpoint video for the transition period of the switching of those videos.

[0064] Also, the generation of the virtual viewpoint during the transition period may be performed for each shooting frame of the real camera 112 during the transition period (or for each frame of the virtual viewpoint video by the first virtual viewpoint), or may be performed at a predetermined time interval (for example, every 0.5 seconds).

[0065] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. Further, it can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0066] The invention is not limited to the above-described embodiments, and various changes and modifications are possible without departing from the spirit and scope of the invention. Therefore, claims are attached to disclose the scope of the invention.

Explanation of Reference Numerals

[0067] 101: Camera group, 102: Camera control unit, 103: Image processing device, 104: Image acquisition unit, 105: Background image memory unit, 106: Separation unit, 107: Foreground image memory unit, 108: 3D model generation unit, 109: 3D model memory unit, 110: Virtual camera information generation unit, 111: Virtual viewpoint video generation unit, 112: Real camera, 113: Real camera information acquisition unit, 114: Video decision unit, 115: Video switching unit, 116: Video output unit, 117: Virtual camera information automatic generation unit

Claims

1. An acquisition unit that acquires information related to a first video and a second video, at least one of which is a captured video obtained by an imaging device, the acquisition unit acquiring information on a first viewpoint for obtaining the first video and information on a second viewpoint for obtaining the second video at a time corresponding to the time of the first video; A setting unit that sets a period from the end of the output of the first video to the start of the output of the second video when switching the output video from the first video to the second video; A first generation unit that generates information on a virtual viewpoint during the period based on the information on the first viewpoint during the period and the information on the second viewpoint during the period; A second generation unit that generates a virtual viewpoint video for the period based on the information on the virtual viewpoint during the period; An output unit that switches and outputs in the order of the first video, the virtual viewpoint video for the period, and the second video; characterized in that the first generation unit generates the virtual viewpoint for the period based on the information on the first viewpoint, the information on the second viewpoint, and the ratio of the elapsed time from the start of the period to the total time of the period. An image processing apparatus.

2. The first generation unit generates information on the virtual viewpoint for the period based only on the information on the first viewpoint at the start time of the period. The image processing apparatus according to claim 1, characterized in that.

3. An acquisition unit that acquires information related to a first video and a second video, at least one of which is a captured video obtained by an imaging device, the acquisition unit acquiring information on a first viewpoint for obtaining the first video and information on a second viewpoint for obtaining the second video at a time corresponding to the time of the first video; A setting unit that sets a period from the end of the output of the first video to the start of the output of the second video when switching the output video from the first video to the second video; A first generation unit that generates information on a virtual viewpoint during the period based on the information on the first viewpoint during the period and the information on the second viewpoint during the period; A second generation unit that generates a virtual viewpoint video for the period based on the information on the virtual viewpoint during the period; An output unit that switches and outputs in the order of the first video, the virtual viewpoint video for the period, and the second video; A setting unit that sets a ratio according to a user operation received during the period; characterized in that The first generation means generates a virtual viewpoint for the period based on the information of the first viewpoint, the information of the second viewpoint, and the ratio set by the setting means. An image processing apparatus characterized by the above.

4. The first generation means generates a virtual viewpoint for the period by weighted-averaging the information of the first viewpoint and the information of the second viewpoint based on the ratio. The image processing apparatus according to any one of claims 1 to 3, characterized by the above.

5. The first generation means generates a virtual viewpoint at each time in the period based on the information of the first viewpoint at the start time of the period and the information of the second viewpoint at each time. The image processing apparatus according to any one of claims 1 to 4, characterized by the above. An acquisition means for acquiring information related to a first video and a second video, at least one of which is a captured video obtained by an imaging device, the acquisition means acquiring information of a first viewpoint for obtaining the first video and information of a second viewpoint for obtaining the second video at a time corresponding to the time of the first video, A setting means for setting a period from the end of the output of the first video to the start of the output of the second video when switching the output video from the first video to the second video, A first generation means for generating information of a virtual viewpoint in the period based on the information of the first viewpoint in the period and the information of the second viewpoint in the period, A second generation means for generating a virtual viewpoint video for the period based on the information of the virtual viewpoint in the period, An output means for switching and outputting the first video, the virtual viewpoint video for the period, and the second video in this order, comprising The first generation means generates a virtual viewpoint at each time in the period based on the information of the first viewpoint at each time and the information of the second viewpoint at each time. An image processing apparatus characterized by the above.

7. Further comprising an identification means for identifying a subject from the video captured from the second viewpoint, The first generation means generates information on the direction of the line of sight included in the information of the virtual viewpoint for the period based on the position of the subject identified by the identification means. The image processing apparatus according to any one of claims 1 to 6, characterized by the above.

8. The first generation means generates information on the direction of the line of sight of the virtual viewpoint included in the information of the virtual viewpoint for the period based on the direction of the line of sight of the virtual viewpoint for obtaining an image of a shooting range in which the position of the subject reflected in the virtual viewpoint image is the same as the position of the subject reflected in the image obtained by the second viewpoint of the subject, and the direction of the line of sight of the first viewpoint at the start of the period. The image processing apparatus according to claim 7, characterized in that.

9. The first generation means generates information on the focal length of the virtual viewpoint for obtaining an image of a shooting range in which the size of the subject reflected in the virtual viewpoint image is the same as the size of the subject reflected in the image obtained by the second viewpoint of the subject, and the focal length of the line of sight of the first viewpoint at the start of the period, and generates information on the focal length of the virtual viewpoint for the period. The image processing apparatus according to claim 7 or 8, characterized in that.

10. One of the first video and the second video is a virtual viewpoint video generated based on a plurality of images captured by a plurality of imaging devices and a virtual viewpoint. The image processing apparatus according to any one of claims 1 to 9, characterized in that.

11. The second generation means further includes connection means for connecting to a plurality of imaging devices for obtaining a plurality of images for generating a virtual viewpoint image, The virtual viewpoint video is generated based on the plurality of images. The image processing apparatus according to claim 10, characterized in that.

12. An acquisition step of acquiring information related to a first video and a second video, at least one of which is an imaging video obtained by an imaging device, the acquisition step of acquiring information on a first viewpoint for obtaining the first video and information on a second viewpoint for obtaining the second video at a time corresponding to the time of the first video; A setting step of setting a period from the end of the output of the first video to the start of the output of the second video when switching the output video from the first video to the second video; A first generation step of generating information on a virtual viewpoint for the period based on the information on the first viewpoint for the period and the information on the second viewpoint for the period; A second generation step of generating a virtual viewpoint video for the period based on the information on the virtual viewpoint for the period; An output step of switching and outputting the first video, the virtual viewpoint video for the period, and the second video in this order; having In the first generation step, a virtual viewpoint of the period is generated based on the information of the first viewpoint, the information of the second viewpoint, and the ratio of the elapsed time from the start of the period to the total time of the period. A control method for an image processing apparatus, characterized in that.

13. An acquisition step of acquiring information related to a first video and a second video, at least one of which is an imaging video obtained by an imaging device, the acquisition step of acquiring information of a first viewpoint for obtaining the first video and information of a second viewpoint for obtaining the second video at a time corresponding to the time of the first video. A setting step of setting a period from the end of the output of the first video to the start of the output of the second video when switching the output video from the first video to the second video. A first generation step of generating information of a virtual viewpoint in the period based on the information of the first viewpoint in the period and the information of the second viewpoint in the period. A second generation step of generating a virtual viewpoint video of the period based on the information of the virtual viewpoint in the period. An output step of switching and outputting the first video, the virtual viewpoint video of the period, and the second video in this order. A setting step of setting a ratio according to a user operation received during the period. having In the first generation step, a virtual viewpoint of the period is generated based on the information of the first viewpoint, the information of the second viewpoint, and the ratio set by the setting step. A control method for an image processing apparatus, characterized in that.

14. An acquisition step of acquiring information related to a first video and a second video, at least one of which is an imaging video obtained by an imaging device, the acquisition step of acquiring information of a first viewpoint for obtaining the first video and information of a second viewpoint for obtaining the second video at a time corresponding to the time of the first video. A setting step of setting a period from the end of the output of the first video to the start of the output of the second video when switching the output video from the first video to the second video. A first generation step of generating information of a virtual viewpoint in the period based on the information of the first viewpoint in the period and the information of the second viewpoint in the period. A second generation step of generating a virtual viewpoint video of the period based on the information of the virtual viewpoint in the period. An output step of switching and outputting the first video, the virtual viewpoint video of the period, and the second video in this order. having In the first generation step, a virtual viewpoint at each time during the period is generated based on the information of the first viewpoint and the information of the second viewpoint at each time. A method for controlling an image processing apparatus, characterized by the above. **Claim 15** A program for causing a computer to function as the image processing apparatus according to any one of Claims 1 to 11.

Citation Information

Patent Citations

  • Method, device and program for generating free visual point image by local area division

    JP2008015756A

  • Video receiver and control method thereof

    JP2012015788A

  • Information processing apparatus, control method thereof, and program

    JP2020042665A

  • Information processing device, method, and recording media

    JP2020150417A