Video processing device, control method and program thereof

By introducing a line of sight indicator and a line of sight information generation unit into the video processing device, the actor's line of sight is detected in real time and a virtual viewpoint image is generated to make it consistent with the actor's line of sight direction, and the problem of line of sight delay and fine adjustment in the virtual viewpoint image is solved, and the accurate matching of line of sight direction is achieved.

JP7672283B2Active Publication Date: 2025-05-07CANON KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021090584
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-28
Publication Date
2025-05-07
Estimated Expiration
2041-05-28

AI Technical Summary

Technical Problem

When generating virtual viewpoint images, it is difficult to accurately point the actor's sight to the virtual viewpoint, especially during delayed three-dimensional model generation and video generation.

Method used

By introducing a line of sight indicator and a line of sight information generation unit into the video processing device, the line of sight of the actor is detected in real time and a virtual viewpoint image is generated so that it is consistent with the line of sight direction of the actor.

Benefits of technology

It realizes that when generating virtual viewpoint images, accurately match the actor's line of sight direction, solving the problems of line of sight delay and fine adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672283000001
    Figure 0007672283000001
  • Figure 0007672283000002
    Figure 0007672283000002
  • Figure 0007672283000003
    Figure 0007672283000003
Patent Text Reader

Abstract

To generate virtual viewpoint video to an object's sight line.SOLUTION: A video processing device for generating virtual viewpoint video on the basis of video obtained by multiple imaging parts includes: a creation part for creating information for locating a position that an object should face; an imaging control part for acquiring video from multiple imaging parts and storing it in a predetermined storage part while a symbol for expressing a position that an object should face is under display according to information created by the creation part; and a generation part for locating a position of an object from video stored in the storage part, setting a position for displaying a symbol to a position of a virtual viewpoint, setting a direction pointed to a position of the virtual viewpoint and a position of an object to a sight line direction from a virtual viewpoint, and generating virtual viewpoint video for expressing appearance from a virtual viewpoint on the basis of a position of the virtual viewpoint and a sight line direction from the virtual viewpoint.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a technology for generating a virtual viewpoint video. [Background technology]

[0002] There is a virtual viewpoint image generating system that can create an image seen from a virtual viewpoint designated by a user from images captured by multiple cameras and play it as a virtual viewpoint image. For example, see Patent Document 1. The system disclosed in this document transmits images captured by multiple cameras, and then extracts images with large changes from the captured images as foreground images and images with small changes as background images in an image computing server (image processing device). The system then estimates and generates the shape of a three-dimensional model of the subject based on the extracted foreground images, and stores the shape in a storage device together with the foreground image and background image. These foreground image and background image are called materials for generating a virtual viewpoint image. Then, appropriate data is obtained from the storage device based on the virtual viewpoint designated by the user, and a virtual viewpoint image is generated.

[0003] One example of content using virtual viewpoint video is filming a subject, such as a singer, and distributing the video to end users in real time. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2019-050593 A Summary of the Invention [Problem to be solved by the invention]

[0005] When shooting and distributing such video of a subject, it is common to have the subject look in the direction of the camera. However, in the case of virtual viewpoint video, the subject (performer) does not know the location of the virtual viewpoint, so it is difficult for him or her to look in that direction.

[0006] On the other hand, a method such as using a display that shows the position of the virtual viewpoint within the shooting area can also be considered. However, to generate a virtual viewpoint image, a three-dimensional model is generated from shooting data, and then an image is generated based on the virtual viewpoint information, so there is a delay of several seconds between shooting and the generation of the virtual viewpoint image. Therefore, even if the virtual viewpoint from which the image is actually generated is displayed, the line of sight would be directed to a virtual viewpoint that is several seconds behind, making it difficult to direct the line of sight.

[0007] Another possible method would be to use image recognition technology to detect the subject's line of sight and move the virtual viewpoint in that direction, but there is a risk that the virtual viewpoint will react to small movements of the subject's line of sight, and it would be difficult to accommodate line of sight movements such as turning around. [Means for solving the problem]

[0008] The present disclosure provides a technique for solving the above problem. To this end, a video processing device according to the present disclosure has the following configuration. A video processing device that generates a virtual viewpoint video based on videos obtained by a plurality of imaging means, A creating means for creating information that identifies a position to which a subject should be directed; an imaging control means for acquiring images from the plurality of imaging means and storing the images in a predetermined storage means while displaying a symbol representing a position to which the subject should be directed in accordance with the information created by the creation means; The apparatus further comprises a generating means for identifying a position of the subject from the image stored in the memory means, defining a position at which the symbol is displayed as a virtual viewpoint position, defining the position of the virtual viewpoint and a direction directed toward the position of the subject as a line of sight direction from the virtual viewpoint, and generating a virtual viewpoint image representing the appearance from the virtual viewpoint based on the position of the virtual viewpoint and the line of sight direction from the virtual viewpoint. Effect of the Invention

[0009] According to the present disclosure, it is possible to generate a virtual viewpoint image that matches the line of sight of a subject. [Brief description of the drawings]

[0010] [Figure 1] FIG. 1 is a system configuration diagram according to a first embodiment. [Diagram 2] 5A and 5B are diagrams showing the relationship between a subject, a gaze instruction, and gaze information in the first embodiment. [Diagram 3] FIG. 4 is a diagram showing a processing flow for generating gaze information. [Figure 4] FIG. 11 is a diagram showing a processing flow for generating an image using line-of-sight information. [Diagram 5] FIG. 13 is a diagram showing a virtual viewpoint that connects a current virtual viewpoint and a virtual viewpoint based on viewpoint information. [Figure 6] 11A and 11B are diagrams showing the relationship between line-of-sight information and the angle of view at a virtual viewpoint for a plurality of subjects. [Figure 7] FIG. 13 is a diagram showing a plurality of gaze instructions and a display example of the gaze instructions. [Figure 8] FIG. 13 is a block diagram showing an example of the hardware configuration of an applicable computer. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.

[0012] [First embodiment] <Basic system configuration and operation> The configuration of a video processing system for generating a virtual viewpoint video according to this embodiment is shown in Fig. 1. This system includes a plurality of imaging units 1, a synchronization unit 2, a three-dimensional shape estimation unit 3, an accumulation unit 4, a virtual viewpoint operation unit 5, a virtual viewpoint setting unit 6, a video generation unit 7, a display unit 8, a gaze instruction unit 9, a subject position detection unit 10, and a gaze information generation unit 11. This system may be formed by one electronic device or may be formed by multiple electronic devices.

[0013] Next, an outline of the operation of each component of the present system will be described, followed by a detailed description of the characteristic components of this embodiment. First, the multiple imaging units 1 capture images in high-precision synchronization with one another based on a synchronization signal from the synchronization unit 3. For example, if the synchronization unit 3 outputs a synchronization signal at 1 / 30 second intervals, the imaging units 1 will capture images in synchronization at 1 / 30 second intervals. The imaging units 1 output the captured images to the three-dimensional shape estimation unit 3. The multiple imaging units 1 are installed to surround the subject in order to capture the subject from multiple viewpoints and multiple directions.

[0014] The three-dimensional shape estimation unit 3 uses the input captured images from multiple viewpoints to extract, for example, a silhouette of the subject, and then generates a three-dimensional model of the subject using a volume intersection method or the like. The three-dimensional shape estimation unit 3 then outputs the generated three-dimensional model of the subject and the captured images of the subject to the storage unit 4. The subject in this case includes people for whom the three-dimensional shape model is to be generated and items handled by these people, but for simplicity, the present embodiment will be described assuming that the subject is only people.

[0015] The storage unit 4 saves and accumulates the following data group as virtual viewpoint video material. Specifically, these are the photographed images of the subject and the three-dimensional model of the subject input from the three-dimensional shape estimation unit 3. Also, these are camera parameters such as the position and orientation and optical characteristics of each imaging unit. Also, a background model and background texture image may be recorded in advance to render the background, but as this is not the main focus of this embodiment, the background will not be mentioned here.

[0016] The virtual viewpoint operation unit 5 includes a physical user interface such as a joystick and an operation unit for numerical data, and supplies these input information to a virtual viewpoint setting unit 6 .

[0017] The virtual viewpoint setting unit 6 sets virtual viewpoint position information based on input information from the virtual viewpoint operation unit 5, and supplies it to the image generation unit 7. The virtual viewpoint position information is composed of information equivalent to external parameters of the camera, such as the position and orientation of the virtual viewpoint, information equivalent to internal parameters of the camera, such as the focal length and angle of view, and information related to time.

[0018] Based on the input information on the virtual viewpoint position and time, the video generation unit 7 acquires data for the corresponding time from the storage unit 4. The video generation unit 7 renders the subject at the virtual viewpoint from the subject 3D model and the subject photographed image among the acquired data, generates a virtual viewpoint video, and outputs it to the display unit 8. At this time, the background at the virtual viewpoint may be rendered from a background model and a background texture image together with the subject.

[0019] <Operation of the characteristic configuration of this embodiment> Next, the characteristic components of the embodiment and their specific operations will be described. The description will be divided into two parts: a description of the generation and storage of line-of-sight information, and a description of image generation using line-of-sight information.

[0020] First, the functions of the gaze instruction unit 9, the subject position detection unit 10, and the gaze information generation unit 11, which are characteristic components of this embodiment related to gaze information generation, will be described.

[0021] The gaze instruction unit 9 in this embodiment is composed of a display device such as a display or projector, and a display control device that manages and controls the images displayed on the display device in a time series. Hereinafter, unless otherwise specified, the gaze instruction unit 9 refers to the display device. The gaze instruction unit 9 is a display device that is placed in or around the shooting area, and displays a symbol that indicates the position of the viewpoint instruction that is visible from the subject. The displayed symbol may be, for example, a circle, a cross, or other mark.

[0022] The subject position detection unit 10 detects the position of the subject within the shooting area. In this embodiment, the subject position is detected using a three-dimensional model of the subject, which is the shape estimation result by the three-dimensional shape estimation unit 3. Specifically, the subject position, particularly the head position, is detected by detecting approximately 25 cm above the three-dimensional model of the subject as the head and determining the center of gravity.

[0023] The gaze instruction unit 9 instructs the direction and position of the subject's gaze. The display position of the gaze instruction and the subject position detected by the subject position detection unit 10 are calibrated in advance with high accuracy to the three-dimensional shape estimation unit 3 and the imaging unit 1, and the position information indicated by each piece of information can be handled in the same coordinate space.

[0024] The gaze information generating unit 11 records the gaze information in the storage unit 4 based on the information of the gaze instruction position indicated by the gaze instruction unit 9 and the position information of the subject detected by the subject position detecting unit 10. The gaze information at this time is recorded as information of a straight line (line segment) connecting the subject position and the gaze instruction position as shown in FIG.

[0025] Next, the line-of-sight information generation process of the line-of-sight information generating unit 11 will be described with reference to the flowchart of FIG.

[0026] In S300, the line-of-sight information generating unit 11 makes a plan regarding the content to be shot, the time from the reference time at which to display the line-of-sight instruction in accordance with the music, and the position and direction in which to display it.

[0027] Next, in S301, the gaze information generating unit 11 registers the planned number of gaze instructions in the gaze instructing unit 9 with respect to the planned gaze instructions.

[0028] Next, in S302, the gaze information generating unit 11 starts the operation of the gaze instruction unit 9 in synchronization with a reference time for starting shooting. For example, in the case of a music video or concert video, this reference time means starting the operation of the gaze instruction unit 9 in synchronization with the start time of a song.

[0029] Thereafter, in S303, at the scheduled time, the gaze information generation unit 11 displays a mark on the display unit of the gaze instruction unit 9 at a position registered in advance. The mark displayed at this time can be moved based on the registration information. In addition, in S304, the gaze information generation unit 11 causes the subject position detection unit 10 to detect the position of the subject. Then, in S305, the gaze information generation unit 11 generates gaze information based on the gaze instruction position and the subject position, and records it in the storage unit 4. By repeating the operations from S303 to S305 above for the scheduled time and number of times, the gaze information is generated and recorded in addition to the recording of the virtual viewpoint video material.

[0030] Next, a process flow regarding the operation of the virtual viewpoint setting unit 6 and the image generating unit 7 using the line of sight information will be described with reference to the flowchart of FIG.

[0031] First, the virtual viewpoint setting unit 6 determines whether or not line-of-sight information at a certain time (at the initial stage, at the start of video recording) is stored in the storage unit 4. If the virtual viewpoint setting unit 6 determines that line-of-sight information does not exist, it proceeds to processing at S400. In this S400, the virtual viewpoint setting unit 6 generates virtual viewpoint information for that time (period) based on the input from the virtual viewpoint operation unit 5. This virtual viewpoint information includes the position of the virtual viewpoint, and the optical axis direction and angle of view at the virtual viewpoint.

[0032] On the other hand, when the virtual viewpoint setting unit 6 determines that the line of sight information at a certain time (at the start of video recording in the initial stage) is stored in the storage unit 4, the process proceeds to S401. In this S401, the virtual viewpoint setting unit 6 generates virtual viewpoint information based on the line of sight information stored in the storage unit 4. In the case of the embodiment, this virtual viewpoint information includes information regarding the position of a mark displayed to specify the direction of the subject's line of sight as the position of the virtual viewpoint, information regarding the line of sight direction from the position of the virtual viewpoint toward the subject (the head) as the optical axis, and information indicating a predetermined angle of view. Then, in S402, the virtual viewpoint setting unit 6 performs a correction process of the virtual viewpoint information generated in S401. Specifically, the virtual viewpoint setting unit 6 generates virtual viewpoint information to which a change in the angle of view or a linear forward / backward movement is applied within a permitted range by the input of the virtual viewpoint operation unit 5. This allows the virtual viewpoint position to be brought closer to the subject, or the angle of view can be changed while maintaining the state of the camera's line of sight, such as zooming.

[0033] Then, in S403, the virtual viewpoint setting unit 6 sets the generated virtual viewpoint information in the video generating unit 7. The above processes of S400 to S403 are performed at each time during video generation.

[0034] Every time virtual viewpoint information is set by the virtual viewpoint setting unit 6, the video generation unit 7 uses the captured images of the multiple viewpoints at the corresponding time stored in the storage unit 4 to generate a virtual viewpoint image (one frame in a virtual viewpoint video) based on the set virtual viewpoint information, and outputs it to the display unit 8. In the above, the video generation unit 7 has been described as outputting the generated virtual viewpoint image to the display unit 8, but it may also be stored in a storage device as one frame of a virtual viewpoint video file, or distributed over a network.

[0035] As explained above, during the time range stored as virtual information in accumulation unit 4, a virtual viewpoint position is set in the direction in which the subject is looking, and it is possible to generate a virtual viewpoint image facing the subject from that position, i.e., an image from the subject's line of sight.

[0036] By combining these two processing flows, it is possible to give appropriate gaze instructions to a subject in the shooting space of the virtual viewpoint video, and gaze information generated from the positional relationship between the instruction position and the subject at the time the instruction is given is recorded. In addition, by using the gaze information, it is possible to generate a virtual viewpoint video from the direction the subject would be gazing at the time the video is generated, even if there is a delay in generating the virtual viewpoint video.

[0037] [Other embodiments of the present invention] In this embodiment, the gaze instruction unit 9 is not limited to a configuration using a display device such as a display, but may be configured to include a physical mark. For example, it may be configured with a sphere attached to the end of a rod, and the subject is instructed to direct their gaze to the sphere. By configuring it to include a motion tracking means that can detect the sphere, the gaze instruction position can be obtained three-dimensionally. In this case, the device has a tracking unit that records the position of the physical marker at the scheduled time.

[0038] In this embodiment, the subject position detection unit 10 detects the head of the subject 3D model estimated by the 3D shape estimation unit 3, but this is not limited to the above. For example, the subject position detection unit 10 may simply detect the center of gravity of the subject 3D model or the center of gravity of a circumscribed rectangular parallelepiped (bounding box). Alternatively, the subject position may be detected by a sensor such as a GPS without using the subject's 3D shape, or the subject's face and head positions may be detected from the captured image using image recognition technology.

[0039] In addition, in the present embodiment, when a virtual viewpoint is generated based on viewpoint information, it has been described that the movement of the virtual viewpoint in a straight line is permitted, but specifically, a restriction may be imposed such that the virtual viewpoint cannot approach the subject by more than a preset distance, etc. This makes it possible to prevent, for example, the virtual viewpoint from moving around inside or behind the subject, causing the subject to be outside the angle of view.

[0040] Furthermore, when it is desired to input to the virtual viewpoint simulating camera shake, the restriction of the line of sight information to a straight line may be relaxed and slight movement outside the straight line may be permitted, and movement within that range may be made possible by input from the virtual viewpoint operation unit 5. In addition, the simulation of camera shake is not limited to input from the virtual viewpoint operation unit 5, and vibration (movement) simulating camera shake may be applied in the virtual viewpoint setting unit. For example, the position of the mark indicating the line of sight direction set by the gaze instruction unit 9 or the position of the virtual viewpoint set by the virtual viewpoint operation unit 5 may be added with a coordinate change amount equivalent to a vibration that changes randomly on the time axis.

[0041] In the above embodiment, the virtual viewpoint setting unit 6 automatically moves the virtual viewpoint to a virtual viewpoint position based on the line-of-sight information at a time when the line-of-sight information is available, but this is not necessarily limited to this. For example, as shown in Fig. 5, the virtual viewpoint setting unit 6 may generate intermediate virtual viewpoint information, which is virtual viewpoint information that moves continuously and smoothly from the current virtual viewpoint position to the virtual viewpoint position based on the line-of-sight information prior to the time when the viewpoint information is recorded, based on a user instruction. In addition, even if the time is when the line-of-sight information is recorded, it may be possible to select not to move to the virtual viewpoint position based on the line-of-sight information based on a user instruction.

[0042] In addition, in a configuration in which a plurality of image generating units 7 are provided and virtual viewpoint images are simultaneously generated and a plurality of virtual viewpoint images are switched using an image switching device or the like, it is not necessary to set all virtual viewpoints as virtual viewpoints based on line-of-sight information. Specifically, it is possible to set only one selected from the plurality of virtual viewpoint images to generate a virtual viewpoint image based on line-of-sight information, and then switch between the other virtual viewpoint images and the virtual viewpoint based on line-of-sight information using the image switching device to output the image.

[0043] In addition, although the present embodiment has been described with respect to one viewpoint instruction unit 9 and one subject, the present invention is not necessarily limited to this configuration. For example, as shown in FIG. 6(a), a configuration may be adopted in which multiple subjects are the targets of gaze information generation for one gaze instruction. In this case, the subject position detection unit 10 detects the position information of each subject, and the gaze information generation unit 11 generates multiple straight lines connecting the gaze instruction position and the multiple subject positions as gaze information, and records them in the storage unit 4. In addition, when generating a virtual viewpoint video, multiple virtual viewpoints may be set based on each gaze information as shown in FIG. 6(b), and multiple virtual viewpoint videos may be generated. In addition, as shown in FIG. 6(c), one virtual viewpoint may be set at a point (gaze instruction position) where multiple gaze information intersects, and the direction and angle of view of the virtual viewpoint may be set so as to include the direction based on the two gaze information. In FIG. 6(c), an angle of view in which two subjects are evenly arranged is described as an example, but the virtual viewpoint setting unit 6 may be configured to set a virtual viewpoint so that one of the subjects is located in the center, and then adjust the angle of view so that the other person fits within the angle of view.

[0044] In addition, although the configuration using one viewpoint instruction unit 9 has been shown in this embodiment, the present invention is not limited to this. For example, as shown in FIG. 7(a), a configuration may be adopted in which multiple viewpoint instructions are provided and gaze instructions are given to multiple subjects. In this case, the gaze information generating unit 11 generates gaze information for each subject position corresponding to the position indicated by the gaze instruction unit 9, and records multiple gaze information in the storage unit 4 at the same time. This makes it possible to simultaneously generate images of the camera gaze of each subject when performing multi-viewpoint video distribution such as providing multiple virtual viewpoint videos at the same time. Here, multiple viewpoint instruction units 9 are used for convenience, but in the case of a viewpoint instruction unit 9 constituted by a display device such as a display, it is not necessary to have multiple viewpoint instruction units 9. For example, as shown in FIG. 7(b) and (c), a single display device may be configured to indicate multiple gaze instructions with different marks. In addition, in order to enable the subject to identify the mark to which the subject is directing his or her gaze, a configuration may be adopted in which gaze instructions are given to each subject by displaying information for identifying the subject (in the illustrated example, a name) on the mark.

[0045] [Other configurations] In the above embodiment, each processing unit shown in Fig. 1 has been described as being configured with hardware. However, the configuration shown in Fig. 1 except for the imaging unit 1 may be realized with the hardware of an information processing device such as a personal computer and an application program to be executed.

[0046] FIG. 8 is a block diagram showing an example of the hardware configuration of a computer applicable to each of the above-described embodiments.

[0047] The CPU 801 controls the entire computer using computer programs and data stored in the RAM 802 and the ROM 803, and executes the processes described above as being performed by the indirect position estimation device according to each of the above embodiments. That is, the CPU 801 functions as each processing unit except for the imaging unit 1 shown in FIG.

[0048] The RAM 802 has an area for temporarily storing computer programs and data loaded from an external storage device 806, data acquired from the outside via an I / F (interface) 807, etc. Furthermore, the RAM 802 has a work area used when the CPU 801 executes various processes. That is, the RAM 802 can be allocated as a frame memory, for example, or can provide various other areas as appropriate.

[0049] The ROM 803 stores setting data for the computer, a boot program, and the like. The operation unit 804 is composed of a keyboard, a mouse, and the like, and can be operated by a user of the computer to input various instructions to the CPU 801. The output unit 805 displays the results of processing by the CPU 801. The output unit 805 is composed of, for example, a liquid crystal display. For example, the virtual viewpoint operation unit 5 is composed of the operation unit 804, and the display unit 8 is composed of the output unit 805.

[0050] The external storage device 806 is a large-capacity information storage device, such as a hard disk drive. The external storage device 806 stores an OS (operating system) and computer programs for causing the CPU 801 to realize the functions of each unit shown in Fig. 1. Furthermore, the external storage device 806 may store each image data to be processed. The external storage device 806 also functions as the accumulation unit 4 in Fig. 1, and is further used for storing schedule data of the line-of-sight direction along the time axis created by the line-of-sight instruction unit 9.

[0051] Computer programs and data stored in the external storage device 806 are loaded into the RAM 802 as appropriate under the control of the CPU 801, and become the subject of processing by the CPU 801. A network such as a LAN or the Internet, a plurality of image capture devices 1 corresponding to the image capture unit 1 in Fig. 1, and other devices such as a gaze instruction unit 9 can be connected to the I / F 807, and the computer can obtain and transmit various information via this I / F 807. 808 is a bus that connects the above-mentioned units.

[0052] The operation of the above-described configuration is controlled mainly by the CPU 801, as described in the above embodiment.

[0053] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0054] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0055] 1...imaging unit, 2...synchronization unit, 3...three-dimensional shape estimation unit, 4...storage unit, 5...virtual viewpoint operation unit, 6...virtual viewpoint setting unit, 7...image generation unit, 8...display unit, 9...gaze instruction unit, 10...subject position detection unit, 11...gaze information generation unit

Claims

1. A video processing device that generates a virtual viewpoint video based on videos obtained by a plurality of imaging means, A creating means for creating information that identifies a position to which a subject should be directed; an imaging control means for acquiring images from the plurality of imaging means and storing the images in a predetermined storage means while displaying a symbol representing a position to which the subject should be directed in accordance with the information created by the creation means; a generating means for identifying a position of the subject from the image stored in the storage means, defining a position where the symbol is displayed as a virtual viewpoint, defining the position of the virtual viewpoint and a direction toward the position of the subject as a line of sight from the virtual viewpoint, and generating a virtual viewpoint image representing a view from the virtual viewpoint based on the position of the virtual viewpoint and the line of sight from the virtual viewpoint; A video processing device comprising:

2. 2 . The image processing apparatus according to claim 1 , wherein the generating means generates information for specifying a position representing a line-of-sight direction along a time axis to which the subject should be directed.

3. When there are a plurality of subjects, the creating means creates information for identifying the position indicating the line of sight of each of the subjects, and the imaging control means displays a symbol that is identifiable by each of the subjects in accordance with the information created by the creating means.

3. The image processing device according to claim 1, wherein the first and second inputs are connected to the first and second inputs.

4. 4. The image processing device according to claim 1, wherein the generating means generates a virtual viewpoint image having an optical axis in a direction from the position of the virtual viewpoint toward the position of the subject.

5. 4. The image processing device according to claim 1, wherein, when the subjects are a plurality of subjects, the generating means generates a virtual viewpoint image having an angle of view that is a range including the plurality of subjects.

6. a designation means for designating a virtual viewpoint position and an attitude of the virtual viewpoint through a user's operation, The generating means generates a virtual viewpoint image according to the virtual viewpoint position and orientation instructed by the instructing means during a period in which the creating means does not have information specifying the line of sight direction to which the subject should be directed.

6. The image processing device according to claim 1,

7. A control method for a video processing device that generates a virtual viewpoint video based on videos obtained by a plurality of imaging means, comprising the steps of: A creating step of creating information that identifies a position to which the subject should be directed; an imaging control step of acquiring images from the plurality of imaging means and storing the images in a predetermined storage means while a symbol representing a position to which the subject should be directed is displayed in accordance with the information created in the creating step; a generating step of identifying a position of the subject from the image stored in the storage means, determining a position at which the symbol is displayed as a virtual viewpoint position, determining the position of the virtual viewpoint and a direction toward the position of the subject as a line of sight direction from the virtual viewpoint, and generating a virtual viewpoint image representing a view from the virtual viewpoint based on the position of the virtual viewpoint and the line of sight direction from the virtual viewpoint; A control method for a video processing device comprising:

8. A program for causing a computer to execute each step of the method according to claim 7 when the program is read and executed by the computer.

Citation Information

Patent Citations

  • Photographic system

    JP2006217366A

  • Image processing apparatus

    JP2009237899A

  • Image processing system, image processor, control method, and program

    JP2019050593A

  • Information processing apparatus, display control method, and storage medium

    US20190124316A1

  • Image processing device, image processing method for image processing device, and program

    WO2019012817A1