Information processing device, information processing method, data structure and program

The information processing device facilitates precise control over virtual camera paths and subject visibility in virtual viewpoint videos, addressing the challenges of existing technologies to create powerful and desired virtual viewpoint videos.

JP7795931B2Active Publication Date: 2026-01-08CANON KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022013582
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-31
Publication Date
2026-01-08
Estimated Expiration
2042-01-31

AI Technical Summary

Technical Problem

Existing technologies for generating virtual viewpoint videos struggle to precisely control the transition of virtual camera positions, postures, and angles, making it difficult to create powerful and desired virtual viewpoint videos.

Method used

An information processing device that generates control information, including virtual viewpoint and setting information, to specify the virtual camera path and subject display, allowing for precise control over the virtual viewpoint video generation.

Benefits of technology

Enables the easy generation of desired virtual viewpoint videos by accurately controlling the virtual camera path and subject visibility, enhancing the quality and precision of the generated content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007795931000001
    Figure 0007795931000001
  • Figure 0007795931000002
    Figure 0007795931000002
  • Figure 0007795931000003
    Figure 0007795931000003
Patent Text Reader

Abstract

To make it easier to generate a desired virtual viewpoint picture.SOLUTION: There are acquired pieces of information for specifying a virtual viewpoint in a frame of a virtual viewpoint picture and information for specifying an object to display in the frame of the virtual viewpoint picture of a plurality of objects. There is output control information including virtual viewpoint information for specifying a virtual viewpoint for the frame of the virtual viewpoint picture and setting information for specifying the object displayed in the frame of the virtual viewpoint picture.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing method, a data structure, and a program, and in particular to a technology for generating a virtual viewpoint video. [Background technology]

[0002] A technology that generates a virtual viewpoint video using the multiple images obtained by installing multiple imaging devices in different positions and capturing images synchronously has attracted attention. This technology for generating a virtual viewpoint video using images from multiple viewpoints allows a video producer to create powerful content from any viewpoint using footage captured of, for example, a soccer or basketball game. In this case, the video producer specifies the optimal position and posture of the virtual viewpoint (virtual camera path) to generate a powerful video depending on the game scene, such as the movement of the players or the ball. Patent Document 1 discloses a technology for setting the virtual camera path by operating a device or a UI screen. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2017-212592 Summary of the Invention [Problem to be solved by the invention]

[0004] According to the technology described in Patent Document 1, the transition of the position, posture, and angle of view of the virtual viewpoint is specified as the virtual camera path. However, in order to create a powerful virtual viewpoint video, it is necessary to not only generate a virtual viewpoint video from a virtual viewpoint according to these parameters, but also to control the video generation more precisely.

[0005] The present disclosure aims to make it easier to generate a desired virtual viewpoint image. [Means for solving the problem]

[0006] An information processing device according to an embodiment of the present disclosure has the following configuration: a viewpoint acquisition means for acquiring information specifying a virtual viewpoint in a frame of a virtual viewpoint video; a setting acquisition means for acquiring information specifying a subject to be displayed in the frame of the virtual viewpoint video among a plurality of subjects; an output means for outputting control information including virtual viewpoint information for specifying the virtual viewpoint for the frame of the virtual viewpoint video and setting information for specifying the subject displayed in the frame; With death, The output means outputs the control information as virtual camera path data, and the virtual camera path data records the virtual viewpoint information and the setting information for each frame. . [Effects of the Invention]

[0007] According to the present disclosure, it is possible to easily generate a desired virtual viewpoint video. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of a virtual viewpoint image generation system according to an embodiment. [Figure 2] FIG. 10 is a diagram showing an example of a format of sequence data including virtual camera path data. [Figure 3] FIG. 10 is a diagram showing an example of a format of virtual camera path data. [Figure 4] FIG. 10 is a diagram showing an example of a format of virtual camera path data. [Figure 5] 10A and 10B are diagrams for explaining a method of generating an image according to display subject setting information. [Figure 6] 10A and 10B are diagrams for explaining an image generation method according to coloring camera setting information. [Figure 7] 10A and 10B are diagrams for explaining a video generation method according to rendering area setting information. [Figure 8] 1 is a flowchart of an information processing method according to one embodiment. [Figure 9] FIG. 1 is a diagram showing an example of the configuration of an information processing apparatus according to an embodiment. [Figure 10] 1 is a flowchart of an information processing method according to one embodiment. [Figure 11] FIG. 2 is a diagram showing an example of the hardware configuration of a computer used in an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0010] An embodiment of the present disclosure relates to a technology for generating control information used to generate a virtual viewpoint video including a subject viewed from a virtual viewpoint, and a technology for generating a virtual viewpoint video including a subject viewed from a virtual viewpoint in accordance with such control information. According to one embodiment, such control information includes setting information related to video generation, which setting information includes information specifying which of multiple subjects to display in each frame of the virtual viewpoint video. Such setting information can be used to set whether to display or hide a specific subject. With this configuration, for example, it is possible to hide one of multiple subjects and control the visibility of the subject behind it. In particular, when generating a virtual viewpoint video based on images captured by multiple imaging devices, unlike computer graphics video production, it is difficult for a video creator to easily control the positional relationship between each subject. As a result, a desired subject may be hidden by other subjects in the virtual viewpoint video from the desired virtual viewpoint. On the other hand, hiding other subjects using such setting information makes it easier to generate a video of a desired subject from any viewpoint, thereby facilitating the generation of powerful virtual viewpoint video.

[0011] According to one embodiment, the setting information includes information specifying which captured images from multiple positions are used to render the subject in each frame. Such setting information can be used to set the imaging device used to color the subject. With this configuration, for example, the color of the subject in the virtual viewpoint video can be determined according to the color of the subject as seen from a specific imaging device. In particular, when generating a virtual viewpoint video based on images captured by multiple imaging devices, a desired subject may be hidden by other subjects when viewed from one imaging device. Determining the color of the subject in the virtual viewpoint video using images captured by such imaging devices may result in a decrease in the reproducibility of the subject's color. On the other hand, by appropriately selecting the imaging device used to color the subject using such setting information, it becomes easier to more accurately reproduce the subject, thereby facilitating the generation of a more powerful virtual viewpoint video.

[0012] First, an information processing device according to an embodiment of the present disclosure will be described that generates control information used to generate a virtual viewpoint video including a subject viewed from a virtual viewpoint. In the following example, the virtual viewpoint video is generated based on captured images obtained by capturing images of the subject from multiple positions. Hereinafter, this control information will be referred to as virtual camera path data. The virtual camera path data can include information specifying the virtual viewpoint for each frame, i.e., time-series information. This control information can include external parameters such as the position of the virtual viewpoint and the line of sight direction from the virtual viewpoint, and may further include internal parameters such as the angle of view corresponding to the field of view from the virtual viewpoint.

[0013] The captured image used in this embodiment can be obtained by capturing an image of an imaging area where a subject exists from different directions using multiple imaging devices. The imaging area is, for example, an area defined by the plane and height of a stadium where a sport such as rugby or soccer is played. Multiple imaging devices can be installed in different positions and facing different directions so as to surround such an imaging area, and each imaging device captures images synchronously. Note that the imaging devices do not need to be installed along the entire perimeter of the imaging area, and may be installed only near a portion of the imaging area depending on, for example, installation space restrictions. The number of imaging devices is not limited. For example, if the imaging area is a rugby stadium, tens to hundreds of imaging devices may be installed around the stadium.

[0014] Furthermore, multiple imaging devices with different angles of view, such as a telephoto camera and a wide-angle camera, may be installed. For example, using a telephoto camera allows the subject to be captured at high resolution, thereby improving the resolution of the generated virtual viewpoint video. Furthermore, using a wide-angle camera widens the capturing range of a single camera, thereby reducing the number of cameras to be installed. The imaging devices are synchronized using a single time in the real world, and each frame of video captured by each imaging device is assigned with capturing time information.

[0015] Note that one imaging device may be composed of one camera or multiple cameras. Furthermore, the imaging device may include devices other than cameras. For example, the imaging device may include a distance measuring device using laser light or the like.

[0016] When generating a virtual viewpoint video, the state of each imaging device is referenced. The state of the imaging device may include the position, orientation (direction and imaging direction), focal length, optical center, and distortion of the resulting image of the imaging device. The position and orientation (direction and imaging direction) of the imaging device may be controlled by the imaging device itself or by a camera platform that controls the position and orientation of the imaging device. Hereinafter, data indicating the state of the imaging device will be referred to as the camera parameters of the imaging device, and these camera parameters may include data indicating a state controlled by another device such as a camera platform. Camera parameters related to the position and orientation (direction and imaging direction) of the imaging device are so-called extrinsic parameters. Parameters related to the focal length, image center, and image distortion of the imaging device are so-called intrinsic parameters. The position and orientation of the imaging device can be expressed, for example, in a coordinate system having one origin and three orthogonal axes (hereinafter referred to as the world coordinate system).

[0017] A virtual viewpoint image is also called a free viewpoint image. However, the virtual viewpoint image is not limited to an image from a viewpoint freely (arbitrarily) designated by the user; for example, an image from a viewpoint selected by the user from a plurality of candidate viewpoints is also included in the virtual viewpoint image. Furthermore, the virtual viewpoint may be designated by a user operation or automatically based on the results of image analysis, etc. Furthermore, although the present specification mainly describes the case where the virtual viewpoint image is a moving image, the virtual viewpoint image may also be a still image.

[0018] The virtual viewpoint information in this embodiment is information indicating the position and orientation of the virtual viewpoint. Specifically, the virtual viewpoint information includes a parameter indicating the three-dimensional position of the virtual viewpoint and a parameter indicating the line-of-sight direction of the virtual viewpoint in the pan, tilt, and roll directions. However, the virtual viewpoint information may also include a parameter indicating the size of the field of view (angle of view) of the virtual viewpoint.

[0019] The virtual viewpoint information may also be virtual camera path data that specifies a virtual viewpoint for each of a plurality of frames. That is, the virtual viewpoint information may have parameters corresponding to each of a plurality of frames that constitute the moving image of the virtual viewpoint video. Such virtual viewpoint information can indicate the position and orientation of the virtual viewpoint at each of a plurality of consecutive points in time.

[0020] A virtual viewpoint video is generated, for example, by the following method. First, an imaging device captures each imaging region from different directions, thereby obtaining multiple captured images. Next, from each of the multiple captured images, a foreground image extracted from a foreground region corresponding to a subject, such as a person or a ball, and a background image extracted from a background region other than the foreground region are obtained. The foreground image and the background image have texture information (such as color information). Then, a foreground model representing the three-dimensional shape of the subject and texture data for coloring the foreground model are generated based on the foreground image. The foreground model can be obtained using a shape estimation method such as the shape-from-silhouette method. A background model representing the three-dimensional shape of a background, such as a stadium, can be generated by performing three-dimensional measurements of, for example, a stadium or venue in advance. Furthermore, texture data used to color the background model can be generated based on the background image. The texture data is then mapped to the foreground model and the background model, and an image from the virtual viewpoint indicated by the virtual viewpoint information is rendered, thereby generating the virtual viewpoint video. Note that the method for generating the virtual viewpoint video is not limited to this method. For example, various methods can be used, such as a method of generating a virtual viewpoint video by projective transformation of a captured image without using a foreground model and a background model.

[0021] Note that a frame image of one frame of the virtual viewpoint video can be generated using multiple captured images captured synchronously at the same time. Then, by generating a frame image for each frame using an image captured at a time corresponding to each frame, a virtual viewpoint video composed of multiple frames can be generated.

[0022] A foreground image is an image extracted from a subject area (foreground area) among captured images obtained by imaging using an imaging device. A subject extracted as a foreground area is, for example, a dynamic object (moving body) that moves (whose position or shape may change) when images are captured from the same direction in chronological order. In the case of a sport, the subject may include, for example, a person such as a player or referee on the field where the sport is being played, and in the case of a ball game, the subject may include a ball in addition to a person. In a concert or entertainment event, a singer, musician, performer, or presenter is an example of the subject. In addition, if a background is registered in advance by, for example, specifying a background image, a still subject that does not exist in the background may also be extracted as a foreground area.

[0023] A background image is an image extracted from a region (background region) different from the foreground subject. For example, a background image may be an image obtained by removing the foreground subject from a captured image. A background is an imaged object that remains stationary or nearly stationary when images are captured from the same direction in chronological order. Examples of such imaged objects include a stage for a concert, a stadium where an event such as a sport is held, a structure such as a goal used in a ball game, or a field. Of course, while the background is a region different from the subject, objects different from the subject and background may also be present as imaged objects.

[0024] FIG. 1 is a diagram illustrating an example configuration of a virtual viewpoint image generation system according to an embodiment of the present disclosure. This system includes a data processing device 1, which is an information processing device according to an embodiment of the present disclosure, an imaging device 2, a shape estimation device 3, a storage device 4, an image generation device 5, a virtual camera operation device 6, and a data output device 7. Note that FIG. 1 illustrates one imaging device 2, while omitting other imaging devices. Furthermore, two or more of these devices may be integrated into a single device. For example, the data processing device 1 may have the functions of at least one of the image generation device 5 and the virtual camera operation device 6, which will be described below.

[0025] The data processing device 1 generates control information used to generate a virtual viewpoint video including a subject from a virtual viewpoint. In FIG. 1, the data processing device 1 is connected to a virtual camera operation device 6, a storage device 4, and a data output device 7. The data processing device 1 also acquires virtual viewpoint information from the virtual camera operation device 6 and acquires setting information related to video generation from the video generation device 5. The data processing device 1 then generates and outputs control information used to generate the virtual viewpoint video based on the acquired virtual viewpoint information and setting information related to video generation. The control information in this embodiment is virtual camera path data including virtual viewpoint information for each frame and setting information indicating a video generation method for each frame. The virtual camera path data output by the data processing device 1 is then output to the storage device 4 and the data output device 7.

[0026] The virtual camera operation device 6 generates virtual viewpoint information that specifies a virtual viewpoint in order to generate a virtual viewpoint video. The virtual viewpoint is specified by a user (operator) using an input device such as a joystick, a jog dial, a touch panel, a keyboard, or a mouse. The virtual viewpoint information can include information such as the position, attitude, and angle of view of the virtual viewpoint, as well as other information.

[0027] Here, the user can specify a virtual viewpoint while viewing a virtual viewpoint video or frame image generated according to the input virtual viewpoint information. To this end, the virtual camera operation device 6 transmits virtual viewpoint information to the video generation device 5. The virtual camera operation device 6 can also receive a virtual viewpoint video based on the transmitted virtual viewpoint information from the video generation device 5 and display this virtual viewpoint video. The user can consider the position of the virtual viewpoint while referring to the displayed virtual viewpoint video. Note that the method for specifying a virtual viewpoint is not limited to the above method. For example, the virtual camera operation device 6 can read a pre-created virtual camera path file and sequentially specify virtual viewpoints according to this virtual camera path file. The virtual camera operation device 6 may also receive user input specifying the movement of the virtual viewpoint and determine the position of the virtual viewpoint in each frame according to the specified movement. Alternatively, information indicating the movement of the virtual viewpoint may be used as virtual viewpoint information. The virtual camera operation device 6 may also recognize an object and automatically specify a virtual viewpoint based on the position of the recognized object, etc.

[0028] In addition to the virtual viewpoint information, the virtual camera control device 6 can also generate setting information related to image generation used to generate a virtual viewpoint video. Such setting information can also be specified by the user using an input device. For example, the virtual camera control device 6 can present, via a display, a user interface that includes the virtual viewpoint video generated by the image generation device 5 and accepts the user's specification of at least one of the virtual viewpoint information and the setting information. The user can also specify the virtual viewpoint information or the setting information while viewing the virtual viewpoint video or frame image generated according to the input virtual viewpoint information or the setting information. For this purpose, the virtual camera control device 6 can transmit the setting information to the image generation device 5. The virtual camera control device 6 can also receive a virtual viewpoint video based on the transmitted setting information from the image generation device 5 and display this virtual viewpoint video. The user can consider the setting information while referring to the displayed virtual viewpoint video. The virtual camera control device 6 may also automatically specify the setting information. For example, the virtual camera control device 6 can determine whether to display other subjects so that the target subject is not obscured by the other subjects.

[0029] As described above, the image generation device 5 can generate a virtual viewpoint video according to the virtual viewpoint information. The image generation device 5 may also generate a virtual viewpoint video according to setting information. At this time, the image generation device 5 acquires, from the storage device 4, object data used to generate the virtual viewpoint video. This object data may be, for example, an image captured by the imaging device 2, camera calibration information of the imaging device 2, point cloud model data, billboard model data, or mesh model data. As will be described later, the object specified by the virtual camera control device 6 may correspond to the object data acquired from the storage device 4. The image generation device 5 can also transmit setting information acquired from the virtual camera control device 6 to the data processing device 1. For example, the image generation device 5 can transmit a virtual viewpoint video to the virtual camera control device 6 for display, and can also transmit setting information used to generate the virtual viewpoint video displayed on the virtual camera control device 6 to the data processing device 1.

[0030] The storage device 4 stores the object data acquired by and generated by the shape estimation device 3. The storage device 4 may be configured, for example, as a semiconductor memory or a magnetic recording device. Each object data stored in the storage device 4 is associated with image capture time information of the object. The image capture time information can be associated with the object data, for example, by adding the image capture time information to the metadata of the object data. There are no particular limitations on the device that adds such image capture time information; for example, the image capture device 2 or the storage device 4 can add the image capture time information. The storage device 4 also outputs the object data in response to a request.

[0031] The shape estimation device 3 acquires captured images or foreground images from the imaging device 2, estimates the three-dimensional shape of the subject based on these images, and outputs three-dimensional model data indicating the three-dimensional shape of the subject. The three-dimensional model is represented by point cloud model data, billboard model data, mesh model data, or the like, as described above. The three-dimensional model may also include not only shape information but also color information of the subject. Note that if the image generation device 5 generates a virtual viewpoint image without using a foreground model or a background model, the virtual viewpoint image generation system does not need to include the shape estimation device 3.

[0032] The imaging device 2 has a unique identification number for distinguishing it from other imaging devices 2. The imaging device 2 may have other functions, such as a function to extract a foreground image from a captured image, and may also include hardware (such as a circuit or device) for realizing such functions.

[0033] The data output device 7 receives virtual camera path data from the data processing device 1 and object data corresponding to the virtual camera path data from the storage device 4, and stores or outputs the input object data. The format of the data when stored or output will be described later. Note that the data output device 7 does not need to output or store the object data, and the data output device 7 may store or output only the virtual camera path data as sequence data. Furthermore, the data output device 7 may store or output not only one pattern of virtual camera path data, but also multiple patterns of virtual camera path data.

[0034] Next, a description will be given of the configuration of the data processing device 1. The data processing device 1 includes a viewpoint information acquisition unit 101, a setting information acquisition unit 102, a camera path generation unit 103, and a camera path output unit 104.

[0035] The viewpoint information acquisition unit 101 performs a viewpoint acquisition operation to acquire information for specifying a virtual viewpoint in a frame of a virtual viewpoint video. The viewpoint information acquisition unit 101 can acquire information for specifying a virtual viewpoint in each frame. In this embodiment, the viewpoint information acquisition unit 101 acquires virtual viewpoint information specified by the virtual camera operation device 6. Note that the viewpoint information acquisition unit 101 may acquire the virtual viewpoint information for all frames collectively from the virtual camera operation device 6, or may continue to acquire the virtual viewpoint information for each frame that is sequentially specified by operating the virtual camera operation device 6 in real time.

[0036] The setting information acquisition unit 102 performs a setting acquisition operation to acquire setting information used to generate a virtual viewpoint video including a subject from a virtual viewpoint. In this embodiment, the setting information acquisition unit 102 can acquire information specifying a subject to be displayed in each frame of the virtual viewpoint video from among multiple subjects. The setting information acquisition unit 102 may also acquire information for specifying a captured image to be used to determine the color of the subject in the frame of the virtual viewpoint video from among multiple captured images obtained by capturing images of the subject from multiple positions. As described above, the setting information acquisition unit 102 can acquire setting information related to video generation used by the video generation device 5 from the video generation device 5. Note that, similar to the viewpoint information acquisition unit 101, the setting information acquisition unit 102 can collectively acquire setting information for all frames output by the virtual camera operation device 6. The setting information acquisition unit 102 may also continue to acquire virtual viewpoint information for each frame sequentially specified by real-time operation of the virtual camera operation device 6.

[0037] The camera path generation unit 103 outputs control information including virtual viewpoint information for identifying a virtual viewpoint for a frame of a virtual viewpoint video and setting information for identifying a subject to be displayed in the frame. The camera path generation unit 103 can generate control information including virtual viewpoint information indicating a virtual viewpoint for each frame and setting information related to video generation for each frame (e.g., information indicating a subject to be displayed or information indicating a captured image to be used for rendering). In this embodiment, the camera path generation unit 103 outputs this control information as virtual camera path data. The virtual camera path data can indicate an association between information indicating a virtual viewpoint specified for each frame and setting information. For example, the camera path generation unit 103 can generate virtual camera path data by adding control information acquired by the setting information acquisition unit 102 to the virtual viewpoint information acquired by the viewpoint information acquisition unit 101. The camera path generation unit 103 can output the generated control information to the camera path output unit 104.

[0038] The camera path output unit 104 outputs control information including virtual viewpoint information and setting information generated by the camera path generation unit 103. As described above, the camera path output unit 104 can output the control information as virtual camera path data. The camera path output unit 104 may add header information or the like to the virtual camera path data before outputting it. The camera path output unit 104 may output the virtual camera path data as a data file. Alternatively, the camera path output unit 104 may sequentially output multiple packet data representing the virtual camera path data. Furthermore, the virtual camera path data may be output on a frame-by-frame basis, or on a virtual camera path basis or on a group of a certain number of frames.

[0039] FIG. 2A shows an example of the format of sequence data output by the data output device 7, including virtual camera path data output by the camera path output unit 104. In FIG. 2A, the virtual camera path data constitutes sequence data indicating a virtual camera path in one virtual viewpoint video. One piece of sequence data may be generated for each video clip or each shot. Each piece of sequence data includes a sequence header, and the sequence header stores subject sequence data information that identifies the sequence data for the corresponding subject data. This information may be, for example, but is not limited to, a sequence header start code that can uniquely identify the subject data, information about the subject's shooting location and shooting date and time, or path information indicating the location of the subject data. The sequence header may also include information indicating that the sequence data includes virtual camera path data. This information may be, for example, information indicating the data set included in the sequence header or information indicating the presence or absence of virtual camera path data.

[0040] The sequence header then stores information about the entire sequence data. For example, the name of the virtual camera path sequence, information about the creator of the virtual camera path, information about the rights holder, the name of the event in which the subject was captured, the camera frame rate at the time of capture, and reference time information for the virtual camera path can be stored. In addition, the virtual viewpoint video size and background data information assumed when rendering the virtual viewpoint video can be stored. However, the information stored in the sequence header is not limited to these.

[0041] In the sequence data, each piece of virtual camera path data is stored in a unit called a data set. The number N of these data sets is stored in the sequence header. In this embodiment, the sequence data includes two types of data sets: virtual camera path data and subject data. Information for each data set is stored in the portion following the sequence header.

[0042] As information about one dataset in the sequence header, the dataset's identification ID is first stored. A unique ID is assigned to all datasets as the identification ID. Next, the dataset's type code is stored. In this embodiment, this type code indicates whether the dataset represents virtual camera path data or object data. The dataset's type code may be a 2-byte code shown in FIG. 2(B). However, the dataset's type and code are not limited to these. For example, the sequence data may include other types of data used when generating a virtual viewpoint video. Next, a pointer to this dataset is stored. However, other information for accessing the dataset itself may be stored instead of the pointer. For example, a file name in a file system established in the storage device 4 may be stored.

[0043] FIG. 3A shows an example of the configuration of a dataset of virtual camera path data. As described above, the control information in this embodiment may include setting information related to image generation for each frame. The setting information may also include information indicating which of multiple subjects will be displayed in each frame of the virtual viewpoint video. Here, the method for identifying the subjects to be displayed is not particularly limited. For example, the setting information may include display subject setting information, which indicates whether each of multiple subjects is to be displayed. The setting information may also include rendering area setting information, which indicates an area in three-dimensional space to be rendered. In this case, subjects located within this area are displayed in the frame image. On the other hand, the setting information may also include coloring camera setting information, which specifies which of multiple captured images will be used to render the subject in each frame. The setting information may also include other types of data used when generating the virtual viewpoint video. For example, the setting information may include additional information other than the display subject setting information, coloring camera setting information, and rendering area setting information. Examples of additional information include information specifying whether to add a shadow to the subject, information indicating the depth of the shadow, setting information related to the display of a virtual advertisement, and effect information. The setting information may include any of these types of information.

[0044] The virtual camera path data shown in Fig. 3A includes, as setting information, display subject setting information, coloring camera setting information, and rendering area setting information. The virtual camera path data shown in Fig. 3A also includes virtual viewpoint information.

[0045] A virtual camera path data header is stored at the beginning of the data set. At the beginning of this header, information indicating that the data set is a virtual camera path data set and the data size of the data set are stored. Next, the number of frames M of the stored virtual camera path data is described. Then, format information of the virtual camera path data is described. This format information indicates the format of the stored virtual camera path data, and can indicate, for example, whether various data related to the virtual camera path is stored by type or by frame. In the example of FIG. 3(A), each data is stored by type. That is, the virtual camera path data includes multiple data blocks, one data block containing virtual viewpoint information for each frame, and another data block containing setting information for each frame. Next, the number of data L is described in the virtual camera path data header. Subsequent virtual camera path data headers store information for each data included in the virtual camera path data.

[0046] The information for each data in the virtual camera path data header first stores a data type code. In this embodiment, the data type is represented by a virtual camera path data type code. The virtual camera path data type code may be, for example, a 2-byte code as shown in FIG. 3(B). However, the data type and code are not limited to these. For example, the code may be, for example, a code longer than 2 bytes or a code shorter than 2 bytes depending on the information to be described. Next, information for accessing the data body, such as a pointer, is stored. Then, format information corresponding to the data is described. For example, format information for virtual viewpoint information may include information indicating that camera extrinsic parameters representing the position and orientation of the virtual camera are expressed in quaternions.

[0047] After the virtual camera path data header, actual data (data body) of each piece of data related to the virtual camera path is written as virtual camera path data in accordance with the format written in the virtual camera path data header. A start code indicating the start of each piece of data is written at the beginning of that data. In the example of FIG. 3(A), the data body includes virtual viewpoint information, display subject setting information, coloring camera setting information, and rendering area setting information written in that order. Each piece of data includes information about the first to Mth frames. The virtual viewpoint information may include information specifying a virtual viewpoint for each frame, such as internal parameters and / or external parameters. In one embodiment, the virtual viewpoint information includes external parameters indicating the position of the virtual viewpoint and the line of sight from the virtual viewpoint. In another embodiment, the virtual viewpoint information includes internal parameters indicating the angle of view or focal length of the virtual viewpoint.

[0048] The display subject setting information indicates whether or not to display each of a plurality of subjects. Here, subjects to be displayed or not to be displayed can be specified using the model identifier of the target subject. The example in FIG. 3A shows an example in which a method for specifying subjects to be displayed is adopted, and model identifiers 001 and 003 of the subjects to be displayed are specified, and an example in which a method for specifying subjects not to be displayed is adopted, and model identifier 002 of the subject not to be displayed is specified. In both examples, the subject identified by model identifier 002 is not displayed in the virtual viewpoint video. A unique identifier that can uniquely identify a three-dimensional model in one frame can be used to specify the subject. Such an identifier may be defined for each frame, or the same identifier may be used for the same subject in a content data group.

[0049] The coloring camera setting information is information for specifying captured images used to determine the color of a subject in a frame of a virtual viewpoint image. This information can indicate the captured images used to render the subject in each frame of the virtual viewpoint video, and more specifically, the captured images referenced to determine the color of the subject in each frame image. This information can control the selection of imaging devices to be used to add color to a subject or its three-dimensional model. In the example of FIG. 3(A), imaging devices to be used or not used for coloring are specified. The imaging devices to be specified can be specified using a unique identifier that can uniquely identify the imaging device. Such imaging device identifiers can be determined when building an image generation system, and in this case, the same identifier is used for the same imaging device in a content data group. However, an identifier for each imaging device may also be specified for each frame. Since a large number of imaging devices, for example, tens to hundreds, are used to generate a virtual viewpoint video, using a method for specifying imaging devices not to be used for coloring may reduce the burden on the user.

[0050] The rendering area setting information is information that indicates an area in three-dimensional space for which a virtual viewpoint image is to be generated (or for which a rendering target is to be made). In each frame, a subject located within the area set here can be displayed. For example, a coordinate range can be specified, in which case a three-dimensional model not included in the specified coordinate range will not be rendered, i.e., will not be displayed in the virtual viewpoint image. The range can be specified, for example, using a coordinate system that defines the three-dimensional model, such as x, y, and z coordinates according to world coordinates. However, the method for setting the area is not particularly limited, and for example, the area may be set so that all subjects whose x and z coordinates are within a predetermined range are rendered.

[0051] These setting information may be described for each frame. That is, in one embodiment, virtual viewpoint information and setting information are recorded in the virtual camera path data for each frame. On the other hand, common setting information may be used for the entire content represented by the sequence data (e.g., for all frames) or for a part of the content (e.g., for multiple frames). That is, the virtual camera path data may record setting information that is commonly applied to multiple frames. Whether different setting information is described for each frame or common setting information is described for all frames can be determined for each type of data. For example, in the example of FIG. 3(A), the display subject setting information and coloring camera setting information are specified for each frame, and the rendering area setting information is used commonly for the entire content. On the other hand, display subject setting information or coloring camera setting information that is common to the entire content may be specified.

[0052] FIG. 4 shows an example of virtual camera path data when various data related to the virtual camera path are stored for each frame. In this way, the virtual camera path data may include multiple data blocks, and each data block may include virtual viewpoint information and setting information for one frame. When data is stored frame by frame, a frame data header is added to the beginning of each frame data. This frame data header can contain a code indicating the start of frame data, as well as information indicating the type and order of data stored as frame data.

[0053] The control of the virtual viewpoint video using the display subject setting information, coloring camera setting information, and rendering area setting information will be specifically described below.

[0054] Fig. 5 shows an example of control using display subject setting information. Fig. 5(A) shows three-dimensional models of subjects 501, 502, and 503 obtained by capturing an image of the space in which the subjects exist, and a virtual viewpoint 500 specified for generating a virtual viewpoint video. Here, when a virtual viewpoint video is generated according to the three-dimensional models of subjects 501 to 503, subjects 501 to 503 are displayed in the virtual viewpoint video as shown in Fig. 5(B). Here, when a virtual viewpoint video is generated by specifying the three-dimensional model of subject 501 as a non-display subject, subject 501 is not displayed in the virtual viewpoint video as shown in Fig. 5(C), and therefore subject 502 becomes visible.

[0055] FIG. 6 shows an example of control using coloring camera setting information. FIG. 6(A) shows a space in which subjects exist, showing imaging devices 510 and 511 and an obstacle 520. If a three-dimensional model of subjects 501 to 503 is generated using captured images obtained by these imaging devices and other imaging devices (not shown), and a virtual viewpoint video is generated from a virtual viewpoint 500, the virtual viewpoint video shown in FIG. 6(B) is expected to be obtained. In FIG. 6(B), a texture based on an image captured by imaging device 511, which is close to subject 503, is applied to subject 503. However, due to an unexpected obstacle 520, the color of subject 503 differs from that of the original subject. Here, if imaging device 511 is excluded from the imaging devices used for coloring using coloring camera control, the virtual viewpoint video shown in FIG. 6(C) is obtained. In FIG. 6(C), a texture based on an image captured by imaging device 510 is applied to subject 502, and subject 502 is displayed in the correct color.

[0056] There are various algorithms for selecting an imaging device to be used to colorize an object. For example, it is possible to select an imaging device close to the position of the virtual viewpoint, an imaging device close to the line of sight of the virtual viewpoint, or an imaging device close to the object. By using such coloring camera setting information, it is possible to limit the cameras that can be selected when rendering an object. This method can deal with obstacles such as those shown in FIG. 6(A), particularly obstacles located in positions where three-dimensional modeling is not possible. Furthermore, by using this method when generating a virtual viewpoint video in which a virtual viewpoint is rotated around the object at the same time, the sense of incongruity caused by switching cameras used to render the object can be alleviated.

[0057] FIG. 7 shows an example of control based on rendering area setting information. FIG. 7(A) shows three-dimensional models of subjects 501, 502, and 503 obtained by capturing images of the space in which the subjects exist, and a rendering area 530 specified for generating a virtual viewpoint video. The rendering area 530 shown in FIG. 7(A) is the entire space that can be specified by the system. In this case, as shown in FIG. 7(B), all three-dimensional models are displayed in the generated virtual viewpoint video. On the other hand, FIG. 7(C) shows an example in which a rendering area 540 approximately half the size of the rendering area 530 is specified. In this case, the three-dimensional model of subject 503 is outside the rendering area, so subject 503 is not displayed in the virtual viewpoint video, as shown in FIG. 7(D). This type of rendering area control achieves the same effect as the above-described subject display control. On the other hand, with this configuration, if only a portion of a three-dimensional model is within the area, that portion is displayed.

[0058] Thus, a data structure according to an embodiment, such as virtual camera path data, includes first data, such as virtual viewpoint information, for identifying a virtual viewpoint for a frame of a virtual viewpoint video. The data structure according to an embodiment also includes second data, such as display subject setting information or rendering area setting information, for identifying a subject to be displayed among a plurality of subjects for a frame of a virtual viewpoint video. Such a data structure is used by an information processing device that generates a virtual viewpoint video in a process of identifying a subject from a plurality of subjects using the second data. Such a data structure is also used in a process of generating a frame image corresponding to the virtual viewpoint identified by the first data and including the identified subject. Meanwhile, a data structure according to an embodiment also includes second data for identifying a captured image, from a plurality of captured images obtained by capturing images from a plurality of positions, to be used to determine the color of the subject in the frame of the virtual viewpoint video. An example of the second data is the coloring camera setting information described above. Such a data structure is also used by an information processing device that generates a virtual viewpoint video in a process of identifying a captured image from a plurality of captured images using the second data. Such a data structure is also used in a process of generating a frame image corresponding to the virtual viewpoint identified by the first data, based on the identified captured image.

[0059] The sequence data shown in FIG. 2A includes two data sets: virtual camera path data and subject data. However, the method for storing the virtual camera path data and subject data is not limited to this method. For example, the sequence data may include only virtual camera path data. In this case, the subject data may be stored in the storage device 4 separately from the virtual camera path data (or sequence data).

[0060] An example of an information processing method performed by the data processing device 1 as described above will be described with reference to the flowchart in Fig. 8. The processes of S801 to S804 are repeated frame by frame from the start of the virtual camera path until the input of the virtual camera path or frame by frame is completed. For example, the following process can be repeated from the frame where the user starts setting the virtual camera path to the frame where it is completed.

[0061] In S802, the viewpoint information acquisition unit 101 acquires virtual viewpoint information indicating a virtual viewpoint for the frame to be processed from the virtual camera operation device 6. In S803, the setting information acquisition unit 102 acquires the above setting information related to image generation for the frame to be processed from the image generation device 5.

[0062] In S805, the camera path generation unit 103 generates control information including virtual viewpoint information for each frame acquired by the viewpoint information acquisition unit 101 and setting information for each frame acquired by the setting information acquisition unit 102. For example, the camera path generation unit 103 can generate virtual camera path data by adding setting information to the virtual viewpoint information.

[0063] In S806, the camera path output unit 104 outputs the control information generated by the camera path generation unit 103. For example, the camera path output unit 104 can output the virtual camera path data after adding header information or the like to the virtual camera path data.

[0064] According to this embodiment, as described above, it is possible to generate control information including virtual viewpoint information indicating a virtual viewpoint for each frame and setting information related to image generation for each frame. In particular, since the virtual camera path data in this embodiment is provided with not only the virtual viewpoint information but also the setting information, as already described, the degree of freedom in control in generating a virtual viewpoint image is increased, making it easier to generate a desired virtual viewpoint image.

[0065] A method for generating a virtual viewpoint video in accordance with control information generated by the data processing device 1 will now be described. FIG. 9 shows a configuration example of a system including an image generation device, which is an information processing device according to an embodiment of the present disclosure. The image generation device 900 generates a virtual viewpoint video including a subject from a virtual viewpoint. This image generation device 900 can generate the virtual viewpoint video based on captured images obtained by capturing images of the subject from multiple positions. The configurations of the data processing device 1 and the storage device 4 are as already described.

[0066] The video generation device 900 includes a camera path acquisition unit 901 , a video setting unit 902 , a data management unit 903 , a video generation unit 904 , and a video output unit 905 .

[0067] The camera path acquisition unit 901 acquires control information including virtual viewpoint information for identifying a virtual viewpoint for a frame of a virtual viewpoint video and setting information related to video generation for each frame. The camera path acquisition unit 901 can acquire virtual camera path data including such control information output by the above-described data processing device 1. As described above, the setting information may be information for identifying a subject to be displayed in a frame of the virtual viewpoint video. Furthermore, the setting information may be information for identifying a captured image to be used to determine the color of the subject in the frame of the virtual viewpoint video, from among multiple captured images obtained by capturing images of the subject from multiple positions.

[0068] In FIG. 9 , the image generation device 900 is connected to the data processing device 1, but the image generation device 900 may also acquire virtual camera path data via a storage medium. For example, the virtual camera path data from the data processing device 1 may be input to the camera path acquisition unit 901 as a data file or as packet data. The camera path acquisition unit 901 may acquire the virtual camera path data for each frame, for each group of a certain number of frames, or for one or more data sets of virtual camera path data. When multiple data sets of virtual camera path data are acquired, the video output unit 905 can separately output virtual viewpoint videos corresponding to each virtual camera path data set. The data sets of each virtual camera path can be distinguished by an identification ID written in the header of each virtual camera path data set.

[0069] The video setting unit 902 acquires the above setting information used to generate a virtual viewpoint video from the virtual camera path data acquired by the camera path acquisition unit 901. Then, the video setting unit 902 sets a video generation method to be performed by the video generation unit 904 based on the acquired setting information.

[0070] The data management unit 903 acquires object data corresponding to the virtual camera path based on a request from the image generation unit 904. In FIG. 9, the image generation device 900 is connected to the storage device 4, and the data management unit 903 can acquire the object data from the storage device 4. The image generation device 900 may also acquire the object data via a storage medium. For example, the data management unit 903 can acquire the object data included in the sequence data output by the data output device 7. Furthermore, the image generation device 900 may store the same data as the object data stored in the storage device 4.

[0071] The subject data acquired by the data management unit 903 is selected based on the method by which the video generation unit 904 generates the virtual viewpoint video. For example, when using a video generation method based on a foreground model or a background model, the data management unit 903 can acquire point cloud model data or mesh model data of the foreground or background. The data management unit 903 can also acquire texture images corresponding to these models or captured images for generating textures, camera calibration data, etc. On the other hand, when using a video generation method that does not use a foreground model or a background model, the data management unit 903 can acquire captured images and camera calibration data, etc.

[0072] The video generation unit 904 generates a virtual viewpoint video by generating a frame image from a virtual viewpoint indicated by the virtual viewpoint information for each frame of the virtual viewpoint video based on the setting information. In this embodiment, the video generation unit 904 generates the virtual viewpoint video using the virtual viewpoint information acquired by the camera path acquisition unit 901 and the subject data acquired by the data management unit 903. Here, the video generation unit 904 generates the virtual viewpoint video according to the video generation method set by the video setting unit 902. As described above, the video generation unit 904 can generate frame images corresponding to the virtual viewpoint identified by the virtual viewpoint information, including the subject identified by the setting information, according to the setting information for identifying the subject to be displayed in the frame. Furthermore, the video generation unit 904 can generate frame images including the subject corresponding to the virtual viewpoint identified by the virtual viewpoint information for the frame of the virtual viewpoint video, based on the captured image identified by the setting information. The video generation method based on the setting information is as described with reference to FIGS. 5 to 7.

[0073] The video output unit 905 acquires the virtual viewpoint video from the video generation unit 904 and outputs the virtual viewpoint video to a display device such as a display. Note that the video output unit 905 may output the virtual viewpoint video acquired from the video generation unit 904 as a data file or packet data.

[0074] An information processing method performed by the information processing device according to this embodiment will be described with reference to the flowchart in Fig. 10. The processes of S1001 to S1008 are repeated for each frame from the start to the end of the virtual camera path.

[0075] In S1002, the camera path acquisition unit 901 acquires control information including virtual viewpoint information indicating a virtual viewpoint for a frame to be processed and the above-described setting information related to video generation. For example, the camera path acquisition unit 901 can acquire information about the frame to be processed that is included in virtual camera path data acquired from the data processing device 1. The setting information has already been described.

[0076] In S1003, the video setting unit 902 acquires setting information from the camera path acquisition unit 901 and sets the video generation unit 904 to operate in accordance with the setting information. In S1004, the video generation unit 904 acquires virtual viewpoint information from the camera path acquisition unit 901. In S1005, the data management unit 903 acquires subject data from the storage device 4 in accordance with a request from the video generation unit 904.

[0077] In S1006, the video generation unit 904 generates a frame image from the virtual viewpoint indicated by the virtual viewpoint information for the frame to be processed in accordance with the setting information. The video generation unit 904 can generate a virtual viewpoint video based on the subject data acquired in S1005 and the virtual viewpoint information acquired in S1004 in accordance with the setting specified in S1003. The method for generating an image in accordance with the setting information has already been described. In S1007, the video output unit 905 outputs the frame image of the virtual viewpoint video generated in S1006 via a display device such as a display. The video output unit 905 may output the frame image of the virtual viewpoint video as a data file or packet data.

[0078] According to the above-described embodiment, a virtual viewpoint video can be generated based on control information including virtual viewpoint information indicating a virtual viewpoint for each frame and setting information related to video generation for each frame. By using such setting information, the degree of freedom in control of the generation of the virtual viewpoint video is increased, making it easier to output a powerful virtual viewpoint video.

[0079] Furthermore, since setting information can be recorded in control information such as the virtual camera path data described above, it becomes easy for a user to create control information, view virtual viewpoint video according to this control information, and then modify the virtual viewpoint information or setting information. Furthermore, by transmitting such control information created by a video producer along with subject data to a viewer, the viewer can view virtual viewpoint video according to the control information and recommended by the video producer. On the other hand, the viewer can also choose whether to view virtual viewpoint video according to the control information or view virtual viewpoint video from a desired viewpoint without using the control information.

[0080] Each information processing device, such as the data processing device 1 and the image generation device 900, can be realized by a computer including a processor and a memory. However, some or all of the functions of each information processing device may be realized by dedicated hardware. Furthermore, an image processing device according to an embodiment of the present disclosure may be configured by a plurality of information processing devices connected via a network, for example.

[0081] 11 is a block diagram showing an example of the hardware configuration of such a computer. A CPU 1101 controls the entire computer using computer programs or data stored in a RAM 1102 or a ROM 1103, and executes the processes described above as being performed by the information processing device according to the above embodiment. That is, the CPU 1101 can function as each processing unit shown in FIGS. 1 and 9.

[0082] The RAM 1102 is a memory having an area for temporarily storing computer programs or data loaded from an external storage device 1106, and data acquired from the outside via an I / F (interface) 1107. The RAM 1102 also has a work area used by the CPU 1101 when executing various processes. That is, the RAM 1102 can provide, for example, a frame memory and various other areas.

[0083] The ROM 1103 is a memory that stores computer setting data, a boot program, etc. The operation unit 1104 is an input device such as a keyboard or a mouse, which can be operated by a computer user to input various instructions to the CPU 1101. The output unit 1105 is an output device that outputs the results of processing by the CPU 1101, and is, for example, a display device such as a liquid crystal display.

[0084] The external storage device 1106 is a large-capacity information storage device such as a hard disk drive. The external storage device 1106 can store an OS (operating system) and a computer program for causing the CPU 1101 to realize the functions of each unit shown in Fig. 1. The external storage device 1106 may also store image data captured by the imaging device 2 or virtual viewpoint video data generated by the video generation device 5.

[0085] Computer programs or data stored in the external storage device 1106 are loaded into the RAM 1102 as appropriate under the control of the CPU 1101 and become the subject of processing by the CPU 1101. The I / F 1107 can be connected to a network such as a LAN or the Internet, a projection device, or other devices such as a display device, and the computer can obtain and send various information via this I / F 1107. 1108 is a bus that connects the above-mentioned components.

[0086] (Other Examples) The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]

[0087] 101: viewpoint information acquisition unit, 102: setting information acquisition unit, 103: camera path generation unit, 104: camera path output unit, 901: camera path acquisition unit, 902: video setting unit, 903: data management unit, 904: video generation unit, 905: video output unit

Claims

1. a viewpoint acquisition means for acquiring information specifying a virtual viewpoint in a frame of a virtual viewpoint video; a setting acquisition means for acquiring information specifying a subject to be displayed in the frame of the virtual viewpoint video among a plurality of subjects; an output means for outputting control information including virtual viewpoint information for specifying the virtual viewpoint for the frame of the virtual viewpoint video and setting information for specifying the subject displayed in the frame; and An information processing device characterized in that the output means outputs the control information as virtual camera path data, and the virtual camera path data records the virtual viewpoint information and the setting information for each frame.

2. The information processing apparatus according to claim 1 , wherein the setting information is information indicating whether or not each of a plurality of subjects is to be displayed.

3. The information processing device according to claim 1 , wherein the setting information is information indicating an area in a three-dimensional space for which a virtual viewpoint video is to be generated, and a subject located within the area is displayed.

4. a viewpoint acquisition means for acquiring information for specifying a virtual viewpoint in a frame of a virtual viewpoint video; a setting acquisition means for acquiring information for specifying a captured image to be used to determine the color of the subject in the frame of the virtual viewpoint video, from among a plurality of captured images obtained by capturing images of the subject from a plurality of positions; an output means for outputting control information including virtual viewpoint information for specifying the virtual viewpoint for the frame of the virtual viewpoint video and setting information for specifying the captured image used to determine the color of the subject in the frame; and An information processing device characterized in that the output means outputs the control information as virtual camera path data, and the virtual camera path data records the virtual viewpoint information and the setting information for each frame.

5. The information processing apparatus according to claim 1 , wherein the virtual viewpoint information includes external parameters indicating a position of the virtual viewpoint and a line-of-sight direction from the virtual viewpoint.

6. The information processing device according to claim 1 , wherein the virtual viewpoint information includes an internal parameter indicating an angle of view or a focal length of the virtual viewpoint.

7. 7. The information processing device according to claim 1, wherein the virtual camera path data includes a plurality of data blocks, and each data block includes the virtual viewpoint information and the setting information for one frame.

8. 7. The information processing device according to claim 1, wherein the virtual camera path data includes a plurality of data blocks, one data block including the virtual viewpoint information for each frame, and another data block including the setting information for each frame.

9. The information processing device according to claim 1 , wherein the output means outputs the virtual camera path data as a data file, or sequentially outputs a plurality of packet data representing the virtual camera path data.

10. a generating means for generating the virtual viewpoint video based on the virtual viewpoint information and the setting information; a presentation means for presenting a user interface including the generated virtual viewpoint video and for accepting designation by a user of at least one of the virtual viewpoint information and the setting information; 10. The information processing apparatus according to claim 1, further comprising:

11. an acquisition means for acquiring control information including virtual viewpoint information for specifying a virtual viewpoint for a frame of a virtual viewpoint video and setting information for specifying a subject to be displayed in the frame; a generation means for generating a frame image including the subject specified by the setting information and corresponding to the virtual viewpoint specified by the virtual viewpoint information; and An information processing device characterized in that the acquisition means acquires the control information as virtual camera path data, and the virtual camera path data records the virtual viewpoint information and the setting information for each frame.

12. an acquisition means for acquiring control information including virtual viewpoint information for specifying a virtual viewpoint for a frame of a virtual viewpoint video, and setting information for specifying a captured image to be used for determining the color of the subject in the frame from among a plurality of captured images obtained by capturing images of the subject from a plurality of positions; a generation means for generating a frame image including the subject, corresponding to the virtual viewpoint specified by the virtual viewpoint information, for the frame of the virtual viewpoint video, based on a captured image specified by the setting information; and An information processing device characterized in that the acquisition means acquires the control information as virtual camera path data, and the virtual camera path data records the virtual viewpoint information and the setting information for each frame.

13. An information processing device as described in claim 11 or 12, characterized in that the generation means generates the virtual viewpoint image using subject data representing the subject stored in a storage device separately from the virtual camera path data.

14. An information processing method performed by an information processing device, acquiring information specifying a virtual viewpoint in a frame of the virtual viewpoint video; acquiring information specifying a subject to be displayed in the frame of the virtual viewpoint video among a plurality of subjects; outputting control information including virtual viewpoint information for specifying the virtual viewpoint for the frame of the virtual viewpoint video and setting information for specifying the subject to be displayed in the frame; and An information processing method characterized in that in the outputting step, the control information is output as virtual camera path data, and the virtual camera path data records the virtual viewpoint information and the setting information for each frame.

15. An information processing method performed by an information processing device, acquiring information for specifying a virtual viewpoint in a frame of the virtual viewpoint video; acquiring information for specifying a captured image to be used for determining a color of the subject in the frame of the virtual viewpoint video, from among a plurality of captured images obtained by capturing images of the subject from a plurality of positions; outputting control information including virtual viewpoint information for specifying the virtual viewpoint for the frame of the virtual viewpoint video and setting information for specifying the captured image used to determine the color of the subject in the frame; and An information processing method characterized in that in the outputting step, the control information is output as virtual camera path data, and the virtual camera path data records the virtual viewpoint information and the setting information for each frame.

16. An information processing method performed by an information processing device, acquiring control information including virtual viewpoint information for specifying a virtual viewpoint for a frame of a virtual viewpoint video and setting information for specifying a subject to be displayed in the frame; generating a frame image that includes the subject identified by the setting information and corresponds to the virtual viewpoint identified by the virtual viewpoint information; and An information processing method characterized in that in the acquiring step, the control information is acquired as virtual camera path data, and the virtual camera path data records the virtual viewpoint information and the setting information for each frame.

17. An information processing method performed by an information processing device, acquiring control information including virtual viewpoint information for specifying a virtual viewpoint for a frame of a virtual viewpoint video, and setting information for specifying a captured image to be used for determining the color of the subject in the frame from among a plurality of captured images obtained by capturing images of the subject from a plurality of positions; generating a frame image including the subject, corresponding to the virtual viewpoint specified by the virtual viewpoint information, for the frame of the virtual viewpoint video, based on the captured image specified by the setting information; and An information processing method characterized in that in the acquiring step, the control information is acquired as virtual camera path data, and the virtual camera path data records the virtual viewpoint information and the setting information for each frame.

18. A program for causing a computer to function as the information processing device according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Controller, control method, and program

    JP2017212592A

  • Information processing apparatus, information processing method, and program

    JP2018085571A

  • Image processing system, control method therefor, and program

    JP2019079468A

  • Image processing apparatus, image processing method, and program

    JP2019125929A

  • Image generation device, image generation method, image generation system, and program

    JP2020135290A