Information processing apparatus, information processing method, and program
The information processing device addresses the issue of unsupported data formats by generating compatible material data for virtual viewpoint images, ensuring proper processing and output.
Patent Information
- Application Number
- JP2025196828
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-01-23
AI Technical Summary
Existing methods for generating virtual viewpoint images face issues when the data output to a device is not supported by the device, leading to improper processing.
An information processing device that acquires and outputs material data for generating virtual viewpoint images, using multiple imaging devices to capture images and generate data formats compatible with the capabilities of the output device.
Enables the output of material data that can be processed by the device, ensuring proper generation of virtual viewpoint images.
Smart Images

Figure 2026012567000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technique for generating a virtual viewpoint image. [Background technology]
[0002] In recent years, a technology has been attracting attention in which multiple imaging devices are arranged around an imaging area to capture an object, and then multiple captured images acquired by the imaging devices are used to generate an image (virtual viewpoint image) viewed from a specified viewpoint (virtual viewpoint). This technology allows users to view highlight scenes of soccer or basketball games from various angles, providing a more realistic experience than conventional images. Patent Document 1 describes a system that generates point cloud data consisting of multiple points indicating three-dimensional positions as three-dimensional shape data representing the shape of an object. The system described in Patent Document 1 also generates an image (virtual viewpoint image) viewed from an arbitrary viewpoint by performing a rendering process using the generated point cloud data. Patent Document 2 describes a technology that reconstructs a wide-area three-dimensional model of an object, including not only its frontal region but also its lateral region, from information acquired by an RGB-D camera. Patent Document 2 describes a technology that extracts an object region based on depth data from a depth map, and generates a three-dimensional surface model of the object based on the depth data. Patent Document 2 also describes a technology that generates a virtual viewpoint image by performing a rendering process using a three-dimensional mesh model generated using the three-dimensional surface model. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2018 / 147329 [Patent Document 2] Japanese Patent Application Laid-Open No. 2016-71645 Summary of the Invention [Problem to be solved by the invention]
[0004] As described above, there are multiple methods for generating virtual viewpoint images, and each method uses different data. Therefore, if the data output to a device that executes the process related to generating the virtual viewpoint image is data that the device does not support, the process may not be executed properly.
[0005] The present disclosure has been made in view of the above-mentioned problems, and has as its object to enable output of material data related to the generation of a virtual viewpoint image that can be processed by a device to which the data is output. [Means for solving the problem]
[0006] The information processing device according to the present disclosure is characterized by having an acquisition means for acquiring a plurality of material data used to generate a virtual viewpoint image based on a plurality of captured images obtained by a plurality of imaging devices capturing images of a subject, the material data including first material data and second material data different from the first material data, and an output means for outputting to the other device material data determined from the plurality of material data acquired by the acquisition means based on information for identifying a format of the material data that can be processed by the other device to which the material data is output. [Effects of the Invention]
[0007] According to the present disclosure, it is possible to output material data related to the generation of a virtual viewpoint image that can be processed by a device to which the data is output. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a virtual viewpoint image generating system including an information processing device according to a first embodiment. [Figure 2] 10 is an example of a data structure of material data stored in a storage unit. [Figure 3] 10 is an example of a data structure of material data stored in a storage unit. [Figure 4] 10 is an example of a data structure of material data stored in a storage unit. [Figure 5] 10 is an example of a data structure of material data stored in a storage unit. [Figure 6] 10 is an example of a data structure of material data stored in a storage unit. [Figure 7] 10 is an example of a data structure of material data stored in a storage unit. [Figure 8] 4 is a flowchart illustrating an example of processing performed by the virtual viewpoint image generating system according to the first embodiment. [Figure 9] 4 is a flowchart illustrating an example of processing performed by the virtual viewpoint image generating system according to the first embodiment. [Figure 10] FIG. 2 is a diagram illustrating an example of communication performed in the virtual viewpoint image generation system. [Figure 11] 10 is an example of a data structure of material data stored in a storage unit. [Figure 12] 10 is an example of a data structure of material data stored in a storage unit. [Figure 13] 10 is an example of a data structure of material data stored in a storage unit. [Figure 14] 10 is an example of a data structure of material data stored in a storage unit. [Figure 15] 10 is an example of a data structure of material data stored in a storage unit. [Figure 16] FIG. 1 is a diagram illustrating an example of the configuration of a virtual viewpoint image generating system including an information processing device according to a first embodiment. [Figure 17] FIG. 2 is a diagram illustrating an example of communication performed in the virtual viewpoint image generation system. [Figure 18] FIG. 10 is a diagram illustrating an example of the configuration of a virtual viewpoint image generating system including an information processing device according to a second embodiment. [Figure 19] 10 is a flowchart illustrating an example of processing performed by the virtual viewpoint image generating system according to the second embodiment. [Figure 20]FIG. 10 is a diagram illustrating an example of the configuration of a virtual viewpoint image generating system including an information processing device according to a third embodiment. [Figure 21] 10 is a flowchart illustrating an example of processing performed by the virtual viewpoint image generating system according to the third embodiment. [Figure 22] 10 is an example of a data structure of material data stored in a storage unit. [Figure 23] FIG. 2 is a block diagram illustrating an example of a hardware configuration of an information processing device. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the components described in the following embodiments are examples of embodiments, and the present disclosure is not limited to them.
[0010] (First embodiment) In this embodiment, we will describe an information processing device that outputs data (hereinafter referred to as material data) used to generate a virtual viewpoint image based on multiple captured images obtained by multiple imaging devices capturing images of a subject to a display device.
[0011] Fig. 23 is a block diagram showing an example of the hardware configuration of a computer applicable to the information processing device in this embodiment. The information processing device has a CPU 1601, RAM 1602, ROM 1603, an operation unit 1604, an output unit 1605, an external storage device 1606, an I / F 1607, and a bus 1608. Note that the hardware configuration shown in Fig. 23 is applicable to any device included in a virtual viewpoint image generation system described later.
[0012] The CPU 1601 controls the entire computer using computer programs and data stored in the RAM 1602 and ROM 1603. The RAM 1602 has an area for temporarily storing computer programs and data loaded from the external storage device 806, data acquired from the outside via an I / F (interface) 1607, and the like. The RAM 1602 also has a work area used by the CPU 1601 when executing various processes. That is, the RAM 1602 can be allocated as a frame memory, for example, or can provide various other areas as needed.
[0013] The ROM 1603 stores setting data for the device, a boot program, etc. The operation unit 1604 is made up of input devices such as a keyboard, a mouse, and a joystick, and can input various instructions to the CPU 1601 based on user operations using these input devices. The output unit 1605 outputs the results of processing by the CPU 1601. The output unit 1605 is made up of, for example, a liquid crystal display, etc.
[0014] The external storage device 1606 can be an information storage device typified by a hard disk drive. The external storage device 1606 stores an OS (operating system) and computer programs for causing the CPU 1601 to realize the functions of each processing unit of the device in this embodiment. The computer programs and data stored in the external storage device 1606 are loaded into the RAM 1602 as appropriate under the control of the CPU 1601, and become the subject of processing by the CPU 1601. Furthermore, the external storage device 1606 may also store image data to be processed, and various types of information used in the processing.
[0015] The I / F 1607 can be connected to a network such as a LAN or the Internet, or to other devices such as a projector or display device, and the computer can acquire and transmit various information via this I / F 1607. For example, when the information processing device is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 1607. If the information processing device has a function for wireless communication with an external device, the communication I / F 1607 is provided with an antenna. A bus 1608 is a bus that connects the above-mentioned components.
[0016] At least one of the above-described components may be connected to the outside of the information processing device as another device. This also applies to any device included in the virtual viewpoint image generation system described later.
[0017] <Configuration of the virtual viewpoint image generation system> Next, the configuration of the virtual viewpoint image generation system in this embodiment will be described. Note that the virtual viewpoint image in this embodiment is also called a free viewpoint image, but is not limited to an image corresponding to a viewpoint freely (arbitrarily) designated by the user, and for example, an image corresponding to a viewpoint selected by the user from a plurality of candidates is also included in the virtual viewpoint image. In this embodiment, the description will focus on a case where the virtual viewpoint is designated by a user operation, but the virtual viewpoint may also be designated automatically based on the results of image analysis, etc. Furthermore, in this embodiment, the description will focus on a case where the virtual viewpoint image is a moving image, but the virtual viewpoint image may also be a still image.
[0018] 1 is a diagram illustrating the configuration of a virtual viewpoint image generation system 1. The virtual viewpoint image generation system 1 captures images of a subject using multiple imaging devices installed in a facility such as a sports field (stadium) or concert hall, and generates data (material data) for generating a virtual viewpoint image based on the multiple captured images. The virtual viewpoint image generation system 1 includes cameras 101a-z, an input unit 102, a storage unit 103, a first model generation unit 104, a second model generation unit 105, a distance image generation unit 106, a mesh model generation unit 107, an information processing device 100, and terminals 112a-d. The functional configuration of the information processing device 100 will be described later.
[0019] The cameras 101a to 101z are arranged to surround the subject and capture images in synchronization. However, the number of cameras is not limited to that shown in FIG. 1. The cameras 101a to z are connected via a network and to the input unit 102. The cameras 101a to z capture images in synchronization. That is, the frames of images captured by the cameras 101a to z are acquired at the same time. The acquired captured images are assigned time information relating to the time of capture and a frame number, and are transmitted to the input unit 102. The time information can use any format. Each camera is assigned a camera ID, and the camera ID information is assigned to the captured images acquired by each camera. In the following description, unless there is a particular reason to distinguish between them, the cameras 101a to z will be simply referred to as camera 101.
[0020] The input unit 102 receives image data obtained by imaging performed by the cameras 101a to 101z, and outputs the image data to the storage unit 103. The storage unit 103 is a storage unit that temporarily stores the input image data.
[0021] The first model generation unit 104, the second model generation unit 105, the distance image generation unit 106, and the mesh model generation unit 107 are processing units that generate material data used to generate a virtual viewpoint image. Here, the material data in this embodiment refers to data for generating a virtual viewpoint image, and refers to data generated based on a captured image. The material data includes, for example, data of a foreground image and a background image extracted from a captured image, a model representing the shape of a subject in three-dimensional space (hereinafter also referred to as three-dimensional shape data), and texture data for coloring the three-dimensional shape data. Also, different material data may be required depending on the method for generating the virtual viewpoint image. Details will be described later.
[0022] The first model generation unit 104 generates point cloud data representing the three-dimensional shape of the subject and a foreground image based on the input image data as material data. Here, the foreground image is an image obtained by extracting an area corresponding to an object (e.g., a player or a ball) from a captured image. For example, the first model generation unit 104 captures an image in which the subject is not captured as a reference image, and extracts the area corresponding to the subject using the difference between the reference image and the input image data (captured image). Furthermore, based on the subject area extraction result, the first model generation unit 104 generates a silhouette image in which the pixel value of the area corresponding to the subject is set to 1 and the pixel value of other areas is set to 0.
[0023] Furthermore, the first model generating unit 104 uses the generated silhouette image to generate point cloud data using a shape-from-silhouette method. The generated point cloud data is data that represents the three-dimensional shape of the subject using a collection of multiple points each having coordinate information in three-dimensional space. The first model generation unit 104 also calculates a circumscribing rectangle for the extracted subject region, cuts out a region in the captured image that corresponds to the circumscribing rectangle, and generates this as a foreground image. The foreground image is used as texture data for coloring the three-dimensional shape data when generating a virtual viewpoint image. Note that the method for generating the point cloud data and foreground image is not limited to the above-described method, and any generation method can be used.
[0024] The second model generation unit 105 generates, as material data, data including a point cloud representing the three-dimensional shape of the subject and information representing the color of the subject (hereinafter, colored point cloud data). The colored point cloud data can be generated using the point cloud data generated by the first model generation unit 104 and the captured images or foreground images of each camera 101 acquired from the storage unit 103. For each point included in the point cloud data, the second model generation unit 105 identifies the color (pixel value) of the captured image acquired by the camera 101 that captures the three-dimensional position corresponding to the point. The second model generation unit 105 can generate colored point cloud data, which is point cloud data having color information, by associating the identified color with the point. Note that the second model generation unit 105 may generate the colored point cloud data using image data acquired from the storage unit 103 without using the point cloud data generated by the first model generation unit 104.
[0025] The distance image generation unit 106 generates distance images representing the distance from each camera 101 to the subject as raw data. The distance image generation unit 106 generates the distance images by, for example, calculating the parallax between multiple captured images using a stereo matching method, and determining pixel values of the distance image from the parallax data. Note that the method for generating the distance images is not limited to this. For example, the distance image generation unit 106 may generate the distance images by using the point cloud data generated by the first model generation unit 104 and calculating the distance from the three-dimensional position of each point in the point cloud data to each camera 101. Alternatively, the distance image generation unit 106 may be configured to separately acquire distance images using a distance camera with an infrared sensor or the like.
[0026] The mesh model generation unit 107 generates mesh data representing the three-dimensional shape of the subject as material data based on the image data acquired from the storage unit 103. The mesh model generation unit 107 generates mesh data, which is a collection of multiple polygons, using, for example, the method described in Patent Document 2. When using this method, the distance image data generated by the distance image generation unit 106 is used as a depth map. Note that the method for generating mesh data is not limited to this. For example, the mesh model generation unit 107 may generate mesh data by converting point cloud data generated by the first model generation unit 104 and colored point cloud data generated by the second model generation unit 105.
[0027] The first model generation unit 104, the second model generation unit 105, the distance image generation unit 106, and the mesh model generation unit 107 may be included in a single device, or may be connected as separate devices. Alternatively, the system may be configured with a plurality of devices each including any combination of these generation units. For example, a first generation device including the first model generation unit 104 and the second model generation unit 105 may be connected to a second generation device including the distance image generation unit 106 and the mesh model generation unit 107. Furthermore, at least one of these generation units, or a processing unit that executes part of the processing performed by these generation units, may be included in the camera 101 or another device, such as the information processing device 100 described below. For example, the silhouette image generation and / or foreground image generation processing performed by the first model generation unit may be performed by the camera 101. In this case, the camera 101 assigns various information, such as time information, frame numbers, and camera IDs, to the silhouette image and / or foreground image generated based on the captured image and transmits the resulting image to the input unit 102.
[0028] The terminals 112a to 112d are display devices that acquire material data from the information processing device 100 (described later) and generate and display virtual viewpoint images based on the material data. The terminals 112a to 112d may be, for example, PCs, smartphones, tablets, etc. In the following description, the terminals 112a to 112d will be simply referred to as terminals 112 unless there is a particular reason to distinguish them.
[0029] Next, the functional configuration of the information processing device 100 included in the virtual viewpoint image generation system 1 will be described with reference to Fig. 1. The information processing device 100 includes a management unit 108, a storage unit 109, a transmission / reception unit 110, and a selection unit 111.
[0030] Management unit 108 acquires multiple pieces of material data generated by first model generation unit 104, second model generation unit 105, distance image generation unit 106, and mesh model generation unit 107, and stores them in storage unit 108. When storing the data, management is performed so that the data can be read and written in association with time information, frame numbers, etc., by generating a data access table for reading each piece of data. Data is also output based on instructions from selection unit 111, which will be described later.
[0031] The storage unit 109 stores the input material data. The storage unit 109 corresponds to the external storage device 1606 in Fig. 23 and is composed of a semiconductor memory, a magnetic recording device, or the like. The format of the material data when it is stored will be described later. Data is written and read based on instructions from the management unit 108, and the written data is output to the transmission / reception unit 110 in accordance with a read instruction.
[0032] The transmitting / receiving unit 110 communicates with a terminal 112 (described later) and transmits and receives requests from the terminal and data. The selecting unit 111 selects material data to be transmitted to a display device connected to the information processing device 100 from among a plurality of material data. The operation will be described later. The selecting unit 111 transmits an output instruction for the selected material data to the managing unit 108.
[0033] The terminals 112a to 112d acquire material data from the information processing device 100. The terminals 112a to 112d also receive a user's operation to specify a virtual viewpoint, and generate a virtual viewpoint image corresponding to the specified virtual viewpoint based on the material data. The terminals 112a to 112d also display the generated virtual viewpoint image. In this embodiment, the terminal 112a generates the virtual viewpoint image using a foreground image and point cloud data. The terminal 112b generates the virtual viewpoint image using colored point cloud data. The terminal 112c generates the virtual viewpoint image using a foreground image and a distance image. The terminal 112d generates the virtual viewpoint image using mesh data.
[0034] As described above, the material data used to generate a virtual viewpoint image may differ depending on the type of terminal. Factors that may cause the difference in material data used include, for example, the virtual viewpoint image generation method employed by the terminal and the type of software used to generate the virtual viewpoint image. Another factor may be the difference in the format of material data that the terminal can handle. Furthermore, differences in the display and processing capabilities of the terminal may also be a factor, such as the ability to display virtual viewpoint images based on mesh data but not point cloud data. These factors may result in the terminal being unable to properly display a virtual viewpoint image even if the terminal is provided with material data corresponding to a virtual viewpoint image that the terminal cannot display. To solve this problem, the information processing device 100 in this embodiment determines and outputs the material data to be transmitted from among multiple pieces of material data based on the type of virtual viewpoint image that the display device terminal can display. Details of this processing will be described later.
[0035] <Example of material data format> FIG. 2 shows an example of the format of the three-dimensional shape data stored in the storage unit 109. Material data is stored as a sequence corresponding to a single unit, as shown in FIG. 2(a). A sequence can be generated, for example, for a shooting period, a shot of a sport, or even for each event or cut that occurred during the sport. The management unit 108 manages data on a sequence-by-sequence basis. Each sequence includes a sequence header. As shown in FIG. 2(b), the sequence header stores a sequence header start code indicating the start of the sequence. Next, information about the entire sequence is stored. For example, the name of the sequence, the shooting location, the date and time when shooting started, time information indicating the time, frame rate, and image size are stored. The sequence name includes information for identifying the sequence, such as the name of the sport or event.
[0036] Each sequence stores three-dimensional shape data in units called data sets. The sequence header contains the number of data sets (M). Information is then stored for each data set. The data set information begins with an ID for identifying the data set. The ID is assigned by the storage unit 109 or a unique ID across all data sets. Next, the data set type is stored. In this embodiment, data sets include point cloud data, foreground image data, colored point cloud data, range image data, and mesh data. Each data set is represented as a data set class code. The data set class code is represented as a two-byte code shown in FIG. 2(e). However, the data type and code are not limited to this. Data representing other three-dimensional shape data may also be used. Next, a pointer to the data set is stored. However, any information for accessing each data set will suffice, and is not limited to a pointer. For example, a file system may be constructed in the storage unit 109, and the file name may be used.
[0037] Next, the format of each data set included in the sequence will be described. In this embodiment, the types of data sets will be described as point cloud data, foreground image, colored point cloud data, range image data, and mesh data.
[0038] Figure 3 shows an example of the configuration of a foreground image dataset. For the sake of explanation, it is assumed that the foreground image dataset is stored frame by frame, but this is not limiting. As shown in Figure 3(a), a foreground image data header is stored at the beginning of the dataset. The header stores information such as the fact that this dataset is a foreground image dataset and the number of frames.
[0039] The information contained in each frame is shown in Figure 3(b). Each frame stores time information (time information) indicating the time of the first frame of the foreground image data and the data size of that frame. The data size is used to reference the data of the next frame and may be stored together in the header. Next, the number of objects (P) used to generate the virtual viewpoint image at the time indicated by the time information is stored. The number of cameras (C) used for capturing the image at that time is stored. Next, the camera ID of the camera used (Cth Camera) is stored. Next, image data of the foreground image in which the object is captured (1st Foreground Image of Cth Camera) is stored for each object. As shown in Figure 3(c), the foreground image data size, size, pixel bit depth, and pixel values are stored at the beginning of the image data in raster order. The foreground image data from each camera is then stored for each object. Note that if the camera does not capture the object, NULL data can be written, or the number of cameras capturing the object and the corresponding camera ID can be stored.
[0040] Figure 4 shows an example of the configuration of a point cloud data dataset. For the sake of explanation, it is assumed that the point cloud dataset is stored in units of frames, but this is not limited to this. A point cloud data header is stored at the beginning of the dataset. As shown in Figure 4(a), the header stores information such as the fact that this dataset is a point cloud data dataset and the number of frames.
[0041] The information contained in each frame is shown in Figure 4(b). Time information indicating the time of the frame is saved in each frame. Next, the data size of the frame is saved. This is for referencing the data of the next frame, and may be saved together in the header. Next, the objects at the time indicated by the time information, i.e., the number of point cloud data (Number of Objects) P, are saved. Thereafter, point cloud data for each object is saved in order. First, the number of coordinate points making up the point cloud of the first object (Number of Points in 1st Object) is saved. Thereafter, the coordinates of that number of points (Point coordination in 1st Object) are saved. In the same manner, point cloud data for all objects included at that time is saved.
[0042] In this embodiment, the coordinate system is saved as three-axis data, but this is not limited to this and polar coordinates or other coordinate systems may be used. The data length of the coordinate values may be fixed and written in the point cloud data header, or the data length may vary for each point cloud. If the data length varies for each point cloud, the data size is saved for each point cloud. By referring to the data size, the storage position of the next point cloud can be identified from the number of coordinate points.
[0043] FIG. 5 shows an example of the configuration of a colored point cloud data dataset. For the sake of explanation, it is assumed that the colored point cloud dataset is saved on a frame-by-frame basis, but this is not limiting. A colored point cloud data header is saved at the beginning of the dataset. As shown in FIG. 5(a), the header stores information such as the fact that this dataset is a colored point cloud data dataset and the number of frames.
[0044] The information contained in each frame is shown in Figure 5(b). In each frame, time information indicating the time of the frame is stored in the point cloud data of that frame. Next, the data size of that frame is stored. This is for referencing the data of the next frame, and may be stored together in the header. Next, the number of objects at the time indicated by the time information, i.e., the number of colored point cloud data (Number of Objects) P, is stored. Thereafter, the colored point cloud data for each object is stored in order. First, the number of coordinate points that make up the point cloud of the first object (Number of Points in 1st Object) is stored. Subsequently, the coordinates of that number of points and the color information of the points (Point coordination in Pth Object) are stored. Thereafter, the colored point cloud data for all objects included in that time are stored in the same manner.
[0045] In this embodiment, the coordinate system is a three-axis data system, and color information is stored as values of the three primary colors of RGB, but this is not limited to this. Polar coordinates or other coordinate systems may also be used. Color information may also be expressed as information such as a uniform color space, brightness, or chromaticity. The data length of the coordinate values may be fixed and recorded in the colored point cloud data header, or different data lengths may be used for each colored point cloud. If the data length differs for each colored point cloud, the data size is stored for each point cloud. By referencing the data size, the storage position of the next colored point cloud can be identified from the number of coordinate points.
[0046] Figure 6 shows an example of the structure of a range image data dataset. For the sake of explanation, we will assume that the range image dataset is saved in frame units, but this is not limited to this. A range image data header is saved at the beginning of the dataset. As shown in Figure 6(a), the header stores information such as the fact that this dataset is a range image data dataset and the number of frames.
[0047] The information contained in each frame is shown in Figure 6(b). Each frame stores time information (time information), which indicates the frame time. Next, the data size of the frame (data size) is stored. This information is used to reference the data of the next frame and can be stored together in the header. Because range image data is acquired on a per-camera basis, the number of cameras used in the frame (Number of Cameras C) is stored. Next, the camera IDs of each camera (Camera ID of Cth Camera) are stored in order. Next, the range image data corresponding to each camera (Distance Image of Cth Camera) is stored. At the beginning of the range image data, the data size of the range image, the bit depth of the pixel values, and the pixel values are stored in raster order, as shown in Figure 6(c). The range image data from each camera is then stored successively. Note that if the camera does not capture the subject, NULL data can be written, or only the number of cameras C that capture the subject and the corresponding camera ID can be stored.
[0048] Figure 7 shows an example of the structure of a mesh data set. For the sake of explanation, it is assumed that the mesh data set is stored on a frame-by-frame basis, but this is not limiting. A mesh data header is stored at the beginning of the data set. As shown in Figure 7(a), the header stores information such as the fact that this data set is a mesh data set and the number of frames.
[0049] The information contained in each frame is shown in Figure 7(b). Each frame stores time information (Time information) that indicates the time of the frame. Next, the data size of the frame (Data Size) is stored. This is used to reference the data of the next frame, and may be stored together in the header. Next, the number of objects (Number of Objects) P is stored. Below, mesh data for each object is stored. The mesh data for each object stores the number of polygons that make up the mesh data (Number of Points in Pth Object) at the beginning. Furthermore, the mesh data for each object stores data for each polygon, i.e., the coordinates of the polygon vertices and polygon color information (Polygon information in Pth Object). Below, mesh data for all objects included in that time is stored in a similar manner.
[0050] In this embodiment, the coordinate system describing vertices is three-axis data, and color information is stored as RGB primary color values, but this is not limited to this. Polar coordinates or other coordinate systems may also be used. Color information may also be expressed using information such as a uniform color space, brightness, and chromaticity. In this embodiment, the mesh data is in a format that assigns one color to each polygon, but this is not limited to this. For example, mesh data without color information may be generated, and the color of the polygon may be determined using foreground image data. Furthermore, the virtual viewpoint image generation system 1 may generate mesh data without color information and a foreground image as material data, separate from mesh data that assigns one color to each polygon. When a virtual viewpoint image is generated using mesh data in which one color is assigned to each polygon, the color of the polygon is determined regardless of the position of the virtual viewpoint and the viewing direction from the virtual viewpoint. On the other hand, when a virtual viewpoint image is generated using mesh data without color information and a foreground image, the color of the polygon changes depending on the position of the virtual viewpoint and the viewing direction from the virtual viewpoint.
[0051] Alternatively, existing formats such as PLY format or STL format may be used, by providing a dataset class code corresponding to each format.
[0052] <Operation of the virtual viewpoint image generation system> Next, the operation of the virtual viewpoint image generation system 1 will be described with reference to the flowchart in Fig. 8. The processing in Fig. 8 is realized by the CPU of each device in the virtual viewpoint image generation system 1 reading and executing a program stored in a ROM or an external storage device.
[0053] In step S800, the management unit 108 generates a sequence header in a format for storing material data. At this point, a data set is determined to be generated by the virtual viewpoint image generation system 1. In step S801, the camera 101 starts capturing images, and the processes of steps S802 to S812 are repeated for each frame of the captured image.
[0054] In step S802, the input unit 102 acquires frame data of captured images from the cameras 101a to 101z and transmits it to the storage unit 103. In step S803, the first model generation unit 104 generates a foreground image and a silhouette image based on the captured images acquired from the storage unit 103. In step S804, the first model generation unit 104 uses the generated silhouette images to perform shape estimation using, for example, a volume intersection method, and generates point cloud data.
[0055] In step S805, management unit 108 acquires a foreground image from first model generation unit 104 and stores it in storage unit 109 according to the formats shown in FIGS. 2 and 3. Management unit 108 also updates data of values that change during repeated processing, such as the number of frames in the foreground image dataset header that has already been stored. In step S806, management unit 108 acquires point cloud data from first model generation unit 104 and stores each point cloud data according to the formats shown in FIGS. 2 and 4. Management unit 108 also updates data such as the number of frames in the point cloud dataset header.
[0056] In step S807, the second model generation unit 105 generates colored point cloud data using the point cloud data and foreground image data generated by the first model generation unit 104. In step S808, the management unit 108 acquires the colored point cloud data from the second model generation unit 105 and stores it in the storage unit 109 in accordance with the formats shown in Figures 2 and 5. The management unit 108 also updates data such as the number of frames in the colored point cloud dataset header.
[0057] In step S809, the distance image generation unit 106 uses the point cloud data generated by the first model generation unit 104 to generate distance image data corresponding to each camera 101. In step S810, the management unit 108 acquires the distance images from the distance image generation unit 106 and stores them in the storage unit 109 in the formats shown in Figures 2 and 6. The management unit 108 also updates data such as the number of frames in the distance image data set header.
[0058] In step S811, the mesh model generation unit 107 generates mesh data using the distance image data and foreground image data generated by the distance image generation unit 106. In step S812, the management unit 108 acquires the mesh data from the mesh model generation unit 107 and stores it in the storage unit 109 in accordance with the formats shown in Figures 2 and 7. The management unit 108 also updates data such as the number of frames in the mesh data set header. In step S813, the virtual viewpoint image generation system 1 determines whether to end the repetitive processing. For example, the repetitive processing ends when the camera 101 finishes capturing images or when processing for a predetermined number of frames is completed.
[0059] In step S814, a request for material data is received from the terminal. The transmitter / receiver 110 also acquires information for identifying a format of material data that can be processed by the terminal 112 that displays the virtual viewpoint image. Here, being able to process material data by the terminal 112 means that the terminal 112 can interpret the data in the file of the acquired material data and process it appropriately in accordance with the contents described therein. For example, a file format of material data that the terminal 112 does not support and cannot read is deemed to be unprocessable.
[0060] The acquired information may include, for example, the specifications of the terminal, information about the type, information about the processing capacity of the terminal, etc. Furthermore, for example, the acquired information may include, for example, information about a method for generating a virtual viewpoint image used by the terminal, information about software used to generate or play back the virtual viewpoint image, etc. Furthermore, for example, the acquired information may include information regarding the format of a virtual viewpoint image that can be displayed by the terminal. The transmitting / receiving unit 110 transmits the acquired information to the selecting unit 111. Based on the acquired information, the selecting unit 111 determines material data to be output to the display device from among the multiple material data stored in the storing unit 109.
[0061] For example, the information processing device 100 selects a foreground image dataset and a point cloud dataset for a terminal such as the terminal 112a that generates a virtual viewpoint image from foreground image data and point cloud data. Furthermore, the information processing device 100 selects a colored point cloud dataset for a terminal such as the terminal 112b that generates a virtual viewpoint image from colored point cloud data. Furthermore, the information processing device 100 selects a distance image dataset and a foreground image dataset for a terminal such as the terminal 112c that generates a virtual viewpoint image from distance image data and foreground image data. Furthermore, when the information processing device 100 has a renderer that renders mesh data, such as the terminal 112d, it selects a mesh dataset based on the dataset type. In this way, the information processing device 100 operates so that different material data is output depending on the terminal.
[0062] Note that the information acquired by the transmitter / receiver 110 from the terminal is not limited to the information described above, and may be information specifying specific material data such as point cloud data. Furthermore, the information processing device 100 may be configured to previously store in an external storage device or the like a table associating multiple terminal types with material data to be output to each terminal, and the selection unit 111 may acquire this table. In this case, the information processing device 100 selects material data listed in the acquired table upon confirming the connection of the terminal.
[0063] In step S815, the transmitting / receiving unit 110 receives information related to the virtual viewpoint image to be generated. The information related to the virtual viewpoint image includes information for selecting a sequence for generating the virtual viewpoint image. The information for selecting the sequence includes, for example, at least one of the name of the sequence, the date and time of imaging, and the time information of the first frame for generating the virtual viewpoint image to be generated. The transmitting / receiving unit 110 outputs the information related to the virtual viewpoint image to the selecting unit 111.
[0064] The selection unit 111 outputs information about the sequence data selected by the management unit 108, such as the sequence name, to the management unit 108. The management unit 108 selects the sequence data with the input sequence name. The selection unit 111 outputs time information of the first frame from which a virtual viewpoint image is generated to the management unit 108. The management unit 108 compares the time information at the start of imaging of each data set with the time information of the first frame from which the input virtual viewpoint image is generated, and selects the first frame. For example, if the time information at the start of imaging of the sequence is 10:32:42:00 on April 1, 2021, and the time information of the first frame from which a virtual viewpoint image is generated is 10:32:43:40 on April 1, 2021, 100 frames of data will be skipped. This is achieved by calculating the start of the frame data to be read based on the frame data size of each data set.
[0065] In step S816, the information processing device 100 and the terminal 112 repeatedly execute the processes from S817 to S820 until the final frame of the virtual viewpoint image to be generated.
[0066] In step S817, the selection unit 111 selects a data set to be stored in the storage unit 109 based on the type of data set required to generate a virtual viewpoint image input in step S814. In step S818, the management unit 108 selects material data corresponding to the frames to be read in sequence. In step S819, the storage unit 109 outputs the selected material data to the transmission / reception unit 110, which then outputs the material data to the terminal 112 that generates the virtual viewpoint image. In step S820, the terminal 110 generates and displays a virtual viewpoint image based on the received material data and the virtual viewpoint set by user operation. In a terminal that generates a virtual viewpoint image from foreground image data and point cloud data, such as the terminal 112a, the positions of points of the point cloud data in the virtual viewpoint image are first identified. The virtual viewpoint image is then generated by coloring pixel values corresponding to the positions of the points using the projection of the foreground image data. In the method of generating a virtual viewpoint image from foreground image data and point cloud data, the color used to color the points changes depending on the position of the virtual viewpoint and the line of sight from the virtual viewpoint. In a terminal that generates a virtual viewpoint image from colored point cloud data, such as terminal 112b, the positions of the points of the point cloud data in the virtual viewpoint video are first identified. The virtual viewpoint image is generated by coloring the pixel values corresponding to the identified positions of the points with the colors associated with the points. In the method of generating a virtual viewpoint image using a colored point cloud, the colors corresponding to each point are fixed, so the points are colored with a predetermined color regardless of the position of the virtual viewpoint or the line of sight from the virtual viewpoint.
[0067] In a terminal such as terminal 112c that generates a virtual viewpoint image from distance image data and foreground image data, the position of an object on the virtual viewpoint image is identified from multiple pieces of distance image data, and the virtual viewpoint image is generated by texture mapping using the foreground image. In a terminal such as terminal 112d that generates a virtual viewpoint image from polygon data, the virtual viewpoint image is generated by pasting color values and images onto surfaces visible from the virtual viewpoint, similar to ordinary computer graphics.
[0068] In step S821, the virtual viewpoint image generating system 1 determines whether to end the repetitive processing. For example, the repetitive processing ends when the generation of the virtual viewpoint image ends or when the processing of a predetermined number of frames ends.
[0069] 10 is a sequence diagram showing the state of communication between the various units. First, the terminal 112 is started up (F1100), and the terminal 112 acquires information about the renderer for generating virtual viewpoint images that the terminal 112 has. The terminal 112 transmits the acquired terminal information to the information processing device 100 (F1101). Here, the information transmitted from the terminal 112 is assumed to be information about the specifications of the terminal, but any of the information described above may be transmitted.
[0070] The transmitting / receiving unit 110 transmits the terminal specifications to the selecting unit 111 (F1102). Based on the transmitted information, the selecting unit 111 selects a data set of material data that can be rendered by the terminal 112. The selecting unit 111 notifies the management unit 108 of the type of the selected data set (F1103). At this time, the selecting unit 111 stores, for example, a type code of the material data in advance, and transmits the type code of the determined type of material data to the management unit 108. Furthermore, the terminal 112 outputs a data transmission start request to the information processing device 100 (F1104), and the transmitting / receiving unit 110 transmits the transmission start request to the management unit 108 via the selecting unit 111 (F1105).
[0071] The terminal 112 transmits time information of the frame from which the virtual viewpoint image starts to the information processing device 100 (F1106), and inputs it to the management unit 108 via the transmission / reception unit 110 and the selection unit 111 (F1107). The management unit 108 transmits to the storage unit 109 a specification for reading frame data of the data set selected from the type of selected material data and the time information (F1108). The storage unit 109 outputs the specified data to the transmission / reception unit 110 (F1109), and the transmission / reception unit 110 transmits the material data to the terminal 112 that has received the start of transmission (F1110). The terminal 112 receives the material data, uses the material data to generate (render) a virtual viewpoint image, and displays the generated virtual viewpoint image (F1111). Thereafter, the information processing device 100 transmits the material data in frame order until a request to end transmission is received from the terminal 112.
[0072] When the generation of the virtual viewpoint image is completed, the terminal 112 transmits a request to end transmission of the material data to the information processing device 100 (F1112). When the transmission / reception unit 110 receives the request to end transmission, it notifies each processing unit of the information processing device 100 (F1113). This completes the processing. Note that the above-described processing is executed independently for each of the multiple terminals 112.
[0073] With the above configuration and operation, the information processing device 100 outputs to the terminal 112, from among a plurality of pieces of material data containing different material data, material data that is determined based on information for identifying a format of the material data that can be processed by the terminal 112. This makes it possible to appropriately generate and display a virtual viewpoint image on each terminal even when a plurality of terminals 112 are connected.
[0074] (Modification of the first embodiment) In the first embodiment, for example, the second model generation unit 105 generates a colored point cloud using the point cloud data generated by the first model generation unit 105 and the foreground image data. Thus, in the first embodiment, an example has been described in which the first model generation unit 104, the second model generation unit 105, the distance image generation unit 106, and the mesh model generation unit 107 generate material data using data generated by each other. However, this is not limiting. For example, in the first embodiment, the first model generation unit 104, the second model generation unit 105, the distance image generation unit 106, and the mesh model generation unit 107 may each acquire a captured image from the storage unit 103 and independently generate material data. In this case, the processes of steps S803 to S812 in FIG. 8 are executed in parallel by the respective processing units.
[0075] In the first embodiment, the foreground image is divided into individual objects and stored, but the present invention is not limited to this, and the image captured by the camera itself may be stored. In the first embodiment, the point cloud data is stored in units of frames, but the present invention is not limited to this. For example, it is also possible to identify each object and store the data in units of objects.
[0076] In the first embodiment, information for identifying the data is included in the header of the sequence in order to identify the three-dimensional shape data, but this is not limited to this, and the data may be managed as a list associated with the data in the storage unit 109 or the management unit 108. In addition, the point cloud data, foreground image, colored point cloud data, range image data, and mesh data generated by the first model generation unit 104, the second model generation unit 105, the range image generation unit 106, and the mesh model generation unit 107 may be encoded in the management unit 108.
[0077] Furthermore, in the first embodiment, time information is used as information representing time, but the present invention is not limited to this, and it is possible to use a frame ID or information that can identify a frame. In the first embodiment, a method for generating a data set may be added as meta information of the data set, and when data sets of the same type exist but generated by different methods, the information may be sent to the terminal, and the terminal may select the data set. Furthermore, in addition to the present embodiment, a spherical image or a normal two-dimensional image may be generated and saved as a data set.
[0078] It is also possible to simply interleave and alternately transmit frame data of point cloud data and frame data of foreground image data as shown in Figure 12. Similarly, distance image data may be used instead of point cloud data.
[0079] Furthermore, material data may be managed on an object-by-object basis, as shown in Figure 13 as an example.
[0080] An object data set is defined as a new data set for the sequence. Note that there may be multiple object data sets. For example, they may be created for each team to which the subject belongs, or for other predetermined units.
[0081] The object data header at the beginning of an object data set stores an object data header start code (Object Data Header) indicating the beginning of the object data set. The object data header also stores a data size (Data size). Next, the number of objects (Number of Objects) P included in the object data set is stored. The following data information for each object is stored.
[0082] The object data information (Pth Object Data Description) stores the object ID number (Object ID) and the size of the object data information (Data Size). Next, a data set class code (Data set class code) indicating the type of data as a data set is stored. A pointer to the object data of the object is also stored. Metadata related to the object is also stored. For example, in the case of sports, information such as the player's name, team, and uniform number can be stored as metadata.
[0083] Data based on the dataset class code is stored in each object data. Figure 13 is an example of a dataset that integrates a foreground image and point cloud data. At the beginning, a header is stored containing the size of the object dataset, time information for the first frame, frame rate, etc. Data is then stored frame by frame. Frame by frame data contains the time information for that frame (time information) and the data size for that frame (data size). Next, the size of the point cloud data for that frame (data size of point cloud) and the number of points (number of points) are stored. Next, the coordinates of the point cloud points (point coordination) are stored. Next, the number of cameras (number of cameras) that captured the subject is stored. Next, the camera ID of each camera (camera ID of 1st camera) and foreground image data (data size of foreground image of 1st camera) are stored. Next, the camera ID and foreground image data are stored for each camera.
[0084] If the dataset of object data is colored point cloud data, a dataset class code is set as shown in Figure 14(a), and time information, data size, number of points in the colored point cloud, and colored point cloud data are saved as data for each frame. If the dataset of object data is mesh data, a dataset class code is set as shown in Figure 14(b), and time information, data size, number of polygons, and data for each polygon are saved as data for each frame. If the dataset of object data is range images, a dataset class code is set as shown in Figure 13(c), and time information, data size, number of cameras, and data for each range image are saved as data for each frame.
[0085] By managing and outputting data for each subject in this way, it is possible to specify a subject, such as a player or performer, and shorten the time it takes to read an image centered on the specified subject. This is effective when generating an image centered on the movement of a specified player, or a virtual viewpoint image that follows behind the subject. It also makes it easy to search for a specified player.
[0086] 15, the management unit 108 may configure a sequence for each type of material data. In this case, the selection unit 111 selects a sequence that corresponds to the type of material data to be output to the terminal. This configuration allows sequences to be output together in accordance with the rendering method used by the output destination terminal, thereby enabling rapid output.
[0087] In the first embodiment, a configuration has been described in which a virtual viewpoint image is generated by the terminal after the generation of a plurality of pieces of material data has been completed, but the present invention is not limited to this. By performing the operation of Fig. 9, it becomes possible to generate a virtual viewpoint image on the terminal 112 while generating material data. In the figure, steps in which the operations of the respective parts are the same as those in Fig. 8 are assigned the same numbers, and descriptions thereof will be omitted.
[0088] 8, in step S900, the transmitting / receiving unit 110 acquires information for identifying a format of material data that can be processed by the terminal 112 that displays the virtual viewpoint image. Furthermore, the transmitting / receiving unit 110 acquires information such as time information related to the virtual viewpoint image generated by the terminal 112. In S901, the virtual viewpoint image generation system 1 repeatedly executes the processes of S802 to S920 for each frame of the virtual viewpoint image to be generated.
[0089] In step S902, the selection unit 111 outputs information about the sequence data selected based on information such as time information to the management unit 108. The management unit 108 selects the sequence data indicated by the acquired information. In steps S801 to S812, as in FIG. 8, captured images of each frame are acquired, and material data is generated and saved. In step S917, the management unit 108 selects a data set of material data to be transmitted to the terminal 112 based on the type of data set determined in step S902. In step S918, the transmission / reception unit 110 selects frame data of the selected data set from the generated material data and outputs it to the terminal.
[0090] In step S920, the terminal 112 generates and displays a virtual viewpoint image based on the received material data and the virtual viewpoint set based on a user operation or the like.
[0091] In S921, the virtual viewpoint image generation system 1 determines whether to end the repetitive processing. For example, when the information processing device 100 receives a notification to end the generation of the virtual viewpoint image from the terminal 112, the repetitive processing ends. Through the above operations, it becomes possible to generate a virtual viewpoint image on the terminal 112 while generating material data. This makes it possible to view the virtual viewpoint image with low latency while capturing an image with the camera 101, for example. Note that even if the terminal 112 ends the generation of the virtual viewpoint image in S921, the generation and storage of material data shown in S802 to S812 may be continued.
[0092] Furthermore, in the first embodiment, the foreground image data and the point cloud data are stored as separate data sets, but this is not limiting. FIG. 11 shows an example of the configuration of a foreground image & point cloud data set that integrates the foreground image data and the point cloud data. In the format shown in FIG. 11, a foreground image & point cloud data header (Foreground & Point Cloud Model Header) is stored at the beginning of the data set, similar to the foreground image data set of FIG. 3. This data header stores information such as the fact that this data set is a foreground image & point cloud data data set, the number of frames, etc. Next, the foreground image data and point cloud data for each frame are stored. The foreground image & point cloud data for each frame stores time information (time information) indicating the time of the frame and the data size (data size) of the frame. Next, the number of objects (Number of Objects) P corresponding to the time indicated by the time information is stored. Furthermore, the number of cameras (Number of Cameras) C used for image capture at that time is stored. Next, the camera IDs (Camera IDs of 1st to Cth Cameras) of the cameras used are stored. Next, the data for each object is stored. First, the number of coordinate points that make up the point cloud of the subject is saved, just like the point cloud data in Figure 4. The coordinates of that number of points are then saved. Next, the foreground image data from each camera that captured the subject is stored in order, just like in Figure 3. From then on, point cloud data and foreground image data are saved for each subject.
[0093] With the above configuration, the required combination of data sets can be integrated and saved by the terminal that is the renderer. This makes it possible to simplify data reading. Note that although the format integrating point cloud data and foreground image data has been described here, the present invention is not limited to this. For example, it is possible to configure a format integrating any combination of data, such as a format integrating a range image and a foreground image, or a format integrating three or more pieces of material data.
[0094] 16 may be added to the configuration of the virtual viewpoint image generation system 1 in the first embodiment. The CG model generation unit 113 uses CAD, CG software, etc. to generate multiple types of three-dimensional shape data as material data. The management unit 108 can acquire the material data generated by the CG model generation unit 113 and store it in the storage unit 109 in the same way as other material data.
[0095] 17 shows another example of the flow of communication performed by the virtual viewpoint image generation system 1. The processing shown in FIG. 17 is an example in which processing is started with the transmission of material data from the information processing device 100 as a trigger. First, the selection unit 111 determines the type of data set to be transmitted and the time information of the beginning of the data to be transmitted, and sends these to the management unit 108 (F1700). The management unit 108 transmits to the storage unit 109 a command to read frame data of the data set selected from the type of the selected data set and the time information. The selection unit 111 transmits a command to the transmission / reception unit 110 to start transmitting data (F1701). The transmission / reception unit 110 transmits a command to each terminal 112 to notify them that data transmission will start (F1702). In response to this, the terminal 112 prepares to receive data.
[0096] The selection unit 111 transmits the type of material data and the start time information as specifications of the data set to be transmitted to the transmission / reception unit 110 (F1703). The transmission / reception unit 110 transmits this information to each terminal 112 (F1704). Each terminal 112 determines whether a virtual viewpoint image can be generated using the data set of the type transmitted from the information processing device 100 (F1705), and, if possible, prepares to generate the virtual viewpoint image. The terminal 112 also transmits the result of the determination to the information processing device 100 (F1706). If the result of the determination by the terminal 112 is "yes," the selection unit 111 transmits a notification to the management unit 108 indicating the start of data transmission. If the result of the determination is "no," the information processing device 100 transmits specifications of a different data set to the terminal 112 and waits for the result of the determination again. Note that, in F1703, the selection unit 111 may be configured to refer to the type of material data already stored in the storage unit 109 and transmit the specifications of a data set that can be transmitted. In this case, in F1705, the terminal 112 notifies the information processing apparatus 100 of information that can identify the types of material data that the terminal 112 can process.
[0097] The management unit 108 receives the transmission start signal and transmits to the storage unit 109 a specification for reading frame data of the data set selected from the type of selected material data and the time information of the first frame (F1707). The storage unit 109 outputs the specified data to the transmission / reception unit 110 (F1708), and the transmission / reception unit 110 transmits the data to the terminal that has been enabled for reception (F1709). The terminal 112 receives the data, performs rendering, etc., and generates and displays a virtual viewpoint image (F1710). Thereafter, the terminal 112 receives frame data from the selected data set in frame order until generation of the virtual viewpoint image is completed, and generates a virtual viewpoint image.
[0098] By performing the above communication, it becomes possible for the transmitting side to select material data. It is also possible to transmit data unilaterally without transmitting the result of the decision. When the information processing device 100 selects material data, it may be configured to refer to, for example, information on terminals with which it has communicated in the past, or a table that previously associates terminal information with the type of material data.
[0099] In the first embodiment, an example in which point cloud data, a foreground image, colored point cloud data, a distance image, and mesh data are generated as a plurality of pieces of material data has been described, but a configuration in which data other than these pieces of material data is also generated may be used. As an example, a case in which billboard data is generated will be described.
[0100] Billboard data is data obtained by coloring plate-shaped polygons using a foreground image. Here, it is assumed that the data is generated by the mesh model generation unit 107. FIG. 22 shows an example of the configuration of a data set of billboard data stored in the storage unit 109. For the sake of explanation, it is assumed that the billboard data set is stored on a frame-by-frame basis, but this is not limiting. As shown in FIG. 22(a), a billboard data header is stored at the beginning of the data set. The header stores information such as that the data set is a billboard data data set and the number of frames. Time information indicating the time of the frame is stored in the billboard data of each frame. Next, the data size of the frame is stored. Billboard data for each camera is then stored for each subject.
[0101] Billboard data is saved for each subject as a combination of billboard polygon data for each camera and image data to be texture-mapped. First, one polygon data for the billboard (Polygon billboard of Cth Camera for Pth Object) is written. Next, foreground image data for texture mapping (Foreground Image of Cth Camera for Pth Object) is saved. The billboard method treats the subject as a single polygon, which can significantly reduce processing and data volume. Note that the number of polygons representing billboard data is not limited to one, and is assumed to be fewer than the number of polygons in mesh data. This makes it possible to generate billboard data that requires less processing and requires less data volume than mesh data.
[0102] As described above, the virtual viewpoint image generation system 1 according to the present embodiment is capable of generating various types of material data and outputting them to a terminal. It is not necessary to generate all of the above-described material data, and any desired material data may be generated. Furthermore, instead of generating multiple types of material data in advance, the virtual viewpoint image generation system 1 capable of generating multiple types of material data may specify the material data to be generated in accordance with information about the terminal. For example, the selection unit 111 of the information processing device 100 determines the type of material data to be output to the terminal 112 based on information for specifying the type of virtual viewpoint image that the terminal 112 can display. The selection unit 111 notifies the management unit 108 of the determined type of material data, and the management unit 108 transmits an instruction to a processing unit that generates the determined material data to generate the material data. For example, if the determined material data is colored point cloud data, the management unit 108 instructs the second model generation unit 105 to generate colored point cloud data.
[0103] Furthermore, management unit 108 can read image data input from storage unit 103 via a signal line (not shown) and create the determined material data. The read image data is input to a corresponding generation unit, either first model generation unit 104, second model generation unit 105, distance image generation unit 106, or mesh model generation unit 107, to generate the required material data. The generated material data is sent to terminal 112 via management unit 108, storage unit 109, and transmission / reception unit 110.
[0104] Furthermore, in the first embodiment, the information for identifying the format of material data that can be processed by the terminal 112 has been described as being unchanging information such as the type of terminal, but this is not limiting. That is, the information for identifying the format of material data that can be processed by the terminal 112 may be dynamically changing information. For example, if the material data that the terminal 112 can process changes due to an increase in the processing load, etc., the terminal 112 transmits the information again to the information processing device 100. This allows the information processing device 100 to select appropriate material data to output to the terminal 112. Furthermore, the information processing device 100 may dynamically determine the material data to output in accordance with changes in the communication bandwidth of the communication path with the terminal 112.
[0105] Furthermore, in the first embodiment, the generated material data is stored in the same storage unit 109, but this is not limiting. That is, a configuration may be adopted in which different types of material data are stored in different storage units. In this case, the information processing device 100 may be configured to have multiple storage units, or each storage unit may be included in the virtual viewpoint image generation system 1 as a different device. When multiple types of material data are stored in different storage units, the information processing device 100 is configured to output to the terminal 112 information such as a pointer indicating the location where the material data to be output to the terminal 112 is stored. This allows the terminal 112 to access the storage unit in which the material data required for generating a virtual viewpoint image is stored and acquire the material data.
[0106] (Second embodiment) In the first embodiment, the configuration of a system in which the terminal 112 generates a virtual viewpoint image was described. In the present embodiment, a configuration in which an information processing device generates a virtual viewpoint image will be described. Note that the hardware configuration of the information processing device 200 in this embodiment is the same as in the first embodiment, and therefore a description thereof will be omitted. Furthermore, the same reference numerals will be used to designate the same functional configuration as in the first embodiment, and a description thereof will be omitted.
[0107] 18 is a diagram illustrating the configuration of a virtual viewpoint image generation system 2 having an information processing device 200. The information processing device 200 has a virtual viewpoint image generation unit 209, a transmission / reception unit 210, and a selection unit 211 in addition to the same configuration as the information processing device 100 of the first embodiment. The terminals 212a to 212d set a virtual viewpoint based on user operations or the like, and transmit virtual viewpoint information representing the set virtual viewpoint to the information processing device 200. The terminals 212a to 212d do not have a function (renderer) for generating a virtual viewpoint image, and only set a virtual viewpoint and display the virtual viewpoint image. In the following description, unless there is a particular reason to distinguish between them, the terminals 212a to 212d will be simply referred to as terminals 212.
[0108] The transmitting / receiving unit 210 receives virtual viewpoint information from the terminal 212 and transmits it to the virtual viewpoint image generating unit 209. It also has a function of transmitting the generated virtual viewpoint image to the terminal that transmitted the virtual viewpoint information. The virtual viewpoint image generating unit 209 has multiple renderers and generates a virtual viewpoint image based on the input virtual viewpoint information. The multiple renderers each generate a virtual viewpoint image based on different material data. Note that the virtual viewpoint image generating unit 209 can operate multiple renderers simultaneously and generate multiple virtual viewpoint images in response to requests from multiple terminals.
[0109] The selection unit 211 selects a data set required for the virtual viewpoint image generation unit 209 to generate a virtual viewpoint image. The selection is made based on the processing load and the data read load. For example, suppose that a request for image generation is received from terminal 212b while image generation is being performed using a foreground image data set and a point cloud data set in response to a request from terminal 212a. In this embodiment, it is assumed that the foreground image data and point cloud data have larger data volumes than other material data. If the selection unit 211 determines that the read bandwidth from the storage unit 109 is insufficient to read two sets of data, a foreground image data set and a point cloud data set, the selection unit 211 selects a mesh data set with a smaller data volume than the foreground image data and point cloud data. The virtual viewpoint image generation unit 209 sets a renderer for the mesh data and uses the mesh data to generate a virtual viewpoint image for the terminal 212b. Note that the data volume can also be reduced by selecting a colored point cloud data set with a smaller data volume than the foreground image data and point cloud data instead of the mesh data.
[0110] Furthermore, depending on requests from the terminal 212, the processing capacity of the virtual viewpoint image generation unit 209 may be exceeded, or it may become necessary to reduce power consumption. The selection unit 211 selects another data set with a lower processing load so that a renderer with a lighter processing load is used. For example, processing using a foreground image dataset and a point cloud dataset imposes a larger processing load than processing using a colored point cloud dataset. If the number of terminals exceeds the capacity to generate a virtual viewpoint image using the foreground image dataset and the point cloud dataset, the virtual viewpoint image is generated using the colored point cloud dataset. As described above, the information processing device 200 in this embodiment selects material data to be used to generate a virtual viewpoint image depending on the processing load.
[0111] The operation of the virtual viewpoint image generation system 2 will be described using the flowchart in Fig. 19. In Fig. 19, steps in which the operation of each unit is the same as in Fig. 8 are assigned the same numbers, and descriptions thereof will be omitted. In step S930, the transmission / reception unit 212 acquires virtual viewpoint information from the terminal 212 and transmits it to the selection unit 211 and the virtual viewpoint image generation unit 209. In step S931, the selection unit 211 selects a virtual viewpoint image generation method (renderer) for the virtual viewpoint image generation unit 209 based on the load, transmission bandwidth, etc. of the virtual viewpoint image generation unit 209, and transmits information to the management unit 108 to select the necessary material data.
[0112] In step S932, a virtual viewpoint image is generated using the selected virtual viewpoint image generation means and output to the terminal via the transmission / reception unit 210. With the above configuration and operation, the three-dimensional information processing device can adjust the processing load in response to an increase or decrease in the number of terminals, and can transmit virtual viewpoint images to a greater number of terminals.
[0113] Note that, in the present embodiment, a configuration has been described in which a virtual viewpoint image is generated using different material data depending on the processing load and transmission bandwidth of the information processing device 200. However, the present invention is not limited to this. The information processing device 200 may be configured to determine the virtual viewpoint image to be generated based on, for example, the format of the virtual viewpoint image that can be played back by the terminal 212, the software or application used for playback, the display capabilities of the display unit of the terminal 212, and the like.
[0114] Furthermore, if the processing capacity of the virtual viewpoint image generating unit 209 is exceeded, the number of information processing devices 200 may be increased depending on the type of data set requested from the terminal.
[0115] (Third embodiment) In this embodiment, an information processing device that converts generated material data into different material data for use will be described. The hardware configuration of the information processing device 300 of this embodiment is the same as that of the above-described embodiment, and therefore a description thereof will be omitted. Furthermore, the same reference numerals will be used to designate the same functional configuration as that of the above-described embodiment, and a description thereof will be omitted.
[0116] The configuration of a virtual viewpoint image generation system 3 having an information processing device 300 will be described with reference to FIG. 20 . The conversion unit 301 is a conversion unit that converts a data set into a different data set. In addition to the functions of the selection unit 111 in the first embodiment, the selection unit 311 has a function of controlling conversion of the data set using the conversion unit 301 when a data set not in the storage unit 109 is input from the terminal 112. The storage unit 309 has a function of outputting the type of the stored data set to the selection unit 311 and a function of outputting the stored data to the conversion unit 301.
[0117] An example of processing performed by the information processing device 300 will be described. For example, assume that only a foreground image dataset and a point cloud dataset are stored in the storage unit 109, and the terminal 112 requests output of a range image. In this case, the selection unit 311 compares the dataset request from the terminal 212 input via the transmission / reception unit 110 with the contents of the storage unit 309. As a result of the comparison, it determines that a matching dataset is not stored. Furthermore, the selection unit 311 instructs the management unit 108 to read the necessary dataset to be used for conversion by the conversion unit 301. The conversion unit 301 acquires the point cloud dataset from the storage unit 309, calculates the distance from each point in the point cloud data to each camera, generates range image data, and converts it into a range image dataset. That is, the point cloud dataset of FIG. 4 is converted into the range image dataset of FIG. 6. The transmission / reception unit 110 outputs the range image data obtained by the conversion to the terminal 112.
[0118] As another example of the conversion process, the conversion unit 301 converts a dataset into another dataset by rearranging the data in the dataset. For example, assume that only a foreground image dataset and a point cloud dataset are stored in the storage unit 109, and that the terminal 112 requests a data format organized by object. The selection unit 311 compares the dataset request from the terminal 112 input from the transmission / reception unit 110 with the contents of the storage unit 309. As a result of the comparison, the selection unit 311 instructs the management unit 108 to read the required dataset. Based on the instruction from the selection unit 311, the management unit 108 reads the foreground image data and the point cloud dataset from the storage unit 309 and transmits them to the conversion unit 301. The conversion unit 301 generates an object dataset by rearranging the data and generating header data, and outputs the object dataset to the terminal 112 via the transmission / reception unit 110.
[0119] The above is an example of the conversion process performed by the conversion unit 301. Note that the conversion process is not limited to the above example, and any data can be converted into any other data, such as changing point cloud data into mesh data, or integrating or separating multiple pieces of material data.
[0120] Next, the operation of the virtual viewpoint image generation system 3 will be described using the flowchart in Fig. 21. In the figure, steps in which the operation of each unit is the same as in the above-described embodiment are given the same numbers and descriptions thereof will be omitted. In step S1001, the selection unit 311 checks whether the material data stored in the storage unit 309 contains data that matches the request from the terminal 112. If there is no matching data set, the process proceeds to step S1001; if there is, the process proceeds to step S816.
[0121] In step S1001, the selection unit 311 selects a data set stored in the storage unit 309 to be used in order to generate a data set that meets the request from the terminal 112. The conversion unit 301 acquires the selected data set from the storage unit 309 and converts the acquired data set into a data set that conforms to the request of the terminal 212. As described above, by the information processing device 300 converting a data set in response to a request from a terminal, it becomes possible to transmit material data to a wider variety of terminals. Furthermore, because material data is converted and generated in response to a request from the terminal 112, the amount of material data that needs to be stored in advance in the storage unit 309 can be smaller than in the case where many types of material data are stored. This allows the resources of the storage unit 309 to be used more effectively. Note that the data set generated by conversion may of course be stored in the storage unit 309. Furthermore, material data that has been converted and generated once and stored in the storage unit 309 may be read from the storage unit 309 and used in subsequent processing.
[0122] (Other embodiments) In the first, second, and third embodiments described above, examples have been described in which material data is output to a terminal that displays a virtual viewpoint image, but the output destination device is not limited to this. For example, the above-described embodiments can also be applied to cases in which material data is output to another device that acquires the material data and performs predetermined processing. In this case, the information processing device 100 determines the material data to be output to the other device based on information for identifying a format of the material data that the other output destination device can process.
[0123] It is possible to combine and use any of the configurations of the first, second, and third embodiments described above. Furthermore, the types of material data are not limited to those described in the above embodiments, and the above embodiments can be applied to a configuration in which a plurality of pieces of material data including different pieces of material data are generated and a virtual viewpoint image is generated.
[0124] The present disclosure can also be realized by providing a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]
[0125] 100 Information processing device 108 Management Department 110 Transmitter / Receiver
Claims
[Claim 1] an acquisition means for acquiring a plurality of material data used to generate a virtual viewpoint image based on a plurality of captured images obtained by capturing images of a subject by a plurality of imaging devices, the material data including first material data and second material data different from the first material data; output means for outputting, to another device, material data determined based on information for specifying a format of material data that can be processed by the other device to which the material data is output, from among the plurality of material data acquired by the acquisition means; An information processing device comprising:
Citation Information
Patent Citations
Object three-dimensional model restoration method, device, and program
JP2016071645A
Free-viewpoint image generation method and free-viewpoint image generation system
WO2018147329A1