Video generation device and video generation method

The image generation device uses AI models to generate free-viewpoint videos with a small number of cameras, overcoming blind spots and ensuring a realistic viewing experience by integrating frame extraction and complementary image generation techniques.

WO2026004106A1PCT designated stage Publication Date: 2026-01-02NTT DOCOMO INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/023555
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing methods for generating free-viewpoint images face challenges in arranging multiple cameras without creating blind spots, particularly in venues like concerts and sports, limiting the ability to provide a realistic viewing experience.

Method used

An image generation device that uses a small number of cameras to capture images, employing AI models to generate complementary images and free-viewpoint videos by integrating frame extraction, complementary image generation, and rendering units to fill in blind spots and allow arbitrary viewpoint selection.

Benefits of technology

Enables the generation of clear free-viewpoint videos with reduced processing load by using a minimal number of cameras, effectively addressing blind spots and providing a realistic viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024023555_02012026_PF_FP_ABST
    Figure JP2024023555_02012026_PF_FP_ABST
Patent Text Reader

Abstract

A video generation device according to one embodiment of the present invention comprises: an acquisition unit which acquires a plurality of captured videos that are captured by a plurality of cameras and capturing viewpoint information that represents capturing viewpoints of the plurality of cameras; a frame extraction unit which extracts a plurality of frame images at a first time point from the plurality of captured videos; a complementary image generation unit which generates, on the basis of the plurality of frame images at the first time point, a complementary image viewed from a complementary viewpoint; and a video generation unit which generates, on the basis of the plurality of frame images at the first time point, the capturing viewpoint information, the complementary image, and complementary viewpoint information indicating the complementary viewpoint, a free-viewpoint image at the first time point, and generates, on the basis of the free-viewpoint image at the first time point, a free-viewpoint image at a second time point which is after the first time point.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation device and image generation method

[0001] The present disclosure relates to an image generation device and an image generation method.

[0002] There is known a technique for generating a free-viewpoint image, from which any viewpoint can be selected, based on a plurality of images captured from different viewpoints. For example, Non-Patent Document 1 describes a technique called 3D Gaussian splatting. 3D Gaussian splatting is a technique for generating point cloud data from a plurality of images captured from different viewpoints, and reconstructing a free-viewpoint image by applying a Gaussian function to parameters set for each point in the point cloud data and smoothly blending the points with their surroundings.

[0003] Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis, “3D Gaussian Splatting for Real-Time Radiance Field Rendering”, SIGGRAPH 2023 (ACM Transactions on Graphics), Vol.42, No.4, (2023)

[0004] The above-mentioned 3D Gaussian splatting technique uses multiple cameras (approximately 30) positioned along a dome-shaped curved surface to capture images of a subject without creating blind spots, and then reconstructs the captured images, thereby enabling the reconstruction of free viewpoint video from which a user can freely select a viewpoint. This technology makes it possible to provide video of a subject (musician or athlete) viewed from a viewpoint (angle) desired by the user during live broadcasts of concerts, sports, and the like, thereby providing the user with a more realistic experience. However, due to structural constraints of the venue, it is not easy to arrange multiple cameras without creating blind spots during live broadcasts of concerts and sports.

[0005] Therefore, an object of the present disclosure is to generate free viewpoint video using a small number of cameras.

[0006] An image generating device according to one embodiment includes an acquisition unit that acquires a plurality of captured images captured by a plurality of cameras and image capturing viewpoint information indicating the image capturing viewpoints of the plurality of cameras; a frame extraction unit that extracts a plurality of frame images at a first point in time from the plurality of captured images; a complementary image generation unit that generates a complementary image viewed from a complementary viewpoint based on the plurality of frame images at the first point in time; and an image generation unit that generates a free viewpoint image at the first point in time based on the plurality of frame images, the image capturing viewpoint information, the complementary image, and the complementary viewpoint information indicating the complementary viewpoint at the first point in time, and that generates a free viewpoint image at a second point in time later than the first point in time based on the free viewpoint image at the first point in time.

[0007] According to the present disclosure, free viewpoint video can be generated using a small number of cameras.

[0008] Fig. 1 is a block diagram showing the functional configuration of an image generation device according to an embodiment. Fig. 2 is a diagram schematically showing the arrangement of multiple cameras. Fig. 3 is a flowchart showing an image generation method according to an embodiment. Fig. 4 is a block diagram showing the hardware configuration of the image generation device.

[0009] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the description of the drawings, the same elements are designated by the same reference numerals, and duplicated description will be omitted.

[0010] 1 is a block diagram showing the functional configuration of an image generation device 1 according to an embodiment. The image generation device 1 generates a free viewpoint image from a plurality of images captured by a plurality of cameras 2, from which a viewpoint can be arbitrarily selected.

[0011] The multiple cameras 2 are imaging devices that capture images of the subject 3 from different viewpoints. In the following description, "image" or "frame image" refers to still image data, and "video" refers to moving image data. In one embodiment, as shown in Figure 2, the multiple cameras 2 are arranged around the subject 3 and capture images of the subject 3 from different viewpoints.

[0012] 1, the video generation device 1 includes, as functional components, an acquisition unit 11, a frame extraction unit 12, a complementary image generation unit 13, a video generation unit 14, and a rendering unit 15. In FIG. 1, the video generation device 1 includes, as functional components, an acquisition unit 11, a frame extraction unit 12, a complementary image generation unit 13, a video generation unit 14, and a rendering unit 15, which are realized by a single video generation device 1, but these functional components may be distributed across multiple devices.

[0013] Each functional element of the image generation device 1 can access the storage unit 20. The storage unit 20 is a storage device that stores a plurality of captured images Xj (j = 1, 2, ..., L) and shooting viewpoint information indicating the shooting viewpoints Cj (j = 1, 2, ..., L) of the plurality of cameras 2. The storage unit 20 may be disposed in the image generation device 1 as shown in FIG. 1 , or may be disposed outside the image generation device 1 so as to be accessible from the image generation device 1.

[0014] The multiple captured images Xj are two-dimensional video data captured by multiple cameras 2. Each of the multiple captured images Xj includes multiple frame images Fn (n=1, 2, ..., N) arranged in time series. If the captured images Xj are color images, each pixel of the frame image Fn includes information related to RGB pixel values.

[0015] The shooting viewpoint Cj indicates the viewpoint of each camera 2. More specifically, the shooting viewpoint Cj is a parameter representing the position information and attitude information of the camera 2. The position information of the camera 2 indicates the position of the camera 2 in three-dimensional space. The attitude information of the camera 2 indicates, for example, the roll angle, pitch angle, and yaw angle of the camera 2.

[0016] The acquisition unit 11 acquires, from the storage unit 20, a plurality of captured images Xj captured by a plurality of cameras 2 and information indicating the capturing viewpoints Cj corresponding to the plurality of cameras 2. The acquisition unit 11 outputs the acquired plurality of captured images Xj and information indicating the capturing viewpoints Cj to the frame extraction unit 12.

[0017] The frame extraction unit 12 extracts a plurality of frame images Fn from each of the plurality of captured videos Xj. The frame extraction unit 12 outputs a frame image F1 at time t1 (first time point) from the plurality of frame images Fn extracted from each captured video Xj and information indicating the shooting viewpoint Cj to the complementary image generation unit 13 and the video generation unit 14. The frame image F1 at time t1 is, for example, the first frame in the chronological order of the captured video Xj. The frame extraction unit 12 also stores the plurality of frame images Fn of each captured video Xj in the storage unit 20.

[0018] The complementary image generation unit 13 generates a complementary image Fc viewed from a complementary viewpoint Cc based on a plurality of frame images F1 at time t1 extracted from the plurality of captured videos Xj. The complementary viewpoint Cc is one or a plurality of viewpoints virtually set at positions different from the capturing viewpoints Cj of the plurality of cameras 2. As shown in FIG. 2 , the complementary viewpoint generation unit 16 sets a complementary viewpoint Cc at a position between the capturing viewpoints Cj of the plurality of cameras 2 in order to complement blind spots of the subject 3, and outputs complementary viewpoint information indicating the complementary viewpoint Cc to the complementary image generation unit 13.

[0019] The complementary image generation unit 13 receives the multiple frame images F1 at time t1, the shooting viewpoint information, and the complementary viewpoint information, and generates a complementary image Fc using a complementary image generation model 13a that has been machine-trained to output a complementary image Fc corresponding to the complementary viewpoint Cc. Note that the complementary image generation unit 13 may acquire multiple complementary viewpoints Cc and generate multiple complementary images Fc corresponding to the multiple complementary viewpoints Cc.

[0020] The complementary image generation model 13a is an AI model that generates multiple complementary images Fc corresponding to multiple complementary viewpoints Cc. The complementary image generation model 13a may be stored in the image generation device 1, or may be stored in another device connected to the image generation device 1 via a network and configured to be able to exchange information with the image generation device 1.

[0021] The complementary image generation model 13a is constructed by using a diffusion model to learn training data including a plurality of learning frame images captured by a plurality of cameras 2, information indicating the shooting viewpoints Cj of the plurality of cameras 2, and information indicating the complementary viewpoint Cc. The complementary image generation unit 13 outputs the generated one or more complementary images Fc and complementary viewpoint information indicating the complementary viewpoint Cc to the video generation unit 14.

[0022] The video generation unit 14 generates a free-viewpoint video Y based on a plurality of frame images Fn at time tn (n=1, 2, ..., N), shooting viewpoint information, complementary image Fc, and complementary viewpoint information. The free-viewpoint video Y is a three-dimensional model of a stereoscopic subject 3 created on a computer, from which the viewpoint can be arbitrarily selected. The free-viewpoint video Y includes a plurality of free-viewpoint images Yn (n=1, 2, ..., N) arranged in chronological order.

[0023] 1 , the video generation unit 14 includes a first reconstructor 21 and a second reconstructor 22. The first reconstructor 21 generates a free-viewpoint image Y1 at time t1 based on a plurality of frame images F1, shooting viewpoint information, a complementary image Fc, and complementary viewpoint information at time t1. The free-viewpoint image Y1 at time t1 is the first free-viewpoint image in the time series of the free-viewpoint video Y.

[0024] The first reconstruction unit 21 receives a plurality of frame images F1 at time t1, shooting viewpoint information, complementary image Fc, and complementary viewpoint information, and generates a free-viewpoint image Y1 using a first image generation model 21a that has been machine-learned to output a free-viewpoint image Y1. The first image generation model 21a is an AI model for generating the free-viewpoint image Y1, and is constructed by inputting training data including a plurality of frame images for learning at time t1 and information indicating the shooting viewpoint into a neural network and causing it to learn.

[0025] The second reconstructor 22 generates a free viewpoint image Yi (i=2, ..., N) at a time ti (i=2, ..., N) after the time t1, based on the free viewpoint image Y1 at the time t1 generated by the first reconstructor 21, a plurality of frame images Fi at a time ti (i=2, ..., N) (second time point) after the time t1, and shooting viewpoint information. The free viewpoint image Yi at the time ti is a free viewpoint image at a time later in the time series than the free viewpoint image Y1.

[0026] The second reconstruction unit 22 receives a free-viewpoint image Y1, a plurality of frame images Fi at time ti, shooting viewpoint information, and time ti, and generates a free-viewpoint image Yi using a second image generation model 22a that has been machine-learned to output the free-viewpoint image Yi. The second image generation model 22a is an AI model that predicts the movement direction, movement amount, and deformation amount of the subject 3 included in the free-viewpoint image Y1. The second image generation model 22a is constructed by inputting training data including a free-viewpoint image for learning at time t1, time ti, a plurality of frame images Fi at time ti, and shooting viewpoint information into a neural network and allowing the neural network to learn. For example, the second image generation model 22a is trained to output a free-viewpoint image Yi at time ti using a free-viewpoint image Y(i-1) at time t(i-1).

[0027] The second reconstructor 22 generates a free viewpoint video Y by chronologically arranging the free viewpoint image Y1 generated by the first reconstructor 21 and the free viewpoint image Yn including the free viewpoint image Yi generated by the second reconstructor 22. The second reconstructor 22 outputs the generated free viewpoint video Y to the rendering unit 15.

[0028] The rendering unit 15 generates a rendering image Fr of the subject 3 viewed from an arbitrary viewpoint using the free viewpoint video Y. For example, the rendering unit 15 receives information indicating the rendering viewpoint Cr from the rendering viewpoint setting unit 18, and generates a video including time-series rendering images corresponding to the rendering viewpoint Cr using the free viewpoint video Y.

[0029] The rendering viewpoint Cr is specified by the user. That is, the video including the time-series rendering images is a video of the subject 3 viewed from the viewpoint specified by the user. The rendering viewpoint Cr may be a viewpoint different from the shooting viewpoints Cj of the multiple cameras 2. For example, the rendering viewpoint setting unit 18 receives information indicating the viewpoint from a user terminal (not shown) and outputs the received viewpoint as the rendering viewpoint Cr to the rendering unit 15. The rendering unit 15 generates a video including the rendering image corresponding to the rendering viewpoint Cr and transmits the video to the user terminal.

[0030] The first reconstructor 21 and the second reconstructor 22 may be a single functional element. In this case, the first image generation model 21 a and the second image generation model 22 a may be constructed as a single AI model and designed to behave differently at time t1 than at time t2 and thereafter.

[0031] In one embodiment, the image generation device 1 may further include a comparison unit 24 and a learning unit 25. The comparison unit 24 compares the complement image Fc generated by the complement image generation unit 13 with the frame image F1 to determine the difference between them. For example, the complement image generation unit 13 selects one viewpoint Cs from the shooting viewpoints Cj of the multiple cameras 2, and generates a complement image Fc corresponding to the viewpoint Cs. The comparison unit 24 calculates the difference between the complement image Fc corresponding to the viewpoint Cs and the frame image F1 corresponding to the viewpoint Cs.

[0032] For example, the comparison unit 24 calculates the mean squared error (MSE) between pixel values ​​of the complemented image Fc corresponding to the viewpoint Cs and pixel values ​​of the frame image F1 corresponding to the viewpoint Cs. The comparison unit 24 outputs the mean squared error to the learning unit 25 as information indicating the comparison result between the complemented image Fc and the frame image F1.

[0033] The learning unit 25 trains the complement image generation model 13a according to the comparison result between the complement image Fc and the frame image F1. For example, the learning unit 25 generates parameter update information based on the mean square error calculated by the comparison unit 24, and updates the parameters (weighting) of the neural network of the complement image generation model 13a using the parameter update information so as to reduce the mean square error. For example, the learning unit 25 updates the parameters of the neural network using an error backpropagation method that backpropagates the mean square error from the output layer to the input layer of the complement image generation model 13a.

[0034] Furthermore, the comparison unit 24 compares the rendering image Fr generated by the rendering unit 15 with the frame image Fn to determine the difference between them. For example, the rendering unit 15 selects one viewpoint Cs from the shooting viewpoints Cj of the multiple cameras 2, and generates a rendering image Fr corresponding to the viewpoint Cs using the free viewpoint image Y1 at time t1 generated by the first reconstruction unit 21. The comparison unit 24 calculates the difference between the rendering image Fr corresponding to the viewpoint Cs and the frame image F1 at time t1 corresponding to the viewpoint Cs.

[0035] For example, the comparison unit 24 calculates the mean square error between pixel values ​​of the rendering image Fr corresponding to the viewpoint Cs and pixel values ​​of the frame image F1 corresponding to the viewpoint Cs. The comparison unit 24 outputs the mean square error to the learning unit 25 as information indicating the comparison result between the rendering image Fr and the frame image F1.

[0036] The learning unit 25 trains the first image generation model 21a according to the comparison result between the rendering image Fr and the frame image F1. For example, the learning unit 25 generates parameter update information based on the mean squared error calculated by the comparison unit 24, and uses the parameter update information to update the parameters (weighting) of the neural network of the first image generation model 21a so as to reduce the mean squared error. For example, the learning unit 25 updates the parameters of the neural network using an error backpropagation method that backpropagates the mean squared error from the output layer to the input layer of the first image generation model 21a.

[0037] Similarly, the rendering unit 15 generates a rendering image Fr corresponding to the viewpoint Cs using the free viewpoint image Yi at time ti generated by the second reconstruction unit 22. Then, the comparison unit 24 calculates the difference between the rendering image Fr at time ti corresponding to the viewpoint Cs and a frame image Fi (i=2, ..., N) at time ti corresponding to the viewpoint Cs.

[0038] For example, the comparison unit 24 calculates the mean square error between the pixel values ​​of the rendering image Fr at the time ti corresponding to the viewpoint Cs and the pixel values ​​of the rendering image Fr at the time ti corresponding to the viewpoint Cs. The comparison unit 24 outputs the mean square error to the learning unit 25 as information indicating the comparison result between the rendering image Fr and the frame image Fi.

[0039] The learning unit 25 trains the second image generation model 22a according to the comparison result between the rendering image Fr and the frame image Fi. For example, the learning unit 25 generates parameter update information based on the mean squared error calculated by the comparison unit 24, and updates the parameters (weighting) of the neural network of the second image generation model 22a using the parameter update information so as to reduce the mean squared error. For example, the learning unit 25 updates the parameters of the neural network using an error backpropagation method that backpropagates the mean squared error from the output layer to the input layer of the second image generation model 22a.

[0040] As described above, the video generation device 1 generates a complementary image Fc of the subject 3 viewed from a complementary viewpoint Cc based on a plurality of frame images F1 at time t1. As described above, by generating a free-viewpoint image Y1 using the complementary image Fc in addition to the plurality of frame images F1 at time t1, blind spots of the subject 3 can be filled in by the complementary image Fc, and therefore a clear free-viewpoint video can be generated even when the number of cameras is small.

[0041] Furthermore, the video generation device 1 generates a complementary image Fc only at time t1, and does not generate a complementary image Fc at any time after time t1. This is because, if there are images of the subject 3 seen from a sufficient number of viewpoints at time t1, it is possible to generate free viewpoint images Yi from time t2 onwards by predicting the movement of the subject 3 from time t2 onwards. In other words, since there is no need to generate a complementary image at every time tn, the processing load can be reduced.

[0042] Furthermore, in the video generation device 1, the first image generation model 21a and the second image generation model 22a are trained based on the comparison results between the rendering image Fr generated by the rendering unit 15 and the frame image Fn, so that a highly reproducible free viewpoint video Y can be generated.

[0043] Next, an image generation method according to an embodiment will be described. This image generation method is executed by the above-described image generation device 1. Fig. 3 is a flowchart showing the image generation method according to an embodiment.

[0044] As shown in Figure 3, in one embodiment of the video generation method, the acquisition unit 11 of the video generation device 1 acquires multiple captured images Xj captured by multiple cameras 2 and shooting viewpoint information indicating the shooting viewpoints Cj of the multiple cameras 2 (step ST1).

[0045] Next, the frame extracting unit 12 extracts a plurality of frame images Fn at each time tn from the plurality of captured videos Xj (step ST2).

[0046] Next, the complementary image generating unit 13 generates a complementary image Fc of the subject 3 viewed from the complementary viewpoint Cc based on a plurality of frame images F1 at time t1 among the plurality of frame images Fn and the photographing viewpoint Cj (step ST3).

[0047] Next, the first reconstruction unit 21 of the video generation unit 14 inputs multiple frame images F1, information indicating multiple shooting viewpoints Cj, complementary image Fc, and information indicating complementary viewpoint Cc into the first image generation model 21a to generate a free viewpoint image Y1 at time t1 (step ST4).

[0048] Next, the second reconstruction unit 22 of the video generation unit 14 inputs the free-viewpoint image Y(i-1) at time t(i-1), the multiple frame images Fi, the shooting viewpoint information, and time ti (i=2, 3, ..., N) into the second image generation model 22a to generate a free-viewpoint image Yi at time ti (step ST5). The generation of the free-viewpoint image Yi is repeated until i=N (steps ST6 and ST9).

[0049] Next, the video generator 14 arranges the free viewpoint image Y1 and the free viewpoint image Yi, that is, the free viewpoint image Yn, in time series and outputs the result as a free viewpoint video Y (step ST7).

[0050] Next, the rendering unit 15 generates a rendering video corresponding to the rendering viewpoint Cr specified by the user using the free viewpoint video Y (step ST8). The rendering video is generated by generating a rendering image Fr corresponding to the rendering viewpoint Cr at each time tn using the free viewpoint video Yn and arranging the rendering images Fr in chronological order. The generated rendering video is transmitted to a user terminal (not shown).

[0051] The image generation device 1 and image generation method of the present disclosure may have the following configuration.

[0052] [1] An image generation device comprising: an acquisition unit that acquires a plurality of captured images captured by a plurality of cameras and image capture viewpoint information indicating the image capture viewpoints of the plurality of cameras; a frame extraction unit that extracts a plurality of frame images at a first time point from the plurality of captured images; a complementary image generation unit that generates a complementary image viewed from a complementary viewpoint based on the plurality of frame images at the first time point; and an image generation unit that generates a free viewpoint image at the first time point based on the plurality of frame images at the first time point, the image capture viewpoint information, the complementary image, and the complementary viewpoint information indicating the complementary viewpoint, and that generates a free viewpoint image at a second time point after the first time point based on the free viewpoint image at the first time point.

[0053] [2] The video generation device according to [1], wherein the complementary image generation unit receives the plurality of frame images at the first time point, the shooting viewpoint information, and the complementary viewpoint information, and generates the complementary image using a complementary image generation model that has been machine-learned to output the complementary image corresponding to the complementary viewpoint.

[0054] [3] The video generation device according to [2], further comprising a learning unit that trains the complementary image generation model based on a comparison result between the complementary image and a frame image corresponding to the complementary viewpoint among the plurality of frame images at the first time point.

[0055] [4] The video generation device according to any one of [1] to [3], wherein the frame extraction unit extracts a plurality of frame images at a second time point from the plurality of captured videos, and the video generation unit includes: a first reconstruction unit that receives the plurality of frame images, the shooting viewpoint information, the complementary image, and the complementary viewpoint information at the first time point, and generates a free-viewpoint image at the first time point using a first image generation model that has been machine-learned to output a free-viewpoint image at the first time point; and a second image generation unit that receives the free-viewpoint image at the first time point, information indicating the second time point, the plurality of frame images at the second time point, and the shooting viewpoint information, and generates a free-viewpoint image at the second time point using a second image generation model that has been machine-learned to output a free-viewpoint image at the second time point.

[0056] [5] The image generating device according to [4], further comprising a rendering unit that generates a rendering image corresponding to a rendering viewpoint using the free viewpoint image at the first time point.

[0057] [6] The image generation device according to [5], further comprising a learning unit that learns the first image generation model based on a comparison result between the rendering image and a frame image corresponding to the rendering viewpoint among the plurality of frame images at the first time point.

[0058] [7] The image generating device according to any one of [4] to [6], further comprising a rendering unit that generates a rendering image corresponding to a rendering viewpoint using the free viewpoint image at the second time point.

[0059] [8] The frame extraction unit extracts a plurality of frame images at the second time point from the plurality of captured videos, and the image generation device further includes a learning unit that trains the second image generation model based on a comparison result between the rendering image and a frame image corresponding to the rendering viewpoint among the plurality of frame images at the second time point.

[0060] [9] An image generation method including the steps of: acquiring a plurality of captured images captured by a plurality of cameras and camera viewpoint information indicating the camera viewpoints of the plurality of cameras; extracting a plurality of frame images at a first time point from the plurality of captured images; generating a complementary image viewed from a complementary viewpoint based on the plurality of frame images at the first time point; generating a free viewpoint image at the first time point based on the plurality of frame images at the first time point, the camera viewpoint information, the complementary image, and complementary viewpoint information indicating the complementary viewpoint; and generating a free viewpoint image at a second time point after the first time point based on the free viewpoint image at the first time point.

[0061] The block diagram shown in FIG. 1 shows functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (e.g., via wire, wirelessly, etc.) and these multiple devices. The functional block may be realized by combining software with the single device or multiple devices.

[0062] Functions include, but are not limited to, judgment, determination, assessment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.

[0063] For example, the image generation device 1 according to an embodiment of the present invention may function as a computer. Fig. 4 is a diagram showing an example of the hardware configuration of the image generation device 1. The image generation device 1 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.

[0064] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the image generation apparatus 1 may be configured to include one or more of the devices shown in Fig. 4, or may be configured to exclude some of the devices.

[0065] Each function in the video generation device 1 is realized by loading specific software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations and control communication via the communication device 1004 and the reading and / or writing of data in the memory 1002 and storage 1003.

[0066] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, each of the functional elements shown in FIG. 1 may be realized by the processor 1001.

[0067] The processor 1001 also reads programs (program codes), software modules, and data from the storage 1003 and / or the communication device 1004 into the memory 1002 and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, each functional element of the image generation device 1 may be implemented by a control program stored in the memory 1002 and running on the processor 1001. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented on one or more chips. The programs may also be transmitted from a network via a telecommunications line.

[0068] The memory 1002 is a computer-readable recording medium and may be composed of at least one of, for example, a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), and a random access memory (RAM). The memory 1002 may also be called a register, a cache, a main memory (primary storage device), or the like. The memory 1002 can store executable programs (program codes), software modules, and the like for implementing an information processing method according to one embodiment of the present invention.

[0069] Storage 1003 is a computer-readable recording medium, and may be, for example, at least one of an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including memory 1002 and / or storage 1003.

[0070] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via a wired and / or wireless network, and is also called, for example, a network device, a network controller, a network card, or a communication module.

[0071] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. The input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).

[0072] Furthermore, each device such as the processor 1001 and the memory 1002 is connected to a bus 1007 for communicating information. The bus 1007 may be configured as a single bus, or may be configured as different buses between the devices.

[0073] Furthermore, the image generating device 1 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented by at least one of these pieces of hardware.

[0074] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.

[0075] Each aspect / embodiment described in the present disclosure may be applied to at least one of systems using LTE (Long Term Evolution), LTE-Advanced (LTE-A), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (New Radio), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, UWB (Ultra-Wide Band), Bluetooth (registered trademark), or other suitable systems, and next-generation systems enhanced based on these. Furthermore, a combination of multiple systems (e.g., a combination of at least one of LTE and LTE-A with 5G, etc.) may also be applied.

[0076] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.

[0077] In the present disclosure, a specific operation described as being performed by a base station may be performed by its upper node in some cases. In a network consisting of one or more network nodes having a base station, it is clear that various operations performed for communication with a terminal may be performed by at least one of the base station and another network node other than the base station (for example, an MME or an S-GW, etc., but are not limited to these). Although the above example illustrates a case where there is one other network node other than the base station, a combination of multiple other network nodes (for example, an MME and an S-GW) may also be used.

[0078] Information etc. may be output from a higher layer (or a lower layer) to a lower layer (or a higher layer), or may be input / output via multiple network nodes.

[0079] Input and output information may be stored in a specific location (for example, memory) or managed in a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.

[0080] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).

[0081] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).

[0082] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.

[0083] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0084] Software, instructions, etc. may also be transmitted or received over a transmission medium. For example, if the software is transmitted from a website, server, or other remote source using wired technologies such as coaxial cable, fiber optic cable, twisted pair, and Digital Subscriber Line (DSL), and / or wireless technologies such as infrared, radio, and microwave, these wired and / or wireless technologies are included within the definition of transmission media.

[0085] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0086] It should be noted that terms explained in this disclosure and / or terms necessary for understanding this specification may be replaced with terms having the same or similar meanings.

[0087] As used in this disclosure, the terms "system" and "network" are used interchangeably.

[0088] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed as absolute values, relative values ​​from a predetermined value, or other corresponding information. For example, a radio resource may be indicated by an index.

[0089] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.

[0090] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.

[0091] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly specified otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."

[0092] When designations such as "first," "second," etc. are used in this disclosure, any reference to an element does not generally limit the quantity or order of those elements. These designations may be used herein as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed therein or that the first element must precede the second element in some way.

[0093] The "means" in the configuration of each of the above devices may be replaced with "part," "circuit," "device," etc.

[0094] To the extent that the terms "include," "including," and variations thereof are used herein or in the claims, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, the term "or," as used herein or in the claims, is not intended to be an exclusive or.

[0095] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.

[0096] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."

[0097] The image generating device 1 and the information processing method according to the present disclosure may have the following configuration.

[0098] 1...image generation device, 2...camera, 11...acquisition unit, 12...frame extraction unit, 13...complementary image generation unit, 13a...complementary image generation model, 14...image generation unit, 15...rendering unit, 21...first reconstruction unit, 21a...first image generation model, 22a...second image generation model, 25...learning unit, Cc...complementary viewpoint, Cj...shooting viewpoint, Cr...rendering viewpoint, F1...frame image, Fc...complementary image, Fn...frame image, Fr...rendering image, tn...time, Xj...shooting image, Y...free viewpoint image, Yn...free viewpoint image.

Claims

1. A video generation device comprising: an acquisition unit that acquires multiple captured images captured by multiple cameras and camera viewpoint information indicating the camera viewpoints of the multiple cameras; a frame extraction unit that extracts multiple frame images at a first time point from the multiple captured images; a complementary image generation unit that generates a complementary image viewed from a complementary viewpoint based on the multiple frame images at the first time point; and a video generation unit that generates a free viewpoint image at the first time point based on the multiple frame images at the first time point, the camera viewpoint information, the complementary image, and the complementary viewpoint information indicating the complementary viewpoint, and that generates a free viewpoint image at a second time point later than the first time point based on the free viewpoint image at the first time point.

2. The video generation device described in claim 1, wherein the complementary image generation unit inputs the plurality of frame images at the first time point, the shooting viewpoint information, and the complementary viewpoint information, and generates the complementary image using a complementary image generation model that has been machine-trained to output the complementary image corresponding to the complementary viewpoint.

3. The video generation device described in claim 2, further comprising a learning unit that trains the complementary image generation model based on the comparison result between the complementary image and a frame image corresponding to the complementary viewpoint among the plurality of frame images at the first time point.

4. The video generation device of claim 1, wherein the frame extraction unit extracts a plurality of frame images at the second time point from the plurality of captured videos, and the video generation unit includes: a first reconstruction unit that receives the plurality of frame images, the shooting viewpoint information, the complementary image, and the complementary viewpoint information at the first time point, and generates a free-viewpoint image at the first time point using a first image generation model that has been machine-learned to input the plurality of frame images, the shooting viewpoint information, the complementary image, and the complementary viewpoint information, and output a free-viewpoint image at the first time point; and a second image generation unit that receives the free-viewpoint image at the first time point, information indicating the second time point, the plurality of frame images at the second time point, and the shooting viewpoint information, and generates a free-viewpoint image at the second time point using a second image generation model that has been machine-learned to output a free-viewpoint image at the second time point.

5. The image generating device according to claim 4, further comprising a rendering unit that generates a rendering image corresponding to a rendering viewpoint using the free viewpoint image at the first time point.

6. The image generation device described in claim 5, further comprising a learning unit that learns the first image generation model based on the comparison result between the rendering image and a frame image corresponding to the rendering viewpoint among the plurality of frame images at the first time point.

7. The image generating device according to claim 4, further comprising a rendering unit that generates a rendering image corresponding to a rendering viewpoint using the free viewpoint image at the second time point.

8. The image generation device described in claim 7, wherein the frame extraction unit extracts multiple frame images at the second time point from the multiple captured images, and the image generation device further includes a learning unit that learns the second image generation model based on a comparison result between the rendering image and a frame image corresponding to the rendering viewpoint among the multiple frame images at the second time point.

9. A video generation method comprising the steps of: acquiring a plurality of captured videos taken by a plurality of cameras and camera viewpoint information indicating the camera viewpoints of the plurality of cameras; extracting a plurality of frame images at a first time point from the plurality of captured videos; generating a complementary image viewed from a complementary viewpoint based on the plurality of frame images at the first time point; generating a free viewpoint image at the first time point based on the plurality of frame images at the first time point, the camera viewpoint information, the complementary image, and complementary viewpoint information indicating the complementary viewpoint; and generating a free viewpoint image at a second time point after the first time point based on the free viewpoint image at the first time point.

Citation Information

Patent Citations

  • Apparatus and method for processing image, apparatus and method for creating complement image, program, and storage medium

    JP2012244527A

  • Information processing device, information processing method, and computer program

    JP2023183059A

  • Information processing device and information processing method

    WO2022224964A1