Information processing device and method

By using rolling shutter cameras and synchronizing shooting times within frames, the method addresses the cost and complexity of synchronized camera setups, enabling accurate three-dimensional reconstruction and free viewpoint image generation from asynchronous video data.

WO2025154695A1PCT designated stage expired Publication Date: 2025-07-24PREFERRED NETWORKS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/000780
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2025-01-14
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing methods for training a restoration model to generate free viewpoint images require synchronized shooting times across multiple cameras, which is costly and impractical due to the need for expensive equipment and synchronization techniques.

Method used

A method that utilizes rolling shutter cameras to capture asynchronous video data, estimates and synchronizes shooting times within frames, and trains a neural network model to generate detailed free viewpoint images based on this data, reducing costs and equipment complexity.

Benefits of technology

Enables accurate three-dimensional reconstruction of scenes from multiple viewpoints using asynchronous video data, allowing for cost-effective generation of detailed free viewpoint images without the need for expensive synchronization equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025000780_24072025_PF_FP_ABST
    Figure JP2025000780_24072025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device comprises one or more memories and one or more processors. The one or more processors acquire a moving image of a scene captured by using a rolling shutter method for each viewpoint, and train, on the basis of the moving image captured at each viewpoint and information on distortion due to the rolling shutter method when the moving image is captured, a model for three-dimensionally reconstructing a scene including a temporal change during capturing of the moving image.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device and method

[0001] The present disclosure relates to an information processing device and method.

[0002] 2. Description of the Related Art Image generation techniques are known that use a plurality of image capture devices to capture images of the same scene from different viewpoints, and reconstruct the scene based on the captured images.

[0003] For example, Non-Patent Document 1 discloses a restoration model called NeRF (Neural Radiance Fields). A restoration model to which the NeRF technology is applied makes it possible to generate a free viewpoint image.

[0004] Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, Ren Ng, "NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis," Computer Vision - ECCV 2020: 16th European Conference, Proceedings, Part I, Pages 405-421, Aug 2020.

[0005] An object of the present disclosure is to provide a technology capable of accurately reconstructing a scene captured from multiple viewpoints in three dimensions.

[0006] An information processing device according to one aspect of the present disclosure includes one or more memories and one or more processors, and the one or more processors acquire video of a scene captured using a rolling shutter method for each viewpoint, and train a model that three-dimensionally reconstructs the scene, including changes over time during video capture, based on the video captured from each viewpoint and information regarding distortion caused by the rolling shutter method when the video was captured.

[0007] FIG. 1 is a diagram showing an example of a training process for a restoration model. FIG. 2 is a diagram showing an example of a generation process using a restoration model. FIG. 3 is a diagram showing an example of video data in which the shooting times are synchronized. FIG. 4 is a diagram showing an example of video data in which the shooting times are not synchronized. FIG. 5 is a block diagram showing an example of the overall configuration of a video generation system. FIG. 6 is a block diagram showing an example of the functional configuration of a generation device. FIG. 7 is a diagram showing an example of shooting time information after synchronization. FIG. 8 is a flowchart showing an example of a training process. FIG. 9 is a flowchart showing an example of a generation process. FIG. 10 is a block diagram showing an example of the hardware configuration of a generation device.

[0008] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.

[0009] [Embodiment] One embodiment of the present disclosure is a video generation system that generates video data. In this embodiment, the video generation system has a function of generating video data including time-varying changes in a scene viewed from an arbitrary viewpoint. The video data generated by the video generation system includes a plurality of free-viewpoint images that are continuous in the time direction. In other words, the video data is an example of time-series data of free-viewpoint images.

[0010] <Overview of Reconstruction Model> The video generation system may generate time-series data of free-viewpoint images based on a trained reconstruction model. The reconstruction model is trained based on multiple video data of the same scene captured from different viewpoints, and may generate images showing the scene as seen from any viewpoint at any time. In this embodiment, the reconstruction model may be, for example, a neural network to which NeRF technology is applied.

[0011] 1 is a diagram illustrating an example of the training process of the restoration model. As shown in FIG. 1, in the training process of the restoration model, for each pixel of each of a plurality of captured images V1 in which a scene is captured from a plurality of viewpoints, a combination (x, y, z), viewpoint information (θ, φ), and time information t (x, y, z, θ, φ, t) is used to train the restoration model F. θ is entered into

[0012] The coordinate information is information that specifies the coordinates of three-dimensional points in a scene. In FIG. 1, coordinates are expressed in a Cartesian coordinate system as an example, but the three-dimensional points may be specified in any coordinate system. The viewpoint information is information that specifies a direction vector that represents the line of sight from the viewpoint toward the three-dimensional point. In FIG. 1, directions are expressed in a spherical coordinate system as an example, but the line of sight may be specified in any coordinate system. The time information is information that specifies the time of the scene. For example, the time information may be the elapsed time from an arbitrary time origin, and the time origin may be the earliest shooting time among the shooting times of each of the captured images V1.

[0013] Restoration Model F θ The system includes an encoder 1 and a renderer 2. The encoder 1 generates a reconstruction model F θ The renderer 2 outputs a feature value h based on a combination (x, y, z, θ, φ, t) of coordinate information, viewpoint information, and time information input to the encoder 1. The renderer 2 outputs a combination (R, G, B) of color (R, G, B) and opacity σ of the three-dimensional point based on the feature value h output from the encoder 1. The renderer 2 may be a decoder that outputs a combination of color and opacity based on the feature value. That is, the restored model F θ calculates the color and opacity of a 3D point at a certain viewpoint and time.

[0014] In the training process of the restoration model, the restoration model F θ The same process is performed for a plurality of viewpoints and a plurality of time information for the restored model F. θ outputs a plurality of combinations of color and opacity for each of a plurality of three-dimensional points on each line of sight for a combination of viewpoint information and time information.

[0015] Restoration Model F θ A volume rendering process 3 is performed on the multiple combinations of color and opacity output from the volume rendering process 3. The volume rendering process 3 calculates the color of each pixel included in an image seen from a certain viewpoint at a certain time using a volume rendering method. Specifically, the volume rendering process 3 calculates the color of each pixel included in an image seen from a certain viewpoint at a certain time using a restored model F. θ The color of each pixel is calculated by performing a predetermined product-sum operation based on the color and opacity output from the volume rendering process 3. As a result, the volume rendering process 3 generates a plurality of free viewpoint images V2 corresponding to combinations of viewpoint information and time information.

[0016] For the multiple free viewpoint images V2 generated by the volume rendering process 3, a loss calculation process 4 is performed for each viewpoint at each time. The loss calculation process 4 calculates an error at a given viewpoint at a given time by comparing a free viewpoint image V2 corresponding to a given viewpoint at a given time with a captured image V1 of a scene captured from that viewpoint at that time. The error may be, for example, the average of the squared errors of the R value, G value, and B value of each pixel, or some other value, as long as an appropriate value can be defined as the error. In this way, the loss calculation process 4 calculates the error at each of the multiple viewpoints at each time.

[0017] In the training process of the restoration model, the restoration model F is trained based on the error calculated in the loss calculation process 4. θ The model parameters of the restored model F are updated. θ For example, the model parameters of the reconstruction model F may be updated based on the backpropagation method. In the reconstruction model training process, the model parameters may be repeatedly updated until a predetermined convergence condition is satisfied. θ The model parameters of the trained reconstruction model F are updated. θ is generated.

[0018] 2 is a diagram illustrating an example of the generation process using the restoration model. As shown in FIG. 2, in the generation process using the restoration model, a combination (x, y, z), viewpoint information (θ, φ), and time information t (x, y, z, θ, φ, t) is generated as a trained restoration model F θ is entered into

[0019] Trained reconstruction model F θ The combination of coordinate information, viewpoint information, and time information (x, y, z, θ, φ, t) input to the trained encoder 1 is input to the trained encoder 1. The trained encoder 1 outputs a feature value h based on the combination of coordinate information, viewpoint information, and time information (x, y, z, θ, φ, t). The trained renderer 2 outputs a combination of color and opacity (R, G, B, σ) of the 3D point based on the feature value h output from the trained encoder 1. The trained reconstruction model F θ For the combination of color and opacity of each 3D point output from the 3D image processing unit 10, volume rendering processing 3 is performed for each pixel of the free viewpoint image V2 corresponding to a specified viewpoint at a specified time, thereby generating the free viewpoint image V2 corresponding to a certain viewpoint at a certain time.

[0020] In the generation process using the reconstruction model, a trained reconstruction model F θ By sequentially inputting different time information t to the image data processor 10, a plurality of free viewpoint images V2 that are continuous in the time direction are generated. This generates time-series data of the free viewpoint images V2.

[0021] <<Capturing Method of Video Data>> The captured images V1 used in the training process of the restoration model may be captured by multiple image capturing devices. The multiple image capturing devices may be arranged around the subject so as to capture the same scene from different viewpoints. The captured images V1 may be each frame included in the video data captured by the image capturing devices. One frame may be interpreted as one image.

[0022] Conventionally, in order to train a restoration model based on multiple video data captured by multiple imaging devices, it was necessary to synchronize the capture times of each frame between the video data. Furthermore, it was necessary to synchronize the capture times of each pixel included in each frame of each video data. If the capture times of each frame or each pixel are not synchronized, the position of the subject at the time of capture will be different in each frame or each pixel, for example, if the subject is moving. Therefore, a restoration model trained based on video data whose capture times are not synchronized may have problems such as distortion in the free-viewpoint image or inability to correctly restore the scene.

[0023] To synchronize the capture time of each frame of video data, for example, multiple image capture devices are connected with a synchronization cable and controlled to capture the images simultaneously. Furthermore, to synchronize the capture time of each pixel within a frame, for example, a global shutter camera is used. However, equipment for synchronizing multiple image capture devices and global shutter cameras are expensive and not easily introduced.

[0024] Fig. 3 is a diagram showing an example of video data whose capture times are synchronized. In the example shown in Fig. 3, video data α is video data captured by a first image capture device, and video data β is video data captured by a second image capture device. The first image capture device and the second image capture device are both global shutter cameras, and are connected by a synchronization cable so that the capture timing of each frame is synchronized.

[0025] As shown in Fig. 3, in video data with synchronized shooting times, for each frame (frames 1 to 5) included in each video data (video data α to β), multiple pixels (pixels A to D) included in the frame with the same frame number all have the same shooting time. Note that in the example shown in Fig. 3, the shooting time is the elapsed time [seconds] from the start of shooting.

[0026] Fig. 4 shows an example of video data in which the capture times are not synchronized. In the example shown in Fig. 4, the first and second image capture devices are both rolling shutter cameras, and each captures each frame at an independent capture time.

[0027] As shown in Figure 4, in video data whose shooting times are not synchronized, for each frame (frames 1 to 5) included in each video data (video data α to β), the shooting times differ between frames with the same frame number, and the shooting times differ between multiple pixels (pixels A to B or pixels C to D) included in the same frame.

[0028] In this embodiment, a high-resolution free-viewpoint image can be generated based on multiple video data captured asynchronously. If multiple video data captured asynchronously can be used, for example, an inexpensive rolling shutter camera can be used to capture each video data independently (without synchronizing the capture timing). In one aspect, this embodiment reduces the cost of training a restoration model capable of generating free-viewpoint images.

[0029] <Overall Configuration of Moving Image Generation System> The overall configuration of the moving image generation system in this embodiment will be described with reference to Fig. 5. Fig. 5 is a block diagram showing an example of the overall configuration of the moving image generation system.

[0030] 5, the video generation system 1000 includes a plurality of image capture devices 10 (10-1 to 10-3), a generation device 20, and a terminal device 30. Although three image capture devices 10-1 to 10-3 are shown in FIG. 5 as an example, the number of image capture devices 10 is not limited.

[0031] Each of the multiple imaging devices 10 may be electrically connected to the generating device 20. The generating device 20 and the terminal device 30 may be connected to each other so as to be able to communicate data with each other via a communication network such as a local area network (LAN) or the Internet.

[0032] The imaging device 10 is an optical device that captures images of a subject S from a predetermined viewpoint at a predetermined frame rate. Multiple imaging devices 10 may be arranged around the subject S so as to capture the subject S from different viewpoints. The subject S is an object or space captured by the imaging device 10. In FIG. 5 , the subject S is depicted as a single person as an example, but the type and number of subjects are not limited. For example, the number of subjects S may increase or decrease over time, or may be replaced.

[0033] In this embodiment, the imaging device 10 does not need to synchronize the imaging time of each pixel in each frame of video data obtained by imaging. For example, the imaging device 10 may be a rolling shutter camera. Furthermore, the imaging timing of frames of multiple imaging devices 10 does not need to be synchronized with each other.

[0034] The generating device 20 is an information processing device such as a personal computer, a workstation, or a server that generates time-series data of free-viewpoint images. The generating device 20 acquires multiple pieces of video data captured by multiple imaging devices 10 and trains a restoration model based on the video data. Upon receiving a video generation request from the terminal device 30, the generating device 20 generates time-series data of free-viewpoint images based on the trained restoration model and transmits the data to the terminal device 30.

[0035] The terminal device 30 is an information processing terminal such as a personal computer, smartphone, or tablet terminal operated by a user of the video generation system 1000. The terminal device 30 accepts a video generation request from the user and transmits the video generation request to the generation device 20. The terminal device 30 receives time-series data of free viewpoint images from the generation device 20 and presents the data to the user.

[0036] Note that the overall configuration of the video generation system 1000 shown in FIG. 5 is one example, and various system configuration examples are possible depending on the application and purpose. For example, the video generation system 1000 may include multiple generation devices 20 and one or more terminal devices 30. For example, the generation device 20 may be realized by multiple computers, or may be realized as a cloud computing service. For example, the video generation system 1000 may be realized by a standalone computer that combines the functions of the generation device 20 and the terminal device 30. The classification of devices such as the generation device 20 and the terminal device 30 shown in FIG. 5 is one example.

[0037] <Functional Configuration of Generating Device> The functional configuration of the generating device in this embodiment will be described with reference to Fig. 6. Fig. 6 is a block diagram showing an example of the functional configuration of the generating device.

[0038] 6 , the generation device 20 includes a video acquisition unit 101, a video storage unit 102, a time estimation unit 103, a time synchronization unit 104, a model learning unit 105, a model storage unit 106, a request receiving unit 107, an image generation unit 108, and a video transmission unit 109. The generation device 20 functions as the video acquisition unit 101, the video storage unit 102, the time estimation unit 103, the time synchronization unit 104, the model learning unit 105, the model storage unit 106, the request receiving unit 107, the image generation unit 108, and the video transmission unit 109 by executing a generation program installed in advance.

[0039] The video acquisition unit 101 acquires video data from each of the multiple image capture devices 10. The video acquisition unit 101 may acquire the video data by receiving video data output in real time from the image capture devices 10. The video acquisition unit 101 may acquire the video data by reading video data captured by the image capture devices 10 and stored in a portable storage medium.

[0040] The moving image storage unit 102 stores a plurality of moving image data acquired by the moving image acquisition unit 101. The plurality of moving image data stored in the moving image storage unit 102 may include a plurality of images of the subject S captured successively in the time direction from a plurality of viewpoints. The capturing times of the frames of the plurality of moving image data may be asynchronous between the frames. The capturing times of the pixels included in each frame of the moving image data may be asynchronous between the frames.

[0041] The time estimation unit 103 estimates the shooting time of each pixel included in each frame of video data read from the video storage unit 102 and generates shooting time information indicating the shooting time of each pixel. The shooting time may be a relative time. For example, the relative time may be information indicating the time difference from the shooting time of a specific pixel, which is set as the time origin.

[0042] Here, an example has been described in which the time estimation unit 103 estimates the shooting time for each pixel, but the shooting time may also be estimated for each arbitrary portion included in each frame. The unit for estimating the shooting time may be, for example, one pixel, or a partial image including a predetermined number of pixels. In the following description, an example of processing in pixel units will be described, but this can also be interpreted as processing in units of partial images of a predetermined size.

[0043] The time estimation unit 103 may generate shooting time information for each piece of video data. The shooting time information corresponding to each piece of video data may indicate the shooting time of each pixel included in each frame included in the video data. An example of the shooting time information generated by the time estimation unit 103 is shown in FIG. 4.

[0044] Specifically, the time estimation unit 103 may estimate the shooting time using the following method: Note that the difference in shooting time between pixels in the same frame due to the rolling shutter may be obtained in advance by performing calibration.

[0045] As a first technique, the time estimation unit 103 may estimate the difference in shooting time between video data by calculating the correlation between sounds recorded in each video data. When calculating the correlation between sounds, it is preferable to correct for the speed of sound.

[0046] As a second method, the time estimation unit 103 may estimate the difference in shooting time between video data based on a synchronization signal recorded in each video data. In this case, it is preferable to transmit the synchronization signal when shooting a scene.

[0047] As a third technique, the time estimation unit 103 may estimate the difference in shooting time between video data based on a marker or the like recorded in each video data. The marker may be, for example, a stopwatch with a high refresh rate. In this case, it is preferable to record the marker before recording the scene.

[0048] As a fourth technique, the time estimation unit 103 may estimate the difference in shooting time between video data based on the time recorded in each video data. The time recorded in the video data may be the time of a clock built into the image capture device 10. In this case, it is preferable to synchronize the time of the clock built into each image capture device 10 before capturing a scene.

[0049] The time synchronization unit 104 synchronizes the shooting times in the multiple pieces of shooting time information generated by the time estimation unit 103. Note that the time synchronization unit 104 only needs to synchronize the shooting times of pixels included in the same frame number between video data so that they are roughly aligned, and does not need to synchronize all the shooting times so that they are exactly aligned.

[0050] The time synchronization unit 104 may synchronize the shooting times so that the difference in shooting time between pixels is shorter than the shooting interval in the video data. Specifically, the time synchronization unit 104 may shift the frame numbers of each pixel so that the shooting times of all pixels included in each frame of the video data with the same frame number fall within a time interval shorter than the shooting interval.

[0051] Fig. 7 is a diagram showing an example of shooting time information after synchronization. In the example shown in Fig. 7, pixels C and D of video data β are shifted by three frames so that the shooting time of each frame of video data α and the shooting time of each frame of video data β are within the shooting interval (0.020 seconds). Note that in the example shown in Fig. 7, each pixel in the same video data β is shifted by the same shift amount, but if the difference in shooting time between pixels varies due to the influence of a rolling shutter, the shift amount may differ for each pixel.

[0052] When the frame rates of the multiple image capture devices 10 are different from one another, the time synchronization unit 104 may synchronize the capture times based on the video data captured by the image capture device 10 with the highest frame rate (in other words, the video data with the shortest capture interval). For example, when the frame rate of the video data α is higher than the frame rate of the video data β, the time synchronization unit 104 may change the frame number of each frame of the video data β to the frame number of the frame captured closest to the capture time of each frame of the video data α.

[0053] The time synchronization unit 104 may synchronize the shooting time by interpolating frames in the video data read from the video storage unit 102. For example, if the subject S moves little, a frame interpolation method can be applied to the video data. With the frame interpolation method, for example, a frame at time 0.000 and a frame at time 0.020 can be used to generate a frame at any time between time 0.000 and time 0.020 (e.g., time 0.001, 0.002, ..., etc.).

[0054] When synchronizing shooting times using frame interpolation, the time synchronization unit 104 determines reference video data and, for each of the other video data, interpolates frames at times corresponding to the shooting times of each frame of the reference video data. The time synchronization unit 104 also updates the video data stored in the video storage unit 102 with the video data after frame interpolation. Furthermore, the time synchronization unit 104 also updates the shooting times of the video data for which frame interpolation has been performed to the shooting times of the reference video data in the multiple pieces of shooting time information generated by the time estimation unit 103.

[0055] The model learning unit 105 acquires the plurality of video data read from the video storage unit 102 and the shooting time information synchronized by the time synchronization unit 104, and generates a restored model F based on the video data and the shooting time information. θ For example, the model learning unit 105 trains a reconstruction model F by representing a scene in a time interval [t0, t1] from a time specified by time information t0 to a time specified by time information t1 using a neural network. θ We can train a reconstruction model F θ For example, a neural network using NeRF technology may be used. In this case, the model learning unit 105 may train an encoder 1 that outputs a feature value h based on coordinate information (x, y, z) and time information t, and a renderer 2 that outputs a color (R, G, B) and opacity σ based on the feature value h.

[0056] The encoder 1 may be configured to handle continuous time information t. To handle continuous time information t, the encoder 1 may be configured as follows.

[0057] For example, the encoder 1 may be configured by direct encoding. In the case of direct encoding, the encoder 1 may have a structure defined by equation (1), for example.

[0058]

[0059] For example, the encoder 1 may be configured using linear interpolation. When configured using linear interpolation, the encoder 1 may have a structure defined by equation (2), for example.

[0060]

[0061] Here, h0 is the feature value corresponding to the time information t0, and h1 is the feature value corresponding to the time information t1.

[0062] For example, the encoder 1 may be configured by warping. When configured by warping, the encoder 1 may have a structure defined by equation (3), for example.

[0063]

[0064] The encoder 1 may be configured by combining at least two of the above direct encoding, linear interpolation, and warping. For example, the encoder 1 may be configured to obtain feature values ​​h and h for each of the time information t and t by warping, and then obtain the feature value h by linearly interpolating them.

[0065] The encoder 1 may be trained based on a loss function that is differentiable with respect to the time information t. This configuration enables the encoder 1 to backpropagate the error between the captured image V1 and the free-viewpoint image V2 to the time information t, thereby simultaneously optimizing the time information t. By optimizing the time information t, it becomes possible to generate a free-viewpoint image with high accuracy even if the capture time estimated by the time estimation unit 103 contains an error.

[0066] The model storage unit 106 stores a trained reconstruction model F θ is stored. The trained reconstruction model F θ is the reconstruction model F trained by the model learning unit 105. θ is.

[0067] The request receiving unit 107 receives a video generation request from the terminal device 30. The video generation request may include, for example, viewpoint information that identifies a viewpoint specified by the user and time range information that indicates a time range of the video. Furthermore, the video generation request may include, for example, playback information that indicates a video format that can be played on the terminal device 30, a playback magnification of the video, etc.

[0068] The image generating unit 108 receives the video generation request received by the request receiving unit 107 and the trained reconstruction model F read from the model storage unit 106. θ The image generation unit 108 generates a plurality of free viewpoint images that are continuous in the time direction based on the viewpoint information and the time information t that belongs to the time range. θ The image generator 108 generates time information t at a predetermined time interval within the time range and generates a free viewpoint image corresponding to the viewpoint at the time. θ The time interval for generating the time information may be determined based on the video format or playback magnification included in the video generation request, for example.

[0069] The predetermined time interval may be, for example, the same as the shooting interval of the video data acquired by the video acquisition unit 101, or may be shorter than the shooting interval of the video data. If free viewpoint images are generated at a time interval shorter than the shooting interval of the video data, it becomes possible to play back the time-series data of the free viewpoint images in slow motion.

[0070] When generating free viewpoint images at a time interval shorter than the shooting interval of video data, the restoration model F θis a function that can calculate the color and opacity of a 3D point over that short time interval, and volume rendering processing 3 is performed at each time point within that short time interval. The result of calculating the time integral of the amount of light in the line of sight obtained is captured as each frame of the video. The light receiving element of the imaging device 10 does not observe the light intensity at the instant of time information t, but rather integrates the amount of light that arrives within a certain period of time. The time span of the integrated amount of light is also called the shutter speed. The faster the shutter speed of the imaging device 10, the more rapid a phenomenon can be observed. On the other hand, increasing the shutter speed of the imaging device 10 reduces the amount of integrated light, resulting in a lower S / N ratio.

[0071] By performing volume rendering processing 3 at each short time interval and calculating the time integral of the amount of light at each time, it is possible to express the relationship between the movement of the subject and the captured video without increasing the shutter speed, and a restoration model F for short time intervals can be created. θ It is possible to estimate the amount of light. For example, when the amount of light is low, such as at night, increasing the shutter speed reduces the amount of light captured within the shutter time, resulting in a decrease in the S / N ratio. However, by taking a photograph without increasing the shutter speed, the S / N ratio does not decrease. For example, it is assumed that the subject moves in a time shorter than the shutter speed, and that each pixel of the observed image observes the amount of light within the shutter speed of the subject. This amount of light can be found by integrating the amount of light in the optical axis direction relative to the subject over time. By estimating the subject based on images taken in multiple directions and at multiple times, it is possible to estimate changes over time shorter than the shutter speed.

[0072] The video transmitting unit 109 converts the time-series data of the free viewpoint images generated by the image generating unit 108 into a video format that can be played back on the terminal device 30. The video transmitting unit 109 may convert the time-series data of the free viewpoint images into a video format specified in the video generation request. The video transmitting unit 109 transmits the video data obtained by the format conversion to the terminal device 30.

[0073] <Flow of Training Process> The training process executed by the video generation system 1000 will be described with reference to Fig. 8. Fig. 8 is a flowchart showing an example of the training process. The training process is a process of training a restoration model based on multiple video data.

[0074] In step S1, each of the plurality of imaging devices 10 captures an image of a subject S and generates video data. Next, each of the plurality of imaging devices 10 transmits the video data to the generation device 20.

[0075] In the generation device 20, the video acquisition unit 101 receives video data from each of the multiple imaging devices 10. Next, the video acquisition unit 101 stores the received multiple pieces of video data in the video storage unit .

[0076] In step S2, the time estimation unit 103 of the generation device 20 reads out multiple pieces of video data stored in the video storage unit 102. Next, the time estimation unit 103 estimates the shooting time of each pixel included in each frame for each piece of read video data. Subsequently, the time estimation unit 103 generates shooting time information indicating the shooting time of each pixel for each piece of video data. Then, the time estimation unit 103 sends the multiple pieces of shooting time information corresponding to each piece of video data to the time synchronization unit 104.

[0077] In step S3, the time synchronization unit 104 of the generation device 20 receives multiple pieces of shooting time information from the time estimation unit 103. Next, the time synchronization unit 104 synchronizes the shooting times in the received multiple pieces of shooting time information. Then, the time synchronization unit 104 sends the multiple pieces of shooting time information with synchronized shooting times to the model learning unit 105.

[0078] The time synchronization unit 104 may synchronize the image capturing times so that the difference in image capturing time between pixels is shorter than the image capturing interval in the video data. Specifically, the time synchronization unit 104 shifts the frame numbers of each pixel so that the image capturing times of all pixels included in each frame of the video data with the same frame number fall within a time interval shorter than the image capturing interval.

[0079] The time synchronization unit 104 may synchronize the shooting times of the video data by interpolating frames. Specifically, the time synchronization unit 104 first determines reference video data, and for each of the other video data, interpolates frames at times corresponding to the shooting times of each frame of the reference video data. Next, the time synchronization unit 104 stores the video data after frame interpolation in the video storage unit 102. Then, the time synchronization unit 104 updates the shooting times of the video data for which frame interpolation has been performed to the shooting times of the reference video data in the multiple pieces of shooting time information generated in step S2.

[0080] In step S4, the model learning unit 105 of the generating device 20 receives a plurality of pieces of shooting time information from the time synchronization unit 104. Next, the model learning unit 105 generates a restored model F based on the received plurality of pieces of shooting time information and the plurality of pieces of video data read in step S2. θ This shooting time information is shooting time information in which the shooting times of at least two pixels in the same frame (same captured image) of the video data are different. Then, the model learning unit 105 trains the trained restoration model F θ is stored in the model storage unit 106.

[0081] <Flow of Generation Process> The generation process executed by the video generation system 1000 will be described with reference to Fig. 9. Fig. 9 is a flowchart showing an example of the generation process. The generation process is a process of generating time-series data of free viewpoint images based on a trained restoration model.

[0082] In step S11, the user performs an operation to request generation of a video on the terminal device 30. As an example, the operation may be an operation on an application pre-installed on the terminal device 30. The operation may include, for example, an operation to specify a viewpoint, an operation to specify a time range of the video, an operation to specify a video format, and an operation to specify a playback magnification of the video.

[0083] The terminal device 30 generates a video generation request in response to a user operation. The video generation request may include viewpoint information, time range information, and playback information. The terminal device 30 then transmits the video generation request to the generation device 20.

[0084] In the generation device 20, the request receiving unit 107 receives the moving image generation request from the terminal device 30. The request receiving unit 107 sends the received moving image generation request to the image generation unit .

[0085] In step S12, the image generating unit 108 of the generating device 20 receives a video generation request from the request receiving unit 107. Next, the image generating unit 108 uses the trained reconstruction model F stored in the model storage unit 106. θ Next, the image generation unit 108 reads out the video generation request and the trained reconstruction model F θ Then, the image generation unit 108 sends the time series data of the free viewpoint image to the video transmission unit 109.

[0086] In step S13, the video transmission unit 109 of the generation device 20 receives the time-series data of the free-viewpoint images from the image generation unit 108. Next, the video transmission unit 109 converts the time-series data of the free-viewpoint images into the video format specified in the video generation request. Then, the video transmission unit 109 transmits the video data obtained by the format conversion to the terminal device 30.

[0087] The terminal device 30 receives the video data from the generation device 20. Next, the terminal device 30 plays the received video data on a display device. The display device may be a display built into the terminal device 30 or an external display connected to the terminal device 30 via various interfaces.

[0088] <Summary> As is clear from the above explanation, the generation device 20 according to an embodiment of the present disclosure acquires videos of a scene captured using the rolling shutter method for each viewpoint, and trains a model that three-dimensionally reconstructs the scene, including changes over time during video capture, based on the videos captured from each viewpoint and information related to distortion caused by the rolling shutter method when the videos were captured.

[0089] The information relating to distortion caused by the rolling shutter method may be information relating to differences in photographing time within the same frame that occur due to the rolling shutter method.

[0090] The generation device 20 acquires images of the scene taken from each viewpoint, acquires time information for each image indicating the time corresponding to each part in the image, where the corresponding time differs between a first part and a second part in the image, and trains a model that reconstructs the scene in three dimensions based on the images taken from each viewpoint and the time information of each image.

[0091] The first portion and the second portion may be pixels associated with different times. The time information is information relating to the capture time of the image, and the times associated with the first portion and the second portion may be the capture times of the first portion and the second portion, respectively.

[0092] The generation device 20 acquires a plurality of time-series data sets including a plurality of images of a scene captured successively in a time direction for each viewpoint, in which the capture times of a plurality of pixels included in each of the images are asynchronous, generates, as time information, capture time information indicating the capture times of each pixel included in the image for each of the time-series data sets, and generates a restoration model F for generating time-series data of a free viewpoint image based on the time-series data and the capture time information. θ may be trained.

[0093] The plurality of time-series data may be images corresponding to the plurality of viewpoints whose capture times are asynchronous.

[0094] The generating device 20 may synchronize the shooting times in the shooting time information so that the difference in shooting time between images is shorter than the shooting interval between the multiple images.

[0095] When the shooting intervals for each piece of time-series data are different, the generating device 20 may synchronize the shooting times in the shooting time information using the time-series data with the shortest shooting interval as a reference.

[0096] The generating device 20 may synchronize the shooting times in the shooting time information by interpolating images for each piece of time-series data.

[0097] Restoration Model Fθ may include an encoder that generates features based on coordinates and time. The encoder may be configured to handle continuous time.

[0098] The encoder may generate features corresponding to any time by at least one of direct encoding, linear interpolation, and warping.

[0099] The encoder may be trained based on a loss function that is differentiable in time.

[0100] The generating device 20 generates a trained reconstruction model F θ and the trained reconstruction model F θ The generation device 20 may generate time-series data of free viewpoint images based on the trained reconstruction model F θ Based on this, free viewpoint images may be generated at a time interval shorter than the interval at which the multiple images are captured.

[0101] As a result, according to one embodiment of the present disclosure, it is possible to provide a technique capable of accurately reconstructing a 3D image of a scene captured from multiple viewpoints. In one aspect, it is possible to generate a high-resolution free-viewpoint image based on multiple video data captured a scene asynchronously.

[0102] [Hardware Configuration of Information Processing Device] Some or all of the devices (generation device 20 and terminal device 30) in the above-described embodiments may be configured with hardware, or may be configured with software (program) information processing executed by a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or the like. When configured with software information processing, software that realizes at least some of the functions of each device in the above-described embodiments may be stored on a non-transitory storage medium (non-transitory computer-readable medium) such as a CD-ROM (Compact Disc-Read Only Memory) or a USB (Universal Serial Bus) memory, and the software information processing may be executed by loading the software into a computer. The software may also be downloaded via a communication network. Furthermore, all or part of the software processing may be implemented in a circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), thereby allowing the software information processing to be executed by hardware.

[0103] The storage medium that stores the software may be a removable medium such as an optical disk, or a fixed medium such as a hard disk, memory, etc. The storage medium may be provided inside the computer (such as a main storage device or auxiliary storage device) or outside the computer.

[0104] 10 is a block diagram showing an example of the hardware configuration of each device (the generation device 20 and the terminal device 30) in the above-described embodiment. Each device may be realized as a computer 7 including, for example, a processor 71, a main storage device 72 (memory), an auxiliary storage device 73 (memory), a network interface 74, and a device interface 75, which are connected via a bus 76.

[0105] Although the computer 7 in FIG. 10 includes one of each component, it may also include multiple of the same component. Also, while FIG. 10 shows one computer 7, the software may be installed on multiple computers, with each of the multiple computers executing the same or different parts of the software. In this case, a distributed computing configuration may be used in which each computer communicates via a network interface 74 or the like to execute processing. In other words, each device (the generation device 20 and the terminal device 30) in the above-described embodiment may be configured as a system in which one or more computers execute instructions stored in one or more storage devices to realize functions. Furthermore, the system may be configured such that information transmitted from a terminal is processed by one or more computers provided on a cloud, and the processing results are transmitted to the terminal.

[0106] The various calculations of each device (the generation device 20 and the terminal device 30) in the above-described embodiments may be executed in parallel using one or more processors, or using multiple computers via a network. Furthermore, the various calculations may be distributed to multiple processing cores within a processor and executed in parallel. Furthermore, some or all of the processes, means, etc. disclosed herein may be realized by at least one of a processor and a storage device provided on a cloud that can communicate with the computer 7 via a network. Thus, each device in the above-described embodiments may be implemented in the form of parallel computing using one or more computers.

[0107] The processor 71 may be an electronic circuit (processing circuit, processing circuitry, CPU, GPU, FPGA, ASIC, etc.) that at least controls or performs calculations on a computer. The processor 71 may be a general-purpose processor, a dedicated processing circuit designed to perform a specific calculation, or a semiconductor device that includes both a general-purpose processor and a dedicated processing circuit. The processor 71 may also include an optical circuit or a calculation function based on quantum computing.

[0108] The processor 71 may perform arithmetic processing based on data or software input from each device or the like configured internally of the computer 7, and may output arithmetic results or control signals to each device or the like. The processor 71 may control each component constituting the computer 7 by executing the OS (Operating System) of the computer 7, applications, etc.

[0109] Each device (the generation device 20 and the terminal device 30) in the above-described embodiment may be realized by one or more processors 71. Here, the processor 71 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, the electronic circuits may communicate with each other via wire or wirelessly.

[0110] The main memory device 72 may store instructions executed by the processor 71, various data, etc., and information stored in the main memory device 72 may be read by the processor 71. The auxiliary memory device 73 is a memory device other than the main memory device 72. Note that these memory devices refer to any electronic component capable of storing electronic information and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. The memory devices for saving various data, etc. in each device (the generation device 20 and the terminal device 30) in the above-described embodiments may be realized by the main memory device 72 or the auxiliary memory device 73, or may be realized by internal memory built into the processor 71. For example, each memory unit in the above-described embodiments may be realized by the main memory device 72 or the auxiliary memory device 73.

[0111] When each device (the generating device 20 and the terminal device 30) in the above-described embodiment is configured with at least one storage device (memory) and at least one processor connected (coupled) to this at least one storage device, at least one processor may be connected to one storage device. Also, at least one storage device may be connected to one processor. Also, a configuration in which at least one processor among multiple processors is connected to at least one storage device among multiple storage devices may be included. Also, this configuration may be realized by storage devices and processors included in multiple computers. Furthermore, a configuration in which a storage device is integrated with a processor (for example, a cache memory including an L1 cache and an L2 cache) may be included.

[0112] The network interface 74 is an interface for connecting to the communication network 8 wirelessly or via a wire. The network interface 74 may be an appropriate interface, such as one that conforms to an existing communication standard. The network interface 74 may exchange information with an external device 9A connected via the communication network 8. The communication network 8 may be any one of a wide area network (WAN), a local area network (LAN), a personal area network (PAN), etc., or a combination thereof, as long as information is exchanged between the computer 7 and the external device 9A. An example of a WAN is the Internet, an example of a LAN is IEEE 802.11 or Ethernet (registered trademark), and an example of a PAN is Bluetooth (registered trademark) or NFC (Near Field Communication), etc.

[0113] The device interface 75 is an interface such as a USB that directly connects to the external device 9B.

[0114] The external device 9A is a device connected to the computer 7 via a network, and the external device 9B is a device connected directly to the computer 7.

[0115] The external device 9A or the external device 9B may be, for example, an input device. The input device is, for example, a camera, a microphone, a motion capture device, various sensors, a keyboard, a mouse, a touch panel, or the like, and provides acquired information to the computer 7. Alternatively, the external device 9A or the external device 9B may be a device including an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0116] Furthermore, the external device 9A or the external device 9B may be, for example, an output device. The output device may be, for example, a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) panel, or a speaker that outputs sound or the like. Alternatively, the external device 9A or the external device 9B may be a device including an output unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0117] The external device 9A or the external device 9B may be a storage device (memory). For example, the external device 9A may be a network storage or the like, and the external device 9B may be a storage device such as an HDD.

[0118] Furthermore, the external device 9A or the external device 9B may be a device having some of the functions of the components of each device (the generation device 20 and the terminal device 30) in the above-described embodiment. That is, the computer 7 may transmit some or all of the processing results to the external device 9A or the external device 9B, or may receive some or all of the processing results from the external device 9A or the external device 9B.

[0119] In this specification (including the claims), when the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, ab, ac, bc, or abc. It may also include multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it also includes the addition of elements other than the enumerated elements (a, b, and c), such as having d, as in abcd.

[0120] In this specification (including claims), when expressions such as "using / using data as input / based on / according to / in response to data" (including similar expressions) are used, unless otherwise specified, this includes cases where the data itself is used, or where data that has been processed in some way (e.g., data with noise added, normalized data, features extracted from data, intermediate representations of data, etc.) is used. Furthermore, when a statement is made that a result is obtained "using data as input / based on / according to / in response to data" (including similar expressions), this includes cases where the result is obtained based solely on the data, or where the result is influenced by other data, factors, conditions, and / or states other than the data. Furthermore, when a statement is made that "data is output" (including similar expressions), this includes cases where the data itself is used as output, or where data that has been processed in some way (e.g., data with noise added, normalized data, features extracted from data, intermediate representations of various data, etc.) is used as output, unless otherwise specified.

[0121] When the terms "connected" and "coupled" are used in this specification (including the claims), they are intended as open-ended terms that include any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, and physically connection / coupling. These terms should be interpreted appropriately according to the context in which they are used, but any connection / coupling form that is not intentionally or naturally excluded should be interpreted as being included in these terms without limitation.

[0122] In this specification (including the claims), the expression "A configured to B" may include the physical structure of element A having a configuration capable of performing operation B, and the permanent or temporary setting / configuration of element A being configured / set to actually perform operation B. For example, if element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). Also, if element A is a dedicated processor, dedicated arithmetic circuit, etc., it is sufficient that the circuit structure of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.

[0123] Whenever words implying containing or possessing (e.g., "comprising / including," "having," etc.) are used in this specification (including the claims), they are intended to be open-ended terms that include containing or possessing things other than the object designated by the object of the term. When the object of such words implying containing or possessing does not specify a quantity or suggests a singular number (e.g., expressions using the articles "a" or "an"), the expression should be construed as not being limited to a specific number.

[0124] In this specification (including the claims), although expressions such as "one or more" and "at least one" are used in some places and expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") should be interpreted as not necessarily being limited to a specific number.

[0125] In this specification, when a particular advantage / result is described as being obtained with respect to a particular configuration of an embodiment, it should be understood that the same advantage / result can also be obtained with one or more other embodiments having the same configuration, unless otherwise stated. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or situations, and that the effect is not necessarily obtained with the configuration. The effect is merely obtained by the configuration described in the embodiment when various factors, conditions, and / or situations are satisfied, and the effect does not necessarily occur in a claimed invention that defines the same configuration or a similar configuration.

[0126] In this specification (including claims), when multiple pieces of hardware perform a predetermined process, the pieces of hardware may cooperate to perform the predetermined process, or some of the hardware may perform all of the predetermined process. Furthermore, some of the hardware may perform part of the predetermined process, and other hardware may perform the rest of the predetermined process. In this specification (including claims), when an expression such as "one or more pieces of hardware perform a first process, and the one or more pieces of hardware perform a second process" (including similar expressions) is used, the hardware performing the first process and the hardware performing the second process may be the same or different. In other words, it is sufficient that the hardware performing the first process and the hardware performing the second process are included in the one or more pieces of hardware. Note that hardware may include an electronic circuit, a device including an electronic circuit, etc.

[0127] In this specification (including the claims), when multiple storage devices (memories) store data, each of the multiple storage devices may store only a portion of the data, or may store the entire data. Also, a configuration in which only some of the multiple storage devices store data may be included.

[0128] In this specification (including the claims), terms such as "first," "second," etc. are used merely as a way of distinguishing between two or more elements, and are not necessarily intended to impose technical meanings such as temporal aspect, spatial aspect, sequence, quantity, etc. on the subject. Thus, for example, a reference to a first element and a second element does not necessarily mean that only two elements may be employed therein, that the first element must precede the second element, that the first element must be present in order for the second element to be present, etc.

[0129] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, partial deletions, etc. are possible within the scope of the conceptual idea and spirit of the present invention, which is derived from the content defined in the claims and their equivalents. For example, when numerical values ​​or formulas are used in the above-described embodiments, they are shown for illustrative purposes and do not limit the scope of the present disclosure. Furthermore, the order of each operation shown in the embodiments is also illustrative and does not limit the scope of the present disclosure.

[0130] The disclosed technology may take the following forms as described below.

[0131] [Supplementary Note 1] An information processing device comprising: one or more memories; and one or more processors, wherein the one or more processors acquire videos of a scene captured by a rolling shutter method for each viewpoint; and train a model that three-dimensionally reconstructs the scene, including changes over time during capture of the videos, based on the videos captured from each viewpoint and information related to distortion caused by the rolling shutter method when capturing the videos.

[0132] [Supplementary Note 2] The information processing device according to Supplementary Note 1, wherein the information is information relating to a difference in shooting time within the same frame that occurs due to the rolling shutter method.

[0133] [Supplementary Note 3] An information processing device comprising: one or more memories; and one or more processors, wherein the one or more processors acquire images of a scene photographed for each viewpoint; generate, for each of the images of the scene photographed for each viewpoint, time information indicating a time corresponding to each part in the image, where different times are associated between a first part and a second part of the same image; and train a model that three-dimensionally reconstructs the scene based on the images photographed from each viewpoint and the time information of each image.

[0134] [Supplementary Note 4] The information processing device according to Supplementary Note 3, wherein the first portion and the second portion are pixels associated with different times.

[0135] [Supplementary Note 5] The information processing device according to Supplementary Note 3 or 4, wherein the time information is information relating to a shooting time of the image, and the times associated with the first portion and the second portion are shooting times of the first portion and the second portion, respectively.

[0136] [Supplementary Note 6] The information processing device according to any one of Supplementary Notes 3 to 5, wherein the one or more processors acquire a plurality of time series data including a plurality of images of the scene photographed consecutively in a time direction for each viewpoint, wherein the photographing times of a plurality of pixels included in each of the images are asynchronous; generate, as the time information, photographing time information indicating the photographing time of each pixel included in the image for each of the time series data; and train a restoration model for generating time series data of a free viewpoint image based on the time series data and the photographing time information.

[0137] [Supplementary Note 7] The information processing device according to Supplementary Note 6, wherein the plurality of time-series data are asynchronous in terms of the image capturing times of the images corresponding to the plurality of viewpoints.

[0138] [Supplementary Note 8] The information processing device according to Supplementary Note 6 or 7, wherein the one or more processors synchronize the shooting times in the shooting time information so that a difference in shooting time between the pixels is shorter than a shooting interval at which the multiple images were taken.

[0139] [Supplementary Note 9] The information processing device according to Supplementary Note 8, wherein, when the shooting intervals for each of the time series data are different, the one or more processors synchronize the shooting times in the shooting time information based on the time series data with the shortest shooting interval.

[0140] [Supplementary Note 10] The information processing device according to Supplementary Note 8 or 9, wherein the one or more processors synchronize the image capture times in the image capture time information by interpolating the images for each of the time series data.

[0141] [Supplementary Note 11] The information processing device according to any one of Supplementary Notes 6 to 10, wherein the restoration model includes an encoder that generates features based on coordinates and time, and the encoder is configured to be able to handle continuous time.

[0142] [Supplementary Note 12] The information processing device according to Supplementary Note 11, wherein the encoder generates the feature corresponding to any time by at least one of direct encoding, linear interpolation, and warping.

[0143] [Supplementary Note 13] The information processing device according to Supplementary Note 11 or 12, wherein the encoder is trained based on a loss function that is differentiable in time.

[0144] [Supplementary Note 14] The information processing device according to any one of Supplementary Notes 6 to 13, wherein the one or more memories store the trained reconstruction model, and the one or more processors generate time-series data of the free viewpoint image based on the trained reconstruction model.

[0145] [Supplementary Note 15] The information processing device according to Supplementary Note 14, wherein the one or more processors generate the free viewpoint images at a time interval shorter than a capturing interval at which the plurality of images are captured, based on the trained restoration model.

[0146] [Supplementary Note 16] A method for generating a model for three-dimensionally reconstructing a scene, executed by one or more processors, comprising: acquiring videos of a scene captured using a rolling shutter method for each viewpoint; and training a model for three-dimensionally reconstructing the scene, including changes over time during the capture of the videos, based on the videos captured at each viewpoint and information related to distortions caused by the rolling shutter method when the videos were captured.

[0147] [Supplementary Note 17] A method for generating a model for three-dimensionally reconstructing a scene, executed by one or more processors, comprising: acquiring images of the scene taken from each viewpoint; generating time information for each of the images of the scene taken from each viewpoint, the time information indicating a time corresponding to each part in the image; training a model for three-dimensionally reconstructing the scene based on the images taken from each viewpoint and the time information of each image; and the time information associating different times between a first part and a second part of the same image.

[0148] This application claims priority from Japanese Patent Application No. 2024-4608, filed with the Japan Patent Office on January 16, 2024, the entire contents of which are incorporated herein by reference.

[0149] REFERENCE SIGNS LIST 10 Imaging device 20 Generation device 30 Terminal device 101 Video acquisition unit 102 Video storage unit 103 Time estimation unit 104 Time synchronization unit 105 Model learning unit 106 Model storage unit 107 Request reception unit 108 Image generation unit 109 Video transmission unit 1000 Video generation system

Claims

1. An information processing apparatus comprising: one or more memories and one or more processors, wherein the one or more processors acquire a video of a scene captured in a rolling shutter manner for each viewpoint, and train a model for three-dimensionally reconstructing the scene including temporal changes during the shooting of the video based on the video captured at each viewpoint and information regarding distortion due to the rolling shutter manner when the video was captured.

2. The information processing apparatus according to claim 1, wherein the information is information regarding a difference in shooting times within the same frame caused by the rolling shutter manner.

3. An information processing apparatus comprising: one or more memories and one or more processors, wherein the one or more processors acquire an image of a scene captured for each viewpoint, and for each of the images of the scene captured for each viewpoint, acquire time information indicating a time corresponding to each part in the image, the time information being such that the times corresponding to a first part and a second part in the image are different, and train a model for three-dimensionally reconstructing the scene based on the images captured at each viewpoint and the time information of each image.

4. The information processing apparatus according to claim 3, wherein the first part and the second part are pixels associated with different times.

5. The information processing apparatus according to claim 3 or 4, wherein the time information is information regarding the shooting time of the image, and the times associated with the first part and the second part are respectively the shooting times of the first part and the second part.

6. The one or more processors acquire a plurality of time-series data including a plurality of images in which the scene is continuously captured in the time direction for each viewpoint, and the shooting times of a plurality of pixels included in each of the images are asynchronous with each other, generate shooting time information indicating the shooting time of each pixel included in the image as the time information for each of the time-series data, and train a restoration model for generating time-series data of a free viewpoint image based on the time-series data and the shooting time information.

7. The information processing apparatus according to claim 6, wherein the shooting times of the plurality of time-series data are asynchronous with each other among the images corresponding to a plurality of viewpoints.

8. The one or more processors synchronize the shooting times in the shooting time information so that a difference in shooting times between the pixels is shorter than a shooting interval at which the plurality of images are shot, in the information processing apparatus according to claim 6 or 7.

9. The one or more processors synchronize the shooting times in the shooting time information based on the time-series data having the shortest shooting interval when the shooting intervals for each of the time-series data are different, in the information processing apparatus according to claim 8.

10. The one or more processors synchronize the shooting times in the shooting time information by interpolating the images for each of the time-series data, in the information processing apparatus according to claim 8 or 9.

11. The restoration model includes an encoder that generates feature amounts based on coordinates and time, and the encoder is configured to handle continuous time, in the information processing apparatus according to any one of claims 6 to 10.

12. The encoder generates the feature amounts corresponding to arbitrary times by at least one of direct encoding, linear interpolation, or warping, in the information processing apparatus according to claim 11.

13. The encoder is trained based on a loss function differentiable with respect to time, in the information processing apparatus according to claim 11 or 12.

14. The one or more memories store the trained restoration model, and the one or more processors generate time-series data of the free viewpoint image based on the trained restoration model, in the information processing apparatus according to any one of claims 6 to 13.

15. The one or more processors generate the free viewpoint image at a time interval shorter than the shooting interval at which the plurality of images are shot, based on the trained restoration model, in the information processing apparatus according to claim 14.

16. A method of generating a model for three-dimensionally reconstructing a scene, which is executed by one or more processors, the method including: acquiring videos of a scene shot in a rolling shutter method for each viewpoint; and training a model for three-dimensionally reconstructing the scene including a time change during shooting of the videos, based on the videos shot at each viewpoint and information regarding distortion due to the rolling shutter method when the videos are shot.

17. A method for generating a model for three-dimensional reconstruction of a scene, which is executed by one or more processors, the method comprising: obtaining images of the scene taken for each viewpoint; generating time information indicating a time corresponding to each part in the image for each of the images of the scene taken for each viewpoint; training a model for three-dimensional reconstruction of the scene based on the images taken at each viewpoint and the time information of each image; wherein the time information has different times associated with a first part and a second part of the same image.

Citation Information

Patent Citations

  • Training device, training method, and prediction device

    JP2020135141A

  • Three-dimensional model generation method

    JP2022131971A

  • Information processing device and method

    WO2024203013A1