Information processing apparatus, information processing method, and computer-readable non-transitory storage medium

By using model estimation and difference map model to generate switching maps in the rendering results of new viewpoint images, the real-time rendering problem of super-resolution processing in new viewpoint images is solved, achieving efficient image sharpness improvement and time reduction.

CN121773449APending Publication Date: 2026-03-31SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve real-time super-resolution processing in the rendering results of new viewpoint images, especially at high resolutions, and existing super-resolution processing methods cannot be effectively applied to the rendering results of new viewpoint images.

Method used

The new viewpoint image is estimated using a first model, and the target and non-target regions are determined using a second model. Super-resolution processing is performed only on the target regions, and a switching map is generated by combining the difference map model to accelerate the super-resolution processing.

Benefits of technology

It accelerates super-resolution processing in the rendering results of new viewpoint images, improves image clarity and reduces processing time, thus meeting the requirements of real-time rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121773449A_ABST
    Figure CN121773449A_ABST
Patent Text Reader

Abstract

The information processing apparatus includes an estimation unit and a super-resolution processing unit. The estimation unit estimates a new viewpoint image corresponding to an arbitrary line-of-sight direction by using a first model trained to estimate the new viewpoint image based on a multi-viewpoint image including a plurality of two-dimensional images. A super-resolution processing unit performs super-resolution processing on the estimated new viewpoint image. The estimation unit estimates a processing target region in which the super-resolution processing is performed and a non-processing target region in which the super-resolution processing is not performed by using a second model trained to estimate the processing target region and the non-processing target region for each new viewpoint image in accordance with the line-of-sight direction. The super-resolution processing unit performs super-resolution processing on the estimated processing target region and does not perform super-resolution processing on the estimated non-processing target region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to information processing apparatus, information processing methods, and computer-readable non-transitory storage media. Background Technology

[0002] In recent years, a technique called Neural Radiation Field (NeRF) has been proposed and developed, which is capable of reproducing images from any new viewpoint (new viewpoint or new viewpoint image) using multi-viewpoint input images as well as the position and posture of the camera device as input (see, for example, Non-Patent Literature 1).

[0003] Initially, NeRF required tens of seconds to generate a new viewpoint image (new viewpoint synthesis), and real-time rendering for the new viewpoint direction of the input was difficult. However, in recent years, various techniques have been proposed to improve real-time performance to, for example, approximately 60 fps in full HD (see, for example, non-patent literature 2). However, NeRF suffers from the problem that the output image is more blurry than the input image, and it is difficult to reproduce sharp images such as real-time pictures.

[0004] As a signal processing solution to this problem, there are methods that apply learning-based super-resolution (SR) techniques to the NeRF output image to achieve the restoration of a sharper image that is closer to the input image. On the other hand, a method combining "direct voxel mesh optimization" (a NeRF technique) with learning-based super-resolution processing has also been proposed to demonstrate that a sharper output image than that obtained by conventional NeRF can be obtained (see, for example, Non-Patent Literature 3).

[0005] However, since super-resolution processing also takes processing time, there is a problem: high-resolution real-time rendering (e.g., at 60 fps or higher in 4K or 8K) is difficult when combined with NeRF.

[0006] Here, the following method has been proposed as a method for accelerating super-resolution processing: it determines the complexity of each image block as a predetermined pixel unit, and switches between lightweight super-resolution processing and heavyweight super-resolution processing for each image block based on the determination result (see, for example, Non-Patent Literature 4).

[0007] However, real-time rendering is difficult to achieve using this method because determining the complexity of each image patch also consumes processing time. To address this, it is conceivable to perform the determination process in advance and store the determination results. However, in the case of new viewpoint images, unlike the case of two-dimensional images, storing the determination results of the viewpoint to be output is impractical in terms of data capacity.

[0008] Note that, as another previous example of switching super-resolution processing for each image patch, techniques have been proposed to switch super-resolution processing based on the precision of motion vectors in the video (see, for example, Patent Document 1).

[0009] Citation List

[0010] Non-patent literature

[0011] Non-Patent Reference 1: Mildenhall, Ben, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. 2020. “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis.” arXiv [cs.CV]. arXiv. http: / / arxiv.org / abs / 2003.08934.

[0012] Non-Patent Literature 2: Wizadwongsa, Suttisak, Pakkapon Phongthawee, Jiraphon Yenphraphai, and Supasorn Suwajanakorn. 2021. “NeX: Real-Time View Synthesis with Neural Basis Expansion.” arXiv [cs.CV]. arXiv. http: / / arxiv.org / abs / 2103.05606.

[0013] Non-Patent Literature 3: Wang, Zhongshu, Lingzhi Li, Zhen Shen, Li Shen and Liefeng Bo. 2022. “4K-NeRF: High Fidelity Neural Radiance Fields at Ultra HighResolutions.” arXiv [cs.CV]. arXiv. http: / / arxiv.org / abs / 2212.04701.

[0014] Non-patent literature 4: Kong, Xiangtao, Hengyuan Zhao, Yu Qiao and Chao Dong. 2021. "ClassSR: A General Framework to Accelerate Super-Resolution Networks by DataCharacteristic." https: / / openaccess.thecvf.com / content / CVPR2021 / papers / Kong_ClassSR_A_General_Framework_to_Accelerate_Super-Resolution_Networks_by_Data_CVPR_2021_paper.pdf.

[0015] Patent documents

[0016] Patent Document 1: JP 2018-023034 A Summary of the Invention

[0017] Technical issues

[0018] However, although the aforementioned prior art can accelerate super-resolution processing by switching super-resolution processing for each image patch based on the accuracy of motion prediction, the prior art is a technique used to improve the quality of 2D video images using codec information of 2D video, and therefore cannot be applied as is to perform super-resolution processing on the rendering results of new viewpoint images.

[0019] Therefore, this disclosure provides an information processing apparatus, an information processing method, and a computer-readable non-transitory storage medium that can accelerate super-resolution processing when performing super-resolution processing on the rendering results of a new viewpoint image.

[0020] Solution to the problem

[0021] To address the above problems, one aspect of the information processing apparatus according to this disclosure includes an estimation unit and a super-resolution processing unit. The estimation unit uses a first model to estimate a new viewpoint image corresponding to any viewing direction, the first model being learned to estimate the new viewpoint image based on a multi-viewpoint image comprising multiple two-dimensional images. The super-resolution processing unit performs super-resolution processing on the estimated new viewpoint image. Furthermore, the estimation unit uses a second model to estimate a processing target region for which super-resolution processing is to be performed and a non-processing target region for which super-resolution processing is not performed, the second model being learned to estimate the processing target region and non-processing target region according to the viewing direction for each region in the new viewpoint image. Additionally, the super-resolution processing unit performs super-resolution processing on the estimated processing target region and does not perform super-resolution processing on the estimated non-processing target region. Attached Figure Description

[0022] Figure 1 This is a diagram illustrating the flow of basic processing prior to super-resolution processing according to an embodiment of the present disclosure.

[0023] Figure 2 This is a conceptual diagram of MPI.

[0024] Figure 3 This is a concept map of formula (b).

[0025] Figure 4 yes Figure 3 Supplementary illustration.

[0026] Figure 5 This is a conceptual diagram of the information processing method according to this embodiment.

[0027] Figure 6 Examples of NeRF output images, super-resolution processed images, and difference images are shown.

[0028] Figure 7 This is a diagram used to illustrate an overview of the information processing method according to this embodiment.

[0029] Figure 8 This is a block diagram illustrating an example configuration of an information processing apparatus according to an embodiment of the present disclosure.

[0030] Figure 9 This is a flowchart illustrating the learning process performed by an information processing device to generate a difference graph model for a switching graph.

[0031] Figure 10 This is a graph showing the parameters added during the learning process in the difference graph model.

[0032] Figure 11This is a flowchart illustrating the process of real-time rendering using a switching graph, performed by an information processing device.

[0033] Figure 12 An example of a difference plot is shown.

[0034] Figure 13 An example of a super-resolution processing switching graph is shown.

[0035] Figure 14 This is a hardware configuration diagram illustrating an example of a computer that performs the functions of an information processing device. Detailed Implementation

[0036] In the following description, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Note that in the following embodiments, the same parts are indicated by the same reference numerals, so that redundant descriptions can be omitted.

[0037] Furthermore, in the following description, the information processing apparatus according to the embodiments of this disclosure (hereinafter, as appropriate, referred to as "this embodiment") is the following information processing apparatus 100 (see Figure 8 The method uses NeRF to render a new viewpoint image corresponding to the viewing direction, and performs super-resolution processing on the rendered output image (hereinafter, where appropriate, referred to as the "NeRF output image") to output a super-resolution processed image. Furthermore, in the following description, the information processing method according to embodiments of this disclosure is an information processing method executed by the information processing apparatus 100.

[0038] The contents of this disclosure will be described in the following order.

[0039] 1. Overview

[0040] 1-1. Description of NeRF

[0041] 1-2. Description of NeX rendering process

[0042] 1-3. Description of super-resolution processing and its switching process

[0043] 2. Configuration example of information processing device

[0044] 2-1. Flowchart

[0045] 2-2. The learning process for the difference graph model used to generate the switching graph

[0046] 2-3. Processing steps for real-time rendering using switching graphs

[0047] 3. Modification

[0048] 4. Hardware Configuration

[0049] 5. Conclusion

[0050] <<1. Overview>>

[0051] <1-1. Description of NeRF>

[0052] Figure 1 This is a diagram illustrating the flow of basic processing prior to super-resolution processing according to an embodiment of this disclosure. (As shown) Figure 1 As shown, "NeRF" is a technique used to generate new viewpoint images from any new viewpoint in a three-dimensional space where the object is the subject, using multi-viewpoint images as a group of two-dimensional images captured from various angles by camera device 50 based on deep learning.

[0053] like Figure 1 As shown, in the information processing method according to this embodiment, NeRF deep learning (hereinafter, as appropriate, referred to as "NeRF learning") using multi-viewpoint images as a dataset is performed to generate a NeRF model 104a. The NeRF model 104a is a neural network that is learned to estimate a new viewpoint image observed from the viewpoint direction from an arbitrary viewpoint location when the viewpoint direction is input. In other words, the NeRF model 104a is a model for synthesizing radiation fields using a neural network.

[0054] Subsequently, in the information processing method according to this embodiment, when an arbitrary viewing direction is input, the NeRF model 104a is used to estimate a new viewpoint image observed from the viewing direction (hereinafter, as appropriate, referred to as "NeRF estimation"), and a NeRF output image is generated as the NeRF estimation result, that is, the rendering result. Figure 1 The arrows in the NeRF output image shown schematically indicate how the position and orientation of the object (in this case, the cat) change depending on the viewpoint position and the direction of the line of sight.

[0055] Then, in the information processing method according to this embodiment, super-resolution processing is performed on the NeRF output image, and the super-resolution processed image is displayed on a display device (such as a display).

[0056] Note that while photogrammetry techniques, which conventionally employ model-based methods, exist, they struggle to reproduce information such as the degree of light reflection and the appearance of a landscape as seen through an object. However, in NeRF, by using neural networks, the presence and color distribution of an object at each point in three-dimensional space (hereafter referred to as the "radiation field") can be estimated with high accuracy based on information about position and viewing angle. This radiation field is then transformed into a two-dimensional image using a rendering method called volume rendering, allowing for the representation of a new viewpoint image.

[0057] <1-2. Description of NeX rendering process>

[0058] Incidentally, while NeRF can typically render high-resolution new viewpoint images, it suffers from the problem of requiring significant processing time for volume rendering using neural networks. Specifically, even when using a GPU (Graphics Processing Unit), processing for each viewpoint takes approximately 30 seconds.

[0059] On the other hand, recently, many techniques have been proposed to shorten rendering time by caching the computation results of neural networks in various formats. In the technique known as "NeX" disclosed in Non-Patent Document 2, intermediate data of NeRF model 104a is cached in the form of a multi-plane image (MPI) with multiple images arranged in the depth direction, thereby accelerating the volume rendering process by about 1,000 times compared to NeRF and achieving real-time rendering.

[0060] Figure 2 This is a conceptual diagram of MPI. (Example) Figure 2 As shown, MPI typically has a structure including RGBA layer groups in which translucent RGBA images are stacked in a frustocon shape in three-dimensional space to cover from the object closest to the camera device 50 capturing the scene to a sufficiently far depth. New viewpoint images are generated by superimposing the RGBA images that constitute the MPI while shifting them from each viewpoint according to the viewing direction. Note that... Figure 2 The first viewpoint is, for example, the shooting position of the camera device 50. The second viewpoint is, for example, any new viewpoint position that is not the shooting position of the camera device 50.

[0061] The radiation field in NeX is learned so that the RGB values ​​obtained by volume rendering of these MPIs are optimized to be close to the RGB values ​​captured by camera device 50.

[0062] Then, the rendering process performed by NeX takes the following form: Rays are traced from the view direction, and the color at the point where the ray intersects with the MPI is multiplied by the probability of existence at that point. And the probability of existence added to that point. The following formula (a) represents the probability of existence. And input the color components. The formula for calculating the RGB values ​​of the output image under certain conditions.

[0063] ... (a)

[0064] By analyzing color components that do not change according to the angle at which the image is observed. (That is, angle-independent color components) and color components that vary depending on the angle at which the image is observed. to (That is, the angle-dependent (viewpoint-dependent, referred to as "VD" where appropriate in the following text) color components) are weighted and summed to calculate the value in formula (a). The following formula (b) is used to calculate... The formula.

[0065] ... (b)

[0066] Figure 3 This is a concept map of formula (b). Furthermore, Figure 4 yes Figure 3 Supplementary illustration. For example... Figure 3 As shown, any point in a layer of NeX's NeRF model 104a (here, a point in the upper right corner) has the existence probability for each point as described above. of and ,……and The reflection coefficient, Is when It is obtained when observed from different angles.

[0067] Then, the basis functions learned from the neural network ,……and Depending on the viewing angle ,……and The values ​​are weighted differently and linearly combined. Therefore, the RGB values ​​at a given point are determined when viewed from a specific angle. This modeling allows us to reproduce the phenomenon that objects appear different depending on the angle, even in the same location.

[0068] Note that this is for reference. Figure 4 right Figure 3 The description of each parameter shown is supplemented, which in part includes the overlaps mentioned above. It represents the probability of existence in three-dimensional coordinates and corresponds to the "geometric component" (see "1" Alpha in the figure). Furthermore, It is angle-independent color information and corresponds to the “texture component” (see “2” Color in the figure).

[0069] also, to ...This is angle-dependent color information, and corresponds to the "VD component" (see "3" Color_VD" in the figure). Furthermore, to ...is angle-dependent color information and corresponds to the "basis function" (see "4" Basis in the figure).

[0070] <1-3. Description of Super-Resolution Processing and its Switching Process>

[0071] Next, we will refer to Figures 5 to 7 Describe super-resolution processing and its switching process. Figure 5 This is a conceptual diagram of the information processing method according to this embodiment. Furthermore, Figure 6 Examples of NeRF output images, images after super-resolution processing, and difference images are shown. Furthermore, Figure 7 This is a diagram used to illustrate an overview of the information processing method according to this embodiment.

[0072] Super-resolution processing is a technique used to obtain a high-resolution output image from a low-resolution input image. Learning-based super-resolution processing (one type of super-resolution processing) can generate a high-resolution image from a single input image and is known as a method with high restoration accuracy. For example... Figure 5 As shown in the "NeRF Output Image" and "Image After Super-Resolution Processing (100%)", super-resolution processing can be performed on the NeRF output image, for example, in Figure 5 Medium-sharpen the doll's hair to improve the image's resolution.

[0073] However, since super-resolution processing also incurs computational costs, it is necessary to accelerate super-resolution processing in order to achieve real-time rendering via NeRF. To achieve this... Figure 5 The “Image after super-resolution processing (100%)” in the text refers to the processing time spent performing super-resolution processing on all pixels of the NeRF output image (that is, 100%).

[0074] In the information processing method according to this embodiment, a switching process for accelerating super-resolution processing has been studied. For example... Figure 5 As shown in the “Switch Image (50% Open)”, the switching process in super-resolution processing refers to the following process: it is used to divide the image into predetermined pixel units (e.g., 128). The image is divided into 128-pixel blocks, and super-resolution processing is toggled for each block. Image blocks for which super-resolution processing is performed correspond to "processing target regions." Image blocks for which super-resolution processing is not performed correspond to "non-processing target regions." Note that in this embodiment, performing super-resolution processing corresponds to "on," and not performing super-resolution processing corresponds to "off."

[0075] In the information processing method according to this embodiment, NeRF learning is performed on the difference image between the super-resolution image before and after super-resolution processing. Figure 6 Examples of NeRF output images, super-resolution processed images, and difference images are shown. In the information processing method according to this embodiment, such as... Figure 6 As shown, NeRF learning is performed on the difference image between the NeRF output image and the image after super-resolution processing for each viewpoint of the multiple 2D images included in the above multi-viewpoint image. Then, when a new viewpoint and a new viewing direction are input, image patches estimated to have a large amount of 3D difference between before and after super-resolution processing are set to open, and image patches estimated to have a small amount of difference are set to close.

[0076] For example, even if image patches are set to be enabled to perform super-resolution processing on half (that is, 50%) of the total number of image patches (such as... Figure 5 In the case of “Image after super-resolution processing (50%)”, it is also possible to generate an output image that looks basically the same as “Image after super-resolution processing (100%)”.

[0077] To achieve this processing, a high-precision on / off switching graph of super-resolution processing (hereinafter referred to as a "switching graph") is needed to implement the super-resolution processing. The switching graph corresponds to an example of "graph information." In the prior art, networks for complexity determination are used to switch processing for each image patch, or based on the reliability of motion vectors. In this embodiment, a method for generating a switching graph corresponding to a 3D representation using NeRF will be described.

[0078] Note that although an example using image blocks has been described here, the predetermined pixel unit used as the unit for switching on / off for super-resolution processing can be each pixel. Such a unit for switching on / off can be appropriately selected based on, for example, the performance of the GPU included in the information processing device 100.

[0079] An overview of the information processing method related to the generation of the switching diagram according to this embodiment will be described. For example... Figure 7As shown, in the information processing method according to this embodiment, a multi-view difference image is acquired, which is a difference image between the NeRF output image and the image after super-resolution processing for each viewpoint of the plurality of two-dimensional images. Then, NeRF learning is performed using the multi-view difference image as a dataset to generate a difference map model 110a. The difference map model 110a is a neural network that is learned to estimate the difference image observed from the viewing direction from any viewpoint position when the input viewing direction is given. In other words, the difference map model 110a is a model for synthesizing radiation fields using a neural network.

[0080] At this point, the difference map model 110a can be learned through transfer learning based on the NeRF model 104a. This is because the NeRF model 104a has already learned the existence probability of objects in three-dimensional space. Therefore, transfer learning can be used to achieve high-speed convergence. (See below for further details.) Figure 10 Details describing this point.

[0081] Subsequently, in the information processing method according to this embodiment, when an arbitrary viewing direction is input, NeRF estimation using the difference map model 110a is performed, and a difference image observed from the viewing direction is rendered and output as a difference map.

[0082] Then, in the information processing method according to this embodiment, the switching map is generated based on the difference map. For example, for each of the above image blocks, the switching map is generated by reducing the resolution of the difference map.

[0083] Then, in the information processing method according to this embodiment, the above-described on / off settings are performed for each image block based on the generated switching map, and super-resolution processing is performed only for image blocks that have been set to be on.

[0084] <<2. Configuration Example of Information Processing Device>>

[0085] <2-1. Flowchart>

[0086] Next, a configuration example of the information processing apparatus 100 that performs the information processing method according to this embodiment will be described. Figure 8 This is a block diagram illustrating an example configuration of an information processing apparatus 100 according to an embodiment of the present disclosure. Note that in Figure 8 In this document, only the components required to describe this embodiment are shown as functional blocks, and descriptions of general components are omitted.

[0087] In addition, when referring to Figure 8 When providing a description, simplify or omit details about the components that have already been described.

[0088] like Figure 8 As shown, the information processing device 100 includes a camera device position and posture input unit 101, a multi-view image input unit 102, a NeRF learning unit 103, a NeRF model storage unit 104, a NeRF estimation unit 105, a NeRF output image storage unit 106, and a super-resolution processing unit 107.

[0089] In addition, the information processing device 100 includes a super-resolution processed image storage unit 108, a difference image calculation unit 109, a difference map model storage unit 110, a difference map storage unit 111, a switching map generation unit 112, a gaze direction input unit 113, and a switching threshold setting unit 114.

[0090] Note that, although Figure 8 Although not shown, the information processing device 100 may be connected to an input device (such as a keyboard or mouse) or an output device (such as a display) that receives general input operations from users, etc.

[0091] Furthermore, although not shown, the information processing apparatus 100 includes a storage unit and a control unit. The storage unit is implemented by, for example, a storage device (such as random access memory (RAM), read-only memory (ROM), flash memory, or hard disk drive (HDD)).

[0092] The storage units include, for example, a NeRF model storage unit 104, a NeRF output image storage unit 106, a super-resolution processed image storage unit 108, a difference map model storage unit 110, and a difference map storage unit 111.

[0093] The control unit controls each unit of the information processing device 100. The control unit executes a program (not shown) stored in a storage unit according to this embodiment, using RAM or the like as a working area, via a GPU, central processing unit (CPU), microprocessor unit (MPU), or similar device. Furthermore, the control unit can be implemented using an integrated circuit, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0094] The control unit includes, for example, a camera device position and posture input unit 101, a multi-view image input unit 102, a NeRF learning unit 103, a NeRF estimation unit 105, a super-resolution processing unit 107, a difference image calculation unit 109, a switching image generation unit 112, a gaze direction input unit 113, and a switching threshold setting unit 114, and implements or performs the information processing functions and actions described below.

[0095] <2-2. Processing steps for learning the difference map model used to generate the switching map>

[0096] Next, we will refer to the references that have already been referenced. Figure 8 as well as Figure 9 and Figure 10 The process described is the learning process performed by the information processing device 100 to generate the difference graph model 110a of the switching graph. Figure 9 This is a flowchart illustrating the learning process performed by the information processing device 100 to generate the difference map model 110a of the switching map. Furthermore, Figure 10 This is a graph showing the parameters added in the learning process of the difference graph model 110a.

[0097] The switching graph required according to this embodiment is graph information capable of determining whether the difference between the input image and the output image in super-resolution processing is large or small. Furthermore, it is desirable that the graph information can represent values ​​that change according to viewpoint transformation, and that it can be obtained with low computational cost and high speed.

[0098] Therefore, in this embodiment, NeRF learning is performed on the difference between the NeRF rendered image (that is, the NeRF output image) and the super-resolution processed image, which is the result of super-resolution processing of the NeRF output image, to generate graph information that enables real-time 3D rendering based on the difference.

[0099] First, in typical NeRF learning, such as Figure 8 As shown, the NeRF learning unit 103 learns color information and 3D presence probability based on viewpoint changes from the camera device position and posture input unit 101 and the multi-view image input unit 102, along with the corresponding multi-view images. Then, the NeRF learning unit 103 outputs the obtained NeRF model 104a to the NeRF model storage unit 104.

[0100] Furthermore, the NeRF estimation unit 105 estimates the image observed from the viewpoint and the viewing direction input from the viewing direction input unit 113 based on the NeRF model 104a stored in the NeRF model storage unit 104, and outputs the image as a NeRF output image to the NeRF output image storage unit 106.

[0101] Furthermore, when super-resolution processing is subsequently performed, the super-resolution processing unit 107 performs super-resolution processing on the NeRF output image stored in the NeRF output image storage unit 106, and outputs the image as the super-resolution processed image to the super-resolution processed image storage unit 108.

[0102] Following the standard NeRF learning and super-resolution processing described above, the learning process for generating the difference graph model 110a of the switching graph is performed as follows. For example... Figure 9 As shown, firstly, the difference image calculation unit 109 calculates the difference image between the NeRF output image and the super-resolution processed image obtained by performing super-resolution processing on the NeRF output image (step S101). That is, the difference image calculation unit 109 calculates the difference image between the new viewpoint image and the image after super-resolution processing of each viewpoint of the multiple two-dimensional images included in the above-mentioned multi-viewpoint images for NeRF learning of NeRF model 104a.

[0103] The NeRF output image is stored in the NeRF output image storage unit 106, and the super-resolution processed image is stored in the super-resolution processed image storage unit 108. Furthermore, as... Figure 6 As already shown, the difference image calculated at this time tends to have a small amount of difference in flat regions with few edges (see the portion of the difference image with strong blackness), and a large amount of difference in non-flat regions with many edges (see the portion of the difference image with strong whiteness).

[0104] Then, the NeRF learning unit 103 uses the difference image calculated in step S101 and the camera device position and posture input from the camera device position and posture input unit 101 as input to perform NeRF learning (step S102).

[0105] Note that at this point, a completely different neural network than NeRF model 104a can be used to perform NeRF learning of the difference images. On the other hand, as... Figure 10 As shown, a model using the common existence probability of NeRF model 104a can be created by adding new parameters to the original NeRF model 104a (see “5” Map_Color and “6” Map_VD in the figure). The neural network is designed to allow it to perform transfer learning. When performing transfer learning, the probability of the existence of objects in three-dimensional space is already learned. Therefore, learning can converge at a high speed.

[0106] exist Figure 10 In the example, This is graph basis information. Furthermore, It is angle-dependent graph basis information, and corresponds to the "VD component". Furthermore, This is angle-dependent color information, and corresponds to the "basis function" (see "7" Map_Basis" in the figure). In Figure 10For ease of description, only one component is shown in the text. As VD components, but two or more components can also be used. That is, the parameters newly added to the difference graph model 110a indicate the appearance of the differences when the angle-independent graph basis information is viewed from various angles. These new parameters can also be described as representing the differences before and after super-resolution processing as three-dimensional semi-transparent edge information using MPI.

[0107] Then, the NeRF learning unit 103 outputs the difference map model 110a of the NeRF model as the learned difference image to the difference map model storage unit 110 for storing the difference map model 100a (step S103).

[0108] At this point, the NeRF learning unit 103 converts the difference map model 110a into a format suitable for real-time rendering and outputs the model. For example, in the case of NeX, the NeRF learning unit 103 will convert the color components that do not change according to the viewing direction. Color components that change according to the direction of the line of sight , ……and and color components , ……and Corresponding weighted basis function , ……and Output in MPI format to the difference graph model storage unit 110.

[0109] Here, by sharing the probability of existence When performing reinforcement learning using the original NeRF model 104a, no new output is required. Furthermore, the MPI resolution in the difference map model 110a is not necessarily the same as the MPI resolution of the original NeRF model 104a. For example, the resolution may be 1 / 2, 1 / 4, or smaller than the original MPI resolution.

[0110] <2-3. Processing steps for real-time rendering using switching graphs>

[0111] Next, we will refer to the references that have already been referenced. Figure 8 and Figures 11 to 13 The process of real-time rendering using a switching graph, performed by the information processing device 100, is described. Figure 11 This is a flowchart illustrating the real-time rendering process performed by the information processing device 100 using a switching graph. Furthermore, Figure 12 An example of a difference plot is shown. Furthermore, Figure 13 An example of a switching graph is shown.

[0112] In real-time rendering processing using switching graphs, such as Figure 11 As shown, the NeRF estimation unit 105 estimates the NeRF output image and the difference map based on the NeRF model 104a in the NeRF model storage unit 104, the gaze direction obtained from the gaze direction input unit 113, and the difference map model 110a in the difference map model storage unit 110 (step S201). Then, the NeRF estimation unit 105 outputs the estimated NeRF output image to the NeRF output image storage unit 106 and outputs the difference map to the difference map storage unit 111. Figure 12 An example of the estimated difference plot is shown in the figure.

[0113] Subsequently, the switching map generation unit 112 generates a super-resolution processing switching map based on the difference map (step S202). First, the switching map generation unit 112 generates a switching map for each of the aforementioned image blocks by reducing the resolution of the difference map. Figure 13 An example of the switching diagram at this stage is shown in the figure.

[0114] Then, the switching image generation unit 112 performs threshold processing on the switching image of each of the aforementioned image blocks according to a threshold preset by the switching threshold setting unit 114 (e.g., 50% as mentioned above), and performs on / off settings for super-resolution processing for each of the image blocks. The image at this time corresponds to, for example, the previously referenced... Figure 5 The setting is "Switch Graph (50% On)". Note that the graph setting value does not have to be a binary value for on / off, but can be three or more settings corresponding to first super-resolution processing, second super-resolution processing, off, etc. Furthermore, the graph setting value can be a setting value indicating at least one of the processing target area and non-processing target areas.

[0115] Then, the super-resolution processing unit 107 performs super-resolution processing only on the valid areas in the switching map (that is, the image blocks set to be open) according to the setting value of the switching map (step S203).

[0116] Finally, the super-resolution processing unit 107 synthesizes the NeRF output image stored in the NeRF output image storage unit 106 and the super-resolution processed image (S204). At this time, the super-resolution processing unit 107 performs synthesis according to the difference map in the difference map storage unit 111, and outputs the synthesized image to the super-resolution processed image storage unit 108.

[0117] Here, in the synthesis, for example, it can be performed based on the values ​​in the difference plot as shown in the following formula (c). mix.

[0118] ... (c)

[0119] Here, x and y indicate the xy coordinates of the image. out is the output image. Furthermore, This is the difference map, nerf is the NeRF output image, and sr is the image after super-resolution processing. This is a function that performs scaling to match the scaling of Nerf and SR, and It is to combine the difference diagram with The function associated with the blending values ​​(0 to 1). In regions where no super-resolution processing was performed, the output is out = s(nerf). The composited actual image corresponds to, for example, the referenced image. Figure 5 The image after super-resolution processing (50%).

[0120] <<3. Revision>>

[0121] For example, in the processes described in the above embodiments of this disclosure, all or part of the processes described as automatically executed can be executed manually, or all or part of the processes described as manually executed can be executed automatically by known methods. Furthermore, unless otherwise specified, the processes, specific names, and information including various data and parameters shown in the above documents and figures can vary arbitrarily. For example, the various types of information shown in each figure are not limited to the information shown.

[0122] Furthermore, each component of each device shown in the accompanying drawings is conceptual in function and is not necessarily physically configured as shown in the drawings. That is, the specific form of distribution and integration of each device is not limited to the form shown, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads, usage conditions, etc.

[0123] For example, the NeRF estimation unit 105 and the switching graph generation unit 112 can be integrated into a single processing unit.

[0124] Furthermore, the above-described embodiments of this disclosure can be appropriately combined to the extent that they do not contradict the processing content. Additionally, the order of each step shown in the sequence diagrams or flowcharts of this embodiment can be appropriately varied.

[0125] In the embodiments described above in this disclosure, the super-resolution processed image is displayed on a display device (e.g., a monitor). However, the monitor may be a light field display that uses parallax images to express stereoscopic vision. In this case, the gaze direction input unit 113 receives the user's gaze direction as input from, for example, a camera device mounted on the light field display. Furthermore, the NeRF model 104a and the difference map model 110a render a left-eye image for the user's left eye and a right-eye image for the user's right eye.

[0126] <<4. Hardware Configuration>>

[0127] For example, the information processing apparatus 100 according to the above embodiments of this disclosure uses, for example... Figure 14 The computer 1000 configured as shown in the figure is used to implement this. Figure 14 This is a hardware configuration diagram illustrating an example of a computer 1000 that implements the functions of the information processing device 100. The computer 1000 includes a CPU 1100, RAM 1200, ROM 1300, auxiliary storage device 1400, communication interface 1500, and input / output interface 1600. The units of the computer 1000 are connected to each other via a bus 1050.

[0128] The CPU 1100 operates based on programs stored in ROM 1300 or auxiliary storage device 1400 and controls each unit. For example, the CPU 1100 deploys programs stored in ROM 1300 or auxiliary storage device 1400 to RAM 1200 and executes processing corresponding to various programs.

[0129] ROM 1300 stores boot programs, such as the Basic Input / Output System (BIOS) executed by CPU 1100 when computer 1000 starts, and programs that depend on the hardware of computer 1000.

[0130] The auxiliary storage device 1400 is a computer-readable recording medium that non-transitorily records programs executed by the CPU 1100, data used by the programs, etc. Specifically, the auxiliary storage device 1400 is a recording medium that records a program used as an example of program data 1450 according to this embodiment.

[0131] Communication interface 1500 is an interface used to connect computer 1000 to external network 1550. For example, CPU 1100 receives data from another device or transmits data generated by CPU 1100 to another device via communication interface 1500.

[0132] Input / output interface 1600 is an interface for connecting input / output device 1650 and computer 1000. For example, CPU 1100 receives data from input devices (such as keyboards and mice) via input / output interface 1600. Furthermore, CPU 1100 transmits data to output devices (such as monitors, speakers, or printers) via input / output interface 1600. Additionally, input / output interface 1600 can be used as a media interface for reading programs recorded on a predetermined recording medium (medium). The medium is, for example, an optical recording medium (such as a digital multifunction disc (DVD) or phase-change rewritable disc (PD)), a magneto-optical recording medium (such as a magneto-optical disc (MO)), magnetic tape, a magnetic recording medium, semiconductor memory, etc.

[0133] For example, when the computer 1000 is used as an information processing device 100, the CPU 1100 of the computer 1000 performs the aforementioned functions of the control unit by executing a program loaded on the RAM 1200. Furthermore, the auxiliary storage device 1400 stores the program according to this embodiment, various models including the NeRF model 104a and the difference map model 110a, and various data including the NeRF output image and the image after super-resolution processing. The CPU 1100 reads program data 1450 from the auxiliary storage device 1400 and executes the program data; however, as another example, these programs can be obtained from another device via an external network 1550.

[0134] <<5. Conclusion>>

[0135] As described above, according to embodiments of this disclosure, the information processing apparatus 100 includes a NeRF estimation unit 105 (corresponding to an example of "estimation unit") and a super-resolution processing unit 107. The NeRF estimation unit 105 uses a NeRF model 104a (corresponding to an example of "first model") to estimate a NeRF output image (corresponding to an example of "new viewpoint image") corresponding to any viewing direction. The NeRF model 104a is learned to estimate the NeRF output image based on a multi-viewpoint image comprising multiple two-dimensional images. The super-resolution processing unit 107 performs super-resolution processing on the estimated NeRF output image. Furthermore, the NeRF estimation unit 105 uses a difference map model 110a (corresponding to an example of "second model") to estimate a processing target region for which super-resolution processing is to be performed and a non-processing target region for which super-resolution processing is not performed. The difference map model 110a is learned to estimate the processing target region and the non-processing target region for each NeRF output image based on the viewing direction. Furthermore, the super-resolution processing unit 107 performs super-resolution processing on the estimated processing target region but does not perform super-resolution processing on the estimated non-processing target region. Therefore, super-resolution processing can be accelerated when performing super-resolution processing on the rendering results of the new viewpoint image.

[0136] Although embodiments of the present disclosure have been described above, the technical scope of the present disclosure is not limited to the embodiments described above, and various modifications can be made without departing from the essential points of the present disclosure. Furthermore, different embodiments and modified components can be appropriately combined.

[0137] Furthermore, the effects of the embodiments described in this specification are merely illustrative and not limiting, and other effects may be provided.

[0138] This technology can also have the following configurations.

[0139] (1) An information processing device, comprising:

[0140] An estimation unit, wherein the estimation unit uses a first model to estimate a new viewpoint image, the first model being learned to estimate the new viewpoint image corresponding to any viewing direction based on a multi-viewpoint image comprising multiple two-dimensional images; and

[0141] A super-resolution processing unit performs super-resolution processing on the estimated new viewpoint image, wherein...

[0142] The estimation unit uses a second model to estimate the processing target region for which the super-resolution processing will be performed and the non-processing target region for which the super-resolution processing will not be performed. The second model is learned to estimate the processing target region and the non-processing target region for each of the new viewpoint images based on the viewing direction.

[0143] The super-resolution processing unit performs the super-resolution processing on the estimated target region, but does not perform the super-resolution processing on the estimated non-target region.

[0144] (2) The information processing device according to (1) further includes:

[0145] The learning unit learns the first model based on the multi-viewpoint images and learns the second model based on the difference image between the new viewpoint image of each viewpoint of the plurality of two-dimensional images and the super-resolution processed image.

[0146] (3) The information processing device according to (2), wherein,

[0147] The estimation unit uses the second model to estimate the difference image corresponding to the viewing direction, and generates map information indicating at least one of the processing target region and the non-processing target region based on the estimated difference image.

[0148] The super-resolution processing unit performs the super-resolution processing on the processing target region indicated by the image information.

[0149] (4) The information processing device according to (3), wherein,

[0150] The estimation unit generates the graph information, in which regions with relatively large differences in the difference image estimated using the second model are set as the processing target regions, and regions with relatively small differences are set as the non-processing target regions.

[0151] (5) The information processing device according to (4), wherein,

[0152] The estimation unit sets the processing target region and the non-processing target region for each predetermined pixel unit.

[0153] (6) An information processing apparatus according to any one of (2) to (5), wherein,

[0154] The estimation unit sets the processing target region based on a threshold, which indicates the proportion of the processing target region in the difference image estimated using the second model.

[0155] (7) An information processing apparatus according to any one of (2) to (6), wherein,

[0156] The learning unit uses transfer learning based on the first model to learn the second model.

[0157] (8) The information processing device according to (7), wherein,

[0158] The first model and the second model cache intermediate data in the form of a multi-layer image in which multiple images are arranged in the depth direction from the viewpoint position of the two-dimensional image, and share the probability of the existence of objects at the intersection of the light rays tracked from the line of sight and the multi-layer image.

[0159] (9) An information processing apparatus according to any one of (1) to (8), wherein,

[0160] The first model and the second model are models used to synthesize radiation fields using neural networks.

[0161] (10) An information processing apparatus according to any one of (1) to (9), wherein,

[0162] The super-resolution processing unit synthesizes the super-resolution processed image with the new viewpoint image and outputs the synthesized image.

[0163] (11) An information processing method, comprising the following steps:

[0164] A first model is used to estimate a new viewpoint image, the first model being learned to estimate the new viewpoint image corresponding to any viewing direction based on a multi-viewpoint image comprising multiple two-dimensional images;

[0165] Perform super-resolution processing on the estimated new viewpoint image;

[0166] A second model is used to estimate the target region for which the super-resolution processing will be performed and the non-target region for which the super-resolution processing will not be performed. This second model is learned to estimate the target region and the non-target region for each of the new viewpoint images based on the viewing direction.

[0167] The super-resolution processing is performed on the estimated target region, but not on the estimated non-target region.

[0168] (12) A computer-readable non-transitory storage medium storing a program for causing a computer to perform the following steps:

[0169] A first model is used to estimate a new viewpoint image, the first model being learned to estimate the new viewpoint image corresponding to any viewing direction based on a multi-viewpoint image comprising multiple two-dimensional images;

[0170] Perform super-resolution processing on the estimated new viewpoint image;

[0171] A second model is used to estimate the target region for which the super-resolution processing will be performed and the non-target region for which the super-resolution processing will not be performed. The second model is learned to estimate the target region for processing and the non-target region for each of the new viewpoint images based on the viewing direction.

[0172] The super-resolution processing is performed on the estimated target region, but not on the estimated non-target region.

[0173] List of reference numerals

[0174] 50 camera devices

[0175] 100 Information Processing Device

[0176] 101 Camera device position and posture input unit

[0177] 102 multi-view image input units

[0178] 103 NeRF learning units

[0179] 104 NeRF model storage units

[0180] 104a NeRF model

[0181] 105 NeRF estimation units

[0182] 106 NeRF output image storage units

[0183] 107 super-resolution processing units

[0184] 108 super-resolution image storage units

[0185] 109 difference image computing units

[0186] 110 difference graph model storage units

[0187] 110a Difference Map Model

[0188] 111 difference graph storage units

[0189] 112 Switching Diagram Generation Unit

[0190] 113 Eye Direction Input Unit

[0191] 114 Switching Threshold Setting Unit

[0192] Probability of existence

Claims

1. An information processing apparatus, comprising: An estimation unit uses a first model to estimate a new viewpoint image, the first model being learned to estimate the new viewpoint image corresponding to any viewing direction based on a multi-viewpoint image comprising multiple two-dimensional images; as well as A super-resolution processing unit performs super-resolution processing on the estimated new viewpoint image, wherein... The estimation unit uses a second model to estimate the processing target region for which the super-resolution processing will be performed and the non-processing target region for which the super-resolution processing will not be performed. The second model is learned to estimate the processing target region and the non-processing target region for each of the new viewpoint images based on the viewing direction. The super-resolution processing unit performs the super-resolution processing on the estimated target region and does not perform the super-resolution processing on the estimated non-target region.

2. The information processing apparatus according to claim 1, further comprising: The learning unit learns the first model based on the multi-viewpoint images and learns the second model based on the difference image between the new viewpoint image of each viewpoint of the plurality of two-dimensional images and the super-resolution processed image.

3. The information processing apparatus according to claim 2, wherein, The estimation unit uses the second model to estimate the difference image corresponding to the viewing direction, and generates map information indicating at least one of the processing target region and the non-processing target region based on the estimated difference image. The super-resolution processing unit performs the super-resolution processing on the processing target region indicated by the image information.

4. The information processing apparatus according to claim 3, wherein, The estimation unit generates the graph information, in which regions with relatively large differences in the difference image estimated using the second model are set as the processing target regions, and regions with relatively small differences are set as the non-processing target regions.

5. The information processing apparatus according to claim 4, wherein, The estimation unit sets the processing target region and the non-processing target region for each predetermined pixel unit.

6. The information processing apparatus according to claim 2, wherein, The estimation unit sets the processing target region based on a threshold, which indicates the proportion of the processing target region in the difference image estimated using the second model.

7. The information processing apparatus according to claim 2, wherein, The learning unit uses transfer learning based on the first model to learn the second model.

8. The information processing apparatus according to claim 7, wherein, The first model and the second model cache intermediate data in the form of a multi-layer image in which multiple images are arranged in the depth direction from the viewpoint position of the two-dimensional image, and share the probability of the existence of objects at the intersection of the light rays tracked from the line of sight and the multi-layer image.

9. The information processing apparatus according to claim 1, wherein, The first model and the second model are models for synthesizing radiation fields using neural networks.

10. The information processing apparatus according to claim 1, wherein, The super-resolution processing unit synthesizes the super-resolution processed image with the new viewpoint image and outputs the synthesized image.

11. An information processing method, comprising the following steps: A first model is used to estimate a new viewpoint image, the first model being learned to estimate the new viewpoint image corresponding to any viewing direction based on a multi-viewpoint image comprising multiple two-dimensional images; Perform super-resolution processing on the estimated new viewpoint image; A second model is used to estimate the target region to be processed and the non-processed target region to be processed, and the second model is learned to estimate the target region to be processed and the non-processed target region according to the viewing direction for each of the new viewpoint images. as well as The super-resolution processing is performed on the estimated target region, but not on the estimated non-target region.

12. A computer-readable non-transitory storage medium storing a program for causing a computer to perform the following steps: A first model is used to estimate a new viewpoint image, the first model being learned to estimate the new viewpoint image corresponding to any viewing direction based on a multi-viewpoint image comprising multiple two-dimensional images; Perform super-resolution processing on the estimated new viewpoint image; A second model is used to estimate the target region to be processed and the non-processed target region to be processed, and the second model is learned to estimate the target region to be processed and the non-processed target region according to the viewing direction for each of the new viewpoint images. as well as The super-resolution processing is performed on the estimated target region, but not on the estimated non-target region.

Citation Information

Patent Citations

  • Super-resolution frame selection apparatus, super-resolution device, and program

    JP2018023034A