Image generation device, image generation method, and program

The image generation device with a learned optical learning inference unit and reflection direction calculator addresses the inflexibility of NeRF by allowing control of illumination and gloss, enabling flexible two-dimensional image generation from fixed lighting environments.

JP2025109387APending Publication Date: 2025-07-25NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024003243
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing three-dimensional modeling technologies like NeRF require fixed lighting environments for capturing teacher images, making it difficult to reset lighting settings and necessitate professional cameramen, limiting flexibility in generating two-dimensional images of free viewpoints.

Method used

An image generation device equipped with a learned optical learning inference unit and a reflection direction calculator that allows control of illumination environment and gloss by inputting azimuth information and angle control parameters, enabling generation of two-dimensional images from free viewpoints.

Benefits of technology

Enables control of illumination environment and gloss on two-dimensional images generated from fixed lighting conditions, facilitating flexible image generation without requiring expert intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025109387000001_ABST
    Figure 2025109387000001_ABST
Patent Text Reader

Abstract

To provide an image generation device which contributes to making it possible to control an illumination environment and glossiness on a two-dimensional image when generating the two-dimensional image of a three-dimensional scene with a free viewpoint after photographing a teacher image under a fixed illumination environment.SOLUTION: An image generation device for generating a two-dimensional image of a three-dimensional scene with a free viewpoint which has already learned a space structure and optical information, includes a learned optical learning inference unit and a reflection direction calculator. The optical learning inference unit accepts direction information and extracts optical information by using at least the direction information. The reflection direction calculator accepts an angle control parameter as input, generates the direction information by calculating an observation direction and the angle control parameter to at least a partial region in a re-construction space based on the learned space structure, and outputs the direction information to the optical learning inference unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image generation device, an image generation method, and a program.

Background Art

[0002] Regarding the display of a two-dimensional image of a three-dimensional scene, the following documents can be cited.

[0003] Patent Document 1 relates to an image processing apparatus that performs three-dimensional modeling by machine learning.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] The following analysis is provided by the inventor of the present invention.

[0006] NeRF (Neural Radiance Fields), which is a learning-based three-dimensional modeling technology, is explosively spreading in a wide field dealing with three-dimensional (3D) shapes. In NeRF, a two-dimensional image of a three-dimensional scene captured using a camera is input as a teacher image. Usually, however, the shooting environment of the teacher image, such as the position of lighting in the three-dimensional space at the time of shooting, is fixed.

[0007] In addition, setting the shooting environment, for example, lighting settings or camera angle settings, etc., takes time, so it is difficult to reset the shooting environment and reshoot, and in many cases, an expert such as a professional cameraman is required.

[0008] The present invention contributes to enabling control of the illumination environment and gloss on a two-dimensional image when generating a two-dimensional image of a free viewpoint of a three-dimensional scene after photographing a teacher image in a fixed illumination environment, and aims to provide an image generation device, an image generation method, and a program.

Means for Solving the Problems

[0009] According to a first aspect of the present invention, there is provided an image generation device that generates a two-dimensional image of a free viewpoint of a three-dimensional scene, which has learned a spatial structure and optical information, including the learned optical learning inference unit and a reflection azimuth calculator, wherein the optical learning inference unit receives azimuth information, extracts optical information using at least the azimuth information, and the reflection azimuth calculator receives an angle control parameter as an input, performs an operation of the observation azimuth and the angle control parameter on at least a partial region in a reconstruction space based on the learned spatial structure to generate the azimuth information, and outputs the azimuth information to the optical learning inference unit. An image generation device can be provided.

[0010] According to a second aspect of the present invention, a computer of an image generation device that generates a two-dimensional image of a free viewpoint of a three-dimensional scene, which has learned a spatial structure and optical information, receives azimuth information, extracts optical information using at least the azimuth information, receives an angle control parameter as an input, performs an operation of the observation azimuth and the angle control parameter on at least a partial region in a reconstruction space based on the learned spatial structure to generate the azimuth information, and provides an image generation method including this. This method is associated with a specific machine, namely a computer that executes the above method.

[0011] According to a third aspect of the present invention, a computer of an image generation device that generates a two-dimensional image of a free viewpoint of a three-dimensional scene, which has learned a spatial structure and optical information, a process of receiving orientation information, a process of extracting optical information using at least the orientation information, a process of receiving an angle control parameter as an input, a process of performing an operation of an observation orientation and the angle control parameter on at least a partial region in a reconstructed space based on the learned spatial structure to generate the orientation information, A program for causing the above to be executed can be provided.

[0012] These programs can be recorded on a computer-readable storage medium. The storage medium can be a non-transitory one such as a semiconductor memory, a hard disk, a magnetic recording medium, an optical recording medium, etc. The present invention can also be embodied as a computer program product.

Advantages of the Invention

[0013] According to the present invention, after photographing a teacher image in a fixed illumination environment, when generating a two-dimensional image of a free viewpoint of a three-dimensional scene, it contributes to making it possible to control the illumination environment and gloss on the two-dimensional image, and an image generation device, an image generation method, and a program can be provided.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

[0015] In the present disclosure, the drawings may be associated with one or more embodiments. Further, each of the embodiments described below can be combined with other embodiments as appropriate, and the present invention is not limited by each embodiment.

[0016] First, an outline of an embodiment will be described with reference to the drawings. Note that the reference numerals in the drawings appended to this outline are for convenience of each element as an example for assisting understanding, and are not intended to limit the present invention to the illustrated aspects. Also, the connection lines between blocks in the drawings and the like referred to in the following description include both bidirectional and unidirectional ones. The one-way arrow schematically shows the flow of main signals (data), and does not exclude bidirectionality.

[0017] FIG. 1 is a block diagram showing an example of the configuration of an image generation device according to the present disclosure. Referring to FIG. 1, an image generation device 101 generates a two-dimensional image of a free viewpoint of a three-dimensional scene that has learned a spatial structure and optical information. The image generation device 101 includes a learned optical learning inference unit 102 and a reflection direction calculator 103.

[0018] The optical learning inference unit 102 receives the orientation information 106 and extracts optical information 107 using at least the orientation information 106.

[0019] The reflection direction calculator 103 receives the angle control parameter 105 as an input, and performs an operation of the observation direction 104 and the angle control parameter 105 on at least a partial region in the reconstructed space based on the learned spatial structure to generate the orientation information 106. Then, the reflection direction calculator 103 outputs the orientation information 106 to the optical learning inference unit 102.

[0020] When the angle control parameter 105 input to the reflection direction calculator 103 is changed, the illumination environment and gloss can be changed on the two-dimensional image when generating a two-dimensional image of a free viewpoint of the three-dimensional scene.

[0021] This is because, for the observation direction 104, by inputting the orientation information 106 obtained by calculating the angle control parameter 105 to the learned optical learning inference unit 102, the viewpoint for generating a two-dimensional image of a free viewpoint, that is, the reflected light when viewed from another direction with respect to the free viewpoint, that is, the optical information 107, or the RGB value, can be extracted from the learned optical learning inference unit 102.

[0022] Therefore, according to one embodiment, it is possible to provide an image generation device, an image generation method, and a program that contribute to making it possible to control the illumination environment and gloss on a two-dimensional image when generating a two-dimensional image of a free viewpoint of a three-dimensional scene after photographing a teacher image under a fixed illumination environment.

[0023] [First Embodiment] Next, the first embodiment will be described in detail with reference to the drawings. FIG. 3 is a block diagram showing an example of the configuration of an image generation system according to the present disclosure. Further, FIG. 2 is a block diagram showing an example of the configuration of a conventional image generation system. Referring to FIG. 2, the conventional image generation system 10 includes an image generation device 100, a rendering device 200, a two-dimensional camera (also referred to as a sensor) 300, and a two-dimensional image display device 400. The two-dimensional camera 300 acquires a two-dimensional image of the three-dimensional scene 1000.

[0024] First, an example of the operation of the conventional image generation system 10 during learning and inference will be described. During the learning of the conventional image generation system 10, the image generation device 100 learns the spatial structure and optical information using the two-dimensional image of the three-dimensional scene acquired by the two-dimensional camera 300 as teacher data by the optical information learning inference unit 110 and the error detection unit 150. Although the two-dimensional camera 300 is connected to the error detection unit 150 by a line, it is not necessary for the two-dimensional camera 300 and the error detection unit 150 to be directly connected, and any configuration is acceptable as long as the image captured by the two-dimensional camera 300 is input to the error detection unit 150.

[0025] (1) Operation of the conventional image generation system 10 during learning First, an example of the operation during the learning of the conventional image generation system 10 will be described. During learning, a two-dimensional image (also referred to as a camera image) 301 captured by a two-dimensional camera 300 of a three-dimensional scene 1000 is input as a teacher image to an error detection unit 150. Also, observation information 302 obtained by capturing the two-dimensional image 301, for example, the camera origin and the camera orientation, is input to an observation information acquisition unit 120. The observation information 302 input to the observation information acquisition unit 120 is output from the observation information acquisition unit 120 as an observation viewpoint 1201 to the L input of a switch (SW or also referred to as a switching device) 190. Note that the switch 190 outputs the information input to the L input as output 1901 during learning, and outputs the information input to the G input as output 1901 during inference. Therefore, the switch 190 outputs the observation information 302 input to the L input as output 1901 to the spatial sampling unit 135 of a sampler 130 during learning. Note that the switch 190 may be configured to provide the observation viewpoint 1201 input to the L input as output 1901 to the spatial sampling unit 135 during learning, and provide the free viewpoint 1411 output from a free viewpoint input unit 141 as output 1901 to the spatial sampling unit 135 during inference, and it does not have to be a specific switch.

[0026] To the sampler 130, information 201 regarding pixels to be rendered (also referred to as volume rendering) is sent from the rendering device 200. The spatial sampling unit 135 generates the position (x, y, z) 131 and the observation direction (θ, φ) 132 necessary to generate the pixels of the two-dimensional image 202 corresponding to the information 201 regarding pixels. It outputs the position (x, y, z) 131 as three-dimensional position information to the density learning inference unit 111 of the optical information learning inference device 110, and outputs the observation direction (θ, φ) 132 as direction information to the optical learning inference unit 112 of the optical information learning inference device 110. Also, the density learning inference unit 111 outputs an intermediate representation 1112 to the optical learning inference unit 112. Note that the density learning inference unit 111 is a subordinate concept of the intermediate representation learning inference unit that performs learning and inference of the intermediate representation, and any configuration may be used as long as it outputs the intermediate representation 1112. Note that a configuration may be used to convert the position 131 and the observation direction 132 output by the sampler 130 into a position embedding representation and a direction embedding representation. In this case, the same configuration shall be used also during inference.

[0027] The density learning inference unit 111 uses the position (x, y, z) 131 to generate a density 1111 and outputs the density 1111 to the rendering device 200. Also, the optical learning inference unit 112 uses the observation direction (θ, φ) 132 and the intermediate representation 1112 to generate an RGB value 1121 and outputs the RGB value 1121 to the rendering device 200.

[0028] The rendering device 200 uses the density 1111 and the RGB value 1121 to generate the pixels of the two-dimensional image 202 corresponding to the information 201 regarding pixels. In this way, all the pixels of the two-dimensional image 202 are generated, and the two-dimensional image 202 is rendered and output. The rendered two-dimensional image 202 is sent to the two-dimensional image display device 400 and the error detection unit 150.

[0029] The error detection unit 150 detects the error between the two-dimensional image 202 rendered by the rendering device 200 and the two-dimensional image 301 of the three-dimensional scene acquired by the two-dimensional camera 300. Based on the error 1501 output by the error detection unit 150, the optical information learning inference device 110 adjusts the parameters of the density learning inference unit 111 and the parameters of the optical learning inference unit 112 to generate a learned model. Note that the density learning inference unit 111 and the optical learning inference unit 112 may include a neural network. Also, the adjustment of the parameters of the density learning inference unit 111 and the adjustment of the parameters of the optical learning inference unit 112 may be executed by adjusting the parameters of the neural network included in the density learning inference unit 111 and the optical learning inference unit 112. The generated learned model includes a first learned model in the density learning inference unit 111 and a second learned model in the optical learning inference unit 112.

[0030] (2) Operation during inference of the conventional image generation system 10 Next, an example of the operation during inference of the conventional image generation system 10, that is, the operation during generation of a two-dimensional image of a free viewpoint of the three-dimensional scene 1000, will be described. During inference of the conventional image generation system 10, as an example, as the user's observation viewpoint with respect to the three-dimensional scene 1000, the free viewpoint 1410 is input by the user to the free viewpoint input unit 141 of the GUI (Graphic User Interface) input unit 140 of the image generation system 10, for example. The free viewpoint 1410 input to the free viewpoint input unit 141 is output from the free viewpoint input unit 141 as the free viewpoint 1411 to the G input of the switch 190. The switch 190 outputs the free viewpoint 1411 input to the G input as the output 1901 to the spatial sampling unit 135 of the sampler 130 during inference. Note that the free viewpoint 1411 is the same as the free viewpoint 1410.

[0031] Information 201 regarding the pixels to be rendered is sent from the rendering device 200 to the sampler 130, and the spatial sampling unit 135 generates the positions (x, y, z) 131 and the observation azimuth (θ, φ) 132 necessary for generating the pixels of the two-dimensional image 202 corresponding to the information 201 regarding the pixels.

[0032] Next, the spatial sampling unit 135 of the sampler 130 outputs the position (x, y, z) 131 as three-dimensional position information to the first learned model of the density learning inference unit 111 of the optical information learning inference device 110. Also, the spatial sampling unit 135 outputs the observation azimuth (θ, φ) 132 as azimuth information to the second learned model of the optical learning inference unit 112 of the optical information learning inference device 110. Further, the first learned model of the density learning inference unit 111 outputs the intermediate representation 1112 to the optical learning inference unit 112.

[0033] The first learned model of the density learning inference unit 111 generates the density 1111 using the position (x, y, z) 131 and outputs the density 1111 to the rendering device 200. Also, the second learned model of the optical learning inference unit 112 generates the RGB value 1121 using the observation azimuth (θ, φ) 132 and the intermediate representation 1112 and outputs the RGB value 1121 to the rendering device 200.

[0034] The rendering device 200 generates the pixels of the two-dimensional image 202 corresponding to the information 201 regarding the pixels using the density 1111 and the RGB value 1121. In this way, all the pixels of the two-dimensional image 202 are generated, the two-dimensional image 202 is rendered and output. The rendered two-dimensional image 202 is sent to the two-dimensional image display device 400. In this way, the two-dimensional image of the three-dimensional scene 1000 from the free viewpoint 1410 input by the user is displayed on the two-dimensional image display device 400.

[0035] (3) Configuration and Operation of the Image Generation System According to the First Embodiment Next, the configuration and operation of the image generation system according to the first embodiment will be described with reference to FIG. 3. In FIG. 3, the components denoted by the same reference numerals as in FIG. 2 indicate the same components.

[0036] Referring to FIG. 3, the image generation system 10A according to the present disclosure includes an image generation device 100A, a rendering device 200, a two-dimensional camera (also referred to as a sensor) 300, and a two-dimensional image display device 400. The two-dimensional camera 300 acquires a two-dimensional image of the three-dimensional scene 1000. The image generation device 100A according to the present disclosure includes a reflection direction calculator 160 and a switch 191 with respect to the image generation device 100 of the conventional image generation system 10. The switch 191 outputs the information input to the L input during learning and outputs the information input to the G input during inference. Further, the GUI input unit 140 of the image generation device 100A of the image generation system 10A according to the present disclosure includes an angle control parameter input unit 142.

[0037] (4) Operation during learning of the image generation system according to the first embodiment During the learning of the image generation system according to the first embodiment, the L input of the switch 191 of the image generation device 100A of the image generation system 10A according to the present disclosure shown in FIG. 3 is output from the switch 191. That is, the observation direction 132 is input to the L input of the switch 191 and output from the switch 191 to the optical learning inference unit 112 as the direction information 1911. The operation during learning of the first embodiment is the same as the operation during learning of the conventional image generation system 10 described with reference to FIG. 2, except that the observation direction 132 is output from the switch 191 to the optical learning inference unit 112 as the direction information 1911 via the switch 191. Note that during learning, the direction information 1911 is the same as the observation direction 132. Note that a configuration for converting the position 131 and the observation direction 132 output by the sampler 130 into a position embedding representation and a direction embedding representation may be used.

[0038] (5) Operation during inference of the image generation system according to the first embodiment An example of the operation during inference of the image generation system according to the first embodiment, that is, the operation during generation of a free-viewpoint two-dimensional image of the three-dimensional scene 1000, will be described with reference to an example of the configuration of the image generation system 10A according to the present disclosure shown in FIG. 3.

[0039] During the inference of the image generation system according to the first embodiment, as an example, as the user's observation viewpoint for the three-dimensional scene 1000, the free viewpoint 1410 is input by the user to the free viewpoint input unit 141 of the GUI (Graphic User Interface) input unit 140 of the image generation system 10A according to the present disclosure shown in FIG. 3, for example. The free viewpoint 1410 input to the free viewpoint input unit 141 is output from the free viewpoint input unit 141 to the G input of the switch 190 as the free viewpoint 1411. During inference, the switch 190 outputs the free viewpoint 1411 input to the G input as the output 1901 to the spatial sampling unit 135 of the sampler 130. Note that the free viewpoint 1411 is the same as the free viewpoint 1410.

[0040] Information 201 regarding the pixels to be rendered is sent from the rendering device 200 to the sampler 130, and the spatial sampling unit 135 of the sampler 130 generates the positions (x, y, z) 131 and the observation directions (θ, φ) 132 necessary to generate the pixels of the two-dimensional image 202 corresponding to the information 201 regarding the pixels.

[0041] Next, the spatial sampling unit 135 of the sampler 130 outputs the position (x, y, z) 131 as three-dimensional position information to the first learned model of the density learning inference unit 111 of the optical information learning inference device 110. Also, the spatial sampling unit 135 outputs the observation direction 132 to the reflection direction calculator 160.

[0042] On the other hand, the angle control parameter 1420 is input to the angle control parameter input unit 142 of the GUI input unit 140, and is output from the angle control parameter input unit 142 to the reflection direction calculator 160 as the angle control parameter 1421. Note that the angle control parameter 1421 is the same as the angle control parameter 1420.

[0043] The angle control parameter 1421 input to the reflection azimuth calculator 160 is calculated with the observation azimuth 132 input from the spatial sampling unit 135 of the sampler 130, and the calculation output 1601 of the reflection azimuth calculator 160 is input to the G input of the switch 191. The calculation output 1601 is output from the switch 191 as azimuth information 1911 to the second learned model of the optical learning inference unit 112 of the optical information learning inference device 110. Note that at the time of inference, the azimuth information 1911 is the same as the calculation output 1601.

[0044] In addition, when a configuration for converting the output position 131 and the observation azimuth 132 output by the sampler 130 into a position embedding representation and an azimuth embedding representation, respectively, is used, as an example, a reflection azimuth calculator 160 capable of performing calculations in the azimuth embedding representation is used to calculate the angle control parameter 1421 in the azimuth embedding representation with the observation azimuth 132 in the azimuth embedding representation, and the azimuth information 1911 in the azimuth embedding representation input to the optical learning inference unit 112 may be generated. Further, in the case of the azimuth embedding representation, the reflection azimuth calculator 160 may include an array concatenation operation or the like.

[0045] Note that the operation of the first learned model of the density learning inference unit 111, the operation of the second learned model of the optical learning inference unit 112, and the operation of the rendering device 200 are the same as the operation at the time of inference of the conventional image generation system 10 described with reference to FIG. 2.

[0046] When the angle control parameter 1421 input to the reflection azimuth calculator 160 is changed, the illumination environment and the gloss can be changed on the two-dimensional image when generating a two-dimensional image of a free viewpoint of the three-dimensional scene.

[0047] This is because, for the observation direction 132, the angle control parameter 1421 is calculated, and the direction information 1911 output from the switch 191 is input to the second learned model of the learned optical learning inference unit 112, so that the viewpoint for generating a two-dimensional image of a free viewpoint, that is, the reflected light when viewed from another direction with respect to the free viewpoint, that is, the RGB value (or optical information) 1121, can be extracted from the second learned model of the learned optical learning inference unit 112.

[0048] Therefore, according to the first embodiment, after photographing a teacher image in a fixed illumination environment, when generating a two-dimensional image of a free viewpoint of a three-dimensional scene, it contributes to making it possible to control the illumination environment and gloss on the two-dimensional image, and an image generation apparatus, an image generation method, and a program can be provided.

[0049] [Second Embodiment] Next, the second embodiment will be described in detail with reference to the drawings. FIG. 4 is a block diagram showing an example of the configuration of an image generation system according to the present disclosure. In FIG. 4, components denoted by the same reference numerals as those in FIG. 3 represent the same components.

[0050] Referring to FIG. 4, an image generation system 10B according to the present disclosure includes an image generation apparatus 100B, a rendering apparatus 200, a two-dimensional camera (also referred to as a sensor) 300, and a two-dimensional image display apparatus 400. The two-dimensional camera 300 acquires a two-dimensional image of a three-dimensional scene 1000. The image generation apparatus 100B according to the present disclosure further includes a region determination unit 170 and an execution determination unit 180 with respect to the image generation apparatus 100A according to the present disclosure described in FIG. 3. In addition, the GUI input unit 140 of the image generation apparatus 100B further includes a valid region designation input unit 143 with respect to the GUI input unit 140 of the image generation apparatus 100A described in FIG. 3.

[0051] The operation during learning in the second embodiment is the same as the operation during learning in the first embodiment described with reference to FIG. 3. Note that a configuration may be used to convert the position 131 and the observation direction 132 output by the sampler 130 into a position embedding expression and an orientation embedding expression, respectively.

[0052] An example of the operation during inference of the image generation system in the second embodiment, that is, the operation during generation of a free viewpoint two-dimensional image of the three-dimensional scene 1000, will be described with reference to an example of the configuration of the image generation system 10B according to the present disclosure in FIG. 4.

[0053] During inference of the image generation system in the second embodiment, as an example, the free viewpoint 1410 is input by the user as an observation viewpoint for the three-dimensional scene 1000 to the free viewpoint input unit 141 of the GUI input unit 140 of the image generation system 10B according to the present disclosure shown in FIG. 4, for example. The free viewpoint 1410 input to the free viewpoint input unit 141 is output from the free viewpoint input unit 141 as the free viewpoint 1411 to the G input of the switch 190. The switch 190 outputs the free viewpoint 1411 input to the G input as the output 1901 to the spatial sampling unit 135 of the sampler 130 during inference. Note that the free viewpoint 1411 is the same as the free viewpoint 1410.

[0054] Information 201 regarding the pixels to be rendered is sent from the rendering device 200 to the sampler 130, and the spatial sampling unit 135 of the sampler 130 generates the position (x, y, z) 131 and the observation direction (θ, φ) 132 necessary for generating the pixels of the two-dimensional image 202 corresponding to the information 201 regarding the pixels.

[0055] Next, the spatial sampling unit 135 of the sampler 130 outputs the position (x, y, z) 131 as three-dimensional position information to the first learned model of the density learning inference unit 111 of the optical information learning inference device 110. Also, the spatial sampling unit 135 outputs the observation direction 132 to the reflection direction calculator 160.

[0056] On the one hand, the angle control parameter 1420 is input to the angle control parameter input section 142 of the GUI input section 140, and is output from the angle control parameter input section 142 to the execution determination section 180 as the angle control parameter 1421.

[0057] Until the valid area designation 1431 is input, the area determination section 170 outputs an instruction indicating outside the area as the determination result 1701.

[0058] As an example, when an instruction indicating outside the area is input to the execution determination section 180 from the area determination section 170 as the determination result 1701, the execution determination section 180 outputs, as the output 1801, a parameter for which the reflection direction calculator 160 does not perform calculations, to the reflection direction calculator 160. Note that, as an example, when the reflection direction calculator 160 performs addition or subtraction, the parameter for which the reflection direction calculator 160 does not perform calculations is "0 (zero)", and as another example, when the reflection direction calculator 160 performs multiplication or division, the parameter is "1". When such a parameter for which calculations are not performed is input to the reflection direction calculator 160, the reflection direction calculator 160 outputs the observation direction 132 input from the spatial sampling section 135 of the sampler 130 as the calculation output 1601 to the G input of the switch 191.

[0059] As another example, when an instruction indicating outside the area is input to the execution determination section 180 from the area determination section 170 as the determination result 1701, the execution determination section 180 may output, as the output 1801, a specific control code or the like for controlling the reflection direction calculator 160 so that it does not execute calculations, to the reflection direction calculator 160. In this case, the reflection direction calculator 160 outputs the observation direction 132 input from the spatial sampling section 135 of the sampler 130 as the calculation output 1601 to the G input of the switch 191.

[0060] The calculation output 1601 is output from the switch 191 as the azimuth information 1911 to the second learned model of the optical learning inference section 112 of the optical information learning inference device 110. Note that during inference, the azimuth information 1911 is the same as the calculation output 1601.

[0061] In addition, when a configuration for converting the output position 131 and observation direction 132 of the sampler 130 into position-embedded representation and direction-embedded representation respectively is used, as an example, the reflection direction calculator 160 that can perform operations on the direction-embedded representation is used to calculate the angle control parameter 1421 in the direction-embedded representation with the observation direction 132 in the direction-embedded representation, and generate the direction information 1911 in the direction-embedded representation to be input to the optical learning inference unit 112. Also, in the case of the direction-embedded representation, the reflection direction calculator 160 may include operations such as concatenation of arrays. Further, when the position-embedded representation is used, the region determination unit 170 that can make determinations in the position-embedded representation is used to perform determinations using the valid region specification 1431 in the position-embedded representation.

[0062] Note that the operation of the first learned model of the density learning inference unit 111, the second learned model of the optical learning inference unit 112, and the rendering device 200 is the same as the operation during inference of the conventional image generation system 10 described with reference to FIG. 2.

[0063] In this way, the rendering device generates a two-dimensional image of the free viewpoint 1411 of the three-dimensional scene 1000 and displays it on the two-dimensional image display device 400. The two-dimensional image of the free viewpoint 1411 of the three-dimensional scene 1000 displayed on the two-dimensional image display device 400 corresponds to the reconstructed space based on the learned spatial structure.

[0064] Furthermore, on the two-dimensional image of the free viewpoint 1411 of the three-dimensional scene 1000 rendered by the rendering device 200 and displayed on the two-dimensional image display device 400, the user inputs a valid region specification 1430 to the valid region specification input section 143 of the GUI input section 140. The input valid region specification 1430 is output from the valid region specification input section 143 as the valid region specification 1431 to the region determination unit 170.

[0065] The position (x, y, z) 131 output by the sampler 130 is input to the area determination unit 170. The area determination unit 170 determines whether the position (x, y, z) 131 input from the sampler 130 is within the valid area specified by the valid area specification 1431. When the area determination unit 170 determines that the position (x, y, z) 131 is within the valid area specified by the valid area specification 1431, it outputs an instruction indicating within the area as the determination result 1701 to the execution permission determination unit 180. Also, when the area determination unit 170 determines that the position (x, y, z) 131 is not within the valid area specified by the valid area specification 1431, it outputs an instruction indicating outside the area as the determination result 1701 to the execution permission determination unit 180. As an example, the boundary of the area may be determined to be within the area. Alternatively, the boundary of the area may be determined to be outside the area.

[0066] When an instruction indicating within the area is input as the determination result 1701 from the area determination unit 170, the execution permission determination unit 180 outputs the input angle control parameter 1421 as the output 1801 to the reflection direction calculator 160. Note that the angle control parameter 1421 is the same as the angle control parameter 1420.

[0067] The output 1801 (i.e., the angle control parameter 1421) input to the reflection direction calculator 160 is calculated with the observation direction 132 input from the spatial sampling unit 135 of the sampler 130, and the calculation output 1601 of the reflection direction calculator 160 is input to the G input of the switch 191.

[0068] The calculation output 1601 is output from the switch 191 as the direction information 1911 to the second learned model of the optical learning inference unit 112 of the optical information learning inference device 110. Also, at the time of inference, the direction information 1911 is the same as the calculation output 1601.

[0069] Note that the operations of the first trained model of the density learning inference unit 111, the second trained model of the optical learning inference unit 112, and the rendering device 200 are the same as the operations during inference of the conventional image generation system 10 described with reference to FIG. 2.

[0070] In this way, by rendering a two-dimensional image of the free viewpoint 1411 of the three-dimensional scene again by the rendering device 200, the angle control parameter 1421 can be applied only to a specific region to generate a two-dimensional image of the free viewpoint 1411 of the three-dimensional scene.

[0071] Note that when the angle control parameter 1421 input to the reflection azimuth calculator 160 is changed as the output 1801 via the execution determination unit 180, the illumination environment and gloss can be changed on the two-dimensional image when generating a two-dimensional image of the free viewpoint of the three-dimensional scene.

[0072] This is because, for the observation azimuth 132, the angle control parameter 1421 is calculated, and the azimuth information 1911 output from the switch 191 is input to the second trained model of the trained optical learning inference unit 112, so that the viewpoint for generating a two-dimensional image of the free viewpoint, that is, the reflected light when viewed from another azimuth with respect to the free viewpoint, that is, the RGB value (or optical information) 1121, can be taken out from the second trained model of the trained optical learning inference unit 112.

[0073] Also, by specifying the effective region to which the angle control parameter 1421 is applied, when generating a two-dimensional image of the free viewpoint of the three-dimensional scene by applying the angle control parameter 1421 only to a specific effective region, the illumination environment and gloss can be changed only in a specific designated region on the two-dimensional image.

[0074] In the above description, the effective area has been described as a configuration specified by the user using the effective area specification 1430. However, on the two-dimensional image of the free viewpoint of the three-dimensional scene first rendered by the rendering device, a partial area including the portion clicked with a pointing device or the like may be automatically specified using a function or tool that automatically selects the partial area. In this case, for the automatically specified area, the rendering device 200 renders a two-dimensional image of the free viewpoint of the three-dimensional scene to which the angle control parameter 1421 is applied again, and the illumination environment and gloss can be changed only in a specific area on the two-dimensional image of the free viewpoint of the three-dimensional scene.

[0075] Also, if the user designates the entire area of the two-dimensional image of the free viewpoint of the three-dimensional scene as the effective area by the effective area specification 1430, the two-dimensional image of the free viewpoint of the generated three-dimensional scene to which the input angle control parameter 1421 is applied is the same as the two-dimensional image of the free viewpoint of the three-dimensional scene generated in the first embodiment. Note that the user inputs a set of the angle control parameter 1420 and the effective area specification 1430 for a plurality of areas of the two-dimensional image of the free viewpoint of the three-dimensional scene. For example, the GUI input unit 140 stores these sets in the memory, and outputs a set of the angle control parameter 1421 and the effective area specification 1431 stored in the memory from the GUI input unit 140, and for each of the plurality of areas specified by the effective area specification 1431, a corresponding separate angle control parameter 1421 may be applied.

[0076] As described above, according to the second embodiment, after photographing the teacher image under the fixed illumination environment, when generating the two-dimensional image of the free viewpoint of the three-dimensional scene, it is possible to contribute to enabling control of the illumination environment and gloss only in a specific designated area on the two-dimensional image, and an image generation device, an image generation method, and a program can be provided.

[0077] [Third Embodiment] Next, the third embodiment will be described in detail with reference to the drawings. FIG. 5 is a diagram showing an example of the configuration of an input screen of an angle control parameter input unit of a GUI input unit of an image generation apparatus according to the present disclosure.

[0078] Referring to FIG. 5, on an input screen 500 of an angle control parameter input unit 142 of a GUI input unit 140, a dθ axis 501 of dθ of an angle control parameter 1420, a dφ axis 502 of dφ of the angle control parameter 1420, and sliders 503 and 504 are shown. The dθ and dφ of the angle control parameter 1420 move the sliders 503 and 504 on the dθ axis 501 and the dφ axis 502, respectively, with a pointing device or the like, and select values dθ1 and dφ1. Thereby, the angle control parameter 1420 can be input via the angle control parameter input unit 142.

[0079] [Fourth Embodiment] Next, the fourth embodiment will be described in detail with reference to the drawings. FIG. 6 is a diagram showing an example of the configuration of an input screen of an angle control parameter input unit of a GUI input unit of an image generation apparatus according to the present disclosure.

[0080] Referring to FIG. 6, a dθdφ plane 603 of an angle control parameter 1420 is displayed on an input screen 600 of an angle control parameter input unit 142 of a GUI input unit 140. Note that, as an example, a color map indicating each position on the dθdφ plane 603 by color may be displayed.

[0081] On the dθdφ plane 603, a dθ axis 601 of dθ of the angle control parameter 1420 and a dφ axis 602 of dφ of the angle control parameter 1420 are shown. The dθ and dφ of the angle control parameter 1420 are selected on the dθdφ plane 603 by an operation such as pointing and clicking on a point 605 with a pointing device or the like, thereby selecting values dθ1 and dφ1. Thereby, the selected angle control parameter 1420 can be input via the angle control parameter input unit 142.

[0082] [Fifth Embodiment] Next, the fifth embodiment will be described in detail with reference to the drawings. FIG. 7 is a block diagram showing an example of the configuration of an input screen of an angle control parameter input unit of a GUI input unit of an image generation apparatus according to the present disclosure.

[0083] Referring to FIG. 7, a sphere 703 showing the dθdφ region of the angle control parameter 1420 is displayed on the input screen 700 of the angle control parameter input unit 142 of the GUI input unit 140. As an example, a color map showing each position on the surface of the sphere 703 indicating the dθdφ region may be displayed in color.

[0084] On the surface of the sphere 703 showing the dθdφ region, a dθ axis 701 of dθ of the angle control parameter 1420 and a dφ axis 702 of dφ of the angle control parameter 1420 are shown. The dθ axis 701 and the dφ axis 702 are assumed to be fixed even when the sphere 703 rotates. The sphere 703 can perform an operation of rotating 704 by a finger or the like on the displayed screen, and by making a point 705 coincide with the origin 707 of the dθ axis 701 and the dφ axis 702, for example, the values dθ1 and dφ1 indicated by the point 705 can be selected. Thereby, the angle control parameter 1420 can be input via the angle control parameter input unit 142.

[0085] As described above, according to the third to fifth embodiments, the angle control parameter 1420 can be input by the GUI input unit 140.

[0086] [Sixth Embodiment] Next, the sixth embodiment will be described in detail with reference to the drawings. FIG. 8 is a diagram schematically showing an example of information obtained by the optical learning inference unit 112 of the optical information learning inference device 110 shown in FIG. 3 according to the present disclosure. Referring to FIG. 8, it is assumed that the three-dimensional scene 1000 is irradiated with light from the light source 800, and the incident light 801 is incident on the reflection position 1001 of the three-dimensional scene 1000. Note that although the reflection position 1001 is indicated by a black circle, it only schematically shows the position and does not intend the existence of an object in the form of a black circle.

[0087] The incident light 801 is reflected at the reflection position 1001. The reflected light is divided into an isotropic component schematically shown by arrows within the range surrounded by the broken line 850 and a total reflection component schematically shown by arrows within the range surrounded by the broken line 851. The azimuth at which the magnitude of the reflected light among the total reflection components peaks is defined as the reflection peak azimuth 811.

[0088] Among the information obtained by the optical learning inference unit 112, the reflected light component in the reflection peak azimuth 811 is important for obtaining the gloss of the two-dimensional image rendered by the rendering device 200. Therefore, when the optical learning inference unit 112 performs learning and generates a second learned model, it is important that the teacher image used contains a large amount of the reflected light component in the reflection peak azimuth 811.

[0089] FIG. 9 is a diagram schematically showing an example of a method for photographing a teacher image according to the present disclosure. In FIG. 9, components denoted by the same reference numerals as in FIG. 8 indicate the same components. When photographing a two-dimensional image including the reflection position 1001 as a teacher image, from the azimuth that matches the azimuth of the reflection peak 811 of the total reflection component irradiated from the light source 800 and reflected at the reflection position 1001, the two-dimensional camera 300 shown in FIG. 3 is used to photograph a two-dimensional image including the reflection position 1001. Also, from the azimuth that matches the azimuth of the reflection peak 812 of the total reflection component irradiated from the light source 800 and reflected at the reflection position 1002, the two-dimensional camera 300 shown in FIG. 3 is used to photograph a two-dimensional image including the reflection position 1002.

[0090] In this way, by photographing a two-dimensional image with a two-dimensional camera 300 as a teacher image and using it for the learning of the optical learning inference unit 112, it is possible to generate a two-dimensional image of a free viewpoint three-dimensional scene with higher reproducibility.

[0091] As described above, each embodiment of the present invention has been explained. However, the present invention is not limited to the above-described embodiments, and further modifications, substitutions, and adjustments can be made without departing from the basic technical idea of the present invention. For example, the network configuration shown in each drawing, the configuration of each element, and the expression form of the message are examples for assisting the understanding of the present invention, and are not limited to the configurations shown in these drawings. Also, "A and / or B" is used to mean at least one of A or B.

[0092] Also, the procedures shown in the first to sixth embodiments described above can be realized by a program that causes a computer (9000 in FIG. 10) that functions as an image generation device according to the present invention to realize the functions of the image generation device. Such a computer is exemplified by a configuration including a CPU (Central Processing Unit) 9010, a communication interface 9020, a memory 9030, and an auxiliary storage device 9040 in FIG. 10. That is, the CPU 9010 in FIG. 10 may execute a control program of the image generation device and perform update processing of each calculation parameter held in the auxiliary storage device 9040 or the like.

[0093] The memory 9030 is a RAM (Random Access Memory), a ROM (Read Only Memory), or the like.

[0094] That is, each part (processing means, function) of the image generation device shown in the first to sixth embodiments described above can be realized by a computer program that causes a processor of the above computer to execute each of the above-described processes using its hardware.

[0095] Finally, the preferred forms of the present invention are summarized. [First Form] The image generation device may generate a two-dimensional image of a free viewpoint of a three-dimensional scene that has learned the spatial structure and optical information. The image generation device may include the learned optical learning inference unit and a reflection azimuth calculator. The optical learning inference unit of the image generation device receives azimuth information, and may extract optical information using at least the azimuth information. The reflection azimuth calculator of the image generation device receives an angle control parameter as an input, performs an operation of the observation azimuth and the angle control parameter on at least a partial region in the reconstructed space based on the learned spatial structure to generate the azimuth information, and may output the azimuth information to the optical learning inference unit. [Second form] In the image generation device according to the first form, it is preferable that the angle control parameter specifies the amount of deviation of the angle from the reference angle. [Third form] In the image generation device according to the first form, it is preferable that the angle control parameter is input from the outside. [Fourth form] In the image generation device according to the first form, it is preferable that the angle control parameter is input from the outside via a graphic user interface. [Fifth form] In the image generation device according to the first form, it is preferable that the effective region designation for designating the partial region is input from the outside. [Sixth form] The image generation device according to the fifth form preferably further includes a region determination unit that determines the partial region based on the effective region designation. [Seventh form] In the image generation device according to the first form, it is preferable that the partial region is automatically determined. [Eighth form] In the image generation device according to the first form, it is preferable that the operation is an addition of the observation azimuth and the angle control parameter. [Form 9] In an image generation method, a computer of an image generation device that generates a two-dimensional image of a free viewpoint of a three-dimensional scene, which has learned a spatial structure and optical information, receives orientation information, and may extract optical information using at least the orientation information. In the image generation method, the computer receives an angle control parameter as an input, and may generate the orientation information by performing an operation between the observation orientation with respect to at least a partial region in a reconstructed space based on the learned spatial structure and the angle control parameter. [Form 10] In a program, a computer of an image generation device that generates a two-dimensional image of a free viewpoint of a three-dimensional scene, which has learned a spatial structure and optical information, is caused to perform a process of receiving orientation information, and may be caused to execute a process of extracting optical information using at least the orientation information. The program causes the computer to perform a process of receiving an angle control parameter as an input, and a process of generating the orientation information by performing an operation between the observation orientation with respect to at least a partial region in a reconstructed space based on the learned spatial structure and the angle control parameter, and may be caused to execute the process. Note that, similar to Form 1, Forms 9 and 10 can be expanded to Forms 2 to 8. [Form 11] An image generation device that generates a two-dimensional image of a free viewpoint of a three-dimensional scene, which has learned a spatial structure and optical information, includes the learned optical learning inference unit and a reflection orientation calculator, wherein the optical learning inference unit receives orientation information, extracts optical information using at least the orientation information, and the reflection orientation calculator receives an angle control parameter as an input, Based on the learned spatial structure, calculate the observation orientation and the angle control parameter for at least a partial region in the reconstructed space to generate the orientation information. Output the orientation information to the optical learning inference unit. The angle control parameter is input by the GUI input unit to an image generation device. [12th form] The GUI input unit displays a slider for specifying an angle control parameter, and the angle control parameter indicated by the slider is input to the image generation device of the 11th form. [13th form] The GUI input unit displays an angle control parameter plane indicating an angle control parameter, and by specifying a position on the angle control parameter plane with a pointing device, the angle control parameter corresponding to the position is input to the image generation device of the 11th form. [14th form] The GUI input unit displays a sphere including a surface indicating an angle control parameter, and by rotating the sphere, the angle control parameter indicated by a point on the surface of the sphere that coincides with a predetermined origin on the surface of the sphere is input to the image generation device of the 11th form. [15th form] An image generation method in which a computer of an image generation device executes the processing of the 11th form to the 13th form. [16th form] A program that causes a computer of an image generation device to execute the processing of the 15th form. [17th form] When learning the spatial structure and optical information, an image including a two-dimensional image taken according to the reflection peak orientation of each point of the three-dimensional scene from a light source is used as a teacher image for learning in the image generation device of the 1st form. [18th form] An image generation method in which a computer of an image generation device executes the processing of the 17th form. [19th form] A program that causes a computer of an image generation device to execute the processing of the 18th form.

[0096] Note that the disclosures of the above patent documents are incorporated herein by reference. Within the scope of the entire disclosure of the present invention (including the claims), modifications and adjustments of the embodiments or examples can be made based on the basic technical idea. Also, within the scope of the disclosure of the present invention, various combinations or selections of various disclosure elements (including each element of each claim, each element of each embodiment or example, each element of each drawing, etc.) are possible. That is, the present invention naturally includes all the disclosures including the claims, as well as various modifications and corrections that those skilled in the art could make according to the technical idea. In particular, for the numerical ranges described in this document, any numerical value or small range included within the range should be construed as specifically described even without separate description. Further, each disclosure item of the above-cited documents, as needed and in accordance with the spirit of the present invention, is regarded as included in the disclosure of the present application as a part of the disclosure of the present invention, and can be used in combination with the description items of this document, either in part or in whole.

Explanation of Signs

[0097] 10, 10A, 10B Image generation system 100, 100A, 100B Image generation device 101 Image generation device 102 Optical learning inference unit 103 Reflection azimuth calculator 110 Optical information learning inference device 111 Density learning inference unit 112 Optical learning inference unit 120 Observation information acquisition unit 130 Sampler 131 Position 132 Observation azimuth 135 Spatial sampling unit 140 GUI input unit 141 Free viewpoint input unit 142 Angle control parameter input unit 143 Effective area designation input unit 150 Error detection unit 160 Reflection azimuth calculator 170 Region determination unit 180 Execution feasibility determination unit 190, 191 Switches 200 Rendering device 300 2D camera (sensor) 400 2D image display device 500, 600, 700 Input screens 800 Light source 1000 3D scene 9000 Computer 9010 CPU 9020 Communication interface 9030 Memory 9040 Auxiliary storage device

Claims

1. An image generation device that generates a two-dimensional image of a free viewpoint of a three-dimensional scene, which has learned a spatial structure and optical information, comprising: the learned optical learning inference unit and a reflection azimuth calculator; the optical learning inference unit: receives azimuth information; extracts optical information using at least the azimuth information; the reflection azimuth calculator: receives an angle control parameter as an input; performs an operation of an observation azimuth and the angle control parameter on at least a partial region in a reconstructed space based on the learned spatial structure to generate the azimuth information; and outputs the azimuth information to the optical learning inference unit. An image generation device.

2. The image generation device according to claim 1, wherein the angle control parameter specifies an amount of deviation of an angle from a reference angle.

3. The image generation device according to claim 1, wherein the angle control parameter is input from the outside.

4. The image generation device according to claim 1, wherein the angle control parameter is input from the outside via a graphic user interface.

5. The image generation device according to claim 1, wherein a valid region designation for designating the partial region is input from the outside.

6. The image generation device according to claim 5, further comprising a region determination unit that determines the partial region based on the valid region designation.

7. The image generation device according to claim 1, wherein the partial region is automatically determined.

8. The image generation device according to claim 1, wherein the operation is an addition of the observation azimuth and the angle control parameter.

9. A computer of an image generation device that generates a two-dimensional image of a free viewpoint of a three-dimensional scene, which has learned a spatial structure and optical information, receives azimuth information; extracts optical information using at least the azimuth information; receives an angle control parameter as an input; and performs an operation of an observation azimuth and the angle control parameter on at least a partial region in a reconstructed space based on the learned spatial structure to generate the azimuth information. An image generation method comprising this.

10. In a computer of an image generation device that generates a two-dimensional image of a free viewpoint of a three-dimensional scene, which has learned a spatial structure and optical information, a process of receiving azimuth information; a process of extracting optical information using at least the azimuth information; a process of receiving an angle control parameter as an input; A process of calculating an observation orientation and the angle control parameter for at least a partial region in the reconstructed space based on the learned spatial structure, and generating the orientation information, A program for execution.

Citation Information

Patent Citations

  • Image processing device, image processing method, and program

    JP2023066705A