Learning method, learning device, and learning program

By calculating positions within images with reduced peripheral light intensity and training the learning model to compensate for light falloff, the method generates high-quality free-viewpoint images, overcoming the limitations of conventional models.

WO2026034220A1PCT designated stage Publication Date: 2026-02-12SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/026291
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2025-07-24
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Conventional learning models for generating free-viewpoint images do not adequately address the reduction in peripheral light intensity of images, leading to reduced quality in generated images.

Method used

A learning method that calculates positions within images with reduced peripheral light intensity and inputs these positions into a learning model to train it, compensating for light falloff, thereby generating high-quality free-viewpoint images.

Benefits of technology

The method enables the generation of high-quality free-viewpoint images by effectively addressing the reduction in peripheral light intensity, improving image quality and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025026291_12022026_PF_FP_ABST
    Figure JP2025026291_12022026_PF_FP_ABST
Patent Text Reader

Abstract

A learning method according to the present disclosure involves a computer acquiring a plurality of images including a prescribed object, calculating a position within each image on the basis of the plurality of images, and inputting the calculated position within each image to a learning model that generates a free-viewpoint image of a prescribed object, thereby training the learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Learning method, learning device, and learning program

[0001] The present disclosure relates to a learning method, a learning device, and a learning program.

[0002] Conventionally, a technique is known in which a correction amplification amount is acquired to correct the decrease in peripheral illumination, which occurs when the amount of light decreases from the center of a photographed image to the periphery of the photographed image, and peripheral illumination correction processing is performed on the photographed image based on the correction amplification amount.

[0003] Japanese Patent Application Laid-Open No. 2021-052282

[0004] However, the above-mentioned conventional techniques only correct the reduction in peripheral light intensity in a photographed image, and therefore the range of application is limited.

[0005] For example, in general, the learning process of a learning model that generates a free-viewpoint image from an image may not take into account the reduction in peripheral light intensity of the image being learned. For example, if the image being learned by the learning model is an image with reduced peripheral light intensity, the quality of the generated free-viewpoint image may be reduced.

[0006] Therefore, the present disclosure proposes a learning method, a learning device, and a learning program that are capable of generating high-quality free viewpoint images.

[0007] In order to solve the above problem, one form of a learning method according to the present disclosure includes a computer acquiring a plurality of images including a predetermined object, calculating a position within each image based on the plurality of images, and inputting the calculated positions within each image into a learning model that generates a free viewpoint image of the predetermined object, thereby training the learning model.

[0008] 1 is a diagram illustrating an example of an image with reduced peripheral illumination. FIG. 1 is a schematic diagram illustrating an example of light from a central light beam being received by an image sensor. FIG. 2 is a schematic diagram illustrating an example of light from an oblique light beam being received by an image sensor. FIG. 3 is a schematic diagram illustrating an example of an optical system in the case of vignetting by a lens barrel. FIG. 4 is a diagram illustrating a learning model based on Neural Radiance Fields (NeRF). FIG. 5 is a conceptual diagram illustrating the relationship between an image and a free viewpoint image. FIG. 6 is a diagram illustrating an overview of a learning system according to an embodiment. FIG. 7 is a diagram illustrating an example configuration of a learning device according to an embodiment. FIG. 8 is a schematic diagram illustrating pixel positions. FIG. 9 is a schematic diagram illustrating positions based on regions divided into a grid shape. FIG. 10 is a schematic diagram illustrating positions based on distances from the center of an image. FIG. 11 is a schematic diagram illustrating positions based on regions divided by a plurality of circles. FIG. 12 is a diagram illustrating a learning model according to the present disclosure. FIG. 13 is a diagram illustrating a learning model with a changed number of multilayer perceptrons. FIG. 14 is a diagram illustrating a learning model using an embedding vector. FIG. 15 is a diagram illustrating another example of a learning model using an embedding vector. FIG. 16 is an example of a flowchart illustrating the flow of a generation process according to an embodiment. FIG. 17 is an example of a flowchart illustrating the flow of a learning process according to an embodiment. FIG. 18 is a diagram illustrating a learning model based on Instant-NGP (Instant Neural Graphics Primitives) according to a modified example. FIG. 1 is a diagram showing a learning model in which one multilayer perceptron is added according to a modified example. FIG. 2 is a diagram showing a learning model in which two multilayer perceptrons are added according to a modified example. FIG. 3 is a diagram showing an example (1) of a learning model based on 3D Gaussian Splatting (3DGS) in accordance with a modified example. FIG. 4 is a diagram showing an example (2) of a learning model based on 3DGS in accordance with a modified example. FIG. 5 is a diagram showing an example (3) of a learning model based on 3DGS in accordance with a modified example. FIG. 6 is a hardware configuration diagram showing an example of a computer that realizes the functions of a learning device.

[0009] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0010] The present disclosure will be described in the following order of items. 1. Embodiment 1-1. Overview of embodiment 1-1-1. Reduction in peripheral light intensity 1-1-2. Impact of learning an image with reduced peripheral light intensity 1-2. Overview of learning system according to embodiment 1-3. Configuration of learning device according to embodiment 1-3-1. Example of calculation of position within image (1) 1-3-2. Example of calculation of position within image (2) 1-3-3. Example of calculation of position within image (3) 1-3-4. Example of calculation of position within image (4) 1-3-5. Learning example (1) 1-3-6. Learning example (2) 1-3-7. Learning example (3) 1-3-8. Learning example (4) 1-4. Flowchart showing the procedure of generation processing according to embodiment 1-5. Flowchart showing the procedure of learning processing according to embodiment 1-6. Modified example according to embodiment 1-6-1. Application to other learning models (1) 1-6-2. Application to other learning models (2) 1-6-3. Other learning methods 2. Other embodiments 3. Effects of the learning method according to the present disclosure 4. Hardware configuration

[0011] (1. Embodiments) (1-1. Overview of the Embodiments) Techniques related to learning models (e.g., NeRF) that generate free-viewpoint images of a predetermined object from images that include the object are known. For example, learning models such as NeRF can reconstruct a three-dimensional model of a predetermined object from two-dimensional images that include the object. In this way, such learning models can render and generate photorealistic free-viewpoint images.

[0012] In such a learning model, the quality of the base image may affect the quality of the generated free-viewpoint image. For example, in an image captured by a photographing device such as a camera, the peripheral light intensity of the image may be lower than the light intensity at the center of the image (hereinafter, sometimes referred to as the central light intensity). If an image with such low peripheral light intensity is used for training the learning model, the quality of the generated free-viewpoint image may be reduced.

[0013] Below, we will first explain the reduction in the amount of peripheral light in an image, and then explain the influence of learning an image with reduced peripheral light.

[0014] (1-1-1. Reduction of peripheral illumination) First, reduction of peripheral illumination in an image will be described. Reduction of peripheral illumination in a captured image will be described using FIG. 1. FIG. 1 is a diagram showing an example of an image IM11 with reduced peripheral illumination.

[0015] Image IM11 is a captured image of a landscape including a Ferris wheel. In image IM11, the first end CE11 and the second end CE11, which are the edges of image IM11, appear darker than the center of image IM11. As such, the amount of light may decrease around the periphery of the image, such as around the first end CE11 and the second end CE11. This decrease in peripheral light is called vignetting.

[0016] The peripheral light amount is generally expressed by the following mathematical formula: When vignetting occurs, the peripheral light amount is known to be reduced in amount relative to the central light amount of the image.

[0017]

[0018] Causes of vignetting include, for example, the cosine fourth power law, vignetting, aperture efficiency, etc. As an example of the causes of vignetting, the cosine fourth power law will be described. In general, an image capturing device generates an image by collecting light coming from an object or a light source using a lens and converting the light into an electrical signal using an image sensor. Here, an example of the optical system of an image capturing device will be described with reference to FIGS. 2 and 3 .

[0019] FIG. 2 shows an example in which light from a central light beam is received by an image sensor. FIG. 2 is a schematic diagram showing an example in which light from a central light beam is received by an image sensor. FIG. 2 shows an optical axis OA21, a central light beam LF21, and an oblique light beam LF22. The area of ​​the light from the central light beam LF21 is area LA21. The area of ​​the light from the oblique light beam LF22 is area LA22. In FIG. 2, light from the central light beam LF21 is collected by a lens L21. The collected light is then received by a sensor at the center of the image sensor IS21.

[0020] Fig. 3 shows an example in which light from an oblique light beam is received by an image sensor. Fig. 3 is a schematic diagram showing an example in which light from an oblique light beam is received by an image sensor. In Fig. 3, light from an oblique light beam LF22 is incident on lens L21 at an angle AN31 of θ and is collected by lens L21. The collected light is then received by a sensor at the end of image sensor IS21.

[0021] Here, the area LA21 is larger than the area LA22. Therefore, the light from the oblique light beam LF22 has a small luminous flux density on the surface of the image sensor IS21. In this way, the amount of light incident at an angle AN31 of θ is cos 4 It is known that the angle θ is proportional to the cosine fourth power law. This phenomenon is called the cosine fourth power law. For example, because the cosine fourth power law depends on the angle θ, it can have a significant effect on images captured with a short focal length or using a wide-angle lens.

[0022] Next, vignetting will be described as an example of a cause of vignetting. Vignetting caused by a lens barrel will be described using FIG. 4. FIG. 4 is a schematic diagram showing an example of an optical system in the case of vignetting caused by a lens barrel. FIG. 4 shows an example in which light from a central light beam is received by an image sensor, and an example in which light from an oblique light beam is received by an image sensor. FIG. 4 shows a central light beam LF41, an oblique light beam LF42, lenses L41 and L42, a diaphragm D41, and an image sensor IS41.

[0023] In Fig. 4, the light area of ​​the central light beam LF41 is area LA41. The light area of ​​the oblique light beam LF42 is area LA42. In Fig. 4, light from the central light beam LF41 is collected by lenses L41 and L42. The collected light is then received by the central sensor of the image sensor IS41.

[0024] Meanwhile, light from the oblique light beam LF42 is condensed at the lower end of lens L41 and the upper end of lens L42. As a result, the area of ​​the oblique light beam LF42 passing through lenses L41 and L42 is reduced at the lower end of lens L41 and the upper end of lens L42. The condensed light is then received by a sensor at the end of image sensor IS41. In this manner, the amount of light from the oblique light beam LF42 is reduced. While FIG. 4 generally illustrates an example in which the amount of light is reduced at the end of a lens, other factors, such as the aperture D41 or accessories like a lens hood, can also reduce the amount of light. Furthermore, dust or other particles adhering to the lens can also reduce the amount of light. As described above, when vignetting occurs, the amount of peripheral light may be reduced relative to the amount of light at the center.

[0025] (1-1-2. Effects of Learning Images with Reduced Peripheral Illumination) Next, the effects of learning images with reduced peripheral illumination will be described. First, the learning model will be described. In the following, NeRF will be used as an example of the learning model.

[0026] NeRF is a technology that uses multiple images of an object taken from various angles as training data to learn the spatial representation of a scene containing the object and generate images from new viewpoints. By applying NeRF, users can obtain free-viewpoint images of the object.

[0027] Furthermore, image generation (image rendering) in NeRF is performed by estimating what color a certain coordinate on a certain ray has in the NeRF space. Specifically, in NeRF, when an image capturing device virtually captures an image of the NeRF space from a certain direction, the object is reconstructed in the NeRF space based on the probability of what color exists at what position (coordinate) on the ray.

[0028] Here, the learning process in NeRF will be described with reference to Fig. 5. Fig. 5 is a diagram showing a learning model M1 based on NeRF. The learning model M1 shown in Fig. 5 includes at least three multilayer perceptrons MLP1 to MLP3.

[0029] In the example of FIG. 5, the position coordinate x of the predetermined object 3d is subjected to positional encoding before being input to the multilayer perceptron MLP1 (corresponding to the PE shown in FIG. 5). For example, the position coordinate x of a given object 3d Coordinate data (x, y, z) indicating a position on a light ray is used as the coordinate data. The coordinate data here includes spatial coordinate values ​​that specify a position in the NeRF space.

[0030] Next, the position-encoded position coordinates are input to the multilayer perceptron MLP1. In this case, the position-encoded position coordinates are input to the previous layer of the multilayer perceptron MLP1, and the position-encoded position coordinates are further input to the subsequent layer of the multilayer perceptron MLP1.

[0031] Then, the value output from the multilayer perceptron MLP1 is input to the multilayer perceptron MLP2, which outputs the density σ. Next, the value output from the multilayer perceptron MLP1 and the angle d based on the direction of the ray are input to the multilayer perceptron MLP3. 3d For example, the angle d based on the direction of the light ray is input. 3d is expressed by the angles (θ, φ) that specify the ray. Then, the multi-layer perceptron MLP3 outputs a color C. For example, the color C is an RGB value.

[0032] In this way, in the learning model M1, the position coordinate x of a predetermined object 3d and angle d based on the direction of the light ray 3d By inputting the above, the learning device learns to output density σ and color C. Here, an overview of the learning process of the learning model M1 will be described. For example, assume that a learning device has the learning model M1. In this case, the learning device performs rendering based on the colors of multiple points on the light ray. Next, the learning device compares the inferred value, which is the obtained rendered image, with ground truth data, which is an image actually captured of the target. The learning device then determines the difference between the inferred value and the ground truth data as a loss, and updates the coefficients of the learning model to reduce this loss. In this way, the learning device learns the learning model.

[0033] Next, a case where the quality of a base image affects the quality of a generated free-viewpoint image in a learning model will be described with reference to Fig. 6. Fig. 6 is a conceptual diagram showing the relationship between an image and a free-viewpoint image.

[0034] In FIG. 6 , an example will be described in which the predetermined object is a bulldozer. Here, it is assumed that images IM61 and IM62 are used for learning. For example, images IM61 and IM62 are images in which peripheral light intensity is reduced. In the example of FIG. 6 , the region in image IM61 in which peripheral light intensity is reduced is indicated as region CE61. For example, in image IM61, region CE61 occurs at the four corners of image IM61. Similarly, in image IM62, region CE61 occurs at the four corners of image IM62.

[0035] Furthermore, images IM61 and IM62 include the rotating light of the bulldozer MO61. A light ray BE61 in image IM61 corresponds to the rotating light of the bulldozer MO61. A light ray BE62 in image IM62 corresponds to the rotating light of the bulldozer MO61.

[0036] In image IM61, the rotating light is located near the center of the image. Furthermore, the rotating light in image IM61 is located away from area CE61. On the other hand, in image IM62, the rotating light is located near the edge of the image. Furthermore, the rotating light in image IM62 is located in a position that overlaps area CE61. Therefore, the amount of light of the rotating light in image IM61 is greater than the amount of light of the rotating light in image IM62. For example, if the rotating light is red, the rotating light in image IM61 will appear bright red. On the other hand, if the rotating light is red, the rotating light in image IM62 will appear dark red. In this way, colors with different amounts of light are learned during learning.

[0037] In NeRF training, free-viewpoint images are updated to minimize loss. However, if there is a reduction in peripheral illumination, the loss increases, which may result in an incorrect free-viewpoint image being updated. Furthermore, updating with an incorrect loss may not only result in incorrectly learned colors, but may also affect the accuracy of the shape. For example, such training may result in degradation of the image quality of the free-viewpoint image, such as incorrect colors, incorrect shapes, or floating artifacts, which resemble smoke floating in the air. Thus, a method for training while compensating for the reduction in peripheral illumination of an image is desired.

[0038] In response to the above problem, the present disclosure proposes a learning method for training a learning model by inputting positions within each image into the learning model that generates a free-viewpoint image of a predetermined target. As a result, the present disclosure trains the learning model based on images with reduced peripheral light falloff, thereby enabling the generation of high-quality free-viewpoint images.

[0039] (1-2. Overview of the Learning System According to the Embodiment) Next, an overview of the learning system 1 according to the embodiment will be described with reference to Fig. 7. Fig. 7 is a diagram showing an overview of the learning system 1 according to the embodiment.

[0040] 7, learning system 1 includes a photographing device 10 and a learning device 100. The photographing device 10 and the learning device 100 are connected, for example, by wire or wirelessly via a network N. Note that learning system 1 shown in FIG. 7 may include multiple photographing devices 10.

[0041] However, when generating a single learning model, it is desirable to use multiple images with the same characteristics, such as peripheral light falloff, for learning. Therefore, it is preferable to use images captured using the same image capture device 10. Furthermore, it is preferable that each of the multiple images be captured using an optical system (optical characteristics) with the same lens aperture, zoom, etc. Furthermore, when using multiple images captured using multiple different image capture devices 10, it is preferable that each image capture device has the same optical system (optical characteristics), such as the same lens, image sensor, and housing.

[0042] Alternatively, if there are multiple imaging devices 10, each with different optical characteristics and different characteristics such as the reduction in peripheral light intensity of the captured images, a different learning model can be generated for each set of images captured by each imaging device 10.

[0043] The photographing device 10 is an information processing device having a photographing function. For example, the photographing device 10 is a camera. Note that the photographing device 10 may also be a terminal device used by a user as long as it has a photographing function. The terminal device referred to here is, for example, a mobile phone, a tablet terminal, or the like.

[0044] The learning device 100 is, for example, an information processing device such as a server device, and executes the learning process according to the embodiment. For example, the learning device 100 acquires multiple images including a predetermined object. Next, the learning device 100 calculates a position within each image based on the multiple images. The learning device 100 then inputs the calculated positions within each image into a learning model that generates a free viewpoint image of the predetermined object, thereby training the learning model.

[0045] (1-3. Configuration of the Learning Device According to the Embodiment) Next, the configuration of the learning device 100 according to the embodiment will be described with reference to Fig. 8. Fig. 8 is a diagram showing an example of the configuration of the learning device 100 according to the embodiment.

[0046] 8, learning device 100 has communication unit 110, memory unit 120, control unit 130, input unit 140, and display unit 150. Learning device 100 may also have an input unit (e.g., a touch panel) that accepts various operations from a user operating learning device 100, and a display unit (e.g., a liquid crystal display) that displays various information.

[0047] The communication unit 110 is realized by, for example, a network interface card (NIC), etc. The communication unit 110 is connected to a network N (the Internet, near field communication (NFC), Bluetooth (registered trademark), etc.) via a wired or wireless connection, and transmits and receives information to and from the imaging device 10, etc. via the network N.

[0048] The storage unit 120 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. In particular, this storage device may include a computer-readable non-transitory recording medium such as a flash memory, a hard disk, or an optical disk. As shown in FIG. 9 , the storage unit 120 includes an image storage unit 121, a learning model storage unit 122, and a free viewpoint image storage unit 123. Each of these storage units may be configured as a single common storage device (or non-transitory recording medium), or may be configured as an independent individual storage device (or non-transitory recording medium).

[0049] The image storage unit 121 stores images captured by the image capturing device 10. Specifically, the image storage unit 121 stores a plurality of images including a first image captured by the image capturing device 10 of a predetermined object from a first direction and a second image captured by the image capturing device 10 of the predetermined object from a second direction.

[0050] The learning model storage unit 122 stores a learning model learned by the learning process according to the embodiment. For example, a learning model is generated for each object photographed by the image capture device 10. Furthermore, when there are images photographed by multiple image capture devices 10 with different optical characteristics, a learning model may be generated for each set of images photographed by each of the image capture devices 10.

[0051] The generated learning model can be stored in the non-transitory recording medium of the learning model storage unit 122 described above. Alternatively, the learning model can be used as a data asset by copying it to another portable non-transitory storage medium or to an external non-transitory recording medium via an online line. For example, if there is an image generation device other than the learning device 100 according to this embodiment that can generate free viewpoint video based on a similar learning model, it is possible to manufacture an image generation device that can generate similar 3D images using the learning model by copying the above-described learning model to the non-transitory recording medium of that image generation device.

[0052] The control unit 130 is realized, for example, by a central processing unit (CPU) or a micro processing unit (MPU) executing a program (e.g., a learning program) stored inside the learning device 100 using a random access memory (RAM) or the like as a work area. The control unit 130 is a controller and may be realized, for example, by an integrated circuit such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). The control unit 130 includes an acquisition unit 131, a calculation unit 132, a learning unit 133, and a generation unit 134. These units may be configured such that a single common CPU executes a program to perform the functions of each unit, or may be realized such that multiple CPUs share and execute the program. When multiple CPUs are used, each unit does not necessarily have to correspond one-to-one to each CPU; the multiple CPUs may cooperate to perform the functions of the multiple units.

[0053] The acquisition unit 131 acquires various types of information. The acquisition unit 131 acquires a plurality of images including a predetermined object from the image capturing device 10. For example, the acquisition unit 131 acquires a plurality of images from the image capturing device 10, including a first image in which the predetermined object is captured by the image capturing device 10 from a first direction, and a second image in which the predetermined object is captured by the image capturing device 10 from a second direction. The acquisition unit 131 then stores the acquired plurality of images in the image storage unit 121.

[0054] The calculation unit 132 calculates the position within each image based on the multiple images. Then, the calculation unit 132 stores the calculated position within each image in the storage unit 120. The calculation process executed by the calculation unit 132 will be described in detail in 1-3-1 to 1-3-4 below.

[0055] The learning unit 133 inputs the positions in each image calculated by the calculation unit 132 into a learning model that generates a free viewpoint image of a predetermined target, thereby learning the learning model. For example, the learning unit 133 inputs the positions in each image into a learning model based on NeRF, which is an example of a learning model, thereby learning the NeRF-based learning model. The learning unit 133 then stores the learned learning model in the learning model storage unit 122.

[0056] For example, the learning unit 133 inputs the position coordinates of a predetermined object, an angle based on the direction of a light ray calculated from an image, and a position within the image into a learning model based on NeRF, and learns the model to output density and a color compensated for the amount of light in the image.The learning unit 133 then stores this learning model in the learning model storage unit 122.The learning process performed by the learning unit 133 will be described in detail in 1-3-5 to 1-3-8 below.

[0057] The generation unit 134 generates a free viewpoint image based on each image in which changes in the amount of light indicated by each image have been compensated for, using the learning model stored in the learning model storage unit 122. The generation unit 134 then stores the generated free viewpoint image in the free viewpoint image storage unit 123.

[0058] In this embodiment, the generation unit 134, the learning model storage unit 122, and the free viewpoint image storage unit 123 are configured as parts of the learning device 100. However, these units may also configure an image generation device independent of the learning device 100.

[0059] That is, in this case, an image generation device including components similar in configuration to the generation unit 134, learning model storage unit 122, and free viewpoint image storage unit 123 of the present embodiment is provided independently of the learning device 100. The learned learning model generated by the learning unit 133 is then copied to the learning model storage unit 122 of the image generation device via a portable non-transitory storage medium, an online line, or the like. This makes it possible to manufacture an image generation device capable of generating free viewpoint images based on the learning model, separate from the learning device 100, and to arbitrarily generate similar free viewpoint images even in remote locations, for example.

[0060] Furthermore, the image generation device may be connected to a computer or media player equipped with a display device such as a normal display, a head-mounted display, a large screen, etc. This makes it possible to provide free viewpoint images generated using this technology to viewers as part of content such as a game, or as background images when producing other video content such as a movie.

[0061] The input unit 140 accepts input of various types of information. For example, the input unit 140 accepts input of various types of information via an input device such as a user interface (UI) or keyboard that can be operated by the user.

[0062] Display unit 150 displays information output by study device 100. For example, display unit 150 is a liquid crystal display or the like built into study device 100 or a liquid crystal display or the like connected to study device 100.

[0063] (1-3-1. Example of Calculating Position Within an Image (1)) Next, the calculation process of a position within an image according to this embodiment will be described with reference to FIG. 9. FIG. 9 is a schematic diagram showing pixel positions. FIG. 9 shows image IM91 as one image out of multiple images. Edge BO91 is the edge of image IM91. Furthermore, in the example of FIG. 9, the position of a pixel will be described as an example of a position within an image.

[0064] In this case, the calculation unit 132 calculates the position PO91 of the pixel in the image IM91 as the position within the image. In the example of Fig. 9, the position PO91 is represented by image coordinates (pixel coordinates) as (u, v). For example, in image coordinates, the upper left corner of the image IM91 is the origin (0, 0), u is the horizontal axis, and v is the vertical axis, and the position PO91 is represented by (u, v).

[0065] In this way, the calculation unit 132 can indicate the position within the image based on the image coordinates, and therefore can identify the location where the light amount has decreased with sufficient resolution. For example, if the shape of the lens hood is asymmetric, the decrease in light amount may not be symmetric. In such a case, the calculation unit 132 can identify the location where the light amount has decreased with sufficient resolution, even if the decrease in light amount is asymmetric.

[0066] Furthermore, for example, if dust or the like adheres to the lens, the amount of light may decrease due to the dust or the like. In such a case, the calculation unit 132 can identify with sufficient resolution the location where the amount of light has decreased due to the dust or the like.

[0067] Furthermore, for example, compensation for the amount of light may be partial or the symmetry of the image may be distorted depending on the signal processing within the image capturing device 10. Even in such cases, the calculation unit 132 can identify the area where the amount of light has decreased with sufficient resolution.

[0068] Furthermore, for example, in an actual product, the amount of light may change depending on the position in the image due to the characteristics of the imaging device 10. Even in such a case, the calculation unit 132 can identify the location where the amount of light has decreased with sufficient resolution.

[0069] (1-3-2. Example of Calculating Position in Image (2)) In the above embodiment, an example of calculating the position of a pixel in an image has been described, but the present invention is not limited to this. For example, the calculation unit 132 may calculate, as the position in each image, a position based on regions obtained by dividing each image into a grid of a predetermined size.

[0070] The calculation process of a position based on areas divided into a grid pattern will be described with reference to Fig. 10. Fig. 10 is a schematic diagram showing a position based on areas divided into a grid pattern. Fig. 10 shows an image IM101 as one of a plurality of images. An edge BO101 is an edge of the image IM101.

[0071] In this case, the calculation unit 132 calculates a position PO101 based on regions obtained by dividing the image IM101 into a grid shape having a predetermined size as a position within the image. In the example of Fig. 10, the image IM101 is shown divided into a grid shape with a square DA101 as the smallest unit.

[0072] In the example of FIG. 10, the position PO101 is, for example, the origin (0, 0) at the top left of the image IM101, b (u) is the horizontal axis, v b (v) is the vertical axis, and (u b , v b ) is shown by

[0073] In this way, the calculation unit 132 can reduce the memory size for storing positions within an image by calculating positions based on an area having a predetermined size larger than the pixel position, which allows the learning unit 133 to more easily execute the learning process.

[0074] In the above embodiment, an example was given of an area divided into a grid pattern with the square DA101 as the smallest unit, but for example, such an area may be divided into an area of ​​any shape.

[0075] (1-3-3. Example of Calculating Position in Image (3)) In the above embodiment, an example of calculating the position of a pixel in an image has been described, but the present invention is not limited to this. For example, the calculation unit 132 may calculate a position based on the distance from the center of each image as the position in each image.

[0076] The calculation process of a position based on the distance from the center of an image will be described using Fig. 11. Fig. 11 is a schematic diagram showing a position based on the distance from the center of an image IM111. Fig. 11 shows an image IM111 as one of a plurality of images. An edge BO111 is an edge of the image IM111.

[0077] In this case, the calculation unit 132 calculates the position PO111 based on the distance r from the center CN111 of the image IM111 as the position within the image. For example, if u is the horizontal axis and v is the vertical axis, and the coordinates of the center CN111 of the image IM111 are (C x , C y ), if the coordinates of the position PO111 are represented by r(u, v), the position PO111 is represented by the following mathematical formula: where (u, v) represents the position of the pixel.

[0078]

[0079] Generally, changes in the amount of peripheral light in an image depend on the distance from the center of the image. Therefore, the calculation unit 132 can reduce the degree of freedom of variables by indicating a position in the image as a distance from the center of the image. This allows the learning unit 133 to restrict the learning parameters of a multilayer perceptron or the like so that they do not fall into a local optimum.

[0080] Furthermore, the calculation unit 132 can reduce the dimensions of the input variables, which allows the learning unit 133 to more easily execute the learning process.

[0081] (1-3-4. Example of Calculating Position Within an Image (4)) In the above embodiment, an example of calculating the position of a pixel within an image has been described, but this is not limiting. For example, the calculation unit 132 may calculate, as the position within each image, a position based on an area formed by the circumference of a first circle and the circumference of a second circle, the area being divided by a plurality of circles having a predetermined radius with the center of each image as the center of the circle.

[0082] The process of calculating a position based on an area divided by a plurality of circles will be described with reference to Fig. 12. Fig. 12 is a schematic diagram showing a position based on an area divided by a plurality of circles. Fig. 12 shows an image IM121 as one of a plurality of images. An edge BO121 is an edge of the image IM121.

[0083] In this case, the calculation unit 132 calculates a position PO121 as a position within the image based on an area divided by multiple circles having a predetermined radius with the center of the image as the center of the circle and formed by the circumference of a first circle and the circumference of a second circle.

[0084] In the example of Fig. 12, a plurality of circles having a predetermined radius and centered at the center of the image IM121 are drawn on the image IM121. The area formed by the circumference CI121 of the first circle and the circumference CI122 of the second circle is the area AR121 (the shaded area in Fig. 12). In this case, the coordinates of the position PO121 based on the area AR121 are calculated by dividing the distance r from the center of the image IM121 by the distance r r Based on r It is denoted as (u, v).

[0085] Furthermore, the same amount of light is set for each area formed by the circumference of the first circle and the circumference of the second circle. For example, if the calculated position in the image is included in area AR121, the same amount of light is set.

[0086] In this way, the calculation unit 132 can reduce the memory size required for calculating positions within an image because the calculated positions are discretized, which allows the learning unit 133 to more easily execute the learning process.

[0087] (1-3-5. Learning Example (1)) Next, the learning process of the learning model according to the embodiment will be described with reference to FIG. 13. FIG. 13 is a diagram showing the learning model according to the present disclosure. The learning model shown in FIG. 13 includes at least three multilayer perceptrons MLP1 to MLP3. In the following, 2d is the position of a pixel in each image.

[0088] In the example of FIG. 13, the position coordinate x of the predetermined object3d are position-encoded before being input to the multilayer perceptron MLP1. Subsequently, the learning unit 133 inputs the position-encoded position coordinates to the multilayer perceptron MLP1. In this case, the learning unit 133 inputs the position-encoded position coordinates to the previous layer of the multilayer perceptron MLP1. Furthermore, the learning unit 133 further inputs the position-encoded position coordinates to the subsequent layer of the multilayer perceptron MLP1.

[0089] The learning unit 133 then inputs the value output from the multilayer perceptron MLP1 to the multilayer perceptron MLP2, thereby outputting the density σ. Next, the learning unit 133 calculates the density σ for the multilayer perceptron MLP3 by calculating the value output from the multilayer perceptron MLP1 and the angle d based on the direction of the light ray. 3d and the position d in the image 2d Here, the position d in the image is input. 2d is position-encoded before being input to the multilayer perceptron MLP3. Then, the learning unit 133 calculates the color C C Output.

[0090] In this way, when the learning unit 133 learns the learning model, it inputs the position in the image to the multilayer perceptron MLP3 that outputs the color. As a result, the learning unit 133 outputs the color C C The learning model is trained to output

[0091] (1-3-6. Learning Example (2)) In the above embodiment, an example of training a learning model including at least three multilayer perceptrons has been described, but this is not limiting. For example, the learning model may include at least four multilayer perceptrons.

[0092] The learning process of the learning model according to the embodiment will be described with reference to Fig. 14. Fig. 14 is a diagram showing a learning model in which the number of multilayer perceptrons is changed. The learning model shown in Fig. 14 includes at least four multilayer perceptrons MLP1 to MLP4.

[0093] In the learning model shown in Fig. 14, a multi-layer perceptron MLP4 is added compared to the learning model shown in Fig. 13. 2d is input to the multilayer perceptron MLP4, and the value output from the multilayer perceptron MLP3 is multiplied by the value output from the multilayer perceptron MLP4 to output a color in which the amount of peripheral illumination has been compensated.

[0094] The learning model will be explained in more detail with reference to Fig. 14. In the example of Fig. 14, the position coordinate x of a predetermined object is 3d are position-encoded before being input to the multilayer perceptron MLP1. Subsequently, the learning unit 133 inputs the position-encoded position coordinates to the multilayer perceptron MLP1. In this case, the learning unit 133 inputs the position-encoded position coordinates to the previous layer of the multilayer perceptron MLP1. Furthermore, the learning unit 133 further inputs the position-encoded position coordinates to the subsequent layer of the multilayer perceptron MLP1.

[0095] The learning unit 133 then inputs the value output from the multilayer perceptron MLP1 to the multilayer perceptron MLP2, thereby outputting the density σ. Next, the learning unit 133 calculates the density σ for the multilayer perceptron MLP3 by calculating the value output from the multilayer perceptron MLP1 and the angle d based on the direction of the light ray. 3d In this case, the learning unit 133 outputs the color C.

[0096] Then, the learning unit 133 calculates the position d 2d is input to the multilayer perceptron MLP4. Here, the position d 2d is position-encoded before being input to the multilayer perceptron MLP4. Next, the learning unit 133 multiplies the value output from the multilayer perceptron MLP3 by the value output from the multilayer perceptron MLP4 to obtain the color C C In this case, the value output from the multilayer perceptron MLP4 becomes a compensation coefficient for compensating for the amount of light. In this way, the learning unit 133 outputs a color C with the amount of peripheral light compensated. C The learning model is trained to output

[0097] In this way, in the learning model shown in FIG. 14, the learning unit 133 calculates the compensation coefficient c l The compensation coefficient c can be calculated. l For example, the color may be expressed in common by RGB values. C = c l C = (c l R, c l G, c l B) and the RGB values ​​and the compensation coefficient c l and the color C with the peripheral illumination compensated. C In this way, the learning unit 133 can reduce the number of coefficients to one, which makes it easier to learn the learning model.

[0098] In addition, the compensation coefficient c l may be different compensation coefficients for each RGB value. In this way, even if the degree of contribution of the reduction in peripheral illumination, etc., to the R, G, and B values ​​differs, the learning unit 133 can obtain a color C with the peripheral illumination compensated. C can be calculated with high accuracy.

[0099] (1-3-7. Learning Example (3)) In the above embodiment, the position d 2d Although the above description is given by way of example in which the image is position-encoded before being input to the multi-layer perceptron, this is not limiting. For example, the position within the image may be calculated as an embedding vector.

[0100] The learning process of the learning model according to the embodiment will be described with reference to Fig. 15. Fig. 15 is a diagram showing a learning model using embedding vectors. The learning model shown in Fig. 15 includes at least four multilayer perceptrons MLP1 to MLP4.

[0101] 15, a multi-layer perceptron MLP4 is added to the learning model shown in FIG. 13. In addition, the learning unit 133 calculates the position d 2d is input to the multilayer perceptron MLP4. The learning unit 133 also inputs the position d 2dThe vector value selected based on the above is input to the multilayer perceptron MLP4. The input vector value is a vector value selected from the latent code LC1. The latent code LC1 is, for example, a table in which positions in an image are associated with vector values. Note that the latent code LC1 is assumed to be generated in advance.

[0102] The learning unit 133 then multiplies the value output from the multilayer perceptron MLP3 by the value output from the multilayer perceptron MLP4, thereby outputting a color in which the amount of peripheral illumination has been compensated.

[0103] The learning model will be explained in more detail with reference to Fig. 15. In the example of Fig. 15, the position coordinate x of a predetermined object is 3d are position-encoded before being input to the multilayer perceptron MLP1. Subsequently, the learning unit 133 inputs the position-encoded position coordinates to the multilayer perceptron MLP1. In this case, the learning unit 133 inputs the position-encoded position coordinates to the previous layer of the multilayer perceptron MLP1. Furthermore, the learning unit 133 further inputs the position-encoded position coordinates to the subsequent layer of the multilayer perceptron MLP1.

[0104] The learning unit 133 then inputs the value output from the multilayer perceptron MLP1 to the multilayer perceptron MLP2, thereby outputting the density σ. Next, the learning unit 133 calculates the density σ for the multilayer perceptron MLP3 by calculating the value output from the multilayer perceptron MLP1 and the angle d based on the direction of the light ray. 3d In this case, the learning unit 133 outputs the color C.

[0105] Then, the learning unit 133 calculates the position d 2d is input to the multilayer perceptron MLP4. Also, the learning unit 133 inputs the position d 2d The learning unit 133 then multiplies the value output from the multilayer perceptron MLP3 by the value output from the multilayer perceptron MLP4 to obtain a color C CIn this case, the value output from the multilayer perceptron MLP4 becomes a compensation coefficient for compensating for the amount of light. In this way, the learning unit 133 outputs a color C with the amount of peripheral light compensated. C The learning model is trained to output

[0106] In this way, when the multilayer perceptron becomes complicated because it has to predict a change in light amount, the learning unit 133 can simplify the multilayer perceptron by using vector values ​​selected based on the pixel position. In the example of Fig. 15, the learning unit 133 can input optimized vector values ​​to the multilayer perceptron MLP4, making it possible to easily predict the compensation coefficients even when the calculation performance of the multilayer perceptron MLP4 is low.

[0107] (1-3-8. Learning Example (4)) In the above embodiment, the position d 2d Although the above description is based on an example in which the image is position-encoded before being input to the multilayer perceptron, this is not limiting. For example, the position within the image may be calculated as an embedding vector. In this case, the learning model includes at least three multilayer perceptrons.

[0108] The learning process of the learning model according to the embodiment will be described with reference to Fig. 16. Fig. 16 is a diagram showing another example of a learning model using embedding vectors. The learning model shown in Fig. 16 includes at least three multilayer perceptrons MLP1 to MLP3.

[0109] In the learning model shown in FIG. 16, compared to the learning model shown in FIG. 13, the learning unit 133 uses a compensation coefficient for compensating for the amount of light at the position d 2d Specifically, the learning unit 133 uses a vector value selected based on the position d 2d The vector value selected based on the above is multiplied by the value output from the multilayer perceptron MLP4 to output a color compensated for peripheral illumination. The vector value to be multiplied is a vector value selected from the latent code LC1.

[0110] The learning model will be explained in more detail with reference to Fig. 16. In the example of Fig. 16, the position coordinate x of a predetermined object is3d are position-encoded before being input to the multilayer perceptron MLP1. Subsequently, the learning unit 133 inputs the position-encoded position coordinates to the multilayer perceptron MLP1. In this case, the learning unit 133 inputs the position-encoded position coordinates to the previous layer of the multilayer perceptron MLP1. Furthermore, the learning unit 133 further inputs the position-encoded position coordinates to the subsequent layer of the multilayer perceptron MLP1.

[0111] The learning unit 133 then inputs the value output from the multilayer perceptron MLP1 to the multilayer perceptron MLP2, thereby outputting the density σ. Next, the learning unit 133 calculates the density σ for the multilayer perceptron MLP3 by calculating the value output from the multilayer perceptron MLP1 and the angle d based on the direction of the light ray. 3d In this case, the learning unit 133 outputs the color C.

[0112] Then, the learning unit 133 calculates the position d 2d The vector value selected based on the above is multiplied by the value output from the multilayer perceptron MLP3 to obtain a color C with the amount of light compensated. C In this way, the learning unit 133 outputs the color C C The learning model is trained to output

[0113] In this way, the learning unit 133 uses vector values ​​selected based on pixel positions as compensation coefficients, making the multilayer perceptron MLP4 unnecessary in the case of a learning model such as that shown in Fig. 15. This allows the learning unit 133 to further simplify the learning model.

[0114] Note that a learning model that uses latent code LC1 may require a large memory size. Generally, the memory size increases with the fineness (resolution) of the position within the image. Therefore, the calculation unit 132 may adjust the memory size to be smaller by using the method for calculating the position within the image shown in FIG. 10 or the method for calculating the position within the image shown in FIG. 12.

[0115] 1-4. Flowchart showing the procedure of the generation process according to the embodiment Next, the procedure of the generation process executed by the learning device 100 according to the embodiment will be described with reference to Fig. 17. Fig. 17 is an example of a flowchart showing the flow of the generation process according to the embodiment.

[0116] 17 , the acquisition unit 131 acquires a plurality of images including a predetermined target (step S101). Subsequently, the learning unit 133 inputs the plurality of images acquired by the acquisition unit 131 into a learning model, thereby learning the learning model (step S102). Then, the generation unit 134 generates a free viewpoint image of the predetermined target using the learning model learned by the learning unit 133 (step S103).

[0117] 1-5. Flowchart showing the procedure of the learning process according to the embodiment Next, the procedure of the learning process executed by the learning device 100 according to the embodiment will be described with reference to Fig. 18. Fig. 18 is an example of a flowchart showing the flow of the learning process according to the embodiment.

[0118] 18, the calculation unit 132 calculates a position within an image (step S201). Subsequently, the learning unit 133 inputs the position within the image calculated by the calculation unit 132 into a learning model, thereby learning the learning model (step S202).

[0119] Then, the learning unit 133 determines whether the difference between the color at the position in the image and the color predicted by the learning model is less than a predetermined threshold (step S203).

[0120] For example, if the difference between the color at a position in the image and the color predicted by the learning model is not less than a predetermined threshold (step S203; No), the learning unit 133 determines that the difference is not less than the predetermined threshold. In this case, the learning unit 133 updates the coefficients of the learning model so that the difference between the color at a position in the image and the color predicted by the learning model becomes smaller (step S204). Then, the process returns to before step S203, and the learning unit 133 updates the coefficients of the learning model until the difference between the color at a position in the image and the color predicted by the learning model becomes less than the predetermined threshold.

[0121] On the other hand, if the difference between the color at the position in the image and the color predicted by the learning model is less than the predetermined threshold (step S203; Yes), the learning unit 133 determines that the difference is less than the predetermined threshold. In this case, the learning unit 133 ends the learning process.

[0122] (1-6. Modifications of the Embodiment) The information processing according to the embodiment described above may be modified in various ways. Modifications of the embodiment will be described below.

[0123] (1-6-1. Application to Other Learning Models (1)) In the above embodiment, an example has been described in which the learning unit 133 learns an NeRF-based learning model by inputting positions within each image to the NeRF-based learning model, but the present invention is not limited to this. For example, the learning model may be any learning model as long as it is a learning model that can replace the NeRF-based learning model.

[0124] For example, the learning model may be a learning model based on Instant-NGP (Instant Neural Graphics Primitives) instead of the learning model based on NeRF. In this case, the learning unit 133 learns the learning model based on Instant-NGP by inputting positions in each image to the learning model based on Instant-NGP.

[0125] Below, a learning example of a learning model based on Instant-NGP will be described. First, the learning process of a learning model based on Instant-NGP according to a modified example will be described using Fig. 19. Fig. 19 is a diagram showing a learning model based on Instant-NGP according to a modified example. The learning model M2 shown in Fig. 19 will be described as a learning model based on Instant-NGP.

[0126] In the example of FIG. 19, the learning unit 133 calculates the angle d based on the direction of the light ray. 3d and the position d in the image 2d and are input to the learning model M2. In this case, the angle d based on the direction of the light ray 3d and position d in the image 2dare position-encoded before being input to the learning model M2.

[0127] Then, the learning unit 133 calculates the density σ and the color C C In this way, the learning unit 133 outputs the color C C The learning model is trained to output

[0128] Next, another example of the learning process of the learning model according to the modified example will be described with reference to Fig. 20. Fig. 20 is a diagram showing a learning model according to the modified example to which one multilayer perceptron has been added. In the example of Fig. 20, the learning model includes a learning model M2 and a multilayer perceptron MLP4.

[0129] In the example of FIG. 20, the learning unit 133 calculates the angle d based on the direction of the light ray. 3d is input to the learning model M2. In this case, the angle d based on the direction of the light ray 3d is position-encoded before being input to the training model M2.

[0130] Then, the learning unit 133 outputs the density σ and the color C. Next, the learning unit 133 calculates the density σ and the color C at the position d 2d is input to the multilayer perceptron MLP4. Here, the position d 2d is position-encoded before being input to the multi-layer perceptron MLP4.

[0131] Then, the learning unit 133 multiplies the value output from the learning model M2 by the value output from the multilayer perceptron MLP4 to obtain a color C with the light amount compensated. C In this way, the learning unit 133 outputs the color C C The learning model is trained to output

[0132] Next, another example of the learning process of the learning model according to the modified example will be described with reference to Fig. 21. Fig. 21 is a diagram showing a learning model according to the modified example to which two multilayer perceptrons have been added. In the example of Fig. 21, the learning model includes a learning model M2, a multilayer perceptron MLP2, and a multilayer perceptron MLP3.

[0133] In the example of FIG. 21, the learning unit 133 calculates the position coordinate x of a predetermined object. 3d and angle d based on the direction of the light ray 3d and are input to the learning model M2. The learning unit 133 then inputs the value output from the learning model M2 to the multilayer perceptron MLP2. The learning unit 133 then outputs the density σ. The learning unit 133 then inputs the value output from the learning model M2 to the multilayer perceptron MLP3. The learning unit 133 then outputs the color C.

[0134] Then, the learning unit 133 calculates the position d 2d The vector value selected based on the above is multiplied by the value output from the multilayer perceptron MLP3 to obtain a color C with the amount of light compensated. C In this way, the learning unit 133 outputs the color C C The learning model is trained to output

[0135] In this way, the learning unit 133 can reduce the amount of calculation required for learning by using a learning model based on Instant-NGP, thereby shortening the learning time.

[0136] (1-6-2. Application to Other Learning Models (2)) Furthermore, for example, the learning model may be a 3DGS-based learning model instead of the NeRF-based learning model. In this case, the learning unit 133 learns the 3DGS-based learning model by inputting positions in each image to the 3DGS-based learning model.

[0137] 3DGS represents a scene using primitives, which are three-dimensional representations with dimensions called Gaussians. In this case, it starts with sparse points captured by the image capture device 10. 3DGS then generates a free-viewpoint image by adjusting the Gaussians. In the following, an example of a learning model based on 3DGS will be described, in which a process of inputting positions within each image is applied to the rasterization process.

[0138] The learning process of a learning model according to a modified example will be described with reference to Fig. 22. Fig. 22 is a diagram showing an example (1) of a learning model based on 3DGS according to a modified example. The learning model M3 shown in Fig. 22 will be described as a learning model based on 3DGS.

[0139] For example, in the learning model M3, (1) a fragment is generated based on an image (2). Then, in the learning model M3, (3) a rendered image is generated, and (4) the coefficients of the learning model M3 are updated based on a loss function. An overview of this processing flow is described below.

[0140] For example, in (1), a Gaussian is generated. Here, the Gaussian has parameters: mean μ, which indicates position; variance Σ, which indicates size and shape; opacity α; and color c. Next, by rasterization, once the camera viewpoint for rendering is determined, the Gaussian projected onto the image is converted into pixels. In this case, the converted pixels are called (2) fragments.

[0141] Then, a fragment is generated for each Gaussian. Next, a rendered image (3) is generated using a shader and blending based on the fragment. The difference between the ground truth image and the rendered image is calculated as a loss, and the Gaussian parameters are updated (loss function (4)) using backpropagation or the like to reduce this loss.

[0142] The following describes a case where a multilayer perceptron MLP5 is added between (3) and (4) with reference to FIG. 22. In this case, the learning unit 133 performs the learning for the position d 2d is input to the multilayer perceptron MLP5. At this time, the position d 2d is position-encoded before being input to the multi-layer perceptron MLP5. The learning unit 133 performs learning with the configuration of the learning model M3. In this way, the learning unit 133 obtains a color C C The learning model M3 is trained to output

[0143] This allows the learning unit 133 to learn to compensate for the peripheral illumination of the image. Furthermore, when learning, the learning unit 133 calculates the difference from the correct data using the compensated pixel values, for example, and therefore can reduce the updating of the Gaussian parameters due to backpropagation of loss caused by a decrease in the peripheral illumination of the image.

[0144] 22, the learning unit 133 adds processing to the rendering image, so that the number of changes to the configuration of the learning model M3 can be reduced. Furthermore, because the learning unit 133 only needs to change a small number of parts, the additional amount of processing can be reduced.

[0145] Next, a case where a multilayer perceptron MLP5 is added between (2) and (3) will be described with reference to FIG. 23. FIG. 23 is a diagram showing an example (2) of a learning model based on a 3DGS according to a modified example. In the example of FIG. 23, the learning unit 133 calculates a position d 2d is input to the multilayer perceptron MLP5. In this case, the position d 2d is position-encoded before being input to the multi-layer perceptron MLP5. The learning unit 133 performs learning with the configuration of the learning model M3. In this way, the learning unit 133 obtains a color C C 23, the learning unit 133 can learn to compensate for the amount of peripheral light in an image.

[0146] Next, a case where a multilayer perceptron MLP5 is added between (1) and (2) will be described with reference to FIG. 24. FIG. 24 is a diagram showing an example (3) of a learning model based on a 3DGS according to a modified example. In the example of FIG. 24, the learning unit 133 calculates a position d 2d is input to the multilayer perceptron MLP5. In this case, the position d 2d is position-encoded before being input to the multi-layer perceptron MLP5. The learning unit 133 performs learning with the configuration of the learning model M3. In this way, the learning unit 133 obtains a color C CThe learning model M3 is trained to output the following. In this way, in the learning model shown in FIG. 24, the reduction in peripheral illumination is compensated for even in fragment images. Therefore, the learning unit 133 can learn to compensate for the peripheral illumination of the image.

[0147] (1-6-3. Other Learning Methods) In the above embodiment, an example has been described in which the learning unit 133 learns a learning model based on NeRF by inputting positions within each image to the learning model based on NeRF. However, the present invention is not limited to this. For example, the learning unit 133 may learn so as not to use a portion of an image where peripheral light intensity is reduced in the learning.

[0148] For example, the calculation unit 132 calculates a position within the image. Next, the calculation unit 132 identifies a position within the image where a reduction in peripheral light intensity occurs, based on the correspondence between the position within the image and the light intensity at that position. Then, the learning unit 133 may train the learning model without using the position within the image identified by the calculation unit 132 for learning.

[0149] (2. Other Embodiments) The processing according to each of the above-described embodiments may be implemented in various different forms other than the above-described embodiments.

[0150] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0151] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0152] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.

[0153] Furthermore, the effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0154] (3. Effects of the Learning Method According to the Present Disclosure) As described above, the learning method according to the present disclosure includes a computer (in an embodiment, the learning device 100) acquiring multiple images including a specified object, calculating a position within each image based on the multiple images, and inputting the calculated positions within each image into a learning model that generates a free viewpoint image of the specified object, thereby learning a learning model.

[0155] In this way, the learning method can predict and compensate for vignetting in an image, and can avoid vignetting from becoming an obstacle to training the learning model when using actual images, thereby improving the accuracy of the free-viewpoint images generated by the learning model.

[0156] Furthermore, the learning method can improve the quality of free-viewpoint images by reducing image quality degradation. Generally, the influence of the cosine fourth power law is significant when shooting using a wide-angle lens. Even in such cases, the learning method can take the influence of the cosine fourth power law into account, thereby improving the quality of free-viewpoint images.

[0157] Generally, peripheral light falloff in the image capture device 10 may be automatically adjusted by processing within the image capture device 10. However, this adjustment is often not complete, so the learning method can provide a solution to eliminate peripheral light falloff. Furthermore, when RAW data is used, automatic adjustment of peripheral light falloff within the image capture device 10 is not performed, so the learning method can provide a solution to eliminate peripheral light falloff. Furthermore, the learning method can also intentionally simulate peripheral light falloff to generate a free viewpoint image.

[0158] Furthermore, when one image capturing device 10 is used, the learning method uses multiple images, but because the image capturing device 10 is common, only one parameter for the position within the image is required. Therefore, the learning method can reduce the number of parameters required for learning.

[0159] Furthermore, the computer calculates the position of a pixel in each image as the position in each image, and inputs the pixel position into the learning model, thereby learning the learning model.

[0160] This allows the learning method to identify the location of the light dropoff with sufficient resolution because it can indicate the location within the image based on the image coordinates. For example, if the shape of the lens hood is asymmetric, the light dropoff may not be symmetric. In this case, the learning method can identify the location of the light dropoff with sufficient resolution even if the light dropoff is asymmetric.

[0161] In addition, the computer calculates positions within each image based on areas into which each image is divided into a grid of a predetermined size, and inputs the positions based on the areas into the learning model, thereby learning the learning model.

[0162] This allows the learning method to reduce the memory size used, for example, the memory size used when calculating positions based on areas divided into a grid pattern can be smaller than the memory size used when calculating pixel positions.

[0163] Furthermore, the computer calculates a position within each image based on the distance from the center of each image, and inputs the distance-based positions into the learning model, thereby learning the learning model.

[0164] This allows the learning method to reduce the degree of freedom of variables by representing positions within an image as distances from the center of the image, thereby reducing the dimensions of the input variables and making the learning process easier to execute.

[0165] The computer also calculates positions within each image based on an area divided by multiple circles having a predetermined radius with the center of each image as the center of the circle, the area being formed by the circumference of a first circle and the circumference of a second circle, and inputs the positions based on the areas into the learning model, thereby learning the learning model.

[0166] As a result, the learning method can reduce the memory size used to calculate positions within an image because the calculated positions are discretized. For example, the learning method can reduce the memory size used when calculating positions based on areas divided by multiple circles compared to the memory size used when calculating positions based on distances from the center of the image.

[0167] Furthermore, the computer inputs positions within each image into the NeRF-based learning model as a learning model, thereby learning the NeRF-based learning model.

[0168] This allows the learning method to perform learning while compensating for a decrease in light amount due to peripheral light intensity of an image, etc., in the learning process of a learning model based on NeRF.

[0169] In addition, the computer inputs the position coordinates of a specified object, an angle based on the direction of a light ray calculated from the image, and a position within the image into a learning model based on NeRF, and learns to output density and a color compensated for the amount of light in the image.

[0170] As a result, the learning method can perform learning while compensating for the reduction in light intensity due to peripheral light intensity, etc., by further inputting the position within the image along with the position coordinates of a specified object and an angle based on the direction of the light ray in the learning process of the NeRF-based learning model.

[0171] In addition, in a learning model based on NeRF in which a computer includes at least three multilayer perceptrons, the computer learns to input the position coordinates of a predetermined object to a first multilayer perceptron, input the value output from the first multilayer perceptron to a second multilayer perceptron to output density, and input the value output from the first multilayer perceptron, an angle based on the direction of the light ray, and a position within the image to a third multilayer perceptron to output a color compensated for the amount of light in the image.

[0172] This allows the learning method to perform learning while compensating for reductions in light intensity due to peripheral light, etc., by further inputting the position within the image to the third multilayer perceptron, along with the value output from the first multilayer perceptron and an angle based on the direction of the light ray.

[0173] Furthermore, in a learning model based on NeRF in which a computer includes at least four multilayer perceptrons, the computer learns to input the position coordinates of a predetermined object to a first multilayer perceptron, input the value output from the first multilayer perceptron to a second multilayer perceptron to output density, input the value output from the first multilayer perceptron and an angle based on the direction of a light ray to a third multilayer perceptron, input the position within the image to a fourth multilayer perceptron, and multiply the value output from the third multilayer perceptron by the value output from the fourth multilayer perceptron to output a color compensated for the amount of light in the image.

[0174] For example, the training method may further add a multilayer perceptron for inputting a position within an image. In this way, the training method may add, for example, a fourth multilayer perceptron having a lower computational capability than the third multilayer perceptron. This allows the training method to select an appropriate multilayer perceptron according to the input information, thereby enabling a suitable training model to be trained.

[0175] Furthermore, in a learning model based on NeRF in which a computer includes at least four multilayer perceptrons, the computer learns to input position coordinates of a predetermined object to a first multilayer perceptron, input the value output from the first multilayer perceptron to a second multilayer perceptron to output density, input the value output from the first multilayer perceptron and an angle based on the direction of a light ray to a third multilayer perceptron, input the position within the image and a vector value selected based on the position within the image to a fourth multilayer perceptron, and multiply the value output from the third multilayer perceptron by the value output from the fourth multilayer perceptron to output a color compensated for the amount of light in the image.

[0176] This allows the training method to simplify the multi-layer perceptron, which would otherwise be complicated due to predicting changes in light levels, by using vector values ​​that are selected based on location in the image.

[0177] Furthermore, in a NeRF-based learning model in which a computer includes at least three multilayer perceptrons, the computer learns to input the position coordinates of a predetermined object to a first multilayer perceptron, input the value output from the first multilayer perceptron to a second multilayer perceptron to output density, input the value output from the first multilayer perceptron and an angle based on the direction of a light ray to a third multilayer perceptron, and multiply the value output from the third multilayer perceptron by a vector value selected based on the position within the image to output a color compensated for the amount of light in the image.

[0178] In this way, the learning method can reduce the number of multilayer perceptrons included in the learning model by using vector values ​​selected based on the position in the image as compensation coefficients for compensating for the amount of light, thereby simplifying the learning model.

[0179] Furthermore, the computer learns the learning model based on Instant-NGP by inputting the positions in each image to the learning model based on Instant-NGP as the learning model.

[0180] This allows the learning method to learn other learning models, such as a learning model based on Instant-NGP, while compensating for the reduction in light intensity due to peripheral light intensity in the image, etc., thereby providing a highly scalable learning method.

[0181] Furthermore, the computer inputs positions within each image into a learning model based on 3D Gaussian splatting as a learning model, thereby learning the learning model based on 3D Gaussian splatting.

[0182] This allows the learning method to learn other learning models, such as a learning model based on 3D Gaussian Splatting, while compensating for reductions in light intensity due to peripheral light intensity, etc., thereby providing a highly scalable learning method.

[0183] The computer also acquires a plurality of images including a first image of the predetermined object photographed from a first direction and a second image of the predetermined object photographed from a second direction.

[0184] This allows the learning method to acquire images of a predetermined object taken from multiple directions, and also allows the learning method to acquire multiple images necessary for a learning model that generates a free viewpoint image of the predetermined object.

[0185] The computer also trains the learning model by inputting the locations within each image to compensate for the changes in the amount of light each image exhibits.

[0186] As a result, the learning method can learn the learning model while compensating for the reduction in the amount of light in the image, thereby generating high-quality free-viewpoint images.

[0187] In addition, the computer learns a learning model by inputting positions within each image to compensate for changes in peripheral illumination shown in each image, which is the change in light intensity from the center of the image to the periphery of the image.

[0188] As a result, the learning method can learn the learning model while compensating for the reduction in light amount due to the peripheral light amount of the image, thereby generating high-quality free viewpoint images.

[0189] In addition, the computer learns the learning model by inputting the positions within each image to compensate for changes in the amount of light due to vignetting contained in each image.

[0190] As a result, the learning method can learn the learning model while compensating for the reduction in light amount due to vignetting, thereby generating high-quality free-viewpoint images.

[0191] The method further includes the computer using the learning model to generate a free viewpoint image based on each image in which changes in the amount of light shown in each image have been compensated for.

[0192] As a result, the learning method trains the learning model based on images in which reductions in light intensity due to peripheral light intensity, vignetting, etc. are reduced, making it possible to generate high-quality free viewpoint images.

[0193] (4. Hardware Configuration) Information devices such as the learning device 100 and imaging device 10 according to each of the above-described embodiments are realized by a computer 1000 configured as shown in FIG. 25, for example. The following description will use the learning device 100 according to the embodiments as an example. FIG. 25 is a hardware configuration diagram showing an example of a computer 1000 that realizes the functions of the learning device 100. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0194] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0195] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .

[0196] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records a conversion program according to the present disclosure, which is an example of program data 1450.

[0197] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0198] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, or semiconductor memory.

[0199] For example, when computer 1000 functions as learning device 100 according to an embodiment, CPU 1100 of computer 1000 executes a learning program loaded onto RAM 1200 to realize the functions of control unit 130, etc. Also, HDD 1400 stores the learning program according to the present disclosure and data in storage unit 120. Note that CPU 1100 reads and executes program data 1450 from HDD 1400, but as another example, these programs may be acquired from another device via external network 1550.

[0200] Note that the present technology can also be configured as follows. (1) A learning method including a computer acquiring a plurality of images including a predetermined object, calculating a position within each image based on the plurality of images, and inputting the calculated position within each image into a learning model that generates a free viewpoint image of the predetermined object, thereby learning the learning model. (2) The learning method according to (1), in which the computer calculates, as the position within each image, a position of a pixel within the image, and inputs the pixel position into the learning model, thereby learning the learning model. (3) The learning method according to (1) or (2), in which the computer calculates, as the position within each image, a position based on regions obtained by dividing each image into a grid of a predetermined size, and inputs the position based on the region into the learning model, thereby learning the learning model. (4) The learning method according to any one of (1) to (3), in which the computer calculates, as the position within each image, a position based on a distance from the center of each image, and inputs the position based on the distance into the learning model, thereby learning the learning model. (5) The learning method according to any one of (1) to (4), wherein the computer calculates, as positions within each of the images, positions based on an area divided by a plurality of circles having a predetermined radius with the center of each of the images as the center of the circle, the area being formed by the circumference of a first circle and the circumference of a second circle, and inputs the positions based on the areas into the learning model, thereby training the learning model. (6) The learning method according to any one of (1) to (5), wherein the computer inputs positions within each of the images to an Neural Radiance Fields (NeRF)-based learning model as the learning model, thereby training the NeRF-based learning model. (7) The learning method according to (6), wherein the computer inputs position coordinates of the predetermined object, an angle based on the direction of a light ray calculated from the image, and a position within the image into the NeRF-based learning model, thereby training the computer to output density and a color compensated for the amount of light of the image.(8) The learning method according to (1), wherein the multiple images are captured by the same camera. (9) The learning method according to (1), wherein the multiple images are captured by one or more camera devices having the same optical characteristics. (10) The learning method according to (1), wherein, after the learning, the learning model is stored in a non-transitory storage medium. (11) The learning method according to (10), wherein the non-transitory storage medium is part of an image generation device capable of generating free-viewpoint images based on a predetermined learning model, and after the learning, the learning model is stored in the non-transitory storage medium to manufacture the image generation device that generates the free-viewpoint images based on each of the images. (12) The learning method according to any one of (1) to (5), wherein the computer learns a learning model based on Instant-NGP (Instant Neural Graphics Primitives) by inputting positions in each of the images to the learning model. (13) The learning method according to any one of (1) to (5), wherein the computer learns a learning model based on 3D Gaussian splatting by inputting positions within each of the images to the learning model based on 3D Gaussian splatting. (14) The learning method according to any one of (1) to (13), wherein the computer acquires the plurality of images including a first image in which the predetermined object is photographed from a first direction and a second image in which the predetermined object is photographed from a second direction. (15) The learning method according to any one of (1) to (14), wherein the computer learns the learning model by inputting positions within each of the images to compensate for changes in the amount of light shown in each of the images. (16) The learning method described in (15) in which the computer learns the learning model by inputting positions within each image to compensate for changes in peripheral illumination shown in each image, where the amount of light changes from the center of the image to the periphery of the image.(17) The learning method according to (15) or (16), wherein the computer learns the learning model by inputting positions within each of the images to compensate for changes in light intensity due to vignetting included in each of the images. (18) The learning method according to any one of (1) to (17), further comprising using the learning model to generate the free-viewpoint images based on each of the images in which changes in light intensity shown in each of the images have been compensated for. (19) The learning method according to (7), wherein the computer learns the NeRF-based learning model including at least three multilayer perceptrons by inputting position coordinates of the predetermined object to a first multilayer perceptron, inputting values ​​output from the first multilayer perceptron to a second multilayer perceptron to output a density, and inputting the values ​​output from the first multilayer perceptron, an angle based on the direction of the light ray, and a position within the image to a third multilayer perceptron to output a color in which the light intensity of the image has been compensated for. (20) The learning method described in (7), wherein the computer, in the NeRF-based learning model including at least four multilayer perceptrons, inputs the position coordinates of the predetermined object to a first multilayer perceptron, outputs density by inputting the value output from the first multilayer perceptron to a second multilayer perceptron, inputs the value output from the first multilayer perceptron and an angle based on the direction of the light ray to a third multilayer perceptron, and inputs the position within the image to a fourth multilayer perceptron and multiplies the value output from the third multilayer perceptron by the value output from the fourth multilayer perceptron, thereby learning to output a color compensated for the amount of light in the image.(21) The learning method described in (7), in which the computer learns to output a color compensated for the amount of light in the image by inputting the position coordinates of the specified object to a first multilayer perceptron in the NeRF-based learning model including at least four multilayer perceptrons, inputting the value output from the first multilayer perceptron to a second multilayer perceptron to output a density, inputting the value output from the first multilayer perceptron and an angle based on the direction of the light ray to a third multilayer perceptron, and inputting the position in the image and a vector value selected based on the position in the image to a fourth multilayer perceptron, and multiplying the value output from the third multilayer perceptron by the value output from the fourth multilayer perceptron, so as to output a color compensated for the amount of light in the image. (22) The learning method according to (7), wherein the computer learns to: input position coordinates of the predetermined object to a first multilayer perceptron in the NeRF-based learning model including at least three multilayer perceptrons; input a value output from the first multilayer perceptron to a second multilayer perceptron to output a density; input the value output from the first multilayer perceptron and an angle based on the direction of the light ray to a third multilayer perceptron; and multiply the value output from the third multilayer perceptron by a vector value selected based on the position within the image to output a color in which the amount of light of the image is compensated. (23) A learning device comprising: an acquisition unit that acquires multiple images including the predetermined object; a calculation unit that calculates a position within each image based on the multiple images; and a learning unit that trains a learning model that generates a free viewpoint image of the predetermined object by inputting the positions within each image calculated by the calculation unit to the learning model. (24) A learning program that causes a computer to function as a learning device having an acquisition unit that acquires multiple images including a specified object, a calculation unit that calculates a position within each image based on the multiple images, and a learning unit that learns a learning model that generates a free viewpoint image of the specified object by inputting the positions within each image calculated by the calculation unit into the learning model.(25) A method for manufacturing an image generation device that generates a free viewpoint image based on a predetermined learning model stored in a non-transitory storage device, comprising: acquiring a plurality of images including a predetermined object; calculating a position within each image based on the plurality of images; inputting the calculated position within each image into a learning model that generates a free viewpoint image of the predetermined object to train the learning model; and storing the trained learning model in the non-transitory storage device. (26) The manufacturing method according to (25), further comprising providing a display device that can display an image generated based on the learning model stored in the non-transitory storage device.

[0201] REFERENCE SIGNS LIST 1 Learning system 10 Imaging device 100 Learning device 110 Communication unit 120 Storage unit 121 Image storage unit 122 Learning model storage unit 123 Free viewpoint image storage unit 130 Control unit 131 Acquisition unit 132 Calculation unit 133 Learning unit 134 Generation unit 140 Input unit 150 Display unit

Claims

1. A learning method including: a computer acquiring a plurality of images including a predetermined object; calculating a position within each image based on the plurality of images; and inputting the calculated positions within each image into a learning model that generates a free viewpoint image of the predetermined object, thereby learning the learning model.

2. The learning method according to claim 1, wherein the computer calculates the position of a pixel in each of the images as the position in each of the images, and inputs the pixel position into the learning model, thereby learning the learning model.

3. The learning method according to claim 1, wherein the computer calculates positions within each image based on areas obtained by dividing each image into a grid of a predetermined size, and inputs the positions based on the areas into the learning model, thereby learning the learning model.

4. The learning method according to claim 1, wherein the computer calculates a position within each image based on a distance from the center of each image, and inputs the distance-based position into the learning model, thereby learning the learning model.

5. The learning method according to claim 1, wherein the computer calculates, as a position within each of the images, a region divided by a plurality of circles having a predetermined radius with the center of each of the images as the center of the circle, the region being formed by the circumference of a first circle and the circumference of a second circle, and inputs the positions based on the regions into the learning model, thereby learning the learning model.

6. The learning method according to claim 1, wherein the computer learns a learning model based on Neural Radiance Fields (NeRF) by inputting positions within each image into the learning model.

7. The learning method described in claim 6, wherein the computer learns to output density and a color compensated for the amount of light in the image by inputting the position coordinates of the specified object, an angle based on the direction of a light ray calculated from the image, and a position within the image into the learning model based on the NeRF.

8. The learning method according to claim 1, wherein the plurality of images are taken by the same photographing device.

9. The learning method according to claim 1, wherein the plurality of images are taken by one or more image capturing devices having the same optical characteristics.

10. The learning method according to claim 1, wherein after said learning, said learning model is stored in a non-transitory storage medium.

11. The learning method described in claim 10, wherein the non-transitory storage medium is part of an image generation device capable of generating free viewpoint images based on a predetermined learning model, and after the learning, the learning model is stored in the non-transitory storage medium to manufacture the image generation device that generates the free viewpoint images based on each of the images.

12. The learning method according to claim 1, wherein the computer learns a learning model based on Instant-NGP (Instant Neural Graphics Primitives) by inputting positions within each image into the learning model based on Instant-NGP as the learning model.

13. The learning method according to claim 1, wherein the computer learns a learning model based on 3D Gaussian splatting by inputting positions within each of the images to the learning model based on 3D Gaussian splatting.

14. The learning method of claim 1, wherein the computer acquires the plurality of images including a first image of the specified object photographed from a first direction and a second image of the specified object photographed from a second direction.

15. The learning method of claim 1, wherein the computer trains the learning model by inputting positions within each of the images to compensate for changes in the amount of light shown in each of the images.

16. The learning method of claim 15, wherein the computer learns the learning model by inputting positions within each of the images to compensate for changes in peripheral illumination shown in each of the images, where the amount of light changes from the center of the image to the periphery of the image.

17. The learning method according to claim 15, wherein the computer learns the learning model by inputting positions within each of the images to compensate for changes in the amount of light due to vignetting contained in each of the images.

18. The learning method of claim 1, further comprising using the learning model to generate the free viewpoint images based on each image in which changes in the amount of light shown in each image have been compensated for.

19. A learning device comprising: an acquisition unit that acquires multiple images including a specified object; a calculation unit that calculates a position within each image based on the multiple images; and a learning unit that trains a learning model that generates a free viewpoint image of the specified object by inputting the positions within each image calculated by the calculation unit into the learning model.

20. A learning program that causes a computer to function as a learning device having an acquisition unit that acquires multiple images including a specified object, a calculation unit that calculates a position within each image based on the multiple images, and a learning unit that learns a learning model that generates a free viewpoint image of the specified object by inputting the positions within each image calculated by the calculation unit into the learning model.

Citation Information

Patent Citations

  • Imaging apparatus, control method of the same, and program

    JP2021052282A

  • Information processing device and information generation method

    JP2023079022A