Information processing apparatus, information processing method, and storage medium
By learning the model to estimate the potential variables of the captured image and recover high-dimensional parameters, the problem of difficulty in estimating and capturing the physical characteristics of the environment in the prior art is solved, and the natural degree of image synthesis is improved.
Patent Information
- Application Number
- CN202380079915.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-25
- Filing Date
- 2023-11-08
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to estimate the physical characteristics of the capture environment with high accuracy, making it difficult to generate natural images when combining the captured image and the CG image.
By estimating potential variables of captured images using a learning model, compressing high-dimensional parameters to low-dimensionality, recovering high-dimensional parameters to indicate the physical characteristics of the capture environment, and using them to combine the captured image and CG images.
The physical characteristics of the capture environment are achieved with high-precision estimation, which improves the naturalness of image synthesis, making the generated image look more realistic.
Smart Images

Figure CN120226053A_ABST
Abstract
Description
Technical Field
[0001] The present technology relates to an information processing apparatus, an information processing method, and a storage medium, and particularly to an information processing apparatus, an information processing method, and a storage medium capable of estimating information on physical characteristics of a capture environment with high accuracy. Background Art
[0002] In recent years, in order to train AI (Artificial Intelligence) and the like, it has been necessary to prepare a large number of images. The accuracy of AI depends on the quality and quantity of the images used in training, and therefore, it is desirable to have various realistic images as the images for training.
[0003] Collecting a large number of images for training from captured images obtained by photographing an actual scene or the like is troublesome. At the same time, generating a large number of realistic images using CG (Computer Graphics) as the images for training takes time and is troublesome. Since it is troublesome to collect images for training using these techniques, a technique for easily generating a large number of images by combining a captured image and a CG image as an image generated using CG has been proposed (for example, see NPL1).
[0004] Citation List
[0005] Non-Patent Literature
[0006] Non-Patent Literature 1
[0007] Zhengqin Li,Mohammad Shafiei,Ravi Ramamoorthi,Kalyan Sunkavalli,Manmohan Chandraker,“Inverse Rendering for Complex Indoor Scenes:Shape,Spatially-Varying Lighting and SVBRDF from a Single Image”,CVPR 2020 Summary of the Invention
[0008] Technical Problem
[0009] In the technique described in Non-Patent Literature 1, it is necessary to combine a captured image and a CG image so that the resulting image looks natural. In order to combine a captured image and a CG image so that the resulting image looks natural, information on the physical characteristics of the capture environment of the captured image is required. When the physical characteristics of the capture environment are unknown, it is necessary to estimate the physical characteristics based on the captured image. In the technique described in Non-Patent Literature 1, machine learning is used to estimate the physical characteristics.
[0010] Although the physical characteristics of the capture environment are indicated by numerous parameters, the number of parameters that can be estimated using machine learning is limited. Thus, with the technique described in Non-Patent Document 1, it is not possible to estimate the physical characteristics of the capture environment with high accuracy, and it is difficult to combine the captured image and the CG image so that the resulting image looks natural.
[0011] In view of the above, the present technique has been proposed, and the present technique enables the physical characteristics of the capture environment to be estimated with high accuracy.
[0012] Solution to the problem
[0013] An aspect of the present technique provides an information processing apparatus including: an estimation unit that estimates a latent variable of a capture environment regarding a captured image using a learning model, the captured image being input to the learning model and outputting the latent variable, the latent variable being obtained by compressing high-dimensional parameters to a lower dimension compared to the parameters, the parameters indicating physical characteristics of the capture environment regarding the captured image; and a recovery unit that recovers the parameters from the latent variable estimated by the estimation unit.
[0014] An aspect of the present technique provides an information processing method including: an information processing apparatus that estimates a latent variable of a capture environment regarding a captured image using a learning model, the captured image being input to the learning model and outputting the latent variable, the latent variable being obtained by compressing high-dimensional parameters to a lower dimension compared to the parameters, the parameters indicating physical characteristics of the capture environment regarding the captured image; and recovering the parameters from the estimated latent variable.
[0015] An aspect of the present technique provides a storage medium storing a program for executing a process including: estimating a latent variable of a capture environment regarding a captured image using a learning model, the captured image being input to the learning model and outputting the latent variable, the latent variable being obtained by compressing high-dimensional parameters to a lower dimension compared to the parameters, the parameters indicating physical characteristics of the capture environment regarding the captured image; and recovering the parameters from the estimated latent variable.
[0016] According to an aspect of the present technique, a latent variable of a capture environment regarding a captured image is estimated using a learning model, the captured image being input to the learning model and outputting the latent variable, the latent variable being obtained by compressing high-dimensional parameters to a lower dimension compared to the parameters, the parameters indicating physical characteristics of the capture environment regarding the captured image; and the parameters are recovered from the latent variable estimated by the estimation unit. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a block diagram showing an example of the configuration of an information processing apparatus according to an embodiment of the present technique.
[0018] Figure 2Shows an example of the synthesis of a captured image and a CG image.
[0019] Figure 3 Shows an example of learning data for training a learning model for estimating latent variables.
[0020] Figure 4 Shows an example of spectral reflectance characteristics.
[0021] Figure 5 Shows a technique for compressing high-dimensional parameters of 31 wavelengths into N latent variables.
[0022] Figure 6 Shows a process for recovering spectral reflectance characteristics from N latent variables.
[0023] Figure 7 Shows an example of a method for representing the incident direction and reflection direction of light.
[0024] Figure 8 Shows an example of a two-dimensional mapping.
[0025] Figure 9 Shows a technique for further compressing 324 high-dimensional parameters into (N + 2) latent variables using principal component analysis, which uses a second-variable BRDF model obtained by modeling the reflectance characteristics of an object.
[0026] Figure 10 Shows a process for recovering a two-dimensional mapping from (N + 2) latent variables.
[0027] Figure 11 Shows a technique for compressing 6,220,800 high-dimensional parameters into latent variables using principal component analysis.
[0028] Figure 12 Shows a technique for compressing 6,220,800 high-dimensional parameters into latent variables using principal component analysis.
[0029] Figure 13 Shows a process for recovering 3600 units from (5328 × N) latent variables.
[0030] Figure 14 Shows a process of a conventional technique for estimating physical properties.
[0031] Figure 15 Shows an example of learning data for training a learning model for estimating physical properties of a capture environment.
[0032] Figure 16 Shows an example of the synthesis of a captured image and a CG image according to a conventional technique.
[0033] Figure 17 This is a flowchart showing the processing executed by the information processing device.
[0034] Figure 18 This shows the flow of a technique for estimating high-dimensional parameters when performing scene recognition.
[0035] Figure 19 This shows the flow of a technique for estimating high-dimensional parameters when gradually estimating latent variables.
[0036] Figure 20 This is a block diagram showing an example of the hardware configuration of a computer. Detailed Implementation Modes
[0037] The modes for executing the present technology will be described below. The description will proceed in the following order.
[0038] 1. Configuration of the information processing device
[0039] 2. Examples of techniques for compressing high-dimensional parameters
[0040] 3. Operations of the information processing device
[0041] 4. Variations
[0042] <1. Configuration of the information processing device>
[0043] Figure 1 This is a block diagram showing an example of the configuration of the information processing device 1 according to an embodiment of the present technology.
[0044] Figure 1 The information processing device 1 in [description] is a device that estimates the physical characteristics of the capture environment of a capture image obtained by capturing an actual scene or the like and combines the capture image and a CG image generated using CG based on the estimated physical characteristics.
[0045] As Figure 1 shown, the information processing device 1 includes a latent variable estimation unit 11, a restoration unit 12, and a synthesis unit 13. As used herein, a "unit" refers to a circuit that can be configured by executing computer-readable instructions, and the circuit may include one or more local processors (e.g., CPUs), and / or one or more remote processors (such as cloud computing resources), or any combination thereof. One or more of these circuits may be on the same chip, on separate chips, or the functions of these circuits may be spread across different devices or platforms. As used herein, a "latent variable" is a variable that can be indirectly inferred from other observable variables (e.g., physical characteristics) in an image.
[0046] The latent variable estimation unit 11 estimates a latent variable for each pixel of the captured image based on the captured image. The latent variable is a compressed indication of the physical characteristics of the capture environment of the captured image. The physical characteristics of the capture environment (real world) include, for example, the shape of the surface of an object, the spectral reflectance characteristics of the surface of the object, the reflectance distribution characteristics of the surface of the object, and the spectral radiation characteristics of a light source, and are indicated by higher-dimensional parameters (hereinafter referred to as "high-dimensional parameters"). The latent variable is a parameter that can recover the high-dimensional parameters indicating the physical characteristics of the capture environment and is obtained by compressing the high-dimensional parameters to a lower dimension compared to the high-dimensional parameters.
[0047] Specifically, the latent variable estimation unit 11 estimates a latent variable for each of a plurality of types of physical characteristics using a learning model into which the captured image is input and which outputs the latent variable. The latent variable estimation unit 11 supplies the estimated latent variable to the recovery unit 12.
[0048] The recovery unit 12 recovers high-dimensional parameters (estimated physical characteristics) from the latent variable provided by the latent variable estimation unit 11, and supplies the high-dimensional parameters to the synthesis unit 13.
[0049] The synthesis unit 13 generates a composite image by combining the captured image and a CG image obtained by processing CG data using the high-dimensional parameters provided by the recovery unit 12. For example, the CG data is data of a three-dimensional model of an object to be virtually arranged in the capture environment of the captured image. The synthesis unit 13 can also reproduce the captured image based on the high-dimensional parameters, and combine the reproduced captured image with the CG image.
[0050] Figure 2 An example of the synthesis of the captured image and the CG image is shown.
[0051] In Figure 2 the example, the captured image P1 of a part of a room and the CG image P2 obtained by processing the CG data of a rabbit are combined to generate a composite image P3 that looks as if the rabbit is arranged on the floor of the room photographed in the captured image P1.
[0052] In the information processing apparatus 1 according to the present technology, the physical characteristics of the object and the light source in the capture environment of the captured image P1 are indicated by high-dimensional parameters. The synthesis unit 13 can process the captured image P1 and the CG data using a fine calculation formula based on the high-dimensional parameters so that the rabbit in the CG data looks natural when combined with the captured image P1.
[0053] For example, a shadow of the rabbit matching the direction of the light source can be added to a part of the captured image P1, or the color of the rabbit can be changed to a color matching the color temperature of the light source.
[0054] Figure 3 An example of learning data for training a learning model for estimating latent variables is shown.
[0055] For example, the learning model used by the latent variable estimation unit 11 to estimate latent variables can be trained using student data for learning and teacher data (expected values) for learning.
[0056] As Figure 3 shown, the student data can be, for example, an RGB image generated using CG.
[0057] Meanwhile, the teacher data can be, for example, a latent variable obtained by compressing high-dimensional parameters indicating the spectral reflectance characteristics or reflectance distribution characteristics of the surface of an object captured in the RGB image, a latent variable obtained by compressing high-dimensional parameters indicating the coordinates in the three-dimensional space of an object captured in the RGB image, and a latent variable obtained by compressing high-dimensional parameters indicating the spectral radiation characteristics of a light source in the RGB image.
[0058] The RGB image as student data is generated using CG, and thus, together with the RGB image, the physical characteristics of the object and the light source in the capture environment for the RGB image can be easily obtained as teacher data. By associating the physical attributes of a set of RGB images with latent variables having dimensions lower than the high-dimensional physical attributes, a training set including the set of images and the set of latent variables can be generated for training a machine, such as a neural network, which then generates a learning model for use in the latent variable estimation unit 11.
[0059] <2. Examples of techniques for compressing high-dimensional parameters>
[0060] Techniques for compressing high-dimensional parameters indicating the physical characteristics of the capture environment into lower-dimensional latent variables will be described below.
[0061] - Spectral reflectance characteristics
[0062] Figure 4 An example of spectral reflectance characteristics is shown. In Figure 4 , the horizontal axis represents the wavelength of the radiation light radiated onto the surface of the object, and the vertical axis represents the spectral reflectance (intensity of the reflected light).
[0063] As Figure 4 shown, an object (subject) existing in the capture environment has spectral reflectance characteristics, where the intensity of the reflected light reflected from the surface of the object varies according to the wavelength of the illumination light. For example, the spectral reflectance characteristics are different between the colors of the surface of the subject.
[0064] Generally, for R, G, and B, the spectral reflectance characteristics are usually represented by three parameters. However, in order to accurately represent the spectral reflectance characteristics, 31 parameters are required for wavelengths obtained by sampling at 10 nm intervals in the range from 400 nm to 700 nm.
[0065] Next, reference will be made to Figure 5 Describe a technique for compressing high-dimensional parameters for 31 wavelengths into N latent variables using principal component analysis.
[0066] The color of the surface of an object is represented by, for example, 1569 color colors in the Munsell color system as a global standard. Therefore, there are also 1569 types of spectral reflectance characteristics.
[0067] First, the 1569 types of spectral reflectance characteristics are subjected to principal component analysis to obtain 1569 principal components, as shown by the arrow #1 in Figure 5 For 31 wavelengths, the principal components are represented by high-dimensional parameters.
[0068] In addition, the 1569 types of spectral reflectance characteristics are subjected to principal component analysis to obtain 1569 coefficients, which can recover the high-dimensional parameters of a specific color when the 1569 principal components are multiplied by the 1569 coefficients respectively, as shown by the arrow #2 in Figure 5 A coefficient group consisting of 1569 coefficients is obtained for each color, thereby generating 1569 types of coefficient groups.
[0069] Next, as shown by the arrow #3 in Figure 5 , N principal components with high contribution rates are extracted from the 1569 principal components. In addition, as shown by the arrow #4 in Figure 15 , N coefficients to be multiplied by the N principal components are extracted from the 1569 coefficients.
[0070] By multiplying the N principal components by the corresponding coefficients, parameters close to the high-dimensional parameters of a specific color can be recovered. The 1569 types of coefficient groups each consisting of N coefficients are regarded as composite coefficients for reproducing the 1569 types of spectral reflectance characteristics. In this technique, the N coefficients are regarded as latent variables obtained by compressing the high-dimensional parameters of 31 wavelengths. The high-dimensional parameters of 31 wavelengths can be replaced by N latent variables in the above manner.
[0071] Figure 6 Shows the process of recovering the spectral reflectance characteristics from N latent variables.
[0072] First, the learning model of the latent variable estimation unit 11 outputs a group of latent variables (N coefficients) that match the color of the object captured in each pixel of the captured image from the 1569 types of latent variable groups. Next, as shown in Figure 6As shown, the restoration unit 12 can obtain high-dimensional parameters indicating the spectral reflectance characteristics of the surface of the object 31 by combining the N principal components extracted by principal component analysis using the N latent variables output from the learning model.
[0073] When directly estimating the parameters of R, G, and B using the learning model, three parameters are obtained as a result of estimating the spectral reflectance characteristics. On the other hand, in the present technology, when N is 3, three latent variables are output from the learning model, but since the spectral reflectance characteristics are estimated, 31 parameters are obtained. Therefore, by using the information processing device 1 according to the present technology, compared with directly estimating the parameters of R, G, and B using the learning model, it is possible to capture an image and a CG image using a parameter combination that represents (i.e., more accurately and in more detail) the spectral reflectance characteristics.
[0074] For example, it is conceivable to preset the number N of latent variables to a value of 3 or 4 in order to ensure the accuracy of estimating the reflectance distribution characteristics.
[0075] - Reflectance distribution characteristics
[0076] An object existing in the capture environment has a reflectance distribution characteristic in which the intensity of the reflected light reflected from the surface of the object varies according to the incident direction and the reflection direction of the light. For example, between objects, the reflectance distribution characteristics are different.
[0077] For example, each of the incident direction and the reflection direction of the light is represented by an azimuth angle obtained by sampling in the range from 0° to 360° at intervals of 10° and a zenith angle obtained by sampling in the range from 0° to 90° at intervals of 10°. In order to accurately represent the reflectance distribution characteristics, 104976 parameters are required for the combination of the incident direction and the reflection direction of the light.
[0078] Figure 7 An example of a method for representing the incident direction and the reflection direction of the light is shown. In Figure 7 , the vector n is the normal vector perpendicular to the surface of the object, and the vector t is the tangent vector. In addition, the vector l is the vector representing the incident direction of the light, and the vector v is the vector representing the reflection direction (the direction of the viewpoint of the camera).
[0079] In Figure 7 example A, the vector l is represented by the azimuth angle and the zenith angle θ l , and the vector v is represented by the azimuth angle and the zenith angle θ v .
[0080] In Figure 7 example B, assuming azimuthal isotropy, the vector v and the vector l are represented by the zenith angle of the vector h equidistant from the vector l and the vector v and the rotation angle from vector l to vector v is represented, where vector h is defined as the rotation axis. With this representation method, the incident direction and reflection direction of light can be represented two-dimensionally. Therefore, in order to represent the reflectance distribution characteristics, for the zenith angles obtained by sampling the range from 0° to 90° at intervals of 10° and the rotation angles
[0081] obtained by sampling the range from 0° to 90° at intervals of 10°, 324 parameters are required.
[0082] In this way, the number of parameters required to represent the reflectance distribution characteristics can be reduced from 104976 to 324, for example, by changing the geometric transformation of the method for representing the incident direction and reflection direction of light. Figure 8 and Figure 9 The following will describe the technique of further compressing 324 high-dimensional parameters into (N + 2) latent variables using principal component analysis, which uses the second variable BRDF (bidirectional reflectance distribution function) model obtained by modeling the reflectance characteristics of an object.
[0083] In this technique, for example, the reflectance distribution characteristics of a certain object are represented in the form of a two-dimensional graph that represents the BRDF calculated for the combination of the zenith angle and the rotation angle A two-dimensional mapping of the reflectance distribution characteristics with significant angular dependence is shown in A of
[0084] and a two-dimensional mapping of the reflectance distribution characteristics without angular dependence is shown in B of Figure 8 . In the two-dimensional mapping of Figure 8 , the magnitude of the BRDF is represented by the shade of color. The BRDF varies according to the combination of the zenith angle Figure 8 and the rotation angle Figure 8 in the two-dimensional image of A of and the BRDF is constant regardless of the combination of the zenith angle and the rotation angle Figure 8 in the two-dimensional image of B of . The reflectance distribution characteristics can be represented in the form of a two-dimensional graph that indicates the BRDF matching the combination of the azimuth angle
[0085] indicating the incident direction of light and the zenith angle θ l and the azimuth angle indicating the reflection direction of light v and the zenith angle θ.
[0086] In order to compress the above-mentioned two-dimensional map (324 high-dimensional parameters) to (N+2) latent variables, for example, 1000 types of two-dimensional maps indicating reflectance distribution characteristics of 1000 types of objects are prepared.
[0087] First, if Figure 9 As shown by arrow #21, 1000 types of two-dimensional mappings are subjected to principal component analysis to obtain 1000 principal components. In addition, 1000 types of two-dimensional mappings are subjected to principal component analysis to obtain 1000 coefficients, such as Figure 9 As shown by arrow #22 in FIG. 1 , when 1000 principal components are multiplied by 1000 coefficients, the coefficients can restore a two-dimensional mapping of a certain object. A coefficient group consisting of 1000 coefficients is obtained for each type of object, thereby generating 1000 types of coefficient groups.
[0088] Next, if Figure 9 As shown by arrow #23 in , N principal components with high contribution rates are extracted from 1000 principal components. Figure 9 As shown by arrow #24 in , N coefficients to be multiplied by the N principal components are extracted from the 1000 coefficients.
[0089] By multiplying the N principal components by the corresponding coefficients, a two-dimensional map close to the two-dimensional map for a specific object can be restored. Thus, 1000 types of coefficient groups each consisting of N coefficients are regarded as synthetic coefficients for reproducing 1000 types of two-dimensional maps. In the present technology, the N coefficients are regarded as latent variables obtained by compressing 324 high-dimensional parameters. In addition, the zenith angle required to extract the BRDF from the two-dimensional map is and rotation angle Also considered as latent variables. The 324 high-dimensional parameters can be replaced by (N+2) latent variables in the above manner.
[0090] Figure 10 The process of recovering a two-dimensional mapping from (N+2) latent variables is shown.
[0091] First, the learning model of the latent variable estimation section 11 outputs a latent variable group (N coefficients) that matches the object captured in each pixel of the captured image among 1000 types of latent variable groups. In addition, the learning model of the latent variable estimation section 11 estimates the incident direction and the reflected direction of the light of the object captured in each pixel of the captured image, and outputs the zenith angle and rotation angle As an estimated result.
[0092] Next, the restoration unit 12 can obtain a two-dimensional map represented by 324 high-dimensional parameters by combining the N principal components extracted by principal component analysis using the N coefficients output from the learning model, as Figure 10 shown by arrow #31 in
[0093] Next, the restoration unit 12 can extract the BRDF (reflectance) from the two-dimensional map using the zenith angle and the rotation angle output from the learning model, as Figure 10 shown by arrow #32 in
[0094] For example, when N is 2, the information processing apparatus 1 according to the present technology can estimate the reflectance distribution characteristics of the surface of an object whose large number of parameters cannot be routinely estimated, and the reflectance matching the incident direction and the reflection direction of the principal ray, by simply estimating four latent variables.
[0095] For example, in order to ensure the accuracy of estimating the reflectance distribution characteristics, it can be considered to set in advance the number N of latent variables other than the zenith angle and the rotation angle to a value of 100.
[0096] - Coordinates in three-dimensional space
[0097] The shape of an object existing in the capture environment is represented by coordinates (X, Y, Z) in the three-dimensional space of the surface of the object, for example. It is not easy to estimate the coordinates in the three-dimensional space of the surface of the object captured in each pixel of a two-dimensional RGB image (capture image) based on the RGB image. For example, when the RGB image has an image size of 1920×1080 and the coordinates in the three-dimensional space of the surface of the object captured in each pixel are to be estimated, the number of parameters to be estimated is 1920×1080×3 = 6220800.
[0098] Hereinafter, a technique of compressing high-dimensional parameters into latent variables using principal component analysis will be described with reference to Figure 11 and Figure 12 First, as
[0099] shown, the coordinates in the three-dimensional space of the surface of the object captured in each of the 1920×1080 pixels are divided into units of (32×18) pixels × 3 coordinates, generating 3600 units. The number of high-dimensional parameters after division is (32×18×3×3600). Figure 11 Next, the 3600 units are subjected to principal component analysis to obtain 3600 principal components, as
[0100] shown in Figure 12as indicated by arrow #31 in. In addition, 3600 units are subjected to principal component analysis to obtain 3600 coefficients, as Figure 12 indicated by arrow #32 in. When the 3600 principal components are multiplied by the 3600 coefficients respectively, the 3600 coefficients can restore a certain unit. A coefficient group composed of 3600 coefficients is obtained for the unit to be restored, and 3600 types of coefficient groups are generated.
[0101] Next, as Figure 12 indicated by arrow #33 in, N principal components with high contribution rates are extracted from the 3600 principal components. In addition, as Figure 12 indicated by arrow #34 in, N coefficients multiplied by the N principal components are extracted from the 3600 coefficients.
[0102] By multiplying the N principal components by the corresponding coefficients, a unit close to a specific unit can be restored. Thus, each coefficient group composed of N coefficients is regarded as a synthesis coefficient for reproducing 3600 types of units. To restore the coordinates in the three-dimensional space close to the surface of the object captured in each of 1920×1080 pixels, all 3600 units need to be restored, so 3600×N synthesis coefficients are required.
[0103] In this technology, (3600×N) coefficients are regarded as latent variables obtained by compressing 6,220,800 high-dimensional parameters. In an RGB image, the N principal components are different, so the N principal components are also regarded as latent variables. Each of the N principal components is composed of (32×18×3) parameters (coordinates).
[0104] Therefore, the number of parameters as latent variables is N×(3600 + 1728)=(5328×N). For example, when N is 100, 6,220,800 high-dimensional parameters can be replaced by 532,800 latent variables, which reduces the number of parameters to 8.5%.
[0105] Figure 13 shows the process of restoring 3600 units from (5328×N) latent variables.
[0106] First, the learning model of the latent variable estimation unit 11 estimates N principal components and 3600 types of coefficient groups (each coefficient group is composed of N coefficients) based on the captured image, and outputs the principal components and the coefficient groups.
[0107] Next, as Figure 13 shown, the restoration unit 12 can obtain 3600 units, each unit composed of 32×18×3 parameters, by using the N coefficients output from the learning model for 3600 types to combine with the N principal components.
[0108] For example, in order to ensure the accuracy of estimating the coordinates of the surface of a subject in three-dimensional space, a quantity N related to the number of latent variables is considered to be set in advance to the value 100.
[0109] - Spectral radiance characteristics
[0110] The light sources in the capture environment have spectral radiance characteristics, where the intensity of the radiated light varies according to the wavelength. The spectral radiance characteristics are different between light sources, for example.
[0111] In order to accurately represent the spectral reflectance characteristics, 31 parameters are required for wavelengths obtained by sampling the range from 400 nm to 700 nm at intervals of 10 nm. A technique for compressing the high-dimensional parameters of 31 wavelengths into a single latent variable using principal component analysis will be described below.
[0112] For example, when shooting is performed outdoors, the sun is used as a light source. The spectral radiance characteristics of sunlight are represented by, for example, standard light determined by the CIE (International Commission on Illumination) as a global standard. The standard light is obtained by performing principal component analysis on a database of the spectral radiance characteristics of sunlight measured under different latitudes, longitudes, and times. By multiplying two principal components with high contribution rates in the principal components obtained by performing principal component analysis by corresponding coefficients, characteristics close to the spectral radiance characteristics of sunlight measured under a specific condition can be restored.
[0113] The two coefficients for multiplying the two principal components are obtained based on, for example, a single parameter indicating the color temperature. Thus, the spectral reflectance characteristics of sunlight can be restored from a single parameter indicating the color temperature. In the present technique, the single parameter indicating the color temperature is regarded as a latent variable obtained by compressing the high-dimensional parameters of 31 wavelengths.
[0114] In order to estimate the spectral radiance characteristics, first, the learning model of the latent variable estimation unit 11 estimates the color temperature of the light source in the capture environment of the captured image based on the captured image, and outputs the estimated color temperature.
[0115] Next, the restoration unit 12 converts the color temperature output from the learning model into two coefficients. Then, the restoration unit 12 can obtain 31 high-dimensional parameters representing the spectral radiance characteristics of the light source by combining the two principal components extracted by principal component analysis using the two coefficients.
[0116] - Conventional techniques for estimating physical properties
[0117] Figure 14 The flow of a conventional technique for estimating physical properties is shown.
[0118] In the conventional technique, first, as Figure 14As shown in #101, physical characteristics of a capture environment with respect to a captured image are estimated based on the captured image. For example, the shape of an object captured in the captured image (distance from the camera, perpendicular to the reflection surface), the influence of a light source on the object (direction, color, and degree of influence of the light source), and reflection information about the object (color and unevenness of the surface) can be estimated as physical characteristics of the capture environment.
[0119] For example, a learning model can be used to estimate the physical characteristics of the capture environment, and the captured image is input to the learning model and multiple physical characteristics of the capture environment are output.
[0120] Next, as Figure 14 shown in #102, the captured image and a CG image obtained by processing CG data are combined using the estimated physical characteristics to generate a synthetic image.
[0121] Figure 15 An example of learning data for training a learning model for estimating physical characteristics of a capture environment is shown.
[0122] For example, student data used for learning and teacher data used as a teacher (expected value) for learning can be used to train the learning model for estimating physical characteristics of the capture environment in Figure 14 #101.
[0123] As Figure 15 shown, the student data can be, for example, an RGB image generated using CG.
[0124] Meanwhile, the teacher data can be, for example, the color of the surface of an object captured in the RGB image, the normal of the reflection surface, the unevenness of the surface, the distance from the camera, and the effect of the light source as shown in Figure 15 #.
[0125] In the conventional technique, as described above, a learning model obtained by performing machine learning using learning data (such as those described above) directly estimates the physical characteristics of the capture environment with respect to the captured image.
[0126] However, when using CG to obtain learning data, there are countless CG parameters corresponding to the physical characteristics of the capture environment, and such parameters affect the RGB image while being associated with each other, and thus it is difficult to learn the physical characteristics individually. Therefore, the physical characteristics to be learned are limited to the shape of the object, the effect of the light source, reflectance characteristics, etc.
[0127] The physical characteristics of an object and a light source in a capture environment in the real world (such as spectral reflectance characteristics and reflectance distribution characteristics) cannot be correctly estimated by simply estimating a limited type of physical characteristics.
[0128] On the other hand, when obtaining student data for learning by capturing an actual scene, it is desirable to use measurement data obtained by sampling the physical properties of an object and a light source with sufficient accuracy as teacher data in order to improve the accuracy of estimating the physical properties using a learning model. However, a large amount of data is required as learning data, and thus a large amount of measurement data needs to be prepared.
[0129] Accordingly, with conventional techniques, it may not be possible to accurately estimate the physical properties of the capture environment regarding a captured image using a learning model.
[0130] Figure 16 An example of the synthesis of a captured image and a CG image according to a conventional technique is shown.
[0131] In Figure 16 the example, a captured image P11 that captures a part of a room and a CG image P12 obtained by processing CG data of a rabbit are combined to generate a composite image P13 that looks as if the rabbit is arranged on the floor of the captured room in the captured image P11.
[0132] When the physical properties of the capture environment regarding a captured image cannot be accurately estimated, the captured image P11 and the CG data may not be carefully (i.e., in sufficient detail or with sufficient accuracy) processed, and a natural-looking composite image P13 may not be obtained.
[0133] When the rabbit is represented in gray in the Figure 16 composite image P13, the rabbit looks unnatural.
[0134] With the present technique, a latent variable regarding the capture environment of a captured image is estimated using a learning model. The captured image is input to the learning model and outputs a latent variable. The latent variable is obtained by compressing a high-dimensional parameter to a lower dimension compared to the high-dimensional parameter, where the high-dimensional parameter represents the physical properties of the capture environment regarding the captured image, and the high-dimensional parameter is recovered from the latent variable.
[0135] A high-dimensional parameter that accurately represents the physical properties of the capture environment can be obtained based on the captured image, and thus the information processing device 1 can carefully combine the captured image and the CG image.
[0136] <3. Operation of the Information Processing Device>
[0137] Reference will be made to Figure 17 the flowchart in to describe the processing performed by the information processing device 1 configured as described above.
[0138] In step S1, the latent variable estimation unit 11 estimates a latent variable regarding the capture environment of a captured image using a learning model. The captured image is input to the learning model and outputs a latent variable.
[0139] In step S2, the restoration unit 12 restores the high-dimensional parameters from the latent variables estimated in step S1.
[0140] In step S3, the synthesis unit 13 generates a composite image by combining the captured image and the CG image using the high-dimensional parameters.
[0141] Through the above processing, high-dimensional parameters that accurately represent the physical characteristics of the objects and light sources in the capture environment of the captured image can be estimated. In addition, the synthesis unit 13 can combine the captured image and the CG image using a fine calculation formula based on the high-dimensional parameters, making the resulting image look natural.
[0142] <4. Modification Example>
[0143] - Examples of scene recognition or segmentation
[0144] The accuracy of the high-dimensional parameters can be improved by performing scene recognition or segmentation.
[0145] Figure 18 The flow of a technique for estimating high-dimensional parameters when performing scene recognition is shown.
[0146] First, as Figure 18 shown in #151, the capture scene (situation) of the captured image is recognized based on the captured image. For example, it is recognized whether the captured image is captured outdoors or indoors. For example, scene recognition is performed by the latent variable estimation unit 11 of the information processing device 1 ( Figure 1 ).
[0147] Next, as Figure 18 shown in #152, latent variables regarding the capture environment of the captured image are estimated based on the captured image and the result of scene recognition. Specifically, the captured image and the result of scene recognition are input to the learning model, and latent variables are output from the learning model. When estimating latent variables based on the result of scene recognition, an RGB image and information indicating the scene of the captured RGB image are used as student data when training the learning model of the latent variable estimation unit 11.
[0148] Next, as Figure 18 shown in #153, high-dimensional parameters are restored from the latent variables based on the result of scene recognition. For example, in the case where the color temperature of the light source is estimated as a latent variable, the restoration unit 12 can switch the types of light sources such as sunlight, halogen lamps, and LED lamps according to whether the captured image is taken outdoors or indoors, thereby restoring the high-dimensional parameters. By switching the type of light source, the accuracy of estimating the high-dimensional parameters used to combine the captured image and the CG image can be improved.
[0149] Next, as Figure 18As shown in #154, a high-dimensional parameter combination recovered from a latent variable is used to capture an image and a CG image obtained by processing CG data to generate a synthetic image.
[0150] Segmentation may be performed on the captured image instead of scene recognition. Specifically, the captured image and the segmentation result are input to a learning model, and a latent variable is output from the learning model. When estimating the latent variable based on the segmentation result, the RGB image and the segmentation result of the RGB image are used as student data when training the learning model of the latent variable estimation unit 11.
[0151] Segmentation on the captured image clarifies the boundaries of the objects captured in the captured image. At the boundaries of the subject, the coordinates and reflectance distribution characteristics in the three-dimensional space of the surface of the subject are likely to change sharply. Therefore, the accuracy of the high-dimensional parameters recovered from the latent variable can be improved by estimating the latent variable based on whether each pixel of the captured image corresponds to the boundary of the subject.
[0152] - Example of step-by-step estimation of latent variable
[0153] The accuracy of the high-dimensional parameters can be improved by step-by-step estimation of the latent variable.
[0154] Figure 19 The figure shows the process of a technique for estimating high-dimensional parameters when estimating the latent variable step by step.
[0155] First, as Figure 19 shown in #201, a first latent variable regarding the capture environment of the captured image is estimated based on the captured image, and as shown in #202, a first high-dimensional parameter is recovered from the first latent variable.
[0156] Next, as Figure 19 shown in #203, a second latent variable regarding the capture environment of the captured image is estimated based on the captured image and the first high-dimensional parameter recovered in #202, and as shown in #202, a second high-dimensional parameter is recovered from the second latent variable.
[0157] The second high-dimensional parameter is a parameter representing physical characteristics different from the first high-dimensional parameter. For example, by estimating the first high-dimensional parameter representing the shape of the subject captured in the captured image, and then estimating the second high-dimensional parameter representing the reflectance characteristics of the surface of the subject based on the shape of the subject, the estimation accuracy of the reflectance characteristics can be improved.
[0158] The second high-dimensional parameter may be a parameter that represents the same physical characteristics in detail as the first high-dimensional parameter. For example, certain physical characteristics are roughly estimated by estimating the first latent variable and recovering the first high-dimensional parameter, and these physical characteristics are estimated in detail by estimating the second latent variable and recovering the second high-dimensional parameter.
[0159] Next, as shown in #205 of Figure 19 , at least one combination of the first high-dimensional parameter and the second high-dimensional parameter is used to capture an image and a CG image obtained by processing CG data to generate a composite image.
[0160] - Computer
[0161] The series of processes discussed above can be executed by hardware and can also be executed by software. When the series of processes are executed by software, the program constituting the software is installed from a program storage medium (e.g., a non-volatile computer-readable storage medium) in a computer incorporated with dedicated hardware, a general-purpose personal computer, etc.
[0162] Figure 20 is a block diagram showing an example of the hardware configuration of a computer that executes the above series of processes using a program.
[0163] A CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are connected to each other via a bus 504.
[0164] An input / output interface 505 is further connected to the bus 504. An input unit 506 composed of a keyboard, a mouse, etc. and an output unit 507 composed of a display, a speaker, etc. are connected to the input / output interface 505. In addition, a storage unit 508 composed of a hard disk, a non-volatile memory, etc., a communication unit 509 composed of a network interface, etc., and a drive 510 that drives a removable medium 511 are connected to the input / output interface 505.
[0165] In the computer configured as described above, for example, the program stored in the storage unit 508 is loaded into the RAM 503 via the input / output interface 505 and the bus 504 by the CPU 501, and the program is executed to perform the above series of processes.
[0166] For example, the program executed by the CPU 501 is installed in the storage unit 508 by being stored in the removable medium 511 or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting.
[0167] The program executed by the computer can be a program in which the processes are executed in chronological order as described herein, or can be a program in which the processes are executed simultaneously or at a necessary timing such as when a call is made.
[0168] The effects described herein are merely exemplary and not restrictive, and other effects can be achieved.
[0169] Embodiments of the present technology are not limited to the above-described embodiments, and various changes can be made without departing from the scope and spirit of the present technology.
[0170] For example, the present technology may take the form of cloud computing, in which a single function is processed in a distributed and shared manner by multiple devices via a network.
[0171] Each step described with reference to the above flowchart can be executed by multiple devices in a distributed manner in addition to being executed by a single device.
[0172] In addition, when multiple processes are included in a single step, the multiple processes included in the single step can also be executed by multiple devices in a distributed manner in addition to being executed by a single device.
[0173] <Configuration combination example>
[0174] The present technology can adopt the following configurations. (1)
[0176] An information processing apparatus, comprising:
[0177] An estimation unit that estimates a latent variable regarding a capture environment of a captured image using a learning model, the captured image is input to the learning model and outputs the latent variable, the latent variable is obtained by compressing a high-dimensional parameter to a lower dimension compared to the parameter, and the parameter indicates a physical characteristic regarding the capture environment of the captured image; and
[0178] A recovery unit that recovers a parameter from the latent variable estimated by the estimation unit. (2)
[0180] The information processing apparatus according to (1), further comprising:
[0181] A synthesis unit that combines the captured image and a CG image using the parameter recovered by the recovery unit. (3)
[0183] The information processing apparatus according to (1) or (2), wherein:
[0184] The latent variable includes coefficients for combining a predetermined number of principal components obtained by performing principal component analysis on the physical characteristics; and
[0185] The recovery unit recovers the parameter by multiplying the principal components by the corresponding coefficients. (4)
[0187] The information processing apparatus according to any one of (1) to (3), wherein
[0188] The physical characteristics are compressed into the latent variable by a geometric transformation. (5)
[0190] The information processing apparatus according to (3), wherein,
[0191] The physical characteristics include the spectral reflectance characteristics of the surface of the object captured in the captured image. (6)
[0193] The information processing apparatus according to (5), wherein,
[0194] The latent variables include coefficients for combining the principal components obtained by principal component analysis with the spectral reflectance characteristics of each color of the object. (7)
[0196] The information processing apparatus according to any one of (3) to (6), wherein,
[0197] The physical characteristics include the reflectance distribution characteristics of the object captured in the captured image. (8)
[0199] The information processing apparatus according to (7), wherein,
[0200] The latent variables include coefficients for combining the principal components obtained by principal component analysis of the reflectance distribution characteristics of each object, and information indicating the incident direction and reflection direction of the light of the object for extracting the reflectance from the reflectance distribution characteristics. (9)
[0202] The information processing apparatus according to any one of (3) to (8), wherein,
[0203] The physical characteristics include the shape of the surface of the object captured in the captured image. (10)
[0205] The information processing apparatus according to (9), wherein,
[0206] The latent variables include the principal components obtained by principal component analysis on the coordinates indicating the shape of the surface of the object and coefficients for combining the principal components of each unit obtained by dividing the captured image. (11)
[0208] The information processing apparatus according to any one of (3) to (10), wherein,
[0209] The physical characteristics include the spectral radiation characteristics of the light source in the capture environment. (12)
[0211] The information processing apparatus according to (11), wherein,
[0212] The latent variables include the color temperature of the light source. (13)
[0214] An information processing apparatus according to any one of (1) to (12), wherein,
[0215] An estimation unit performs scene recognition on a captured image of a scene and estimates a latent variable by inputting a result of the scene recognition into a learning model. (14)
[0217] An information processing apparatus according to any one of (1) to (13), wherein,
[0218] An estimation unit performs segmentation on a captured image and estimates a latent variable by inputting a result of the segmentation into a learning model. (15)
[0220] An information processing apparatus according to any one of (1) to (14), wherein:
[0221] An estimation unit estimates a second latent variable of a latent variable obtained by compressing a second parameter in compressed parameters by inputting a first parameter of parameters restored from a first latent variable of the latent variable by a restoration unit into a learning model; and
[0222] A restoration unit restores the second parameter from the second latent variable estimated by the estimation unit. (16)
[0224] An information processing apparatus according to (15), wherein,
[0225] The second parameter indicates a physical property different from the first parameter. (17)
[0227] An information processing apparatus according to (15), wherein,
[0228] The second parameter indicates a physical property the same as the physical property of the first parameter in more detail than the first parameter. (18)
[0230] An information processing method, comprising:
[0231] Using an information processing apparatus:
[0232] Using a learning model to estimate a latent variable regarding a capture environment of a captured image, the captured image being input into the learning model and the learning model outputting the latent variable, the latent variable being obtained by compressing a high-dimensional parameter to a lower dimension compared with the parameter, the parameter indicating a physical property regarding the capture environment of the captured image; and
[0233] Restoring a parameter from the estimated latent variable. (19)
[0235] A computer-readable storage medium stores a program for executing a process, the process including:
[0236] Estimating a latent variable regarding a capture environment of a captured image using a learning model, the captured image being input to the learning model and the latent variable being output, the latent variable being obtained by compressing a high-dimensional parameter to a lower dimension compared to the parameter, the parameter indicating a physical property of the capture environment regarding the captured image; and
[0237] Recovering the parameter from the estimated latent variable. (20)
[0239] An information processing apparatus includes:
[0240] A circuit configured to:
[0241] Estimate a latent variable regarding a capture environment of a captured image using a learning model to which the captured image is input, and
[0242] Output the estimated latent variable, the latent variable being obtained by compressing a high-dimensional parameter to a lower dimension, the high-dimensional parameter indicating a physical property of the capture environment of the captured image. (21)
[0244] The information processing apparatus according to (20), wherein the circuit is further configured to:
[0245] Reconstruct a high-dimensional parameter from the estimated latent variable for the captured image; and
[0246] Combine the captured image and a computer-generated image using the reconstructed high-dimensional parameter. (22)
[0248] The information processing apparatus according to (20) or (21), wherein the latent variable includes coefficients for combining a predetermined number of principal components obtained by principal component analysis of the physical property, and the circuit is further configured to:
[0249] Reconstruct the high-dimensional parameter by multiplying the principal components by the corresponding coefficients. (23)
[0251] The information processing apparatus according to any one of (20) to (22), wherein the high-dimensional parameter is compressed into the latent variable by a geometric transformation. (24)
[0253] The information processing apparatus according to (22), wherein the physical property includes spectral reflectance characteristics of a surface of an object captured in the captured image. (25)
[0255] The information processing apparatus according to (24), wherein
[0256] The latent variables include coefficients for combining the principal components obtained by principal component analysis with the spectral reflectance characteristics of each color of the object. (26)
[0258] The information processing apparatus according to (22), wherein
[0259] The physical characteristics include the reflectance distribution characteristics of the object captured in the captured image. (27)
[0261] The information processing apparatus according to (26), wherein
[0262] The latent variables include coefficients for combining the principal components obtained by principal component analysis of the reflectance distribution characteristics of each object, and information indicating the incident direction and reflection direction of light of the object for extracting the reflectance from the reflectance distribution characteristics. (28)
[0264] The information processing apparatus according to (22), wherein the physical characteristics include the shape of the surface of the object captured in the captured image. (29)
[0266] The information processing apparatus according to (28), wherein the latent variables include the principal components obtained by principal component analysis on the coordinates indicating the shape of the surface of the object and coefficients for combining the principal components of each unit obtained by dividing the captured image. (30)
[0268] The information processing apparatus according to claim (22), wherein the physical characteristics include the spectral radiation characteristics of the light source in the capture environment. (31)
[0270] The information processing apparatus according to (30), wherein the latent variable includes the color temperature of the light source. (32)
[0272] The information processing apparatus according to any one of (20) to (31), wherein the circuit is further configured to:
[0273] Perform scene recognition in which the captured image is captured, and
[0274] Estimate the latent variables by inputting the result of the scene recognition into a learning model. (33)
[0276] The information processing apparatus according to any one of (20) to (32), wherein the circuit is further configured to:
[0277] Perform segmentation on the captured image, and
[0278] Estimate a latent variable by inputting the result of segmentation into a learning model. (34)
[0280] The information processing apparatus according to any one of (20) to (33), wherein the circuit is further configured to:
[0281] Reconstruct a first parameter in the high-dimensional parameters from a first estimated latent variable among the estimated latent variables,
[0282] Estimate a second latent variable among the estimated latent variables, the second latent variable being obtained by inputting a first parameter among the parameters restored from a first latent variable among the latent variables into a learning model to compress a second parameter in the high-dimensional parameters; and
[0283] Reconstruct a second parameter from the second latent variable. (35)
[0285] The information processing apparatus according to (34), wherein the second parameter indicates a physical characteristic different from that of the first parameter. (36)
[0287] The information processing apparatus according to (34), wherein
[0288] The second parameter indicates the same physical characteristic as the first parameter in more detail than the first parameter. (37)
[0290] An information processing method, comprising:
[0291] Estimate a latent variable of a capture environment of a captured image using a learning model input with the captured image, and
[0292] Output a latent variable obtained by compressing high-dimensional parameters to a lower dimension compared to the parameters, the high-dimensional parameters indicating physical characteristics of a capture environment of the captured image. (38)
[0294] A non-transitory computer-readable storage medium storing a program for executing a process, the process comprising:
[0295] Estimate a latent variable of a capture environment of a captured image using a learning model input with the captured image, and
[0296] Output a latent variable obtained by compressing high-dimensional parameters to a lower dimension, the high-dimensional parameters representing physical characteristics of a capture environment of the captured image. (39)
[0298] A computer-implemented method for generating a training set for training a machine for estimating a physical property in a captured image, the method comprising:
[0299] Collecting a set of images having known physical characteristics;
[0300] Associating the physical property of each image in the set of images with a latent variable having a lower dimension than the high-dimensional parameter of the physical property; and
[0301] Creating a training set including the set of images and the set of latent variables for each image.
[0302] List of reference numerals
[0303] 1 Information processing device
[0304] 11 Latent variable estimation unit
[0305] 12 Restoration unit
[0306] 13 Synthesis unit
Claims
1. An information processing apparatus, comprising: a circuit configured to: estimate a latent variable of a capture environment of the captured image using a learning model to which the captured image is input, and output the estimated latent variable, the latent variable being obtained by compressing a high-dimensional parameter to a lower dimension, the high-dimensional parameter indicating a physical characteristic of the capture environment of the captured image.
2. The information processing apparatus according to claim 1, wherein, The circuit is further configured to: reconstruct the high-dimensional parameter from the estimated latent variable for the captured image; and combine the captured image and a computer-generated image using the reconstructed high-dimensional parameter.
3. The information processing apparatus according to claim 1, wherein, The latent variable includes coefficients for combining a predetermined number of principal components obtained by principal component analysis of the physical characteristic, and the circuit is further configured to: reconstruct the high-dimensional parameter by multiplying the principal components by the corresponding coefficients.
4. The information processing apparatus according to claim 1, wherein the high-dimensional parameter is compressed into the latent variable by a geometric transformation.
5. The information processing apparatus according to claim 3, wherein the physical characteristic includes a spectral reflectance characteristic of a surface of an object captured in the captured image.
6. The information processing apparatus according to claim 5, wherein the latent variable includes coefficients for combining the principal components obtained by the principal component analysis with the spectral reflectance characteristics of each color of the object.
7. The information processing apparatus according to claim 3, wherein the physical characteristic includes a reflectance distribution characteristic of an object captured in the captured image.
8. The information processing apparatus according to claim 7, wherein the latent variable includes coefficients for combining the principal components obtained by the principal component analysis of the reflectance distribution characteristic of each object, and information indicating an incident direction and a reflection direction of light of the object for extracting a reflectance from the reflectance distribution characteristic.
9. The information processing apparatus according to claim 3, wherein the physical characteristic includes a shape of a surface of an object captured in the captured image.
10. The information processing apparatus according to claim 9, wherein the latent variable includes principal components obtained by the principal component analysis on coordinates indicating the shape of the surface of the object and coefficients for combining the principal components of each unit obtained by dividing the captured image.
11. The information processing apparatus according to claim 3, wherein the physical characteristic includes a spectral radiation characteristic of a light source in the capture environment.
12. The information processing apparatus according to claim 11, wherein the latent variable includes a color temperature of the light source.
13. The information processing apparatus according to claim 1, wherein the circuit is further configured to: perform scene recognition in which the captured image is captured, and estimate the latent variable by inputting a result of the scene recognition to the learning model.
14. The information processing apparatus according to claim 1, wherein the circuit is further configured to: perform segmentation on the captured image, and Estimate the latent variable by inputting the result of segmentation into the learning model.
15. The information processing apparatus according to claim 1, wherein the circuit is further configured to: reconstruct a first parameter among the high-dimensional parameters from a first estimated latent variable among the estimated latent variables; estimate a second latent variable among the estimated latent variables, the second latent variable being obtained by inputting the first parameter among the parameters restored from the first latent variable among the latent variables into the learning model to compress a second parameter among the high-dimensional parameters; and reconstruct the second parameter from the second latent variable.
16. The information processing apparatus according to claim 15, wherein, The second parameter indicates a physical characteristic different from that of the first parameter.
17. The information processing apparatus according to claim 15, wherein the second parameter indicates the same physical characteristic as the first parameter in more detail than the first parameter.
18. An information processing method, comprising: estimating a latent variable of a capture environment of a captured image using a learning model into which the captured image is input; and outputting the latent variable, the latent variable being obtained by compressing a high-dimensional parameter to a lower dimension compared to the parameter, the high-dimensional parameter indicating a physical characteristic of the capture environment of the captured image.
19. A non-transitory computer-readable storage medium storing a program for executing a process, the process comprising: estimating a latent variable of a capture environment of a captured image using a learning model into which the captured image is input; and outputting the latent variable, the latent variable being obtained by compressing a high-dimensional parameter to a lower dimension, the high-dimensional parameter indicating a physical characteristic of the capture environment of the captured image.
20. A computer-implemented method for generating a training set for training a machine for estimating a physical property in a captured image, the method comprising: collecting a set of images having known physical characteristics; associating a physical property of each image in the set of images with a latent variable having a lower dimension than a high-dimensional parameter of the physical property; and creating a training set including the set of images and the set of latent variables of each image.