Optical estimation method and apparatus
The method and apparatus enhance light estimation for augmented reality and computer graphics by using a model-based approach with self-supervised learning and known object databases to accurately render virtual objects, addressing the challenges of data requirements and shadow analysis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2022-06-22
- Publication Date
- 2026-07-22
AI Technical Summary
Existing light estimation methods struggle to accurately estimate light information for rendering virtual objects in augmented reality and computer graphics, requiring extensive training data and time, and face challenges in precise shadow analysis.
A method and apparatus for light estimation that includes estimating light information using a model, detecting reference objects and planes, rendering virtual objects, and updating the model based on comparison results, utilizing self-supervised learning and known object databases for improved accuracy.
Enhances the realism of virtual objects by accurately estimating light information, reducing the need for extensive training data and improving shadow rendering precision.
Smart Images

Figure 0007893547000001 
Figure 0007893547000002 
Figure 0007893547000003
Abstract
Description
Technical Field
[0001] The following embodiments relate to a light estimation method and apparatus.
Background Art
[0002] Light estimation is a technique for estimating the light illuminating a scene. The estimated light information is used to render virtual objects in the video or space. For example, the estimated light information may be applied to virtual objects such as augmented reality (AR) and computer graphics (CG). The more accurately the light information is estimated, the higher the realism of the virtual object. A machine learning-based model may be used for light estimation. The model can be trained using clues such as ambient light, shading, specular highlights, and reflection.
Summary of the Invention
Problems to be Solved by the Invention
[0003] An object of the present invention is to provide a light estimation method and apparatus.
Means for Solving the Problems
[0004] According to one embodiment, a method for light estimation includes the steps of: estimating light information corresponding to an input image using a light estimation model; detecting a reference object in the input image; determining object information of the reference object and planar information of a reference plane supporting the reference object; rendering a virtual object corresponding to the reference object based on the light information, the object information, and the planar information; and updating and training the light estimation model based on the comparison result between the reference object and the virtual object.
[0005] The step of rendering the virtual object may include the step of rendering the shading and shadow of the virtual object. The comparison result may include the result of comparing the pixel data representing the shading and shadow of the reference object with the pixel data representing the shading and shadow of the virtual object.
[0006] The object information includes at least one of the pose, shape, and material of the reference object, and the plane information may include at least one of the pose, shape, and material of the reference plane.
[0007] The step of determining the object information and the plane information may include, if the reference object is a known object, the step of determining at least a portion of the object information using an object database. The step of determining the object information and the plane information may include, if the reference object and the reference plane are a known combined structure, the step of determining at least a portion of the object information and the plane information using an object database. A spherical substructure corresponding to the reference object and a flat substructure corresponding to the reference plane can be combined to form the structure.
[0008] The step of determining the object information and the plane information may include, if the reference plane is an unknown plane, the step of detecting the reference plane from the input video and determining the plane information. The step of determining the object information and the plane information may include, if the reference object is an unknown object, the step of determining the object information based on information of a predetermined proxy object.
[0009] The step of rendering the virtual object may include the steps of fusing light information and visibility information at each sampling point of the reference plane to determine shadow information for each sampling point, and projecting the shadow information of each sampling point of the reference plane onto the captured view of the input image to render the shadow of the virtual object. If the reference object and the reference plane are a known combined structure, the visibility information may be predetermined for the shadow rendering of the combined structure.
[0010] The step of training the light estimation model may include updating the light estimation model so that the difference between the reference object and the virtual object becomes smaller.
[0011] The method for light estimation may further include the steps of acquiring other images and estimating light information corresponding to the other images using the trained light estimation model.
[0012] An apparatus for optical estimation according to one embodiment includes a processor and a memory containing instructions that can be executed by the processor. When the instructions are executed by the processor, the processor estimates optical information corresponding to an input image using an optical estimation model, detects reference objects in the input image, determines object information of the reference objects and planar information of a reference plane supporting the reference objects, renders a virtual object corresponding to the reference objects based on the optical information, object information, and planar information, and updates the optical estimation model based on the comparison result between the reference objects and the virtual objects.
[0013] The processor renders the shading and shadow of the virtual object, and the comparison result may include a comparison result between pixel data representing the shading and shadow of the reference object and pixel data representing the shading and shadow of the virtual object.
[0014] According to one embodiment, the electronic device includes a camera that generates an input image, an optical estimation model that estimates optical information corresponding to the input image, detects a reference object in the input image, determines object information of the reference object and planar information of a reference plane supporting the reference object, renders a virtual object corresponding to the reference object, the shading of the virtual object, and the shadow of the virtual object based on the optical information, the object information, and the planar information, and updates the optical estimation model based on the comparison result between pixel data showing the shading and shadow of the reference object and pixel data showing the shading and shadow of the virtual object.
[0015] According to one embodiment, a method for light estimation includes the steps of acquiring a second image and estimating second light information corresponding to the second image using a trained light estimation model, wherein the light estimation model estimates light information corresponding to an input image using the light estimation model, detects a reference object in the input image, determines object information of the reference object and planar information of a reference plane supporting the reference object, renders a virtual object corresponding to the reference object based on the light information, the object information, and the planar information, and updates and trains the light estimation model based on the results of comparing the reference object and the virtual object.
[0016] The method for light estimation includes the steps of rendering a second virtual object corresponding to a second object in the second image based on the second light information, The method may further include the step of generating augmented reality by superimposing the second virtual object onto the second image.
[0017] A method for light estimation according to one embodiment includes the steps of: estimating light information corresponding to an input image using a light estimation model; determining information of a reference object in the input image based on whether the stored information corresponds to a reference object; rendering a virtual object corresponding to the reference object based on the light information and the information of the reference object; and comparing the reference object with the rendered virtual object, updating the light estimation model, and training the light estimation model.
[0018] The stored information is a predetermined object, and the step of determining the information of the reference object may include determining the stored information as the information of the reference object in response to the reference object corresponding to the predetermined object.
[0019] The step of determining the information of the reference object may include the step of determining the stored information as the information of the reference object in response to the stored information corresponding to the reference object.
[0020] The information of the reference object may include at least a part of the object information of the reference object and the plane information of the reference plane supporting the reference object.
Advantages of the Invention
[0021] According to the present invention, an optical estimation method and apparatus can be provided.
Brief Description of the Drawings
[0022] [Figure 1] Shows the optical estimation operation using the optical estimation model according to an embodiment. [Figure 2] Shows the operation of training the optical estimation model according to an embodiment. <00An electronic device according to one embodiment is shown. [Modes for carrying out the invention]
[0023] The specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and can be modified in various ways. Therefore, the embodiments are not limited to any particular disclosure, and the scope of this specification includes modifications, equivalents, or substitutions that are part of the technical concept.
[0024] Terms such as "first" or "second" may be used to describe multiple components, but such terms should be interpreted solely for the purpose of distinguishing one component from others. For example, the first component may be named the second component, and similarly, the second component may also be named the first component.
[0025] When it is mentioned that one component is “linked” or “connected” to another component, it should be understood that it is directly linked to or connected to the other component, but that other components may be present in between.
[0026] A singular expression includes plural expressions unless the context clearly indicates otherwise. In this specification, terms such as “includes” or “has” indicate the presence of features, figures, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood not to presuppose the existence or addition of one or more other features, figures, steps, actions, components, parts, or combinations thereof.
[0027] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as those generally understood by a person of ordinary skill in the art to which this embodiment belongs. Commonly used, predefined terms should be interpreted as having the meaning consistent with their meaning in the context of the relevant art, and not as ideal or overly formal unless expressly defined herein.
[0028] The embodiments will be described in detail below with reference to the attached drawings. When describing with reference to the drawings, the same components will be given the same reference numerals regardless of the reference numerals used in the drawings, and redundant explanations will be omitted.
[0029] Figure 1 shows the operation of light estimation using a light estimation model according to one embodiment. Referring to Figure 1, the light estimation device 100 receives an input video 101 and outputs light information 102 corresponding to the input video 101. The input video 101 may show a scene from a specific shooting view seen from a specific shooting location. The input video 101 may be a scene image or a front image.
[0030] The light information 102 includes information about all the light that affected the scene in the input video 101. The light information 102 can represent the light in the scene in various forms. For example, the light information 102 may represent the light in the form of an environment map, or it may represent the light using predefined attributes (e.g., direction, color, brightness, width, etc.).
[0031] Light information 102 can be applied to virtual objects such as augmented reality (AR) and computer graphics (CG). For example, when AR is provided by overlaying a virtual object onto an input image 101, applying light information 102 to the virtual object allows the virtual object to be displayed in the input image 101 without any sense of incongruity. The more accurately light information 102 represents the light in the input image 101, the higher the realism of the virtual object.
[0032] The optical estimation device 100 can generate optical information 102 using the optical estimation model 110. The optical estimation model 110 may be a machine learning model. For example, the optical estimation model 110 may include a deep neural network (DNN) based on deep learning.
[0033] A DNN may include multiple layers, at least some of which can consist of various networks such as fully connected networks (FCNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs). A DNN may have the generalization ability to map non-linearly related input and output data to each other based on deep learning.
[0034] Model training requires a large amount of training data. For example, to give a model the ability to derive light information 102 from input video 101, various scene videos and ground truth data representing the light in those scenes are required. Securing such a large amount of training data requires considerable time and effort.
[0035] According to one embodiment, the light estimation model 110 can be trained through self-supervised learning of the video rendering infrastructure. Once the light estimation model 110 generates a light estimation result for a given scene image, video rendering is performed based on the light estimation result, and the light estimation model 110 can be trained in a direction that reduces the difference between the scene image and the rendering result. Such a method does not require the use of ground truth data, and the effort and time required to secure training data can be greatly reduced.
[0036] According to the embodiment, image rendering is performed using objects, planes, and light, and known objects, known planes, shadow information, etc., may be used for image rendering and model training. Such a training method improves the training effect, and the light estimation model 110 can estimate light information 102 that is close to reality.
[0037] Figure 2 shows the operation of training a light estimation model according to one embodiment. Referring to Figure 2, the light estimation model 210 estimates light information 202 corresponding to the input video 201. The light estimation model 210 may be a neural network model including multiple layers. At least some of the layers are CNNs, and at least some of the other layers are FCNs. The light information 202 may represent light in the form of an environment map, or it may represent light using predefined attributes (e.g., direction, color, brightness, width, etc.).
[0038] The rendering model 220 generates a rendering result 205 based on light information 202, object information 203, and plane information 204. For example, the rendering model 220 may perform neural rendering. The object information 203 is information about an object in the input image 201 and may include at least one of the object's pose, shape, and material. The material refers to texture, color, etc. The object in the input image 201 may be referred to as a reference object. The plane information 204 is information about a reference plane supporting the reference object and may include at least one of the reference plane's pose (e.g., normal orientation), shape, and material. The reference object is detected from the input image 201 via an object detection method, and at least a portion of the object information 203 and plane information 204 can be determined based on the object detection result.
[0039] The rendering result 205 may include a virtual object corresponding to the reference object. The rendering model 220 can render the virtual object based on the optical information 202, object information 203, and plane information 204. A comparison result 206 is determined by comparing the reference object and the virtual object, and the optical estimation model 210 is updated based on the comparison result 206. The difference between the reference object and the virtual object corresponds to the training loss, and the optical estimation model 210 can be updated to reduce the loss, in other words, to reduce the difference between the reference object and the virtual object. The parameters of the optical estimation model 210 (e.g., weights) may be updated through training.
[0040] The rendering model 220 renders the virtual object along with its shading and shadow. The comparison result 206 includes a comparison between pixel data showing the shading and shadow of a reference object and pixel data showing the shading and shadow of a virtual object. The rendering model 220 may perform shading rendering and shadow rendering simultaneously or separately, and the shading rendering result and the shadow rendering result can be merged to generate the rendering result 205. According to the embodiment, a known rendering technique may be applied to shading rendering, and the rendering technique according to the embodiments of Figures 7 and 8 may be applied to shadow rendering.
[0041] Shadows are used as clues about light, but shadow analysis is a task that requires precision, making accurate shadow analysis difficult. Even rough analysis of shadows, such as their generation direction analysis, can provide relatively accurate clues about light. The accuracy of shadow rendering is improved through planar information 204. According to the embodiment, a rendering result 205 that is close to reality is generated through shadow rendering, and the model can be updated in a direction closer to reality by comparing the shadows.
[0042] According to the embodiment, information on known objects and / or known planes may be pre-stored in an object database, and if the target object and / or reference plane corresponds to a known object and / or known plane, at least a portion of the object information 203 and / or plane information 204 is determined using the information in the object database. In this way, a rendering result 205 that is close to reality can be derived by utilizing the pre-stored information. For example, if the reference object corresponds to a known object, at least a portion of the object information 203 can be determined using an object database that stores information on known objects. An input image 201 is generated by photographing a known structure in which a specific object and a specific plane are combined. If the reference object and reference plane are such a known combined structure, at least a portion of the object information 203 and plane information 204 can be determined using the object database. If the reference plane is an unknown plane, the plane information 204 may be determined via a plane detection method. If the reference object is an unknown object, the object information 203 can be determined based on information of a pre-determined proxy object (e.g., cube, cylinder, point cloud, etc.).
[0043] Figure 3 shows the comparison operation between an input image and a rendering result according to one embodiment. Referring to Figure 3, an object region 315 containing a reference object 311 is detected from the input image 310. Object detection may be performed by an object detection model. The object detection model may be a machine learning model (e.g., a neural network model) that has been pre-trained to generate object detection results based on the input image 310. The object region 315 includes the reference object 311, the shading 314 of the reference object 311, and the shadows 312, 313 of the reference object 311.
[0044] Object information for the reference object 311 can be determined based on the object detection results. The object information may include at least one of the pose, shape, and material of the reference object 311. The pose is derived from the object region 315. If the reference object 311 is known, the shape and material may be determined via an object database; if it is unknown, the shape and material may be determined based on information from a proxy object.
[0045] If the reference plane of the reference object 311 is known, the shape and material of the reference plane are determined via the object database. In the case of a combined structure, the pose of the plane may be determined via the relationship between the object and the plane stored in the object database. The reference object 311 and reference plane shown in Figure 3 are examples of a combined structure. A spherical substructure corresponding to the reference object 311 and a flat substructure corresponding to the reference plane can be combined to form such a structure. If the reference plane is unknown, the pose, shape, and material are determined via plane detection. Plane detection may be performed by a plane detection model. The plane detection model may be a machine learning model (e.g., a neural network model) that has been pre-trained to generate plane detection results based on the input image 310. If the plane shape is assumed to be flat, its shape is not considered as plane information.
[0046] A rendering result 320 can be derived based on light information, object information, and plane information. The rendering result 320 may include a virtual object 321, shading 324 of the virtual object 321, and shadows 322, 323 of the virtual object 321. Shading 324 is derived through a shading rendering method, and shadows 322, 323 are derived through a shadow rendering method. The rendering result 320 can be formed by the fusion of shading 324 and shadows 322, 323.
[0047] The rendering result 320 may be compared with the input image 310, and the light estimation model may be updated based on the comparison result (e.g., the difference). Here, the pixel data of a certain area (e.g., an object area) of the input image 310 is compared with the pixel data of the corresponding area of the rendering result 320. For example, the pixel data showing the shading 314 and shadows 312,313 of the reference object 311 may be compared pixel by pixel with the pixel data showing the shading 324 and shadows 322,323 of the virtual object 321.
[0048] Figure 4 shows a training operation using a known reference object and a known reference plane according to one embodiment. Referring to Figure 4, the object detection model 410 detects the reference object from the input video 401 and generates an object detection result 402. The object detection model 410 may be a machine learning model (e.g., a neural network model) that has been pre-trained to generate the object detection result 402 based on the input video 401. The known object and the known plane are combined to form a known structure, and such a known structure may be filmed as the reference object in the input video 401.
[0049] The object detection result 402 may include information about the object region corresponding to the reference object (e.g., the location and size of the object region). Information about the connected structure corresponding to the reference object may be pre-stored in the object database 430. For example, the shape and material of the connected structure may be stored. The shape and material of the connected structure may include the connection relationship between the reference object and the reference plane (e.g., the connection point), the shape and material of the reference object (e.g., a spherical substructure), and the shape and material of the reference plane (e.g., a flat substructure).
[0050] Object information 404 and plane information 405 can be determined based on the object detection result 402 and the object database 430. The object information 404 may include at least one of the pose, shape, and material of the reference object, and the plane information 405 may include at least one of the pose, shape, and material of the reference plane. For example, the pose of the reference object and / or reference plane may be determined based on the object detection result 402 and / or the connection relationship, and the shape and material of the reference object and / or reference plane may be determined based on the object database 430.
[0051] The light estimation model 420 can estimate light information 403 corresponding to the input image 401. The rendering model 440 can generate a rendering result 406 including virtual objects based on the light information 403, object information 404, and plane information 405. The light estimation model 420 can be trained based on the comparison results between the reference object in the input image 401 and the virtual object in the rendering result 406. Since the object information 404 and plane information 405 provide near-factual data about cues such as objects, shading, planes, and shadows based on known information, the accuracy of the light information 403 can be improved by gradually reducing the difference between the reference object and the virtual object.
[0052] Figure 5 shows a training operation using a known reference object and an unknown reference plane according to one embodiment. Referring to Figure 5, the object detection model 510 can detect a reference object in the input video 501 and generate an object detection result 502. The object detection model 510 may be a machine learning model (e.g., a neural network model) that has been pre-trained to generate an object detection result 502 based on the input video 501. A known object (e.g., a known spherical object) may be captured as the reference object in the input video 501.
[0053] The object detection result 502 may include information about the object region corresponding to the reference object (e.g., the location and size of the object region). Information about the reference object may be pre-stored in the object database 540. For example, the shape and material of the reference object may be stored.
[0054] Object information 503 is determined based on the object detection result 502 and the object database 540. The object information 503 may include at least one of the pose, shape, and material of the reference object. For example, the pose of the reference object may be determined based on the object detection result 502, and the shape and material of the reference object may be determined based on the object database 540.
[0055] The plane detection model 520 can detect a reference plane from the input video 501 and generate plane information 504 based on the plane detection result. The plane detection model 520 may be a machine learning model (e.g., a neural network model) that has been pre-trained to generate a plane detection result from the input video 501. The plane information 504 may include at least one of the pose, shape, and material of the reference plane.
[0056] The light estimation model 530 estimates light information 505 corresponding to the input image 501. The rendering model 550 can generate a rendering result 506 including virtual objects based on the light information 505, object information 503, and plane information 504. The light estimation model 530 can be trained based on the comparison result between the reference object in the input image 501 and the virtual object in the rendering result 506. Since the object information 503 and plane information 504 provide factual data about cues such as objects, shading, planes, and shadows based on known information and plane detection, the accuracy of the light information 505 can be improved by gradually reducing the difference between the reference object and the virtual object.
[0057] Figure 6 shows a training operation using an unknown reference object and an unknown reference plane according to one embodiment. Referring to Figure 6, the object detection model 610 can detect a reference object from the input video 601 and generate an object detection result 602. The object detection model 610 may be a machine learning model (e.g., a neural network model) that has been pre-trained to generate an object detection result 602 based on the input video 601. The input video 601 may include an unknown reference object and an unknown reference plane. In other words, the input video 601 may be a general video and may be collected through various channels such as the web, video platforms, and photo communities.
[0058] The object detection result 602 may include information about the object region corresponding to the reference object (e.g., the location and size of the object region). A proxy object is generated based on the object detection result 602. Block 640 shows the proxy object generation operation. For example, the proxy object may be a cube, a cylinder, a point cloud, etc. Information about the proxy object may be stored in advance in the object database. For example, the shape and material of the proxy object may be stored.
[0059] Object information 603 is determined based on the information of the proxy object. Object information 603 may include at least one of the proxy object's pose, shape, and material. For example, the proxy object's pose may be determined based on the object detection result 602, and the shape and material of the reference object may be determined based on the proxy object's shape and material. In other words, the shape and material of the proxy object may be substituted for the shape and material of the reference object. Using the information of a known proxy object instead of trying to grasp the information (shape, material, etc.) of an unknown target object is effective for rendering and training. In some embodiments, shadows are used, but the benefits of using proxy objects are even greater because shadows can also provide crucial clues about rough characteristics such as direction.
[0060] The plane detection model 620 can detect a reference plane from the input video 601 and generate plane information 604 based on the plane detection result. The plane detection model 620 may be a machine learning model (e.g., a neural network model) that has been pre-trained to generate plane detection results from the input video 601. The plane information 604 may include at least one of the pose, shape, and material of the reference plane.
[0061] The light estimation model 630 estimates light information 605 corresponding to the input image 601. The rendering model 650 generates a rendering result 606 including virtual objects based on the light information 605, object information 603, and planar information 604. The rendering model 650 may also render virtual objects, shading, and shadows corresponding to proxy objects. The light estimation model 630 can be trained based on the comparison result between the reference object in the input image 601 and the virtual object in the rendering result 606. The reference object and the virtual object may have different shapes and materials, but meaningful training directions may be presented through the use of proxy objects and shadows.
[0062] Figure 7 shows shadow rendering operation according to one embodiment. The shadow of a virtual object can be rendered through the fusion of light information and visibility information. If the reference object and reference plane are a known combined structure, the visibility information may be predetermined in the preprocessing stage 710. Then, the visibility information is fused with the light information in the execution stage 720 where rendering is performed. By pre-determining the visibility information, rendering calculations and time can be reduced.
[0063] Referring to Figure 7, in step S711, visibility rays are sampled for each sampling point. The sampling points are defined on a reference plane, and visibility rays incident in various directions are defined for each sampling point. In step S712, a visibility map is generated based on the sampling results.
[0064] A visibility map includes the occlusion ratio of visibility rays for each of several sampling points. For example, if a reference object blocks the incidence direction of any ray at any of the sampling points, the occlusion ratio of that ray may be 1.0, and if there is no obstacle blocking the incidence of other rays at the sampling point, the occlusion ratio of that ray may be 0.0. The occlusion ratio is defined as a value between 1.0 and 0.0. However, such a number is merely an example, and the occlusion ratio may be limited to other numerical ranges.
[0065] In the execution stage 720, given an input image, in step S721, light information corresponding to the input image is estimated. The light information may be estimated via a light estimation model. In step S722, calculations are performed based on the light information and visibility information, and in step S723, the shadow of each sampling point is determined. More specifically, the shadow information of each sampling point can be determined by fusing (e.g., multiplying) the light information and visibility information of each sampling point in the reference plane.
[0066] Graph 701 shows the light information for one of the sampling points, Graph 702 shows the visibility information for that sampling point, and Graph 703 shows the result of fusing the light information and visibility information. Graphs 701-703 show cross-sections of hemispheres centered on the sampling points. White areas in the hemispheres indicate areas where light is present, and black areas indicate areas where light is absent. Graph 701 shows that illumination is present in a certain range at the 12 o'clock position, and Graph 702 shows that an obstacle (reference object) is present in a certain range at the 1 o'clock position.
[0067] When the light information in graph 701 and the visibility information in graph 702 are combined, the shadow information in graph 703 is derived. The shadow information in graph 703 indicates that a portion of the illumination area is blocked by an obstacle (reference object), and light is supplied to the sampling point from the remaining area. The blocked portion can form a shadow. Once the shadow information for each sampling point is determined in this way, in step S724, the shadow information is projected onto the shooting view of the input video, shadow rendering is performed, and the rendering result 725 is determined.
[0068] Figure 8 shows the operation of determining a visibility map according to one embodiment. Referring to Figure 8, a reference object 810 and a reference plane 820 are combined to form a known structure. A sampling point 801 is defined on the reference plane 820, and visibility rays 802 corresponding to multiple incident directions are defined with respect to the sampling point 801. In addition, an occlusion ratio table 803 is defined for the sampling point 801.
[0069] Multiple sampling points are defined on the reference plane 820, and a occlusion ratio table 803 is defined for each sampling point. Each element of the occlusion ratio table 803 represents the occlusion ratio for a specific direction of incidence. For example, the center of the occlusion ratio table 803 corresponds to sampling point 801, and the position of each element corresponds to each direction of incidence when sampling point 801 is viewed vertically from above.
[0070] Depending on the position of the sampling point 801 on the reference plane 820 and the direction of incidence of visibility rays, some elements may have complete visibility while others have no visibility at all. For example, a ratio of 0.4 for a table element indicates that 0.4 of the rays in that incident direction are blocked by an obstacle (reference object 810). A blockage ratio table 803 for all sampling points can be pre-calculated and stored in the visibility map. As described above, the visibility map may be calculated in a preprocessing step before model training and then used in the model training step. This can reduce the time required for shadow rendering.
[0071] Figure 9 shows a training method according to one embodiment. Referring to Figure 9, in step S910, optical information corresponding to the input image is estimated using an optical estimation model. In step S920, a reference object in the input image is detected. In step S930, object information of the reference object and planar information relating to the reference plane supporting the reference object are determined, and in step S940, a virtual object is rendered based on the optical information, object information, and planar information. In step S950, the optical estimation model is updated based on the comparison result between the reference object and the virtual object. In addition, the explanations in Figures 1 to 8, Figure 10, and Figure 11 may be applied to the training method.
[0072] Figure 10 shows a training device and a user terminal according to one embodiment. Referring to Figure 10, the training device 1010 can train the optical estimation model 1001. At least a portion of the optical estimation model 1001 can be implemented as a hardware module and / or a software module. The training device 1010 may perform the training operations shown in Figures 1 to 9.
[0073] For example, the training device 1010 can estimate light information corresponding to the input image using the light estimation model 1001, detect reference objects in the input image, determine object information of the reference objects and planar information of the reference plane supporting the reference objects, render a virtual object corresponding to the reference objects based on the light information, object information, and planar information, and update the light estimation model based on the comparison results between the reference objects and the virtual objects. The training device 1010 can render the shading and shadows of the virtual objects and compare the pixel data showing the shading and shadows of the reference objects with the pixel data showing the shading and shadows of the virtual objects.
[0074] Once training of the optical estimation model 1001 is complete, the optical estimation model 1001 can be installed on the user terminal 1020. The user terminal 1020 can use the optical estimation model 1001 to estimate optical information corresponding to the input video and use the optical information to give a sense of realism to virtual objects. For example, the optical information may be applied to virtual objects such as AR and CG. In addition, the explanations in Figures 1 to 9 and Figure 11 apply to the training device 1010, the optical estimation model 1001, and the user terminal 1020.
[0075] Figure 11 shows an electronic device according to one embodiment. Referring to Figure 11, the electronic device 1100 includes a processor 1110, a memory 1120, a camera 1130, a storage device 1140, an input device 1150, an output device 1160, and a network interface 1170, which can communicate via a communication bus 1180. For example, the electronic device 1100 can be implemented as at least part of a mobile device such as a mobile phone, smartphone, PDA, netbook, tablet computer, or laptop computer; a wearable device such as a smartwatch, smart band, or smart glasses; a computing device such as a desktop or server; a home appliance such as a television, smart TV, or refrigerator; a security device such as a door rack; or a vehicle such as an autonomous vehicle or smart vehicle. The electronic device 1100 may structurally and / or functionally include the training device 1010 and / or user terminal 1020 shown in Figure 10.
[0076] The processor 1110 executes functions and instructions for execution within the electronic device 1100. For example, the processor 1110 processes instructions stored in the memory 1120 or the storage device 1140. The processor 1110 may perform one or more operations as described with reference to Figures 1 to 10. The memory 1120 may include a computer-readable storage medium or a computer-readable storage device. The memory 1120 stores instructions for execution by the processor 1110 and related information while the software and / or applications are executed by the electronic device 1100.
[0077] Camera 1130 takes photographs and / or videos. The photographs and / or videos may constitute the input images for the optical estimation model. The storage device 1140 includes a computer-readable storage medium or a computer-readable storage device. The storage device 1140 can store a larger amount of information than memory 1120 and can store information for a longer period of time. For example, the storage device 1140 may include a magnetic hard disk, an optical disk, flash memory, a floppy disk, or other forms of non-volatile memory known in the art.
[0078] The input device 1150 can receive input from the user via traditional input methods such as a keyboard and mouse, and newer input methods such as touch input, voice input, and image input. For example, the input device 1150 may include a keyboard, mouse, touchscreen, microphone, or any other device that can detect input from the user and transmit the detected input to the electronic device 1100. The output device 1160 can provide the user with the output of the electronic device 1100 via a visual, auditory, or tactile channel. The output device 1160 may include, for example, a display, touchscreen, speaker, vibration generator, or any other device that can provide output to the user. The network interface 1170 can communicate with external devices via a wired or wireless network.
[0079] The embodiments described above are embodied in hardware components, software components, or combinations of hardware and software components. For example, the devices and components described in these embodiments are embodied using one or more general-purpose or special-purpose computers, such as a processor, controller, ALU (arithmetic logic unit), digital signal processor, microcomputer, FPA (field programmable array), PLU (programmable logic unit), microprocessor, or different devices that execute and respond to instructions. The processing device executes an operating system (OS) and one or more software applications that run on the OS. The processing device also accesses, stores, manipulates, processes, and generates data in response to the execution of the software. For convenience of understanding, the processing device may sometimes be described as being used as a single unit, but a person with ordinary skill in the art will understand that the processing device includes multiple processing elements and / or multiple types of processing elements. For example, the processing device includes multiple processors or one processor and one controller. Other processing configurations are also possible, such as a parallel processor.
[0080] Software includes computer programs, code, instructions, or a combination of one or more of these, which can configure a processing unit to operate as desired, or instruct the processing unit independently or in combination. Software and / or data can be permanently or temporarily embodied in any type of machine, component, physical device, virtual device, computer storage medium or device, or transmitted signal wave, for interpretation by a processing unit or for providing instructions or data to a processing unit. Software can be distributed across a network of computer systems and stored and executed in a distributed manner. Software and data can be stored on a recording medium readable by one or more computers.
[0081] The method according to this embodiment is embodied in the form of program instructions that are implemented via various computer means and recorded on a computer-readable recording medium. The recording medium includes program instructions, data files, data structures, etc., individually or in combination. The recording medium and program instructions may be specifically designed and configured for the purposes of the present invention, or they may be known and usable by those skilled in the art who have technology in the field of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floppy disks, and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memory. Examples of program instructions include not only machine code generated by a compiler, but also high-level language code executed by a computer using an interpreter or the like.
[0082] The hardware device described above may be configured to operate as one or more software modules to perform the operations shown in the present invention, and vice versa.
[0083] As described above, although embodiments have been illustrated with limited drawings, a person with ordinary skill in the art can apply various technical modifications and variations based on the above description. For example, the described techniques may be performed in a different order than described, and / or the components of the described systems, structures, devices, circuits, etc. may be combined or assembled in a different manner than described, or replaced or substituted with other components or equivalents, and still achieve suitable results. [Explanation of Symbols]
[0084] 100: Light estimation device 101: Input video 102: Optical information 110: Light Estimation Model
Claims
1. A method for light estimation, The steps include: estimating the optical information corresponding to the input video using an optical estimation model; The steps include detecting a reference object in the input video, The steps include determining object information of the reference object and planar information of the reference plane supporting the reference object, A step of rendering a virtual object corresponding to the reference object based on the light information, the object information, and the planar information, The steps include updating the light estimation model and training the light estimation model based on the comparison result between the reference object and the virtual object, Includes, The object information includes the material of the reference object, The planar information includes the material of the reference plane, and the method.
2. The method according to claim 1, wherein the step of rendering the virtual object includes the step of rendering the shading and shadows of the virtual object.
3. The method according to claim 1, wherein the comparison result includes a comparison result between pixel data showing the shading and shadow of the reference object and pixel data showing the shading and shadow of the virtual object.
4. The object information further includes at least one of the pose and shape of the reference object, The method according to claim 1, wherein the planar information further includes at least one of the pose and shape of the reference plane.
5. The method according to claim 1, wherein the step of determining the object information and the planar information includes, if the reference object is a known object, the step of determining at least a portion of the object information using an object database.
6. The method according to claim 1, wherein the step of determining the object information and the plane information includes, if the reference object and the reference plane are a known combined structure, the step of determining at least a portion of the object information and the plane information using an object database.
7. The method according to claim 6, wherein a spherical substructure corresponding to the reference object and a flat substructure corresponding to the reference plane are combined to form the structure.
8. The method according to claim 1, wherein the step of determining the object information and the plane information includes, if the reference plane is an unknown plane, the step of detecting the reference plane from the input video and determining the plane information.
9. The method according to claim 1, wherein the step of determining the object information and the planar information includes, if the reference object is an unknown object, the step of determining the object information based on information of a predetermined proxy object.
10. The step of rendering the virtual object is: The steps include: determining the shadow information of each sampling point by fusing the light information and visibility information of each sampling point on the aforementioned reference plane; The steps include: projecting the shadow information of each sampling point on the reference plane onto the shooting view of the input video to render the shadow of the virtual object; The method according to claim 1, including the method described in claim 1.
11. The method according to claim 10, wherein, if the reference object and the reference plane are a known combined structure, the visibility information is predetermined for the shadow rendering of the combined structure.
12. The method according to claim 1, wherein the step of training the light estimation model includes updating the light estimation model so that the difference between the reference object and the virtual object becomes smaller.
13. The aforementioned method for light estimation is: Steps to acquire other video footage, The steps include: estimating the optical information corresponding to the other image using the trained optical estimation model; The method according to claim 1, further comprising:
14. A computer program stored on a computer-readable recording medium for use in conjunction with hardware to perform the method according to any one of claims 1 to 13.
15. A device for light estimation, Processor and A memory containing an instruction word that can be executed by the aforementioned processor, Includes, When the instruction word is executed by the processor, the method according to any one of claims 1 to 13 is performed. Device.
16. A camera that generates input video, Using an optical estimation model, the optical information corresponding to the input video is estimated. The reference object in the input video is detected, The object information of the reference object and the planar information of the reference plane supporting the reference object are determined. Based on the light information, object information, and planar information, a virtual object corresponding to the reference object, the shading of the virtual object, and the shadow of the virtual object are rendered. A processor that updates and trains the light estimation model based on the comparison result between pixel data representing the shading and shadow of the reference object and pixel data representing the shading and shadow of the virtual object. Includes, The object information includes the material of the reference object, The aforementioned planar information includes the material of the reference plane, and is an electronic device.
17. The electronic device according to claim 16, wherein the processor determines the object information based on predetermined proxy object information.
18. A method for light estimation, Steps to acquire the second video, The steps include: estimating second optical information corresponding to the second image using a trained optical estimation model; Includes, The aforementioned optical estimation model is, Using the aforementioned optical estimation model, optical information corresponding to the input image is estimated. The reference object is detected in the aforementioned input video, The object information of the reference object and the planar information of the reference plane supporting the reference object are determined. Based on the aforementioned light information, object information, and planar information, a virtual object corresponding to the reference object is rendered. Based on the results of comparing the reference object and the virtual object, the light estimation model is updated and trained. The object information includes the material of the reference object, The planar information includes the material of the reference plane, and the method.
19. The steps include rendering a second virtual object corresponding to a second object in the second image based on the second optical information, The steps include generating augmented reality by superimposing the second virtual object onto the second image, The method according to claim 18, further comprising:
20. A method for light estimation, The steps include: estimating the optical information corresponding to the input video using an optical estimation model; A step of determining the information of a reference object in the input video based on whether the stored information corresponds to a reference object, The steps include rendering a virtual object corresponding to the reference object based on the optical information and the information of the reference object, The steps include comparing the reference object with the rendered virtual object, updating the light estimation model, and training the light estimation model, Includes, A method wherein the information of the reference object includes at least a portion of the object information of the reference object and the planar information of a reference plane supporting the reference object, the object information includes the material of the reference object, and the planar information includes the material of the reference plane.
21. The stored information is, It is a predetermined object, The method according to claim 20, wherein the step of determining the information of the reference object includes determining the stored information as the information of the reference object in response that the reference object corresponds to the predetermined object.
22. The method according to claim 20, wherein the step of determining the information of the reference object includes determining the stored information as the information of the reference object in response that the stored information corresponds to the reference object.