A method and device for drawing shadows of virtual objects
Optimize the light source position through the HDR panoramic image generation model and the light spherical harmony feature generation model, solving the lighting consistency problem of virtual object shadow drawing in AR, and improving the realism and user experience of virtual objects.
Patent Information
- Application Number
- CN202210310337.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-03-28
AI Technical Summary
In the existing AR technology, when drawing shadows of virtual objects, the light source position estimate is not accurate enough, resulting in poor lighting consistency between virtual objects and real scenes, affecting the user experience.
The HDR panoramic image generation model and the light spherical harmonic feature generation model are used to optimize the spherical Gaussian intensity distribution of the light spherical harmonic feature model through the KL divergence loss function, and the light source position is determined using the HDR panoramic image, and the virtual object shadow is drawn according to the light source position.
It improves the lighting consistency between virtual objects and real scenes, improves the realism and AR experience of virtual objects, and saves time and resources in the model training process.
Smart Images

Figure CN114638950B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of Augmented Reality (AR), and particularly to a method and device for rendering shadows of virtual objects. Background Art
[0002] AR technology is a new technology developed on the basis of Virtual Reality (VR). It is a technology that enhances users' perception of the real world through information provided by a computer system, and superimposes virtual objects, virtual scenes, system prompt information, or non-geometric information about real objects generated by the computer system onto the real world, thereby realizing the "enhancement" of the real world and has been widely used in various industries.
[0003] In AR experiences, users have a subtle sense of lighting, and usually regard lighting consistency as an important indicator of virtual-real fusion. Lighting consistency means that virtual objects have the same lighting effects as real objects. The goal of lighting consistency is to make the lighting conditions of virtual objects consistent with those in the real scene, that is, virtual objects and real objects have consistent light and shadow effects to enhance the realism of virtual objects.
[0004] Currently, when rendering shadows for virtual objects in AR scenes, most related technologies use a deep neural network to learn the original images captured by a camera, output an environment texture map, determine the light source position based on the environment texture map, and then render the shadow effect for the virtual object according to the light source position. However, since the original images captured by the camera of an AR device are generally Low Dynamic Range (LDR) images with less lighting information, the estimated light source position is not accurate enough, thereby reducing the lighting consistency between virtual objects and the real scene and affecting the user's AR experience. Summary of the Invention
[0005] Embodiments of this application provide a method and device for rendering shadows of virtual objects to improve the lighting consistency between virtual objects and the real scene.
[0006] In a first aspect, an embodiment of this application provides a method for rendering shadows of virtual objects, which is applied to an AR device and includes:
[0007] Obtain any low dynamic range (LDR) image of a real scene captured by the camera of the AR device;
[0008] Input the LDR image into a high dynamic range (HDR) panoramic image generation model to obtain an HDR panoramic image, where the HDR panoramic image contains the true spherical Gaussian intensity distribution of ambient light;
[0009] Input the HDR panoramic image into the lighting spherical harmonic feature generation model, and optimize the original spherical Gaussian intensity distribution of the lighting spherical harmonic feature generation model by using the true spherical Gaussian intensity distribution to obtain the target spherical Gaussian intensity distribution;
[0010] Determine the position of the light source in the virtual scene according to the true spherical Gaussian intensity distribution and the target spherical Gaussian intensity distribution;
[0011] Overlay and display the virtual object in the LDR image, and draw a shadow for the overlaid and displayed virtual object according to the position of the light source.
[0012] In a second aspect, an AR device provided by an embodiment of the present application includes a processor, a memory, a camera, and a display screen. The display screen, the camera, the memory, and the processor are connected through a bus:
[0013] The memory stores a computer program, and the processor performs the following operations according to the computer program:
[0014] Obtain a low dynamic range (LDR) image of any real scene collected by the camera and display it on the display screen;
[0015] Input the LDR image into a high dynamic range (HDR) panoramic image generation model to obtain an HDR panoramic image, where the HDR panoramic image includes the true spherical Gaussian intensity distribution of the ambient light;
[0016] Input the HDR panoramic image into the lighting spherical harmonic feature generation model, and optimize the original spherical Gaussian intensity distribution of the lighting spherical harmonic feature generation model by using the true spherical Gaussian intensity distribution to obtain the target spherical Gaussian intensity distribution;
[0017] Determine the position of the light source in the virtual scene according to the true spherical Gaussian intensity distribution and the target spherical Gaussian intensity distribution;
[0018] Overlay and display the virtual object in the LDR image, and draw a shadow for the virtual object overlaid and displayed on the display screen according to the position of the light source.
[0019] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores computer-executable instructions for causing a computer to execute a method for drawing a shadow of a virtual object.
[0020] In the above embodiments of the present application, for any real scenario, an LDR image collected by a camera of an AR device is obtained, and an HDR panoramic image is obtained through an HDR panoramic image generation model. Among them, the HDR panoramic image contains the true spherical Gaussian intensity distribution of environmental light. In this way, after the HDR panoramic image is input into the illumination spherical harmonic feature generation model, the true spherical Gaussian intensity distribution can be used to optimize the original spherical Gaussian intensity distribution of the illumination spherical harmonic feature generation model to obtain a target spherical Gaussian intensity distribution, and based on the true spherical Gaussian intensity distribution and the target spherical Gaussian intensity distribution, the position of the light source in the virtual scene is determined. Further, according to the position of the light source, shadows are drawn for the virtual objects superimposed and displayed on the LDR image. Since the true spherical Gaussian intensity distribution is used to optimize the original spherical Gaussian intensity distribution of the illumination spherical harmonic feature generation model, there is no need to train the illumination spherical harmonic feature generation model, which saves time and effort; moreover, since the HDR panoramic image contains the true spherical Gaussian intensity distribution of environmental light, the optimized target spherical Gaussian intensity distribution can truly reflect the position of the light source in the virtual space. Therefore, when shadows are drawn for virtual objects according to the position of the light source in the virtual scene, the virtual objects and the real scene have a consistent lighting effect, improving the authenticity of the virtual-real fusion and further enhancing the AR experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 Exemplarily shows a flowchart of a training method for an HDR panoramic image generation model provided by an embodiment of the present application;
[0023] Figure 2 Exemplarily shows a flowchart of a method for drawing shadows of virtual objects provided by an embodiment of the present application;
[0024] Figure 3 Exemplarily shows a flowchart of a method for determining a target spherical Gaussian intensity distribution provided by an embodiment of the present application;
[0025] Figure 4 Exemplarily shows a flowchart of a fine-tuning method for an illumination spherical harmonic feature generation model provided by an embodiment of the present application;
[0026] Figure 5 Exemplarily shows a flowchart of a method for determining the position of a light source in a target spherical Gaussian intensity distribution provided by an embodiment of the present application;
[0027] Figure 6 The structural diagram of the AR device provided by the embodiments of the present application is exemplarily shown. Detailed implementation manners
[0028] To make the objectives, implementation manners and advantages of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part rather than all of the embodiments of the present application.
[0029] Based on the exemplary embodiments described in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope protected by the appended claims of the present application. In addition, although the disclosed content in the present application is introduced according to one or several exemplary instances, it should be understood that each aspect of these disclosed contents can also constitute a complete implementation manner alone.
[0030] To clearly describe the embodiments of the present application, the following gives explanations of the terms in the present application.
[0031] Low Dynamic Range (LDR) image: An image in which each of the red (R), green (G), and blue (B) color channels is described by an 8-bit integer. The storage formats of LDR images include JPG and PNG, etc.
[0032] High Dynamic Range (HDR) image: An image in which each of the red (R), green (G), and blue (B) color channels is described by a 32-bit floating-point number. Compared with LDR images, HDR images can provide more dynamic range and image details. The storage formats of HDR images include HDR, TIF, EXR, and RAW, etc.
[0033] HDR panoramic image: Refers to an HDR image that conforms to the normal effective viewing angle of a person's two eyes or above the peripheral vision angle of the two eyes, and even up to 360 degrees.
[0034] In an AR scenario, in order to naturally integrate the simulated virtual objects with the real scene, the virtual objects should have the same lighting effect as the real scene. This requires that the surface brightness of the virtual objects should be the same as the brightness of the ambient light, and the shadows of the virtual objects should be in the same position as the light source. For example, in a darker environment, the surface brightness of the virtual objects should be darker; in a brighter environment, the surface brightness of the virtual objects should be brighter; when the light source is in the front left of the virtual object, the shadow is in the back right of the virtual object, and when the distance between the light source and the virtual object is relatively close, the shadow area is larger.
[0035] Currently, most deep neural networks estimate the light source position based on the original images captured by cameras. However, due to the time-consuming and laborious training process of deep neural networks, and the fact that the original images captured by cameras are generally LDR images with less light source information, the estimation of the light source position is not accurate enough, which in turn reduces the lighting consistency between virtual objects and the real scene and affects the user's AR experience.
[0036] In view of this, the embodiments of the present application provide a method and device for drawing virtual object shadows. The LDR image captured by the camera is input into the HDR panoramic image generation model to obtain the corresponding HDR panoramic image, and the HDR panoramic image is input into the lighting spherical harmonic feature model. The HDR panoramic image is used to supervise the lighting spherical harmonic feature model, and the KL (Kullback-Leibler Divergence) loss function is used to fine-tune the lighting spherical harmonic feature model, so as to optimize the original spherical Gaussian intensity distribution of the lighting spherical harmonic feature model. The brightest value in the optimized spherical Gaussian intensity distribution is determined as the light source position, and the shadow of the virtual object is drawn according to the light source position. Since the HDR panoramic image completely records the ambient light information and can represent the real spherical Gaussian intensity distribution of the ambient light information, the accuracy of determining the light source position is improved, the drawing effect of the virtual object shadow is enhanced, and thus the lighting consistency between the virtual object and the real scene is improved. Moreover, by supervising the lighting spherical harmonic feature model with the HDR panoramic image, the cumbersome training process of the lighting spherical harmonic feature model is omitted, and a better spherical Gaussian intensity distribution can be obtained through fine-tuning, reducing the direction error of the light source position, making the shadow of the virtual object softer, improving the authenticity of the virtual object, and thus enhancing the user's AR experience.
[0037] In the embodiments of the present application, a HDR panoramic image generation model and a lighting spherical harmonic feature generation model are designed using the principle of deep learning. Among them, the HDR panoramic image generation model is used to generate the HDR panoramic image according to the LDR image in the real scene, and the lighting spherical harmonic feature generation model is used to estimate the position of the light source in the virtual scene according to the HDR panoramic image.
[0038] Among them, the training process of the HDR panoramic image generation model is referred to Figure 1 and mainly includes the following steps:
[0039] S101: Obtain multiple LDR images in each real scene and the corresponding real HDR panoramic image of each LDR image to obtain a training sample set.
[0040] In an alternative embodiment, a color transfer algorithm can be used to obtain the true HDR image corresponding to each LDR image, and the true HDR panoramic image can be obtained through transfer learning.
[0041] S102: Based on the obtained training sample set, perform multiple rounds of iterative training on the initial HDR panoramic image generation model until a preset condition is met, and obtain the final HDR panoramic image generation model.
[0042] Among them, the HDR panoramic image generation model is composed of multiple Convolutional Neural Networks (CNNs). The process of each round of iterative training is as follows:
[0043] S1021: According to the initial HDR panoramic image generation model, determine the predicted HDR panoramic image corresponding to each LDR image.
[0044] S1022: According to each predicted HDR panoramic image and the corresponding true HDR panoramic image, determine the prediction loss value.
[0045] S1023: Adjust the parameters of the initial HDR panoramic image generation model according to the prediction loss value.
[0046] It should be noted that the training process of the HDR panoramic image generation model is not the key content of this application, and no detailed description will be given here.
[0047] After obtaining the trained HDR panoramic image generation model, for any input LDR image, the output HDR panoramic image can be obtained.
[0048] Next, based on the trained HDR panoramic image generation model, the method flow for drawing virtual object shadows provided in the embodiments of this application will be described. Specifically, refer to Figure 2 This process is executed by the AR device and mainly includes the following steps:
[0049] S201: Obtain an LDR image of any real scene captured by the camera of the AR device.
[0050] In the embodiments of this application, the AR device is equipped with a camera, and an LDR image of any real scene under the current Field of View (FOV) can be captured thereby.
[0051] S202: Input the LDR image into the HDR panoramic image generation model to obtain the corresponding HDR panoramic image, where the HDR panoramic image contains the true spherical Gaussian intensity distribution of the ambient light.
[0052] Since the HDR panoramic image records the complete environmental light information and contains the true spherical Gaussian intensity distribution of the environmental light, the spherical Gaussian intensity distribution of the light source in the virtual scene can be obtained by using the HDR panoramic image to supervise the light spherical harmonic feature generation model.
[0053] S203: Input the HDR panoramic image into the light spherical harmonic feature generation model, and use the true spherical Gaussian intensity distribution to optimize the original spherical Gaussian intensity distribution of the light spherical harmonic feature generation model to obtain the target spherical Gaussian intensity distribution.
[0054] In S203, input the HDR panoramic image into the light spherical harmonic feature generation model as the supervisor of the light spherical harmonic feature generation model. Using the Fine-tuning method, the original spherical Gaussian intensity distribution of the light spherical harmonic feature generation model is optimized by using the true spherical Gaussian intensity distribution, saving the training process of the light spherical harmonic feature generation model and saving manpower and material resources. Moreover, since the HDR panoramic image contains the true spherical Gaussian intensity distribution of the environmental light, the optimized target spherical Gaussian intensity distribution can truly reflect the position of the light source in the virtual space, improving the accuracy of light source estimation.
[0055] For the specific optimization process, see Figure 3 :
[0056] S2031: Determine the KL divergence loss value between the true spherical Gaussian intensity distribution and the original spherical Gaussian intensity distribution.
[0057] Among them, the KL divergence, also known as relative entropy, is a measure of the asymmetry of the difference between two distributions P and Q. The KL divergence is used to measure the additional average number of bits required to encode the distribution obeying P using the distribution based on Q. Generally, P represents the true distribution of the data, and Q represents the estimated model distribution. Since the KL divergence can effectively express the closeness of two distributions and can measure how much information the original spherical Gaussian intensity distribution loses relative to the true spherical Gaussian intensity distribution, it can be used as the loss function of the light spherical harmonic feature generation model.
[0058] In the embodiments of the present application, after obtaining the true spherical Gaussian intensity distribution, using the original spherical Gaussian intensity distribution different from the true spherical Gaussian intensity distribution will definitely result in a loss of coding efficiency, and the additional average information amount increased during transmission is at least equal to the KL divergence between the two distributions. Therefore, the loss value of the light spherical harmonic feature generation model can be calculated based on the KL divergence, thereby fine-tuning the light spherical harmonic feature generation model.
[0059] S2032: Determine whether the KL divergence loss value is greater than or equal to the preset divergence threshold. If so, execute S2033; otherwise, execute S2035.
[0060] When the KL divergence loss value is greater than or equal to the preset divergence threshold, it indicates that there is a large difference between the true spherical Gaussian intensity distribution and the original spherical Gaussian intensity distribution. In this case, it is necessary to adjust the lighting spherical harmonic feature generation model to make the two distributions as close as possible and ensure lighting consistency. When the KL divergence loss value is less than the preset divergence threshold, it indicates that the difference between the true spherical Gaussian intensity distribution and the original spherical Gaussian intensity distribution is small, and there is no need to adjust the lighting spherical harmonic feature generation model.
[0061] Optionally, the preset divergence threshold in the embodiments of the present application is set to 0.3.
[0062] S2033: Fine-tune the parameters of the lighting spherical harmonic feature generation model to obtain the spherical Gaussian intensity distribution of the fine-tuned lighting spherical harmonic feature generation model.
[0063] In S2033, the process of fine-tuning the parameters of the lighting spherical harmonic feature generation model can be regarded as the process of predicting the spherical Gaussian intensity distribution in the virtual scene. By fine-tuning the spherical harmonic function of the lighting spherical harmonic feature generation model through the KL divergence loss value, the lighting consistency between the fine-tuned spherical Gaussian intensity distribution and the true spherical Gaussian intensity distribution is effectively ensured.
[0064] S2034: After calculating the KL divergence loss value between the fine-tuned spherical Gaussian intensity distribution and the true spherical Gaussian intensity distribution, return to S2032.
[0065] In the embodiments of the present application, after fine-tuning the parameters of the lighting spherical harmonic feature generation model, it is necessary to recalculate the KL divergence loss value between the fine-tuned spherical Gaussian intensity distribution and the true spherical Gaussian intensity distribution, and compare the new KL divergence loss value with the preset divergence threshold until the KL divergence loss value is less than the preset divergence threshold.
[0066] S2035: Use the current spherical Gaussian intensity distribution of the lighting spherical harmonic feature generation model as the target spherical Gaussian intensity distribution.
[0067] In S2035, when the KL divergence loss value is less than the preset divergence threshold, it indicates that the difference between the current spherical Gaussian intensity distribution and the true spherical Gaussian intensity distribution is small, and the current spherical Gaussian intensity distribution can be used as the target spherical Gaussian intensity distribution to determine the position of the light source in the virtual scene.
[0068] In an alternative embodiment, the lighting spherical harmonic feature generation model may adopt the Darknet121 network.
[0069] S204: Determine the position of the light source in the virtual scene according to the true spherical Gaussian intensity distribution and the target spherical Gaussian intensity distribution.
[0070] In the embodiments of the present application, the true spherical Gaussian intensity distribution reflects the light source position in the real scene. The light source position in the target spherical Gaussian intensity distribution can be determined using the light source position in the real scene, thereby obtaining the light source position in the virtual scene. For the specific implementation process, refer to Figure 4 :
[0071] S2041: Determine the first position of the maximum brightness in the true spherical Gaussian intensity distribution and, determine the second position of the maximum brightness in the target spherical Gaussian intensity distribution.
[0072] The position with the maximum brightness value in a distribution is most likely the light source position. Assume that the first position of the maximum brightness in the true spherical Gaussian intensity distribution is denoted as p1, and the second position of the maximum brightness in the target spherical Gaussian intensity distribution is denoted as p2.
[0073] S2042: Determine whether the position deviation between the first position and the second position is within a preset value range. If it is, execute S2043; otherwise, execute S2044.
[0074] Calculate the position deviation between p1 and p2 and determine whether the position deviation is within the preset value range. If it is, it indicates that the estimated second position is accurate, and the three-dimensional coordinates of the light source can be obtained using p2; otherwise, it indicates that the second position is inaccurate and the target spherical Gaussian intensity distribution needs to be re-determined.
[0075] Optionally, the preset value range is [16°, 52°].
[0076] S2043: Determine the three-dimensional coordinates of the light source in the virtual scene according to the second position.
[0077] Among them, for the calculation process of the three-dimensional coordinates, refer to Figure 5 :
[0078] S2043_1: Initialize the initial positions of each point in the target spherical Gaussian intensity distribution.
[0079] Taking the spherical Gaussian intensity distribution with 128 cores (i.e., this distribution contains 128 points) as an example, initialize the target spherical Gaussian intensity distribution to determine the initial positions of each point. The initialization formula is as follows:
[0080]
[0081]
[0082]
[0083]
[0084] α = sin -1u Formula 5
[0085] β = θ Formula 6
[0086] Where Δ represents the golden ratio, i represents the i-th point in the target spherical Gaussian intensity distribution, represents the golden section point in the target spherical Gaussian intensity distribution, θ represents the golden section angle in the target spherical Gaussian intensity distribution, u represents dividing the unit sphere into 128 equal parts, (α, β) represents the initial position of each point, α represents the pitch angle of each point in the target spherical Gaussian intensity distribution, and β represents the yaw angle of each point in the target spherical Gaussian intensity distribution.
[0087] S2043_2: Determine the target position of the center point among each point according to the second position of the maximum brightness in the target spherical Gaussian intensity distribution.
[0088] In S2043_2, according to the brightness function, the brightness of each point in the target spherical Gaussian intensity distribution can be obtained, and the formula is as follows:
[0089] I = Intensity[0] * 0.11 + Intensity[1] * 0.59 + Intensity[2] * 0.3 Formula 7
[0090] Where Intensity[0], Intensity[1], and Intensity[2] respectively represent the values of the R, G, and B components corresponding to each point in the target spherical Gaussian intensity distribution, and I represents the brightness value.
[0091] According to the brightness value I of each point in the target spherical Gaussian intensity distribution, the point with the maximum brightness (i.e., the second position) can be obtained. This point can represent the light source in the virtual scene, and taking this point as the center point, the target position of this point is determined. The formula for the target position is as follows:
[0092] α' = Evevation(I), β' = Azimuth(I) Formula 8
[0093] Where α' represents the target pitch angle of the center point, β' represents the target yaw angle of the center point, Evevation() is the pitch angle function, and Azimuth() is the yaw angle function.
[0094] S2043_3: Determine the three-dimensional coordinates corresponding to the target position of the center point, and use the three-dimensional coordinates as the position of the light source in the virtual scene.
[0095] In S2043_3, using the conversion relationship between polar coordinates and three-dimensional coordinates, the target position of the center point is converted into three-dimensional coordinates. The conversion formula is as follows:
[0096] X = cosβ′ Formula 9
[0097] Y = sinα′ Formula 10
[0098] z = cosβ′ Formula 11
[0099] After obtaining the three-dimensional coordinates corresponding to the target position, use the three-dimensional coordinates as the position of the light source in the virtual scene.
[0100] S2044: Reverse fine-tune the parameters of the light spherical harmonic feature generation model, re-determine the target spherical Gaussian intensity distribution, and execute S2041.
[0101] In S2044, when the position deviation between the first position and the second position is not within the preset value range, it indicates that the estimated light source position is inaccurate. It is necessary to reverse fine-tune the parameters of the light spherical harmonic feature generation model and re-determine the target spherical Gaussian intensity distribution until the position deviation between the second position and the first position in the new target spherical Gaussian intensity distribution is within the preset value range.
[0102] S205: Superimpose and display the virtual object in the LDR image, and draw a shadow for the superimposed virtual object according to the position of the light source in the virtual scene.
[0103] In an alternative embodiment, the user can select a virtual object by touching the display screen of the AR device and determine the placement position of the virtual object. After the user selects the virtual object, the virtual object is superimposed and displayed at the target position in the LDR image. During the display process, to improve the authenticity of the virtual-real fusion, a shadow can be drawn for the virtual object according to the determined light source position.
[0104] Specifically in implementation, assign the three-dimensional coordinates of the light source to the Direction variable in the Light function in the Unity engine, and the Unity engine renders the virtual object according to the value of Direction, so that the shadow of the virtual object is consistent with the illumination in the real scene.
[0105] Measured by experimental data, the position error of the light source in the virtual scene estimated by the HDR panoramic image generation model and the light spherical harmonic feature generation model provided in the embodiments of the present application is about 47.6°, which satisfies the angular difference within the visual range of the human eye, and visually shows that the virtual object and the real scene have a consistent illumination effect.
[0106] In a method for rendering shadows of virtual objects provided by an embodiment of the present application, an HDR panoramic image corresponding to an LDR image in a real scene is obtained through an HDR panoramic image generation model. Since the HDR panoramic image contains the true spherical Gaussian intensity distribution of environmental illumination, it can be used as a supervisor for a light spherical harmonic feature generation model. The spherical Gaussian intensity distribution of the light spherical harmonic feature generation model is optimized using the Fine-tunging method. During the optimization process, the KL divergence is used as a loss function to fine-tune the light spherical harmonic feature generation model to obtain a target spherical Gaussian intensity distribution. Since the light spherical harmonic feature generation model is supervised using the HDR panoramic image, the model training process is omitted, saving manpower and material resources. Moreover, since the HDR panoramic image contains the true spherical Gaussian intensity distribution of environmental illumination, the optimized target spherical Gaussian intensity distribution can truly reflect the position of the light source in the virtual space, and the position with the maximum brightness in the target spherical Gaussian intensity distribution is used as the light source position. Therefore, when rendering shadows for virtual objects based on the position of the light source in the virtual scene, the virtual objects and the real scene have a consistent lighting effect, improving the authenticity of the virtual-real fusion and thus enhancing the AR experience.
[0107] Based on the same technical concept, an embodiment of the present application provides an AR device. This AR device can implement the method steps of a method for rendering shadows of virtual objects in the above embodiment and achieve the same technical effects.
[0108] See Figure 6 , this AR device includes a processor 601, a memory 602, a camera 603, and a display screen 604. Among them, the display screen 604, the camera 603, the memory 602 are connected to the processor 601 through a bus 605;
[0109] The memory 602 stores a computer program, and the processor 601 performs the following operations according to the computer program stored in the memory 602:
[0110] Obtain an LDR image of any real scene collected by the camera 603 and display it on the display screen 604;
[0111] Input the LDR image into the HDR panoramic image generation model to obtain an HDR panoramic image, and the HDR panoramic image contains the true spherical Gaussian intensity distribution of environmental illumination;
[0112] Input the HDR panoramic image into the light spherical harmonic feature generation model, and use the true spherical Gaussian intensity distribution to optimize the original spherical Gaussian intensity distribution of the light spherical harmonic feature generation model to obtain a target spherical Gaussian intensity distribution;
[0113] Determine the position of the light source in the virtual scene according to the true spherical Gaussian intensity distribution and the target spherical Gaussian intensity distribution;
[0114] Superimpose and display the virtual object in the LDR image, and draw a shadow for the virtual object superimposed and displayed in the display screen 604 according to the position of the light source.
[0115] Optionally, the processor 601 optimizes the original spherical Gaussian intensity distribution of the illumination spherical harmonic feature generation model by using the true spherical Gaussian intensity distribution to obtain the target spherical Gaussian intensity distribution. The specific operation is as follows:
[0116] Determine the KL divergence loss value between the true spherical Gaussian intensity distribution and the original spherical Gaussian intensity distribution;
[0117] If the KL divergence loss value is greater than or equal to the preset divergence threshold, fine-tune the parameters of the illumination spherical harmonic feature generation model until the KL divergence loss value between the spherical Gaussian intensity distribution of the fine-tuned illumination spherical harmonic feature generation model and the true spherical Gaussian intensity distribution is less than the preset divergence threshold, and obtain the target spherical Gaussian intensity distribution corresponding to the illumination spherical harmonic feature generation model.
[0118] Optionally, the processor 601 determines the position of the light source in the virtual scene according to the true spherical Gaussian intensity distribution and the target spherical Gaussian intensity distribution. The specific operation is as follows:
[0119] Determine the first position of the maximum brightness in the true spherical Gaussian intensity distribution and, determine the second position of the maximum brightness in the target spherical Gaussian intensity distribution;
[0120] Determine whether the position deviation between the first position and the second position is within the preset value range;
[0121] If so, determine the three-dimensional coordinates of the light source in the virtual scene according to the second position.
[0122] Optionally, the processor 601 determines the three-dimensional coordinates of the light source in the virtual scene according to the second position. The specific operation is as follows:
[0123] Initialize the initial positions of each point in the target spherical Gaussian intensity distribution;
[0124] Determine the target position of the center point among the points according to the second position of the maximum brightness in the target spherical Gaussian intensity distribution;
[0125] Determine the three-dimensional coordinates corresponding to the target position of the center point, and use the three-dimensional coordinates as the position of the light source in the virtual scene.
[0126] Optionally, when determining whether the position deviation between the first position and the second position is within a preset value range, the processor 601 further performs the following operations:
[0127] Fine-tune the parameters of the illumination spherical harmonic feature generation model in the reverse direction until the position deviation between the first position and the second position is within the preset value range.
[0128] The embodiments of the present application Figure 6 The processor involved may be a central processing unit (CPU), a general-purpose processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in conjunction with the disclosure of the present application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. Among them, the memory can be integrated in the processor or can be separately provided from the processor.
[0129] It should be noted that Figure 6 is only an example, which shows the necessary hardware for the AR device to execute the method steps for drawing the shadow of a virtual object provided by the embodiments of the present application. Those not shown, the AR device also includes common hardware of human-computer interaction devices, such as speakers, microphones, left and right lens pieces, display screens, etc.
[0130] The embodiments of the present application also provide a computer-readable storage medium for storing some instructions, which, when executed, can complete the methods of the foregoing embodiments.
[0131] The embodiments of the present application also provide a computer program product for storing a computer program, and the computer program is used to execute the methods of the foregoing embodiments.
[0132] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0133] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0134] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0135] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0136] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A method for drawing the shadow of a virtual object, characterized in that Applied to an AR device, including: Obtain a low-dynamic range (LDR) image of any real scene captured by the camera of the AR device; Input the LDR image into a high-dynamic range (HDR) panoramic image generation model to obtain an HDR panoramic image, where the HDR panoramic image contains the true spherical Gaussian intensity distribution of the ambient light; Input the HDR panoramic image into a lighting spherical harmonic feature generation model to obtain the original spherical Gaussian intensity distribution of the lighting spherical harmonic feature generation model, and determine the KL divergence loss value between the true spherical Gaussian intensity distribution contained in the HDR panoramic image and the original spherical Gaussian intensity distribution. When the KL divergence loss value is greater than or equal to a preset divergence threshold, fine-tune the lighting spherical harmonic feature generation model until the KL divergence loss value is less than the preset divergence threshold, and use the spherical Gaussian intensity distribution of the HDR panoramic image output by the lighting spherical harmonic feature generation model after the fine-tuning stops as the target spherical Gaussian intensity distribution; wherein, each spherical Gaussian intensity distribution is the brightness intensity distribution of multiple points on the sphere; Determine the position of the light source in the virtual scene according to the true spherical Gaussian intensity distribution and the target spherical Gaussian intensity distribution; Overlay and display a virtual object in the LDR image, and draw a shadow for the overlaid and displayed virtual object according to the position of the light source.
2. The method according to claim 1, wherein The determining the position of the light source in the virtual scene according to the true spherical Gaussian intensity distribution and the target spherical Gaussian intensity distribution includes: Determine the point with the maximum brightness in the true spherical Gaussian intensity distribution, and determine the point with the maximum brightness in the target spherical Gaussian intensity distribution; Determine whether the position deviation between the two points with the maximum brightness is within a preset value range; If so, determine the three-dimensional coordinates of the light source in the virtual scene according to the point with the maximum brightness in the target spherical Gaussian intensity distribution.
3. The method according to claim 2, characterized in that, The determining the three-dimensional coordinates of the light source in the virtual scene according to the point with the maximum brightness in the target spherical Gaussian intensity distribution includes: Initialize the initial positions of each point in the target spherical Gaussian intensity distribution; Obtain the point with the maximum brightness in the target spherical Gaussian intensity distribution according to the brightness of the points at each initial position in the target spherical Gaussian intensity distribution; Use the brightness of the point with the maximum brightness in the target spherical Gaussian intensity distribution to determine the target position of the center point of each point; Determine the three-dimensional coordinates corresponding to the target position of the center point, and use the three-dimensional coordinates as the position of the light source in the virtual scene.
4. The method according to claim 2, characterized in that, When the position deviation between the two points with the maximum brightness is not within the preset value range, the method further includes: Reverse fine-tune the parameters of the lighting spherical harmonic feature generation model until the position deviation between the point with the maximum brightness in the true spherical Gaussian intensity distribution and the point with the maximum brightness in the target spherical Gaussian intensity distribution is within the preset value range.
5. An AR device, characterized in that, Including a processor, a memory, a camera, and a display screen, where the display screen, the camera, the memory, and the processor are connected through a bus: The memory stores a computer program, and the processor performs the following operations according to the computer program: Obtain a low-dynamic-range (LDR) image of any real scene collected by the camera and display it on the display screen; Input the LDR image into a high-dynamic-range (HDR) panoramic image generation model to obtain an HDR panoramic image, where the HDR panoramic image contains the true spherical Gaussian intensity distribution of the ambient light; Input the HDR panoramic image into a lighting spherical harmonic feature generation model to obtain the original spherical Gaussian intensity distribution of the lighting spherical harmonic feature generation model, and determine the KL divergence loss value between the true spherical Gaussian intensity distribution included in the HDR panoramic image and the original spherical Gaussian intensity distribution. When the KL divergence loss value is greater than or equal to a preset divergence threshold, fine-tune the lighting spherical harmonic feature generation model until the KL divergence loss value is less than the preset divergence threshold, and use the spherical Gaussian intensity distribution of the HDR panoramic image output by the lighting spherical harmonic feature generation model after the fine-tuning stops as the target spherical Gaussian intensity distribution; wherein, each spherical Gaussian intensity distribution is the brightness intensity distribution of multiple points on the sphere; Determine the position of the light source in the virtual scene according to the true spherical Gaussian intensity distribution and the target spherical Gaussian intensity distribution; Superimpose and display a virtual object on the LDR image, and draw a shadow for the virtual object superimposed and displayed on the display screen according to the position of the light source.
6. The AR device according to claim 5, wherein, The processor determines the position of the light source in the virtual scene according to the true spherical Gaussian intensity distribution and the target spherical Gaussian intensity distribution. The specific operation is as follows: Determine the point with the maximum brightness in the true spherical Gaussian intensity distribution, and determine the point with the maximum brightness in the target spherical Gaussian intensity distribution; Determine whether the position deviation between the two points with the maximum brightness is within a preset value range; If so, determine the three-dimensional coordinates of the light source in the virtual scene according to the point with the maximum brightness in the target spherical Gaussian intensity distribution.
7. The AR device according to claim 6, characterized in that, The processor determines the three-dimensional coordinates of the light source in the virtual scene according to the point with the maximum brightness in the target spherical Gaussian intensity distribution. The specific operation is as follows: Initialize the initial positions of the points in the target spherical Gaussian intensity distribution; Obtain the point with the maximum brightness in the target spherical Gaussian intensity distribution according to the brightness of the points at the respective initial positions in the target spherical Gaussian intensity distribution; Use the brightness of the point with the maximum brightness in the target spherical Gaussian intensity distribution to determine the target position of the center point of the respective points; Determine the three-dimensional coordinates corresponding to the target position of the center point and use the three-dimensional coordinates as the position of the light source in the virtual scene.
8. The AR device according to claim 6, characterized in that, When the position deviation between the two points with the maximum brightness is not within the preset value range, the processor further performs the following operations: Reverse fine-tune the parameters of the lighting spherical harmonic feature generation model until the position deviation between the point with the maximum brightness in the true spherical Gaussian intensity distribution and the point with the maximum brightness in the target spherical Gaussian intensity distribution is within the preset value range.
Citation Information
Patent Citations
Illumination rendering method and device, computer equipment and storage medium
CN112927341A
Method and device for drawing virtual object shadow based on light source position
CN113538704A