Light source prediction method and electronic device
By determining the light source prediction point in the image and processing it using a regression encoder and a fully connected layer, the direction of the light source is accurately determined, solving the problem of inaccurate light source direction in existing technologies and improving the realism and accuracy of augmented reality.
Patent Information
- Application Number
- CN202210031003.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-12
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-01-12
AI Technical Summary
Existing technologies cannot accurately determine the direction of light sources in images, resulting in unrealistic augmented reality effects.
By identifying multiple light source prediction points in the image to be predicted, a regression encoder is used for feature extraction. Multiple fully connected layers are used to determine the target spherical color information of the light source prediction points, and finally the direction of the light source is determined.
Accurately determining the direction of the light source improves the realism and precision of augmented reality effects, especially when adding virtual objects with reflective capabilities, resulting in a more realistic effect.
Smart Images

Figure CN116486045B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of augmented reality technology, and in particular to a light source prediction method and electronic device. Background Technology
[0002] Augmented Reality (AR) is a new technology developed on the basis of virtual reality. It is a technology that enhances a user's perception of the real world by using information provided by a computer system. It overlays computer-generated virtual objects, scenes, or system prompts onto images obtained by photographing real scenes, thereby "enhancing" reality.
[0003] Therefore, for AR technology, in order to make virtual objects more closely resemble real-world scenes, it is usually necessary to determine the light source information in the image. Accurate lighting information can achieve more realistic augmented reality effects.
[0004] Currently, existing technologies for analyzing images and estimating the direction of light sources cannot accurately determine the direction of the light source. Summary of the Invention
[0005] To address the problems in the prior art, embodiments of this application provide a light source prediction method and electronic device that can accurately predict the direction of a light source.
[0006] In a first aspect, embodiments of this application provide a light source prediction method, the method comprising:
[0007] Determine multiple light source prediction points in the image to be predicted;
[0008] A regression encoder is used to extract features from the image to be predicted, resulting in feature vectors of the image to be predicted at multiple stages; wherein the number of light source prediction points corresponding to the feature vectors of the image to be predicted at each stage is different.
[0009] Based on the feature vector of the image to be predicted at each stage, the target sphere color information corresponding to the multiple light source prediction points is determined by using the fully connected layer corresponding to each stage.
[0010] The direction of the light source is determined based on the target spherical color information corresponding to the multiple light source prediction points.
[0011] This application provides a light source prediction method. It involves determining multiple light source prediction points in an image to be predicted, and using a regression encoder to extract features from the image, obtaining feature vectors for multiple stages. Based on the feature vectors of each stage, a fully connected layer corresponding to the feature vectors of each stage is used to determine the target spherical color information corresponding to the multiple light source prediction points. Finally, based on the target spherical color information corresponding to the multiple light source prediction points, the light source direction is determined. By performing multi-stage processing on the image to be predicted, accurate target spherical color information corresponding to multiple light source prediction points representing the content of the image to be predicted is obtained, thus accurately determining the light source direction.
[0012] In one possible implementation, based on the feature vector of the image to be predicted at each stage, a fully connected layer corresponding to each stage is used to determine the target spherical color information corresponding to the multiple light source prediction points, including:
[0013] The feature vector of the image to be predicted in the first stage is input into the first fully connected layer to obtain the spherical color information output by the first fully connected layer; the first fully connected layer is the fully connected layer corresponding to the first stage.
[0014] For each regression fully connected layer, the image feature vector to be predicted at the stage corresponding to the regression fully connected layer, and the spherical color information output by the previous fully connected layer are input into the regression fully connected layer to obtain the spherical color information output by the regression fully connected layer; the regression fully connected layer refers to the fully connected layer other than the first fully connected layer.
[0015] The spherical color information output by the last fully connected regression layer is used as the target spherical color information.
[0016] In one possible implementation, determining the light source direction based on the target spherical color information corresponding to the plurality of light source prediction points includes:
[0017] Based on the target spherical color information corresponding to the multiple light source prediction points, the light source intensity information of the multiple light source prediction points is determined respectively;
[0018] The direction of the light source is determined based on the light source intensity information and the target spherical color information of the multiple light source prediction points; or, the direction of the light source is determined based on the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage and the light source intensity information of the multiple light source prediction points.
[0019] In one possible implementation, the target spherical color information includes three color information parameter values; the step of determining the light source intensity information of the multiple light source prediction points based on the target spherical color information corresponding to the multiple light source prediction points includes:
[0020] For any given light source prediction point, the light source intensity information of that prediction point is determined as follows:
[0021] The light source intensity information of the light source prediction point is obtained by multiplying the three color information parameter values corresponding to the light source prediction point by the corresponding preset weight values.
[0022] In one possible implementation, determining the light source direction based on the light source intensity information of the plurality of light source prediction points and the target spherical color information includes:
[0023] Based on the target spherical color information of the multiple light source prediction points, fit the spherical texture corresponding to the multiple light source prediction points respectively;
[0024] The spherical textures corresponding to the multiple light source prediction points are stitched together to obtain the first target spherical texture.
[0025] Based on the first target spherical texture and the light intensity information of the multiple light source prediction points, the position of the first target light source is determined; the position of the first target light source is the position with the maximum light intensity information on the first target sphere corresponding to the first target spherical texture.
[0026] The direction from the position of the first target light source to the center of the first target sphere is taken as the direction of the light source.
[0027] In one possible implementation, after determining the direction of the light source, the method further includes:
[0028] Obtain the first virtual object;
[0029] Based on the light source direction, the first virtual object is added to the first target spherical texture to determine the first virtual object to be added;
[0030] The first virtual object to be added is added to the image to be predicted to obtain the first augmented reality image.
[0031] In one possible implementation, determining the light source direction based on the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage and the light source intensity information of the multiple light source prediction points includes:
[0032] Using the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage, the stage spherical texture corresponding to the spherical color information output by each fully connected layer is determined respectively.
[0033] The image to be predicted is input into the enhancement encoder to determine the enhancement feature vector;
[0034] Based on the enhanced feature vector and the stage spherical texture corresponding to the spherical color information output by each fully connected layer, a spherical convolution module is used to perform a spherical convolution operation to obtain the second target spherical texture.
[0035] Based on the second target spherical texture and the light intensity information of the multiple light source prediction points, the position of the second target light source is determined; the position of the second target light source is the position with the maximum light intensity information on the second target sphere corresponding to the second target spherical texture.
[0036] The direction from the position of the second target light source to the center of the second target sphere is taken as the light source direction.
[0037] In the above method, the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage can be used to determine the stage spherical map corresponding to the spherical color information output by each fully connected layer. The image to be predicted is then input into the enhancement encoder to determine the enhancement feature vector. A spherical convolution module is then used to perform spherical convolution processing on the spherical map of each stage and the enhancement feature vector to obtain the second target spherical map. The second target spherical map obtained in this way better conforms to the characteristics of a sphere, is more realistic, has higher accuracy, and includes more high-frequency image information. The location of the light source with the highest intensity information is more accurately determined in the second target spherical map, and the direction of the light source is more accurately determined using the second target spherical map.
[0038] In one possible implementation, after determining the direction of the light source, the method further includes:
[0039] A second virtual object is added to the target spherical texture. Based on the light source direction, a second virtual object to be added is determined. The second virtual object is a virtual object with reflective capabilities. The second virtual object to be added is a virtual object whose surface exhibits a reflective effect after the second virtual object is added to the target spherical texture.
[0040] The second virtual object to be added is added to the image to be predicted to obtain the second augmented reality image.
[0041] In the above method, since the second spherical texture is more realistic, has higher precision, and includes more high-frequency image information, placing the second virtual object with reflective capabilities on the second spherical texture, based on the target spherical information of multiple light source prediction points, can reflect some scene image information contained in the second spherical texture onto the second virtual object, presenting a more realistic effect. The resulting second virtual object to be added is also more realistic and specific. Placing the second virtual object to be added in the image to be predicted can provide users with a better user experience.
[0042] In one possible implementation, the training process for the regression encoder and multiple fully connected layers is as follows:
[0043] Obtain an image dataset; wherein the image dataset includes multiple image samples to be trained; each image sample to be trained has a corresponding image information label;
[0044] The regression network model is iteratively trained based on the image dataset; the regression network model includes the regression encoder and the multiple fully connected layers; wherein, one iteration of training includes:
[0045] Image samples to be trained are extracted from the image dataset and input into the regression network model;
[0046] A regression encoder is used to extract features from the image samples to be trained, resulting in image sample feature vectors at multiple stages;
[0047] Based on the image sample feature vector of each stage, the spherical color information of each stage is obtained by using the fully connected layer corresponding to the image sample feature vector of each stage.
[0048] Feature extraction is performed based on the image information labels corresponding to the image samples to be trained, and the spherical color information labels corresponding to each stage are obtained.
[0049] The first loss value is determined based on the spherical color information of the last stage and the corresponding spherical color information label;
[0050] Based on the spherical color information and corresponding spherical color information labels of the stages other than the last stage, multiple second loss values are determined respectively;
[0051] Based on the first loss value and the plurality of second loss values, the network parameters of the regression network model are adjusted until the training termination condition is met, thus obtaining the trained regression network model.
[0052] Secondly, embodiments of this application provide a light source prediction device, the device comprising:
[0053] The first determining unit is used to determine multiple light source prediction points in the image to be predicted;
[0054] The feature extraction unit is used to extract features from the image to be predicted using a regression encoder to obtain feature vectors of the image to be predicted at multiple stages; wherein the number of light source prediction points corresponding to the feature vectors of the image to be predicted at each stage is different.
[0055] The second determining unit is used to determine the target spherical color information corresponding to the multiple light source prediction points based on the feature vector of the image to be predicted at each stage and using the fully connected layer corresponding to the feature vector of the image to be predicted at each stage.
[0056] The prediction unit is used to determine the direction of the light source based on the target spherical color information corresponding to the multiple light source prediction points.
[0057] In one possible implementation, the second determining unit is further configured to:
[0058] The feature vector of the image to be predicted in the first stage is input into the first fully connected layer to obtain the spherical color information output by the first fully connected layer; the first fully connected layer is the fully connected layer corresponding to the first stage.
[0059] For each regression fully connected layer, the image feature vector to be predicted at the stage corresponding to the regression fully connected layer, and the spherical color information output by the previous fully connected layer are input into the regression fully connected layer to obtain the spherical color information output by the regression fully connected layer; the regression fully connected layer refers to the fully connected layer other than the first fully connected layer.
[0060] The spherical color information output by the last fully connected regression layer is used as the target spherical color information.
[0061] In one possible implementation, the prediction unit is further configured to:
[0062] Based on the target spherical color information corresponding to the multiple light source prediction points, the light source intensity information of the multiple light source prediction points is determined respectively;
[0063] The direction of the light source is determined based on the light source intensity information and the target spherical color information of the multiple light source prediction points; or, the direction of the light source is determined based on the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage and the light source intensity information of the multiple light source prediction points.
[0064] In one possible implementation, the prediction unit is further configured to:
[0065] For any given light source prediction point, the light source intensity information of that prediction point is determined as follows:
[0066] The light source intensity information of the light source prediction point is obtained by multiplying the three color information parameter values corresponding to the light source prediction point by the corresponding preset weight values.
[0067] In one possible implementation, the prediction unit is further configured to:
[0068] Based on the target spherical color information of the multiple light source prediction points, fit the spherical texture corresponding to the multiple light source prediction points respectively;
[0069] The spherical textures corresponding to the multiple light source prediction points are stitched together to obtain the first target spherical texture.
[0070] Based on the first target spherical texture and the light intensity information of the multiple light source prediction points, the position of the first target light source is determined; the position of the first target light source is the position with the maximum light intensity information on the first target sphere corresponding to the first target spherical texture.
[0071] The direction from the position of the first target light source to the center of the first target sphere is taken as the direction of the light source.
[0072] In one possible implementation, the light source prediction device further includes:
[0073] Imaging unit, used to acquire the first virtual object;
[0074] Based on the light source direction, the first virtual object is added to the first target spherical texture to determine the first virtual object to be added;
[0075] The first virtual object to be added is added to the image to be predicted to obtain the first augmented reality image.
[0076] In one possible implementation, the prediction unit is further configured to:
[0077] Using the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage, the stage spherical texture corresponding to the spherical color information output by each fully connected layer is determined respectively.
[0078] The image to be predicted is input into the enhancement encoder to determine the enhancement feature vector;
[0079] Based on the enhanced feature vector and the stage spherical texture corresponding to the spherical color information output by each fully connected layer, a spherical convolution module is used to perform a spherical convolution operation to obtain the second target spherical texture.
[0080] Based on the second target spherical texture and the light intensity information of the multiple light source prediction points, the position of the second target light source is determined; the position of the second target light source is the position with the maximum light intensity information on the second target sphere corresponding to the second target spherical texture.
[0081] The direction from the position of the second target light source to the center of the second target sphere is taken as the light source direction.
[0082] In one possible implementation, the imaging unit is further configured to:
[0083] A second virtual object is added to the second target spherical texture. Based on the light source direction, a second virtual object to be added is determined. The second virtual object is a virtual object with reflective capabilities. The second virtual object to be added is the virtual object whose surface exhibits a reflective effect after the second virtual object is added to the second target spherical texture.
[0084] The second virtual object to be added is added to the image to be predicted to obtain the second augmented reality image.
[0085] In one possible implementation, the light source prediction device further includes:
[0086] A training unit is used to acquire an image dataset; wherein the image dataset includes multiple image samples to be trained; each image sample to be trained has a corresponding image information label;
[0087] The regression network model is iteratively trained based on the image dataset; the regression network model includes the regression encoder and the multiple fully connected layers; wherein, one iteration of training includes:
[0088] Image samples to be trained are extracted from the image dataset and input into the regression network model;
[0089] A regression encoder is used to extract features from the image samples to be trained, resulting in image sample feature vectors at multiple stages;
[0090] Based on the image sample feature vector of each stage, the spherical color information of each stage is obtained by using the fully connected layer corresponding to the image sample feature vector of each stage.
[0091] Feature extraction is performed based on the image information labels corresponding to the image samples to be trained, and the spherical color information labels corresponding to each stage are obtained.
[0092] The first loss value is determined based on the spherical color information of the last stage and the corresponding spherical color information label;
[0093] Based on the spherical color information and corresponding spherical color information labels of the stages other than the last stage, multiple second loss values are determined respectively;
[0094] Based on the first loss value and the plurality of second loss values, the network parameters of the regression network model are adjusted until the training termination condition is met, thus obtaining the trained regression network model.
[0095] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and when the computer program is executed by the processor, it implements the steps of any of the light source prediction methods in the first aspect described above.
[0096] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the light source prediction methods in the first aspect described above.
[0097] The technical effects achieved by the second to fourth aspects provided in the embodiments of this application are the same as those achieved by the light source prediction method provided in the first aspect, and will not be repeated here. Attached Figure Description
[0098] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0099] Figure 1 A schematic flowchart illustrating a light source prediction method provided in an embodiment of this application;
[0100] Figure 2 This application provides a schematic diagram of the process of a regression encoder extracting features from an image to be predicted.
[0101] Figure 3 A schematic diagram illustrating the process of a regression network model for processing an image to be predicted, provided in an embodiment of this application.
[0102] Figure 4 A schematic diagram illustrating a process for generating a second target spherical texture using an enhanced network, provided as an embodiment of this application;
[0103] Figure 5 A flowchart illustrating the training process of the regression network module provided in this embodiment of the application;
[0104] Figure 6 This is a schematic diagram of the structure of a light source prediction device provided in an embodiment of this application;
[0105] Figure 7 This is a schematic diagram of another light source prediction device provided in an embodiment of this application;
[0106] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0107] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0108] It should be noted that the terms "comprising" and "having" and their variations used in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or device.
[0109] To address the problem in existing technologies where the direction of light sources cannot be accurately determined when adding virtual objects to images captured from real-world scenes, this application provides a light source prediction method. This method involves identifying multiple light source prediction points in the image to be predicted and using a regression encoder to extract features from the image, obtaining feature vectors for multiple stages. Based on the feature vectors for each stage, a fully connected layer corresponding to each stage is used to determine the target spherical color information corresponding to the multiple light source prediction points. Finally, the direction of the light source is determined based on the target spherical color information corresponding to the multiple light source prediction points. By performing feature extraction at multiple stages on the image to be predicted and using fully connected layers at multiple stages to determine the target spherical color information corresponding to the multiple light source prediction points, the direction of the light source can be accurately determined.
[0110] Figure 1 This illustration shows a flowchart of a light source prediction method provided in an embodiment of this application, applied to electronic devices. For example... Figure 1 As shown, the light source prediction method provided in this application includes the following steps:
[0111] Step S101: Determine multiple light source prediction points in the image to be predicted.
[0112] In one possible embodiment, after acquiring the image to be predicted, multiple light source prediction points can be selected from the image. Different light source prediction points can be selected for different images to be predicted. The image to be predicted can be a low dynamic range imaging (LDR) image, such as an image taken with a mobile phone.
[0113] For example, for a given image to be predicted, 128 light source prediction points can be selected from the image.
[0114] Step S102: Use a regression encoder to extract features from the image to be predicted, and obtain feature vectors of the image to be predicted at multiple stages.
[0115] In one possible embodiment, the image to be predicted is input into a regression network model. This regression network model includes a regression encoder and a regression decoder. The regression decoder includes fully connected layers with multiple stages.
[0116] Figure 2 This diagram illustrates the workflow of a regression encoder when extracting features from an image to be predicted. Figure 2 As shown.
[0117] The image to be predicted is input into a regression encoder for feature extraction, which yields feature vectors for the image at multiple stages. The number of light source prediction points corresponding to the feature vectors at each stage varies.
[0118] For example, a regression encoder may include multiple dense blocks, each of which may be a convolutional layer. Taking three dense blocks as an example, namely a first dense block, a second dense block, and a third dense block, the regression encoder may use a DenseNet-121 encoder.
[0119] The input to the first dense block is the image to be predicted. Feature encoding, or feature extraction, is performed on the image to be predicted, and the output is the first dense feature vector.
[0120] The input to the second dense block is the first dense feature vector, and the output is the second dense feature vector.
[0121] The third dense block takes the first and second dense feature vectors as input and outputs the third dense feature vector.
[0122] The first dense feature vector, the second dense feature vector, and the third dense feature vector need to be superimposed to obtain the final feature vector of the image to be predicted.
[0123] As can be seen, the dense feature vector output by each preceding dense block will serve as the input to the subsequent dense block. Therefore, the final output feature vector of the image to be predicted already has multi-level features.
[0124] The number of light source prediction points corresponding to the image feature vector to be predicted varies in each stage. In the first stage, the image feature vector to be predicted can correspond to 8 light source prediction points, resulting in an image feature vector to be predicted corresponding to 8 light source prediction points. In the second stage, it can correspond to an image feature vector to be predicted corresponding to 32 light source prediction points, and in the third stage, it can correspond to an image feature vector to be predicted corresponding to 128 light source prediction points.
[0125] The ultimate goal is to obtain the target sphere color information corresponding to the 128 light source prediction points. However, directly processing the 128 light source prediction points and determining the target sphere color information corresponding to them is very difficult. Therefore, a pyramid-structured regression decoder is chosen to process the image to be predicted, which means processing the feature vector of the image to be predicted at each stage through multiple fully connected layers.
[0126] Alternatively, a small number of light source prediction points can be selected, such as 8 light source prediction points. If only 8 light source prediction points are selected, the spherical color information of these 8 light source prediction points can be obtained directly. However, since the number of light source prediction points is small, 8 light source prediction points cannot accurately predict the direction of the light source in the entire image to be predicted. Therefore, in order to obtain accurate results, a larger number of light source prediction points are usually selected.
[0127] Step S103: Based on the feature vector of the image to be predicted at each stage, use the fully connected layer corresponding to the feature vector of the image to be predicted at each stage to determine the target sphere color information corresponding to multiple light source prediction points respectively.
[0128] In one possible embodiment, the processing of the image to be predicted by the regression network model is described in detail. Figure 3 This diagram illustrates a detailed flowchart of a regression network model processing an image to be predicted. Figure 3 As shown.
[0129] Take a fully connected layer in a regression network model, which includes three stages, as an example.
[0130] It should be noted that, for the image feature vector to be predicted at each stage, before inputting the image feature vector to be predicted at each stage into the fully connected layer of the corresponding stage, the image feature vector to be predicted at each stage can also be input into the global average pooling layer for processing, so as to obtain the 1024-dimensional pooling feature vector corresponding to the image feature vector to be predicted at each stage.
[0131] Therefore, the regression network model processes the image to be predicted as follows: The feature vector of the image to be predicted in the first stage is input into the global average pooling layer in the first stage to obtain a 1024-dimensional pooling feature vector. This 1024-dimensional pooling feature vector is then input into the first fully connected layer to obtain the spherical color information output by the first fully connected layer, which can be called the first-stage spherical color information. The spherical color information output by the first fully connected layer consists of the spherical color information corresponding to the 8 light source prediction points.
[0132] The first fully connected layer is the fully connected layer corresponding to the feature vector of the image to be predicted in the first stage. The spherical color information can be referred to as the amplitude parameters of the spherical Gaussian, i.e., the color system (RGB) information, and can be simply called color information.
[0133] All fully connected layers other than the first fully connected layer can be called regressive fully connected layers. Therefore, the second-stage fully connected layer can be called the first regressive fully connected layer, and the third-stage fully connected layer can be called the second regressive fully connected layer.
[0134] The second-stage processing involves inputting the image feature vector to be predicted into the global average pooling layer to obtain a 1024-dimensional pooling feature vector. The resulting spherical color information from the first fully connected layer, along with the 1024-dimensional pooling feature vector from the second stage, is then input into the first regression fully connected layer to obtain the spherical color information output by the first regression fully connected layer, which can be called the second-stage spherical color information. The spherical color information output by the first regression fully connected layer corresponds to the spherical color information of the 32 predicted light source points.
[0135] It should be noted that the process of inputting the spherical color information output by the first fully connected layer and the 1024-dimensional pooling feature vector from the second stage into the first regression fully connected layer involves first performing a concatenation operation between the spherical color information output by the first fully connected layer and the 1024-dimensional pooling feature vector from the second stage to achieve feature fusion, resulting in a 1048-dimensional fused feature vector, which is then input into the first regression fully connected layer.
[0136] The second-stage processing utilizes the spherical color information corresponding to the eight predicted light source points obtained in the first stage. Using a pyramidal network structure, coarse processing can be performed in the first stage, followed by refined processing in the second stage, achieving a coarse-to-fine structure. This continuously refines the spherical color information of the predicted light source points, generating the desired high-precision result. Furthermore, locating the eight predicted light source points in the first stage helps the fully connected layers in the second stage to pre-locate the spherical color information of these eight points from the first stage.
[0137] The third stage of processing involves inputting the spherical color information output from the first fully connected regression layer and the 1024-dimensional pooled feature vector from the third stage into the second fully connected regression layer. This results in the spherical color information output by the second fully connected regression layer, which can be called the third-stage spherical color information. The spherical color information output by the second fully connected regression layer corresponds to the spherical color information of the 128 predicted light source points. Since the second fully connected regression layer is the last one, the spherical color information output by the second fully connected regression layer, corresponding to the 128 predicted light source points, is the target spherical color information.
[0138] Each light source prediction point has corresponding spherical color information, which means that the final result is 128 sets of spherical color information.
[0139] Step S104: Determine the direction of the light source based on the target spherical color information corresponding to multiple light source prediction points.
[0140] In one possible embodiment, the light intensity information of multiple light source prediction points can be determined by using the target spherical color information corresponding to multiple light source prediction points, and then the light source direction can be determined based on the light intensity information of multiple light source prediction points and the target spherical color information.
[0141] For example, after obtaining the target spherical color information of 128 light source prediction points, the 128 light source prediction points can be processed. The target spherical color information corresponding to any light source prediction point includes three color information parameter values. The light source intensity information of the light source prediction point can be determined by the three color information parameter values of each light source prediction point.
[0142] The light source intensity information of the light source prediction point is obtained by multiplying the three color information parameter values corresponding to the light source prediction point by the corresponding preset weight values.
[0143] The light source intensity information for each light source prediction point can be obtained by multiplying the corresponding three color information parameter values by weights of (0.3, 0.59, 0.11).
[0144] For each spherical color information of a light source prediction point, a spherical texture can be fitted. Each spherical texture includes the light intensity information of a light source prediction point. In any spherical texture, since there is a specific light intensity information for a light source prediction point, and no light intensity information is calculated for other positions, it is equivalent to only having light intensity information at one position on this spherical texture. Therefore, the light source prediction point on this spherical texture can be used as the light source position.
[0145] The spherical texture can be fitted using the Gaussian distribution formula, and the light intensity information at any location in the spherical texture can be calculated. The Gaussian distribution formula is shown below:
[0146] G(v;α,λ,μ)=αexp(λ(μ·v-1))
[0147] α represents the spherical color information of the light source prediction point; λ represents the bandwidth of the spherical Gaussian, λ∈(0,+∞); μ represents the direction vector from the light source prediction point to the center of the sphere; v represents a point at any position on the sphere; G(v;α,λ,μ) represents the light source intensity information of a point at any position on the sphere.
[0148] In the above process, the bandwidth of the spherical Gaussian is a preset value, which can be 33.45.
[0149] Using the Gaussian distribution formula for a sphere, the light source intensity at any point on the sphere can be determined. Essentially, by using the known light source intensity information from a predicted point as a reference, the light source intensity at any point on the sphere can be determined. The closer the light source is to the predicted point, the greater the light source intensity; the farther away, the smaller the light source intensity.
[0150] By fitting the spherical color information of 128 light source prediction points into 128 spherical texture maps using the above method, the first target spherical texture map can be obtained by stitching together the 128 spherical texture maps.
[0151] The first target spherical texture contains light intensity information corresponding to 128 light source prediction points, which can determine the light intensity information of any point on the target sphere. The position with the highest light intensity information on the target sphere is the target light source position to be determined.
[0152] We can assume the location of the light source on the target sphere, where the light source intensity is highest, is the light source position. Then, the direction from the target light source position to the center of the target sphere is the predicted light source direction. This light source direction is the light source direction in the first target sphere texture, which is also the light source direction of the image to be predicted.
[0153] After determining the direction of the light source, a first virtual object can be added to the image to be predicted. Once this first virtual object is obtained, it can be added to the fitted first target spherical texture. The effect of the first virtual object in the first target spherical texture can be determined. Since the position, direction, and intensity of the light source in the first target spherical texture are known, the effect of the first virtual object in the first target spherical texture can be determined well, which is the first virtual object to be added.
[0154] After obtaining the first virtual object to be added, it is directly added to the image to be predicted to obtain the first augmented reality image. During user interaction, the interface image presented to the user will not show the first target spherical texture. After the user selects the first virtual object, the first virtual object to be added is directly determined, and the effect displayed on the user interface is the first augmented reality image.
[0155] In the above method, since the image to be predicted is generally an image obtained by taking pictures with a regular camera, directly placing the first virtual object into the image to be predicted will not produce good results. Therefore, by fitting a first target spherical texture, a first target spherical texture containing high-frequency image information is obtained. Since the first target spherical texture is a spherical image, placing the first virtual object into the first target spherical texture allows for prediction of the direction of the light source based on multiple visual angles determined in the first target spherical texture, resulting in more accurate results.
[0156] The advantage of directly predicting the target spherical color information of multiple light source prediction points using the above spherical Gaussian distribution formula is that the light intensity information of a single light source prediction point can be represented by a small number of parameters, i.e., representing a single light source. However, the disadvantage is that it can only represent the light intensity information of a single light source prediction point and cannot fully represent the detailed information of the light source, such as whether the light source is an incandescent lamp or a light bulb. In order to make the illumination estimation effect more accurate, we use an augmentation network to supplement and correct the predicted light source intensity information and light source direction to a certain extent.
[0157] Alternatively, the second target spherical map can be determined in the following manner, and the direction of the light source can be predicted using the second target spherical map determined in the following process.
[0158] Figure 4 This diagram illustrates a flowchart for generating a second target spherical texture using an augmentation network, which includes an augmentation encoder and a spherical convolution module. The specific process is as follows:
[0159] The image to be predicted is input into the augmentation encoder to determine the augmentation feature vector. The spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage is used to determine the stage spherical texture corresponding to the spherical color information output by each fully connected layer. Based on the augmentation feature vector and the stage spherical texture corresponding to the spherical color information output by each fully connected layer, a spherical convolution module is used to perform spherical convolution operation to obtain the second target spherical texture.
[0160] For example, using the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage, if the regression network model includes three stages of fully connected layers, then the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage is as follows: the spherical color information output by the first fully connected layer, which is the spherical color information corresponding to the 8 light source prediction points; the spherical color information output by the first regression fully connected layer, which is the spherical color information corresponding to the 32 light source prediction points; and the spherical color information output by the second regression fully connected layer, which is the spherical color information corresponding to the 128 light source prediction points, which is the target spherical color information of the 128 light source prediction points. This serves as auxiliary information in the process of determining the second target spherical texture in this manner.
[0161] By using the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage, the stage spherical texture corresponding to the spherical color information output by each fully connected layer is determined.
[0162] That is, the first spherical texture is determined based on the first stage spherical color information, the second spherical texture is determined based on the second stage spherical color information, and the third spherical texture is determined based on the third stage spherical color information.
[0163] If the regression network model includes three fully connected layers, then the spherical convolution module will include three spherical convolution layers.
[0164] The detailed process of using the spherical convolution module for spherical convolution operation is as follows: The first-stage spherical texture and enhanced feature vector, which are fitted with the spherical color information corresponding to the 8 light source prediction points, are used as the input of the first spherical convolution layer in the spherical convolution module. In the first spherical convolution layer, the spherical convolution operation is performed to determine the first spherical convolution feature map.
[0165] The input to the second spherical convolutional layer is the second-stage spherical texture map fitted with the spherical color information corresponding to the 32 light source prediction points and the first spherical convolutional feature map. The output of the second spherical convolutional layer is the second spherical convolutional feature map.
[0166] The input to the third spherical convolutional layer is the second spherical convolutional feature map and the third-stage spherical texture fitted with the spherical color information corresponding to 128 light source prediction points. The output of the third spherical convolutional layer is the third spherical convolutional feature map, which is used as the second target spherical texture.
[0167] The above method resamples and filters the input stage spherical maps at each stage using spherical convolution. Because spherical convolution is used, the resulting second target spherical map avoids the drawbacks of ordinary convolution operations. A spherical map is a parametrically constructed image and will have deviations in different dimensions. With ordinary convolution operations, because the kernel translation is only in a two-dimensional plane, some areas of the second target spherical map will appear larger or smaller than their actual size. Therefore, ordinary convolution operations will produce deviations in the second target spherical map across different dimensions.
[0168] The above method determines the second target spherical texture. The enhancement network that determines the second target spherical texture can be used as a generator, and a discriminator can be added to judge the final second target spherical texture. The Patch-GAN network can be used.
[0169] The Patch-GAN network used in the above process includes a generator and a discriminator, with the generator being the augmentation network.
[0170] The training process of the Patch-GAN network can utilize three loss functions, which supervise the training process. The first can be the feature loss, used to train the discriminator. The second can be the generator's loss function, i.e., the augmentation network's loss function, which supervises the augmentation network by judging the cosine similarity between the generated second target spherical texture and the actual spherical texture. The third is the Patch-GAN network's loss function, which can be any loss function available for the Patch-GAN network.
[0171] The second target spherical texture determined in this way has a better effect and is closer to the desired effect in use. The second target spherical texture determined in this way is more realistic, has higher precision, and includes more high-frequency image information.
[0172] Therefore, if a second virtual object is added to the second target spherical texture determined in this way, and if the second virtual object is a virtual object with reflective capabilities, such as a mirror-like virtual object, then the second virtual object can reflect some real scene images in the second target spherical texture. The second virtual object to be added can be determined, and the second virtual object to be added can be added to the image to be predicted to obtain the second augmented reality image.
[0173] The second virtual object to be added is the virtual object whose surface exhibits a reflective effect after the second virtual object is added to the second target spherical texture.
[0174] The process of determining the direction of the light source also utilizes a regression network model, which can be implemented using methods such as... Figure 5 The training method shown yields, as follows: Figure 5 As shown, the training process of a regression network model includes the following steps:
[0175] Step S501: Obtain the image dataset.
[0176] The image dataset includes multiple training image samples, each with a corresponding image information label.
[0177] The training image samples can be low dynamic range (LDR) imaging images, such as images taken with a mobile phone.
[0178] Image information tags can be images that include high-frequency image information, such as high dynamic range imaging (HDR), which can be understood as panoramic images that include high-frequency image information.
[0179] Step S502: Extract image samples to be trained from the image dataset and input them into the regression network model. Use a regression encoder to extract features from the image samples to be trained, and obtain image sample feature vectors at multiple stages.
[0180] Specifically, the process of using a regression encoder to extract features from the training image samples and obtaining image sample feature vectors at multiple stages is similar to the process in step S102, and the specific process will not be repeated here.
[0181] Step S503: Based on the image sample feature vector of each stage, use the fully connected layer corresponding to the image sample feature vector of each stage to obtain the spherical color information of each stage.
[0182] Specifically, the process of obtaining the spherical color information of each stage by using the fully connected layer corresponding to the image sample feature vector of each stage is similar to the process in step S103, and the specific process will not be repeated here.
[0183] Step S504: Extract features based on the image information labels corresponding to the image samples to be trained, and obtain the spherical color information labels corresponding to each stage.
[0184] In one possible embodiment, due to the high dynamic range imaging type of the image information tag, a spherical color information tag with good performance can be obtained by feature extraction of the image information tag.
[0185] For example, if the fully connected layer corresponding to the image sample feature vector in the first stage obtains spherical color information corresponding to 8 light source prediction points, then 8 points can be determined at the same position in the image information label, and the corresponding spherical color information label can be determined based on the 8 points in the image information label.
[0186] Step S505: Determine the first loss value based on the spherical color information and the corresponding spherical color information label of the last stage, and determine multiple second loss values based on the spherical color information and the corresponding spherical color information label of the other stages besides the last stage.
[0187] In one possible embodiment, the loss value is obtained through two loss functions. The first loss value is determined based on the spherical color information of the last stage and the corresponding spherical color information label. The first loss function is as follows:
[0188]
[0189] Among them, L 2-masked Let A be the first loss function; A represents the spherical color information obtained from the last stage through the regression network model; A gt M represents the spherical color information label corresponding to the spherical color information in the last stage; M represents the mask matrix corresponding to the spherical texture fitted by the spherical color information label. This is represented as L2 paradigm.
[0190] M is obtained by multiplying the three color information parameter values in the spherical color information label with weights of (0.3, 0.59, 0.11) and then summing them to obtain the light source intensity information corresponding to 128 spherical maps. Then, according to the size of the light source intensity information, the mask value of the index corresponding to the first 5% of the light source intensity information is set to 1, and the mask value of the index corresponding to the other light source intensity information is set to 0.
[0191] For the spherical color information in all stages except the last stage, multiple second loss values are calculated. In other stages, a second loss value is calculated for the spherical color information of each stage and the corresponding spherical color information label. The formula for calculating the second loss value is the conventional L2 loss function.
[0192] After obtaining multiple second loss values and the first loss value, proceed to step S506.
[0193] Step S506: Determine whether the first loss value meets the first preset value, and whether the multiple second loss values meet the second preset values. If they meet the requirements, proceed to step S508; if they do not meet the requirements, proceed to step S507.
[0194] Step S507: Adjust the network parameters of the regression network module based on the first loss value and multiple second loss values.
[0195] Step S508: Use the current network parameters as the network parameters of the regression network model to obtain the trained regression network model.
[0196] In one possible embodiment, network parameters can be adjusted using a first loss value and multiple second loss values. The spherical color information and corresponding spherical color information labels at each stage can characterize the high-frequency information in the corresponding image. The first loss value and multiple second loss values are calculated based on the spherical color information and corresponding spherical color information labels at each stage. Therefore, the spherical color information output by the trained regression network model can, to a certain extent, characterize the high-frequency image information contained in the image information labels. Finally, the spherical color information output by the trained regression network model can be well fitted to the spherical texture to obtain either a first target spherical texture or a second spherical texture, accurately predicting the light source direction.
[0197] It can also accurately generate the virtual objects to be added corresponding to the virtual objects, enabling end users to add virtual objects during the process and achieve better visual effects for the virtual objects to be added.
[0198] Based on the same concept, embodiments of this application also provide a light source prediction device. Figure 6 A schematic diagram of a light source prediction device according to an embodiment of this application is shown. This light source prediction device is applied to electronic devices, such as… Figure 6 As shown, the light source prediction device includes:
[0199] The first determining unit 601 is used to determine multiple light source prediction points in the image to be predicted;
[0200] The feature extraction unit 602 is used to extract features from the image to be predicted using a regression encoder to obtain feature vectors of the image to be predicted at multiple stages; wherein the number of light source prediction points corresponding to the feature vectors of the image to be predicted at each stage is different.
[0201] The second determining unit 603 is used to determine the target sphere color information corresponding to multiple light source prediction points based on the feature vector of the image to be predicted at each stage and using the fully connected layer corresponding to the feature vector of the image to be predicted at each stage.
[0202] The prediction unit 604 is used to determine the direction of the light source based on the target spherical color information corresponding to multiple light source prediction points.
[0203] In one possible implementation, the second determining unit 603 is further configured to:
[0204] The feature vector of the image to be predicted in the first stage is input into the first fully connected layer to obtain the spherical color information output by the first fully connected layer; the first fully connected layer is the fully connected layer corresponding to the first stage.
[0205] For each regression fully connected layer, the feature vector of the image to be predicted at the corresponding stage of the regression fully connected layer, and the spherical color information output by the previous fully connected layer are input into the regression fully connected layer to obtain the spherical color information output by the regression fully connected layer; the regression fully connected layer refers to the fully connected layer other than the first fully connected layer.
[0206] The spherical color information output by the last fully connected regression layer is used as the target spherical color information.
[0207] In one possible implementation, the prediction unit 604 is further configured to:
[0208] Based on the target spherical color information corresponding to multiple light source prediction points, the light source intensity information of multiple light source prediction points is determined respectively;
[0209] The direction of the light source is determined based on the light source intensity information of multiple light source prediction points and the target spherical color information; or, the direction of the light source is determined based on the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage and the light source intensity information of multiple light source prediction points.
[0210] In one possible implementation, the prediction unit 604 is further configured to:
[0211] For any given light source prediction point, the light source intensity information is determined as follows:
[0212] The light source intensity information of the light source prediction point is obtained by multiplying the three color information parameter values corresponding to the light source prediction point by the corresponding preset weight values.
[0213] In one possible implementation, the prediction unit 604 is further configured to:
[0214] Based on the target spherical color information of multiple light source prediction points, fit the spherical texture corresponding to the multiple light source prediction points respectively.
[0215] The spherical textures corresponding to multiple light source prediction points are stitched together to obtain the first target spherical texture.
[0216] Based on the first target spherical texture and the light intensity information of multiple light source prediction points, the position of the first target light source is determined; the position of the first target light source is the position with the maximum light intensity information on the first target sphere corresponding to the first target spherical texture.
[0217] The direction from the position of the first target light source to the center of the first target sphere is taken as the direction of the light source.
[0218] In one possible implementation, the prediction unit 604 is further configured to:
[0219] Using the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage, the stage spherical texture corresponding to the spherical color information output by each fully connected layer is determined respectively.
[0220] The image to be predicted is input into the enhancement encoder to determine the enhancement feature vector;
[0221] Based on the enhanced feature vector and the stage spherical texture corresponding to the spherical color information output by each fully connected layer, a spherical convolution module is used to perform spherical convolution operation to obtain the second target spherical texture.
[0222] Based on the second target spherical texture and the light intensity information of multiple light source prediction points, the position of the second target light source is determined; the position of the second target light source is the position with the maximum light intensity information on the second target sphere corresponding to the second target spherical texture.
[0223] The direction from the position of the second target light source to the center of the second target sphere is taken as the direction of the light source.
[0224] In one possible implementation, Figure 7 This application illustrates another light source prediction device provided in an embodiment of the present application, which further includes:
[0225] Training unit 701 is used to acquire an image dataset; wherein, the image dataset includes multiple image samples to be trained; each image sample to be trained has a corresponding image information label;
[0226] The regression network model is iteratively trained based on an image dataset; the regression network model includes a regression encoder and multiple fully connected layers; one iteration of training includes:
[0227] Extract image samples to be trained from the image dataset and input them into the regression network model;
[0228] A regression encoder is used to extract features from the training image samples to obtain image sample feature vectors at multiple stages.
[0229] Based on the image sample feature vector of each stage, the fully connected layer corresponding to the image sample feature vector of each stage is used to obtain the spherical color information of each stage.
[0230] Feature extraction is performed based on the image information labels corresponding to the image samples to be trained, and the spherical color information labels corresponding to each stage are obtained.
[0231] The first loss value is determined based on the spherical color information of the last stage and the corresponding spherical color information label;
[0232] Based on the spherical color information and corresponding spherical color information labels of the stages other than the last stage, multiple second loss values are determined respectively;
[0233] Based on the first loss value and multiple second loss values, adjust the network parameters of the regression network model until the training termination condition is met, thus obtaining the trained regression network model.
[0234] In one possible implementation, the light source prediction device further includes an imaging unit 702, for:
[0235] Obtain the first virtual object;
[0236] Based on the direction of the light source, the first virtual object is added to the first target spherical texture to determine the first virtual object to be added;
[0237] The first virtual object to be added is added to the image to be predicted to obtain the first augmented reality image.
[0238] In one possible implementation, the imaging unit 702 is further configured to:
[0239] A second virtual object is added to the second target spherical texture. Based on the direction of the light source, the second virtual object to be added is determined. The second virtual object is a virtual object with reflective capabilities. The second virtual object to be added is the virtual object whose surface exhibits a reflective effect after the second virtual object is added to the second target spherical texture.
[0240] The second virtual object to be added is added to the image to be predicted to obtain the second augmented reality image.
[0241] This application also provides an electronic device that can be used to execute the process of a light source prediction method. This electronic device can be a server or a terminal device. The electronic device includes at least a memory for storing data and a processor. The processor for data processing can be a microprocessor, CPU, GPU (Graphics Processing Unit), DSP, or FPGA. The memory stores operation instructions, which can be computer-executable code, to implement the various steps in the light source prediction method of this application.
[0242] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 8 As shown, the electronic device 800 includes a memory 801, a processor 802, a data acquisition module 803, and a bus 804. The memory 801, processor 802, and data acquisition module 803 are all connected via the bus 804, which is used for data transmission between the memory 801, processor 802, and data acquisition module 803.
[0243] The memory 801 can be used to store software programs and modules. The processor 802 executes various functional applications and data processing of the electronic device 800 by running the software programs and modules stored in the memory 801, such as the light source prediction method provided in the embodiments of this application. The memory 801 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs of at least one application, etc.; the data storage area may store data created according to the use of the electronic device 800, etc. In addition, the memory 801 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0244] The processor 802 is the control center of the electronic device 800. It connects various parts of the electronic device 800 via the bus 804 and various interfaces and lines. It executes various functions and processes data of the electronic device 800 by running or executing software programs and / or modules stored in the memory 801, and by calling data stored in the memory 801. Optionally, the processor 802 may include one or more processing units, such as a CPU, GPU (Graphics Processing Unit), or digital processing unit.
[0245] The data acquisition module 803 is used to acquire data, such as virtual objects, image datasets, and images to be predicted.
[0246] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, can be used to implement the light source prediction method described in any embodiment of this application.
[0247] In some possible implementations, various aspects of the light source prediction method provided in this application can also be implemented as a program product comprising program code. When the program product is run on a computer device, the program code causes the computer device to perform the steps of the light source prediction method according to the various exemplary embodiments of this application described above. For example, the computer device can perform actions such as... Figure 1 The flowchart of the light source prediction method in steps S101 to S104 is shown.
[0248] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0249] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0250] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0251] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0252] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A light source prediction method, characterized in that, The method includes: Determine multiple light source prediction points in the image to be predicted; A regression encoder is used to extract features from the image to be predicted, resulting in feature vectors of the image to be predicted at multiple stages; wherein the number of light source prediction points corresponding to the feature vectors of the image to be predicted at each stage is different. Based on the feature vector of the image to be predicted at each stage, the target sphere color information corresponding to the multiple light source prediction points is determined by using the fully connected layer corresponding to each stage. Based on the target spherical color information corresponding to the multiple light source prediction points, the light source intensity information of the multiple light source prediction points is determined respectively; Based on the light source intensity information and target spherical color information of the multiple light source prediction points, the light source direction is determined; Alternatively, the direction of the light source can be determined based on the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage, as well as the light source intensity information of the multiple light source prediction points.
2. The method according to claim 1, characterized in that, The method, based on the feature vector of the image to be predicted at each stage, uses the fully connected layer corresponding to each stage to determine the target spherical color information corresponding to the multiple light source prediction points, including: The feature vector of the image to be predicted in the first stage is input into the first fully connected layer to obtain the spherical color information output by the first fully connected layer; the first fully connected layer is the fully connected layer corresponding to the first stage. For each regression fully connected layer, the image feature vector to be predicted at the stage corresponding to the regression fully connected layer, and the spherical color information output by the previous fully connected layer are input into the regression fully connected layer to obtain the spherical color information output by the regression fully connected layer; the regression fully connected layer refers to the fully connected layer other than the first fully connected layer. The spherical color information output by the last fully connected regression layer is used as the target spherical color information.
3. The method according to claim 1, characterized in that, The target spherical color information includes three color information parameter values; based on the target spherical color information corresponding to the multiple light source prediction points, the light source intensity information of each of the multiple light source prediction points is determined. include: For any given light source prediction point, the light source intensity information of that prediction point is determined as follows: The light source intensity information of the light source prediction point is obtained by multiplying the three color information parameter values corresponding to the light source prediction point by the corresponding preset weight values.
4. The method according to claim 1, characterized in that, The step of determining the direction of the light source based on the light source intensity information and the target spherical color information of the multiple light source prediction points includes: Based on the target spherical color information of the multiple light source prediction points, fit the spherical texture corresponding to the multiple light source prediction points respectively; The spherical textures corresponding to the multiple light source prediction points are stitched together to obtain the first target spherical texture. Based on the first target spherical texture and the light intensity information of the multiple light source prediction points, the position of the first target light source is determined; the position of the first target light source is the position with the maximum light intensity information on the first target sphere corresponding to the first target spherical texture. The direction from the position of the first target light source to the center of the first target sphere is taken as the direction of the light source.
5. The method according to claim 4, characterized in that, After determining the direction of the light source, the method further includes: Obtain the first virtual object; Based on the light source direction, the first virtual object is added to the first target spherical texture to determine the first virtual object to be added; The first virtual object to be added is added to the image to be predicted to obtain the first augmented reality image.
6. The method according to claim 1, characterized in that, The determination of the light source direction based on the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage and the light source intensity information of the multiple light source prediction points includes: Using the spherical color information output by the fully connected layer corresponding to the feature vector of the image to be predicted at each stage, the stage spherical texture corresponding to the spherical color information output by each fully connected layer is determined respectively. The image to be predicted is input into the enhancement encoder to determine the enhancement feature vector; Based on the enhanced feature vector and the stage spherical texture corresponding to the spherical color information output by each fully connected layer, a spherical convolution module is used to perform spherical convolution operation to obtain the second target spherical texture. Based on the second target spherical texture and the light intensity information of the multiple light source prediction points, the position of the second target light source is determined; the position of the second target light source is the position with the maximum light intensity information on the second target sphere corresponding to the second target spherical texture. The direction from the position of the second target light source to the center of the second target sphere is taken as the light source direction.
7. The method according to claim 6, characterized in that, After determining the direction of the light source, the method further includes: A second virtual object is added to the second target spherical texture. Based on the light source direction, a second virtual object to be added is determined. The second virtual object is a virtual object with reflective capabilities. The second virtual object to be added is the virtual object whose surface exhibits a reflective effect after the second virtual object is added to the second target spherical texture. The second virtual object to be added is added to the image to be predicted to obtain the second augmented reality image.
8. The method according to claim 1, characterized in that, The training process for the regression encoder and multiple fully connected layers is as follows: Obtain an image dataset; wherein the image dataset includes multiple image samples to be trained; each image sample to be trained has a corresponding image information label; The regression network model is iteratively trained based on the image dataset; the regression network model includes the regression encoder and the multiple fully connected layers; wherein, one iteration of training includes: Image samples to be trained are extracted from the image dataset and input into the regression network model; A regression encoder is used to extract features from the image samples to be trained, resulting in image sample feature vectors at multiple stages; Based on the image sample feature vector of each stage, the spherical color information of each stage is obtained by using the fully connected layer corresponding to the image sample feature vector of each stage. Feature extraction is performed based on the image information labels corresponding to the image samples to be trained, and the spherical color information labels corresponding to each stage are obtained. The first loss value is determined based on the spherical color information of the last stage and the corresponding spherical color information label; Based on the spherical color information and corresponding spherical color information labels of the stages other than the last stage, multiple second loss values are determined respectively; Based on the first loss value and the plurality of second loss values, the network parameters of the regression network model are adjusted until the training termination condition is met, thus obtaining the trained regression network model.
9. An electronic device, characterized in that, The system includes a memory and a processor, wherein a computer program is mounted on the memory and can run on the processor, and when the computer program is executed by the processor, implements the method of any one of claims 1 to 8.
Citation Information
Patent Citations
Indoor scene illumination estimation model, method and device, storage medium and rendering method
CN110910486A
Multi-light-source prediction method
CN112819787A