Illumination estimation image generation method based on elimination of low-activation-rate neurons

By designing a light estimation image generation network model, eliminating low activation rate neurons, and combining with multiple loss function training, the artifact problem in light estimation is solved, and a more realistic and accurate light estimation generation image is achieved.

CN120374824APending Publication Date: 2025-07-25BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510450141.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the existing light estimation methods, low activation rate neurons lead to artifacts in generated images, affecting the authenticity of the image and the accuracy of the light estimation results.

Method used

The light estimation image generation network model is designed, and by counting the neuron activation rate and eliminating the low activation rate neurons, it is trained in combination with the generation of adversarial networks, mean square variance loss, learning-perceived image block similarity loss and image text comparison loss to generate more realistic light estimation images.

Benefits of technology

Improve the authenticity and accuracy of the generated image in the lighting estimation, reduce artifacts, and improve the similarity between the generated image and the real scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374824A_ABST
    Figure CN120374824A_ABST
Patent Text Reader

Abstract

The invention relates to an illumination estimation image generation method based on elimination of low-activation-rate neurons. At present, most research on illumination estimation ignores the influence of the neuron activation rate, and low-activation-rate neurons may cause artifacts and illumination estimation result errors on a generated image. According to the method, the illumination estimation image generation network model is designed for artifact restoration in the generated image, so that low-activation-rate neurons are eliminated, and the authenticity of the generated image and the accuracy of an illumination estimation result can be improved. The method comprises the following steps: firstly, designing a local image illumination estimator; then generating images for multiple times, counting the activation rate of neurons and recording neurons with low activation rate; and finally, when the image is generated, the low-activation-rate neurons are eliminated. Experiments prove that the FID of the generated image is 104.1, and compared with other existing illumination estimation image generation models, the performance is better; compared with the prior art that the low-activation-rate neurons are removed, the SSIM and the PSNR are improved, the error indexes such as Angular Error are reduced, and the overall performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and image processing, and particularly to a method for generating illumination estimation images based on eliminating neurons with low activation rates( Figure 1 ). Specifically, by eliminating neurons with low activation rates in the synthesizer network model, the generation of unwanted artifacts in the generated images is eliminated, ultimately improving the quality of the generated images, making the illumination estimation results more accurate and realistic, and more in line with actual application scenarios. Background Art

[0002] As a branch of 3D scene reconstruction, illumination estimation is a long-term research topic in the fields of computer vision and computer graphics. It extracts illumination information from existing real photos or rendering results, and then uses this information to reconstruct a new scene with changed viewing angles and light source positions, and finally renders a new result. Virtual clothing fitting, augmented reality navigation, movie post-production, virtual military exercises, video game design, etc. all require inserting virtual objects into physical scenes.

[0003] In recent years, the wide application of deep learning technology in the computer field has brought breakthrough progress to illumination estimation technology. However, due to factors such as variable illumination conditions, the influence of reflection and refraction, and shadow problems, this remains a very challenging topic. Illumination estimation methods can be divided into methods based on physical modeling, methods based on statistics, and methods based on learning.

[0004] The complexity of the illumination problem makes it more appropriate to perform illumination estimation using data-driven priors learned from millions of images. Methods based on learning need to learn statistical models from training data. Methods based on learning can be divided into methods based on traditional machine learning and methods based on deep learning. Gamut mapping is a common method based on traditional machine learning. Methods based on deep learning can be divided into regression-based methods and generation-based methods.

[0005] Methods based on physical modeling mainly model through the physical interaction between light and the object surface, model the propagation, reflection, refraction, etc. of light in the scene, and analyze the relationship between illumination and the scene by establishing a physical model to obtain illumination-related parameters. There are two main modeling methods, one is based on diffuse reflection, and the other is to consider both diffuse reflection and specular reflection. Physical modeling is only applicable to ideal conditions and is easily affected by noise in actual modeling.

[0006] Statistical methods mainly utilize the statistical characteristics of object surface colors under different lighting conditions to model and statistically analyze the color distribution under common lighting conditions, thereby inferring the lighting parameters in the input image. These methods include the gray world assumption method, the gray edge assumption method, and the white patch assumption method. The gray world assumption method assumes that the average value of reflection differences in the scene is colorless; the gray edge assumption method assumes that the average reflection difference in the edge region of the scene is colorless and performs better than the gray world assumption method when considering local information; the white patch assumption method assumes that the object with the maximum reflectivity in the scene is a perfect white reflector.

[0007] Traditional machine learning models can be trained to utilize low-level image features to predict lighting parameters. Compared with traditional machine learning models, convolutional neural network models and their variants in deep learning, such as convolutional neural network-based regression, bias correction, and fast Fourier integration models, can provide better prediction accuracy.

[0008] Deep learning models can effectively learn features from massive data. The advantages of deep learning models include not requiring manual feature design, less manual intervention, high accuracy, being able to process unstructured data, and being suitable for solving underconstrained problems. The application of deep learning solves the problem of some difficult-to-model parameters in light estimation, such as sensor parameters, scene lighting, object material properties, etc.; the semantic segmentation technology in deep learning can present the recognition of indoor scene lighting types in a more intuitive way. The models used in deep learning-based methods in the field of light estimation include, but are not limited to, Convolutional Neural Networks, Graph Convolution Neural Networks, Generative Adversarial Networks, and Recurrent Neural Network.

[0009] The deep learning framework takes images with a limited field of view and low dynamic range as input and outputs panoramic images with a high dynamic range. Commonly used sensing devices often lack panoramic cameras and only have cameras with a limited field of view, where the limited field of view is usually less than 100 degrees. The process from a limited field of view to a panoramic view is called out-of-field extrapolation, and the process from low dynamic range to high dynamic range is called inverse gamut mapping. The model estimates the global information and lighting information that it believes to be the most probable given local information to achieve the two operations of out-of-field extrapolation and inverse gamut mapping.

[0010] Learning-based methods are most relevant to the work of the present invention, and their relevant research backgrounds are introduced in turn as follows.

[0011] (1) Traditional machine learning methods

[0012] Color gamut mapping is a classic learning method. It processes by selecting a scene that conforms to the lighting conditions and collecting as many object surfaces under the given lighting conditions as data samples. Then, for an input image with a known light source position, it estimates the lighting parameters from the canonical color gamut to generate a rich and realistic scene. Its main steps include: first, selecting a standard color card; then, using a tongue diagram and a convex hull algorithm to compare and match the color information of the captured scene with the color gamut of the standard color card; finally, estimating the lighting conditions of the scene based on this matching relationship.

[0013] This method relies on a color gamut mapping algorithm, a convex hull algorithm, and a color temperature matching algorithm. Among them, the most important is the color gamut mapping algorithm. Secondly, the convex hull algorithm is used to process the shape of the color gamut. Finally, the color temperature matching algorithm is used to obtain the color temperature matched by the estimated color gamut.

[0014] (2) Deep learning methods

[0015] Deep learning-based methods mainly include regression-based methods and generation-based methods.

[0016] Regression-based methods. Obtain the lighting representation of the light source, and then use the lighting representation to render the scene. The lighting representation is predefined by the researcher and then regression learning is performed using the model. A better lighting representation can better fit the information of the input data. A better fitting effect represents a more accurate lighting estimate, and a more accurate lighting estimate achieves a better rendering effect.

[0017] There are many forms of lighting representation. It can be represented by spherical harmonics, spherical gaussian, needlets, multi scale volume of implicit feature, full neural light fields, image based lighting, etc., or it can be represented by a parameterized light source, such as position, direction, area, intensity, color, etc. It can even combine an environment map with a parameterized light source: such as normal, roughness, reflectivity, opacity of the shadow texture, etc.

[0018] As shown in Equation (1), spherical harmonics are a special class of functions in mathematics. The calculation of spherical harmonic coefficients is based on the idea of probability theory, that is, using the "finite" to estimate the "infinite", and is usually used to represent physical phenomena in three-dimensional space, such as electromagnetic waves, elastic waves, etc. The linear combination of spherical harmonics is still a spherical harmonic, and spherical harmonics have orthogonality and completeness. Spherical harmonics are suitable for low-frequency illumination, but as the order of spherical harmonics increases, false shadows will appear in the generated images. Spherical Gaussian distribution is a Gaussian distribution defined in spherical coordinates, and spherical Gaussian distribution has isotropy and rotational invariance. Wavelet transform is a mathematical transform and a multi-scale analysis method of Fourier transform, which can decompose image signals into components of different frequencies and positions, including Discrete Wavelet Transform, Continuous Wavelet Transform (CWT), Muti-Resolution Analysi, Lifting Wavelet Transform, and Wavelet Packet Transformation. Needlets is a kind of spherical wavelet transform.

[0019]

[0020] In addition to parameterizing the light source direction, the above light representation methods usually need to find the position of the light distribution, that is, the center of the light distribution. Since the distance between indoor light sources and objects in the scene is relatively close, the high-dynamic-range light radiation field in the indoor scene changes rapidly, and the light near the light source is very different from the light in the center of the room, which brings certain challenges to capturing the light source position.

[0021] Common methods for capturing the light source position include: directly capturing by inserting a light probe; capturing by using a light classifier; assuming the number and position of light sources; separating the light source by selecting a part of pixels with the highest brightness value; extracting the light source by methods such as median cut, variance minimization, and peak finding.

[0022] Generation-based methods. This method does not require researchers to pre-define the light representation, but directly generates a high-dynamic-range panoramic image by skipping this step. Generation-based methods usually decompose light estimation into several sub-tasks. Light estimation can be decomposed into three tasks: geometric estimation, scene completion, and low-dynamic-range to high-dynamic-range conversion, or it can be decomposed into three tasks: learning the light distribution, learning the light intensity, and learning the environment term. Geometric estimation can also be omitted, and light estimation can be decomposed into two tasks: panoramic completion of the scene from a limited-field-of-view image and conversion of the low-dynamic-range panoramic image to a high-dynamic-range panoramic image.

[0023] In the process of generating a high-dynamic-range panoramic image from an image with a low dynamic range and a limited field of view, the geometric structure of most of the invisible scenes and the illumination of most of the areas not in the current view must be inferred. Illumination estimation is a severely under-constrained problem, so extensive context information is required, and it is necessary to enhance the extraction rate of information from the data and the utilization rate of the extracted information.

[0024] In 2017, Gardner et al. proposed a simple transformation to address the spatial differences between the observation position and the rendering position. However, the proposed transformation does not use depth information, which may lead to image distortion; moreover, the model outputs only one estimate for one image and lacks the ability to handle spatially varying illumination information. (Reference: Gardner, M.-A., Sunkavalli, K., Yumer, E., Shen, X., Gambaretto, E., Gagné, C., and Lalonde, J.-F. Learning to Predict Indoor Illumination from a Single Image, ACM Transactions on Graphics (SIGGRAPH Asia), 9(4), 2017)

[0025] In 2018, Wang et al. proposed a Global Illumination-Aware and Detail-preserving Network (GLADNet). The input image is resized, and the resized image is input into a network with an encoder-decoder architecture to generate the global prior knowledge of the illumination. Based on the global prior knowledge and the original input image, a convolutional neural network is used for detail reconstruction. (Reference: Wang W, Wei C, Yang W, et al. Gladnet: Low-light enhancement network with global awareness [C] / / 2018 13th IEEE international conference on automatic face&gesture recognition (FG 2018). IEEE, 2018: 751-755.)

[0026] In 2019, Gardner et al. proposed a method for estimating illumination from a single indoor image. This method considered the local nature of indoor lighting, manually performed depth annotation for the dataset, and predicted light source parameters: position, area, intensity, and color. These parameters can be used to render the shadows of inserted objects, producing realistic rendering results. However, the prerequisite for estimating the illumination parameters is the need for 3D scene reconstruction, and there are also problems such as the inability to model directional light or focused beams. As shown in Equation (2), the spherical Gaussian function was used. (Reference document: Gardner M A, Hold-Geoffroy Y, Sunkavalli K, et al. Deep parametric indoor lighting estimation[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019:7175-7183)

[0027]

[0028] In 2021, Zhang et al. proposed a new end-to-end model based on the illumination estimation of a single outdoor image. The brightness channel of the sky region was used in the input to enhance the extracted image features, and pruning and quantization were used to compress the network, thus greatly reducing the network parameters and storage space, while only suffering a slight loss in accuracy. (Reference document: Zhang K, Li X, Jin X, et al. Outdoor illumination estimation via all convolutional neural networks[J]. Computers&Electrical Engineering, 2021, 90(4):106987.)

[0029] The appearance of outdoor scene objects changes with the variation of light and seasons. To save the expensive cost of re - performing 3D modeling, in 2021, Xiong et al. proposed using the Deep Shadow Network (DSNet) to estimate the light source position, which can accurately estimate the illumination of the input image, enhance the sun position estimation using an illumination - based dataset, and a method for rendering according to the light source position. Strategies for runtime rendering and optimization were also discussed. (Reference: Xiong Y, Chen H, Wang J, et al. DSNet: deep shadow network for illumination estimation[C] / / 2021 IEEE Virtual Reality and 3D User Interfaces (VR). IEEE, 2021:179 - 187.)

[0030] In 2022, Wang et al. proposed StyleLight, which estimates high - dynamic - range panoramic illumination from low - dynamic - range limited - field - of - view images and realizes two operations of light addition and deletion. Any operation of editing light can be converted into operations of deleting light and adding light. The dual - coupled StyleGAN network integrates the two steps of inferring the panorama from a limited field of view and inferring the high - dynamic - range from the low - dynamic - range in a unified framework, thus greatly improving the performance and editability of light estimation. (Reference: Wang G, Yang Y, Loy C C, et al. Stylelight: HDR panorama generation for lighting estimation and editing[C] / / European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022:477 - 492.)

[0031] In 2022, Huan et al. proposed a multi-task end-to-end neural network GeoRec for geometrically enhanced semantic 3D reconstruction of RGB-D indoor scenes, constructed a geometry extractor that can learn geometrically enhanced representations from depth data to improve the estimation accuracy of scene layout, sensor position, and object edges, and also introduced a new object mesh generator to enhance the reconstruction robustness of GeoRec to indoor occlusions through geometrically enhanced implicit shape embedding. (Reference: Huan L, Zheng X, Gong J. GeoRec: Geometry-enhanced semantic 3D reconstruction of RGB-D indoor scenes[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2022, 186: 301-314.)

[0032] In 2022, Zhang et al. proposed a new dataset containing real and rendered images, as well as a new cascaded network to study semantic segmentation in weakly lit indoor environments. The network decomposes weakly lit images into two components, illumination and reflectance, and then establishes a multi-task learning scheme by two branches. The reflectance recovery branch learns to reduce noise and recover reflectance information, and the semantic segmentation branch learns to segment the reflectance map. The convolutional neural network features output by the two tasks are concatenated together to improve segmentation accuracy by embedding illumination-invariant features. (Reference: Zhang N, Nex F, Kerle N, et al. LISU: Low-light indoor scene understanding with joint learning of reflectance restoration[J]. ISPRS journal of photogrammetry and remote sensing, 2022, 183: 470-481.)

[0033] In 2023, Bai et al. proposed the first graph learning-based strategy, an indoor light estimation framework DSGLight based on graph learning, which combines physical models and learning models, and is more reasonable and accurate in estimating direct and indirect ambient light. DSGLight directly constructs hundreds of uniformly distributed spherical Gaussian distributions on the indoor panorama, and assumes that each spherical Gaussian distribution has a fixed position. Each spherical Gaussian distribution represents the light and depth information of surrounding nodes. A graph convolutional neural network is used as the model to train the light representation. The adjacency matrix of the graph convolutional neural network is constructed using the K-Nearest Neighbor clustering algorithm. By leveraging the prior knowledge of the non-Euclidean nature of the graph structure and the large number of potential relationships between nodes, specific indoor perception features are extracted. (Reference document: Bai J, Guo J, Wang C, et al. Deep graph learning for spatially-varying indoor lighting prediction[J]. Science China Information Sciences, 2023, 66(3): 132106.)

[0034] In 2024, Phongthawee et al. proposed a technique for estimating the light of a single image using a pre-trained diffusion model, which is achieved by rendering a chrome ball in the image and converting it into an environment map, and performs well in various scenarios. The noise addition principle of the diffusion model is shown in Equation (3). (Reference document: Phongthawee P, Chinchuthakun W, Sinsunthithet N, et al. Diffusion light: Light probes for free by painting a chrome ball[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024: 98-108.)

[0035]

[0036] In summary, most current generative light estimations improve the accuracy of light estimation by using a new generation model, but relatively little research has been done on the essential reasons for artifacts generated in light estimation, and relatively little research has been done on the neuron activation rate. Using neurons with a low activation rate will result in more artifacts in the generated images. The present invention mainly solves this problem, which can not only improve the quality of the generated images of light estimation, but also improve the accuracy of light estimation. Summary of the Invention

[0037] The present invention aims to overcome the influence of artifacts generated in light estimation on the authenticity of images. By counting the number of neuron activations, the activation rate of each neuron is obtained, and then the outputs of neurons with low activation rates are eliminated, thereby improving the authenticity of light estimation images and the accuracy of light estimation results, providing a research basis for image generation in the field of light estimation and other computer vision tasks.

[0038] The present invention mainly studies the problem of light estimation from the perspective of neuron activation rates. Neurons with high activation rates can generate a relatively complete panoramic scene. Neurons with low activation rates vary, but usually are responsible for generating the details of the image and cannot generate a relatively complete panoramic scene.

[0039] The features required for generating an image are highly correlated with neurons having high activation rates, and it can be considered an essential general component. During the process of generating an image, neurons with low activation rates play a role in enriching the types of images and can generate more diverse images. However, this diversity is often uncontrolled and may be the reason for generation failure and the occurrence of artifacts. Artifacts affect the authenticity of the generated image, and overexposed sources affect the details of the image. Therefore, neurons with low activation rates not only affect the authenticity of the generated image but also affect the accuracy of the light estimation result.

[0040] As Figure 2 shown, it is an image generated by neurons with high activation rates. As Figure 3 shown, it is an image generated by neurons with low activation rates. Neurons with high activation rates contain most of the information required in image generation, while neurons with low activation rates are prone to generating artifacts.

[0041] Based on the above considerations, the present invention proposes a method for generating light estimation images based on eliminating neurons with low activation rates to reduce the artifacts generated in light estimation generated images. The present invention designs a network model for generating light estimation images, and based on this model, neurons with low activation rates are eliminated. It has three core functions, namely, counting the activation rate of neurons, eliminating neurons with low activation rates, and generating light estimation images. This has strong practicality for subsequent research and development.

[0042] As Figure 4 shown, it is a light estimation generated image before eliminating neurons with low activation rates. The result of the present invention is as Figure 5 shown, with fewer artifacts and better authenticity.

[0043] Next, the main content of the present invention will be introduced in detail, specifically including the following steps:

[0044] Step 1: Design a network model for generating light estimation images

[0045] The generated light estimation method needs to learn to obtain the estimated light distribution from a random standard Gaussian distribution. Therefore, the present invention first designs a generative network model based on a generative adversarial network for generating light estimation images. The generative network includes two parts: a generator and a discriminator. The generator obtains a high-dynamic-range panoramic image by inputting information of a low-dynamic-range limited-field-of-view image (as shown in Figure 6 ), and the discriminator is responsible for judging whether the generated image is similar to the image of the real scene. The two perform adversarial training and learn from each other. Eventually, the generator can generate an image that infinitely approaches the real scene, thereby obtaining a relatively accurate light estimation result.

[0046] The generated light estimation is an image generation task. In order to improve the authenticity and editability of light estimation, the present invention proposes a generative network model, specifically as shown in Figure 1 . As shown in Figure 4 , it is the panoramic image generated by using the low-dynamic-range limited-field-of-view image as the input of the generative network model. Figure 6 The present invention uses a double-coupled generative adversarial network. As shown in Equation (4), the generator G hopes that the value of this formula is as small as possible, and the discriminators D and D' hope that the value of this formula is as large as possible. The two confront each other and finally converge.

[0047]

[0048]

[0049] The advantage of this method is that it can not only obtain the panoramic image of light estimation, but also edit the light information on the image, which means that the number, intensity, and position of light sources can be adjusted according to actual needs to meet the diverse needs of different scenarios and different applications, greatly expanding the applicability and practicality of this method in practical applications.

[0050] Finally, the light estimation generative network model used in the present invention is trained on the Laval Indoor HDR dataset to generate high-dynamic-range panoramic images with good authenticity, providing a basis for subsequent elimination of neurons with low activation rates. Experiments prove that the light estimation panoramic images generated by this method have good authenticity and are superior to most methods.

[0051] Step 2: Statistically analyze the neuron activation rate and the serial numbers of the eliminated neurons

[0052] First, in order to facilitate the description of the content of the present invention, some symbols are introduced and the activation rate of neurons is defined.

[0053] ​It represents the output of the idx-th neuron in the synthesized image of size res×(2*res). When it is considered that this neuron has a greater impact on the output image, that is, this neuron is in an activated state at this time; when it is considered that this neuron has a smaller impact on the output image, that is, this neuron is in a non-activated state at this time.

[0054] It represents the neuron numbers with neuron activation rate ≥ fre in the synthesized image of size res×(2*res). It represents the neuron numbers with neuron activation < fre in the synthesized image of size res×(2*res).

[0055] The second step of the present invention requires counting the activation rate of neurons during the process of generating images, that is, calculating all The generated images belong to independent experiments. Based on the Bernoulli's law of large numbers, let be the number of times the event occurs in n times of generating images, and p is the probability of the event occurring in each generation of images. Then for any positive number ∈, there is This means that when the number of generated images n is large enough, the probability that the deviation between the probability of the event occurring and the probability p is greater than any given positive number ∈ approaches zero, that is, convergence in probability.

[0056] Based on the Bernoulli's law of large numbers, a standard normal distribution is randomly generated as the input of the illumination estimation generation network model, and a forward propagation hook function is registered in each resolution synthesis layer of the generator. Under multiple independent repeated experiments, the activation situation of neurons is recorded to obtain the activation rate of each neuron. Finally, the neuron numbers with activation rate < fre are counted to obtain These data lay the foundation for subsequent experiments.

[0057] The advantages of this method for counting the activation rate of neurons are as follows: From the perspective of the reliability of the theoretical basis, the Bernoulli's law of large numbers is a solid theoretical basis, and the obtained neuron numbers have extremely high credibility; the forward propagation hook function can accurately capture the activation situation of each neuron, greatly avoiding the errors and subjectivity that may be brought by manual intervention, and improving the efficiency and accuracy of data collection; this method also has good scalability and adaptability, and is applicable to different illumination estimation generation network models.

[0058] Step 3: Design of the illumination estimation image generation network model based on removing neurons with low activation rates

[0059] Based on the design of the illumination estimation image generation model in the first two steps and the statistics of neuron activation rates and the serial numbers of neurons to be removed, the third step of the present invention designs an illumination estimation image generation network model based on removing neurons with low activation rates. It mainly includes the following three aspects:

[0060] 3.1 Design of the network structure

[0061] The schematic diagram of the overall network structure of the present invention is as shown in Figure 1 As shown, the main purpose of this network is to overcome the influence of artifacts generated in illumination estimation on the authenticity of the image, that is, to make the generated image have fewer artifacts and be closer to the image of the real scene. This network consists of two parts, one is the Mapping Network, and the other is the Synthesis Network. The Mapping Network receives the input from the standard Gaussian distribution, which has the characteristics of simplicity, standardization, and easy processing, but there is a large difference from the distribution of real images. The Mapping Network is responsible for decoupling the information of the standard Gaussian distribution function and converting it into an image feature distribution function; the Synthesis Network receives the image feature distribution function output by the Mapping Network, and this information contains illumination and scene geometry. The Synthesis Network is responsible for extracting this information from the image feature distribution function to synthesize an image highly close to the real scene.

[0062] First, based on the neuron activation rates and the serial numbers of neurons to be removed statistically in the second step, during the process of the Synthesis Network synthesizing images, when synthesizing images of different resolution layers, by each neuron serial number recorded in, these neurons are removed, that is, the output of the neurons is set to zero. In this way, the influence of neurons with activation rate <fre is cleared in this step.

[0063] As shown in Figure 5 As shown, it is the image generated after removing neurons with low activation rates. Compared with Figure 4 it can be observed that Figure 5 the artifacts on the ceiling are significantly reduced.

[0064] 3.2 Network joint training strategy

[0065] In the practice of generative adversarial networks, in addition to the network structure having a great impact on the generated images, another factor is the loss function. The loss function is directly related to the learning effect of the network and usually affects the quality of the generated images. The Mean Squared Error Loss uses the squared error. Larger errors will be amplified after the squaring operation, which can prevent small outliers from appearing in the images. The Learned Perceptual Image Patch Similarity Loss can measure the perceptual similarity between two images. Based on the perceptual characteristics of the human visual system, it can learn the semantic and structural information of the images and reduce the perceptual differences between the generated images and the real images. The image-text contrast loss can measure the distance between the image and the text in the feature space. Its core purpose is to enable the generative network model to learn the semantic alignment relationship between the image and the text, and make the distances between the corresponding images and texts in the feature space as close as possible.

[0066] Next, a detailed introduction to the hybrid training strategy adopted in the present invention is given. This strategy is composed of three parts: the Mean Squared Error Loss, the Learned Perceptual Image Patch Similarity Loss, and the image-text contrast loss. This is also one of the key factors for the present invention to achieve good results in addition to the network structure. Many studies have shown that joint training of these multiple losses often can achieve better results. Therefore, most current studies adopt the hybrid training strategy.

[0067] Mean Squared Error Loss. The Mean Squared Error Loss measures the difference between the generated image and the real image by calculating the average of the squares of the differences between the corresponding pixel values. Assume that the generated image is G, the target image is T, the size of the image is H×W×C, where H is the height, W is the width, and C is the number of channels. The pixel value of the generated image is (h, w, c), and the pixel value of the target image is (h, w, c). Then the formula for the Mean Squared Error Loss is shown in Equation (5).

[0068]

[0069] Learned Perceptual Image Patch Similarity Loss. First, use the pre-trained convolutional neural network AlexNet to extract the feature map of the image. To eliminate the influence of different scales, then normalize the extracted feature map. Finally, calculate the cosine similarity for each feature layer of the generated image and the real image, and perform a weighted summation operation on the cosine similarities of different layers to obtain the final value of the Learned Perceptual Image Patch Similarity Loss. Assume that x and y are two images to be compared, L is the set of layers in the pre-trained network used to calculate the similarity, w l is the weight of the l-th layer, and φ l (x) and φ l(y) are the feature maps of image x and image y at the l-th layer, and N l is the number of channels of the feature map at the l-th layer, <.,.> represents the inner product of vectors, and ∥.,.∥ represents the norm of vectors. Then, the learning perception image patch similarity loss formula is shown in Equation (6).

[0070]

[0071] Image-text contrast loss. The Clip model can map images and texts to the same feature space and calculate the similarity between them in this space. The similarity calculated by Clip is used as the image-text contrast loss, making the images generated by the model more relevant to the texts. Suppose the image is I and the text is T. The image encoder encodes the image I into a feature vector v img =E img (I), and the text encoder E text encodes the text T into a feature vector v text =E text (I). Then, the cosine similarity is used to calculate the similarity between these two feature vectors. The specific formula is shown in Equation (7).

[0072]

[0073] For a single image-text pair (I, T), the calculation of the specific loss value is shown in Equation (8).

[0074] L Clip =1 - sim(v img , v text ) (8)

[0075] In summary, during the training stage, the network is trained jointly with the above three losses. The total loss function can be expressed by Equation (9), where α, β, and γ are three different weight parameters used to control the influence weights of different losses on the network learning results.

[0076] L=αL MSE +βL PIPIS +γL Clip (9)

[0077] As Figure 7 shown, it is the image generated after inputting the text prompt "an image of a corridor".

[0078] Generally speaking, the main contribution of the present invention is to design a method for generating illumination estimation images based on eliminating neurons with low activation rates, which is a relatively classic topic in the field of computer vision. Starting from the perspective of neuron activation rates, considering the problem that neurons with low activation rates are more likely to generate artifacts, the present invention separately considers neurons with high activation rates and neurons with low activation rates, and to a certain extent overcomes the influence of neurons with low activation rates on the generated images of illumination estimation. The proposed scheme for eliminating neurons with low activation rates is not involved in other existing illumination estimation methods, and has certain innovation.

[0079] In addition, in order to cooperate with the proposed method for eliminating neurons with low activation rates, the present invention also proposes a hybrid training strategy based on mean squared error loss, learning perceptual image patch similarity loss, and image-text contrast loss, which is also one of the main contributions of the present invention. Experimental results prove that the proposed illumination estimation image generation scheme of the present invention is superior to most existing methods in terms of the authenticity of the generated images, and is also more applicable in practical scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 is the overall structural diagram of the illumination estimation generation network model proposed by the present invention, which is introduced in detail in step three.

[0081] Figure 2 is the image generated by neurons with high activation rates.

[0082] Figure 3 is the image generated by neurons with low activation rates.

[0083] Figure 4 is the image generated before eliminating neurons with low activation rates.

[0084] Figure 5 is the image generated after eliminating neurons with low activation rates.

[0085] Figure 6 is the low-dynamic-range limited field of view image.

[0086] Figure 7 is the image generated by the image-text contrast loss.

[0087] Figure 8 is the redrawn image by the diffusion model.

[0088] Figure 9 is the data augmentation image generated by the diffusion model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0089] The technical solutions, experimental methods, and test results of the present invention will be further described in detail below in conjunction with the drawings and specific experimental embodiments.

[0090] The present invention designs a topic of light estimation in the field of computer vision, and proposes a method for generating light estimation images based on removing neurons with low activation rates, which includes three main steps, namely, designing a network model for generating light estimation images, statistically calculating the activation rates of neurons and the serial numbers of neurons to be removed, and designing a network model for generating light estimation images based on removing neurons with low activation rates.

[0091] The experimental steps are specifically described below.

[0092] Step 1: Based on the generative adversarial network model, build a generative adversarial network that can generate a panoramic view output according to the input of limited field-of-view images.

[0093] Step 2: According to the law of large numbers of Bernoulli, statistically calculate the activation rates of all neurons in the generative network model, and obtain a set of serial numbers of neurons to be removed according to the activation rates.

[0094] Step 3: According to the serial numbers of neurons, remove neurons with low activation rates to obtain an optimized generated image, and simultaneously calculate the corresponding metrics.

[0095] The experimental situation and the obtained conclusions of the present invention are specifically described below.

[0096] (1) Light Estimation Dataset

[0097] The present invention uses the Laval Indoor HDR dataset, which is also used in many literatures related to light estimation. This dataset has a total of 2,233 images, including more than 1,000 images with different exposure times in different indoor scenes. These images have high resolution and high dynamic range, with a width of 7,768 pixels and a height of 3,384 pixels. Each picture contains approximately 66 MB of information. In addition, this dataset also provides the camera parameters and scene information of each image. This dataset was published in SIGGRAPHAsia in 2017 and is widely used in the fields of computer vision and computer graphics. In the present invention, 1,719 images are used for training and 289 images are used for testing.

[0098] (2) Experimental Details and Main Parameter Configurations

[0099] The present invention uses Stylegan2 as the backbone network, which consists of eight fully connected layers and eight upsampling convolutional layers.

[0100] For all input data, this method resets the size of all images to 256×512, sets the batch size of each batch to 16. Then, horizontal flipping operations are performed on the images with the vertical central axis as the axis of symmetry. Finally, the diffusion model is also used to expand the size of the training dataset. As Figure 8As shown in the figure, for the redrawn image, based on the diffusion model, 75% of the pixels were redrawn to generate the data augmentation image, as Figure 9 shown, which is used as the new training dataset.

[0101] During training, the network was trained for a total of 125,000 generations. Both the generator and the discriminator saw 2,000,000 pictures. The initial learning rate was set to 2×10 -3 , and the Adam optimizer was used with β1 = 0 and β2 = 0.99. In the design of the loss function, the corresponding weight parameters were selected through experiments, which were α = 0.5, β = 0.5, and γ = 0.1 in sequence.

[0102] (3) Evaluation Metrics

[0103] In the task of light estimation based on generation, during the test process, usually an image with a limited field of view is input, and a panoramic field of view image is generated as the output, and then the similarity between the real scene where the limited field of view image is located and the generated panoramic field of view image is calculated. To evaluate the performance of the light estimation method based on generation, the current approach is to calculate the Fréchet Inception Distance (FID), Peak Signal-to-Noise Ratio (PSNR), and Structural Similarity Index (SSIM).

[0104] The present invention calculates three metrics: FID, PSNR, and SSIM. Among them, the smaller the FID, the higher the similarity between the real image and the generated image; the higher the PSNR value, the smaller the difference between the processed image and the original image; the closer the SSIM value is to 1, the more similar the two images are, and the closer the value is to 0, the greater the difference between the two images.

[0105] (4) Experimental Results of Removing Neurons with Low Activation Rates

[0106] Based on the above evaluation metrics and experimental details, the present invention was tested on the Laval Indoor HDR dataset, and the corresponding experimental results were obtained.

[0107] As shown in Table 1, in this experiment, the method of the present invention was compared with other relatively advanced network models for FID. In the comparison, some methods that are most closely related to the present invention were selected. First, Stylegan2 was used as the backbone network, and the mean squared error loss, perceptual similarity loss, and image-text contrast loss were jointly used to train the network.

[0108] Table 1 Comparison Results of FID between the Present Invention and Other Methods (Laval Indoor HDR Dataset)

[0109] Method Name FID Gardner2017 307.5 Gardner2019 344.1 EMlight 263.8 StyleLight 137.7 Ours 104.1

[0110] As can be seen from Table 1, after removing neurons with an activation rate lower than 30% in the b4 and b8 modules of the images generated by the present invention, the FID is only 104.1, showing better performance compared to other light estimation image generation network models.

[0111] To further prove the effectiveness of the output of removing neurons with low activation rates proposed by the present invention, an ablation experiment was designed based on the neuron activation rate. First, Stylegan2 was also used as the backbone network, and the mean squared error loss, perceptual similarity loss of learning image patches, and image-text contrast loss were jointly used to train the network, and the test results obtained were used as the BaseLine.

[0112] Next, only neurons with a specific activation rate were used for image generation, and three metrics, namely FID, PSNR, and SSIM, were used for evaluation respectively.

[0113] The specific ablation experiment results are shown in Tables 2, 3, and 4.

[0114] Table 2 Comparison of ablation experiment results (SSIM↑)

[0115]

[0116] Table 3 Comparison of ablation experiment results (PSNR↑)

[0117]

[0118] Table 4 Comparison of ablation experiment results (FID↓)

[0119]

[0120] As can be seen from Tables 2, 3, and 4, after removing neurons with low activation rates, the SSIM increased significantly, the PSNR increased slightly, and the FID decreased slightly under specific parameters. This shows that removing neurons with activation rates can improve the authenticity of the light estimation generated images.

[0121] To further measure the impact of removing neurons with low activation rates on the accuracy of light estimation, three material spheres were rendered based on the generated panoramic images, and the three materials were Mirror Silver, MatteSilver, and Diffuse Grey respectively.

[0122] Render the images of real scenes with three kinds of material spheres as well, and calculate the error between the rendered spheres in the generated images and those in the real scene renderings, which is used to evaluate the accuracy of light estimation. It is mainly evaluated by the following four metrics, namely Angular Error, Root Mean Square Error (RMSE), Scale-Invariant Root Mean Square Error (SI-RMSE), and Normalized Root Mean Square Error (NRMSE).

[0123] As shown in Table 5, Table 6, and Table 7, they are the light estimation errors when rendering the images of specular silver, matte silver, and diffuse gray material spheres respectively.

[0124] Table 5 Light Estimation Error when Rendering the Image of Specular Silver Material Sphere

[0125]

[0126] Table 6 Light Estimation Error when Rendering the Image of Matte Silver Material Sphere

[0127]

[0128] Table 7 Light Estimation Error when Rendering the Image of Diffuse Gray Material Sphere

[0129]

[0130] It can be seen from Table 5, Table 6, and Table 7 that after removing the neurons with low activation rates, the light estimation error metrics Angular Error, RMSE, SI-RMSE, and NRMSE all decrease on the three different material spheres. This shows that removing the neurons with activation rates can improve the accuracy of light estimation.

[0131] The above comparative experiments and ablation experiment results show that the method of the present invention is superior to most of the existing methods in light estimation.

[0132] In summary, the present invention proposes a method for generating illumination estimation images based on eliminating neurons with low activation rates. By eliminating neurons with low activation rates in the synthesizer network model, the adverse artifacts of the generated images are eliminated, and finally the quality of the generated images is improved, the illumination estimation results are more accurate and realistic, and more in line with the actual application scenarios. The generated images of the proposed network have an FID of 104.1 on the LavalIndoor dataset, which is better than most existing methods; the optimal SSIM is only 0.4791, and the optimal PSNR is only 28.5784. At the same time, the present invention proposes a hybrid training strategy based on mean square error loss, learned perceptual image patch similarity loss, and image-text contrast loss, which can provide some references for future research on illumination estimation.

Claims

1. An image generation method for light estimation based on eliminating neurons with low activation rates, characterized in that: Through a light estimation network model, the network weight parameters are obtained by jointly training with three parts of losses, the serial numbers of the eliminated neurons are counted, and the neurons with low activation rates are eliminated; The implementation steps are as follows: S1. Design of the light estimation network: The light estimation network consists of two parts, one is a mapper (Mapping Network), and the other is a synthesizer (Synthesis Network); the mapper receives the input from the standard Gaussian distribution, and the mapper is responsible for decoupling the information of the standard Gaussian distribution function and converting it into an image feature distribution function; the synthesizer receives the image feature distribution function output by the mapper, and the synthesizer is responsible for extracting this information from the image feature distribution function to synthesize an image highly similar to the real scene; S2. Network joint training strategy: In the specific training strategy, a hybrid training strategy is adopted, which is composed of three parts: mean square error loss, learning perceptual image patch similarity loss, and image-text contrast loss; Mean square error loss: The mean square error loss measures the difference between the generated image and the real image by calculating the average of the squares of the differences between the corresponding pixel values of the two. Assume the generated image is G, the target image is T, the size of the image is H×W×C, where H is the height, W is the width, and C is the number of channels, the pixel value of the generated image is (h, w, c), and the pixel value of the target image is (h, w, c), then the mean square error loss formula is shown in Equation (1): Learning perceptual patch similarity loss: First, use the pre-trained convolutional neural network AlexNet to extract the image feature map; then normalize the extracted feature map to eliminate the influence of different scales; finally, calculate the cosine similarity for each feature layer of the generated image and the real image, and perform a weighted summation operation on the cosine similarities of different layers to obtain the final learning perceptual patch similarity loss value; assume that x and y are two images to be compared, L is the set of layers in the pre-trained network for calculating similarity, and w l is the k-th layer of weights, and φ l (x) and φ l (y) are the feature maps of image x and image y at the l-th layer respectively, N l is the number of channels of the feature map at the l-th layer, <.,.> represents the inner product of vectors, and ||.,.|| represents the norm of vectors. Then the learning perceptual patch similarity loss formula is shown in Equation (2): Image-text contrastive loss: The Clip model can map images and text into the same feature space, calculate their similarity in this space, and use the similarity calculated by Clip as the image-text contrastive loss to make the images and text generated by the model more relevant. Assume the image is I and the text is T. The image encoder encodes the image I into a feature vector v img = E img (I), and the text encoder E text encodes the text T into a feature vector v text = E text (I). Then, the cosine similarity is used to calculate the similarity between these two feature vectors. The specific formula is shown in Equation (3): For a single image-text pair (I, T), the calculation of the specific loss value is shown in Equation (4): L Clip = 1 - sim(v img , v text ) (4) In short, in the training stage, the network is trained jointly by the above three losses, and the total loss function can be expressed by Equation (5), where α, β, γ are three different weight parameters used to control the influence weights of different losses on the network learning results; L = αL MSE + βL PIPIS + γL Clip (5) S3. Count the serial numbers of the eliminated neurons and eliminate the neurons with low activation rates: Based on the law of large numbers of Bernoulli, a standard normal distribution is randomly generated as the input of the light estimation generation network model. Before registering the forward propagation hook function in each resolution synthesis layer of the generator, under multiple independent repeated experiments, the activation situation of the neurons is recorded to obtain the activation rate of each neuron, and finally the serial numbers of the neurons with low activation rates are counted; During the process of the synthesizer synthesizing the image, when synthesizing images at different resolution layers, according to the recorded serial number of each neuron, these neurons are eliminated, that is, the output of the neurons with low activation rates is set to zero, and the influence of the neurons with low activation rates is cleared in this step.