Image Processing Method, Apparatus and System

A neural network-based method for predicting lighting parameters in images accurately addresses the limitations of existing methods, enhancing the realism of lighting effects in image rendering.

CN114998092BActive Publication Date: 2025-07-15ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110224606.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-01
Publication Date
2025-07-15
Estimated Expiration
2041-03-01

AI Technical Summary

Technical Problem

The lighting prediction method in the prior art cannot accurately predict the direction, intensity, color and high-frequency components of the light, resulting in poor lighting prediction effects.

Method used

The illumination prediction technology based on spherical distribution is adopted, combined with parameter regression and direct generation methods, and the illumination panorama is trained through a neural network model to predict the illumination panorama parameters of the image to be predicted, including the spherical distribution of the main light source, the overall illumination intensity and the ambient light intensity.

Benefits of technology

It realizes accurate prediction of lighting direction, intensity and high-frequency components, improves the effect of lighting prediction, and meets the real-life rendering needs of virtual object implantation and image synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114998092B_ABST
    Figure CN114998092B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method, device, and system. Among them, the method includes: obtaining at least one sample image, where the sample image is a light panoramic image; training a neural network model using the sample image to obtain a regression network model; using the regression network model to perform light prediction on the image to be predicted, and predicting the light panoramic image parameters of the image to be predicted, where the light panoramic image parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall light intensity of the main light source, and the ambient light intensity. The present application solves the technical problem of poor light prediction effect of pictures in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and in particular, to a method, apparatus, and system for processing images. Background Art

[0002] In the fields of computer vision and computer graphics, accurate prediction of global illumination conditions is one of the important factors for achieving realistic rendering. Especially when virtual objects are implanted or image synthesis is performed, it is more necessary to accurately infer the correct illumination conditions in the scene, so that a realistic illumination rendering can be applied to the target object to obtain a real visual effect. Currently, the illumination prediction methods provided in the related art cannot accurately predict the direction, intensity, color, and high-frequency components of the illumination, resulting in poor illumination prediction effects.

[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of this application provide a method, apparatus, and system for processing images to at least solve the technical problem of poor illumination prediction effect of pictures in the related art.

[0005] According to one aspect of the embodiments of this application, a method for processing an image is provided, including: obtaining at least one sample image, where the sample image is a panoramic illumination image; training a neural network model using the sample image to obtain a regression network model; using the regression network model to perform illumination prediction on an image to be predicted, and predicting panoramic illumination image parameters of the image to be predicted, where the panoramic illumination image parameters at least include: spherical distribution of the main light source, overall illumination intensity of the main light source, and ambient light intensity in the image to be predicted.

[0006] According to another aspect of the embodiments of this application, a method for processing an image is further provided, including: displaying a playing video and a target object to be added to the video in an interface, where the video is composed of multiple frames of images, and each frame of image forms an image sequence according to the loaded timestamp; using the regression network model to perform illumination prediction on each frame of image, and predicting panoramic illumination image parameters of each frame of image, where the regression network model is generated by training a neural network model using a panoramic illumination image as a sample; when it is detected that the target object is loaded into any frame of the image, using the panoramic illumination image parameters of this frame of image to perform illumination rendering on the target object, where the panoramic illumination image parameters at least include: spherical distribution of the main light source, overall illumination intensity of the main light source, and ambient light intensity.

[0007] According to another aspect of the embodiments of the present application, there is also provided a method for processing an image, including: displaying a picture and a target object for loading the picture in an interface; predicting the illumination of the picture by using a regression network model to obtain the illumination panorama parameters of the picture, where the regression network model is generated by training a neural network model with the illumination panorama as a sample; and performing illumination rendering on the target object by using the illumination panorama parameters of the picture, where the illumination panorama parameters at least include: the spherical distribution of the main light source in the picture, the overall illumination intensity of the main light source, and the ambient light intensity.

[0008] According to another aspect of the embodiments of the present application, there is also provided a method for processing an image, including: the cloud server receives a prediction request sent by a client, where the prediction request at least carries identification information characterizing the image to be predicted; the cloud server obtains the image to be predicted based on the identification information; the cloud server performs illumination prediction on the image to be predicted by using a regression network model to obtain the illumination panorama parameters of the image to be predicted, where the illumination panorama parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall illumination intensity of the main light source, and the ambient light intensity; where the regression network model is a model generated by training a neural network model with sample images, and the sample images are at least a part of the illumination panorama.

[0009] According to another aspect of the embodiments of the present application, there is also provided a method for processing an image, including: receiving an image to be predicted sent by a client; performing illumination prediction on the image to be predicted by using a regression network model to obtain the illumination panorama parameters of the image to be predicted, where the illumination panorama parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall illumination intensity of the main light source, and the ambient light intensity; reconstructing the illumination panorama of the image to be predicted by using the illumination panorama parameters; outputting the illumination panorama to the client; where the regression network model is a model generated by training a neural network model with sample images, and the sample images are at least a part of the illumination panorama.

[0010] According to another aspect of the embodiments of the present application, there is also provided an image processing apparatus, including: an acquisition module, configured to acquire at least one sample image, where the sample image is an illumination panorama; a training module, configured to train a neural network model by using the sample image to obtain a regression network model; and a prediction module, configured to perform illumination prediction on an image to be predicted by using the regression network model to obtain the illumination panorama parameters of the image to be predicted, where the illumination panorama parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall illumination intensity of the main light source, and the ambient light intensity.

[0011] According to another aspect of the embodiments of the present application, there is also provided an image processing apparatus, including: a display module, configured to display a playing video and a target object to be added to the video in an interface, where the video is composed of multiple frames of images, and each frame of image forms an image sequence according to the loaded timestamp; a prediction module, configured to perform illumination prediction on each frame of image by using a regression network model, and predict illumination panoramic map parameters of each frame of image, where the regression network model is generated by training a neural network model with an illumination panoramic map as a sample; a rendering module, configured to perform illumination rendering on the target object by using the illumination panoramic map parameters of the frame of image when it is detected that the target object is loaded into any frame of image, where the illumination panoramic map parameters at least include: the spherical distribution of the main light source, the overall illumination intensity of the main light source, and the ambient light intensity.

[0012] According to another aspect of the embodiments of the present application, there is also provided an image processing apparatus, including: a display module, configured to display a picture and a target object that loads the picture in an interface; a prediction module, configured to perform illumination prediction on the picture by using a regression network model, and predict illumination panoramic map parameters of the picture, where the regression network model is generated by training a neural network model with an illumination panoramic map as a sample; a rendering module, configured to perform illumination rendering on the target object by using the illumination panoramic map parameters of the picture, where the illumination panoramic map parameters at least include: the spherical distribution of the main light source in the picture, the overall illumination intensity of the main light source, and the ambient light intensity.

[0013] According to another aspect of the embodiments of the present application, there is also provided an image processing apparatus, including: a receiving module, configured to receive a prediction request sent by a client, where the prediction request at least carries identification information characterizing an image to be predicted; an obtaining module, configured to obtain the image to be predicted based on the identification information; a prediction module, configured to perform illumination prediction on the image to be predicted by using a regression network model, and predict illumination panoramic map parameters of the image to be predicted, where the illumination panoramic map parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall illumination intensity of the main light source, and the ambient light intensity; where the regression network model is a model generated by training a neural network model with a sample image, and the sample image is at least a part of an illumination panoramic map.

[0014] According to another aspect of the embodiments of the present application, there is also provided an image processing apparatus, including: a receiving module, configured to receive a to-be-predicted image sent by a client; a prediction module, configured to perform illumination prediction on the to-be-predicted image by using a regression network model, and predict illumination panorama parameters of the to-be-predicted image, where the illumination panorama parameters at least include: the spherical distribution of the main light source in the to-be-predicted image, the overall illumination intensity of the main light source, and the ambient light intensity; a reconstruction module, configured to reconstruct an illumination panorama of the to-be-predicted image by using the illumination panorama parameters; an output module, configured to output the illumination panorama to the client; where the regression network model is a model generated by training a neural network model with sample images, and the sample images are at least a part of the illumination panorama.

[0015] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned image processing method.

[0016] According to another aspect of the embodiments of the present application, there is also provided a computer terminal, including a memory and a processor, where the processor is configured to run a program stored in the memory, and when the program runs, it executes the above-mentioned image processing method.

[0017] According to another aspect of the embodiments of the present application, there is also provided an image processing system, including: a processor; and a memory, connected to the processor, and configured to provide instructions for the processor to perform the following processing steps: obtaining at least one sample image, where the sample image is an illumination panorama; training a neural network model with the sample image to obtain a regression network model; performing illumination prediction on a to-be-predicted image by using the regression network model, and predicting illumination panorama parameters of the to-be-predicted image, where the illumination panorama parameters at least include: the spherical distribution of the main light source in the to-be-predicted image, the overall illumination intensity of the main light source, and the ambient light intensity.

[0018] In an embodiment of the present application, at least one sample image of a light illumination panorama can be obtained to train a neural network model, and a regression network model can be obtained. Then, the regression network model is used to perform light illumination prediction on an image to be predicted, and the light illumination panorama parameters of the image to be predicted are predicted. The light illumination panorama parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall light intensity of the main light source, and the ambient light intensity, so as to accurately predict the spherical distribution of the main light source, the overall light intensity of the main light source, and the ambient light intensity in the light illumination panorama. It is easy to think that multiple sample images with light illumination distributions can be used to train the neural network model, so that the obtained regression network model can distinguish the spherical distribution of the main light source, the overall light intensity of the main light source, and the ambient light intensity in the picture, and the trained regression network model can distinguish the spherical distribution of the main light source, the overall light intensity of the main light source, and the ambient light intensity in the picture. By performing light illumination prediction on the image to be predicted through the trained regression network model, the predicted light illumination panorama parameters can be more in line with the actual situation, thereby solving the technical problem of poor light illumination prediction effect of pictures in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0020] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing an image processing method according to an embodiment of the present application;

[0021] Figure 2 is a flowchart of an image processing method according to Embodiment 1 of the present application;

[0022] Figure 3 is an interactive interface according to Embodiment 1 of the present application;

[0023] Figures 4a to 4c is a schematic diagram of a scene according to Embodiment 1 of the present application;

[0024] Figures 5a to 5c is another schematic diagram of a scene according to Embodiment 1 of the present application;

[0025] Figures 6a to 6c is a light illumination panorama with different resolutions according to Embodiment 1 of the present application;

[0026] Figure 7 is a flowchart of another image processing method according to Embodiment 1 of the present application;

[0027] Figure 8a is a schematic diagram of a deep neural network model according to Embodiment 1 of the present application;

[0028] Figure 8b It is a schematic diagram of an overall link of a neural mapping according to Embodiment 1 of the present application;

[0029] Figure 9 It is a flowchart of a method for processing an image according to Embodiment 2 of the present application;

[0030] Figure 10 It is a flowchart of a method for processing an image according to Embodiment 3 of the present application;

[0031] Figure 11 It is a schematic diagram of an image processing apparatus according to Embodiment 4 of the present application;

[0032] Figure 12 It is a schematic diagram of an image processing apparatus according to Embodiment 5 of the present application;

[0033] Figure 13 It is a schematic diagram of an image processing apparatus according to Embodiment 6 of the present application;

[0034] Figure 14 It is a block diagram of the structure of a computer terminal according to an embodiment of the present application;

[0035] Figure 15 It is a flowchart of a method for processing an image according to Embodiment 7 of the present application;

[0036] Figure 16 It is a schematic diagram of an image processing apparatus according to Embodiment 8 of the present application;

[0037] Figure 17 It is a flowchart of a method for processing an image according to Embodiment 9 of the present application;

[0038] Figure 18 It is a schematic diagram of an image processing apparatus according to Embodiment 10 of the present application. Detailed implementation manners

[0039] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0040] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0041] First, some nouns or terms that appear during the description of the embodiments of this application are applicable to the following explanations:

[0042] Panorama: It can be an image with a viewing angle covering 180° + / - horizontally and 90° + / - vertically, and a length-width ratio of 2:1.

[0043] Illumination panorama: It can be a panorama representing the illumination intensity of the global environment.

[0044] Earth mover's mover distance: It can refer to a metric for measuring the distance between two distributions.

[0045] Rendering: It can be a process of using computer graphics technology to generate a two-dimensional image from a virtual three-dimensional scene (including information such as geometry, viewing angle, texture material, illumination, etc.).

[0046] Generative adversarial network: It can be a neural network that uses the mutual game between a set of generators and discriminators to generate pictures.

[0047] High-dynamic range image (HDR): An image with a high pixel dynamic range.

[0048] Low-dynamic range image (LDR): An image with a low pixel dynamic range.

[0049] Spherical Gaussian function: It can be a two-dimensional Gaussian function defined on the unit sphere.

[0050] Spherical harmonic function: It can be the angular part of the solution of the Laplace equation in spherical coordinates and is a generalization of the Fourier basis function on the sphere.

[0051] Currently, there are mainly two methods to predict lighting conditions. The first method is the parameter regression-based method, which can represent the lighting panorama with basis functions to obtain a set of lighting parameters, then use a neural network to regress the lighting parameters from the image, and finally reconstruct the lighting panorama using the regressed parameters. The other method is the direct generation method, which can directly use a neural network to generate the lighting panorama. However, the parameter regression-based method can only reconstruct a rough lighting map and it is difficult to accurately predict lighting details, especially high-frequency components; while the direct generation method cannot accurately predict basic lighting attributes such as lighting direction and lighting intensity.

[0052] To solve the above problems, this application provides the following implementation solution. The lighting prediction technology based on spherical distribution combines the parameter regression method and the direct generation method, which can accurately predict the lighting direction, intensity, color, and high-frequency components, and can achieve a leading lighting prediction effect. At the same time, the physical meaning of the algorithm is clear and there is room for continuous improvement and refinement.

[0053] Embodiment 1

[0054] According to the embodiments of this application, an embodiment of a method for processing images is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0055] The method embodiments provided by the embodiments of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the method for processing images is shown. As Figure 1 shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0056] It should be noted that one or more of the above-mentioned processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10 (or mobile device). The data processing circuit serves as a processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0057] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the image processing method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned image processing method. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0058] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0059] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0060] It should be noted here that in some alternative embodiments, the above-mentioned Figure 1 shown computer device (or mobile device) can include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1This is just an example of a specific concrete instance and is intended to illustrate the types of components that may exist in the above computer device (or mobile device).

[0061] In Figure 1 the operating environment shown, the present application provides a method for processing an image as shown in Figure 2 the figure. Figure 2 is a flowchart of a method for processing an image according to Embodiment 1 of the present application. As shown in Figure 2 the figure, the method may include the following steps:

[0062] Step S202, obtain at least one sample image.

[0063] Wherein, the sample image is a panoramic illumination image.

[0064] The sample image in the above step may be a panoramic illumination image in various scenarios; the sample image may be a panoramic illumination image of indoor lighting, or a panoramic illumination image of indoor natural light, or a panoramic illumination image of outdoor sunlight, or a panoramic illumination image of outdoor natural light.

[0065] The panoramic illumination image in the above step may be a panoramic image used to represent the illumination intensity of the global environment; it should be noted that the panoramic image may be an image with a viewing angle covering 180° + / - each of the horizon and 90° + / - each of the vertical, and a length-width ratio of 2:1.

[0066] In an alternative embodiment, in order to better process the sample image, the at least one obtained sample image may be transmitted to a corresponding processing device for processing. For example, it may be directly transmitted to the user's computer terminal (such as a laptop, a personal computer, etc.) for processing, or transmitted to a cloud server through the user's computer terminal for processing. It should be noted that since the processing of the sample image requires a large amount of computing resources, in the embodiments of the present application, the cloud server is taken as an example of the processing device for description.

[0067] For example, in order to facilitate the user to upload the sample image, an interactive interface may be provided to the user. As shown in Figure 3 the figure, the user can sequentially select the sample images for training the model from a large number of stored sample images by clicking the "Select Image" button, or batch-select multiple sample images for training the model, and then upload the sample images for training the model to the cloud server for training by clicking the "Upload" button. In addition, in order to facilitate the user to confirm whether the sample image uploaded to the cloud server is the sample image for training the model, the selected sample image may be displayed in the "Image Display" area, and after the user confirms that it is correct, the data can be uploaded by clicking the "Upload" button.

[0068] Step S204: Train a neural network model using sample images to obtain a regression network model.

[0069] The regression network model in the above steps is a neural network model trained with sample images.

[0070] The neural network model in the above steps can be composed of densenet-121 (dense convolutional network), or can be composed of other types of neural networks. There is no limitation on the composition method of the neural network model here.

[0071] In an optional embodiment, the light panoramic image can be processed, and local images with a limited view are randomly cropped at different positions of the light panoramic image, and the local images are used to train the neural network model.

[0072] In another optional embodiment, various types of light panoramic images or local images of the light panoramic image can be used to train the neural network model, so that the parameters of the neural network model reach a certain accuracy, thereby obtaining a regression network model, enabling users to use the regression network model to perform light prediction on the image to be predicted.

[0073] In another optional embodiment, the sample image can be processed to obtain the overall light intensity of the main light source, the ambient light intensity, and the spherical distribution of the main light source in the sample image; specifically, the pixels in the sample image can be classified according to a preset threshold, the pixels less than the preset threshold are used as ambient light pixels, and the pixels greater than or equal to the preset threshold are used as main light source pixels. A preset number of evenly distributed anchor points can be set on the sphere, and the light intensity corresponding to the main light source pixels is matched and summed according to the distance between the main light source pixels and the anchor points to obtain the light intensity value on each anchor point. The light intensity values on each anchor point are summed to obtain the overall light intensity of the main light source; the average value of the light intensity corresponding to the ambient light pixels can be obtained as the ambient light intensity; the light intensity values in a preset number of anchor points can be normalized to obtain the spherical distribution of the main light source; the overall light intensity of the main light source, the ambient light intensity, and the spherical distribution of the main light source obtained from the sample image can be recorded in the attributes of the sample image.

[0074] It should be noted that the normalization process can be to divide the light intensity value of each anchor point by the overall light intensity to obtain the proportion of the light intensity value of each anchor point in the overall light intensity, that is, to obtain the spherical distribution of the main light source in the sample image.

[0075] In yet another alternative embodiment, the loss function used in training the neural network model can be a spherical earth mover's distance loss function or an Euclidean distance loss function, which can make the spherical distribution of the main light source obtained by training the neural network model closer to the true spherical distribution, so that the obtained spherical distribution of the main light source has better effects in practical applications.

[0076] Step S206: Use the regression network model to perform light prediction on the image to be predicted, and predict the light panorama parameters of the image to be predicted. The light panorama parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall light intensity of the main light source, and the ambient light intensity.

[0077] In an alternative embodiment, the backbone network of the regression network model can be densenet-121 (dense convolutional network), and the output layer of the regression network model can be 3 fully connected layers (fully connected or fc), namely the fully connected layer for light distribution, the fully connected layer for light intensity, and the fully connected layer for ambient light intensity.

[0078] Exemplarily, the light panorama can be randomly cropped to obtain a local image with a limited view angle, and then the local image is input into the dense convolutional network. The spherical distribution of the main light source in the light panorama is output through the fully connected layer for light distribution, the overall light intensity of the main light source is output through the fully connected layer for light intensity, and the ambient light intensity is output through the fully connected layer for ambient light intensity. Among them, the fully connected layer for light distribution has 3*N values, the fully connected layer for light intensity has 3 values, and the fully connected layer for ambient light intensity has 3 values; it should be noted that 3 represents the three channels of RGB of the color, and N represents the number of anchor points. The number of anchor points can be set to 128.

[0079] In another alternative embodiment, after the cloud server outputs the light panorama parameters of the image to be predicted, the light panorama to be predicted and the light panorama parameters can be output to a user-operable display screen, enabling the user to actively modify the light panorama parameters. The light panorama to be predicted and the modified light panorama parameters can also be transmitted to the user-operable display screen, enabling the user to verify whether the modified light panorama parameters are accurate. If accurate, the user can press the confirmation button on the display screen; if inaccurate, the user can press the change button on the display screen. At this time, the user can continue to modify the light panorama parameters with reference to the light panorama to be predicted and send the modified light panorama parameters to the cloud server.

[0080] In yet another alternative embodiment, in a scenario where light prediction is required, a regression network model can be used to perform light prediction on the image to be predicted, and the light panorama parameters of the image to be predicted are obtained, so as to adjust the light intensity in the scenario by using the light panorama parameters, thereby meeting the user's requirements.

[0081] Exemplarily, as Figure 4a shown, in a prepared bedding picture, a pillow needs to be added for display. However, directly adding the picture of the pillow to the bedding picture will make the whole picture very uncoordinated. Therefore, it is necessary to predict the light of the bedding picture to obtain the light panorama parameters of the bedding picture, so as to adjust the light intensity near the newly added pillow picture according to the light panorama parameters of the bedding picture, making the whole picture very coordinated, as Figure 5a shown, the light panorama of the predicted bedding picture is shown in the upper left box, and the pillow in the picture is the newly added pillow picture.

[0082] Exemplarily, as Figure 4b shown, in a taken indoor space picture, a fire hydrant needs to be added for safety learning. However, directly adding the picture of the fire hydrant to the indoor space picture will make the whole picture very uncoordinated. Therefore, it is necessary to predict the light of the indoor space picture to obtain the light panorama parameters of the indoor space picture, so as to adjust the light intensity near the newly added fire hydrant according to the light panorama parameters of the indoor space picture, thereby coordinating the whole picture, as Figure 5b shown, the light panorama of the predicted indoor space picture is shown in the upper left box, and the fire hydrant in the picture is the newly added fire hydrant picture.

[0083] Exemplarily, as Figure 4c shown, in a washbasin picture, a water cup needs to be added. However, directly adding the picture of the water cup to the indoor space picture will make the whole picture very uncoordinated. Therefore, it is necessary to predict the light of the washbasin picture to obtain the light panorama parameters of the washbasin picture, so as to adjust the light intensity near the newly added cup according to the light panorama parameters of the washbasin picture, making the whole picture very coordinated, as Figure 5c shown, the light panorama of the predicted washbasin picture is shown in the upper left box, and the water cup in the picture is the newly added water cup picture.

[0084] Through the solution provided by the above embodiments of the present application, at least one sample image of the illumination panoramic image can be obtained to train a neural network model, and a regression network model can be obtained. Then, the regression network model is used to perform illumination prediction on the image to be predicted, and the illumination panoramic image parameters of the image to be predicted are predicted. The illumination panoramic image parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall illumination intensity of the main light source, and the ambient light intensity, so as to accurately predict the spherical distribution of the main light source, the overall illumination intensity of the main light source, and the ambient light intensity in the illumination panoramic image. It is easy to think that multiple sample images with illumination distributions can be used to train the neural network model, so that the trained regression network model can distinguish the spherical distribution of the main light source, the overall illumination intensity of the main light source, and the ambient light intensity in the picture. By using the trained regression network model to perform illumination prediction on the image to be predicted, the predicted illumination panoramic image parameters can be more in line with the actual situation, thereby solving the technical problem of poor illumination prediction effect of pictures in the related art.

[0085] In the above embodiments of the present application, before training the neural network model with the sample image to obtain the regression network model, the method further includes: obtaining the image data of the main light source in the sample image; obtaining a plurality of anchor points set on the unit sphere; matching the image data of the main light source with each anchor point, and based on the matching result, obtaining the spherical distribution of the main light source in the sample image.

[0086] The matching result in the above step may be that a unique anchor point is assigned to each pixel of the main light source.

[0087] The image data of the main light source in the above step may be the main light source pixels.

[0088] In an alternative embodiment, the illumination panoramic image can be binary-classified using a preset threshold. The pixels with illumination intensity greater than or equal to the preset threshold are used as the main light source pixels, and the pixels with illumination intensity less than the preset threshold are used as the ambient light pixels. Among them, the main light source pixels can be represented as a discrete probability distribution function and the overall illumination intensity on the unit sphere, and the preset threshold can be the light source threshold. A preset number of uniformly distributed anchor points can be predefined on the sphere, and a unique anchor point is assigned to each main light source pixel according to the distance between the anchor point and the main light source pixel. The light intensity value of the anchor point can be the sum of the illumination intensities of the corresponding main light source pixels. It can be assumed that the light intensity value of each anchor point is a spherical Gaussian function. Therefore, the overall illumination intensity of the main light source can be approximately represented as a panoramic Gaussian illumination map.

[0089] Furthermore, the overall light intensity of the main light source can be obtained by summing up the light intensity values of all anchor points. Then, the light intensity value of each anchor point is normalized. Specifically, the light intensity value of each anchor point is divided by the overall light intensity of the main light source to obtain the proportion of the light intensity value of each anchor point to the overall light intensity, that is, the spherical discrete distribution probability of the light source is defined, and each anchor point represents the ratio of the incident light intensity value from the direction it represents to the overall light intensity of the main light source, thereby obtaining the spherical distribution of the main light source in the image to be predicted.

[0090] In yet another optional embodiment, the number of pre-set anchor points may be 128.

[0091] In the above embodiments of the present application, pixels in the sample image are extracted according to the light intensity, wherein pixels with light intensity greater than or equal to the light source threshold belong to the image data of the main light source, and pixels with light intensity less than the light source threshold belong to the image data of the ambient light.

[0092] The light source threshold in the above steps may be set by the user, or may be a light source threshold obtained through experiments.

[0093] In an optional embodiment, pixels in the sample image are extracted according to the light intensity. Since the main light source is generally a light source with stronger light seen by the human eye, the pixels with light intensity greater than or equal to the light source threshold can be set as image data of the main light source; and the ambient light source is generally a light source with more natural or darker light seen by the human eye. Therefore, the pixels with light intensity less than the light source threshold can be set as image data of the ambient light.

[0094] In the above embodiment of the present application, the image data of the main light source is matched with each anchor point, and based on the matching result, the spherical distribution of the main light source is obtained, including: clustering based on the distance between the image data of the main light source and each anchor point, and obtaining the image data of the main light source matching any anchor point; summing the light intensity of the image data of the main light source matching the anchor point to obtain the light intensity value of the anchor point; normalizing the light intensity value of each anchor point to generate the spherical distribution of the main light source in the sample image.

[0095] The clustering in the above steps is the process of dividing a set of physical or abstract objects into multiple classes consisting of similar objects.

[0096] In an alternative embodiment, the main light source pixels near the anchor point whose distance from the anchor point is less than a preset distance can be clustered on the anchor point, or the main light source pixels in a preset area near the anchor point can be clustered on the anchor point. Among them, the shape and size of the preset area can be set by the user; the preset area can be a circular area with a diameter of 1 cm, or a square area with a length and width of 1 cm, or an irregular area. Here, no limitations are imposed on the shape and size of the preset area.

[0097] In another alternative embodiment, the pixels of the main light source matching any one anchor point can be obtained, and then the illumination intensities corresponding to the pixels of the main light source matching the anchor point are summed to obtain the light intensity value of the anchor point. Then, the light intensity value of each anchor point is divided by the overall illumination intensity of the main light source to obtain the proportion of the illumination intensity of each anchor point in the overall illumination intensity of the main light source, that is, the spherical distribution of the main light source in the sample image is obtained.

[0098] In the above embodiments of the present application, the first loss function is used to perform regression summation on the illumination intensities of the pixels belonging to the main light source to obtain the overall illumination intensity of the main light source in the sample image; the second loss function is used to perform regression averaging on the illumination intensities of the pixels belonging to the ambient light to obtain the ambient light intensity of the ambient light in the sample image.

[0099] The first loss function in the above steps can be a spherical earth mover's distance loss function or an Euclidean distance loss function, and the second loss function can be a spherical earth mover's distance loss function or an Euclidean distance loss function. In the embodiments of the present application, the spherical earth mover's distance loss function is taken as an example for illustration.

[0100] In an alternative embodiment, an N×N distance matrix C can be defined, and the element C[i, j] of the matrix C represents the arc length between the anchor point i and the anchor point j. Among them, the object to be optimized can be an N×N matrix T, and the element T[i, j] of the matrix T represents the probability of transferring the light intensity ratio at the anchor point i to the anchor point j. Summing the matrix by rows, an N-dimensional vector is obtained, which is the spherical distribution probability U of the light source obtained during training. Summing the matrix by columns, an N-dimensional vector is also obtained, which is the true spherical distribution probability V of the light source. Thus, the spherical earth mover's distance loss function L can be derived. sml Actually, the spherical distribution probability of the predicted light source can be ensured to be as close as possible to the true spherical distribution by optimizing the network.

[0101] Furthermore, it can be represented by the following formula:

[0102]

[0103]

[0104] Using the above-mentioned earth mover's loss function to perform regression summation on the illumination intensities of the pixels belonging to the main light source, and obtaining the overall illumination intensity of the main light source in the sample image, can make the overall illumination intensity of the main light source more closely match the actual overall light intensity.

[0105] Using the above-mentioned earth mover's loss function to perform regression averaging on the illumination intensities of the pixels belonging to the ambient light, and obtaining the ambient light intensity of the ambient light in the sample image, can make the ambient light intensity of the main light source more closely match the actual ambient light intensity.

[0106] In the above embodiments of the present application, after using the regression network model to perform illumination prediction on the image to be predicted and predicting the illumination panorama of the image to be predicted, the method further includes: using the illumination panorama parameters to reconstruct the illumination panorama of the image to be predicted.

[0107] In an alternative embodiment, since the input local image with a limited view also contains partial scene information, therefore, the regression network can be further fine-tuned by means of the semantic content of the input image to obtain the illumination panorama.

[0108] In the above embodiments of the present application, the adversarial network is used to optimize the illumination panorama of the image to be predicted to generate a high-precision illumination panorama.

[0109] The adversarial network in the above steps is mainly a neural network that uses the mutual game between a group of generators and discriminators to generate pictures.

[0110] In the above embodiments of the present application, using the adversarial network to optimize the illumination panorama of the image to be predicted to generate a high-precision illumination panorama includes: obtaining the local image of the image to be predicted; using the encoder to vectorize the local image of the image to be predicted to obtain a group of latent space vectors; inputting the latent space vectors into the adversarial network, and superimposing and optimizing the resolution of the illumination panorama of the image to be predicted to generate a high-precision illumination panorama.

[0111] The latent space vectors in the above steps can be the feature vectors of the image semantic content.

[0112] In an alternative embodiment, the input image can be sent into the encoder to obtain the feature vectors of the image semantic content, and then the approximate illumination panoramas at different resolutions are obtained and used as input conditions to guide the generation of the refined illumination panorama. First, a low-resolution refined panorama is generated, as Figure 6a shown; then the resolution is continuously increased, as Figure 6b and Figure 6c shown, until it is consistent with the resolution of the approximate illumination panorama. In addition, a spherical convolution kernel module can be used to ensure that the generated image does not deform when approaching the poles.

[0113] In another alternative embodiment, during the training process of the adversarial network, the loss function adopted is defined as follows:

[0114]

[0115] Let \(x\) represent the approximate illumination panorama used to guide the generation process, \(x'\) represent the more refined illumination panorama generated, and \(y\) represent the true illumination panorama.

[0116] Among them, \(L\) feat The loss mainly hopes that the intermediate layer features of the discriminator \(D\) are as similar as possible, which is expressed as follows:

[0117]

[0118] Among them, \(L\) cos The loss mainly hopes that the cosine distance between the generated \(x'\) and the true value \(y\) is similar, so as to ensure that the overall distribution of light intensity is similar, which is expressed as follows:

[0119]

[0120] \(L\) feat The loss and \(L\) cos The loss is the normal generator loss and discriminator loss, and its role is to ensure that the generated \(x'\) is as real as possible.

[0121] In the above embodiments of the present application, after using the adversarial network to optimize the illumination panorama of the image to be predicted and generate a high-precision illumination panorama, the method further includes: using the high-precision illumination panorama as a sample image to perform regression training on the neural network model, and optimizing the training of the neural network model.

[0122] In an alternative embodiment, after using the high-precision illumination panorama as a sample image to perform regression training on the neural network model, the process of optimizing the training of the neural network model can be made more accurate, so that the obtained regression network model is also more accurate.

[0123] Next, in combination with Figure 7 、 Figure 8a 、 Figure 8b A preferred embodiment of the present application will be described in detail. This method can be executed by a mobile terminal or a server. In the embodiments of the present application, an example in which this method is executed by a server will be described. As Figure 7 shown, the method may include the following steps:

[0124] Step S701, select 5% of the maximum intensity value in the illumination panorama as the critical threshold, and distinguish it into two parts: the main light source and the ambient light;

[0125] Step S702: Set 128 evenly distributed anchor points on the unit sphere, and perform matching summation for each pixel in the light panorama that is higher than the threshold according to the distance to obtain the light intensity value at each anchor point.

[0126] Specifically, the intensity value of the main light source is the sum of the intensity values of all the anchor points, while the ambient light intensity is the average value of the pixel values lower than the critical threshold.

[0127] Step S703: Normalize the 128 anchor points to obtain the spherical distribution of the main light source.

[0128] Among them, the light panorama is decomposed into the spherical distribution, the overall light intensity, and the ambient light intensity. These three items are the targets to be predicted in the light panorama. By decomposing the light panorama, the difficulty of regressing the parameters of the high-dynamic-range image can be alleviated, thereby improving the efficiency of regressing the parameters.

[0129] Step S704: Design a deep neural network to regress the spherical distribution, the overall light intensity, and the ambient light intensity.

[0130] As Figure 8a shown in the part above the dashed line, each pixel of the light panorama can be binary-classified with a certain threshold. The pixels lower than the threshold are ambient light pixels, and the pixels higher than the threshold are main light source pixels. The main light source can be further represented as a discrete probability distribution function on the unit sphere and the overall light intensity. 128 evenly distributed anchor points are predefined on the sphere, and a unique anchor point can be assigned to each pixel representing the main light source according to the distance. The light intensity value of the anchor point i is equal to the sum of the intensity values of the corresponding main light source pixels. At the same time, it is assumed that the light source of each anchor point is a spherical Gaussian function, so the main light source is approximately represented as a "panoramic Gaussian illumination map".

[0131] Furthermore, summing up the light intensity values of all 128 anchor points gives the overall light intensity (lightintensity), and then normalizing the light intensity value of each anchor point (dividing the light intensity value of each anchor point by the overall light intensity) to obtain the proportion of the light intensity value of each anchor point in the overall light intensity, thereby obtaining the discrete distribution probability of the light source on the sphere (light distribution), that is, the spherical distribution of the light. Among them, each anchor point represents the ratio of the incident light intensity in the direction it represents to the overall light intensity.

[0132] For all the ambient light pixels, their average can be obtained to get the ambient light intensity (ambient term).

[0133] Figure 8aThe part below the dotted line can represent the training samples and the neural network structure. As shown in the part below the dotted line, images with a limited view can be randomly cropped at different positions of the light panoramic image as the network input, while the network output is the spherical distribution of the light source, the overall light intensity, and the ambient light intensity obtained in the previous steps. The main part of the network structure is the dense convolutional network densenet-121 (which can be composed of multiple dense convolutional modules). After passing through a fully connected layer, the output layer of the network finally consists of 3 fully connected layers (fully connected or fc), which respectively output the spherical distribution of the light source (3*N values), the overall light intensity (3 values), and the ambient light intensity (3 values). Among them, 3 represents the three RGB channels of the color, and N represents the number of anchor points, which can be 128.

[0134] In the above network training, the Euclidean distance loss (L2 loss) can be used to regress the overall light intensity and the ambient light intensity. However, for the spherical distribution of the light source, the biggest problem with the L2 loss function is that it independently penalizes the probability values of each anchor point without considering the 128 anchor points as a whole (i.e., a spherical distribution and the sum of the probability values of all the anchor points on the sphere is 1).

[0135] Therefore, the above problem can be solved by using the spherical earth mover's distance corresponding spherical earth mover's loss function. By taking the spherical distance between the 128 anchor points as the cost in the earth mover's distance, the geometric information of the spherical distribution can be effectively utilized. And the earth mover's distance itself is a measure of the distance between distributions. Therefore, the property that the sum of the anchor point probability values is 1 can also be utilized.

[0136] Step S705, the spherical distribution of the light source, the overall light intensity, and the ambient light intensity can be obtained through the regression network model;

[0137] Step S706, using the spherical Gaussian function, the light map can be reconstructed from the parameters.

[0138] Figure 8a The shown process can obtain various parameters, that is, the spherical distribution of the light source, the overall light intensity, and the ambient light intensity. After obtaining the above parameters, assuming that the light source at each anchor point is a spherical Gaussian function, the approximate expression of the panoramic lighting conditions can be directly obtained using the parameters, that is, the panoramic Gaussian light map, as shown in Figure 8a "Panoramic Gaussian light map (two-dimensional representation)" and "Spherical Gaussian light map (three-dimensional representation)" in.

[0139] Since the input image with a limited view also contains partial scene information, the light panoramic image obtained by the regression network can be further refined by leveraging the semantic content of the input image. Specifically, a generative adversarial network can be used to further guide and optimize the generation of the light panoramic image, reconstruct the light map from the predicted parameters, and use the predicted parameters as a guide to further generate a high-precision light map using the generative adversarial network, thereby enabling the restoration of real high-frequency details. The overall processing flow is as follows Figure 8b shown. First, the input image is fed into the encoder to obtain the semantic features of the image. At the same time, panoramic Gaussian light maps at different resolutions can be obtained and used as input conditions to guide the generation of a refined light panoramic image. The generative adversarial network can first generate a low-resolution refined panoramic image and then continuously increase the resolution until it matches the resolution of the panoramic Gaussian light map, thereby obtaining a high-precision light panoramic image, that is, obtaining the "spherical generated image (three-dimensional representation)" and "panoramic generated image (two-dimensional representation)" as shown in Figure 8b .

[0140] It should be noted that during the process of continuously increasing the resolution to generate the light panoramic image, a spherical convolution module can be adopted to ensure that the generated image does not deform when approaching the poles. As shown in Figure 8b , multiple different spherical convolution modules can be used. Combining Figures 6a to 6c it can be known that the larger the volume of the spherical convolution module, the more pixel points it has. Therefore, the higher the resolution of the generated light panoramic image. Figure 8b The number of spherical convolution modules in

[0141] can be increased or decreased according to the actual situation. If two spherical convolution modules can obtain a panoramic generated image that meets the processing accuracy, then two spherical convolution modules can be used; if five or more spherical convolution modules can obtain a panoramic generated image that meets the processing accuracy, then five or more spherical convolution modules can be used.

[0142] In the above steps, by combining the advantages of the regression network model and the generative model, the spherical distribution of the light source, the overall illumination intensity, and the ambient light intensity in the scene are accurately predicted using the regression network first. Through the normalized spherical distribution, the influence of the large difference in the numerical range of the illumination map is avoided, thereby improving the stability and accuracy of the model regression. Then, the parameters obtained by regression can be used to reconstruct the illumination panorama as a guide, and further, a generative adversarial network is used to generate a high-precision illumination map, and the high-frequency information is also reconstructed through adversarial learning.

[0143] Furthermore, through the illumination prediction technology, an automatically and intelligently predicted global illumination map with a high dynamic range is obtained from a single image, eliminating the cumbersome and time-consuming cost of interactive adjustment of illumination effects in the implantation process. In addition, theoretical knowledge such as global illumination is complex and professional, and designers mostly rely on years of accumulated experience for design and development. The technology in this proposal provides a high-quality illumination prediction tool, helping designers quickly and efficiently perform design tasks such as 3D implantation, greatly improving the work efficiency of designers.

[0144] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0145] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a more preferred implementation. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.

[0146] Embodiment 2

[0147] According to an embodiment of this application, an embodiment of an image processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0148] Figure 9 It is a flowchart of a method for processing images according to Embodiment 2 of the present application. As Figure 9 shown, the method may include the following steps:

[0149] Step S902, display the playing video and the target object of the video to be added in the interface.

[0150] Among them, the video is composed of multiple frames of images, and each frame of image forms an image sequence according to the loaded timestamp.

[0151] The interface in the above steps may be the display interface of terminal devices such as mobile phones, computers, tablets, etc. that can display images.

[0152] The playing video in the above steps may be a previously recorded video, which may be a previously recorded advertisement video, teaching video, film and television video, etc.

[0153] The target object of the video to be added in the above steps may be a product to be displayed, a specific person, a prop, etc.

[0154] In an optional embodiment, the playing video displayed in the interface may be an advertisement video, and the target object of the video to be added may be a product to be displayed in the advertisement, a spokesperson, etc.

[0155] In another optional embodiment, the playing video displayed in the interface may be a teaching video, and the target object of the video to be added may be the main teacher, teaching aids, etc.

[0156] In yet another optional embodiment, the playing video displayed in the interface may be a film and television video, and the target object of the video to be added may be an actor, a prop, etc.

[0157] Step S904, use a regression network model to perform illumination prediction on each frame of image, and predict the illumination panorama parameters of each frame of image.

[0158] Among them, the regression network model is generated by training a neural network model with the illumination panorama as a sample.

[0159] Step S906, when it is detected that the target object is recorded in any frame of image, use the illumination panorama parameters of this frame of image to perform illumination rendering on the target object.

[0160] Among them, the illumination panorama parameters at least include: the spherical distribution of the main light source, the overall illumination intensity of the main light source, and the ambient light intensity.

[0161] The rendering in the above steps is a process of using computer graphics technology to generate a two-dimensional image from a virtual three-dimensional scene (including information such as geometry, perspective, texture material, illumination, etc.).

[0162] In an alternative embodiment, in the application scenario of an advertising video, when it is detected that a product to be displayed is recorded in any frame of the image, the light intensity near the product can be adjusted using the light panoramic map parameters of this frame of the image, so that the product is displayed more realistically in the advertising video.

[0163] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0164] Embodiment 3

[0165] According to an embodiment of this application, an embodiment of a method for processing an image is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0166] Figure 10 is a flowchart of a method for processing an image according to Embodiment 3 of this application. As Figure 10 shown, the method may include the following steps:

[0167] Step S1002, display a picture and the target object for loading the picture in the interface.

[0168] The picture displayed in the above step may be an advertising picture, a poster picture, an indoor design schematic diagram, etc.

[0169] The target object in the above step may be a product to be displayed, a star, furniture to be added, etc.

[0170] In an alternative embodiment, the picture displayed in the interface may be a picture of a kitchen, and the target object may be a picture of a pot to be sold, etc.

[0171] Step S1004, use a regression network model to predict the illumination of the picture, and predict the light panoramic map parameters of the picture.

[0172] Among them, the regression network model is generated by training a neural network model using the light panoramic map as a sample.

[0173] Step S1006, use the light panoramic map parameters of the picture to perform light rendering on the target object.

[0174] Among them, the light panoramic map parameters at least include: the spherical distribution of the main light source in the picture, the overall illumination intensity of the main light source, and the ambient light intensity.

[0175] In an alternative embodiment, the lighting panoramic parameters of the kitchen picture are used to perform lighting rendering on the added cookware picture, so that the cookware picture can be more realistically displayed in the kitchen picture.

[0176] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0177] Embodiment 4

[0178] According to an embodiment of the present application, there is also provided an image processing device for implementing the above image processing method, as Figure 11 shown. The device 1100 includes: an acquisition module 1102, a training module 1104, and a prediction module 1106.

[0179] Among them, the acquisition module is used to acquire at least one sample image, where the sample image is a lighting panoramic image; the training module is used to train a neural network model using the sample image to obtain a regression network model; the prediction module is used to perform lighting prediction on the image to be predicted using the regression network model, and predict the lighting panoramic parameters of the image to be predicted. The lighting panoramic parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall lighting intensity of the main light source, and the ambient light intensity.

[0180] It should be noted here that the above acquisition module 1102, training module 1104, and prediction module 1106 correspond to steps S202 to S206 in Embodiment 1. The functions and application scenarios of the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0181] In the above embodiments of the present application, the device further includes: an acquisition module and a matching module.

[0182] Among them, the acquisition module is further used to acquire the image data of the main light source in the sample image; the acquisition module is further used to acquire a plurality of anchor points set on the unit sphere; the matching module is used to match the image data of the main light source with each anchor point, and based on the matching result, obtain the spherical distribution of the main light source in the sample image.

[0183] In the above embodiments of the present application, the device further includes: an extraction module.

[0184] Among them, the extraction module is used to extract the pixels in the sample image according to the lighting intensity. Among them, the pixels with a lighting intensity greater than or equal to the light source threshold belong to the image data of the main light source, and the pixels with a lighting intensity less than the light source threshold belong to the image data of the ambient light.

[0185] In the above embodiments of the present application, the matching module includes: a clustering unit, a summing unit, and a first processing unit.

[0186] Among them, the clustering unit is used to perform clustering based on the distances between the image data of the main light source and each anchor point, and obtain the image data of the main light source that matches any one of the anchor points; the summing unit is used to sum the illumination intensities of the image data of the main light source that matches the anchor point to obtain the light intensity value of the anchor point; the first processing unit is used to perform normalization processing on the light intensity values of each anchor point to generate the spherical distribution of the main light source in the sample image.

[0187] In the above embodiments of the present application, the device further includes: a summing module and an averaging module.

[0188] Among them, the summing module is used to perform regression summation on the illumination intensities of the pixels belonging to the main light source by using the first loss function to obtain the overall illumination intensity of the main light source in the sample image; the averaging module is used to perform regression averaging on the illumination intensities of the pixels belonging to the ambient light by using the second loss function to obtain the ambient light intensity of the ambient light in the sample image.

[0189] In the above embodiments of the present application, the device further includes: a reconstruction module.

[0190] Among them, the reconstruction module is used to reconstruct the illumination panorama of the image to be predicted by using the illumination panorama parameters.

[0191] In the above embodiments of the present application, the device further includes: a processing module.

[0192] Among them, the processing module is used to optimize the illumination panorama of the image to be predicted by using an adversarial network to generate a high-precision illumination panorama.

[0193] In the above embodiments of the present application, the processing module includes: an acquisition unit, a second processing unit, and an optimization unit.

[0194] Among them, the acquisition unit is used to acquire the local image of the image to be predicted; the second processing unit is used to perform vectorization processing on the local image of the image to be predicted by using an encoder to obtain a group of latent space vectors; the optimization unit is used to input the latent space vectors into the adversarial network to superimpose and optimize the resolution of the illumination panorama of the image to be predicted to generate a high-precision illumination panorama.

[0195] In the above embodiments of the present application, the device further includes: an optimization training module.

[0196] Among them, the optimization training module is used to use the high-precision illumination panorama as a sample image to perform regression training on the neural network model to optimize and train the neural network model.

[0197] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0198] Embodiment 5

[0199] According to an embodiment of the present application, there is also provided an image processing device for implementing the above-mentioned image processing method. As Figure 12 shown, the device 1200 includes: a display module 1202, a prediction module 1204, and a rendering module 1206.

[0200] Among them, the display module is used to display the playing video and the target object to be added to the video on the interface. The video is composed of multiple frames of images, and each frame of image forms an image sequence according to the loaded timestamp. The prediction module is used to perform light prediction on each frame of image by using a regression network model, and predict the light panorama parameters of each frame of image. The regression network model is generated by training a neural network model with the light panorama as a sample. The rendering module is used to perform light rendering on the target object by using the light panorama parameters of the frame of image when it is detected that the target object is loaded into any frame of image. The light panorama parameters at least include: the spherical distribution of the main light source, the overall light intensity of the main light source, and the ambient light intensity.

[0201] It should be noted here that the above display module 1202, prediction module 1204, and rendering module 1206 correspond to steps S902 to S906 in Embodiment 2. The instances and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules can run in the computer terminal 10 provided in Embodiment 1 as part of the device.

[0202] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0203] Embodiment 6

[0204] According to an embodiment of the present application, there is also provided an image processing device for implementing the above-mentioned image processing method. As Figure 13 shown, the device 1300 includes: a display module 1302, a prediction module 1304, and a rendering module 1306.

[0205] Among them, a display module is configured to display an image and a target object for loading the image on an interface; a prediction module is configured to perform illumination prediction on the image by using a regression network model, and obtain illumination panoramic image parameters of the image, where the regression network model is generated by training a neural network model with an illumination panoramic image as a sample; a rendering module is configured to perform illumination rendering on the target object by using the illumination panoramic image parameters of the image, where the illumination panoramic image parameters at least include: the spherical distribution of the main light source in the image, the overall illumination intensity of the main light source, and the ambient light intensity.

[0206] It should be noted here that the above display module 1302, prediction module 1304, and rendering module 1306 correspond to steps S1002 to S1006 in Embodiment 3. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0207] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0208] Embodiment 7

[0209] According to an embodiment of the present application, there is also provided an embodiment of a method for processing an image. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0210] Figure 15 is a flowchart of a method for processing an image according to Embodiment 7 of the present application. As Figure 15 shown, the method may include the following steps:

[0211] Step S1502, the cloud server receives a prediction request sent by the client.

[0212] Among them, the prediction request carries at least identification information characterizing the image to be predicted.

[0213] The cloud server in the above step is a simple, efficient, secure, reliable, and computationally scalable computing server. Its management method is simpler and more efficient than that of a physical server. Users can quickly create or release any number of cloud servers without having to purchase hardware in advance.

[0214] The identification information in the above step may be the address information of the image to be predicted.

[0215] In an alternative embodiment, the cloud server may obtain the image to be predicted from the database according to the identification information carried in the prediction request, which characterizes the image to be predicted.

[0216] The client in the above steps may be a terminal device, an interactive flat panel, a computer device, etc.

[0217] Step S1504, the cloud server obtains the image to be predicted based on the identification information.

[0218] In an alternative embodiment, the storage location of the image to be predicted may be determined according to the address information recorded in the identification information, and then the image to be predicted is obtained.

[0219] Step S1506, the cloud server uses a regression network model to perform illumination prediction on the image to be predicted, and predicts the illumination panorama parameters of the image to be predicted.

[0220] Among them, the illumination panorama parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall illumination intensity of the main light source, and the ambient light intensity.

[0221] Among them, the regression network model is a model generated by training a neural network model with sample images, and the sample images are at least a part of the illumination panorama.

[0222] In an alternative embodiment, after the cloud server performs illumination prediction on the image to be predicted, it may feedback the obtained illumination panorama parameters to the client so that the user can view them.

[0223] Furthermore, the user may modify the illumination panorama parameters on the client and upload the modified illumination panorama parameters to the cloud server, so that the cloud server can optimize the regression network model according to the modified illumination panorama parameters.

[0224] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0225] Embodiment 8

[0226] According to an embodiment of the present application, there is also provided an image processing device for implementing the above image processing method, as Figure 16 shown, the device 1600 includes: a receiving module 1602, an obtaining module 1604, and a prediction module 1606.

[0227] Among them, a receiving module is configured to receive a prediction request sent by a client, where the prediction request carries at least identification information characterizing an image to be predicted; an obtaining module is configured to obtain the image to be predicted based on the identification information; a prediction module is configured to perform illumination prediction on the image to be predicted by using a regression network model, and predict illumination panorama parameters of the image to be predicted, where the illumination panorama parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall illumination intensity of the main light source, and the ambient light intensity; among them, the regression network model is a model generated by training a neural network model with sample images, and the sample images are at least a part of the illumination panorama.

[0228] It should be noted here that the above receiving module 1602, obtaining module 1604, and prediction module 1606 correspond to steps S1502 to S1506 in Embodiment 7. The instances and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0229] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0230] Embodiment 9

[0231] According to an embodiment of the present application, an embodiment of an image processing method is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0232] Figure 17 is a flowchart of an image processing method according to Embodiment 9 of the present application. As Figure 17 shown, the method may include the following steps:

[0233] Step S1702, receive an image to be predicted sent by a client.

[0234] The above steps can be executed by a cloud server. The client in the above steps can be a terminal device, an interactive tablet, a computer device, etc., but is not limited thereto.

[0235] Data interaction can be performed between the cloud server and the client through a specific interface. The client can pass the image to be predicted, or the identification information of the image to be predicted, or the storage address of the image to be predicted into the interface function as a parameter to achieve the purpose of uploading the image to be predicted to the cloud server.

[0236] For example, to facilitate the user to upload the image to be predicted, the client can provide an interactive interface as shown in Figure 3 Figure, and the user can operate on the interactive interface, select the image to be predicted and upload it. At this time, the image to be predicted can be displayed in the "Image Display" area for the user to view.

[0237] Step S1704: Use the regression network model to perform illumination prediction on the image to be predicted, and predict the illumination panorama parameters of the image to be predicted.

[0238] Among them, the illumination panorama parameters at least include: the spherical distribution of the main light source, the overall illumination intensity of the main light source, and the ambient light intensity in the image to be predicted.

[0239] In an alternative embodiment, the cloud server can use the regression network model to perform illumination prediction on the image to be predicted, and decompose the illumination panorama into three parameters: the spherical distribution of the main light source, the overall illumination intensity, and the ambient light intensity.

[0240] Step S1706: Use the illumination panorama parameters to reconstruct the illumination panorama of the image to be predicted.

[0241] Among them, the regression network model is a model generated by training a neural network model with sample images, and the sample images are at least a part of the illumination panorama.

[0242] In an alternative embodiment, after predicting the illumination panorama parameters, the cloud server can directly use the predicted parameters to automatically reconstruct the illumination panorama.

[0243] In another alternative embodiment, in order to reconstruct a high-precision illumination panorama, the cloud server can also use the adversarial network to optimize the illumination panorama of the image to be predicted. The specific processing is the same as that in the above embodiment and will not be elaborated here.

[0244] Step S1708: Output the illumination panorama to the client.

[0245] In an alternative embodiment, the cloud server can return the reconstructed illumination panorama to the client, and the client can display it for the user to view, so that the user can simultaneously see the image to be predicted and the illumination panorama in the interactive interface.

[0246] In another alternative embodiment, the cloud server can not only return the light panoramic image to the client, but also return the light panoramic image parameters to the client, so that the user can adjust or modify the light panoramic image parameters, further upload the modified light panoramic image parameters to the cloud server, and the cloud server can reconstruct a new light panoramic image based on the modified light panoramic image parameters, and the cloud server can optimize the regression network model according to the modified light panoramic image parameters to achieve the purpose of improving the performance of the cloud server.

[0247] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0248] Embodiment 10

[0249] According to an embodiment of the present application, there is also provided an image processing device for implementing the above image processing method, as Figure 18 shown. The device 1800 includes: a receiving module 1802, a prediction module 1804, a reconstruction module 1806, and an output module 1808.

[0250] Among them, the receiving module is used to receive the image to be predicted sent by the client; the prediction module is used to perform light prediction on the image to be predicted by using a regression network model, and predict the light panoramic image parameters of the image to be predicted, where the light panoramic image parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall light intensity of the main light source, and the ambient light intensity; the reconstruction module is used to reconstruct the light panoramic image of the image to be predicted by using the light panoramic image parameters; the output module is used to output the light panoramic image to the client; among them, the regression network model is a model generated by training a neural network model with sample images, and the sample images are at least a part of the light panoramic image.

[0251] It should be noted here that the above receiving module 1802, prediction module 1804, reconstruction module 1806, and output module 1808 correspond to steps S1702 to S1708 in Embodiment 9. The instances and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as a part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0252] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0253] Embodiment 11

[0254] Embodiments of the present application may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above computer terminal may also be replaced with a terminal device such as a mobile terminal.

[0255] Optionally, in this embodiment, the above computer terminal may be located in at least one of multiple network devices in a computer network.

[0256] In this embodiment, the above computer terminal may execute the program code of the following steps in the image processing method: obtaining at least one sample image, where the sample image is a light panoramic image; training a neural network model using the sample image to obtain a regression network model; using the regression network model to perform light prediction on the image to be predicted, and predicting the light panoramic image parameters of the image to be predicted, where the light panoramic image parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall light intensity of the main light source, and the ambient light intensity.

[0257] Optionally, Figure 14 is a structural block diagram of a computer terminal according to an embodiment of the present application. As Figure 14 shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 1402, and 1404.

[0258] Among them, the memory may be used to store software programs and modules, such as the program instructions / modules corresponding to the image processing method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above image processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the terminal A through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0259] The processor may call the information and application programs stored in the memory through the transmission device to execute the following steps: obtaining at least one sample image, where the sample image is a light panoramic image; training a neural network model using the sample image to obtain a regression network model; using the regression network model to perform light prediction on the image to be predicted, and predicting the light panoramic image parameters of the image to be predicted, where the light panoramic image parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall light intensity of the main light source, and the ambient light intensity.

[0260] Optionally, the above-mentioned processor may also execute the program code of the following steps: obtaining the image data of the main light source in the sample image; obtaining a plurality of anchor points set on the unit sphere; matching the image data of the main light source with each anchor point, and based on the matching result, obtaining the spherical distribution of the main light source in the sample image.

[0261] Optionally, the above-mentioned processor may also execute the program code of the following steps: extracting the pixels in the sample image according to the illumination intensity, wherein the pixels with the illumination intensity greater than or equal to the light source threshold belong to the image data of the main light source, and the pixels with the illumination intensity less than the light source threshold belong to the image data of the ambient light.

[0262] Optionally, the above-mentioned processor may also execute the program code of the following steps: clustering based on the distances between the image data of the main light source and each anchor point to obtain the image data of the main light source that matches any one of the anchor points; summing up the illumination intensities of the image data of the main light source that matches the anchor point to obtain the light intensity value of the anchor point; normalizing the light intensity value of each anchor point to generate the spherical distribution of the main light source in the sample image.

[0263] Optionally, the above-mentioned processor may also execute the program code of the following steps: using the first loss function to perform regression summation on the illumination intensities of the pixels belonging to the main light source to obtain the overall illumination intensity of the main light source in the sample image; using the second loss function to perform regression to obtain the average value of the illumination intensities of the pixels belonging to the ambient light to obtain the ambient light intensity of the ambient light in the sample image.

[0264] Optionally, the above-mentioned processor may also execute the program code of the following steps: using the illumination panorama parameters to reconstruct the illumination panorama of the image to be predicted.

[0265] Optionally, the above-mentioned processor may also execute the program code of the following steps: using an adversarial network to optimize the illumination panorama of the image to be predicted to generate a high-precision illumination panorama.

[0266] Optionally, the above-mentioned processor may also execute the program code of the following steps: obtaining a local image of the image to be predicted; using an encoder to vectorize the local image of the image to be predicted to obtain a set of latent space vectors; inputting the latent space vectors into the adversarial network to superimpose and optimize the resolution of the illumination panorama of the image to be predicted to generate a high-precision illumination panorama.

[0267] Optionally, the above-mentioned processor may also execute the program code of the following steps: using the high-precision illumination panorama as a sample image to perform regression training on the neural network model to optimize the training of the neural network model.

[0268] As an alternative example, the processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: display the video being played and the target object to be added to the video in the interface, where the video is composed of multiple frames of images, and each frame of image forms an image sequence according to the loaded timestamp; use a regression network model to perform light prediction on each frame of image, and predict the light panorama parameters of each frame of image, where the regression network model is generated by training a neural network model with the light panorama as a sample; when it is detected that the target object is loaded into any frame of image, use the light panorama parameters of this frame of image to perform light rendering on the target object, where the light panorama parameters at least include: the spherical distribution of the main light source, the overall light intensity of the main light source, and the ambient light intensity.

[0269] As an alternative example, the processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: display a picture and the target object that loads the picture in the interface; use a regression network model to perform light prediction on the picture, and predict the light panorama parameters of the picture, where the regression network model is generated by training a neural network model with the light panorama as a sample; use the light panorama parameters of the picture to perform light rendering on the target object, where the light panorama parameters at least include: the spherical distribution of the main light source in the picture, the overall light intensity of the main light source, and the ambient light intensity.

[0270] As an alternative example, the processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: the cloud server receives a prediction request sent by the client, where the prediction request at least carries identification information characterizing the image to be predicted; the cloud server obtains the image to be predicted based on the identification information; the cloud server uses a regression network model to perform light prediction on the image to be predicted, and predicts the light panorama parameters of the image to be predicted, where the light panorama parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall light intensity of the main light source, and the ambient light intensity; where the regression network model is a model generated by training a neural network model with sample images, and the sample images are at least a part of the light panorama.

[0271] As an alternative example, the processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: receiving the image to be predicted sent by the client; using a regression network model to perform illumination prediction on the image to be predicted, and predicting the illumination panoramic image parameters of the image to be predicted, where the illumination panoramic image parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall illumination intensity of the main light source, and the ambient light intensity; using the illumination panoramic image parameters to reconstruct the illumination panoramic image of the image to be predicted; outputting the illumination panoramic image to the client; where the regression network model is a model generated by training a neural network model with sample images, and the sample images are at least a part of the illumination panoramic image.

[0272] Those of ordinary skill in the art can understand that Figure 14 the structure shown is only illustrative, and the computer terminal can also be a smart phone (such as an android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a mobile internet device (MID), a PAD and other terminal devices. Figure 14 It does not limit the structure of the above electronic device. For example, computer terminal A may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 14 or have a different configuration from that shown in Figure 14 the figure.

[0273] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.

[0274] Embodiment 12

[0275] The embodiments of the present application further provide a computer-readable storage medium. Optionally, in this embodiment, the above storage medium can be used to store the program code executed by the image processing method provided in the above embodiment.

[0276] Optionally, in this embodiment, the above storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0277] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining at least one sample image, where the sample image is a light panoramic image; training a neural network model using the sample image to obtain a regression network model; using the regression network model to perform light prediction on the image to be predicted, and predicting the light panoramic image parameters of the image to be predicted, where the light panoramic image parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall light intensity of the main light source, and the ambient light intensity.

[0278] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining the image data of the main light source in the sample image; obtaining a plurality of anchor points set on the unit sphere; matching the image data of the main light source with each anchor point, and based on the matching result, obtaining the spherical distribution of the main light source in the sample image.

[0279] Optionally, the storage medium is further configured to store program code for performing the following steps: extracting pixels in the sample image according to the light intensity, where pixels with a light intensity greater than or equal to the light source threshold belong to the image data of the main light source, and pixels with a light intensity less than the light source threshold belong to the image data of the ambient light.

[0280] Optionally, the storage medium is further configured to store program code for performing the following steps: clustering based on the distance between the image data of the main light source and each anchor point, obtaining the image data of the main light source that matches any one of the anchor points; summing the light intensities of the image data of the main light source that matches the anchor point to obtain the light intensity value of the anchor point; normalizing the light intensity value of each anchor point to generate the spherical distribution of the main light source in the sample image.

[0281] Optionally, the storage medium is further configured to store program code for performing the following steps: using a first loss function to perform regression summation on the light intensities of the pixels belonging to the main light source to obtain the overall light intensity of the main light source in the sample image; using a second loss function to perform regression averaging on the light intensities of the pixels belonging to the ambient light to obtain the ambient light intensity of the ambient light in the sample image.

[0282] Optionally, the storage medium is further configured to store program code for performing the following steps: using the light panoramic image parameters to reconstruct the light panoramic image of the image to be predicted.

[0283] Optionally, the storage medium is further configured to store program code for performing the following steps: using an adversarial network to optimize the light panoramic image of the image to be predicted to generate a high-precision light panoramic image.

[0284] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining a local image of the image to be predicted; vectorizing the local image of the image to be predicted by using an encoder to obtain a set of latent space vectors; inputting the latent space vectors into an adversarial network to superimpose and optimize the resolution of the illumination panorama of the image to be predicted, and generating a high-precision illumination panorama.

[0285] Optionally, the storage medium is further configured to store program code for performing the following steps: using the high-precision illumination panorama as a sample image to perform regression training on a neural network model, and optimizing the trained neural network model.

[0286] As an optional example, the storage medium is further configured to store program code for performing the following steps: displaying a playing video and a target object to be added to the video in an interface, where the video is composed of multiple frames of images, and each frame of image forms an image sequence according to the loaded timestamp; using a regression network model to perform illumination prediction on each frame of image, and predicting the illumination panorama parameters of each frame of image, where the regression network model is generated by training a neural network model with the illumination panorama as a sample; when it is detected that the target object is loaded into any frame of image, using the illumination panorama parameters of this frame of image to perform illumination rendering on the target object, where the illumination panorama parameters at least include: the spherical distribution of the main light source, the overall illumination intensity of the main light source, and the ambient light intensity.

[0287] As an optional example, the storage medium is further configured to store program code for performing the following steps: displaying a picture and a target object for loading the picture in an interface; using a regression network model to perform illumination prediction on the picture, and predicting the illumination panorama parameters of the picture, where the regression network model is generated by training a neural network model with the illumination panorama as a sample; using the illumination panorama parameters of the picture to perform illumination rendering on the target object, where the illumination panorama parameters at least include: the spherical distribution of the main light source in the picture, the overall illumination intensity of the main light source, and the ambient light intensity.

[0288] As an optional example, the storage medium is further configured to store program code for performing the following steps: a cloud server receives a prediction request sent by a client, where the prediction request at least carries identification information characterizing the image to be predicted; the cloud server obtains the image to be predicted based on the identification information; the cloud server uses a regression network model to perform illumination prediction on the image to be predicted, and predicts the illumination panorama parameters of the image to be predicted, where the illumination panorama parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall illumination intensity of the main light source, and the ambient light intensity; where the regression network model is a model generated by training a neural network model with a sample image, and the sample image is at least a part of the illumination panorama.

[0289] As an alternative example, the storage medium is further configured to store program code for performing the following steps: receiving a to-be-predicted image sent by a client; performing illumination prediction on the to-be-predicted image by using a regression network model to predict illumination panoramic map parameters of the to-be-predicted image, where the illumination panoramic map parameters at least include: the spherical distribution of the main light source in the to-be-predicted image, the overall illumination intensity of the main light source, and the ambient light intensity; reconstructing an illumination panoramic map of the to-be-predicted image by using the illumination panoramic map parameters; outputting the illumination panoramic map to the client; where the regression network model is a model generated by training a neural network model with sample images, and the sample images are at least a part of the illumination panoramic map.

[0290] Embodiment 13

[0291] According to an embodiment of the present application, there is also provided an image processing system, including:

[0292] a processor; and

[0293] a memory, connected to the processor, for providing instructions for the processor to perform the following processing steps: obtaining at least one sample image, where the sample image is an illumination panoramic map; training a neural network model by using the sample image to obtain a regression network model; performing illumination prediction on a to-be-predicted image by using the regression network model to predict illumination panoramic map parameters of the to-be-predicted image, where the illumination panoramic map parameters at least include: the spherical distribution of the main light source in the to-be-predicted image, the overall illumination intensity of the main light source, and the ambient light intensity.

[0294] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0295] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.

[0296] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0297] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of units or modules may be in an electrical or other form.

[0298] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0299] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0300] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drive, read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk or optical disc and other various media that can store program codes.

[0301] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for processing an image, characterized in that, Comprising: Obtain at least one sample image, wherein the sample image is a panoramic illumination image; Train a neural network model using the sample image to obtain a regression network model; Perform illumination prediction on the image to be predicted using the regression network model, and predict the panoramic illumination image parameters of the image to be predicted. The panoramic illumination image parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall illumination intensity of the main light source, and the ambient light intensity; Wherein, the method further includes: Obtain the image data of the main light source in the sample image and a plurality of anchor points set on the unit sphere; Match the image data of the main light source with each anchor point, and based on the matching result, obtain the spherical distribution of the main light source.

2. The method according to claim 1, wherein: Extract pixels in the sample image according to the illumination intensity. Among them, the pixels with the illumination intensity greater than or equal to the light source threshold belong to the image data of the main light source, and the pixels with the illumination intensity less than the light source threshold belong to the image data of the ambient light.

3. The method according to claim 2, wherein Matching the image data of the main light source with each anchor point, and based on the matching result, obtaining the spherical distribution of the main light source, includes: Cluster based on the distance between the image data of the main light source and each anchor point to obtain the image data of the main light source matching any one anchor point; Sum the illumination intensities of the image data of the main light source matching the anchor point to obtain the light intensity value of the anchor point; Normalize the light intensity value of each anchor point to generate the spherical distribution of the main light source in the sample image.

4. The method according to claim 2, wherein: Use the first loss function to perform regression summation on the illumination intensities of the pixels belonging to the main light source to obtain the overall illumination intensity of the main light source in the sample image; Use the second loss function to perform regression averaging on the illumination intensities of the pixels belonging to the ambient light to obtain the ambient light intensity of the ambient light in the sample image.

5. The method according to any one of claims 1 to 4, characterized in that, After performing illumination prediction on the image to be predicted using the regression network model and predicting the panoramic illumination image parameters of the image to be predicted, the method further includes: Use the panoramic illumination image parameters to reconstruct the panoramic illumination image of the image to be predicted.

6. The method according to claim 5, characterized in that Optimize the panoramic illumination image of the image to be predicted using an adversarial network to generate a high-precision panoramic illumination image.

7. The method according to claim 6, wherein Optimizing the panoramic illumination image of the image to be predicted using an adversarial network to generate a high-precision panoramic illumination image, includes: Obtain a local image of the image to be predicted; Perform vectorization processing on the local image of the image to be predicted using an encoder to obtain a set of latent space vectors; Input the latent space vectors into the adversarial network, and superimpose and optimize the resolution of the panoramic illumination image of the image to be predicted to generate a high-precision panoramic illumination image.

8. The method according to claim 7, wherein After optimizing the panoramic illumination image of the image to be predicted using an adversarial network to generate a high-precision panoramic illumination image, the method further includes: Use the high-precision panoramic illumination image as a sample image to perform regression training on the neural network model to optimize the training neural network model.

9. A method for processing an image, characterized in that Comprising: Display the playing video and the target object to be added to the video in the interface, where the video is composed of multiple frames of images, and each frame of image forms an image sequence according to the loaded timestamp; Use a regression network model to perform light prediction on each frame of the image, and predict the light panorama parameters of each frame of the image, where the regression network model is generated by training a neural network model with a light panorama as a sample; When it is detected that the target object is loaded into any frame of the image, use the light panorama parameters of this frame of the image to perform light rendering on the target object, where the light panorama parameters at least include: the spherical distribution of the main light source, the overall light intensity of the main light source, and the ambient light intensity; Wherein, the method further includes: Obtain the image data of the main light source in the sample image and multiple anchor points set on the unit sphere; Match the image data of the main light source with each anchor point, and based on the matching result, obtain the spherical distribution of the main light source.

10. A method for processing an image, characterized in that, Include: Display a picture and the target object that loads the picture in the interface; Use a regression network model to perform light prediction on the picture, and predict the light panorama parameters of the picture, where the regression network model is generated by training a neural network model with a light panorama as a sample; Use the light panorama parameters of the picture to perform light rendering on the target object, where the light panorama parameters at least include: the spherical distribution of the main light source in the picture, the overall light intensity of the main light source, and the ambient light intensity; Wherein, the method further includes: Obtain the image data of the main light source in the sample image and multiple anchor points set on the unit sphere; Match the image data of the main light source with each anchor point, and based on the matching result, obtain the spherical distribution of the main light source.

11. A method for processing an image, characterized in that, Include: The cloud server receives a prediction request sent by the client, where the prediction request carries at least identification information characterizing the image to be predicted; The cloud server obtains the image to be predicted based on the identification information; The cloud server uses a regression network model to perform light prediction on the image to be predicted, and predicts the light panorama parameters of the image to be predicted, where the light panorama parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall light intensity of the main light source, and the ambient light intensity; Wherein, the regression network model is a model generated by training a neural network model with a sample image, and the sample image is at least a part of the light panorama; Wherein, the method further includes: The cloud server obtains the image data of the main light source in the sample image and multiple anchor points set on the unit sphere; The cloud server matches the image data of the main light source with each anchor point, and based on the matching result, obtains the spherical distribution of the main light source.

12. A method for processing an image, characterized in that, Include: Receive the image to be predicted sent by the client; Use a regression network model to perform illumination prediction on the to-be-predicted image, and predict the illumination panorama parameters of the to-be-predicted image, where the illumination panorama parameters at least include: the spherical distribution of the main light source in the to-be-predicted image, the overall illumination intensity of the main light source, and the ambient light intensity; Use the illumination panorama parameters to reconstruct the illumination panorama of the to-be-predicted image; Output the illumination panorama to the client; Among them, the regression network model is a model generated by training a neural network model with sample images, and the sample images are at least a part of the illumination panorama; Among them, the method further includes: Obtain the image data of the main light source in the sample image and a plurality of anchor points set on the unit sphere; Match the image data of the main light source with each anchor point, and based on the matching result, obtain the spherical distribution of the main light source.

13. An image processing device, characterized in that, Include: An acquisition module, configured to acquire at least one sample image, where the sample image is an illumination panorama; A training module, configured to train a neural network model with the sample image to obtain a regression network model; A prediction module, configured to perform illumination prediction on the to-be-predicted image with the regression network model, and predict the illumination panorama parameters of the to-be-predicted image, where the illumination panorama parameters at least include: the spherical distribution of the main light source in the to-be-predicted image, the overall illumination intensity of the main light source, and the ambient light intensity; Among them, the device is further configured to obtain the image data of the main light source in the sample image and a plurality of anchor points set on the unit sphere; match the image data of the main light source with each anchor point, and based on the matching result, obtain the spherical distribution of the main light source.

14. An image processing apparatus, characterized in that, Include: A display module, configured to display a playing video and a target object to be added to the video in an interface, where the video is composed of multiple frames of images, and each frame of image forms an image sequence according to the loaded timestamp; A prediction module, configured to perform illumination prediction on each frame of image with a regression network model, and predict the illumination panorama parameters of each frame of image, where the regression network model is generated by training a neural network model with an illumination panorama as a sample; A rendering module, configured to, when detecting that the target object is loaded into any frame of image, perform illumination rendering on the target object with the illumination panorama parameters of this frame of image, where the illumination panorama parameters at least include: the spherical distribution of the main light source, the overall illumination intensity of the main light source, and the ambient light intensity; Among them, the device is further configured to obtain the image data of the main light source in the sample image and a plurality of anchor points set on the unit sphere; match the image data of the main light source with each anchor point, and based on the matching result, obtain the spherical distribution of the main light source.

15. An image processing apparatus, characterized in that, Include: A display module, configured to display a picture and a target object that loads the picture in an interface; A prediction module, configured to perform illumination prediction on the picture with a regression network model, and predict the illumination panorama parameters of the picture, where the regression network model is generated by training a neural network model with an illumination panorama as a sample; A rendering module, configured to perform lighting rendering on the target object by using the lighting panoramic image parameters of the picture, where the lighting panoramic image parameters at least include: the spherical distribution of the main light source in the picture, the overall lighting intensity of the main light source, and the ambient light intensity; Wherein, the device is further configured to obtain the image data of the main light source in the sample image and a plurality of anchor points set on the unit sphere; match the image data of the main light source with each anchor point, and based on the matching result, obtain the spherical distribution of the main light source.

16. An image processing apparatus, characterized in that, Comprising: A receiving module, configured to receive a prediction request sent by a client, where the prediction request at least carries identification information characterizing the image to be predicted; An obtaining module, configured to obtain the image to be predicted based on the identification information; A prediction module, configured to perform lighting prediction on the image to be predicted by using a regression network model, and predict the lighting panoramic image parameters of the image to be predicted, where the lighting panoramic image parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall lighting intensity of the main light source, and the ambient light intensity; Wherein, the regression network model is a model generated by training a neural network model with a sample image, and the sample image is at least a part of the lighting panoramic image; Wherein, the device is further configured to obtain the image data of the main light source in the sample image and a plurality of anchor points set on the unit sphere; match the image data of the main light source with each anchor point, and based on the matching result, obtain the spherical distribution of the main light source.

17. An image processing apparatus, characterized in that, Comprising: A receiving module, configured to receive the image to be predicted sent by a client; A prediction module, configured to perform lighting prediction on the image to be predicted by using a regression network model, and predict the lighting panoramic image parameters of the image to be predicted, where the lighting panoramic image parameters at least include: the spherical distribution of the main light source in the image to be predicted, the overall lighting intensity of the main light source, and the ambient light intensity; A reconstruction module, configured to reconstruct the lighting panoramic image of the image to be predicted by using the lighting panoramic image parameters; An output module, configured to output the lighting panoramic image to the client; Wherein, the regression network model is a model generated by training a neural network model with a sample image, and the sample image is at least a part of the lighting panoramic image; Wherein, the device is further configured to obtain the image data of the main light source in the sample image and a plurality of anchor points set on the unit sphere; match the image data of the main light source with each anchor point, and based on the matching result, obtain the spherical distribution of the main light source.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the image processing method according to any one of claims 1 to 12.

19. A computer terminal, characterized in that, Comprising a memory and a processor, the processor is configured to run the program stored in the memory, wherein when the program runs, it executes the image processing method according to any one of claims 1 to 12.

20. An image processing system, characterized in that, Comprising: A processor; And A memory, connected to the processor, for providing instructions for the processor to process the following processing steps: obtaining at least one sample image, where the sample image is a light panoramic image; training a neural network model using the sample image to obtain a regression network model; using the regression network model to perform light prediction on an image to be predicted, and predicting the light panoramic image parameters of the image to be predicted, where the light panoramic image parameters at least include: the spherical distribution of the main light source, the overall light intensity of the main light source, and the ambient light intensity in the image to be predicted; where the step further includes: obtaining the image data of the main light source in the sample image and a plurality of anchor points set on the unit sphere; matching the image data of the main light source with each anchor point, and based on the matching result, obtaining the spherical distribution of the main light source.

Citation Information

Patent Citations

  • Illumination system design method based on optical simulation and experimental facility

    CN106838723A

  • Illumination estimation method and device

    CN108805970A