Image processing method and device

Through the combination of image defog treatment and lighting prediction model, the problem of refined adjustment of portrait images in complex light scenes is solved, the clarity and reality of the image are improved, and the precise lighting effect under different lighting conditions is achieved.

CN120390156APending Publication Date: 2025-07-29VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510409074.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art is difficult to achieve refined adjustment of portrait images in complex light scenes, especially in scenes such as backlight, toplight, and sidelight. The images are prone to blurring and blurring, affecting clarity and color restoration.

Method used

By acquiring image and preset light source information, performing image defog processing, input the image to the light prediction model, use the trained model to predict the stereoscopic information and material information of the photographed object, and simulate lighting processing based on the preset light source information and material information to generate the target image.

Benefits of technology

It realizes fine adjustment of the image in complex light scenes, improves the clarity and reality of the image, meets the visual needs under different lighting conditions, and enhances the artistic effect of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390156A_ABST
    Figure CN120390156A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, and belongs to the technical field of image processing. The method comprises the following steps: acquiring a first image and preset light source information; performing image defogging processing on the first image to obtain a second image; inputting the second image into a lighting prediction model to obtain predicted three-dimensional information and predicted material information of the shot object in the second image; the lighting prediction model is obtained by training according to multiple groups of first training data, and each group of first training data comprises a first sample image, reference stereo information and reference material information; the plurality of first sample images are images corresponding to the same sample object in different illumination environments; and according to the preset light source information, the predicted stereo information and the predicted material information, performing simulated illumination processing on the second image to obtain a target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of communication technologies, and particularly relates to an image processing method and apparatus. Background Art

[0002] In contemporary life, photography with portraits as the main subject widely covers various human activities such as traveling and socializing. Photography is essentially an art of light, and light is of great significance to photographic works. After years of development, the shooting technology of mobile electronic devices has been able to meet the basic needs of portrait shooting, but users have higher-order requirements in terms of shooting records and creation in complex light scenarios.

[0003] Currently, mainstream portrait algorithms have obvious defects and cannot properly handle various complex light scenarios. For example, in backlight scenarios, portraits are prone to being blurred and out of focus. Therefore, it is currently difficult to achieve fine adjustment for complex light scenarios. Summary of the Invention

[0004] The objective of the embodiments of this application is to provide an image processing method and apparatus that can achieve fine adjustment for images captured in complex light scenarios.

[0005] In a first aspect, the embodiments of this application provide an image processing method, which includes:

[0006] Obtain a first image and preset light source information;

[0007] Perform image defogging processing on the first image to obtain a second image;

[0008] Input the second image into a lighting prediction model to obtain predicted three-dimensional information and predicted material information of the shooting object in the second image; the lighting prediction model is trained based on multiple groups of first training data, and each group of the first training data includes: a first sample image, reference three-dimensional information, and reference material information; the multiple first sample images are images corresponding to the same sample object in different lighting environments;

[0009] Perform simulated lighting processing on the second image according to the preset light source information, the predicted three-dimensional information, and the predicted material information to obtain a target image.

[0010] In a second aspect, the embodiments of this application provide an image processing apparatus, which includes:

[0011] An obtaining module, configured to obtain a first image and preset light source information;

[0012] A defogging processing module, configured to perform image defogging processing on the first image to obtain a second image;

[0013] A prediction module, configured to input the second image into a lighting prediction model to obtain predicted three-dimensional information and predicted material information of a shooting object in the second image; the lighting prediction model is trained according to multiple groups of first training data, and each group of the first training data includes: a first sample image, reference three-dimensional information, and reference material information; the multiple first sample images are images corresponding to the same sample object under different lighting environments;

[0014] A lighting processing module, configured to perform simulated lighting processing on the second image according to the preset light source information, the predicted three-dimensional information, and the predicted material information to obtain a target image.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0016] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0017] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.

[0018] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0019] In an embodiment of the present application, by obtaining a first image and preset light source information; performing image defogging processing on the first image to obtain a second image. Image defogging processing can improve the quality of the image, remove the influence of the foggy effect generated when shooting in a complex light scene on the image, make the image clearer, and provide a better basis for subsequent processing. Input the second image into the lighting prediction model. Since the lighting prediction model is trained according to multiple sets of first training data, each set of first training data includes: a first sample image, reference three-dimensional information, and reference material information. Multiple first sample images are images corresponding to the same sample object in different lighting environments. The lighting prediction model establishes a mapping relationship between image features and the three-dimensional information and material information of the shooting object by learning multiple sets of first training data. When the second image is input, the lighting prediction model can accurately predict the three-dimensional information and material information of the shooting object according to the learned rules. According to the preset light source information, predicted three-dimensional information, and predicted material information, performing simulated lighting processing on the second image can make the image present different lighting effects, meet the needs of users in different scenarios, make the obtained target image more in line with the visual perception of the human eye under specific lighting conditions, and achieve fine adjustment of images taken in complex light scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flowchart of an image processing method provided by an embodiment of the present application;

[0021] Figure 2 is a schematic structural diagram of a defogging model provided by an embodiment of the present application;

[0022] Figure 3 is a flowchart of another image processing method provided by an embodiment of the present application;

[0023] Figure 4 is a structural diagram of an image processing device provided by an embodiment of the present application;

[0024] Figure 5 is one of the schematic hardware structures of an electronic device provided by an embodiment of the present application;

[0025] Figure 6 is the second schematic hardware structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] Next, the technical solutions of the embodiments of the present application will be clearly described in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0027] The terms "first," "second," and the like in the specification of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in this specification refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0028] The image processing method provided in the embodiments of the present application can be applied to at least the following application scenarios, which are described below.

[0029] In complex lighting scenarios, such as backlit, toplit, sidelit, or with multiple light sources, halos are prone to occur during shooting due to reflections and scattering within the lens. Halos reduce image contrast and clarity, blurring the edges of subjects and creating a fuzzy effect. Halos also affect color reproduction, blurring the overall image.

[0030] For example, in a backlit scene, light comes from behind the subject, in the opposite direction of the camera. In this situation, tiny particles in the air, such as dust and water vapor, scatter the light. This scattered light creates a hazy effect in the image, making the subject appear blurry. Furthermore, scattered light reduces contrast, making it difficult to clearly render details in the subject, resulting in a blurred appearance.

[0031] To avoid overexposing the background, exposure compensation is often required when shooting against the light, reducing the amount of light entering. However, this can result in underexposing the subject, losing detail in shadows, and making the image appear gray and hazy. Furthermore, underexposure can affect camera image quality, reducing image sharpness and further exacerbating the appearance of blurred portraits.

[0032] The above situation can also be called picture fog. Picture fog describes the overall picture presenting a hazy and blurred effect, as if shrouded by a layer of mist. The details and clarity of the portrait are affected, and it looks blurry.

[0033] In response to the problems arising from related technologies, the embodiments of the present application provide an image processing method and device that can achieve fine-grained adjustment of images captured in complex lighting scenes.

[0034] The image processing method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0035] Figure 1The flowchart of an image processing method provided by an embodiment of this application. As Figure 1 shown, the image processing method may include step 110-step 140. This method is applied to an image processing device, specifically as follows: Step 110, obtain a first image and preset light source information;

[0036] First image: The initially obtained image, which is the starting data of the entire image processing process.

[0037] Preset light source information: The pre-set light source related information, such as the position, intensity, color, etc. of the light source, which is used for subsequent simulated lighting processing.

[0038] Obtain the first image as the basic data for processing, and at the same time obtain the preset light source information to provide parameters for subsequent simulated lighting.

[0039] Step 120, perform image defogging processing on the first image to obtain a second image;

[0040] Image defogging processing: An image processing technology aimed at removing the blur and color distortion effects in the image, and improving the clarity and contrast of the image.

[0041] Image defogging processing can improve the quality of the image, remove the influence of image fogging on the image, make the image clearer, and provide a better basis for subsequent prediction and processing.

[0042] Step 130, input the second image into a lighting prediction model to obtain the predicted three-dimensional information and predicted material information of the photographed object in the second image; the lighting prediction model is trained according to multiple groups of first training data, and each group of the first training data includes: a first sample image, reference three-dimensional information, and reference material information; the multiple first sample images are images corresponding to the same sample object in different lighting environments;

[0043] Lighting prediction model: A trained model used to predict the three-dimensional information and material information of the photographed object according to the input image. By learning the characteristics of the object in different lighting environments through a large amount of training data, accurate prediction can be achieved.

[0044] Predicted three-dimensional information: The three-dimensional characteristics of the photographed object predicted by the lighting prediction model according to the input image, such as the shape of the object, surface undulations, etc.

[0045] Predicted material information: The material attributes of the photographed object predicted by the lighting prediction model, such as whether the object is metal, plastic, or glass, etc. Different materials have different reflection and refraction characteristics of light.

[0046] The lighting prediction model establishes a mapping relationship between image features and the three-dimensional information and material information of the photographed object by learning multiple sets of first training data. When the second image is input, the model predicts the three-dimensional information and material information of the photographed object according to the learned rules.

[0047] Step 140: According to the preset light source information, the predicted three-dimensional information, and the predicted material information, perform simulated lighting processing on the second image to obtain a target image.

[0048] Simulated lighting processing: According to the preset light source information, the predicted three-dimensional information, and the material information, simulate the lighting effect on the image to make the image present the effect under specific lighting conditions.

[0049] Target image: The finally obtained image after processing, with the effect after simulated lighting processing.

[0050] According to the preset light source information, the predicted three-dimensional information, and the predicted material information, use the optical principle to simulate the reflection, refraction, etc. of light on the surface of the photographed object, and adjust the lighting effect of the second image to obtain a target image.

[0051] Through image dehazing processing, the fogging interference in the image is removed, making the image clearer and facilitating subsequent processing and observation. The lighting prediction model can accurately predict the three-dimensional information and material information of the photographed object, providing an accurate basis for simulated lighting processing. Performing simulated lighting processing according to the preset light source information can make the image present different lighting effects, meeting the needs of users in different scenarios, such as simulating special lighting effects such as backlight and side light. Through simulated lighting processing, the image is made more in line with the visual perception of the human eye under specific lighting conditions, enhancing the realism and artistic sense of the image.

[0052] In a possible embodiment, when the first image is a portrait, in step 120, it may specifically include the following steps:

[0053] Step 210: Extract the person mask image and skin mask image of the photographed object in the first image;

[0054] Step 220: Input the first image into the dehazing model to obtain a third image; the dehazing model is trained according to multiple sets of second training data, and each set of the second training data includes: a second sample image and a reference image, and the reference image is an image obtained by performing dehazing processing on the second sample image;

[0055] Step 230: Process the third image according to the first image, the person mask image, and the skin mask image to obtain a second image.

[0056] Person mask image: An image that identifies the subject area in the first image. In the mask image, the person and the background are usually distinguished in a specific way, such as a black-and-white binary image, where white represents the person area and black represents the background area, etc., to facilitate separate processing of the person area.

[0057] Skin mask image: A mask image that identifies the skin area of a person in the first image. It is a further refinement of the person mask image and can more precisely locate the skin part of the person for targeted image processing of the skin area.

[0058] Defogging model: A deep learning-based model that learns the rules and methods of defogging by learning a large amount of second training data, so as to be able to perform defogging processing on the input foggy image.

[0059] Third image: The preliminary defogged result image obtained after the first image is input into the defogging model, but there may still be some areas that need further processing to better meet the subsequent processing requirements.

[0060] Step 210 involved: Through image segmentation techniques, such as deep learning-based semantic segmentation methods or traditional image segmentation algorithms, extract the person mask image and the skin mask image from the first image, respectively identifying the person area and the skin area of the person, providing a basis for subsequent processing of specific areas.

[0061] Step 220 involved: Input the first image into the defogging model that has been trained with multiple groups of second training data. The defogging model processes the first image according to the learned defogging patterns and rules, removing the foggy components in the image to obtain the third image.

[0062] Step 230 involved: Process the third image according to the first image, the person mask image, and the skin mask image. Possible processing methods include fine-tuning the person area and the skin area in the third image according to the mask image, such as further enhancing the contrast of the person area and repairing possible over-defogging or under-defogging problems in the skin area, etc., so as to obtain the final second image, making it more suitable for subsequent processing.

[0063] In the embodiments of this application, the process of processing the third image according to the first image, the person mask image, and the skin mask image to obtain the second image can be seen in formula (1):

[0064] Y = (A*x_out+(1-A)*x_input)*skin_mask + x_out*(1-skin_mask) (1)

[0065] Where Y is the second image;

[0066] x_out is the third image output by the defogging model;

[0067] x_input is the first image;

[0068] skin_mask is the skin mask image;

[0069] A is a parameter with a value range from 0 to 1, and the controllability of the final effect is achieved by adjusting the parameter A.

[0070] By extracting the person mask image and the skin mask image, the person and the person's skin area can be accurately located, enabling subsequent processing to be more targeted, avoiding unnecessary processing of irrelevant areas such as the background, and improving the processing efficiency and effect.

[0071] The defogging model can perform effective defogging processing on the first image according to the defogging mode learned from the training data, obtain a relatively clear third image, and improve the overall quality of the image.

[0072] Further processing the third image based on the first image, the person mask image, and the skin mask image can optimize the person and skin areas, making the performance of the person in the image more natural and clear, enhancing the visual effect of the image, and providing a better base image for subsequent light prediction and simulated lighting processing.

[0073] In a possible embodiment, before step 220, it further includes:

[0074] Obtain multiple groups of second training data;

[0075] Input the second sample image into the second neural network to output a sample defogged image. The second neural network includes: N downsampling convolutional layers and N upsampling convolutional layers, where N is an integer greater than 1;

[0076] Adjust the model parameters of the second neural network according to the sample defogged image and the reference image until the preset training conditions are met to obtain a defogging model.

[0077] Second training data: A dataset for training the defogging model, each group containing a foggy second sample image and the corresponding defogged reference image, providing data support for the model to learn the defogging law.

[0078] Second neural network: A structure composed of multiple layers of neural networks, including N downsampling convolutional layers and N upsampling convolutional layers, used for feature extraction and reconstruction of images to achieve the defogging function.

[0079] Downsampling Convolutional Layer: In a neural network, through operations such as convolution and pooling, the size of the feature map is gradually reduced while the number of channels of the feature map is increased, thereby extracting high-level abstract features of the image.

[0080] Upsampling Convolutional Layer: Contrary to the downsampling convolutional layer, the upsampling convolutional layer is used to gradually increase the size of the feature map, restore the spatial resolution of the image, and combine the feature information extracted during the downsampling process to reconstruct the dehazed image.

[0081] Sample Dehazed Image: The dehazed result image output after the second sample image is processed by the second neural network. During the training process, it will be compared with the reference image to adjust the model parameters.

[0082] Model Parameters: The parameters that need to be learned and adjusted in the second neural network, such as the weights and biases of the convolutional kernels. By continuously adjusting these parameters, the model can better complete the dehazing task.

[0083] Preset Training Conditions: Predetermined training termination conditions, such as reaching a certain number of training epochs, the loss function value being lower than a certain threshold, etc., which are used to control the training process of the model.

[0084] Collect a large number of hazy second sample images and the corresponding dehazed reference images as the basic data for training the dehazing model. These data contain various different scenarios and different degrees of image fogging, enabling the model to learn more comprehensive dehazing rules.

[0085] Image dehazing in backlight environments essentially aims to solve the problem of image quality degradation caused by both image fogging and backlight, such as reduced contrast, color distortion, and blurred details. Image fogging is manifested in the image as the scattering and absorption of light, causing the light reflected by objects to attenuate and scatter during propagation, thereby affecting the clarity and contrast of the image.

[0086] The dehazing process of the second neural network mainly includes: Based on the atmospheric scattering model, which describes the propagation law of light in a foggy environment, by reverse deduction of the model, an attempt is made to restore the image in a fog-free situation. Utilizing the powerful learning ability of the neural network, through a large amount of second training data for training, enabling the model to learn the influence pattern of image fogging on the image and the features and rules of how to restore a clear image from a foggy image.

[0087] Adopting downsampling and upsampling structures in the neural network can extract image features at different scales. Downsampling can extract high-level abstract features of the image, while upsampling restores the spatial resolution of the image. Through feature splicing operations, feature information at different scales is fused, thereby more comprehensively restoring the details and quality of the image.

[0088] The second neural network consists of N downsampling convolutional layers and N upsampling convolutional layers. The downsampling convolutional layers extract features from the input second sample image, gradually extracting high-level abstract features of the image while reducing the size of the feature map. The upsampling convolutional layers then gradually increase the size of the feature map based on the feature information extracted during the downsampling process to reconstruct the sample defogged image.

[0089] The sample defogged image is compared with the reference image, and the difference between them is calculated, usually measured using a loss function, such as the mean square error loss function. According to the value of the loss function, an optimization algorithm is used to adjust the model parameters of the second neural network, making the sample defogged image gradually approach the reference image. This process is continuously iterated until the preset training conditions are met, and the model obtained at this time is the defogging model.

[0090] Through a large amount of training data and the learning ability of the neural network, a model with strong defogging ability can be constructed. This model can effectively defog images with different degrees of fogging in different scenarios, improving the clarity and quality of the images.

[0091] The structures of the N downsampling convolutional layers and N upsampling convolutional layers can fully extract the feature information of the image and combine these features during the reconstruction process, enabling the model to adapt to the characteristics of different types of images and improving the generality and accuracy of defogging. By continuously adjusting the model parameters, the model can automatically learn the rules and patterns of defogging without the need to manually design complex defogging algorithms, improving the efficiency and effect of defogging processing.

[0092] Among them, in the steps of obtaining multiple groups of second training data mentioned above, it can specifically include the following steps:

[0093] Obtain multiple groups of corresponding second sample images and initial reference images, where the initial reference image is an image obtained by performing initial defogging processing on the second sample image;

[0094] Determine the brightness difference parameter value according to the second sample image and the initial reference image;

[0095] Adjust the brightness of the initial reference image according to the brightness difference parameter value to obtain the reference image.

[0096] Initial reference image: An image obtained by performing initial defogging processing on the second sample image. The initial defogging processing can use some basic defogging algorithms, such as defogging algorithms based on physical models.

[0097] Brightness difference parameter value: A numerical value obtained by comparing the second sample image and the initial reference image, which is used to measure the brightness difference between the two. This value reflects the relationship between the brightness of the image after the initial haze removal process and the brightness of the original image.

[0098] Reference image: An image obtained by adjusting the brightness of the initial reference image according to the brightness difference parameter value. It will be used as the target image in the training of the haze removal model to guide the model to learn how to convert a hazy image into a clear and properly bright image.

[0099] Collect a large number of hazy images in backlight environments as the second sample images. The second sample images have different degrees of image fogginess, scenes, and backlight conditions. Perform an initial haze removal process on the second sample images, such as using a traditional haze removal algorithm based on the atmospheric scattering model.

[0100] Perform pixel-by-pixel or region-by-region brightness analysis on the second sample image and the initial reference image. The brightness difference can be measured by calculating their average brightness, histogram, etc. For example, the average brightness of the second sample image and the average brightness of the initial reference image can be calculated, and then their ratio can be used as the brightness difference parameter value. The brightness difference parameter value represents the brightness change ratio of the initial reference image relative to the second sample image.

[0101] Adjust the brightness of the initial reference image according to the brightness difference parameter value. If the brightness difference parameter value is greater than 1, it means the initial reference image is relatively dark and its brightness needs to be increased; if the brightness difference parameter value is less than 1, its brightness needs to be decreased. The specific brightness adjustment method can be to multiply each pixel value of the initial reference image by the brightness difference parameter value, or a more complex image content-based brightness adjustment algorithm can be adopted to ensure that the brightness of the adjusted image is more natural and appropriate.

[0102] By adjusting the brightness of the initial reference image, the reference image is made closer in brightness to the visual effect that the original hazy image should have. This helps the haze removal model learn a more accurate haze removal pattern during training and avoid training biases caused by brightness differences. Considering the complexity of image brightness in backlight environments, this brightness adjustment process can enable the model to better adapt to the haze removal requirements of images under different backlight conditions. The adjusted reference image can more realistically reflect the characteristics of a haze-free image, thereby improving the haze removal effect and generalization ability of the model in practical applications.

[0103] Use the reference image with adjusted brightness for training. When the haze removal model processes new backlight hazy images, it can generate haze-free images with more natural brightness and higher quality, avoiding the problem of the haze-free image being too bright or too dark, and enhancing the user's visual experience.

[0104] In a possible embodiment, the second neural network at least includes: a first downsampling convolutional layer, a second downsampling convolutional layer, a first upsampling convolutional layer, and a second upsampling convolutional layer. Inputting the second sample image into the second neural network and outputting a sample dehazed image includes:

[0105] Input the second sample image into the first downsampling convolutional layer to obtain a downsampled feature map of the first size;

[0106] Input the feature map of the first size into the second downsampling convolutional layer to obtain a downsampled feature map of the second size;

[0107] Input the feature map of the second size into the first upsampling convolutional layer to obtain an upsampled feature map of the first size;

[0108] Perform splicing processing on the downsampled feature map of the first size and the upsampled feature map of the first size to obtain a spliced feature map;

[0109] Input the spliced feature map into the second upsampling convolutional layer to obtain a sample dehazed image.

[0110] First downsampling convolutional layer: The first downsampling convolutional layer in the second neural network, which performs preliminary feature extraction and size reduction operations on the input second sample image to obtain a downsampled feature map of the first size.

[0111] Second downsampling convolutional layer: The downsampling convolutional layer after the first downsampling convolutional layer, which further performs feature extraction and size reduction on the downsampled feature map of the first size to obtain a smaller downsampled feature map of the second size.

[0112] First upsampling convolutional layer: The first convolutional layer to start the upsampling operation, which enlarges the size of the downsampled feature map of the second size and restores it to the first size to obtain an upsampled feature map of the first size.

[0113] Second upsampling convolutional layer: The last upsampling convolutional layer, which processes the spliced feature map, further restores the details and spatial resolution of the image, and outputs a sample dehazed image.

[0114] Downsampled feature map: The feature map obtained after being processed by the downsampling convolutional layer, with a gradually decreasing size and a possible increase in the number of channels, containing high-level abstract features of the image.

[0115] Upsampled feature map: The feature map obtained after being processed by the upsampling convolutional layer, with a gradually increasing size, combining the feature information in the previous downsampling process, and used for image reconstruction.

[0116] First, input the second sample image into the first downsampling convolutional layer. Through operations such as convolutional operations and pooling operations, extract the preliminary features of the image while reducing the size of the feature map to obtain a downsampled feature map of the first size. This process can extract some local features of the image, such as edges and textures.

[0117] Next, input the downsampled feature map of the first size into the second downsampling convolutional layer to further extract more advanced abstract features and reduce the size of the feature map again to obtain a downsampled feature map of the second size. As the downsampling progresses, the size of the feature map continuously shrinks, and the number of channels may increase, enabling the model to learn more macroscopic and abstract image features.

[0118] Input the downsampled feature map of the second size into the first upsampling convolutional layer. Through upsampling methods such as transposed convolution or interpolation, restore the size of the feature map to the first size to obtain an upsampled feature map of the first size. The upsampling process is to gradually restore the spatial resolution of the image.

[0119] Perform splicing processing on the downsampled feature map of the first size and the upsampled feature map of the first size. The downsampled feature map contains rich image feature information but has a low spatial resolution; the upsampled feature map restores a certain spatial resolution but may lose some detailed features. Splice the downsampled feature map of the first size and the upsampled feature map of the first size in the channel dimension, enabling subsequent processing to utilize both the rich features extracted during the downsampling process and the spatial information restored during the upsampling process. Through splicing, the advantages of both can be combined to obtain more comprehensive feature information.

[0120] Input the spliced feature map into the second upsampling convolutional layer to further restore the details and spatial resolution of the image and finally output the sample dehazed image.

[0121] It can be understood that the second neural network may also at least include: the first downsampling convolutional layer, the second downsampling convolutional layer, the third downsampling convolutional layer, the first upsampling convolutional layer, the second upsampling convolutional layer, and the third upsampling convolutional layer. The following will be described in conjunction with Figure 2 for illustration:

[0122] For a second sample image of 1024×1024×3, these three numbers represent the following meanings respectively:

[0123] 1024: It means that the number of pixels contained in the image in the horizontal direction is 1024.

[0124] 1024: It means that the number of pixels contained in the image in the vertical direction is 1024.

[0125] 3: It means that each pixel is composed of 3 channels, corresponding to red, green, and blue respectively, namely the RGB color model. Through different numerical combinations of these three channels, various colors can be represented, thus constituting a rich and colorful image.

[0126] 1024×1024×3 describes the size of the image and the color representation method of each pixel. The entire image is composed of 1024×1024 = 1,048,576 pixel points, and each pixel point has a corresponding RGB value to determine its color.

[0127] Refer to Figure 2 , the step of inputting the second sample image into the second neural network to output a sample defogged image includes:

[0128] Input the second sample image T10 into the first downsampling convolutional layer to obtain a downsampled feature map T11 of the first size;

[0129] Input the downsampled feature map T11 of the first size into the second downsampling convolutional layer to obtain a downsampled feature map T12 of the second size;

[0130] Input the feature map T12 of the second size into the third downsampling convolutional layer to obtain a downsampled feature map T13 of the third size;

[0131] Input the downsampled feature map T13 of the third size into the first upsampling convolutional layer to obtain an upsampled feature map T22 of the second size;

[0132] Perform splicing processing on the downsampled feature map T12 of the second size and the upsampled feature map T22 of the second size to obtain a spliced feature map C1;

[0133] Input the spliced feature map C1 into the second upsampling convolutional layer to obtain an upsampled feature map T21 of the first size;

[0134] Perform splicing processing on the downsampled feature map T11 of the first size and the upsampled feature map T21 of the first size to obtain a spliced feature map C2;

[0135] Input the spliced feature map C2 into the third upsampling convolutional layer to obtain a sample defogged image.

[0136] Through the downsampling process, the model can extract image features at different scales, and sufficient learning can be obtained from local details to macroscopic features. The upsampling process combined with the feature splicing operation fuses feature information at different scales, enabling the model to retain rich feature information while restoring the spatial resolution of the image, thereby improving the defogging effect.

[0137] The feature splicing operation enables the model to utilize more detailed features extracted during the downsampling process when reconstructing the image, which helps to more accurately restore the detailed information in the image, avoid the problem of detail loss during the defogging process, and make the defogged image clearer and more natural.

[0138] This structure of multi-scale feature extraction and fusion enables the model to adapt to images of different types and different degrees of image fogging, improves the generalization ability of the model, and can have good defogging performance in various scenarios.

[0139] In a possible embodiment, before step 130, it may further include:

[0140] Input the first sample image into the first neural network to output sample three-dimensional information and sample material information;

[0141] Adjust the model parameters of the first neural network according to the reference three-dimensional information, the reference material information, the sample three-dimensional information, and the sample material information until the first neural network meets the preset training conditions to obtain the lighting prediction model;

[0142] Among them, the reference three-dimensional information is used to describe the normal vector information corresponding to each pixel point on the surface of the sample object in the three-dimensional space of the first sample image; the reference material information is used to describe the attribute information of the surface material of the sample object in a uniform lighting environment.

[0143] First sample image: Image data used to train the lighting prediction model. Multiple first sample images are images corresponding to the same sample object under different lighting environments, containing rich lighting and object feature information.

[0144] First neural network: A neural network model specifically used to learn to predict the three-dimensional information and material information of the sample object from the image, and continuously adjusts its model parameters to improve the prediction accuracy.

[0145] Sample three-dimensional information: The three-dimensional feature information of the sample object output by the first neural network according to the input first sample image, usually represented by the normal vector information corresponding to each pixel point on the surface of the sample object, reflecting the three-dimensional shape and orientation of the object surface. The sample three-dimensional information is also called Normal, which refers to the normal vector information corresponding to each pixel point in the three-dimensional space of the human portrait surface. This vector is perpendicular to the surface of objects such as the human face and is used to characterize the direction and orientation of the surface, reflecting the geometric shape and concave and convex details of the object surface.

[0146] Sample material information: Information about the surface material properties of a sample object output by the first neural network, which describes the characteristics of the surface material of the sample object in a uniform illumination environment, such as reflectivity, roughness, etc. The sample material information is also called Albedo, which refers to the inherent property that describes the light reflection characteristics of the surface material of a person or an object when the external additional light influence is removed and in a uniform illumination environment. It determines the basic color and brightness of the object surface.

[0147] Reference stereo information: Pre-defined information used to describe the normal vector information corresponding to each pixel point on the surface of the sample object in the three-dimensional space of the first sample image. As the true label during training, it is used to guide the first neural network to learn the correct stereo information.

[0148] Reference material information: Pre-defined information that describes the surface material properties of the sample object in a uniform illumination environment. As the true label during training, it is used to guide the first neural network to learn the correct material information.

[0149] Preset training conditions: Pre-set training termination conditions, such as reaching a certain number of training epochs, the loss function value being lower than a certain threshold, etc., which are used to control the training process of the first neural network to ensure that the model converges to a better state.

[0150] Lighting prediction model: The first neural network obtained after training, which can accurately predict the stereo information and material information of the photographed object based on the input image, providing a basis for subsequent simulated lighting processing.

[0151] Input multiple groups of first sample images into the first neural network. The multiple first sample images are images corresponding to the same sample object in different illumination environments, which contain rich illumination changes and object features, providing diverse data for model learning.

[0152] The first neural network processes the input first sample images. Through its internal convolutional layers, fully connected layers and other structures, it learns the feature patterns in the images and outputs sample stereo information and sample material information. This process is the process of the model analyzing and predicting the input images based on its current parameters.

[0153] Compare the sample stereo information and sample material information output by the lighting prediction model with the pre-defined reference stereo information and reference material information, and calculate the differences between them. According to the value of the loss function, use an optimization algorithm to adjust the model parameters of the first neural network so that the sample stereo information and sample material information gradually approach the reference stereo information and reference material information.

[0154] After each parameter adjustment, check whether the first neural network meets the preset training conditions. If the conditions are met, such as the loss function value has dropped low enough or the preset number of training rounds has been reached, stop the training. At this time, the obtained first neural network is the lighting prediction model. Through a large amount of training data and continuous adjustment of model parameters, the lighting prediction model can accurately predict the three-dimensional information and material information of the photographed object from the input image. This is very important for subsequent simulated lighting processing because different three-dimensional shapes and materials reflect and refract light in different ways, and accurate feature prediction can make the simulated lighting effect more realistic.

[0155] Since the first sample images are taken under different lighting environments, the lighting prediction model can learn the impact of lighting changes on image features, so as to have better prediction performance when processing images under different lighting conditions, improving the generalization ability of the model.

[0156] Accurate prediction of three-dimensional information and material information provides a necessary basis for subsequent simulated lighting processing. Based on this information, the reflection, refraction and other phenomena of light on the surface of the photographed object can be more accurately simulated, making the final generated target image have a more realistic lighting effect.

[0157] In a possible embodiment, the first neural network includes a first sub-convolutional network and a second sub-convolutional network connected in series. In the above-mentioned step of inputting the first sample image into the first neural network and outputting the sample three-dimensional information and sample material information, it may specifically include the following steps:

[0158] Input the first sample image into the first sub-convolutional network to extract the sample three-dimensional information of the first sample image;

[0159] Input the first sample image and the sample three-dimensional information into the second sub-convolutional network to extract the sample material information;

[0160] Adjusting the model parameters of the first neural network according to the reference three-dimensional information, the reference material information, the sample three-dimensional information and the sample material information includes:

[0161] Adjust the model parameters of the first sub-convolutional network according to the reference three-dimensional information and the sample three-dimensional information; and,

[0162] Adjust the model parameters of the second sub-convolutional network according to the reference material information and the sample material information.

[0163] First Sub-convolutional Network: A part of the first neural network, mainly used for feature extraction of the input first sample image to obtain sample three-dimensional information. It consists of multiple convolutional layers and possibly pooling layers, etc., and learns the spatial features in the image through convolutional operations on the image.

[0164] Second Sub-convolutional Network: Another part of the first neural network connected in series with the first sub-convolutional network, which receives the first sample image and the sample three-dimensional information output by the first sub-convolutional network as inputs, and further extracts sample material information. It is also composed of structures such as convolutional layers, and is used to learn image features related to materials.

[0165] Sample Three-dimensional Information: The three-dimensional feature information of the sample object in three-dimensional space output by the first sub-convolutional network after processing the first sample image, usually represented in the form of normal vector information corresponding to each pixel point on the surface of the sample object, reflecting features such as the shape and surface orientation of the object.

[0166] Sample Material Information: The information about the surface material properties of the sample object output by the second sub-convolutional network after receiving the first sample image and the sample three-dimensional information, such as the reflectivity and roughness of the material, etc., used to describe the surface characteristics of the sample object in a uniform illumination environment.

[0167] Input the first sample image into the first sub-convolutional network. The first sub-convolutional network performs convolutional operations on the image through its internal convolutional layers, gradually extracting the spatial features in the image. As the convolutional layers are continuously processed, the network can learn the feature patterns related to the three-dimensional shape of the sample object in the image, and finally output the sample three-dimensional information. For example, by learning features such as edges and contours in the image, the network can infer the normal vector information of each point on the surface of the sample object, thereby obtaining the sample three-dimensional information.

[0168] Input the first sample image and the sample three-dimensional information output by the first sub-convolutional network into the second sub-convolutional network at the same time. The second sub-convolutional network further processes the image on the basis of combining the features of the image itself and the sample three-dimensional information. It learns the feature patterns related to the material of the sample object through convolutional operations, such as the reflective characteristics and textures of the material, etc., and thus outputs the sample material information. Because the sample three-dimensional information can provide information about the shape of the object, this helps the second sub-convolutional network to more accurately judge the material properties of different regions.

[0169] For the first sub-convolutional network, compare the sample three-dimensional information output by it with the reference three-dimensional information, and calculate the difference between the two. According to the value of the loss function, use an optimization algorithm to adjust the model parameters of the first sub-convolutional network to make the sample three-dimensional information closer to the reference three-dimensional information.

[0170] For the second sub-convolutional network, compare the sample material information output by it with the reference material information, calculate the difference in the same way, and use an optimization algorithm to adjust the model parameters to make the sample material information closer to the reference material information.

[0171] By dividing the first neural network into a first sub-convolutional network and a second sub-convolutional network, which are responsible for extracting sample three-dimensional information and sample material information respectively, each sub-network can focus on extracting specific types of features, improving the efficiency and accuracy of feature extraction. The first sub-convolutional network can better learn features related to three-dimensional shapes, while the second sub-convolutional network can more effectively capture features related to materials.

[0172] Adjust the model parameters of the first sub-convolutional network according to the reference three-dimensional information and sample three-dimensional information respectively, and adjust the model parameters of the second sub-convolutional network according to the reference material information and sample material information, making the training process of the model clearer and more targeted. This can avoid the interference of the learning of different types of features during the training process, help the model converge faster, and improve the training efficiency and stability.

[0173] The tandem structure of the first sub-convolutional network and the second sub-convolutional network allows the sample three-dimensional information and the sample material information to fully share the structural information of the first sample image as the input image, first extract the sample material information of the first sample image, which helps the subsequent learning of the sample material information.

[0174] Thus, the structure and training method of the first neural network can enable the first neural network to more accurately predict the sample three-dimensional information and sample material information. Accurate feature prediction provides a more reliable basis for subsequent simulated lighting processing, making the simulated lighting effect more realistic and in line with the actual situation, and improving the performance of the entire image processing system.

[0175] In a possible embodiment, in step 140, it may specifically include the following steps:

[0176] Obtain a unit lighting map according to the preset light source information, the predicted three-dimensional information, and the predicted material information;

[0177] Perform simulated lighting processing on the second image according to the unit lighting map and the lighting weight value, and obtain a target image, where the lighting weight value is determined according to the second image.

[0178] Preset light source information: Various parameter information about the light source set in advance, such as the position, intensity, color, etc. of the light source, which is used to guide the specific method of simulated lighting.

[0179] Predicted three-dimensional information: Three-dimensional stereo feature information of the photographed object obtained by analyzing the second image through a lighting prediction model, usually represented by the normal vector information corresponding to each pixel point on the surface of the photographed object to reflect its shape, orientation, etc.

[0180] Predicted material information: Surface material attribute information predicted by the lighting prediction model for the photographed object in the second image, such as the reflectivity and roughness of the material, which will affect the interaction between light and the object surface.

[0181] Unit lighting map: An image calculated and generated based on the preset light source information, predicted three-dimensional information, and predicted material information. It represents the lighting effect that each pixel point should receive under given conditions and is a quantitative representation of the lighting distribution.

[0182] Lighting weight value: A weight value determined according to the characteristics of the second image. The characteristics of the second image are, for example, the brightness distribution, color information, and position of the object, etc. It is used to measure the application degree of the unit lighting map during simulated lighting processing, and different regions may have different lighting weight values.

[0183] Target image: The finally obtained image after simulated lighting processing, which integrates the original information of the second image and the simulated lighting effect added according to the unit lighting map and the lighting weight value.

[0184] Determine parameters such as the position, intensity, and color of the light source according to the preset light source information. Combine the predicted three-dimensional information and use the principles of geometric optics to calculate the reflection and refraction of light on the surface of the photographed object. For example, according to the normal vector information of each point on the object surface, determine the angle between the light and the surface, and then calculate the direction and intensity of the reflected light.

[0185] Consider the predicted material information. Different materials have different reflection and absorption characteristics for light. Adjust the calculation of light propagation and reflection according to the material attributes to finally obtain the unit lighting map, which reflects the lighting situation of each pixel point in the image under the current light source and object characteristics.

[0186] Exemplarily, in three-dimensional space, if the position of the preset light source is directly in front of the forehead of a human face, the angles are pitch = 45, yaw = 0, roll = 0, and the intensity is 1, that is, unit brightness. In this way, a unit lighting map light_map facing the human face can be obtained.

[0187] Determine the lighting weight value according to the characteristics of the second image. Specifically, the histogram of the second image can be used to determine the brightness interval of the second image, and the lighting weight value can be determined according to the brightness interval.

[0188] The histogram represents the pixel distribution of different brightness values in an image. In the histogram, the horizontal axis usually represents the brightness value, and the vertical axis represents the number of pixels with the corresponding brightness value. According to the shape of the histogram and the pixel distribution, the brightness range of the image can be divided into different intervals, such as the dark interval, the middle tone interval, and the bright interval.

[0189] For example, if it is determined that there are too many pixels in the dark interval of the image and the overall image is too dark, in order to make the image clearer and more natural, it may be necessary to increase the lighting of the dark part. At this time, the value of the lighting weight in the dark interval will be relatively large to emphasize the processing of the dark part. The lighting weight is a set of values separately set for different brightness intervals, and each value represents the degree of adjustment for the corresponding brightness interval or the relative size of the lighting intensity. By reasonably setting the lighting weight, the brightness distribution of the image can be optimized and the visual effect of the image can be improved.

[0190] Combine the unit lighting map with the lighting weight value to perform simulated lighting processing on the second image. Specifically, adjust the lighting effect of the unit lighting map according to the lighting weight value, and then superimpose the adjusted lighting effect on the second image. In this way, while the second image retains the original content, it presents a simulated lighting effect, and finally the target image is obtained.

[0191] The lighting process is described below in combination with the following formula:

[0192] Light_map = albedo+normal*L (2)

[0193] Results = light_map*M+input_image (3)

[0194] Among them, Light_map is the unit lighting map, L is the preset light source information, albedo is the predicted material information, and normal is the predicted three-dimensional information;

[0195] M is the lighting weight value, input_image is the second image, and Results is the target image.

[0196] Thus, by comprehensively considering the preset light source information, the predicted three-dimensional information, and the predicted material information to generate the unit lighting map, it is possible to more accurately simulate the propagation and reflection of light on the surfaces of different objects, thereby adding a realistic lighting effect to the image and making the target image look more realistic.

[0197] Determining the lighting weight value according to the second image and performing targeted simulated lighting processing can highlight the key areas in the image, adjust the overall brightness and contrast of the image, enhance the visual expressiveness of the image, and make it more in line with the aesthetic needs of users or the requirements of specific application scenarios. It can flexibly adjust the effect of simulated lighting according to the characteristics of different second images, and has good adaptability. Whether it is portrait photography, landscape photography or other types of images, satisfactory simulated lighting results can be obtained through reasonable parameter settings.

[0198] The following will describe the embodiments of the present application in conjunction with Figure 3 as follows:

[0199] Obtain the first image 10 and the preset light source information;

[0200] Extract the person mask map and the skin mask map 13 of the photographed object in the first image 10;

[0201] Input the first image 10 into the defogging model 11 to obtain a third image; the defogging model is trained according to multiple sets of second training data, and each set of the second training data includes: a second sample image and a reference image, and the reference image is an image obtained by defogging the second sample image;

[0202] Input the third image into the skin color restoration module 12, and perform skin color restoration processing on the third image according to the first image 10, the person mask map and the skin mask map 13 to obtain a second image;

[0203] Input the second image into the lighting prediction model 14 to obtain the predicted three-dimensional information and predicted material information of the photographed object in the second image; the lighting prediction model 14 is trained according to multiple sets of first training data, and each set of the first training data includes: a first sample image, reference three-dimensional information and reference material information; the multiple first sample images are images corresponding to the same sample object under different lighting environments;

[0204] Input the predicted three-dimensional information and predicted material information of the photographed object in the second image into the rendering module 15, and perform rendering processing according to the preset light source information, the predicted three-dimensional information and the predicted material information to obtain a unit lighting map;

[0205] Among them, the rendering module 15 can be a Physically Based Rendering (PBR) module. The rendering module 15 simulates the interaction between light and the object surface based on the physical laws of the real world. It takes into account the material properties of the object and phenomena such as the propagation, reflection, and refraction of light to generate a more realistic rendering effect. The material system of the rendering module 15 defines the properties of various materials and adopts a lighting model that conforms to physical laws, capable of accurately calculating the reflection and scattering of light on the object surface, including the effects of direct lighting and indirect lighting, and using various textures to add details and variations to the materials. By combining the textures with the material properties, the realism of the object is further enhanced.

[0206] Input the unit lighting map into the lighting module 16, and perform simulated lighting processing on the second image according to the unit lighting map and the lighting weight value to obtain the target image 17.

[0207] In the embodiments of the present application, by obtaining the first image and the preset light source information; performing image defogging processing on the first image to obtain the second image. Image defogging processing can improve the quality of the image, remove the influence of the foggy effect generated when shooting in a complex light scene on the image, make the image clearer, and provide a better basis for subsequent processing. Input the second image into the lighting prediction model. Since the lighting prediction model is trained according to multiple sets of first training data, each set of first training data includes: a first sample image, reference stereo information, and reference material information. The multiple first sample images are images corresponding to the same sample object in different lighting environments. The lighting prediction model establishes a mapping relationship between the image features and the stereo information and material information of the shooting object by learning multiple sets of first training data. When the second image is input, the lighting prediction model can accurately predict the stereo information and material information of the shooting object according to the learned rules. According to the preset light source information, the predicted stereo information, and the predicted material information, performing simulated lighting processing on the second image can make the image present different lighting effects, meet the needs of users in different scenarios, make the obtained target image more in line with the visual perception of the human eye under specific lighting conditions, and realize fine adjustment of the image taken in a complex light scene.

[0208] For the image processing method provided by the embodiments of the present application, the execution subject can be an image processing device. In the embodiments of the present application, taking the image processing device executing the image processing method as an example, the image processing device provided by the embodiments of the present application is described.

[0209] Figure 4 It is a block diagram of an image processing device provided by the embodiments of the present application. The device 400 includes:

[0210] An acquisition module 410, configured to acquire a first image and preset light source information;

[0211] A defogging processing module 420, configured to perform image defogging processing on the first image to obtain a second image;

[0212] A prediction module 430, configured to input the second image into a lighting prediction model to obtain predicted three-dimensional information and predicted material information of a photographed object in the second image; the lighting prediction model is trained according to multiple sets of first training data, and each set of the first training data includes: a first sample image, reference three-dimensional information, and reference material information; the multiple first sample images are images corresponding to the same sample object under different lighting environments;

[0213] A lighting processing module 440, configured to perform simulated lighting processing on the second image according to the preset light source information, the predicted three-dimensional information, and the predicted material information to obtain a target image.

[0214] In a possible embodiment, when the first image is a portrait, the defogging processing module 420 is specifically configured to:

[0215] Extract a person mask image and a skin mask image of the photographed object in the first image;

[0216] Input the first image into a defogging model to obtain a third image; the defogging model is trained according to multiple sets of second training data, and each set of the second training data includes: a second sample image and a reference image, and the reference image is an image obtained by performing defogging processing on the second sample image;

[0217] Process the third image according to the first image, the person mask image, and the skin mask image to obtain a second image.

[0218] In a possible embodiment, the lighting processing module 440 is specifically configured to:

[0219] Obtain a unit lighting map according to the preset light source information, the predicted three-dimensional information, and the predicted material information;

[0220] Perform simulated lighting processing on the second image according to the unit lighting map and a lighting weight value, and the lighting weight value is determined according to the second image to obtain a target image.

[0221] In a possible embodiment, the acquisition module 410 is further configured to acquire multiple sets of the second training data;

[0222] The apparatus 400 may further include:

[0223] A first input module, configured to input the second sample image into a second neural network to output a sample defogged image, where the second neural network includes: N downsampling convolutional layers and N upsampling convolutional layers, and N is an integer greater than 1;

[0224] A first training module, configured to adjust model parameters of the second neural network according to the sample defogged image and the reference image until a preset training condition is satisfied, to obtain a defogging model.

[0225] In a possible embodiment, the second neural network at least includes: a first downsampling convolutional layer, a second downsampling convolutional layer, a first upsampling convolutional layer, and a second upsampling convolutional layer. The first input module is specifically configured to:

[0226] Input the second sample image into the first downsampling convolutional layer to obtain a downsampled feature map of a first size;

[0227] Input the feature map of the first size into the second downsampling convolutional layer to obtain a downsampled feature map of a second size;

[0228] Input the feature map of the second size into the first upsampling convolutional layer to obtain an upsampled feature map of the first size;

[0229] Perform splicing processing on the downsampled feature map of the first size and the upsampled feature map of the first size to obtain a spliced feature map;

[0230] Input the spliced feature map into the second upsampling convolutional layer to obtain a sample defogged image.

[0231] In a possible embodiment, the obtaining module 410 is specifically configured to:

[0232] Obtain multiple groups of corresponding second sample images and initial reference images, where the initial reference image is an image obtained by performing initial defogging processing on the second sample image;

[0233] Determine a brightness difference parameter value according to the second sample image and the initial reference image;

[0234] Adjust the brightness of the initial reference image according to the brightness difference parameter value to obtain the reference image.

[0235] In a possible embodiment, the apparatus 400 may further include:

[0236] A second input module, further configured to input the first sample image into a first neural network to output sample stereo information and sample material information;

[0237] A second training module, configured to adjust model parameters of the first neural network according to the reference three-dimensional information, the reference material information, the sample three-dimensional information, and the sample material information until the first neural network meets a preset training condition, so as to obtain the lighting prediction model;

[0238] Wherein, the reference three-dimensional information is used to describe the normal vector information corresponding to each pixel point on the surface of the sample object in the three-dimensional space of the first sample image; the reference material information is used to describe the attribute information of the surface material of the sample object in a uniform illumination environment.

[0239] In a possible embodiment, the first neural network includes a first sub-convolutional network and a second sub-convolutional network connected in series. The second input module is specifically configured to:

[0240] Input the first sample image into the first sub-convolutional network to extract the sample three-dimensional information of the first sample image;

[0241] Input the first sample image and the sample three-dimensional information into the second sub-convolutional network to extract the sample material information;

[0242] The second training module is specifically configured to:

[0243] Adjust the model parameters of the first sub-convolutional network according to the reference three-dimensional information and the sample three-dimensional information; and

[0244] Adjust the model parameters of the second sub-convolutional network according to the reference material information and the sample material information.

[0245] In an embodiment of the present application, by obtaining a first image and preset light source information; performing image defogging processing on the first image to obtain a second image. Image defogging processing can improve the quality of the image, remove the influence of the foggy effect generated by shooting in a complex light scene on the image, make the image clearer, and provide a better basis for subsequent processing. Input the second image into the lighting prediction model. Since the lighting prediction model is trained based on multiple sets of first training data, each set of first training data includes: a first sample image, reference three-dimensional information, and reference material information. Multiple first sample images are images corresponding to the same sample object in different lighting environments. The lighting prediction model establishes a mapping relationship between image features and the three-dimensional information and material information of the shooting object by learning multiple sets of first training data. When the second image is input, the lighting prediction model can accurately predict the three-dimensional information and material information of the shooting object according to the learned rules. According to the preset light source information, predicted three-dimensional information, and predicted material information, performing simulated lighting processing on the second image can make the image present different lighting effects, meet the needs of users in different scenarios, make the obtained target image more in line with the visual perception of the human eye under specific lighting conditions, and achieve fine adjustment of the image taken in a complex light scene.

[0246] The image processing device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than a terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiment of the present application does not make specific limitations.

[0247] The image processing device in the embodiment of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiment of the present application does not make specific limitations.

[0248] The image processing device provided by the embodiments of the present application can implement each process implemented by the above method embodiments. To avoid repetition, it will not be elaborated here.

[0249] Optionally, as Figure 5 shown, the embodiments of the present application further provide an electronic device 510, including a processor 511, a memory 512, a program or instruction stored on the memory 512 and executable on the processor 511. When the program or instruction is executed by the processor 511, it implements each step of any of the above image processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0250] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0251] Figure 6 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.

[0252] The electronic device 600 includes but is not limited to: a radio frequency unit 601, a network module 602, an audio output unit 603, an input unit 604, a sensor 605, a display unit 606, a user input unit 607, an interface unit 608, a memory 609, and a processor 610, etc.

[0253] Those skilled in the art can understand that the electronic device 600 may further include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 610 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 6 The structure of the electronic device shown in does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0254] Among them, the processor 610 is used to obtain a first image and preset light source information;

[0255] The processor 610 is further used to perform image defogging processing on the first image to obtain a second image;

[0256] The processor 610 is further used to input the second image into a lighting prediction model to obtain predicted three-dimensional information and predicted material information of the photographed object in the second image; the lighting prediction model is trained according to multiple groups of first training data, and each group of the first training data includes: a first sample image, reference three-dimensional information, and reference material information; the multiple first sample images are images corresponding to the same sample object under different lighting environments;

[0257] The processor 610 is further configured to perform simulated illumination processing on the second image according to the preset light source information, the predicted three-dimensional information, and the predicted material information to obtain a target image.

[0258] Optionally, when the first image is a portrait, the processor 610 is further configured to extract a person mask map and a skin mask map of the photographed object in the first image;

[0259] The processor 610 is further configured to input the first image into a defogging model to obtain a third image; the defogging model is trained according to multiple sets of second training data, and each set of the second training data includes: a second sample image and a reference image, and the reference image is an image obtained by performing defogging processing on the second sample image;

[0260] The processor 610 is further configured to process the third image according to the first image, the person mask map, and the skin mask map to obtain a second image.

[0261] In some embodiments, the processor 610 is further configured to obtain a unit lighting map according to the preset light source information, the predicted three-dimensional information, and the predicted material information;

[0262] The processor 610 is further configured to perform simulated illumination processing on the second image according to the unit lighting map and a lighting weight value, which is determined according to the second image, to obtain a target image.

[0263] In some embodiments, the processor 610 is further configured to obtain multiple sets of the second training data;

[0264] The processor 610 is further configured to input the second sample image into a second neural network to output a sample defogged image, and the second neural network includes: N downsampling convolutional layers and N upsampling convolutional layers, where N is an integer greater than 1;

[0265] The processor 610 is further configured to adjust model parameters of the second neural network according to the sample defogged image and the reference image until a preset training condition is satisfied to obtain a defogging model.

[0266] In some embodiments, the second neural network at least includes: a first downsampling convolutional layer, a second downsampling convolutional layer, a first upsampling convolutional layer, and a second upsampling convolutional layer, and the processor 610 is further configured to input the second sample image into the first downsampling convolutional layer to obtain a downsampled feature map of a first size;

[0267] The processor 610 is further configured to input the feature map of the first size into the second downsampling convolutional layer to obtain a downsampled feature map of a second size;

[0268] The processor 610 is further configured to input the feature map of the second size into the first upsampling convolutional layer to obtain an upsampled feature map of the first size;

[0269] The processor 610 is further configured to splice the downsampled feature map of the first size and the upsampled feature map of the first size to obtain a spliced feature map;

[0270] The processor 610 is further configured to input the spliced feature map into the second upsampling convolutional layer to obtain a sample defogged image.

[0271] In some embodiments, the processor 610 is further configured to obtain multiple groups of corresponding second sample images and initial reference images, where the initial reference image is an image obtained by performing initial defogging processing on the second sample image;

[0272] The processor 610 is further configured to determine a brightness difference parameter value according to the second sample image and the initial reference image;

[0273] The processor 610 is further configured to adjust the brightness of the initial reference image according to the brightness difference parameter value to obtain the reference image.

[0274] In some embodiments, the processor 610 is further configured to input the first sample image into a first neural network to output sample three-dimensional information and sample material information;

[0275] The processor 610 is further configured to adjust the model parameters of the first neural network according to the reference three-dimensional information, the reference material information, the sample three-dimensional information, and the sample material information until the first neural network meets a preset training condition to obtain the lighting prediction model;

[0276] Wherein, the reference three-dimensional information is used to describe the normal vector information corresponding to each pixel point on the surface of the sample object in the three-dimensional space of the first sample image; the reference material information is used to describe the attribute information of the surface material of the sample object in a uniform lighting environment.

[0277] In some embodiments, the first neural network includes a first sub-convolutional network and a second sub-convolutional network connected in series. The processor 610 is further configured to input the first sample image into the first sub-convolutional network to extract the sample three-dimensional information of the first sample image;

[0278] The processor 610 is further configured to input the first sample image and the sample three-dimensional information into the second sub-convolutional network to extract the sample material information;

[0279] The processor 610 is further configured to adjust the model parameters of the first sub-convolutional network according to the reference stereo information and the sample stereo information; and adjust the model parameters of the second sub-convolutional network according to the reference material information and the sample material information.

[0280] In the embodiments of the present application, by obtaining a first image and preset light source information; performing image dehazing processing on the first image to obtain a second image. Image dehazing processing can improve the quality of the image, remove the influence of the foggy effect generated by shooting in a complex light scene on the image, make the image clearer, and provide a better basis for subsequent processing. Input the second image into the lighting prediction model. Since the lighting prediction model is trained according to multiple sets of first training data, each set of first training data includes: a first sample image, reference stereo information, and reference material information. The multiple first sample images are images corresponding to the same sample object under different lighting environments. The lighting prediction model establishes a mapping relationship between the image features and the stereo information and material information of the shooting object by learning multiple sets of first training data. When the second image is input, the lighting prediction model can accurately predict the stereo information and material information of the shooting object according to the learned rules. According to the preset light source information, predicted stereo information, and predicted material information, performing simulated lighting processing on the second image can make the image present different lighting effects, meet the needs of users in different scenarios, make the obtained target image more in line with the visual perception of the human eye under specific lighting conditions, and realize fine adjustment of the image captured in a complex light scene.

[0281] It should be understood that in the embodiments of the present application, the input unit 604 may include a Graphics Processing Unit (GPU) 6041 and a microphone 6042. The GPU 6041 processes the image data of static pictures or video images obtained by an image capturing device (such as a camera) in a video image capturing mode or an image capturing mode. The display unit 606 may include a display panel 6061, and the display panel 6061 may be configured in the form of, for example, a liquid crystal display, an organic light emitting diode, etc. The user input unit 607 includes at least one of a touch panel 6071 and other input devices 6072. The touch panel 6071 is also referred to as a touch screen. The touch panel 6071 may include two parts: a touch detection device and a touch controller. The other input devices 6072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an action bar, which will not be elaborated here. The memory 609 may be used to store software programs and various data, including but not limited to application programs and operating systems. The processor 610 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 610.

[0282] The memory 609 can be used to store software programs and various data. The memory 609 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 609 can include volatile memory or non-volatile memory, or the memory 609 can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 609 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memory.

[0283] The processor 610 may include one or more processing units; optionally, the processor 610 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 610 either.

[0284] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above embodiment of the image processing method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0285] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.

[0286] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above embodiment of the image processing method and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0287] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0288] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above embodiment of the image processing method and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0289] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed. It may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described method may be executed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0290] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0291] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining a first image and preset light source information; Performing image dehazing processing on the first image to obtain a second image; Inputting the second image into a lighting prediction model to obtain predicted three-dimensional information and predicted material information of the photographed object in the second image; the lighting prediction model is trained according to multiple groups of first training data, and each group of the first training data includes: a first sample image, reference three-dimensional information, and reference material information; the multiple first sample images are images corresponding to the same sample object under different lighting environments; Performing simulated lighting processing on the second image according to the preset light source information, the predicted three-dimensional information, and the predicted material information to obtain a target image.

2. The method according to claim 1, wherein When the first image is a portrait, the performing image dehazing processing on the first image to obtain a second image includes: Extracting a person mask map and a skin mask map of the photographed object in the first image; Inputting the first image into a dehazing model to obtain a third image; the dehazing model is trained according to multiple groups of second training data, and each group of the second training data includes: a second sample image and a reference image, and the reference image is an image obtained by performing dehazing processing on the second sample image; Processing the third image according to the first image, the person mask map, and the skin mask map to obtain a second image.

3. The method according to claim 1, characterized in that, The performing simulated lighting processing on the second image according to the preset light source information, the predicted three-dimensional information, and the predicted material information to obtain a target image includes: Obtaining a unit lighting map according to the preset light source information, the predicted three-dimensional information, and the predicted material information; Performing simulated lighting processing on the second image according to the unit lighting map and a lighting weight value, and the lighting weight value is determined according to the second image to obtain a target image.

4. The method according to claim 2, wherein Before the inputting the first image into the dehazing model to obtain a third image, the method further includes: Obtaining multiple groups of the second training data; Inputting the second sample image into a second neural network to output a sample dehazed image, and the second neural network includes: N downsampling convolutional layers and N upsampling convolutional layers, and N is an integer greater than 1; Adjusting the model parameters of the second neural network according to the sample dehazed image and the reference image until a preset training condition is met to obtain a dehazing model.

5. The method according to claim 4, wherein The second neural network at least includes: a first downsampling convolutional layer, a second downsampling convolutional layer, a first upsampling convolutional layer, and a second upsampling convolutional layer, and the inputting the second sample image into the second neural network to output a sample dehazed image includes: Inputting the second sample image into the first downsampling convolutional layer to obtain a downsampled feature map of a first size; Inputting the feature map of the first size into the second downsampling convolutional layer to obtain a downsampled feature map of a second size; Inputting the feature map of the second size into the first upsampling convolutional layer to obtain an upsampled feature map of the first size; Perform a splicing process on the downsampled feature map of the first size and the upsampled feature map of the first size to obtain a spliced feature map; Input the spliced feature map into the second upsampling convolutional layer to obtain a sample defogged image.

6. The method according to claim 4, characterized in that, The obtaining of multiple groups of the second training data includes: Obtain multiple groups of corresponding second sample images and initial reference images, where the initial reference image is an image obtained by performing initial defogging processing on the second sample image; Determine a brightness difference parameter value according to the second sample image and the initial reference image; Adjust the brightness of the initial reference image according to the brightness difference parameter value to obtain the reference image.

7. The method according to claim 1, wherein Before inputting the second image into the lighting prediction model to obtain the predicted three-dimensional information and predicted material information of the photographed object in the second image, the method further includes: Input the first sample image into a first neural network to output sample three-dimensional information and sample material information; Adjust the model parameters of the first neural network according to the reference three-dimensional information, the reference material information, the sample three-dimensional information, and the sample material information until the first neural network meets a preset training condition to obtain the lighting prediction model; Wherein, the reference three-dimensional information is used to describe the normal vector information corresponding to each pixel point on the surface of the sample object in the three-dimensional space of the first sample image; the reference material information is used to describe the attribute information of the surface material of the sample object in a uniform lighting environment.

8. The method according to claim 7, characterized in that, The first neural network includes a first sub-convolutional network and a second sub-convolutional network connected in series. The inputting of the first sample image into the first neural network to output sample three-dimensional information and sample material information includes: Input the first sample image into the first sub-convolutional network to extract the sample three-dimensional information of the first sample image; Input the first sample image and the sample three-dimensional information into the second sub-convolutional network to extract the sample material information; The adjusting of the model parameters of the first neural network according to the reference three-dimensional information, the reference material information, the sample three-dimensional information, and the sample material information includes: Adjust the model parameters of the first sub-convolutional network according to the reference three-dimensional information and the sample three-dimensional information; and, Adjust the model parameters of the second sub-convolutional network according to the reference material information and the sample material information.

9. An image processing apparatus, characterized in that, The device includes: An obtaining module, configured to obtain a first image and preset light source information; A defogging processing module, configured to perform image defogging processing on the first image to obtain a second image; A prediction module, configured to input the second image into a lighting prediction model to obtain the predicted three-dimensional information and predicted material information of the photographed object in the second image; the lighting prediction model is trained according to multiple groups of first training data, and each group of the first training data includes: a first sample image, reference three-dimensional information, and reference material information; the multiple first sample images are images corresponding to the same sample object in different lighting environments; A lighting processing module, configured to perform simulated lighting processing on the second image according to the preset light source information, the predicted three-dimensional information, and the predicted material information, to obtain a target image.

10. The device according to claim 9, characterized in that, When the first image is a portrait, the defogging processing module is specifically configured to: Extract a person mask image and a skin mask image of the photographed object in the first image; Input the first image into a defogging model to obtain a third image; The defogging model is trained according to multiple sets of second training data, and each set of the second training data includes: a second sample image and a reference image, where the reference image is an image obtained by performing defogging processing on the second sample image; Process the third image according to the first image, the person mask image, and the skin mask image to obtain a second image.