Image processing methods, models, model training methods, media and devices

By identifying facial regions in images and combining them with lighting scene classification information, and using 3D LUT parameters for partitioning enhancement, the problem of image enhancement for facial lighting effects, which is difficult to address in existing technologies, is solved, thereby improving the lighting effect of images and user experience.

CN119741727BActive Publication Date: 2026-01-06HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311237169.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-21
Publication Date
2026-01-06
Estimated Expiration
2043-09-21

AI Technical Summary

Technical Problem

Existing image enhancement algorithms struggle to effectively enhance lighting conditions in facial areas, resulting in a poor image capture experience.

Method used

By identifying semantic regions of facial images and combining them with lighting scene classification information, 3D LUT parameters are used to partition and enhance different semantic regions, thereby achieving targeted adjustments to the lighting performance of faces.

Benefits of technology

It improves the lighting effects of images, especially the lighting effects in the face area, thus enhancing the user's image viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741727B_ABST
    Figure CN119741727B_ABST
Patent Text Reader

Abstract

The application relates to the field of image processing, and discloses an image processing method, a model, a model training method, a medium and equipment, which can perform image partition image enhancement or image partition image style migration according to image scenes such as light performance of a face, and greatly improve image effects. The method comprises the following steps: acquiring a first image to be processed; determining a first semantic region and a second semantic region in the first image, and determining first scene classification information corresponding to the first image; adjusting colors of the first semantic region and the second semantic region by using first color parameters and second color parameters corresponding to the first scene classification information of the first image; and obtaining a second image based on the adjusted first semantic region and the adjusted second semantic region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image processing method, model, model training method, medium, and device. Background Technology

[0002] With the development of electronic technology, users are paying more and more attention to the image quality of photos, especially the lighting effects in portrait photography. For example, when shooting portraits, users often focus more on the facial features of the subject, while overexposure of the face, uneven lighting distribution on the face, and blurring due to backlighting can all affect the shooting experience.

[0003] Currently, to improve image quality, image enhancement algorithms are commonly used. Conventional image enhancement algorithms apply a set of enhancement curves to the entire image, specifically adjusting pixel values ​​for different ranges of pixel values ​​using the numerical mapping relationships represented by different enhancement curves. In other words, this enhancement method performs pixel-level adjustments across different ranges of pixel values ​​in the entire image. This means that while this method focuses on the overall lighting effect of the image, it struggles to specifically target and enhance the lighting effects of specific areas like faces. Summary of the Invention

[0004] This application provides an image processing method, model, model training method, medium, and device that can enhance images by partitioning them according to image scenes such as the lighting of a face, thereby greatly improving image quality.

[0005] In a first aspect, embodiments of this application provide an image processing method applied to an electronic device. The method includes: acquiring a first image to be processed; determining a first semantic region and a second semantic region in the first image, and determining first scene classification information corresponding to the first image; adjusting the colors of the first semantic region and the second semantic region respectively using the first color parameter and the second color parameter corresponding to the first scene classification information of the first image; and obtaining a second image based on the adjusted first semantic region and the adjusted second semantic region.

[0006] For example, the first image is a portrait image. It is understood that the first semantic region is different from the second semantic region, and the first color parameter is different from the second color parameter. The first image can be a real-time captured image or an image pre-acquired by the terminal device, and the first image can be in RGB format. When the first image is a real-time captured image during shooting, the terminal device can convert the captured image from RAW format to RGB format to obtain the first image. Specifically, the first and second color parameters can be used to perform zonal enhancement of the first image incorporating scene information, such as zonal image enhancement, or zonal image style transfer incorporating scene information. Thus, this application can improve the classification ability for different image scenes by combining scene classification information, for example, improving the classification ability for different lighting conditions of a face in a portrait image, thereby performing targeted zonal enhancement or zonal image style transfer to improve the user's experience of viewing the adjusted image.

[0007] In one possible implementation of the first aspect, the first scene classification information is associated with at least one of the following: the lighting conditions under which the first image was captured, the skin tone of the person in the first image, and the color temperature of the first image. For example, when the first image is a portrait image, the first scene classification information corresponds to the lighting conditions of the face, such as front lighting, backlighting, or colored lighting. As an example, the aforementioned first scene classification information includes scene encoding information corresponding to various lighting conditions.

[0008] In one possible implementation of the first aspect, the first semantic region is a foreground region in the first image and the second semantic region is a background region in the first image; or, the first semantic region is different foreground regions in the first image. It is understood that the first image may have multiple foreground regions and one background region.

[0009] In one possible implementation of the first aspect, determining a first semantic region and a second semantic region in a first image includes: acquiring a first segmentation mask image of the first image, the first segmentation mask image being a segmentation mask image of the foreground in the first image; determining a first semantic region based on the first segmentation mask image, wherein the first semantic region is a foreground region in the first image; and determining a second semantic region based on the first semantic region, wherein the second semantic region is a background region in the first image. For example, this application can identify the first segmentation mask in the first image using a semantic segmentation network, but is not limited thereto. Furthermore, the second semantic region can be any region in the first image excluding one or more of the first semantic regions.

[0010] In one possible implementation of the first aspect, the first image is a portrait image, and the first semantic region includes at least one of a human body region, a face region, and a mouth region.

[0011] In one possible implementation of the first aspect, determining the first scene classification information corresponding to the first image includes: determining the first scene classification information based on image features of the first facial region in the first image. It is understood that this application can crop out a facial image from the first image to obtain the first facial region. For example, the image features corresponding to the first scene classification information can be image latent codes, such as image latent codes corresponding to various lighting conditions (i.e., illumination conditions).

[0012] In one possible implementation of the first aspect, the method further includes: acquiring first depth image features corresponding to the first image; and determining a first color parameter and a second color parameter based on the first depth image features and first scene classification information, wherein the first color parameter and the second color parameter are different 3D LUT parameters. As an example, when the first scene classification information is for different lighting scene classifications, the 3D LUT parameters corresponding to the facial region are also different. It can be understood that the 3D LUT parameters corresponding to different first semantic regions and second semantic regions in the first image are determined in real-time by combining the first scene classification information and the first depth image features. Furthermore, when the first image has multiple first semantic regions (foreground regions), different first semantic regions can correspond to different first color parameters, i.e., different 3D LUT parameters. For example, when the lighting scene classification information indicates backlighting, the enhancement of brightness by the 3D LUT parameters corresponding to the facial region is greater than the enhancement of brightness by the 3D LUT parameters corresponding to the facial region when the lighting scene classification information indicates front lighting.

[0013] In one possible implementation of the first aspect, the colors of the first semantic region and the second semantic region are adjusted using the first color parameter and the second color parameter corresponding to the first scene classification information of the first image, respectively. This includes: adjusting the colors of the first image using the first color parameter and the second color parameter to obtain a third image containing the adjusted first semantic region and a fourth image containing the adjusted second semantic region; and obtaining a second image based on the adjusted first and second semantic regions, including: fusing the adjusted first semantic region in the third image and the adjusted second semantic region in the fourth image to obtain the second image. It can be understood that different semantic regions in the second image are obtained based on different color parameters (3D LUT parameters).

[0014] In one possible implementation of the first aspect, the number of first semantic regions is M, each first semantic region corresponds to a first color parameter, and the colors of the first image are adjusted using the first color parameter and the second color parameter respectively to obtain a third image containing the adjusted first semantic regions and a fourth image containing the adjusted second semantic regions. This includes: adjusting the colors of the first image using the first color parameters corresponding to each of the M first semantic regions to generate a corresponding third image for each first semantic region, resulting in M ​​third images; adjusting the colors of the first image using the second color parameter to obtain a fourth image containing the second semantic regions; and fusing the adjusted first semantic regions in the third images and the adjusted second semantic regions in the fourth images to obtain a second image, including: fusing the adjusted first semantic regions in each third image and the adjusted second semantic regions in the fourth image to obtain a second image. As an example, this application can use the fourth image as a reference and replace the corresponding semantic regions in each of the third images with the adjusted first semantic regions, thereby enabling the second image to contain different semantic regions adjusted by different color parameters. Furthermore, this application can also use a first semantic region in a third image as a reference to replace the corresponding semantic regions in other third images and the second semantic region in a fourth image with the respective semantic regions in the third image.

[0015] In one possible implementation of the first aspect, determining the first color parameter and the second color parameter based on the first depth image features and the first scene classification information includes: obtaining N first weight values ​​corresponding to the first semantic region and N second weight values ​​corresponding to the second semantic region based on the first depth image features and the first scene classification information, wherein the N first weight values ​​are related to the color of the first semantic region and the N second weight values ​​are related to the color of the second semantic region; performing weighted calculations on the N preset N first 3D LUT parameters using the N first weight values ​​and the N second weight values ​​respectively to obtain the third 3D LUT parameter and the fourth 3D LUT parameter; and performing trilinear interpolation on the third 3D LUT parameter and the fourth 3D LUT parameter respectively to obtain the first color parameter and the second color parameter. For example, the preset N first 3D LUT parameters can be a basic 3D LUT with a size of 33*33*33, while the 3D LUT represented by the first color parameter and the second color parameter can be 256*256*256. It is understood that this application can use a pre-trained encoder to predict the weight values ​​corresponding to different semantic regions in the first image, such as predicting T*N weight values, where T is the number of semantic regions (i.e., the first semantic region and the second semantic region) in the first image.

[0016] In one possible implementation of the first aspect, determining the first color parameter and the second color parameter based on the first depth image features and the first scene classification information includes: obtaining N first weight values ​​corresponding to the first semantic region and N second weight values ​​corresponding to the second semantic region, where N is a positive integer, based on the first depth image features and the first scene classification information; obtaining S first sampling intervals corresponding to the first semantic region and S second sampling intervals corresponding to the second semantic region, where S is a positive integer, based on the first depth image features and the first scene classification information; performing weighted calculations on the N preset first 3D LUT parameters using the N first weight values ​​and the N second weight values ​​respectively to obtain the third 3D LUT parameter and the fourth 3D LUT parameter; performing sampling processing on the third 3D LUT parameter using the S first sampling intervals to obtain the fifth 3D LUT parameter; performing sampling processing on the fourth 3D LUT parameter using the S second sampling intervals to obtain the sixth 3D LUT parameter; and performing trilinear interpolation on the fifth 3D LUT parameter and the sixth 3D LUT parameter respectively to obtain the first color parameter and the second color parameter. It is understood that this application can use a pre-trained encoder to predict the sampling intervals corresponding to different semantic regions in the first image. It is understood that this application can predict T*S sampling intervals, where T is the number of semantic regions (i.e., the first semantic region and the second semantic region) in the first image. For example, S = 32*3 = 96, where 32 represents the number of intervals in the basic 3DLUT of size 33*33*33, and 3 represents the number of channels in the three channels corresponding to the RGB format. Furthermore, the aforementioned T*S sampling intervals are non-uniform sampling intervals. Thus, this application can combine non-uniform interval sampling and trilinear interpolation to determine different 3D LUT parameters corresponding to different semantic regions, thereby improving the adaptability of semantic regions to 3D LUT parameters and contributing to improved final portrait partitioning enhancement effects.

[0017] Secondly, embodiments of this application provide an image processing model, including: a scene classification module, used to generate first scene classification information corresponding to the first image based on an image of a first facial region of the first image to be processed; a parameter generation module, used to determine a first semantic region and a second semantic region in the first image; and an image enhancement module, used to adjust the colors of the first semantic region and the second semantic region respectively based on the first color parameters and the second color parameters corresponding to the first scene classification information of the first image, and obtain a second image based on the adjusted first semantic region and the adjusted second semantic region. The image processing model described above can be a neural network model, such as the image enhancement network referred to below. The scene classification module may also include a scene encoding part.

[0018] In one possible implementation of the second aspect, the parameter generation module includes a first encoder. The first encoder is used to input a first image and a first segmentation mask image of the first image. The first segmentation mask image is a segmentation mask image of the foreground in the first image. Based on the first segmentation mask image, a first semantic region is determined, wherein the first semantic region is the foreground region in the first image. Based on the first semantic region, a second semantic region in the first image is determined, wherein the second semantic region is the background region in the first image. For example, the first encoder can be pre-trained. The parameter generation module can also be called a weight prediction part.

[0019] In one possible implementation of the second aspect, the scene classification module includes a second encoder, specifically used to input an image of a first facial region in the first image, and to determine first scene classification information based on the image features of the first facial region in the first image. For example, the second encoder can be pre-trained.

[0020] In one possible implementation of the second aspect, the first color parameter and the second color parameter are determined based on the first depth image features of the first image and the first scene classification information, wherein the first color parameter and the second color parameter are different 3D LUT parameters.

[0021] In one possible implementation of the second aspect, a parameter generation module is used to obtain N first weight values ​​corresponding to a first semantic region and N second weight values ​​corresponding to a second semantic region based on the first depth image features and the first scene classification information. The N first weight values ​​are related to the color of the first semantic region, and the N second weight values ​​are related to the color of the second semantic region. An image enhancement module is used to perform weighted calculations on N preset first 3D LUT parameters using the N first weight values ​​and N second weight values ​​respectively to obtain third and fourth 3D LUT parameters. Trilinear interpolation is then performed on the third and fourth 3D LUT parameters respectively to obtain first and second color parameters. The image enhancement module can also be referred to as the image enhancement part in a neural network model.

[0022] In one possible implementation of the second aspect, the parameter generation module is used to obtain N first weight values ​​corresponding to the first semantic region and N second weight values ​​corresponding to the second semantic region based on the first depth image features and the first scene classification information, wherein the N first weight values ​​are related to the color of the first semantic region and the N second weight values ​​are related to the color of the second semantic region. Based on the first depth image features and the first scene classification information, it obtains S first sampling intervals corresponding to the first semantic region and S second sampling intervals corresponding to the second semantic region, where S is a positive integer. The image enhancement module is used to perform weighted calculations on the N preset first 3D LUT parameters using the N first weight values ​​and the N second weight values ​​respectively to obtain the third 3D LUT parameters and the fourth 3D LUT parameters. It then performs sampling processing on the third 3D LUT parameters using the S first sampling intervals to obtain the fifth 3D LUT parameters. Finally, it performs trilinear interpolation on the fifth 3D LUT parameters and the sixth 3D LUT parameters respectively to obtain the first color parameters and the second color parameters. At this point, the parameter generation module may include an interval prediction component.

[0023] Thirdly, embodiments of this application provide an image processing method applied to an electronic device, the method comprising: inputting a first image to be processed into an image processing model in the second aspect and any possible implementation thereof; and obtaining a second image output by the image processing model.

[0024] Fourthly, embodiments of this application provide an image processing model training method applied to an electronic device. The method includes: inputting training sample data into a neural network model to be trained. The training sample data includes a sample image, the scene type to which the sample image belongs, and a standard enhanced image corresponding to the sample image. The standard enhanced image is an image obtained by image processing multiple semantic regions of the sample image using different color parameters based on the scene type; obtaining a first enhanced image and scene classification result output by the neural network model; and adjusting the parameters of the neural network model based on the first enhanced image, the standard enhanced image, the scene classification result, and the scene type, so as to obtain an image processing model when the first enhanced image and the scene classification result output by the neural network model satisfy a first condition. For example, the sample image can be a portrait image, and the scene type to which the sample image belongs can be a lighting table corresponding to the facial region in the sample image. This lighting table can be information on the lighting conditions of the facial region image of the sample image pre-labeled by the user. The first enhanced image and scene classification are obtained by the neural network model through real-time prediction of the sample image. Furthermore, standard augmented images can also be used to train augmented portrait images. For example, training augmented portrait images involves manually performing portrait augmentation on a training portrait image, focusing on enhancing the portrait based on the lighting conditions of the face.

[0025] In one possible implementation of the fourth aspect, the scene classification result is related to at least one of the following: the lighting conditions in which the sample image was taken, the skin color of the people in the sample image, and the color temperature of the sample image.

[0026] In one possible implementation of the fourth aspect, the parameters of the neural network model are adjusted based on the first enhanced image, the standard enhanced image, the scene classification result, and the scene type. This includes: determining the value of a first loss function based on the scene classification result and the scene type; adjusting the first parameters of the neural network model based on the value of the first loss function; determining the value of a second loss function based on the first enhanced image and the standard enhanced image; and adjusting the second parameters of the neural network model based on the value of the second loss function. It can be understood that the first parameter can be the network parameters corresponding to the first encoder, and the second parameter can be the network parameters corresponding to the second encoder.

[0027] In one possible implementation of the fourth aspect, the first condition includes: the value of the first loss function is within a first numerical range, and the value of the second loss function is within a second numerical range.

[0028] Fifthly, embodiments of this application provide a readable medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the image processing method as described in the first aspect and any possible implementation thereof.

[0029] In a sixth aspect, embodiments of this application provide an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, one of the processors of the electronic device, for performing an image processing method as described in the first aspect and any possible implementation thereof.

[0030] In a seventh aspect, embodiments of this application provide a computer program product, including: a computer program / instructions, which, when executed by a processor, implement the image processing method as described in the first aspect and any possible implementation thereof.

[0031] The beneficial effects of aspects two through seven of this application can be referred to the description in aspect one and its various possible implementations, and will not be repeated here. Attached Figure Description

[0032] Figure 1 illustrates a schematic diagram of an image processing scenario according to some embodiments of this application;

[0033] Figure 2 According to some embodiments of this application, a schematic diagram of an image processing architecture based on an image enhancement network is shown;

[0034] Figure 3 According to some embodiments of this application, a schematic diagram of a trilinear interpolation correlation calculation process is shown;

[0035] Figure 4 According to some embodiments of this application, a schematic diagram of the network architecture of an image enhancement network is shown;

[0036] Figure 5 According to some embodiments of this application, a schematic flowchart of an image processing method is shown;

[0037] Figure 6 According to some embodiments of this application, a schematic flowchart of a trilinear interpolation process is shown;

[0038] Figure 7 According to some embodiments of this application, a flowchart of non-uniform interval processing and trilinear interpolation processing is shown;

[0039] Figure 8 According to some embodiments of this application, a schematic diagram of the training and application process of an image enhancement network is shown;

[0040] Figure 9 According to some embodiments of this application, a schematic diagram of the structure of a mobile phone is shown. Detailed Implementation

[0041] The illustrative embodiments of this application include, but are not limited to, image processing methods, media, and electronic devices.

[0042] First, some terms used in the embodiments of this application will be introduced.

[0043] 3D Lookup Table (LUT): A color mapping technique that corrects the color of an image by mapping input color values ​​to output color values. It can be used for image enhancement and retouching. Specifically, 3D LUT technology can perform color correction on RGB (red, green, blue) format images through mapping relationships. For example, this mapping relationship can be represented as (R, G, B) = f(r, g, b), where (r, g, b) represents the input color and (R, G, B) represents the output color.

[0044] Semantic segmentation is an important branch of image processing and machine vision, aiming to accurately understand the scene and content of an image. For example, the basic structure of a neural network for semantic segmentation can include an encoder and a decoder. The encoder extracts features from the image using filters. The decoder is responsible for generating the final output, typically a segmentation mask containing the outline of an object; that is, the output of image semantic segmentation can be multiple segmentation masks. For example, semantic segmentation of a portrait image can yield a human body mask image (such as a foreground image), a face mask image, a mouth mask image, etc.

[0045] Loss function: An important concept in machine learning and deep learning, used to measure the difference or error between the model's prediction and the actual result. In this application, the loss function is also referred to as the loss function.

[0046] To address the problem that current image enhancement algorithms struggle to focus on facial lighting performance, this application provides an image processing method. For a portrait image, it can identify the scene classification information, such as the lighting scene classification information corresponding to the facial region and lighting conditions like front lighting or backlighting. Furthermore, this method can distinguish semantic regions such as the face, mouth, body, and background regions of the portrait image. Then, combining the determined lighting scene classification information, it performs color enhancement on these semantic regions separately, for example, image enhancement based on 3D LUT parameters. For instance, the method can determine the 3D LUT parameters corresponding to each semantic region, and use these 3D LUT parameters to enhance the portrait image separately, obtaining multiple enhanced images. These multiple enhanced images are then fused with the different semantic regions corresponding to the 3D LUT parameters to form a complete output image, thus achieving image partitioning enhancement combined with scene classification. In this way, this application can focus on the lighting performance of the face and perform partitioning enhancement based on the lighting performance of the face, thereby greatly improving the image quality and making the image effect meet user needs.

[0047] For example, when the lighting scene classification information indicates backlighting, the 3D LUT parameters corresponding to the facial area enhance brightness more than the 3D LUT parameters corresponding to the mouth area. Furthermore, the 3D LUT parameters corresponding to the facial area also differ across lighting scene classifications. For instance, when the lighting scene classification information indicates backlighting, the 3D LUT parameters corresponding to the facial area enhance brightness more than when the lighting scene classification information indicates front lighting.

[0048] In some embodiments, after acquiring an image, this application can first convert the image to RGB format, crop out the facial region image, and perform semantic segmentation on the image to obtain semantic segmentation mask images such as face mask images, mouth mask images, and body mask images. On one hand, it determines the lighting scene classification information corresponding to lighting conditions such as front lighting, backlighting, side lighting, and colored lighting for the facial region image. This lighting scene classification information can be a set of image latent codes (or image features) corresponding to lighting conditions such as front lighting, backlighting, side lighting, and colored lighting. On the other hand, this method can identify the depth image features of the portrait image, such as semantic features or texture features. Furthermore, this method can combine the identified lighting scene classification information and depth image features to determine the 3D LUT parameters corresponding to each semantic region of the portrait image, where these semantic regions include the semantic regions corresponding to each semantic segmentation mask image and the background region of the portrait image. Furthermore, this method can perform color correction on the portrait image using various 3D LUT parameters to obtain multiple enhanced images. Then, based on each semantic segmentation mask image, the different semantic regions corresponding to different 3D LUT parameters in these multiple enhanced images are fused into a complete RGB format output image. Thus, scene-specific and region-specific enhancement effects for portrait images can be achieved.

[0049] Thus, compared with the image enhancement methods in the background art, the image processing method of this application combines the lighting scene classification information of the face lighting performance, and realizes different enhancement algorithms for different regions of the image based on different lighting scene classifications. That is, the facial region in the image can be targeted for enhancement according to the face lighting performance, which is conducive to improving the image enhancement effect.

[0050] It is understood that the portrait enhancement based on 3D LUT parameters in this application can enhance the image in terms of brightness, saturation, contrast, and other dimensions using the RGB format.

[0051] In some embodiments, the image scene classification provided in this application is not limited to classifying the facial region in the image, but can also classify another region, classify multiple regions separately, or classify the entire image.

[0052] In some embodiments, the scene classification provided in this application is not limited to the above-mentioned lighting scene classification, but can also be set according to the user's actual needs. For example, scene classification is used to classify images of people with different skin tones, such as white, yellow, black, and other skin tone scenes. As another example, scene classification is used to classify images with different color temperatures, such as natural light sources, sunlight at sunrise, sunlight half an hour after sunrise, household incandescent lamps, cloudy days, and other color temperature scenes.

[0053] In some embodiments, the image processing method provided in this application can process images of other types, such as animal images, scene images, or landscape images, not limited to human portrait images.

[0054] In some embodiments, the image processing method provided in this application is not limited to image enhancement under different scene classifications, but can also be applied to image style transfer under different scene classifications. In this case, the image style based on 3D LUT parameters refers to the image style in dimensions such as brightness, saturation, and contrast represented in RGB format.

[0055] The image processing method provided in this application can be applied to the post-processing of image materials such as pictures or videos obtained through shooting. In this case, the method can use 3D LUT technology to perform color restoration and other processing on the image materials, such as portrait enhancement or image style transfer.

[0056] Thus, this application can improve the classification ability of different image scenes by combining scene classification information, such as improving the classification ability of faces under different lighting conditions in portrait images, thereby performing targeted image partition enhancement or partition image style transfer to improve the user's experience of viewing the adjusted image.

[0057] In some embodiments, the image processing method provided in this application can be applied to a mobile phone's camera application. For example, the camera application can provide a partition enhancement function and allow users to enable or disable this function. Furthermore, when the camera application enables this function, after the mobile phone captures a frame of image in shooting mode or video recording mode, it can perform post-processing on the captured image to achieve partitioned portrait enhancement or image style transfer combined with scene classification.

[0058] like Figure 1A As shown, the user opens the phone's camera app and displays a shooting preview interface, which includes a "Photo Album" 200, a "Personalization Enhancement" function switch 201, "Settings" 202, shooting controls 203, a viewfinder 204, and shooting mode selection controls such as "Portrait," "Photo," and "Video." The "Photo Album" 200 triggers the phone to display photos or videos taken during the most recent shooting session. The shooting controls 203 trigger the phone to start a shooting operation, such as taking a photo or recording a video. The viewfinder 204 displays the shooting preview image captured in real time by the phone. The "Personalization Enhancement" function switch 201 triggers the phone to turn the "Zone Enhancement" function on or off. Figure 1A The function switch 201 indicates that the "Partition Enhancement" function is enabled. Additionally, Figure 1AIn shooting modes such as "Portrait", "Photo", and "Video", the "Regional Enhancement" function can be turned on to post-process the images captured during the shooting process. That is, the captured images are combined with scene classification to perform portrait enhancement or image style transfer in different semantic regions.

[0059] So, in the user's opinion Figure 1A After the camera control 203 is clicked to trigger a camera shutter, the phone can capture image A1 in the viewfinder 204 and perform the image processing described above on image A1. For example, it can classify lighting scenes based on facial regions, and use different 3DLUT parameters to enhance the facial and background regions in image A1 to obtain photo A2. Photo A2 is then saved to the photo album application. Furthermore, after the user... Figure 1A After clicking on the displayed album 200, the phone can access it. Figure 1B The album interface is shown and photo A2 is displayed, allowing users to view portrait photos with better facial enhancement, i.e., better lighting on the face.

[0060] In some embodiments of this application, the subject performing the image processing method is not limited to the mobile phone exemplified above, but can also be other electronic devices with shooting functions or camera applications. For example, the electronic device can be a tablet computer, wearable electronic device, in-vehicle electronic device, augmented reality (AR) device, virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc.

[0061] In some embodiments, the image processing method provided in this application, which involves cropping a facial region image and segmenting various semantic segmentation mask images from a portrait image after acquisition, can be referred to as image preprocessing.

[0062] In some embodiments, this application may employ a pre-trained semantic segmentation network to segment various semantic segmentation mask images from a portrait image, such as a face mask image, a mouth mask image, a body mask image, etc. Furthermore, the size of these semantic segmentation mask images may be consistent with the size of the original portrait image.

[0063] In some embodiments, this application may employ network models such as neural networks to perform scene classification and recognition and image partitioning enhancement on portrait images. For example, the neural network may be called an image enhancement network or an image processing model.

[0064] Next, refer to Figure 2 The diagram shown is a schematic of an image processing architecture based on an image enhancement network provided in an embodiment of this application. Figure 2 The image augmentation network shown mainly consists of three parts: scene encoding, weight prediction, and image augmentation. The scene encoding part mainly includes encoder 1 and a fully connected (FC) layer; encoder 1 can also be called a scene encoder. The weight prediction part mainly includes encoder 2, which can also be called a weight prediction encoder.

[0065] In some embodiments, Figure 2 The image augmentation network in the model takes T+1 images as input, including a portrait image A1, a body mask image A11, a mouth mask image A12, a face mask image A13, and a facial region image A1'. The output is a portrait image A2 after augmenting the portrait image A1. Encoder 1 takes A1' as input, and encoder 2 takes T images as input (i.e., A1, A11, A12, and A13).

[0066] In some embodiments, Figure 2 The image enhancement network shown can also take N base 3D LUTs as input, or the N base 3D LUTs can be pre-set in the image enhancement network.

[0067] Reference Figure 2 Encoder 1 can identify a set of image latent codes (i.e., image classification information) corresponding to various lighting conditions, such as front lighting, backlighting, side lighting, and colored lighting, for the facial region image A1'. This set of image latent codes can include the image latent codes corresponding to each of these lighting conditions. For example, the image latent codes corresponding to multiple lighting conditions can be image features such as light position features. At this time, after the FC layer passes through this set of images, a scene classification result can be output, such as the scene classification result indicating that A1' is under front lighting.

[0068] In some embodiments, encoder 1 can be implemented using a neural network. Encoder 1 can contain different types of network layers, such as convolutional layers and pooling layers. The convolutional layer extracts information from the input image; this information is called image features, and these image features are represented by each pixel in the image through combinations or independent methods, such as texture features and color features. The pooling layer selects the image features extracted by the convolutional layer. A fully connected (FC) layer typically follows the pooling layer, transforming all the feature matrices from the pooling layer into a one-dimensional feature vector. A fully connected layer is generally placed at the end of the convolutional neural network structure and is used for image classification; that is, the fully connected layer outputs the classification result from the neural network. At this point, the output of encoder 1 before the FC layer is a set of image latent codes corresponding to the image and various lighting conditions.

[0069] In some embodiments, the encoder 1 provided in this application can be trained based on multiple training facial images and a lighting table corresponding to the lighting conditions of the multiple training facial images. The lighting table may include lighting scene classification information corresponding to front lighting, backlighting, side lighting, and colored lighting. Specifically, during training, the encoder 1 takes a training facial image as input and outputs the lighting scene classification result of the training facial image, such as front lighting. The value of the loss function is determined based on the lighting scene classification result and the lighting scene classification information in the lighting table. If the value of the loss function is not within a preset numerical range of 1, the network parameters of the encoder 1 are adjusted. For example, the network parameters may include the learning rate, the number of iterations, and the size of the mini-batch data. If the value of the loss function is within the numerical range of 1, it indicates that the encoder 1 has completed training.

[0070] For example, Figure 2 The mid-face region image A1' can be a training face image. After A1' enters encoder 1 (i.e., the second encoder), the loss function between the scene classification result output by encoder 1 after the FC layer and the illumination table is Loss1 (i.e., the first loss function). Encoder 1 can determine whether to adjust the network parameters (first parameters) of encoder 1 based on whether the region of Loss1 is within the numerical range 1 (i.e., the first numerical range).

[0071] In some embodiments, the encoder 2 provided in this application can be trained based on multiple training portrait images, corresponding semantic segmentation mask images, and training enhanced portrait images. Each training portrait image corresponds to a semantic segmentation mask image including a human body mask image, a face mask image, and a mouth mask image. Each training enhanced portrait image is an image after manually enhancing the portrait image of a training portrait image, focusing on enhancing the portrait image based on the lighting conditions of the face. Specifically, during training, the encoder 2 takes a training portrait image and corresponding image segmentation mask images as input and outputs an enhanced portrait image. This enhanced portrait image is compared with the corresponding training enhanced portrait image of the training portrait image to determine the value of the loss function related to image enhancement. If the value of the loss function is not within a preset value range of 2, the network parameters of the encoder 2 are adjusted. If the value of the loss function is within the value range of 2, it indicates that the encoder 2 has completed training.

[0072] For example, Figure 2Image A1, representing the mid-face region, is a training portrait image, and image A3 is the corresponding training enhanced portrait image. After A1 and its corresponding images A11-A13 enter encoder 2, encoder 2 (the first encoder) performs a series of subsequent processing steps to output the enhanced portrait image A2. The loss function between A2 and A3 is Loss2 (the second loss function). Encoder 2 can determine whether to adjust its network parameters (the second parameters) based on whether the region of Loss2 falls within the numerical range 2 (the second numerical range).

[0073] Next Figure 2 The weight prediction and image enhancement parts are shown in detail.

[0074] Reference Figure 2 In the weight prediction section, the T images—portrait image A1, body mask image A11, mouth mask image A12, and face mask image A13—are downsampled and input into encoder 2. Encoder 2 can determine the T semantic regions corresponding to portrait image A1, namely the body region, mouth region, face region, and background region. Furthermore, encoder 2 can identify the depth image features of portrait image A1 and, based on these depth features and a set of image latent codes from encoder 1 (excluding those passing through the FC layer) corresponding to various lighting conditions, predict T (e.g., T=4) sets of weight values ​​corresponding to N basic 3D LUTs. Each set of weight values ​​includes N (e.g., N=3) weight values ​​corresponding one-to-one with the N basic 3D LUTs. For example, Figure 2 The T weight values ​​in the encoder include W11-W1N, W21-W2N, ..., WT1-WTN, which means that encoder 2 predicts a total of T*N weight values.

[0075] Reference Figure 2In the image enhancement part, T sets of weight values, including W11-W1N, W21-W2N, ..., WT1-WTN, can be weighted and calculated with N basic 3D LUTs respectively, resulting in T weighted 3D LUTs for fusion of the basic 3D LUTs. Furthermore, the T weighted 3D LUTs undergo trilinear interpolation to obtain T fused 3D LUTs, i.e., one fused 3D LUT corresponding to each of the T semantic regions. Thus, adaptive 3D LUTs for different semantic regions can be generated. These T fused 3D LUTs are then used to enhance the T semantic regions in the human image A1. For example, the T sets of fused 3D LUTs can enhance the entire human image A1 to obtain T enhanced images. Then, based on the human body mask image A11, the mouth mask image A12, and the face mask image A13, the enhanced human body region image, the enhanced mouth region image, the enhanced face region image, and the enhanced background region image can be extracted from the T enhanced images respectively, and these images are then merged into a complete enhanced face image A2.

[0076] In some embodiments, the underlying 3D LUT is represented by the following formula (1):

[0077] T = {(T r,(i,j,k) T g,(i,j,k) T b,(i,j,k) )}, i, j, k∈(0, 1, 2,...N-1) (1)

[0078] Among them, T r,(i,j,k) Represents the R channel value, T g,(i,j,k) T represents the value of the G channel. b,(i,j,k) This represents the value of channel B.

[0079] In some embodiments, the weight prediction, i.e., encoder 2, can use MobileNet-V2 (a neural network) to predict a total of T*N weight values. The fused 3D LUT generated based on these weight values ​​is expressed by the following formula (2):

[0080] Lut t =w t1 *Lut1+w t2 *Lut2+…w tN *Lut N (2)

[0081] In formula (2), t ∈ (1, 2, ..., N). At this point, Lut in formula (2) t correspond Figure 2The 3D LUT after weighted calculation.

[0082] In some embodiments, this application searches the input portrait image A1 on T fused 3D LUTs according to the semantic segmentation results to generate T enhanced image regions, denoted as pic0, pic1, pic2, ..., pic(t-1), where t is an integer less than or equal to T. Image fusion is performed according to a certain fusion method to generate an enhanced portrait image, which is then converted from RGB format to Blue-Green-Red (BGR) format and saved to obtain an enhanced portrait image A2 incorporating scene coding information.

[0083] Understandable. Figure 2 In the weight prediction section shown, N basic 3D LUTs are preset, while T sets of weight values ​​are predicted in real time. However, in some other embodiments, the weight prediction section of the image enhancement network can preset T sets of weight values ​​and predict N basic 3D LUTs in real time; this application does not specifically limit this. In this case, Figure 2 The output of encoder 1 shown is N basic 3D LUTs, and the parameters of the N basic 3D LUTs and the predetermined T sets of weight values ​​are still weighted and calculated to fuse the basic 3D LUTs.

[0084] In some embodiments, the base 3D LUT size is typically 33*33*33. Therefore, T weighted 3D LUTs are each subjected to trilinear interpolation to obtain an interpolated 3D LUT of size 256*256*256 (i.e., the fused 3D LUT). Then, the semantic regions of the portrait image are searched on the interpolated 3D LUT for color adjustment. Furthermore, in other embodiments, the base 3D LUT size can be 77*77*77, etc., and no specific limitation is made here.

[0085] It is understandable that trilinear interpolation is mainly used in a 3D cube, where the values ​​of other points in the cube are calculated based on the values ​​of a given vertex. In other words, it is used to calculate the values ​​of the eight surrounding vertices of a cube.

[0086] For example, the weight prediction part and the interval prediction part in the above image enhancement network can be called the parameter generation module, the scene coding part can be called the scene classification module, and the image enhancement part can be called the image enhancement module.

[0087] Reference Figure 3 The diagram shown is a schematic representation of the calculation process related to trilinear interpolation in an embodiment of this application. Figure 3 The cube in (a) can be Figure 2 The image shows a cube within a weighted 3D LUT, specifically a periodic cubic mesh with a step size of 1. Furthermore, Figure 3 The cube shown in (a) has 8 vertices including C. 000 C 100 C 110 C 010 C 001 C 101 C 111 C 011 Trilinear interpolation can be used to obtain... Figure 3 Points C, C0, C1, and C2 in the cube shown in (b) 01 C 11 C 001 C 10 Specifically, assuming Figure 3 In diagram (a), point C has coordinates (x, y, z), and the coordinates (x0, y0, z0) are the points with the smallest relative coordinates among the eight vertices of the cube (e.g., point C). 000 ), and the coordinates of these points are according to Figure 3 The coordinate system (xyz) is shown. Here, x0 represents a grid point below x along the x-axis, x1 represents a grid point above x along the x-axis, and the relationship between y0, y1, z0, and z1 is similar to that between x0 and x1. Furthermore, xd, yd, and zd are taken as the largest integer differences between the points to be calculated (i.e., the interpolation points) less than x, y, and z, respectively.

[0088] In some embodiments, trilinear interpolation can be performed separately along the x-axis, y-axis, and z-axis. First, formula (3-6) can be used to... Figure 3 Interpolate the cube in (a) along the x-direction:

[0089] C 00 =V[x0, y0, z0](1-x d )+V[x0, y0, z0]x d (3)

[0090] C 01 =V[x0, y0, z1](1-x d )+V[x0, y0, z1]x d (4)

[0091] C 10 =V[x0, y1, z0](1-x d )+V[x0, y1, z0]x d (5)

[0092] C 11 =V[x0, y1, z1](1-x d )+V[x0, y1, z1]x d (6)

[0093] Where V[x0, y0, z0] represents the value of the function V[] at the point (x0, y0, z0).

[0094] Then, formula (7-8) can be used to... Figure 3 Interpolate the cube in (a) along the y-direction:

[0095] C0 = C 00 (1-y d )+C 10 y d (7)

[0096] C1 = C 01 (1-y d )+C 11 y d (8)

[0097] Finally, formula (9) can be used to... Figure 3 Interpolate the cube in (a) along the z-direction:

[0098] C = C0(1-z) d )+C1z d (9)

[0099] In some embodiments, the trilinear interpolation in this application can use RGB values ​​for indexing, which is equivalent to... Figure 3 The XYZ axes in the diagram.

[0100] Understandable. Figure 2 and Figure 3 The 3D LUT shown is a 3D LUT with a regular uniform sampling interval.

[0101] In some embodiments, this application performs image enhancement using a 3D LUT with a non-uniform sampling interval. For example, in Figure 2 Based on the network architecture shown, such as Figure 4 The diagram shown is a schematic representation of another image enhancement network architecture provided in an embodiment of this application. Figure 2 Compared to the architecture shown, Figure 4 An interval prediction component was added, and the image enhancement component also needs to be combined with a 3D LUT with a non-uniform sampling interval for trilinear interpolation to obtain T fused 3D LUTs.

[0102] like Figure 4As shown, encoder 2 can predict not only T sets of weight values, but also T sets of sampling intervals, including W11-W1S, W21-W2S, ..., WT1-WTS, meaning encoder 2 predicts a total of T*S sampling intervals. For example, S = 32 * 3 = 96, where 32 represents the number of intervals in the basic 3D LUT of size 33 * 33 * 33, and 3 represents the number of channels in the three channels corresponding to the RGB format. After obtaining T weighted 3D LUTs by weighting the T sets of weight values ​​with N basic 3D LUTs, the T weighted 3D LUTs can be adjusted using the T sets of sampling intervals to obtain T non-uniform sampling interval 3D LUTs. Then, trilinear interpolation can be performed on the T non-uniform sampling interval 3D LUTs to obtain T fused 3D LUTs, which are then used to perform image enhancement on different semantic regions of the input image (such as portrait image A1).

[0103] It is understood that in this application Figure 2 and Figure 4 The relevant content mainly introduces the network architecture for regional portrait enhancement that combines lighting scene classification. Additionally, it discusses the structure of network architectures that combine skin tone scene classification or color temperature scene classification. Figure 2 and Figure 4 Similarly, details will not be provided here. For example, Figure 2 and Figure 4 The training enhanced images shown in the image enhancement section can be manually enhanced by the user, focusing on skin tone or color temperature.

[0104] Next based on Figure 2 or Figure 4 The network architecture shown provides a detailed description of the image processing method provided in the application embodiments.

[0105] like Figure 5 The diagram shown is a flowchart illustrating an image processing method provided in an embodiment of this application. The method can be executed by an electronic device such as a mobile phone. Furthermore, this method can be applied to… Figure 2 The image enhancement network implementation is shown. Specifically, Figure 5 The method shown includes the following steps:

[0106] S501: Acquire a portrait image.

[0107] For example, the first image can be a portrait image, such as... Figure 2 The portrait image shown is A1.

[0108] In some embodiments, the portrait image described above can be a frame captured during mobile phone shooting. In this case, the captured image is usually in RAW format, and the image can be converted from RAW format to RGB format to obtain the portrait image described above.

[0109] S502: Obtain the facial region image of the portrait image and the corresponding N-1 semantic segmentation mask images.

[0110] In some embodiments, this application can crop a facial region image from a portrait image and generate T-1 semantic segmentation mask images corresponding to the portrait image using a semantic segmentation network. For example, the T-1 semantic segmentation mask images may include a human body mask image, a face mask image, and a mouth mask image, such as... Figure 2 The segmentation mask images A11-A13 are shown, where T is 3.

[0111] S503: Determine the lighting scene classification information corresponding to various lighting conditions for the facial region image. For example, the lighting scene classification information can be a set of image latent codes corresponding to various lighting conditions.

[0112] In some embodiments, this application may employ Figure 2 and Figure 4 The encoder 1 shown identifies the lighting scene classification information of the facial region image. This lighting scene classification information can be a set of image latent codes identified by encoder 1 that have not passed through the FC layer.

[0113] S504: Based on the portrait image and T-1 semantic segmentation mask images, determine the depth image features of the portrait image and the number of semantic regions T.

[0114] In some embodiments, the T semantic regions corresponding to the portrait image are the T-1 semantic regions corresponding to the T-1 semantic segmentation mask images, plus the background region in the portrait image. For example, the T-1 semantic regions are the human body region, the face region, and the mouth region. As an example, the human body region may not include the face region and the mouth region, and the face region may be a region that does not include the mouth region. In this case, T is 4, that is, 4 semantic regions corresponding to the face image.

[0115] Deep image features can be understood to be divided into two categories: primary features (shallow features) and high-level features (structural features). For example, primary features include shape features, color features, and texture features. High-level features include semantic features extracted and learned from lower-level features; these are highly abstract, such as facial analysis.

[0116] In some embodiments, this application may employ Figure 2 or Figure 4The encoder 2 shown determines the depth image features of the portrait image.

[0117] S505: Based on the lighting scene classification information and depth image features of the facial region image, determine the T first 3D LUT parameters corresponding to each of the T semantic regions.

[0118] In some embodiments, the T first 3D LUT parameters corresponding to each of the T semantic regions in this application can be: Figure 2 or Figure 4 The diagram shows T fused 3D LUTs. At this point, the T parameters of the first 3D LUTs are adaptive 3D LUTs for T semantic regions.

[0119] S506: Perform image enhancement on the portrait image according to the T first 3D LUT parameters to obtain T enhanced images corresponding to the T semantic regions.

[0120] For example, the size (i.e. dimension) of the first 3D LUT parameter is (256*256*256).

[0121] In some embodiments, this application can find RGB values ​​that have a mapping relationship with the RGB values ​​of each pixel in a portrait image from the first 3D LUT parameters to achieve color adjustment of the portrait image, i.e., image enhancement.

[0122] S507: Based on T-1 semantic segmentation mask images, fuse the images corresponding to the T semantic regions from the T enhanced images to obtain an enhanced portrait image. For example, refer to... Figure 2 or Figure 4 The enhanced portrait image shown can be A2.

[0123] In some embodiments, this application can convert enhanced portrait images in RGB format to BGR format to adapt to image display and other processing of electronic devices such as mobile phones.

[0124] In some embodiments, assuming that the T enhanced images and the enhanced images corresponding to the background region, human body region, face region, and mouth region are denoted as pic0, pic1, pic2, and pic3, respectively, this application can use one enhanced image as a base image and replace the corresponding pixels in the enhanced images corresponding to the other semantic regions in the base image. For example, this application can use pic0 as a base, and replace the pixels in the human body region of pic0 with the pixels in the human body region of pic1 according to the human body mask image, replace the pixels in the face region of pic0 with the pixels in the face region of pic2 according to the face mask image, and replace the pixels in the mouth region of pic0 with the pixels in the mouth region of pic3 according to the mouth mask image, thereby achieving the fusion of the T enhanced images.

[0125] Thus, the image processing method of this application can combine the lighting scene classification information of the face lighting performance, and realize different enhancement algorithms for different regions of the image based on different lighting scene classifications. That is, the facial region in the image can be targeted for enhancement according to the face lighting performance, which is beneficial to improving the image enhancement effect.

[0126] In some embodiments, the image processing method provided in this application can determine T first 3D LUT parameters based on trilinear interpolation. As an example, such as... Figure 6 As shown, the above S505 can be implemented by S601-S603:

[0127] S601: Based on the lighting scene classification information and depth image features of the facial region image, determine the T sets of weight values ​​corresponding to each of the T semantic regions.

[0128] In some embodiments, the T sets of weight values ​​corresponding to each of the T semantic regions in this application can be: Figure 2 or Figure 4 The diagram shows the T sets of weight values ​​predicted by encoder 2. In this application, when N basic 3D LUTs are provided, each of the T sets of weight values ​​includes N weight values.

[0129] S602: Use T sets of weight values ​​to perform weighted calculations on N basic 3D LUTs to obtain T second 3D LUT parameters.

[0130] For example, T second 3D LUTs can be Figure 2 or Figure 4 The T weighted 3D LUTs are shown. Furthermore, this application can use formula (2) to perform weighted calculations on N basic 3D LUTs using T sets of weight values.

[0131] S603: Perform trilinear interpolation on the T second 3D LUT parameters to obtain T first 3D LUT parameters.

[0132] It is understood that the trilinear interpolation of the T second 3D LUT parameters in this application can be referred to the relevant description of formula (3-9) above, and will not be repeated here.

[0133] For example, the size of each base 3D LUT can be (33*33*33), and the size of each first 3D LUT parameter can be (256*256*256).

[0134] Thus, this application can predict N weight values ​​corresponding to different semantic regions in a portrait image, and calculate N basic 3D LUTs based on the N weight values ​​of different semantic regions, thereby obtaining an adaptive 3D LUT corresponding to different semantic regions. Therefore, the adaptive 3D LUTs for different semantic regions help to make the zoning enhancement of portrait images meet the actual needs of users.

[0135] In some embodiments, this application provides an image processing method that can combine non-uniform interval processing and trilinear interpolation processing to determine T first 3D LUT parameters. As an example, such as... Figure 7 As shown, the above S505 can be implemented by S701-S703:

[0136] S701: Based on the lighting scene classification information and depth image features of the facial region image, determine T sets of weight values ​​corresponding to each of the T semantic regions, and T sets of sampling intervals corresponding to each of the T semantic regions. The similarities between S701 and S601 are not repeated here; the difference is the addition of T sets of predicted sampling intervals. For example, the T sets of sampling intervals are T*32*3, or T*96 intervals.

[0137] S702: Use T sets of weight values ​​to perform weighted calculations on N basic 3D LUTs to obtain T second 3D LUT parameters.

[0138] The similarities between S702 and S602 will not be elaborated here.

[0139] S703: Using T sampling intervals, T second 3D LUT parameters are sampled at non-uniform intervals to obtain T third 3D LUT parameters.

[0140] In this application, a set of sampling intervals corresponding to a semantic region is used to sample the second 3DLUT parameters corresponding to the semantic region at non-uniform intervals to obtain the third 3D LUT parameters corresponding to the semantic region.

[0141] S704: Perform trilinear interpolation on the T third 3D LUT parameters to obtain the T first 3D LUT parameters.

[0142] In this application, the third 3D LUT parameter corresponding to a semantic region is trilinearly interpolated to obtain the first 3D LUT parameter corresponding to that semantic region.

[0143] Thus, this application can combine non-uniform interval sampling and trilinear interpolation to determine different 3D LUT parameters corresponding to different semantic regions, thereby improving the adaptability of semantic regions to 3D LUT parameters and improving the final portrait partitioning enhancement effect.

[0144] In some embodiments, the image processing method provided in this application can employ... Figure 2 or Figure 4 The image enhancement network implementation is shown in the figure. (Refer to...) Figure 8 The diagram shows the training and application process of an image enhancement network, which is still executed by a mobile phone or other terminal.

[0145] S801: Based on multiple training facial region images and the illumination table corresponding to the multiple training facial region images, train encoder 1 in the image augmentation network until the value of the loss function of encoder 1 is within the preset value range of 1, and then end the training.

[0146] Among them, the illumination table corresponding to multiple training facial region images can provide information on the illumination conditions of these training facial region images that have been pre-labeled by the user.

[0147] It is understandable that, based on the difference between the scene classification and recognition result (such as front lighting or backlighting) of encoder 1 for a training facial image and the lighting conditions corresponding to that training facial image in the lighting table, a value of the loss function of encoder 1 can be determined. Furthermore, the specific value within the aforementioned range 1 can be set according to actual needs and is not specifically limited here.

[0148] S802: Based on multiple training portrait images, N-1 semantic segmentation mask images corresponding to each training portrait image, lighting scene classification information of multiple training facial images corresponding to multiple training portrait images, and multiple training enhanced portrait images, train encoder 2 in the image enhancement network until the value of the loss function of encoder 2 is within the preset value range 2 and then training ends.

[0149] It can be understood that, based on the difference between the enhanced portrait image obtained by encoder 1 after image enhancement of a portrait image and the corresponding training enhanced image, a value of the loss function of encoder 2 can be determined. Furthermore, the specific value within the aforementioned range 2 can be set according to actual needs and is not specifically limited here.

[0150] S803: Input the image to be processed, the corresponding T-1 segmentation mask images, and the facial region image into the image enhancement network, and output the enhanced image corresponding to the image to be processed through the image enhancement network.

[0151] It is understandable that the image enhancement network in S803 can be based on Figures 5-7 The image processing method described in the article enables scene- and region-specific portrait enhancement of the portrait image to be processed.

[0152] Next, taking a mobile phone as an example of an electronic device that performs the image processing method of this application, the structure of the electronic device will be described.

[0153] like Figure 9 As shown, the mobile phone 10 may include a processor 110, a power module 140, a memory 180, a mobile communication module 130, a wireless communication module 120, a sensor module 190, an audio module 150, a camera 170, an interface module 160, buttons 101, and a display screen 102, etc.

[0154] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the mobile phone 10. In other embodiments of this application, the mobile phone 10 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0155] Processor 110 may include one or more processing units, such as processing modules or processing circuits of a central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), microprocessor (MCU), artificial intelligence (AI) processor, or field programmable gate array (FPGA). Different processing units may be independent devices or integrated into one or more processors. Processor 110 may include storage units for storing instructions and data. In some embodiments, the storage unit in processor 110 is a cache memory 180. For example, processor 110 can... Figure 2 or Figure 4 The image enhancement network shown performs the functions of this application. Figures 5-8 Image processing methods.

[0156] The power module 140 may include a power supply, a power management component, etc. The power supply may be a battery. The power management component manages the charging of the power supply and the power supply to other modules. In some embodiments, the power management component includes a charging management module and a power management module. The charging management module receives charging input from a charger; the power management module connects to the power supply and the processor 110. The power management module receives input from the power supply and / or the charging management module to supply power to the processor 110, the display 102, the camera 170, and the wireless communication module 120, etc.

[0157] The mobile communication module 130 may include, but is not limited to, an antenna, a power amplifier, a filter, and a low-noise amplifier (LNA). The mobile communication module 130 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on the mobile phone 10. The mobile communication module 130 can receive electromagnetic waves via the antenna, filter and amplify the received electromagnetic waves, and then transmit them to a modem processor for demodulation. The mobile communication module 130 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna. In some embodiments, at least some functional modules of the mobile communication module 130 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 130 and at least some modules of the processor 110 may be housed in the same device.

[0158] The wireless communication module 120 may include an antenna, which enables the transmission and reception of electromagnetic waves. The wireless communication module 120 can provide solutions for wireless communication applications on the mobile phone 10, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The mobile phone 10 can communicate with networks and other devices through wireless communication technologies.

[0159] In some embodiments, the mobile communication module 130 and the wireless communication module 120 of the mobile phone 10 may also be located in the same module.

[0160] The display screen 102 is used to display human-computer interaction interfaces, images, videos, etc.

[0161] The sensor module 190 may include proximity sensors, pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.

[0162] The audio module 150 is used to convert digital audio information into analog audio signals for output, or to convert analog audio input into digital audio signals. The audio module 150 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 150 may be located in the processor 110, or some functional modules of the audio module 150 may be located in the processor 110. In some embodiments, the audio module 150 may include a speaker, a handset, a microphone, and a headphone jack.

[0163] Camera 170 is used to capture still images or videos. An object passes through the lens to generate an optical image that is projected onto a photosensitive element. The photosensitive element converts the light signal into an electrical signal, which is then passed to image signal processing (ISP) to be converted into a digital image signal. Mobile phone 10 can perform shooting functions through the ISP, camera 170, video codec, graphics processing unit (GPU), display screen 102, and application processor. For example, in this application, the mobile phone can acquire images through the ISP, camera 170, etc., and perform post-processing on these images according to this application. Figures 5-8 The relevant image processing workflow.

[0164] Interface module 160 includes an external memory interface, a universal serial bus (USB) interface, and a subscriber identification module (SIM) card interface. The external memory interface can be used to connect an external memory card, such as a microSD card, to expand the storage capacity of the mobile phone 10. The external memory card communicates with the processor 110 through the external memory interface to perform data storage. The USB interface is used for communication between the mobile phone 10 and other electronic devices. The SIM card interface is used to communicate with the SIM card installed in the mobile phone 10, for example, to read or write phone numbers stored in the SIM card.

[0165] In some embodiments, the mobile phone 10 further includes buttons 101, a motor, and indicators. The buttons 101 may include volume buttons, a power button, etc. The motor is used to generate a vibration effect in the mobile phone 10, for example, vibrating when the user's mobile phone 10 is called to prompt the user to answer the call. The indicators may include laser indicators, radio frequency indicators, LED indicators, etc.

[0166] In some embodiments, this application provides a readable medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the image processing method described above.

[0167] In some embodiments, this application provides an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, one of the processors of the electronic device, for performing the image processing method described above.

[0168] In some embodiments, this application provides a computer program product, including: a computer program / instructions that, when executed by a processor, implement the image processing method described above.

[0169] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0170] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.

[0171] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0172] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other propagation signals. Therefore, machine-readable media include any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.

[0173] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.

[0174] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.

[0175] It should be noted that in the examples and description of this patent, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0176] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.

Claims

1. An image processing method, characterized by, The method is applied to an electronic device, and the method comprises: obtaining a first image to be processed, the first image being a portrait image; determining a first semantic region and a second semantic region in the first image, the first semantic region comprising at least one of a human body region, a human face region, and a human mouth region; determining first scene classification information corresponding to the first image according to an image feature of a first face region of a person in the first image, the first scene classification information being related to at least one of a lighting condition in which the first image is captured, a skin color of the person in the first image, and a color temperature of the first image; obtaining first depth image features corresponding to the first image; determining a first color parameter and a second color parameter according to the first depth image features and the first scene classification information, wherein the first color parameter and the second color parameter are different 3D LUT parameters; the first color parameter and the second color parameter are both 3D LUT parameters obtained by performing weighted processing and trilinear interpolation on preset N first 3D LUT parameters, or the first color parameter and the second color parameter are both 3D LUT parameters obtained by performing weighted processing, non-uniform interval processing, and trilinear interpolation on the preset N first 3D LUT parameters; adjusting colors of the first semantic region and the second semantic region respectively by using the first color parameter and the second color parameter corresponding to the first scene classification information of the first image; obtaining a second image based on the adjusted first semantic region and the adjusted second semantic region.

2. The method of claim 1, wherein, The first semantic region is a foreground region in the first image, and the second semantic region is a background region in the first image; or The first semantic region is a different foreground region in the first image.

3. The method of claim 2, wherein, The determination of the first semantic region and the second semantic region in the first image comprises: obtaining a first segmentation mask image of the first image, the first segmentation mask image being a segmentation mask image of a foreground in the first image; determining the first semantic region according to the first segmentation mask image, wherein the first semantic region is a foreground region in the first image; determining the second semantic region in the first image according to the first semantic region, wherein the second semantic region is a background region in the first image.

4. The method of claim 1, wherein, The adjustment of the colors of the first semantic region and the second semantic region respectively by using the first color parameter and the second color parameter corresponding to the first scene classification information of the first image comprises: adjusting the colors of the first image respectively by using the first color parameter and the second color parameter to obtain a third image containing the adjusted first semantic region and a fourth image containing the adjusted second semantic region; and obtaining the second image based on the adjusted first semantic region and the adjusted second semantic region comprises: fuse the adjusted first semantic region in the third image and the adjusted second semantic region in the fourth image to obtain the second image.

5. The method of claim 4, wherein, The number of the first semantic regions is M, one of the first semantic regions corresponds to one of the first color parameters, and The adjusting the color of the first image by using the first color parameter and the second color parameter respectively to obtain a third image containing an adjusted first semantic region and a fourth image containing an adjusted second semantic region comprises: adjusting the color of the first image by using the first color parameter corresponding to each of the M first semantic regions respectively to generate a corresponding third image for each of the first semantic regions, and obtaining M third images; adjusting the color of the first image by using the second color parameter to obtain the fourth image containing the second semantic region; and fusing the adjusted first semantic region in the third image and the adjusted second semantic region in the fourth image to obtain the second image comprises: fusing the adjusted first semantic region in each of the third images and the adjusted second semantic region in the fourth image to obtain the second image.

6. The method according to claim 4 or 5, characterized in that, The determining the first color parameter and the second color parameter according to the first depth image feature and the first scene classification information comprises: According to the first depth image feature and the first scene classification information, obtaining N first weight values corresponding to the first semantic region and N second weight values corresponding to the second semantic region, wherein the N first weight values are related to the color of the first semantic region, and the N second weight values are related to the color of the second semantic region; performing weighted calculation on the N first 3D LUT parameters by using the N first weight values and the N second weight values respectively to obtain third 3D LUT parameters and fourth 3D LUT parameters; performing trilinear interpolation on the third 3D LUT parameters and the fourth 3D LUT parameters respectively to obtain the first color parameter and the second color parameter.

7. The method according to claim 4 or 5, characterized in that, The determining the first color parameter and the second color parameter according to the first depth image feature and the first scene classification information comprises: According to the first depth image feature and the first scene classification information, obtaining N first weight values corresponding to the first semantic region and N second weight values corresponding to the second semantic region, N being a positive integer; According to the first depth image feature and the first scene classification information, obtaining S first sampling intervals corresponding to the first semantic region and S second sampling intervals corresponding to the second semantic region, S being a positive integer; performing weighted calculation on the N first 3D LUT parameters by using the N first weight values and the N second weight values respectively to obtain third 3D LUT parameters and fourth 3D LUT parameters; The third 3D LUT parameter is sampled by using the S first sampling intervals to obtain a fifth 3D LUT parameter; The fourth 3D LUT parameter is sampled by using the S second sampling intervals to obtain a sixth 3D LUT parameter; The fifth 3D LUT parameter and the sixth 3D LUT parameter are respectively subjected to trilinear interpolation to obtain the first color parameter and the second color parameter.

8. An apparatus for running an image processing model, the apparatus comprising: Comprise: The scene classification module is used for generating first scene classification information corresponding to the first image according to the image of the first face region of the first image to be processed; The first image is a portrait image, and the first scene classification information is related to at least one of the following: the light condition of the first image, the skin color of the person in the first image, and the color temperature of the first image; The parameter generation module is used for determining a first semantic region and a second semantic region in the first image, wherein the first semantic region comprises at least one of a human body region, a face region, and a mouth region; The image enhancement module is used for adjusting the color of the first semantic region and the second semantic region respectively according to first color parameters and second color parameters corresponding to the first scene classification information of the first image, and obtaining a second image based on the adjusted first semantic region and the adjusted second semantic region; The scene classification module comprises a second encoder, The second encoder is specifically used for inputting the image of the first face region of the person in the first image, and determining the first scene classification information according to the image features of the first face region in the first image; The first color parameter and the second color parameter are determined according to first depth image features of the first image, the first scene classification information, and preset N first 3D LUT parameters, wherein the first color parameter and the second color parameter are different 3D LUT parameters; the first color parameter and the second color parameter are both 3D LUT parameters obtained by weighting processing and trilinear interpolation on the preset N first 3D LUT parameters, or the first color parameter and the second color parameter are both 3D LUT parameters obtained by weighting processing, non-uniform interval processing, and trilinear interpolation on the preset N first 3D LUT parameters.

9. The apparatus of claim 8, wherein The parameter generation module comprises a first encoder, and the first encoder is used for inputting the first image and a first segmentation mask image of the first image, wherein the first segmentation mask image is a segmentation mask image of a foreground in the first image, The first semantic region is determined according to the first segmentation mask image, wherein the first semantic region is a foreground region in the first image, The second semantic region in the first image is determined according to the first semantic region, wherein the second semantic region is a background region in the first image.

10. The apparatus of claim 8, wherein The parameter generation module is configured to acquire N first weight values corresponding to the first semantic region and N second weight values corresponding to the second semantic region according to the first depth image feature and the first scene classification information, wherein the N first weight values are related to the color of the first semantic region, and the N second weight values are related to the color of the second semantic region. The image enhancement module is configured to perform weighted calculation on N first 3D LUT parameters in the preset by using the N first weight values and the N second weight values respectively, to obtain third 3D LUT parameters and fourth 3D LUT parameters, The third 3D LUT parameters and the fourth 3D LUT parameters are respectively subjected to trilinear interpolation to obtain the first color parameter and the second color parameter.

11. The apparatus of claim 8, wherein, The parameter generation module is configured to acquire N first weight values corresponding to the first semantic region and N second weight values corresponding to the second semantic region according to the first depth image feature and the first scene classification information, wherein the N first weight values are related to the color of the first semantic region, and the N second weight values are related to the color of the second semantic region, According to the first depth image feature and the first scene classification information, S first sampling intervals corresponding to the first semantic region and S second sampling intervals corresponding to the second semantic region are acquired, S being a positive integer; The image enhancement module is configured to perform weighted calculation on N first 3D LUT parameters in the preset by using the N first weight values and the N second weight values respectively, to obtain third 3D LUT parameters and fourth 3D LUT parameters, The third 3D LUT parameters are subjected to sampling processing by using the S first sampling intervals to obtain fifth 3D LUT parameters, The fourth 3D LUT parameters are subjected to sampling processing by using the S second sampling intervals to obtain sixth 3D LUT parameters, The fifth 3D LUT parameters and the sixth 3D LUT parameters are respectively subjected to trilinear interpolation to obtain the first color parameter and the second color parameter.

12. An image processing method, characterized by, The method is applied to an electronic device, and the method comprises: inputting a first image to be processed into a device running an image processing model according to any one of claims 8 to 11; obtaining a second image output by the device running the image processing model.

13. An image processing model training method, applied to a device running an image processing model according to any one of claims 8 to 11, the method comprising: inputting training sample data into a neural network model to be trained, the training sample data comprising a sample image, a scene type to which the sample image belongs, and a standard enhanced image corresponding to the sample image, the standard enhanced image being an image obtained by performing image processing on a plurality of semantic regions of the sample image by using different color parameters according to the scene type; obtaining a first enhanced image and a scene classification result output by the neural network model; The scene classification result is related to at least one of the following: an illumination condition in which the sample image is captured, a human skin color in the sample image, and a color temperature of the sample image. Based on the first enhanced image, a standard enhanced image, the scene classification result, and the scene type, parameters of the neural network model are adjusted to obtain the image processing model in a case where the first enhanced image and the scene classification result output by the neural network model satisfy a first condition.

14. The method of claim 13, wherein, The adjusting of the parameters of the neural network model based on the first enhanced image, the standard enhanced image, the scene classification result, and the scene type comprises: determining a value of a first loss function according to the scene classification result and the scene type; adjusting a first parameter of the neural network model according to the value of the first loss function; determining a value of a second loss function according to the first enhanced image and the standard enhanced image; adjusting a second parameter of the neural network model according to the value of the second loss function.

15. The method of claim 14, wherein, The first condition comprises: the value of the first loss function is in a first value range, and the value of the second loss function is in a second value range.

16. A readable medium characterized by The readable medium stores instructions, and the instructions, when executed on an electronic device, cause the electronic device to perform the image processing method according to any one of claims 1 to 7.

17. An electronic device, comprising: comprise: a memory configured to store instructions executed by one or more processors of an electronic device, and a processor, which is one of the processors of the electronic device, configured to perform the image processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Character picture color style converting method based on face semantic analysis

    CN104732506A

  • Image color enhancement method and device, equipment and storage medium

    CN110363720A