Image Processing Method, Apparatus, Electronic Device, and Storage Medium
By performing multiple filtering processing on the image and frequency band image difference processing, the target face image is generated, which solves the problem that the oily light in the hair area in the image is difficult to eliminate, and improves the image display effect.
Patent Information
- Application Number
- CN202210871899.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-22
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-07-22
AI Technical Summary
The prior art is difficult to effectively eliminate oil in the hair area in the image, affecting the display effect of the image.
By performing multiple filtering of the brightness channel image of the initial face image, multiple filtered images are acquired, and target face images are generated based on these images to eliminate highlights in the hair region.
It realizes the removal of oil in the hair area in the image and improves the display effect of face images.
Smart Images

Figure CN115330610B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimedia technologies, and in particular, to an image processing method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of computer technologies, more and more users record their lives by taking pictures of people.
[0003] In related technologies, in order to improve the effect of the captured image, the face region in the image is often beautified, making the face in the image more delicate. However, in some cases, there is still oiliness in the hair region of the image, resulting in a poor effect of the presented portrait in the image. Based on this, there is an urgent need for an image processing method that can eliminate the oiliness in the hair region. Summary of the Invention
[0004] This application provides an image processing method, apparatus, electronic device, and storage medium, which can eliminate the oiliness in the hair region. The technical solution of this application is as follows:
[0005] On the one hand, an image processing method is provided, including the following steps:
[0006] Perform filtering processing on the luminance channel image of the initial face image to obtain a first filtered image of the initial face image, where the initial face image includes a hair region;
[0007] Obtain a second filtered image and a third filtered image of the initial face image, where the second filtered image is obtained by performing filtering processing on the luminance channel image, and the third filtered image is obtained by performing filtering processing on the second filtered image;
[0008] Based on the luminance channel image, the first filtered image, the second filtered image, and the third filtered image, obtain a first initial frequency band image, a second initial frequency band image, and a third initial frequency band image of the initial face image, where the first initial frequency band image is a difference image between the initial face image and the first filtered image, the second initial frequency band image is a difference image between the first filtered image and the second filtered image, and the third initial frequency band image is a difference image between the second filtered image and the third filtered image;
[0009] Generate a target face image based on the first initial frequency band image, the second initial frequency band image, the third initial frequency band image, the target probability image of the initial face image, and the third filtered image. The target probability image is used to represent the positions of the high-brightness points within the hair region of the initial face image, where the high-brightness points are pixel points whose brightness meets a preset brightness condition. The target face image is the face image obtained after eliminating the high-brightness points in the initial face image.
[0010] In a possible implementation manner, the generating the target face image based on the first initial frequency band image, the second initial frequency band image, the third initial frequency band image, the target probability image of the initial face image, and the third filtered image includes:
[0011] Based on the target probability image, perform pixel value adjustment processing on the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image to obtain a first target frequency band image, a second target frequency band image, and a third target frequency band image;
[0012] Fuse the first target frequency band image, the second target frequency band image, the third target frequency band image, and the third filtered image to obtain a fused image;
[0013] Perform color space transformation on the fused image to obtain the target face image, and the target face image and the initial face image are in the same color space.
[0014] In a possible implementation manner, the performing pixel value adjustment processing on the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image based on the target probability image to obtain a first target frequency band image, a second target frequency band image, and a third target frequency band image includes:
[0015] Based on the target probability image, determine a plurality of target pixel points in the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image respectively, and the target pixel points correspond to the high-brightness points in the luminance channel image;
[0016] When the pixel values of the plurality of target pixel points meet a preset pixel value condition, adjust the pixel values of the plurality of target pixel points based on a first preset threshold and a second preset threshold to obtain the first target frequency band image, the second target frequency band image, and the third target frequency band image, and the second preset threshold is greater than the first preset threshold.
[0017] In a possible implementation manner, adjusting the pixel values of the multiple target pixel points based on the first preset threshold and the second preset threshold to obtain the first target frequency band image, the second target frequency band image, and the third target frequency band image includes:
[0018] For any target pixel point in the first initial frequency band image, when the pixel value of the target pixel point is less than or equal to the first preset threshold, keep the pixel value of the target pixel point unchanged;
[0019] When the pixel value of the target pixel point is greater than or equal to the second preset threshold, multiply the pixel value of the target pixel point by a first parameter to obtain the adjusted pixel value of the target pixel point, where the first parameter is a positive number less than 1;
[0020] When the pixel value of the target pixel point is greater than the first preset threshold and less than the second preset threshold, perform a fusion process on the pixel value of the target pixel point, the first parameter, and a second parameter to obtain the adjusted pixel value of the target pixel point, where the second parameter is determined based on the first preset threshold, the second preset threshold, and the pixel value of the target pixel point.
[0021] In a possible implementation manner, the performing a fusion process on the pixel value of the target pixel point, the first parameter, and the second parameter to obtain the adjusted pixel value of the target pixel point includes:
[0022] Process the first preset threshold, the second preset threshold, and the pixel value of the target pixel point using a smooth step model to obtain the second parameter;
[0023] Add the product of the first parameter and the pixel value of the target pixel point and the second parameter, and the product of a third parameter and the pixel value of the target pixel point to obtain the adjusted pixel value of the target pixel point, where the sum of the third parameter and the second parameter is 1.
[0024] In a possible implementation manner, filtering the luminance channel image of the initial face image to obtain the first filtered image of the initial face image includes:
[0025] Slide a Gaussian convolution kernel on the luminance channel image, and perform Gaussian filtering within the area covered by the Gaussian convolution kernel to obtain a first filtered image of the initial face image, where the horizontal step size when the Gaussian convolution kernel slides is the first step size, the vertical step size when the Gaussian convolution kernel slides is the second step size, the first step size is the ratio of the width of the face area in the initial face image to the width of the preset face area, and the second step size is the ratio of the height of the face area in the initial face image to the height of the preset face area.
[0026] In a possible implementation manner, the face area in the initial face image is obtained through the following method:
[0027] Perform key point detection on the initial face image to obtain multiple face key points in the initial face image;
[0028] Determine the minimum bounding rectangle of the multiple face key points as the face area in the initial face image.
[0029] In a possible implementation manner, the obtaining of the second filtered image and the third filtered image of the initial face image includes:
[0030] Taking the luminance channel image as a guidance image, perform guided filtering on the luminance channel image to obtain the second filtered image;
[0031] Taking the second filtered image as a guidance image, perform guided filtering on the luminance channel image to obtain the third filtered image;
[0032] Wherein, the horizontal step size of the two guided filterings on the luminance channel image is the first step size, and the vertical step size is the second step size.
[0033] In a possible implementation manner, before filtering the luminance channel image of the initial face image to obtain the first filtered image of the initial face image, the method further includes:
[0034] Transform the initial face image to the Lab color space;
[0035] Confirm the image of the L channel in the Lab color space as the luminance channel image of the initial face image.
[0036] In a possible implementation manner, before generating the target face image based on the first initial frequency band image, the second initial frequency band image, the third initial frequency band image, the target probability image of the initial face image, and the third filtered image, the method further includes:
[0037] Perform image recognition on the initial face image to obtain an initial probability image of the initial face image, where the initial probability image is used to represent the probability that the pixel points in the initial face image correspond to hair;
[0038] Convert the initial face image into a luminance image;
[0039] Generate a target probability image of the initial face image based on the initial probability image and the luminance image.
[0040] In a possible implementation manner, the converting the initial face image into a luminance image includes:
[0041] Determine the luminance values of multiple pixel points in the initial face image based on the color channel values of the initial face image in the three RGB color channels;
[0042] Generate a luminance image of the initial face image based on the luminance values of multiple pixel points in the initial face image.
[0043] In a possible implementation manner, the generating a target probability image of the initial face image based on the initial probability image and the luminance image includes:
[0044] Determine the hair region on the initial face image based on the initial probability image;
[0045] Determine the average luminance value of the hair region based on the luminance image of the initial face image;
[0046] Binarize the luminance image based on the average luminance value of the hair region to obtain a reference luminance image;
[0047] Perform dilation processing on the reference luminance image, and fuse the reference luminance image obtained after the dilation processing with the initial probability image to obtain a target probability image of the initial face image.
[0048] In a possible implementation manner, the determining the hair region on the initial face image based on the initial probability image includes:
[0049] Binarize the initial probability image based on a probability threshold to obtain a reference probability image of the initial face image;
[0050] Determine the hair region on the initial face image based on the reference probability image of the initial face image.
[0051] In a possible implementation manner, the binarizing the initial probability image based on a probability threshold includes:
[0052] For any pixel point in the initial probability image, when the probability corresponding to the pixel point is greater than or equal to the probability threshold, adjust the probability corresponding to the pixel point to a first value; when the probability corresponding to the pixel point is less than the probability threshold, adjust the probability corresponding to the pixel point to a second value, where the first value is greater than the second value.
[0053] In a possible implementation manner, the binarizing the luminance image based on the average luminance value of the hair region to obtain a reference luminance image includes:
[0054] For any pixel point in the luminance image, when the luminance value of the pixel point is greater than or equal to the average luminance value, adjust the luminance value of the pixel point to a third value; when the luminance value of the pixel point is less than the average luminance value, adjust the luminance value of the pixel point to a fourth value, where the third value is greater than the fourth value.
[0055] In a possible implementation manner, the fusing the reference luminance image obtained after dilation processing with the initial probability image to obtain a target probability image of the initial face image includes:
[0056] Performing Gaussian filtering on the reference luminance image obtained after dilation processing based on the initial probability image to obtain a target luminance image of the initial face image;
[0057] Multiplying the values of the corresponding pixel points in the initial probability image and the target luminance image to obtain a target probability image of the initial face image.
[0058] On the one hand, there is provided an image processing apparatus, including:
[0059] A first filtering unit configured to perform filtering on a luminance channel image of an initial face image to obtain a first filtered image of the initial face image, where the initial face image includes a hair region;
[0060] A second filtering unit configured to obtain a second filtered image and a third filtered image of the initial face image, where the second filtered image is obtained by performing filtering on the luminance channel image, and the third filtered image is obtained by performing filtering on the second filtered image;
[0061] A frequency band image acquisition unit, configured to execute to obtain a first initial frequency band image, a second initial frequency band image, and a third initial frequency band image of the initial face image based on the luminance channel image, the first filtered image, the second filtered image, and the third filtered image, where the first initial frequency band image is a difference image between the initial face image and the first filtered image, the second initial frequency band image is a difference image between the first filtered image and the second filtered image, and the third initial frequency band image is a difference image between the second filtered image and the third filtered image;
[0062] A target face image generation unit, configured to execute to generate a target face image based on the first initial frequency band image, the second initial frequency band image, the third initial frequency band image, a target probability image of the initial face image, and the third filtered image, where the target probability image is used to represent positions of high-brightness points within the hair region in the initial face image, the high-brightness points are pixel points whose luminance meets a preset luminance condition, and the target face image is a face image obtained after eliminating the high-brightness points in the initial face image.
[0063] In a possible implementation manner, the target face image generation unit is configured to execute to perform pixel value adjustment processing on the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image based on the target probability image to obtain a first target frequency band image, a second target frequency band image, and a third target frequency band image; perform fusion processing on the first target frequency band image, the second target frequency band image, the third target frequency band image, and the third filtered image to obtain a fused image; perform color space transformation on the fused image to obtain the target face image, and the target face image and the initial face image are in the same color space.
[0064] In a possible implementation manner, the target face image generation unit is configured to execute to determine a plurality of target pixel points in the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image respectively based on the target probability image, where the target pixel points correspond to the high-brightness points in the luminance channel image; when the pixel values of the plurality of target pixel points meet a preset pixel value condition, adjust the pixel values of the plurality of target pixel points based on a first preset threshold and a second preset threshold to obtain the first target frequency band image, the second target frequency band image, and the third target frequency band image, and the second preset threshold is greater than the first preset threshold.
[0065] In a possible implementation, the target face image generation unit is configured to perform the following operations for any target pixel point in the first initial frequency band image: when the pixel value of the target pixel point is less than or equal to the first preset threshold, keep the pixel value of the target pixel point unchanged; when the pixel value of the target pixel point is greater than or equal to the second preset threshold, multiply the pixel value of the target pixel point by a first parameter to obtain the adjusted pixel value of the target pixel point, where the first parameter is a positive number less than 1; when the pixel value of the target pixel point is greater than the first preset threshold and less than the second preset threshold, perform a fusion process on the pixel value of the target pixel point, the first parameter, and a second parameter to obtain the adjusted pixel value of the target pixel point, where the second parameter is determined based on the first preset threshold, the second preset threshold, and the pixel value of the target pixel point.
[0066] In a possible implementation, the target face image generation unit is configured to perform the following operations: process the first preset threshold, the second preset threshold, and the pixel value of the target pixel point using a smooth step model to obtain the second parameter; add the product of the first parameter and the pixel value of the target pixel point, the product of the second parameter and the pixel value of the target pixel point, and the product of a third parameter and the pixel value of the target pixel point to obtain the adjusted pixel value of the target pixel point, where the sum of the third parameter and the second parameter is 1.
[0067] In a possible implementation, the first filtering unit is configured to perform the following operations: slide a Gaussian convolution kernel on the luminance channel image and perform Gaussian filtering within the area covered by the Gaussian convolution kernel to obtain the first filtered image of the initial face image, where the horizontal step size when the Gaussian convolution kernel slides is the first step size, the vertical step size when the Gaussian convolution kernel slides is the second step size, the first step size is the ratio of the width of the face area in the initial face image to the width of a preset face area, and the second step size is the ratio of the height of the face area in the initial face image to the height of the preset face area.
[0068] In a possible implementation, the device further includes:
[0069] A face area determination unit, configured to perform key point detection on the initial face image to obtain multiple face key points in the initial face image; determine the minimum bounding rectangle of the multiple face key points as the face area in the initial face image.
[0070] In a possible implementation, the second filtering unit is configured to perform guided filtering on the luminance channel image with the luminance channel image as the guidance image to obtain the second filtered image; and perform guided filtering on the luminance channel image with the second filtered image as the guidance image to obtain the third filtered image; wherein, the horizontal step size of the two guided filterings on the luminance channel image is the first step size, and the vertical step size is the second step size.
[0071] In a possible implementation, the device further includes:
[0072] A luminance channel image confirmation unit, configured to perform transforming the initial face image into the Lab color space; and confirming the image of the L channel in the Lab color space as the luminance channel image of the initial face image.
[0073] In a possible implementation, the device further includes:
[0074] A target probability image generation unit, configured to perform image recognition on the initial face image to obtain an initial probability image of the initial face image, where the initial probability image is used to represent the probability that the pixel points in the initial face image correspond to hair; convert the initial face image into a luminance image; and generate a target probability image of the initial face image based on the initial probability image and the luminance image.
[0075] In a possible implementation, the target probability image generation unit is configured to perform determining the luminance values of multiple pixel points in the initial face image based on the color channel values of the initial face image in the three RGB color channels; and generating a luminance image of the initial face image based on the luminance values of multiple pixel points in the initial face image.
[0076] In a possible implementation, the target probability image generation unit is configured to perform determining the hair region on the initial face image based on the initial probability image; determining the average luminance value of the hair region based on the luminance image of the initial face image; performing binarization on the luminance image based on the average luminance value of the hair region to obtain a reference luminance image; performing dilation processing on the reference luminance image, and fusing the reference luminance image obtained after the dilation processing with the initial probability image to obtain the target probability image of the initial face image.
[0077] In a possible implementation, the target probability image generation unit is configured to perform binarization on the initial probability image based on a probability threshold to obtain a reference probability image of the initial face image; and determine a hair region on the initial face image based on the reference probability image of the initial face image.
[0078] In a possible implementation, the target probability image generation unit is configured to perform, for any pixel point in the initial probability image, when the probability corresponding to the pixel point is greater than or equal to the probability threshold, adjusting the probability corresponding to the pixel point to a first value; and when the probability corresponding to the pixel point is less than the probability threshold, adjusting the probability corresponding to the pixel point to a second value, where the first value is greater than the second value.
[0079] In a possible implementation, the target probability image generation unit is configured to perform, for any pixel point in the luminance image, when the luminance value of the pixel point is greater than or equal to the average luminance value, adjusting the luminance value of the pixel point to a third value; and when the luminance value of the pixel point is less than the average luminance value, adjusting the luminance value of the pixel point to a fourth value, where the third value is greater than the fourth value.
[0080] In a possible implementation, the target probability image generation unit is configured to perform Gaussian filtering on the reference luminance image obtained after the dilation processing based on the initial probability image to obtain a target luminance image of the initial face image; and multiply the values of the corresponding pixel points in the initial probability image and the target luminance image to obtain a target probability image of the initial face image.
[0081] On the one hand, an electronic device is provided, including:
[0082] A processor;
[0083] A memory for storing executable instructions of the processor;
[0084] Wherein, the processor is configured to execute the instructions to implement the above image processing method.
[0085] On the one hand, a computer-readable storage medium is provided, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the above image processing method.
[0086] On the one hand, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the above image processing method is implemented.
[0087] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects:
[0088] Through the technical solution provided by the embodiment of the present application, the luminance channel image of the initial face image is filtered multiple times to obtain a first filtered image, a second filtered image, and a third filtered image. Based on the luminance channel image, the first filtered image, the second filtered image, and the third filtered image, three frequency band images are obtained. Subsequently, the three frequency band images are processed, and the idea of frequency division and dimming is adopted to eliminate the highlights in the initial face image, obtaining the target face image. Since these highlights correspond to the oiliness of the hair in the initial face image, the oiliness in the hair area is eliminated after eliminating the highlights, thereby improving the display effect of the face image.
[0089] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application, and do not constitute an improper limitation of the present application.
[0091] Figure 1 is a schematic diagram of an implementation environment of an image processing method shown according to an exemplary embodiment.
[0092] Figure 2 is a flowchart of an image processing method shown according to an exemplary embodiment.
[0093] Figure 3 is a flowchart of another image processing method shown according to an exemplary embodiment.
[0094] Figure 4 is a flowchart of a method for obtaining a target probability map shown according to an exemplary embodiment.
[0095] Figure 5 is a block diagram of an image processing apparatus shown according to an exemplary embodiment.
[0096] Figure 6 is a block diagram of a terminal shown according to an exemplary embodiment.
[0097] Figure 7 is a block diagram of a server shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0098] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0099] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned appended drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are only examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.
[0100] In order to more clearly illustrate the embodiments of this application, some terms related to the embodiments of this application will be described below.
[0101] Gaussian filtering: Gaussian filtering is a linear smoothing filter, suitable for eliminating Gaussian noise and widely used in the denoising process of image processing. Generally speaking, Gaussian filtering is a process of weighted averaging of the entire image. The value of each pixel is obtained by weighted averaging of itself and other pixel values in its neighborhood. The operation of Gaussian filtering is: scan each pixel in the image with a template (or convolution, mask), and replace the value of the pixel at the center of the template with the weighted average gray value of the pixels in the neighborhood determined by the template.
[0102] Edge-preserving filtering: Edge-preserving filtering is a non-linear filtering method that does not erase the edge part after operating on the image. Common edge-preserving filters include bilateral filtering and guided filtering.
[0103] Guided filtering: Guided filtering is an image filtering technique that filters the input image through a guidance map, so that the output image is similar to the initial image on the large image, but the texture part is similar to the guidance map.
[0104] Lab (Lab color space): It is a color-opponent space, with the dimension L representing brightness, and a and b representing color opponent dimensions, based on the non-linearly compressed CIE XYZ color space coordinates. Among them, the value range of L is [0, 100], representing from pure black to pure white. The value range of a is [127, -128], representing from red to green. The value range of b is [127, -128], representing from yellow to blue.
[0105] Binarization: It is the simplest method of image segmentation. Binarization can convert a grayscale image into a binary image. The pixel grayscale values greater than a certain critical grayscale value are set to the maximum grayscale value, and the pixel grayscale values less than this value are set to the minimum grayscale value, thereby achieving binarization.
[0106] It should be noted that the information involved in this application (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0107] Figure 1 It is a schematic diagram of the implementation environment of an image processing method provided by an embodiment of this application. Refer to Figure 1 In this implementation environment, there are a terminal 101 and a server 102.
[0108] The terminal 101 can be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, and a notebook portable computer. An application program supporting image processing can be installed and run on the terminal 101, and the user can process images through this application program.
[0109] The terminal 101 generally refers to one of multiple terminals. In this embodiment, only the terminal 101 is used as an example for illustration. Those skilled in the art can know that the number of the above terminals can be more or less. For example, the above terminal 101 can be only a few, or the above terminal 101 can be dozens or hundreds, or a larger number. The embodiment of this application does not limit the number and device type of the terminal 101. The terminal 101 can be connected to the server 102 through a wireless network or a wired network.
[0110] The server 102 can be at least one of a server, multiple servers, a cloud computing platform, and a virtualization center. The server 102 provides background services for the application programs running on the terminal 101.
[0111] In some embodiments, the number of the above servers 102 can be more or less, and the embodiment of this application does not limit this. Of course, the server 102 can also include other functional servers to provide more comprehensive and diversified services.
[0112] After introducing the implementation environment of the embodiment of this application, the application scenarios of the embodiment of this application will be introduced below.
[0113] The technical solution provided by the embodiment of this application can be applied to scenarios of processing face images, can also be applied to scenarios of processing videos containing human figures, and can also be applied to scenarios of processing live data streams. The embodiment of this application does not limit this.
[0114] When the technical solution provided in the embodiment of the present application is applied to the scenario of processing a face image, the user selects the face image to be processed through the terminal, and the terminal can process the selected face image by using the technical solution provided in the embodiment of the present application. When there is oiliness in the hair area of the initial face image, the terminal can eliminate the oiliness in the hair area of the initial face image, making the processed face image more delicate and fresh.
[0115] When the technical solution provided in the embodiment of the present application is applied to the scenario of processing a video containing a portrait, the user selects the video to be processed through the terminal, and the terminal processes multiple video frames of the selected video by using the technical solution provided in the embodiment of the present application. When there is oiliness in the hair area displayed on any one of the multiple video frames, the terminal can eliminate the oiliness in the hair area of the video frame. The terminal processes multiple video frames, and the oiliness in the hair areas of multiple video frames in the video is eliminated.
[0116] It should be noted that the process of the terminal using the technical solution provided in the embodiment of the present application to process the live data stream belongs to the same inventive concept as the process of processing the video. For the implementation process, refer to the above description and will not be elaborated here.
[0117] After introducing the implementation environment and application scenarios of the embodiment of the present application, the image processing method provided in the embodiment of the present application will be described below. Refer to Figure 2 , taking the server as the execution subject as an example, the method includes the following steps.
[0118] In step S201, the server performs filtering processing on the luminance channel image of the initial face image to obtain the first filtered image of the initial face image, and the initial face image includes a hair area.
[0119] Among them, the initial face image is also the face image to be processed, and the luminance channel image is the image of the I channel of the initial face image in the Lab color space. The first filtering is used to eliminate the noise on the initial face image, and the first filtered image is also the image obtained after the initial face image is filtered. The hair area is the area where the hair is located in the initial face image.
[0120] In step S202, the server obtains the second filtered image and the third filtered image of the initial face image. The second filtered image is obtained by performing filtering processing on the luminance channel image, and the third filtered image is obtained by performing filtering processing on the second filtered image.
[0121] In some embodiments, the server filters the luminance channel image by means of guided filtering to obtain the second filtered image and the third filtered image, and the guided images corresponding to the second filtered image and the third filtered image are different.
[0122] In step S203, the server obtains the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image of the initial face image based on the luminance channel image, the first filtered image, the second filtered image, and the third filtered image. The first initial frequency band image is the difference image between the initial face image and the first filtered image, the second initial frequency band image is the difference image between the first filtered image and the second filtered image, and the third initial frequency band image is the difference image between the second filtered image and the third filtered image.
[0123] The difference image between the initial face image and the first filtered image refers to the image obtained by subtracting the pixel values of the corresponding pixel points in the initial face image and the first filtered image. The corresponding pixel points refer to the pixel points at the same position in the image.
[0124] In step S204, the server generates a target face image based on the first initial frequency band image, the second initial frequency band image, the third initial frequency band image, the target probability image of the initial face image, and the third filtered image. The target probability image is used to represent the positions of the high-brightness points in the hair region of the initial face image. The high-brightness points are pixel points whose brightness meets the preset brightness condition. The target face image is the face image obtained after eliminating the high-brightness points in the initial face image.
[0125] The high-brightness points in the hair region of the initial face image are also the oiliness of the hair in the initial face image. The target probability image is used to represent the positions of the high-brightness points in the initial face image, that is, the positions of the oiliness of the hair.
[0126] Through the technical solution provided by the embodiments of the present application, the luminance channel image of the initial face image is filtered multiple times to obtain the first filtered image, the second filtered image, and the third filtered image. Based on the luminance channel image, the first filtered image, the second filtered image, and the third filtered image, three frequency band images are obtained. Subsequently, the three frequency band images are processed, and the high-brightness points in the initial face image are eliminated by using the idea of frequency division and dimming to obtain the target face image. Since these high-brightness points correspond to the oiliness of the hair in the initial face image, the oiliness in the hair region is eliminated after eliminating the high-brightness points, thereby improving the display effect of the face image.
[0127] It should be noted that the above steps S201 - S204 are described by taking the server as the execution entity as an example. In other possible implementation manners, the terminal can also be used as the execution entity to execute the above steps S201 - S204, and the embodiments of the present application do not limit this.
[0128] The above steps S201 - S204 are a simple description of the image processing method provided by the embodiments of the present application. Below, some examples will be combined to provide a more detailed description of the image processing method provided by the embodiments of the present application. Refer to Figure 3 , taking the server as the execution entity as an example, the method includes the following steps.
[0129] In step S301, the server acquires an initial face image, and the initial face image includes a hair region.
[0130] Among them, the initial face image is an image containing a face. The server needs to obtain the initial face image with the permission of the user. For example, when the user selects the initial face image, the server displays a confirmation control. In response to a click operation on the confirmation control, the server acquires the initial face image. In some embodiments, the initial face image is any one of the following: an image taken by the user, an image in a video, and an image in a live video stream.
[0131] In a possible implementation manner, in response to an operation on the initial face image, the server acquires the initial face image. In this implementation manner, the user can select the initial image by operating on the initial image, and the autonomy of the user in selecting the initial image is relatively high. Simple operations can improve the efficiency of human - machine interaction.
[0132] For example, the terminal displays a face image selection interface, and multiple candidate face images are displayed on the initial face image selection interface. In response to a click operation on the initial face image among the multiple candidate face images, the terminal sends the initial face image to the server, and the server acquires the initial face image. Among them, the multiple candidate face images displayed on the face image selection interface can be either face images stored on the terminal or face images stored on the server, and the embodiments of the present application do not limit this.
[0133] In a possible implementation manner, the server acquires the initial face image from the live data stream. In this implementation manner, the server can acquire the initial face image from the live data stream without the need for the user to manually select, and the efficiency of human - machine interaction is relatively high.
[0134] For example, the server obtains multiple live video frames from the live data stream, and the multiple live video frames are uploaded by the terminal to the server. The server identifies the multiple live video frames and obtains the live video frames containing human faces among the multiple live video frames. The live video frames containing human faces are also the initial face images. Among them, the process of the server identifying the multiple live video frames and obtaining the live video frames containing human faces can be implemented by an image classification model. The image classification model can identify images containing human faces. The output result of the image classification model is either containing a human face or not containing a human face. The image classification model can be a binary classification model with any structure, and the embodiments of the present application do not limit this.
[0135] In step S302, the server obtains the luminance channel image of the initial face image.
[0136] Among them, the luminance channel image is the image of the I channel of the initial face image in the Lab color space. The color space where the initial face image is located is RGB, XYZ, etc., and the embodiments of the present application do not limit this.
[0137] In a possible implementation manner, the server transforms the initial face image into the Lab color space. The server confirms the image of the L channel in the Lab color space as the luminance channel image of the initial face image.
[0138] Among them, the value range of the pixel points in the image of the L channel is [0, 100], indicating from pure black to pure white. The L channel image can reflect the luminance of the initial face image as a whole.
[0139] In this implementation manner, since the Lab color space includes a luminance channel, directly transforming the initial face image into the Lab color space can quickly obtain the luminance channel image of the initial face image, and the efficiency of image processing is relatively high.
[0140] Taking the color space where the initial face image I is located as RGB as an example, the server transforms the initial face image I from the RGB color space to the XYZ color space. The server transforms the initial face image I from the XYZ color space to the Lab color space. The server confirms the image of the L channel of the initial face image I in the Lab color space as the luminance channel image I l .
[0141] In step S303, the server performs filtering processing on the luminance channel image of the initial face image to obtain the first filtered image of the initial face image.
[0142] Among them, the filtering process in step S303 refers to Gaussian filtering, which is used to eliminate Gaussian noise on the initial face image, and the first filtered image is also the image obtained after the initial face image is Gaussian filtered.
[0143] In a possible implementation manner, the server slides a Gaussian convolution kernel on the luminance channel image, and performs Gaussian filtering processing within the area covered by the Gaussian convolution kernel to obtain the first filtered image of the initial face image. Among them, the horizontal step size when the Gaussian convolution kernel slides is the first step size, the vertical step size when the Gaussian convolution kernel slides is the second step size, the first step size is the ratio of the width of the face area in the initial face image to the width of the preset face area, and the second step size is the ratio of the height of the face area in the initial face image to the height of the preset face area.
[0144] Among them, the preset face area is an area designed by technicians according to the actual situation. In some embodiments, the preset face area is a rectangular area, and the length and width of the rectangular area are set by technicians according to the actual situation. In some embodiments, both the length and width of the Gaussian convolution kernel are odd numbers. The size of the Gaussian convolution kernel is closely related to the effect of Gaussian filtering. The smaller the size of the Gaussian convolution kernel, the more details are retained after Gaussian filtering; the larger the size of the Gaussian convolution kernel, the faster the Gaussian filtering speed and the fewer details are retained. The size of the Gaussian convolution kernel is set by technicians according to the actual situation, such as set to 3×3, and the embodiments of the present application do not limit this. In addition, the Gaussian convolution kernel includes multiple weights, and the values of the multiple weights are generated by the server or set by technicians according to the actual situation, and the embodiments of the present application do not limit this.
[0145] In this implementation manner, the server can eliminate Gaussian noise in the luminance channel image of the initial face image through Gaussian filtering, that is, perform smoothing processing on the luminance channel image to improve the effect of subsequent image processing. During the Gaussian filtering process, the horizontal step size and the vertical step size are restricted, and both the horizontal step size and the vertical step size are determined based on the face area in the initial face image. When performing Gaussian filtering, it can focus on the face area and avoid interference from other areas.
[0146] For example, the server slides a Gaussian convolution kernel on the luminance channel image and performs convolution processing on the area covered by the Gaussian convolution kernel, that is, performs weighted summation of the weights in the Gaussian convolution kernel and the pixel values of the corresponding pixel points in the area to obtain a target pixel value, and uses the target pixel value to replace the pixel value of the pixel point corresponding to the point in the Gaussian convolution kernel. By controlling the continuous movement of the Gaussian convolution kernel on the luminance channel image, the server realizes Gaussian filtering of each pixel point on the luminance channel to obtain the first filtered image G1. The horizontal step length when the Gaussian convolution kernel slides is the first step length, and the vertical step length when the Gaussian convolution kernel slides is the second step length.
[0147] To illustrate the above embodiments more clearly, the method for the server to determine the face area in the initial face image will be described below.
[0148] In a possible implementation, the server performs key point detection on the initial face image to obtain multiple face key points in the initial face image. The server determines the minimum bounding rectangle of the multiple face key points as the face area in the initial face image.
[0149] Among them, the server can use any key point detection method to determine multiple face key points, and the embodiments of the present application do not limit this.
[0150] In this implementation, the server performs key point recognition on the initial face image and determines the face area based on the recognized face key points. Compared with overall face recognition, the accuracy of determining face key points is relatively high, and thus the accuracy of the face area determined using face key points is relatively high.
[0151] In a possible implementation, the server performs Fourier transform on the luminance channel image to obtain the frequency domain image of the luminance channel image. The server processes the frequency domain image using a Gaussian filtering function to obtain the filtered frequency domain image. The server performs inverse Fourier transform on the filtered frequency domain image to obtain the first filtered image G1.
[0152] In this implementation, when converting the luminance channel image to the frequency domain for Gaussian filtering processing, the amount of computation is independent of the Gaussian filtering function, and the filtering speed is relatively fast.
[0153] For example, the form of the Gaussian filtering function is shown in the following formula (1).
[0154]
[0155] Among them, H(u, v) is the value of the frequency point in the frequency spectrum diagram, u and υ are the coordinates of the frequency point, D 0 is the cut-off frequency, D(u, v) = u 2+v 2 That is, the spectrogram is also the frequency-domain image.
[0156] In step S304, the server obtains a second filtered image and a third filtered image of the initial face image. The second filtered image is obtained by filtering the luminance channel image, and the third filtered image is obtained by filtering the second filtered image.
[0157] Among them, the filtering process in step S304 refers to edge-preserving filtering. Edge-preserving filtering can filter the image while retaining the image edges. In some embodiments, the server performs edge-preserving filtering on the luminance channel image in the way of guided filtering to obtain the second filtered image and the third filtered image, and the corresponding guidance images of the second filtered image and the third filtered image are different. The purpose of guided filtering is to make the content of the output image similar to the input image, but the texture is similar to the guidance image.
[0158] In a possible implementation manner, the server uses the luminance channel image as the guidance image to perform guided filtering on the luminance channel image to obtain the second filtered image. The server uses the second filtered image as the guidance image to perform guided filtering on the luminance channel image to obtain the third filtered image. Among them, the horizontal step size of the two guided filterings of the luminance channel image is the above first step size, and the vertical step size is the above second step size, that is, the first step size is the ratio of the width of the face area in the initial face image to the width of the preset face area, and the second step size is the ratio of the height of the face area in the initial face image to the height of the preset face area.
[0159] In this implementation manner, the server can perform guided filtering on the luminance channel image with different guidance images, and the obtained second filtered image and third filtered image correspond to different frequencies, which is convenient for subsequent processing using the idea of frequency division and darkening. In addition, the horizontal step size and the vertical step size are restricted during the guided filtering process, and both the horizontal step size and the vertical step size are determined based on the face area in the initial face image. When performing guided filtering, it can focus on the face area and avoid interference from other areas.
[0160] To illustrate the above implementation manner more clearly, the above implementation manner will be described in two parts below.
[0161] The first part: The server uses the luminance channel image as the guidance image to perform guided filtering on the luminance channel image to obtain the second filtered image.
[0162] In a possible implementation manner, the server determines the pixel values of multiple pixels on the second filtered image GF1 based on the pixel values of multiple pixels on the luminance channel image.
[0163] For example, the guided filtering process satisfies the two constraint conditions in the following formula (2).
[0164]
[0165] Among them, q i is the pixel value of the i-th pixel on the second filtered image, and p i is the pixel value of the i-th pixel on the luminance channel image, n i is the noise of the i-th pixel on the luminance channel image, I i is the pixel value of the i-th pixel on the guidance image, and a and b are weights.
[0166] The server needs to solve n i , a, and b in the above formula (2) to complete the operation. The following explains the process of the server solving n i , a, and b.
[0167] Based on the above formula (2), the server derives the following formula (3).
[0168] n i = p i - aI i - b (3)
[0169] The server defines a cost formula (4).
[0170]
[0171] Among them, E(a k , b k ) is the cost function, ∈ is the penalty coefficient, used to penalize too large a k , w k is the pane centered on the k-th pixel. Through the above formula (4), it is hoped that the final output image can reduce noise and reduce the influence of the guidance image on the output image (the coefficient term of the guidance image) at the same time. Therefore, the sum of the squares of each pixel noise and the coefficient term is defined as the value term that must be paid finally. Based on the principle of minimizing the value of the cost function, (4) can be used to find its linear model by linear regression, so as to find the following two parameters a k and b k that make the cost equation have the minimum solution. The determination process is shown in the following formula (5).
[0172]
[0173]
[0174] Among them, μk and respectively represent the mean and variance of the guiding image in the local window, |ω| is the number of pixel points within the window, the horizontal step size of the local window is the first step size, and the vertical step size is the second step size.
[0175] In the second part, the server uses the second filtered image as the guiding image to perform guided filtering on the luminance channel image to obtain the second filtered image.
[0176] In a possible implementation manner, the server determines the pixel values of multiple pixel points on the third filtered image GF2 based on the pixel values of multiple pixel points on the second filtered image GF1 and the pixel values of multiple pixel points on the luminance channel image.
[0177] It should be noted that the method by which the server uses the second filtered image as the guiding image to perform guided filtering on the luminance channel image belongs to the same inventive concept as the description in the above first part. For the implementation process, refer to the description in the above first part, which will not be elaborated here.
[0178] In step S305, the server generates a target probability image of the initial face image.
[0179] Among them, the target probability image is used to represent the positions of the high-brightness points within the hair region of the initial face image, and the high-brightness points are pixel points whose brightness meets the preset brightness condition. In some embodiments, the high-brightness points are also the positions of the oiliness within the hair region.
[0180] In a possible implementation manner, the server performs image recognition on the initial face image to obtain an initial probability image of the initial face image, and the initial probability image is used to represent the probability that the pixel points in the initial face image correspond to hair. The server converts the initial face image into a luminance image. The server generates the target probability image of the initial face image based on the initial probability image and the luminance image.
[0181] In this implementation manner, the server can generate the target probability image based on image recognition, and the efficiency of image processing is relatively high.
[0182] To illustrate the above implementation manner more clearly, the above implementation manner will be described in several parts below.
[0183] In the first part, the server performs image recognition on the initial face image to obtain an initial probability image of the initial face image.
[0184] In a possible implementation, the server inputs the initial face image into an image recognition model, extracts features from the initial face image through the image recognition model, and obtains a feature image of the initial face image. The server maps the feature image of the initial face image through the image recognition model and outputs an initial probability image of the initial face image.
[0185] Among them, the image recognition model is trained through sample face images and has the function of identifying the hair area from face images. In some embodiments, the size of the initial probability image is the same as that of the initial face image, and the pixel value of a pixel point in the initial probability image represents the probability that the corresponding pixel point in the initial face image belongs to the hair area. For example, if the pixel value of a pixel point in the initial probability image is 1, it means that the probability that the corresponding pixel point in the initial face image belongs to the hair area is 1.
[0186] In this implementation, the initial probability image of the initial face image can be quickly obtained by using the image recognition model, and the efficiency of image processing is relatively high.
[0187] For example, the server inputs the initial face image into the image recognition model, extracts features from the initial face image through the feature extraction layer of the image recognition model, and obtains a feature image of the initial face image. Among them, the feature extraction layer of the image recognition model is a convolutional layer, an attention encoding layer, a fully connected layer, etc., and the embodiments of the present application do not limit this. When the feature extraction layer is a convolutional layer, the server convolves the initial face image through the convolution kernel on the convolutional layer to obtain a feature image of the initial face image; when the feature extraction layer is an attention encoding layer, the server embeds and encodes multiple parts of the initial face image through the attention encoding layer to obtain embedding vectors of the multiple parts, where the multiple parts are multiple regions of the initial face image. The server encodes the embedding vectors of the multiple parts based on the attention mechanism through the attention encoding layer to obtain attention encoding vectors of the multiple parts. The attention encoding vectors of the multiple parts form the feature image of the initial face; when the feature extraction layer is a fully connected layer, the server fully connects the initial face image through the fully connected layer to obtain a feature image of the initial face image. The server fully connects the feature image of the initial face image through the fully connected layer of the image recognition model to obtain a fully connected image of the initial face image. The server normalizes the fully connected image of the initial face image through the normalization layer of the image recognition model and outputs the initial probability image, where normalization includes processing the fully connected image by using a Softmax (soft maximization) function or a Sigmoid (S-shaped growth curve) function, and the embodiments of the present application do not limit this.
[0188] Part II: The server converts the initial face image into a luminance image.
[0189] In a possible implementation, the server determines the luminance values of multiple pixel points in the initial face image based on the color channel values of the initial face image in the three RGB color channels. The server generates a luminance image of the initial face image based on the luminance values of multiple pixel points in the initial face image.
[0190] In this implementation, the channel values of the three RGB color channels can be easily obtained, and the luminance is related to the three color channel values. The luminance image of the initial face image can be quickly generated based on the three RGB color channel values, and the efficiency of image processing is relatively high.
[0191] For example, for any pixel point on the initial face image, the server substitutes the color channel values of the pixel point in the three RGB color channels into the target relationship data to obtain the luminance value of the pixel point. The target relationship data is used to represent the conversion relationship between the three RGB color channel values and the luminance value. For example, the server obtains the luminance value of the pixel point based on the color channel values of the pixel point in the three RGB color channels through the following formula (6). The server generates a luminance image of the initial face image based on the luminance values of multiple pixel points in the initial face image.
[0192] Pl i = red * 0.299 + green * 0.587 + blue * 0.114 (6)
[0193] where Pl i is the luminance value of pixel point i, red is the red channel value of pixel point i, green is the green channel value of pixel point i, and blue is the blue channel value of pixel point i.
[0194] Part III: The server generates a target probability image of the initial face image based on the initial probability image and the luminance image.
[0195] In a possible implementation, the server determines the hair region on the initial face image based on the initial probability image. The server determines the average luminance value of the hair region based on the luminance image of the initial face image. The server binarizes the luminance image based on the average luminance value of the hair region to obtain a reference luminance image. The server performs dilation processing on the reference luminance image, and fuses the reference luminance image obtained after the dilation processing with the initial probability image to obtain a target probability image of the initial face image.
[0196] The dilation processing refers to filling 0 or 1 as the pixel value around the image to expand the size of the image.
[0197] In this embodiment, the server first determines the hair region based on the initial probability image, and then obtains a reference luminance image based on the average luminance value of the hair region. The reference luminance image is dilated to enlarge its size to that of the Hechi, and finally, the fused reference luminance image obtained after the dilation process and the initial probability image are used to obtain the final target probability image. This target probability image can reflect the positions of the high-brightness points in the hair region. This process utilizes processes such as average luminance value, binarization, and dilation, making the accuracy of the target category image relatively high.
[0198] To illustrate the above embodiment more clearly, the above embodiment will be divided into several parts for description below.
[0199] A. The server determines the hair region on the initial face image based on the initial probability image.
[0200] In a possible embodiment, the server binarizes the initial probability image based on a probability threshold to obtain a reference probability image of the initial face image. The server determines the hair region on the initial face image based on the reference probability image of the initial face image.
[0201] Among them, the probability threshold is set by the technical personnel according to the actual situation, such as set to 0.7, 0.8, 0.85, or 0.9, etc. The embodiments of the present application do not limit this. The pixel values of the pixel points in the initial probability image are probabilities. The purpose of binarizing the initial probability image is to unify the pixel values of the pixel points in the initial probability image into two values, so as to achieve the segmentation of the hair region.
[0202] For example, for any pixel in the initial probability image, when the probability corresponding to the pixel is greater than or equal to the probability threshold, the server adjusts the probability corresponding to the pixel to a first value. When the probability corresponding to the pixel is less than the probability threshold, the server adjusts the probability corresponding to the pixel to a second value, where the first value is greater than the second value. Herein, the probability corresponding to the pixel is also the pixel value of the pixel. In this implementation manner, the pixel values of the pixels in the initial probability image are all adjusted to the first value and the second value, realizing binarization. That the probability corresponding to a pixel is greater than the probability threshold indicates that the probability that the pixel belongs to the hair region is relatively high; that the probability corresponding to a pixel is less than the probability threshold indicates that the probability that the pixel belongs to the hair region is relatively low. After the server binarizes the pixel values of multiple pixels in the initial probability image based on the probability threshold, the reference probability image is obtained, and the values of the pixels on the reference probability image are the first value or the second value. In some embodiments, the first value is 1 and the second value is 0. In this case, the pixel values of the pixels in the reference probability image are 0 or 1. The server determines the hair region on the initial face image based on the pixel values of multiple pixels in the reference probability image. That is, the server determines the region enclosed by the pixels with the first value as the pixel value in the reference probability image as the hair region on the reference probability image. The server maps the hair region on the reference probability image to the initial face image to obtain the hair region on the initial face image.
[0203] B. The server determines the average brightness value of the hair region based on the brightness image of the initial face image.
[0204] In a possible implementation manner, the server obtains the brightness values of multiple pixels in the hair region from the brightness image of the initial face image. The server determines the average brightness value of the hair region based on the brightness values of multiple pixels in the hair region.
[0205] C. The server binarizes the brightness image based on the average brightness value of the hair region to obtain a reference brightness image.
[0206] Herein, the purpose of binarizing the brightness image is to unify the pixel values of the pixels in the brightness image into two values, so as to obtain the highlight points in the initial face image.
[0207] In a possible implementation manner, for any pixel in the brightness image, when the brightness value of the pixel is greater than or equal to the average brightness value, the server adjusts the brightness value of the pixel to a third value. When the brightness value of the pixel is less than the average brightness value, the server adjusts the brightness value of the pixel to a fourth value, where the third value is greater than the fourth value.
[0208] Among them, the pixel value of a pixel point in the luminance image is a luminance value. The purpose of binarizing the luminance image is to unify the pixel values of the pixel points in the luminance image into two values, so as to determine the positions of the high-brightness points.
[0209] For example, for any pixel point in the luminance image, when the luminance value corresponding to the pixel point is greater than or equal to the average luminance value, the server adjusts the luminance value corresponding to the pixel point to a third value. When the luminance value corresponding to the pixel point is less than the average luminance value, the server adjusts the luminance value corresponding to the pixel point to a fourth value. In some embodiments, the third value is greater than the fourth value. Among them, the luminance value corresponding to the pixel point is also the pixel value of the pixel point. In this implementation manner, the pixel values of the pixel points in the luminance image are all adjusted to the third value and the fourth value, realizing binarization. That the luminance value corresponding to a pixel point is greater than the average luminance value indicates that the luminance value of the pixel point is relatively high; that the luminance value corresponding to a pixel point is less than the average luminance value indicates that the luminance value of the pixel point is relatively low. After the server binarizes the pixel values of multiple pixel points in the luminance image based on the average luminance value, the reference luminance value image is obtained, and the values of the pixel points on the reference luminance value image are the third value or the fourth value. In some embodiments, the third value is 1 and the fourth value is 0. In this case, the pixel values of the pixel points in the reference luminance value image are 0 or 1.
[0210] D. The server fuses the reference luminance image obtained after the dilation process with the initial probability image to obtain the target probability image of the initial face image.
[0211] Among them, the reference luminance image can reflect the positions of the high-brightness points, the initial probability image can reflect the positions of the hair regions, and the target probability image obtained by fusing the reference probability image and the initial probability image can thus reflect the positions of the high-brightness points in the hair regions.
[0212] In a possible implementation manner, the server performs Gaussian filtering on the reference luminance image obtained after the dilation process based on the initial probability image to obtain the target luminance image of the initial face image. The server multiplies the values of the corresponding pixel points in the initial probability image and the target luminance image to obtain the target probability image of the initial face image.
[0213] For example, based on the initial probability image, the server performs Gaussian filtering on the hair region in the reference luminance image obtained after the dilation process to obtain the target luminance image of the initial face image. That is, the server determines the hair region in the reference luminance image obtained after the dilation process based on the initial probability image. The server performs convolution processing on the hair region using a Gaussian convolution kernel to obtain the target luminance image of the initial face image. In some embodiments, when taking neighborhood points within the Gaussian convolution sum for the current center point, if the probability value corresponding to the neighborhood point is 1, the neighborhood point is directly used; if the probability value corresponding to the neighborhood point is 0, the center point value is used to replace the neighborhood point. The server multiplies the values of the corresponding pixel points in the initial probability image and the target luminance image to obtain the target probability image of the initial face image.
[0214] Among them, the purpose of performing Gaussian filtering on the hair region in the reference luminance image is to avoid non-hair pixels from introducing image processing.
[0215] The following will combine Figure 4 to illustrate the above step S305.
[0216] Refer to Figure 4 , the server binarizes the initial probability image to obtain the reference probability image of the initial face image. The server determines the hair region on the initial face image based on the reference probability image. The initial probability image is also referred to as the hair segmentation probability map, and the hair region is also referred to as the hair segmentation result. The server obtains a luminance image based on the initial face image. The initial face image is also referred to as portrait data. The server performs Gaussian filtering on the hair region of the initial face image. The processing server determines the average luminance value of the hair region based on the luminance image. The server binarizes the luminance image based on the average luminance value to obtain a reference luminance image. The server performs dilation processing on the reference luminance image and fuses the reference luminance image obtained after the dilation process with the initial probability image to obtain a target probability image. The target probability image is also referred to as the hair highlight region probability map.
[0217] It should be noted that the above step S305 can be executed either after step S304 or at any position before steps S301 - S304. The embodiments of the present application do not make any limitations in this regard.
[0218] In step S306, the server obtains a first initial frequency band image, a second initial frequency band image, and a third initial frequency band image of the initial face image based on the luminance channel image, the first filtered image, the second filtered image, and the third filtered image. The first initial frequency band image is a difference image between the initial face image and the first filtered image. The second initial frequency band image is a difference image between the first filtered image and the second filtered image. The third initial frequency band image is a difference image between the second filtered image and the third filtered image.
[0219] Among them, the difference image between the initial face image and the first filtered image refers to an image obtained by subtracting the pixel values of the corresponding pixel points in the initial face image and the first filtered image. The corresponding pixel points refer to the pixel points at the same position in the image.
[0220] In a possible implementation, the server subtracts the first filtered image from the luminance channel image to obtain the first initial frequency band image. The server subtracts the second filtered image from the first filtered image to obtain the second initial frequency band image. The server subtracts the third filtered image from the second filtered image to obtain the third initial frequency band image.
[0221] Among them, the luminance channel image, the first filtered image, the second filtered image, and the third filtered image correspond to different frequencies.
[0222] In step S307, the server generates a target face image based on the first initial frequency band image, the second initial frequency band image, the third initial frequency band image, the target probability image of the initial face image, and the third filtered image. The target probability image is used to represent the positions of the high-brightness points in the hair region of the initial face image. The high-brightness points are pixel points whose luminance meets the preset luminance condition. The target face image is a face image obtained after removing the high-brightness points in the initial face image.
[0223] In a possible implementation, the server performs pixel value adjustment processing on the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image based on the target probability image to obtain a first target frequency band image, a second target frequency band image, and a third target frequency band image. The server performs fusion processing on the first target frequency band image, the second target frequency band image, the third target frequency band image, and the third filtered image to obtain a fusion image. The server performs color space transformation on the fusion image to obtain the target face image, and the target face image and the initial face image are in the same color space.
[0224] In this embodiment, the target face image is obtained by processing based on images of different frequency bands, which embodies the idea of frequency division and dimming. The obtained target face image can eliminate the high-brightness points in the initial face image while retaining the face edges.
[0225] To explain the above embodiment more clearly, the above embodiment will be described in several parts below.
[0226] First part: The server processes the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image based on the target probability image to obtain a first target frequency band image, a second target frequency band image, and a third target frequency band image.
[0227] In a possible embodiment, the server determines a plurality of target pixel points in the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image respectively based on the target probability image, and the target pixel points correspond to the high-brightness points in the luminance channel image. When the pixel values of the plurality of target pixel points meet the preset pixel value conditions, the server adjusts the pixel values of the plurality of target pixel points based on a first preset threshold and a second preset threshold to obtain the first target frequency band image, the second target frequency band image, and the third target frequency band image, and the second preset threshold is greater than the first preset threshold.
[0228] Among them, the server determines a plurality of target pixel points in the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image based on the target probability image, which means that the server determines a plurality of target pixel points in the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image respectively based on the target probability image.
[0229] In this embodiment, the server can adjust the pixel values of the target pixel points based on the first preset threshold and the second preset threshold, so as to obtain the first target frequency band image, the second target frequency band image, and the third target frequency band image, which is convenient for the subsequent image processing process of frequency division and dimming.
[0230] For example, the server determines a plurality of target pixel points in the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image respectively based on the pixel values of a plurality of pixel points in the target probability image, wherein the pixel values of the plurality of pixel points in the target probability image can represent the positions of the target pixel points. When the pixel values of the plurality of target pixel points are greater than a target value, the server adjusts the pixel values of the plurality of target pixel points based on a first preset threshold and a second preset threshold, where the pixel values of the plurality of target pixel points refer to the pixel values of the plurality of target pixel points in the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image respectively. In some embodiments, the target value is 0, that is, when the pixel values of the plurality of target pixel points are greater than 0, the server adjusts the pixel values of the plurality of target pixel points based on a first preset threshold and a second preset threshold.
[0231] The method for the server to adjust the pixel values of the plurality of target pixel points based on the first preset threshold and the second preset threshold will be described below.
[0232] In a possible implementation manner, for any target pixel point in the first initial frequency band image, when the pixel value of the target pixel point is less than or equal to the first preset threshold ti1, the server keeps the pixel value of the target pixel point unchanged. When the pixel value of the target pixel point is greater than or equal to the second preset threshold ti2, the server multiplies the pixel value of the target pixel point by a first parameter to obtain the adjusted pixel value of the target pixel point, where the first parameter is a positive number less than 1, and after multiplying the pixel value by the first parameter, the obtained value is less than the pixel value. When the pixel value of the target pixel point is greater than the first preset threshold ti1 and less than the second preset threshold ti2, the server performs a fusion process on the pixel value of the target pixel point, the first parameter, and a second parameter to obtain the adjusted pixel value of the target pixel point, where the second parameter is determined based on the first preset threshold, the second preset threshold, and the pixel value of the target pixel point.
[0233] In this implementation manner, methods for adjusting pixel values using the first preset threshold and the second preset threshold in multiple cases are provided, which can be flexibly selected according to the actual situation during the image processing process, improving the efficiency of image processing.
[0234] The method for the server to perform a fusion process on the pixel value of the target pixel point, the first parameter, and the second parameter to obtain the adjusted pixel value of the target pixel point will be described below.
[0235] In a possible implementation, the server processes the first preset threshold, the second preset threshold, and the pixel value of the target pixel point using a smooth step model to obtain the second parameter. The server adds the product of the first parameter, the pixel value of the target pixel point, and the second parameter, and the product of the third parameter and the pixel value of the target pixel point to obtain the adjusted pixel value of the target pixel point. The sum of the third parameter and the second parameter is 1.
[0236] In this implementation, the use of the smooth step model enables stepped processing, reflecting the idea of frequency division and segmentation, and improving the accuracy of the second parameter.
[0237] For example, the server determines the second parameter through the following formula (7) and obtains the adjusted pixel value of the target pixel point through the following formula (8).
[0238] α = smoothstep(ti1, ti2, p) (7)
[0239] p = α * p * λ i + (1 - α) * p (8)
[0240] Where α is the second parameter, λ i is the first parameter, λ i < 1, ti1 is the first preset threshold, ti2 is the second preset threshold, p is the pixel value of the target pixel point, and smoothstep is the function corresponding to the smooth step model.
[0241] Second part: The server performs fusion processing on the first target frequency band image, the second target frequency band image, the third target frequency band image, and the third filtered image to obtain a fused image.
[0242] In a possible implementation, the server adds the first target frequency band image B1’, the second target frequency band image B2’, the third target frequency band image B3’, and the third filtered image GF2 to obtain a fused image res. The fused image thus carries the information of the three frequency band images and the third filtered image.
[0243] Third part: The server performs color space transformation on the fused image to obtain the target face image.
[0244] Since the fused image is obtained by fusing the first target frequency band image, the second target frequency band image, the third target frequency band image, and the third filtered image, and the first target frequency band image, the second target frequency band image, the third target frequency band image, and the third filtered image are all processed based on the initial face image in the Iab color space, then the fused image is also in the Iab color space. The server transforms the fused image from Iab to the RGB color space to obtain the final target face image, which is the initial face image with high-brightness points eliminated, that is, the face image after eliminating the oiliness of the hair.
[0245] In some embodiments, the technical solution provided by the embodiments of the present application can be implemented by two modules. The first module is the oiliness detection module, which is used to execute the above step S305 to obtain the target probability image. The second module is the oiliness elimination module, which is used to execute the steps other than S305 in the above S301 - S307 to obtain the target face image.
[0246] Through the technical solution provided by the embodiments of the present application, the luminance channel image of the initial face image is filtered multiple times to obtain the first filtered image, the second filtered image, and the third filtered image. Based on the luminance channel image, the first filtered image, the second filtered image, and the third filtered image, three frequency band images are obtained. Subsequently, the three frequency band images are processed, and the high-brightness points in the initial face image are eliminated by using the idea of frequency division and darkening to obtain the target face image. Since these high-brightness points correspond to the oiliness of the hair in the initial face image, the oiliness in the hair area is eliminated after the high-brightness points are eliminated, thereby improving the display effect of the face image.
[0247] Figure 5 is a block diagram of an image processing device shown according to an exemplary embodiment. Refer to Figure 5 The device includes a first filtering unit 501, a second filtering unit 502, a frequency band image acquisition unit 503, and a target face image generation unit 504.
[0248] The first filtering unit 501 is configured to perform filtering on the luminance channel image of the initial face image to obtain the first filtered image of the initial face image, and the initial face image includes a hair area.
[0249] The second filtering unit 502 is configured to obtain the second filtered image and the third filtered image of the initial face image. The second filtered image is obtained by performing filtering on the luminance channel image, and the third filtered image is obtained by performing filtering on the second filtered image.
[0250] The frequency band image acquisition unit 503 is configured to obtain a first initial frequency band image, a second initial frequency band image, and a third initial frequency band image of the initial face image based on the luminance channel image, the first filtered image, the second filtered image, and the third filtered image. The first initial frequency band image is a difference image between the initial face image and the first filtered image. The second initial frequency band image is a difference image between the first filtered image and the second filtered image. The third initial frequency band image is a difference image between the second filtered image and the third filtered image.
[0251] The target face image generation unit 504 is configured to generate a target face image based on the first initial frequency band image, the second initial frequency band image, the third initial frequency band image, the target probability image of the initial face image, and the third filtered image. The target probability image is used to represent the positions of the high-brightness points within the hair region of the initial face image. The high-brightness points are pixel points whose luminance meets the preset luminance condition. The target face image is a face image obtained after eliminating the high-brightness points in the initial face image.
[0252] In a possible implementation manner, the target face image generation unit 504 is configured to perform pixel value adjustment processing on the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image based on the target probability image to obtain a first target frequency band image, a second target frequency band image, and a third target frequency band image. The first target frequency band image, the second target frequency band image, the third target frequency band image, and the third filtered image are subjected to fusion processing to obtain a fused image. The fused image is subjected to color space transformation to obtain the target face image, and the target face image and the initial face image are in the same color space.
[0253] In a possible implementation manner, the target face image generation unit 504 is configured to determine a plurality of target pixel points in the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image respectively based on the target probability image. The target pixel points correspond to the high-brightness points in the luminance channel image. When the pixel values of the plurality of target pixel points meet the preset pixel value condition, the pixel values of the plurality of target pixel points are adjusted based on a first preset threshold and a second preset threshold to obtain the first target frequency band image, the second target frequency band image, and the third target frequency band image, and the second preset threshold is greater than the first preset threshold.
[0254] In a possible implementation, the target face image generation unit 504 is configured to perform the following operations for any target pixel point in the first initial frequency band image: when the pixel value of the target pixel point is less than or equal to the first preset threshold, keep the pixel value of the target pixel point unchanged; when the pixel value of the target pixel point is greater than or equal to the second preset threshold, multiply the pixel value of the target pixel point by a first parameter to obtain the adjusted pixel value of the target pixel point, where the first parameter is a positive number less than 1; when the pixel value of the target pixel point is greater than the first preset threshold and less than the second preset threshold, perform a fusion process on the pixel value of the target pixel point, the first parameter, and a second parameter to obtain the adjusted pixel value of the target pixel point, where the second parameter is determined based on the first preset threshold, the second preset threshold, and the pixel value of the target pixel point.
[0255] In a possible implementation, the target face image generation unit 504 is configured to perform a process on the first preset threshold, the second preset threshold, and the pixel value of the target pixel point by using a smooth step model to obtain the second parameter. Add the product of the first parameter and the pixel value of the target pixel point, the product of the second parameter and the pixel value of the target pixel point, and the product of a third parameter and the pixel value of the target pixel point to obtain the adjusted pixel value of the target pixel point, where the sum of the third parameter and the second parameter is 1.
[0256] In a possible implementation, the first filtering unit 501 is configured to perform a sliding operation on the luminance channel image by using a Gaussian convolution kernel, and perform Gaussian filtering processing within the area covered by the Gaussian convolution kernel to obtain the first filtered image of the initial face image, where the horizontal step size when the Gaussian convolution kernel slides is the first step size, the vertical step size when the Gaussian convolution kernel slides is the second step size, the first step size is the ratio of the width of the face area in the initial face image to the width of the preset face area, and the second step size is the ratio of the height of the face area in the initial face image to the height of the preset face area.
[0257] In a possible implementation, the device further includes:
[0258] A face area determination unit, configured to perform key point detection on the initial face image to obtain multiple face key points in the initial face image. Determine the minimum bounding rectangle of the multiple face key points as the face area in the initial face image.
[0259] In a possible implementation, the second filtering unit 502 is configured to perform guided filtering on the luminance channel image with the luminance channel image as the guidance image to obtain the second filtered image. Then, perform guided filtering on the luminance channel image with the second filtered image as the guidance image to obtain the third filtered image. Among them, the horizontal step size and the vertical step size of the two guided filterings on the luminance channel image are the first step size and the second step size respectively.
[0260] In a possible implementation, the apparatus further includes:
[0261] A luminance channel image confirmation unit, configured to perform transforming the initial face image into the Lab color space. And confirm the image of the L channel in the Lab color space as the luminance channel image of the initial face image.
[0262] In a possible implementation, the apparatus further includes:
[0263] A target probability image generation unit, configured to perform image recognition on the initial face image to obtain an initial probability image of the initial face image, where the initial probability image is used to represent the probability that the pixel points in the initial face image correspond to hair. Convert the initial face image into a luminance image. And generate a target probability image of the initial face image based on the initial probability image and the luminance image.
[0264] In a possible implementation, the target probability image generation unit is configured to perform determining the luminance values of multiple pixel points in the initial face image based on the color channel values of the initial face image in the three RGB color channels. And generate a luminance image of the initial face image based on the luminance values of multiple pixel points in the initial face image.
[0265] In a possible implementation, the target probability image generation unit is configured to perform determining the hair region on the initial face image based on the initial probability image. Determine the average luminance value of the hair region based on the luminance image of the initial face image. Binarize the luminance image based on the average luminance value of the hair region to obtain a reference luminance image. Perform dilation processing on the reference luminance image, and fuse the reference luminance image obtained after the dilation processing with the initial probability image to obtain the target probability image of the initial face image.
[0266] In a possible implementation, the target probability image generation unit is configured to perform binarizing the initial probability image based on a probability threshold to obtain a reference probability image of the initial face image. And determine the hair region on the initial face image based on the reference probability image of the initial face image.
[0267] In a possible implementation, the target probability image generation unit is configured to perform, for any pixel point in the initial probability image, when the probability corresponding to the pixel point is greater than or equal to the probability threshold, adjusting the probability corresponding to the pixel point to a first value; when the probability corresponding to the pixel point is less than the probability threshold, adjusting the probability corresponding to the pixel point to a second value, where the first value is greater than the second value.
[0268] In a possible implementation, the target probability image generation unit is configured to perform, for any pixel point in the luminance image, when the luminance value of the pixel point is greater than or equal to the average luminance value, adjusting the luminance value of the pixel point to a third value; when the luminance value of the pixel point is less than the average luminance value, adjusting the luminance value of the pixel point to a fourth value, where the third value is greater than the fourth value.
[0269] In a possible implementation, the target probability image generation unit is configured to perform Gaussian filtering on the reference luminance image obtained after the dilation process based on the initial probability image to obtain the target luminance image of the initial face image; multiplying the values of the corresponding pixel points in the initial probability image and the target luminance image to obtain the target probability image of the initial face image.
[0270] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0271] Through the technical solution provided by the embodiments of the present application, the luminance channel image of the initial face image is subjected to multiple filtering processes to obtain a first filtered image, a second filtered image, and a third filtered image. Based on the luminance channel image, the first filtered image, the second filtered image, and the third filtered image, three frequency band images are obtained. Subsequently, the three frequency band images are processed, and the high-brightness points in the initial face image are eliminated using the idea of frequency division dimming to obtain the target face image. Since these high-brightness points correspond to the oiliness of the hair in the initial face image, the oiliness in the hair area is eliminated after eliminating the high-brightness points, thereby improving the display effect of the face image.
[0272] In the embodiments of the present application, the electronic device can be implemented as a terminal. The structure of the terminal will be described below:
[0273] Figure 6 It is a block diagram of a terminal shown according to an exemplary embodiment. The terminal 600 can be a terminal used by a user. The terminal 600 may also be referred to by other names such as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0274] Generally, the terminal 600 includes a processor 601 and a memory 602.
[0275] The processor 601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0276] The memory 602 may include one or more storage media, and the storage media may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices.
[0277] In some embodiments, the terminal 600 may further optionally include a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface 603 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 603 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 604, a display screen 605, a camera module 606, an audio circuit 607, a positioning module 608, and a power supply 6013.
[0278] The peripheral device interface 603 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.
[0279] The radio frequency circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 604 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 604 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 604 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, each generation of mobile communication network (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 604 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0280] The display screen 605 is used to display the UI (User Interface). The UI may include images, texts, icons, videos, and any combination thereof. When the display screen 605 is a touch display screen, the display screen 605 also has the ability to collect touch signals on or above the surface of the display screen 605. The touch signals can be input as control signals to the processor 601 for processing. At this time, the display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 605, which is provided on the front panel of the terminal 600; in other embodiments, there may be at least two display screens 605, which are respectively provided on different surfaces of the terminal 600 or are in a foldable design; in some embodiments, the display screen 605 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 600. Even, the display screen 605 can also be set as an irregular image that is not rectangular, that is, a special-shaped screen. The display screen 605 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0281] The camera module 606 is used to collect images or videos. Optionally, the camera module 606 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 606 may further include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. A two-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0282] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 601 for processing, or input to the radio frequency circuit 604 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 600. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 607 may further include a headphone jack.
[0283] The positioning component 608 is used to locate the current geographical location of the terminal 600 to achieve navigation or LBS (Location Based Service). The positioning component 608 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, the GLONASS system of Russia, or the Galileo system of the European Union.
[0284] The power supply 609 is used to supply power to each component in the terminal 600. The power supply 609 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 609 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0285] In some embodiments, the terminal 600 further includes one or more sensors 610. The one or more sensors 610 include but are not limited to: an acceleration sensor 611, a gyroscope sensor 612, a pressure sensor 613, a fingerprint sensor 614, an optical sensor 615, and a proximity sensor 616.
[0286] The acceleration sensor 611 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the terminal 600. For example, the acceleration sensor 611 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 601 can control the display screen 605 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 611. The acceleration sensor 611 can also be used for collecting game or user's motion data.
[0287] The gyroscope sensor 612 can detect the body orientation and rotation angle of the terminal 600. The gyroscope sensor 612 can cooperate with the acceleration sensor 611 to collect the 3D actions of the user on the terminal 600. Based on the data collected by the gyroscope sensor 612, the processor 601 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0288] The pressure sensor 613 can be disposed on the side frame of the terminal 600 and / or the lower layer of the display screen 605. When the pressure sensor 613 is disposed on the side frame of the terminal 600, it can detect the holding signal of the user on the terminal 600, and the processor 601 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 613. When the pressure sensor 613 is disposed on the lower layer of the display screen 605, the processor 601 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 605. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0289] The fingerprint sensor 614 is used to collect the fingerprints of the user. The processor 601 can identify the user's identity according to the fingerprints collected by the fingerprint sensor 614, or the fingerprint sensor 614 can identify the user's identity according to the collected fingerprints. When the identified user identity is a trusted identity, the processor 601 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 614 can be disposed on the front, back, or side of the terminal 600. When there are physical buttons or manufacturer logos on the terminal 600, the fingerprint sensor 614 can be integrated with the physical buttons or manufacturer logos.
[0290] The optical sensor 615 is used to collect the ambient light intensity. In one embodiment, the processor 601 can control the display brightness of the display screen 605 according to the ambient light intensity collected by the optical sensor 615. Specifically, when the ambient light intensity is high, the display brightness of the display screen 605 is increased; when the ambient light intensity is low, the display brightness of the display screen 605 is decreased. In another embodiment, the processor 601 can also dynamically adjust the shooting parameters of the camera module 606 according to the ambient light intensity collected by the optical sensor 615.
[0291] A proximity sensor 616, also known as a distance sensor, is typically disposed on the front panel of the terminal 600. The proximity sensor 616 is used to collect the distance between the user and the front of the terminal 600. In one embodiment, when the proximity sensor 616 detects that the distance between the user and the front of the terminal 600 is gradually decreasing, the processor 601 controls the display screen 605 to switch from the lit state to the off state; when the proximity sensor 616 detects that the distance between the user and the front of the terminal 600 is gradually increasing, the processor 601 controls the display screen 605 to switch from the off state to the lit state.
[0292] Those skilled in the art can understand that Figure 6 the structure shown in does not constitute a limitation on the terminal 600, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.
[0293] The above electronic device can also be implemented as a server. The structure of the server will be introduced below:
[0294] Figure 7 is a block diagram of a server provided by an embodiment of the present application. The server 700 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 701 and one or more memories 702. Among them, at least one computer program is stored in the one or more memories 702, and the at least one computer program is loaded and executed by the one or more processors 701 to implement the methods provided by the above various method embodiments. Of course, the server 700 may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server 700 may also include other components for implementing device functions, which will not be elaborated here.
[0295] In an exemplary embodiment, a non-volatile computer-readable storage medium including instructions is also provided, such as a memory including instructions. The above instructions can be executed by the processor 701 of the server 700 to complete the above image processing method, or by the processor 601 of the terminal 600 to complete the above image processing method. Optionally, the storage medium may be a non-temporary storage medium. For example, the non-temporary storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0296] In an exemplary embodiment, a computer program product including a computer program is also provided. The computer program can be executed by the processor of the electronic device to implement the above image processing method.
[0297] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only illustrative, and the true scope and spirit of the present application are pointed out by the following claims.
[0298] It should be understood that the present application is not limited to the exact structures described above and shown in the attached drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. An image processing method, characterized in that, comprising: filtering a luminance channel image of an initial face image to obtain a first filtered image of the initial face image, the initial face image including a hair region; acquiring a second filtered image and a third filtered image of the initial face image, the second filtered image being obtained by filtering the luminance channel image, and the third filtered image being obtained by filtering the second filtered image; based on the luminance channel image, the first filtered image, the second filtered image, and the third filtered image, acquiring a first initial frequency band image, a second initial frequency band image, and a third initial frequency band image of the initial face image, the first initial frequency band image being a difference image between the initial face image and the first filtered image, the second initial frequency band image being a difference image between the first filtered image and the second filtered image, and the third initial frequency band image being a difference image between the second filtered image and the third filtered image; generating a target face image based on the first initial frequency band image, the second initial frequency band image, the third initial frequency band image, a target probability image of the initial face image, and the third filtered image, the target probability image being used to represent positions of high-brightness points within the hair region of the initial face image, the high-brightness points being pixel points whose brightness meets a preset brightness condition, and the target face image being a face image obtained after eliminating the high-brightness points in the initial face image.
2. The image processing method according to claim 1, characterized in that, the generating a target face image based on the first initial frequency band image, the second initial frequency band image, the third initial frequency band image, a target probability image of the initial face image, and the third filtered image includes: performing pixel value adjustment processing on the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image based on the target probability image to obtain a first target frequency band image, a second target frequency band image, and a third target frequency band image; performing fusion processing on the first target frequency band image, the second target frequency band image, the third target frequency band image, and the third filtered image to obtain a fused image; performing color space transformation on the fused image to obtain the target face image, the target face image and the initial face image being in the same color space.
3. The image processing method according to claim 2, characterized in that, the performing pixel value adjustment processing on the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image based on the target probability image to obtain a first target frequency band image, a second target frequency band image, and a third target frequency band image includes: determining a plurality of target pixel points in the first initial frequency band image, the second initial frequency band image, and the third initial frequency band image respectively based on the target probability image, the target pixel points corresponding to the high-brightness points in the luminance channel image; When the pixel values of the multiple target pixel points meet the preset pixel value conditions, based on a first preset threshold and a second preset threshold, adjust the pixel values of the multiple target pixel points to obtain the first target frequency band image, the second target frequency band image, and the third target frequency band image, where the second preset threshold is greater than the first preset threshold.
4. The image processing method according to claim 3, wherein, the adjusting the pixel values of the multiple target pixel points based on the first preset threshold and the second preset threshold to obtain the first target frequency band image, the second target frequency band image, and the third target frequency band image includes: For any target pixel point in the first initial frequency band image, when the pixel value of the target pixel point is less than or equal to the first preset threshold, keep the pixel value of the target pixel point unchanged; when the pixel value of the target pixel point is greater than or equal to the second preset threshold, multiply the pixel value of the target pixel point by a first parameter to obtain the adjusted pixel value of the target pixel point, where the first parameter is a positive number less than 1; when the pixel value of the target pixel point is greater than the first preset threshold and less than the second preset threshold, perform a fusion process on the pixel value of the target pixel point, the first parameter, and a second parameter to obtain the adjusted pixel value of the target pixel point, where the second parameter is determined based on the first preset threshold, the second preset threshold, and the pixel value of the target pixel point.
5. The image processing method according to claim 4, wherein, the performing a fusion process on the pixel value of the target pixel point, the first parameter, and the second parameter to obtain the adjusted pixel value of the target pixel point includes: processing the first preset threshold, the second preset threshold, and the pixel value of the target pixel point using a smooth step model to obtain the second parameter; adding the product of the first parameter, the pixel value of the target pixel point, and the second parameter, and the product of a third parameter and the pixel value of the target pixel point to obtain the adjusted pixel value of the target pixel point, where the sum of the third parameter and the second parameter is 1.
6. The image processing method according to claim 1, wherein, the filtering the luminance channel image of the initial face image to obtain the first filtered image of the initial face image includes: sliding a Gaussian convolution kernel on the luminance channel image and performing Gaussian filtering within the area covered by the Gaussian convolution kernel to obtain the first filtered image of the initial face image, where the horizontal step size when the Gaussian convolution kernel slides is a first step size, the vertical step size when the Gaussian convolution kernel slides is a second step size, the first step size is the ratio of the width of the face area in the initial face image to the width of a preset face area, and the second step size is the ratio of the height of the face area in the initial face image to the height of the preset face area.
7. The image processing method according to claim 6, wherein, The face region in the initial face image is obtained by the following method: Perform key point detection on the initial face image to obtain multiple face key points in the initial face image; Determine the minimum bounding rectangle of the multiple face key points as the face region in the initial face image.
8. The image processing method according to claim 6, wherein, the obtaining of the second filtered image and the third filtered image of the initial face image includes: Taking the luminance channel image as a guidance image, performing guided filtering on the luminance channel image to obtain the second filtered image; Taking the second filtered image as a guidance image, performing guided filtering on the luminance channel image to obtain the third filtered image; wherein, the horizontal step size and the vertical step size of the two guided filterings of the luminance channel image are the first step size and the second step size respectively.
9. The image processing method according to claim 1, wherein, before performing filtering on the luminance channel image of the initial face image to obtain the first filtered image of the initial face image, the method further includes: Transforming the initial face image into the Lab color space; Identifying the image of the L channel in the Lab color space as the luminance channel image of the initial face image.
10. The image processing method according to claim 1, wherein, before generating the target face image based on the first initial frequency band image, the second initial frequency band image, the third initial frequency band image, the target probability image of the initial face image, and the third filtered image, the method further includes: Performing image recognition on the initial face image to obtain the initial probability image of the initial face image, where the initial probability image is used to represent the probability that the pixel points in the initial face image correspond to hair; Converting the initial face image into a luminance image; Generating the target probability image of the initial face image based on the initial probability image and the luminance image.
11. The image processing method according to claim 10, wherein, the converting the initial face image into a luminance image includes: Determining the luminance values of multiple pixel points in the initial face image based on the color channel values of the initial face image in the three RGB color channels; Generating the luminance image of the initial face image based on the luminance values of multiple pixel points in the initial face image.
12. The image processing method according to claim 10, wherein, the generating the target probability image of the initial face image based on the initial probability image and the luminance image includes: Determining the hair region on the initial face image based on the initial probability image; Determining the average luminance value of the hair region based on the luminance image of the initial face image; Performing binarization on the luminance image based on the average luminance value of the hair region to obtain a reference luminance image; Perform dilation processing on the reference luminance image, and fuse the reference luminance image obtained after the dilation processing with the initial probability image to obtain the target probability image of the initial face image.
13. The image processing method according to claim 12, wherein, determining the hair region on the initial face image based on the initial probability image includes: Performing binarization on the initial probability image based on a probability threshold to obtain a reference probability image of the initial face image; Determining the hair region on the initial face image based on the reference probability image of the initial face image.
14. The image processing method according to claim 13, wherein, performing binarization on the initial probability image based on a probability threshold includes: For any pixel point in the initial probability image, when the probability corresponding to the pixel point is greater than or equal to the probability threshold, adjusting the probability corresponding to the pixel point to a first value; when the probability corresponding to the pixel point is less than the probability threshold, adjusting the probability corresponding to the pixel point to a second value, and the first value is greater than the second value.
15. The image processing method according to claim 12, wherein, performing binarization on the luminance image based on the average luminance value of the hair region to obtain a reference luminance image includes: For any pixel point in the luminance image, when the luminance value of the pixel point is greater than or equal to the average luminance value, adjusting the luminance value of the pixel point to a third value; when the luminance value of the pixel point is less than the average luminance value, adjusting the luminance value of the pixel point to a fourth value, and the third value is greater than the fourth value.
16. The image processing method according to claim 12, wherein, fusing the reference luminance image obtained after the dilation processing with the initial probability image to obtain the target probability image of the initial face image includes: Performing Gaussian filtering on the reference luminance image obtained after the dilation processing based on the initial probability image to obtain the target luminance image of the initial face image; Multiplying the values of the corresponding pixel points in the initial probability image and the target luminance image to obtain the target probability image of the initial face image.
17. An image processing apparatus, wherein, comprising: A first filtering unit configured to perform filtering on the luminance channel image of the initial face image to obtain a first filtered image of the initial face image, and the initial face image includes a hair region; A second filtering unit configured to obtain a second filtered image and a third filtered image of the initial face image, the second filtered image is obtained by performing filtering on the luminance channel image, and the third filtered image is obtained by performing filtering on the second filtered image; A frequency band image acquisition unit, configured to execute acquiring a first initial frequency band image, a second initial frequency band image, and a third initial frequency band image of the initial face image based on the luminance channel image, the first filtered image, the second filtered image, and the third filtered image, where the first initial frequency band image is a difference image between the initial face image and the first filtered image, the second initial frequency band image is a difference image between the first filtered image and the second filtered image, and the third initial frequency band image is a difference image between the second filtered image and the third filtered image; A target face image generation unit, configured to execute generating a target face image based on the first initial frequency band image, the second initial frequency band image, the third initial frequency band image, a target probability image of the initial face image, and the third filtered image, where the target probability image is used to represent positions of high-brightness points within a hair region in the initial face image, the high-brightness points being pixel points whose luminance meets a preset luminance condition, and the target face image is a face image obtained after eliminating the high-brightness points in the initial face image.
18. An electronic device, characterized in that, it includes: a processor; a memory for storing program code executable by the processor; wherein, the processor is configured to execute the program code to implement the image processing method according to any one of claims 1 to 16.
19. A non-volatile storage medium, when program code in the storage medium is executed by a processor of an electronic device, enabling the electronic device to execute the image processing method according to any one of claims 1 to 16.
20. A computer program product, including a computer program, characterized in that, when the computer program is executed by a processor, it implements the image processing method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Image processing method and device, electronic equipment, storage medium and program product
CN110910309A
Image processing method, electronic device, and computer-readable medium
WO2022016326A1