A method for whitening human skin color and a device capable of processing images
Through deep learning of the skin segmentation model and the inverse distance weight compensation algorithm, the problem of inaccurate color adjustment in the skin area in the prior art is solved, and the precise whitening and natural whitening effects of the skin area are achieved.
Patent Information
- Application Number
- CN202411781985.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-12-05
AI Technical Summary
The existing whitening algorithms cannot accurately adjust the color of the skin area, resulting in a decrease in the naturalness of the portrait image, and ignore the whitening effect of areas other than the face, making it difficult to maintain the consistency of the whitening between the skin area and other areas.
The skin area is extracted using a deep learning-based skin segmentation model, the average value and distance weight of the RGB three-channel value are calculated, and the inverse distance weight is compensated to generate an automatic skin tone whitening image.
Accurate color adjustments to the skin area are achieved, natural effects are maintained, and excessive whitening is avoided, which improves the consistency and naturalness of the whitening effect of portrait images.
Smart Images

Figure CN119722526B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method for whitening the skin color of a portrait and a device capable of processing images. Background Art
[0002] With the rapid development of image processing and computer vision technologies, the whitening processing of portrait images has become a research direction that has received extensive attention.
[0003] Most of the existing whitening algorithms are based on globally adjusting the brightness and contrast of the image, and cannot perform precise color adjustment on the skin area, often resulting in a reduction in the naturalness of portrait images. Some other algorithms are based on regions, such as whitening the facial region, and similarly cannot perform precise color adjustment on the skin areas of other parts.
[0004] In addition, during the whitening process, the existing technology often ignores the whitening effect of regions other than the face. Even if some technologies whiten the above regions, it is difficult to maintain the consistency of whitening between the skin region and the regions other than the skin. Summary of the Invention
[0005] To alleviate or solve at least one aspect or at least one point of the above problems, the present invention is proposed. A method for whitening the skin color of a portrait according to the present invention is characterized in that it includes the following steps;
[0006] S1: Input or collect an original portrait image, and extract the skin region based on a pre-trained deep learning skin segmentation model;
[0007] S2: Calculate the average value of the RGB three-channel values of the pixel points within the skin region;
[0008] S3: Set a preset target value, and calculate the three-channel compensation value according to the average value of the RGB three-channel values;
[0009] S4: Traverse each pixel point of the original portrait image, and calculate the distance weight between the RGB three-channel values of each pixel point of the original portrait image and the average value of the RGB three-channel values according to the average value of the RGB three-channel values;
[0010] S5: Perform compensation according to the RGB three-channel compensation value and the distance weight to obtain a whitened image of the automatic skin color of the original portrait.
[0011] Preferably, step S4 includes: S41: Obtain the RGB three-channel values of each pixel point, and calculate the distance value between it and the average value of the RGB three-channels of the skin region;
[0012] S42: According to the calculated distance value, use the inverse distance method to calculate the weight of each channel of the RGB of the pixel, and perform normalization processing to generate a normalized distance weight. The normalized distance weight is calculated using the following formula:
[0013] Where represents the absolute difference between the R, G, and B channels of the pixel and the average value; represents a small constant used to avoid division by zero; represents using the inverse distance to obtain the weight of each channel of the pixel point; represents the normalized distance weight.
[0014] Preferably, in step S5, it includes: according to the normalized distance weight, increase the compensation values of the three channels to the RGB channel values of each pixel point according to the ratio of the distance weight to form a whitened image.
[0015] Preferably, in step S1, it further includes: step S11: collection and preprocessing of the portrait image dataset for training the skin segmentation model, and the portrait image includes a complete face area;
[0016] Step S12: construction of the skin segmentation model, and the skin segmentation model uses a lightweight semantic segmentation model based on a convolutional neural network;
[0017] Step S13: analyze and optimize the skin segmentation model based on a convolutional neural network;
[0018] Step S14: input the input or acquired original portrait image into the trained skin segmentation model to generate the required skin segmentation map and extract the skin area.
[0019] Preferably, in step S12, it further includes: extracting high-level semantic features from the input image, gradually reducing the spatial dimension of the image through a series of convolutional layers, pooling layers, and activation functions, and enhancing the semantic information in the image.
[0020] Preferably, in step S12, it further includes: during training, calculate the loss for the feature maps output in the previous several stages to improve the ability of the low-level convolution to extract image details.
[0021] Preferably, in step S12, it further includes: during inference, directly fuse the feature maps of the low-level convolution with the high-level convolution feature maps containing more semantic information. As shown in formula (1), its calculation process is:
[0022]
[0023] (1)
[0024] represents the detail loss function; represents the detail information of each pixel or region in the predicted distribution or feature map; represents the true label or target value; represents the Dice loss function; represents the binary cross-entropy loss; represents the predicted value at the i-th pixel position; the true label value at the i-th pixel position; d represents the detail layer; represents the resolution of the feature map; represents that the smoothing term prevents the denominator from being zero.
[0025] Preferably, step S13 further includes: using the sorted portrait image and the annotated image as training data, combining the learning rate and the batch size to accelerate model convergence, and adopting a dynamic learning rate adjustment strategy to adapt to the training stage; combining data augmentation techniques such as random cropping, rotation, and flipping to increase the diversity of training data and improve the generalization ability of the model.
[0026] Preferably, step S14 includes: inputting the original portrait image into the trained skin segmentation model, and the model performs forward inference on the image using a convolutional neural network structure, and extracts multi-scale features of the image through multiple convolutional layers to highlight the details of the edges and skin regions.
[0027] Preferably, step S14 includes: the skin segmentation model reduces the resolution of the feature map through downsampling to obtain more extensive context information, and then restores it to the resolution of the original image through upsampling to accurately locate the skin region.
[0028] Preferably, step S2 includes: step S21: According to a preset condition, used to accurately locate the pixels representing the skin region in the skin segmentation map; find the pixel coordinates that meet the condition according to the preset condition; according to these coordinates, create a boolean mask map to mark the skin region in the segmentation map;
[0029] Step S22: Convert the mask map to a grayscale map, and then use a conditional selection function to find the pixel positions in the original picture with grayscale values; in RGB mode, extract all pixel values within the skin region and calculate the average values of the R, G, and B channels of the pixels in the skin region.
[0030] The present invention also provides a device capable of processing images, which is provided with a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the above is implemented.
[0031] The present invention proposes an automatic skin color whitening method based on portrait images. An accurate color correction and compensation algorithm can adaptively adjust according to the local information of the image, so as to achieve directional optimization and whitening of the skin area, thereby generating an automatic skin color whitening image. By introducing the concept of distance weight, the color adjustment of the image can avoid over-whitening while maintaining a natural effect, making the portrait whitening effect more realistic and natural. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic flowchart of a portrait skin color whitening processing method according to an exemplary embodiment of the present invention.
[0033] Figure 2 It is a schematic flowchart of establishing and training a deep learning skin segmentation model according to an exemplary embodiment of the present invention.
[0034] Figure 3 It is a schematic flowchart of calculating the average value of the RGB three-channel values of the pixel points within the skin area according to an exemplary embodiment of the present invention.
[0035] Figure 4 It is a schematic flowchart of calculating the distance weight between the RGB three-channel values of each pixel point in the original portrait image and the average value of the RGB three-channel values according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The following description of the embodiments of the present invention with reference to the accompanying drawings is intended to explain the overall inventive concept of the present invention and should not be construed as a limitation of the present invention. In the present invention, the same reference numerals represent the same or similar components.
[0037] The features described herein can be implemented in different forms and should not be construed as limited to the examples described herein. On the contrary, the examples provided herein are only intended to illustrate some of the many possible ways of implementing the methods, devices, and / or systems described herein, which will be apparent after understanding the disclosure of the present invention.
[0038] The terms used herein are only for describing various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising", "including", and "having" specify the presence of the recited features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0039] To enable those skilled in the art to use the content of the present invention, the following exemplary embodiments may be given in combination with specific application scenarios, parameters of specific systems, devices, and components, and specific connection methods in the following text. However, for those skilled in the art, these embodiments are only examples, and the general principles defined herein can be applied to other embodiments and application scenarios without departing from the spirit and scope of the present invention.
[0040] According to an exemplary embodiment of the present invention: As Figures 1-4 shown, a method for whitening the skin color of a portrait. As Figure 1 shown, it includes the following steps: Step S1: Input or collect a portrait image; based on a pre-trained deep learning skin segmentation model, extract the skin area. The above skin segmentation model is a computer vision model used to automatically detect and extract the skin area from a portrait image. The above model can be trained with a large number of images with skin annotations so that the model can identify the pixel areas belonging to the skin in the image.
[0041] Exemplarily, in combination with Figure 2 illustrate the method for segmenting the skin area of a portrait image based on a skin segmentation model in an exemplary embodiment of the present invention. Step S11: First is the collection of the portrait image dataset and the preprocessing of the dataset. The portrait image contains a complete face area. Use image standard tools on the portrait image to annotate the background, glasses, eyebrows, hair, skin, clothes, and mouth areas with the colors specified for the corresponding areas. This step ensures that the model can learn the features of the skin area and saves the annotation results. Process the annotated image, set the annotated image to the specified pixel value according to the corresponding color, and the finally generated mask image will be used to train the model. Organize the original portrait image and the corresponding annotated image into a training dataset. To evaluate the performance of the model, divide the dataset into a training set and a validation set, which can be divided according to a ratio of 8:2 to ensure that the training set contains enough samples to improve the generalization ability of the model.
[0042] Step S12: Construction of the skin segmentation model. According to the annotation accuracy and consistency of the training dataset, optionally, the skin segmentation model is trained using real-time semantic segmentation technology to accurately extract the skin area. By adopting a real-time semantic segmentation model, real-time processing can be achieved while ensuring accuracy. During the training process, the model learns based on a large amount of annotated data, enabling it to accurately identify and segment the boundaries of the skin area and adapt to different individuals and various postures. This design not only improves the efficiency of segmentation but also ensures the fluency and reliability of the model in practical applications. Real-time semantic segmentation technology is a computer vision technology aimed at classifying each pixel in an image while maintaining a sufficient processing speed for real-time applications. Semantic segmentation refers to dividing each pixel in an image into a certain class of objects and generating a corresponding segmented image. Real-time means that the segmentation processing speed of the model is fast enough to respond in milliseconds.
[0043] The real-time semantic segmentation model is a lightweight semantic segmentation model based on a convolutional neural network, which achieves real-time segmentation through streamlining the network structure and innovative module combinations. The structure of this lightweight semantic segmentation model includes an encoder, a decoder, and a segmentation head. The lightweight semantic segmentation model adopts a simple pyramid pooling module to aggregate global context information at a lower computational cost, thereby improving the model's perception ability of the environmental background and further enhancing the segmentation performance. At the same time, a flexible lightweight decoder is used to reduce the computational burden of the traditional decoder and improve the inference efficiency. In addition, the lightweight semantic segmentation model also integrates a unified attention fusion module as a key component for enhancing feature representation. This module uses spatial and channel attention mechanisms to generate weights and fuse the input features with these weights to help the model more effectively capture key semantic information.
[0044] Exemplarily: The encoder uses the backbone network as the basic backbone network to extract high-level semantic features from the input image. Through a series of convolutional layers, pooling layers, and activation functions, the spatial dimensions [H, W] of the image are gradually reduced, where H and W represent height and width information respectively, and the semantic information in the image is enhanced, that is, the number of channels of the feature map increases. The task of the encoder is to convert the original image into a low-resolution feature map with rich semantic information, including information such as the objects, textures, and structures in the image. The backbone network is a short-term densely connected network, which is a lightweight convolutional neural network designed to ensure the effectiveness of feature extraction while improving the model inference speed. The backbone network of the short-term densely connected network is divided into 6 stages, where the first 5 stages are mainly used as the backbone for feature extraction and segmentation. The size of the output feature map of each stage gradually decreases, while the number of channels gradually increases. The size of the output feature map of the i-th stage is of the original input H×W. The number of channels of the feature map increases with the increase of the stage, and the number of channels of the i-th stage is 。The number of channels gradually expands in each stage to enhance the network's ability to represent features. When the number of channels of the output feature map in the second stage is greater than 64, a "final convolutional layer" is added to the output of the fifth stage. The role of this layer is to ensure that the number of channels of the output feature map in the fifth stage is not less than 1024 for better subsequent processing. A short dense connection module is used inside each stage, and the core design of this module is to achieve more efficient information transfer through dense connections. The receptive field of the low-level convolution in the module is small, so more channels are needed to extract detailed information. While the receptive field of the high-level convolution is large, and only a few channels are required to extract sufficient semantic information. A detailed guidance training branch is added. Preferably, during training, the loss is calculated for the feature maps output in the previous stages to enhance the ability of the low-level convolution to extract image details. During inference, the feature maps of the low-level convolution are directly fused with the high-level convolution feature maps containing more semantic information. As shown in formula (1) representing its calculation process:
[0045]
[0046] (1)
[0047] represents the detail loss function; represents the detailed information of each pixel or region in the predicted distribution or feature map; represents the true label or target value; represents the Dice loss function; represents the binary cross-entropy loss; represents the predicted value at the i-th pixel position; the true label value at the i-th pixel position; d represents the detail layer; represents the resolution of the feature map; represents the smoothing term to prevent the denominator from being zero.
[0048] Exemplarily, the decoder is the corresponding part to the encoder, responsible for restoring the low-resolution feature map generated by the encoder to the resolution of the original image and performing pixel-level classification. The decoder usually includes an upsampling layer and a convolutional layer, which are used to gradually restore the details of the feature map and generate a feature map with the same size as the input image. The decoder maps the semantic features extracted by the encoder to the pixel level for classifying each pixel.
[0049] Exemplarily, the segmentation head is a fully-connected layer or a convolutional layer, which is used to convert the feature map output by the decoder into the final segmentation structure. The simple pyramid pooling module generates feature maps of various sizes by performing pooling operations on the feature map at different scales. The simple pyramid pooling module uses two different scales of pooling operations, namely average pooling and max pooling. These pooled feature maps are then fused to provide rich context information and enhance the robustness of the network to changes in the shape and size of the target. The flexible lightweight decoder is a module for generating high-quality output feature maps, aiming to accelerate the inference of the model by reducing the amount of computation and the number of parameters, while being able to adapt to various input conditions and task requirements. The main role of the flexible lightweight decoder is to gradually increase the spatial size [H, W] of the feature map and gradually reduce the number of channels of the feature during the operation of the decoder, balancing the computational complexity of the encoder and the decoder and making the overall model more efficient. Exemplarily, the model also includes a unified attention fusion module, which is a structure for fusing multiple feature maps. Using spatial and channel attention mechanisms, it generates a weight matrix , and multiplies the input feature by the weight matrix to obtain the fused feature. This enables the model to better capture the key semantic information in the image and improve the segmentation accuracy. Formula (2) represents its calculation process:
[0050]
[0051] (2)
[0052]
[0053] represents the output of the deep module; represents the corresponding part of the encoder; represents the upsampled feature; represents the attention weight coefficient; represents the upsampling operation on the high-resolution feature map ; The reverse attention weight represents the importance weight of the low-resolution feature map during the fusion process, Represents the attention formula. The attention mechanism is a dynamic selection process for important information in the image input, and this process is achieved by adaptive weights for features. When using attention to process tasks, the importance of different information is reflected by the weights. The distinction in the attention of the attention mechanism to different information is reflected in the weight assignment. The attention mechanism can be regarded as a multi-layer perceptron composed of a query matrix, keys, and weighted averages. The spatial attention mechanism generates a weight using the spatial relationship between pixels, and this weight represents the importance of each pixel in the input feature. The channel attention mechanism generates weights using the relationship between channels, and this weight indicates the importance of each channel in the input feature. The channel attention module uses average pooling and max pooling operations to compress the spatial dimension of the input feature.
[0054] Step S13: This method is a method for analyzing and optimizing a skin segmentation model based on a convolutional neural network. This method evaluates the size and computational efficiency of the overall model by calculating the number of parameters and floating-point operations per layer. In addition, the sorted portrait images and annotated pictures are used as training data, combined with optimized hyperparameters such as the learning rate and batch size to accelerate model convergence, and a dynamic learning rate adjustment strategy is adopted to adapt to the training stage. Combining data augmentation techniques such as random cropping, rotation, and flipping increases the diversity of the training data and improves the generalization ability of the model. Through the above method of analyzing convolutional network parameters, the present invention can effectively optimize the performance of the skin segmentation model in skin segmentation tasks, reduce the computational complexity, improve the segmentation accuracy, and thus achieve real-time and efficient skin segmentation.
[0055] Step S14: Input the portrait image into the trained skin segmentation model. The model uses the convolutional neural network structure to perform forward inference on the image, extracts multi-scale features of the image through multiple convolutional layers, and highlights the details of the edges and skin regions. The skin segmentation model reduces the resolution of the feature map through downsampling to obtain more extensive context information, and then restores it to the resolution of the original image through upsampling to accurately locate the skin region. After a series of convolutional, downsampling, and upsampling operations, the model outputs a segmentation map with the same size as the input image, where the value of each pixel represents the probability that the position belongs to the skin region. Finally, by setting a threshold, the probability map is converted into a binary segmentation map to generate the required skin segmentation map.
[0056] Step S2: Calculate the average values of the RGB three channels within the skin region. Specifically: Step S21: Define a preset condition for accurately locating the pixels representing the skin region in the skin segmentation map. Exemplarily, the above preset condition is usually achieved by comparing the color channel values of each pixel. For example, knowing that the skin region is represented by a specific color value in the segmentation map, such a condition can be set to filter the pixels in the image. Use a conditional selection function to find the pixel coordinates that meet the previously defined condition according to the condition. This function will return the indices of all pixels that meet the condition. These indices will provide us with the specific positions of the skin region in the image. Based on these coordinates, create a boolean mask to mark the skin region in the segmentation map. Initialize a zero-filled array of the same size as the segmentation map, and use the previously found pixel coordinates to set the values at the corresponding positions to 255 to mark the skin region. The above conditional selection function is a function used to select or find array elements according to conditions.
[0057] Step S22: Detect the pixel values that meet the conditions in the mask. By converting the mask image to a grayscale image, it is convenient for detection. Then find the pixel positions in the original image corresponding to the grayscale values. Among them, the above pixel positions can be found using a conditional selection function. In the RGB mode, we extract all the pixel values within the skin region and calculate the average values of the R, G, and B channels of the skin region pixels.
[0058] Step S3: Set a preset target value and calculate the RGB three-channel compensation values based on the average values of the RGB three channels. Exemplarily, the above preset target value can be selected from the better skins in the manual screening as the standard.
[0059] Step S31: First, set the target value of the standard skin color after whitening. This target value represents the ideal skin color characteristics. Step S32: Obtain the R-channel compensation value, G-channel compensation value, and B-channel compensation value according to the differences between the average values of the R channel, G channel, and B channel of the skin region and the set target value. For each channel, calculate the difference between the target value and the current channel average value to obtain the R-channel compensation value, G-channel compensation value, and B-channel compensation value.
[0060] By calculating the multiple relationship between the current skin region color value and the target skin color, the adjustment intensity is controlled. After obtaining the compensation values, in order to achieve a more adaptive whitening effect, it is also necessary to calculate the multiple relationship between the current skin region color value and the target skin color, in order to understand the deviation degree of the current skin region from the target skin color, so as to control the adjustment intensity. Appropriately scaling the compensation values through the multiple relationship can avoid problems such as excessive whitening or uneven whitening of the skin color.
[0061] Step S4: Traverse each pixel of the original portrait image, and calculate the distance weights of the RGB three channels from the average value for each pixel according to the average value of the RGB three channels.
[0062] Step S41: When traversing each pixel of the original portrait image, first obtain the RGB value of each pixel and calculate the distance between it and the average value of the RGB three channels in the skin area. For each pixel in the image, obtain its RGB value and calculate the absolute difference between the RGB value of this pixel and the average value of the skin area to get the distance of the three channels. The calculation process is shown in formula (3):
[0063] (3)
[0064] , , represents the absolute difference between the R, G, and B three channels of this pixel and the average value; The values corresponding to R, G, and B of this pixel point; represents the average value of the R, G, and B three channels of the indicated skin area;
[0065] Step S42: Generate weights for each channel of the RGB of this pixel according to the calculated distance value of the pixel point. Use the inverse distance method to calculate the weights of each channel, and use normalized weights to ensure that the smaller the distance, the larger the weight, and use the sum of the inverse distances to be 1 so that the influence can be reasonably distributed in the subsequent compensation calculation. Obtain the distance weights of the RGB three channels of this pixel point from the average value. The inverse distance is often used in image processing to calculate the weighting coefficient, and is used in this example to measure the distance weight of each pixel from the average value, so as to assign a higher weight to the pixels with a closer distance; the normalized weight is a method of adjusting a group of weights, normalizing their sum to 1. This can ensure that the relative ratio of each weight remains consistent, but the value is in the range of 0 to 1. The normalized weights are often used to calculate the weighted average or balance the contributions of multiple influencing factors. When using the inverse distance to calculate weights in this example, normalization can ensure that the sum of the weights of all channels is 1 to avoid a single channel having too much influence on the adjustment effect; the calculation process is shown in formula (4):
[0066] (4)
[0067] Where represents the absolute difference between the R, G, and B three channels of this pixel and the average value; represents a small constant used to avoid division by zero. Exemplarily, it can take a value of 0.1 - 0.9, and preferably 0.9; represents obtaining the weight of each channel of this pixel point using the inverse distance; represents the normalized weight;
[0068] Step S5: Perform compensation according to the distance weights of the RGB three-channel compensation values and the average value to obtain an automatically skin-color whitened image of the final original image portrait.
[0069] Step S51: Perform weighted adjustment according to the distance weights between the compensation values of each channel of each pixel and the average value obtained. According to the distance weights from the average value, add the compensation values of each channel to the original pixel values in proportion. This way of distributing compensation values according to distance weights can effectively adjust the skin-color whitening effect, making the skin-color adjustment more delicate and uniform. The whitening adjustment of each pixel is based on its distance weight from the average value, which helps to achieve delicate progressive whitening. To prevent the compensated pixel values from exceeding the image display range, ensure that the final value of each channel is limited between 0 and 255, avoiding overflows or undercuts. This limitation ensures the integrity of the generated image in the color space, avoiding situations of over-saturation or incomplete colors. Combine the compensated R, G, and B channel values of each pixel to generate the processed image. Since the concept of distance weights is introduced in the adjustment process, the overall whitening effect is more visually uniform and natural, without obvious hue mutations. The distance weight is a weighting factor used to measure the influence degree of each pixel relative to the reference value during the adjustment process. The magnitude of the distance weight is usually determined based on the distance between the pixel value and the reference value. The smaller the distance, the larger the weight, indicating that the pixel is closer to the target color value and requires more compensation or adjustment; the larger the distance, the smaller the weight, indicating that the pixel has a larger difference from the target value and may require less adjustment; in addition, this weighted compensation method not only optimizes the whitening effect but also improves the operation efficiency and speeds up the processing speed. The finally output image will present a portrait image processed by automatic skin-color whitening, showing a more uniform and natural skin-color effect.
[0070] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that these embodiments can be changed and element combinations can be made without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for whitening the skin color of a portrait, characterized in that: It includes the following steps; S1: Input or collect the original portrait image, and extract the skin area based on a pre-trained deep learning skin segmentation model; S2: Calculate the average value of the RGB three-channel values of the pixel points within the skin area; S3: Set a preset target value, and calculate the three-channel compensation value according to the average value of the RGB three-channel values; S4: Traverse each pixel point of the original portrait image, and calculate the distance weight between the RGB three-channel value of each pixel point in the original portrait image and the average value of the RGB three-channel values according to the average value of the RGB three-channel values; S5: Perform compensation according to the RGB three-channel compensation value and the distance weight to obtain the whitened image of the automatic skin color of the original portrait image; Step S4 includes: S41: Obtain the RGB three-channel value of each pixel point, and calculate the distance value between it and the average value of the RGB three-channel values of the skin area; S42: According to the calculated distance value, use the inverse distance method to calculate the weight of each channel of the RGB of the pixel, and perform normalization processing to generate a normalized distance weight. The normalized distance weight is calculated using the following formula: ; Among them represents the absolute difference between the R, G, and B channels of the pixel and the average value; represents a small constant used to avoid division by zero; represents using the inverse distance to obtain the weight of each channel of the pixel point; represents the normalized distance weight.
2. The processing method according to claim 1, characterized in that: Step S5 includes: According to the normalized distance weight, increase the compensation value of the three channels to the RGB three-channel value of each pixel point according to the proportion of the distance weight to form the whitened image.
3. The processing method according to claim 1, wherein: Step S1 also includes: Step S11: Collection and preprocessing of the portrait image dataset for training the skin segmentation model, and the portrait image includes the complete face area; Step S12: Construction of the skin segmentation model, and the skin segmentation model adopts a lightweight semantic segmentation model based on a convolutional neural network; Step S13: Analyze and optimize the skin segmentation model based on a convolutional neural network; Step S14: Input the input or collected original portrait image into the trained skin segmentation model to generate the required skin segmentation map and extract the skin area.
4. The processing method according to claim 3, wherein: Step S12 also includes: Extract high-level semantic features from the input image, gradually reduce the spatial dimension of the image through a series of convolutional layers, pooling layers and activation functions, and enhance the semantic information in the image.
5. The processing method according to claim 3, wherein: Step S13 also includes: Using the sorted portrait image and the annotated picture as training data, combining the learning rate and batch size to accelerate the model convergence, adopting a dynamic learning rate adjustment strategy to adapt to the training stage; combining data augmentation techniques such as random cropping, rotation and flipping to increase the diversity of training data and improve the generalization ability of the model.
6. The processing method according to claim 3, wherein: Step S14 includes: Input the original portrait image into the trained skin segmentation model, and the model performs forward inference on the image using the convolutional neural network structure, and extracts multi-scale features of the image through multiple convolutional layers to highlight the details of the edges and skin area.
7. The processing method according to claim 6, characterized in that: Step S14 includes: The skin segmentation model reduces the resolution of the feature map through downsampling to obtain more extensive context information, and then restores it to the resolution of the original image through upsampling to accurately locate the skin area.
8. The processing method according to claim 1, wherein: Step S2 includes: Step S21: According to a preset condition, accurately locate the pixels representing the skin area in the skin segmentation map; find the pixel coordinates that meet the condition according to the preset condition; according to these coordinates, create a boolean mask map to mark the skin area in the segmentation map; Step S22: Convert the mask image into a grayscale image, and then use a conditional selection function to find the pixel positions in the original image where the grayscale values are located; in the RGB mode, extract all the pixel values within the skin region and calculate the average values of the R, G, and B channels of the pixels in the skin region.
9. An image - processing device, which is internally provided with a readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed, it implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method for synthesizing hair style, device for synthesizing hair style, and program for synthesizing hair style
JP2013226286A
Apparatus and method for adjusting white balance
US20130128073A1