Face fusion preprocessing method and system based on hair repairing and lightweight skin beautifying
By combining generative adversarial networks and lightweight student networks, the problem of collaborative control between hair occlusion repair and beautification modules was solved, achieving coordination and consistency between hair repair and lightweight skin beautification, ensuring the realism and naturalness of the image, and improving the efficiency and robustness of the processing flow.
Patent Information
- Application Number
- CN202511586966.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-11-03
AI Technical Summary
In existing technologies for repairing hair-covered areas and facial skin beautification, the lack of coordinated control between the repair and beautification modules leads to damage to the hairline boundary structure. The beautification process is prone to introducing unnatural distortions and skin color deviations. Furthermore, the lack of adaptive risk assessment and dynamic adjustment affects the realism and robustness of the image.
We employ generative adversarial networks (GANs) to repair hair-occluded areas, combined with lightweight student networks for skin beautification, and use multi-dimensional indicators to assess risks, dynamically adjust the beautification effect, optimize network training using weighted models and knowledge distillation techniques, and construct a closed-loop mechanism for risk assessment and dynamic intervention.
It effectively prevents artifact superposition, ensures that the hairline area and skin area are consistent, and the skin color is within the physiologically acceptable range, thereby enhancing the realism and naturalness of the image and improving the efficiency and robustness of the processing workflow.
Smart Images

Figure CN121032822A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of face image preprocessing and computer vision, in particular to a face fusion preprocessing method and system based on hair repair and light skin beautification. BACKGROUND
[0002] The present technical solution relates to the image preprocessing link before face fusion or virtual dressing and other applications, and the core goal is to enhance the aesthetic quality of the image while ensuring the authenticity, naturalness and cultural suitability of the output image; in the processing flow of hair occlusion area repair and face skin beautification, the prior art often connects the two as isolated steps, and this processing mechanism has the following key technical defects and challenges:
[0003] Lack of collaborative control between repair and beautification modules: light skin beautification directly acts on the repaired image, which can easily cause secondary damage to the previously accurately reconstructed hairline boundary structure, such as blurring or distortion, resulting in superimposed artifacts, and it is difficult to ensure the coordination between the two modules; lack of quantitative control of excessive beautification risk: if there is no restriction on beautification, it is easy to introduce unnatural distortion, which is manifested as excessive loss of skin texture details, causing unnatural smoothing of the skin area, and excessive deviation of skin color, which exceeds the natural range acceptable by human physiology, causing potential unrealistic or cultural discomfort; lack of adaptive safety degradation logic: existing processing solutions use fixed parameters, which are difficult to dynamically and finely adjust the beautification effect intensity according to the potential repair and beautification risk of the image; once the risk is too high, the quality and naturalness of the processing result will decrease significantly, affecting the robustness of the system;
[0004] Therefore, how to build a risk assessment and dynamic intervention closed-loop mechanism containing multiple indicators, actively identify and quantify the risk of unnatural smoothing, boundary structure distortion and skin color deviation that may be introduced by beautification processing, and dynamically adjust the beautification effect intensity based on the risk measurement, so as to realize efficient aesthetic improvement while ensuring that the output image is both beautiful and authentic, has become a technical problem that needs to be solved. SUMMARY
[0005] To solve the above technical problems, the present application provides a face fusion preprocessing method and system based on hair repair and light skin beautification, specifically, the technical solution of the present application is as follows:
[0006] A face fusion preprocessing method based on hair repair and light skin beautification, comprising:
[0007] Repairing the hair occlusion area of the original input image to obtain a repaired image;
[0008] Light skin beautification of the repaired image to obtain a beautified image;
[0009] determine a skin region unnatural smoothness index based on the beautified image and a preset reference variance;
[0010] determine a hair repair boundary structure distortion index based on a hairline boundary region of the repaired image and the beautified image;
[0011] determine a skin color non-physiological deviation index based on a color histogram of the beautified image and a preset natural skin color space reference probability distribution;
[0012] combine the skin region unnatural smoothness index, the hair repair boundary structure distortion index, and the skin color non-physiological deviation index, and calculate a risk measure through a preset weighting model;
[0013] determine a dynamic fusion coefficient based on the risk measure, a preset warning threshold, and a preset critical threshold;
[0014] weight and fuse the repaired image and the beautified image according to the dynamic fusion coefficient to generate a final image.
[0015] Preferably, determining the dynamic fusion coefficient comprises:
[0016] when the risk measure is less than the preset warning threshold, setting the dynamic fusion coefficient to a preset maximum value;
[0017] when the risk measure is greater than or equal to the preset warning threshold and less than or equal to the preset critical threshold, causing the dynamic fusion coefficient to linearly decrease from the preset maximum value to zero based on the risk measure;
[0018] when the risk measure is greater than the preset critical threshold, setting the dynamic fusion coefficient to zero.
[0019] Preferably, the hair occlusion region repair processing is implemented by using a preset generative adversarial network.
[0020] wherein a composite loss function is formed by combining pixel loss and edge loss to optimize and constrain the generative adversarial network.
[0021] Preferably, the lightweight skin beautification is implemented by using a preset lightweight student network.
[0022] wherein a preset teacher network is used to guide the training of the lightweight student network through a knowledge distillation technique.
[0023] Preferably, the knowledge distillation technique uses soft label loss and feature map alignment loss for constraint.
[0024] Preferably, determining the skin region unnatural smoothness index comprises:
[0025] calculating the texture variance of the skin region in the beautified image and comparing it with a preset reference variance.
[0026] A face fusion preprocessing system based on hair repair and lightweight skin beautification, comprising:
[0027] A hair occlusion area repair module is configured to perform hair occlusion area repair processing on the original input image to obtain a repaired image.
[0028] A lightweight skin beautification module is configured to perform lightweight skin beautification on the repaired image to obtain a beautified image.
[0029] A smoothness index determination module is configured to determine a skin area unnatural smoothness index based on the beautified image and a preset reference variance.
[0030] A distortion index determination module is configured to determine a hair repair boundary structure distortion index based on a hairline boundary area of the repaired image and the beautified image.
[0031] An offset index determination module is configured to determine a skin color non-physiological offset index based on a color histogram of the beautified image and a preset natural skin color space reference probability distribution.
[0032] A risk metric calculation module is configured to calculate a risk metric by a preset weighting model in combination with the skin area unnatural smoothness index, the hair repair boundary structure distortion index, and the skin color non-physiological offset index.
[0033] A coefficient generation module is configured to determine a dynamic fusion coefficient based on the risk metric, a preset warning threshold, and a preset critical threshold.
[0034] An image synthesis module is configured to perform weighted fusion on the repaired image and the beautified image according to the dynamic fusion coefficient to generate a final image.
[0035] Preferably, the coefficient generation module is configured to:
[0036] When the risk metric is less than the preset warning threshold, the dynamic fusion coefficient is set to a preset maximum value.
[0037] When the risk metric is greater than or equal to the preset warning threshold and less than or equal to the preset critical threshold, the dynamic fusion coefficient is linearly decreased from the preset maximum value to zero based on the risk metric.
[0038] When the risk metric is greater than the preset critical threshold, the dynamic fusion coefficient is set to zero.
[0039] Compared with the prior art, the present application has the following advantages:
[0040] 1. By introducing the hair repair boundary structure distortion index, the method solves the technical problem of lack of collaborative control between repair and beautification modules; the index can actively evaluate the degree of secondary damage caused by the skin beautification process to the previously accurately reconstructed hairline boundary structure; based on the evaluation result, the system can dynamically adjust the beautification effect intensity, effectively prevent artifact superposition, ensure that the repaired hairline area and the beautified skin area are consistent, and greatly improve the realism and naturalness of the output image;
[0041] 2. The method constructs a multi-dimensional risk evaluation mechanism including a skin area unnatural smoothness index and a skin color non-physiological deviation index; by calculating the three indexes into a unified risk measure, and based on the measure, a segmented dynamic fusion coefficient is designed; this adaptive intervention loop can gradually weaken the beautification effect according to the increase of the risk measure; it avoids the one-size-fits-all processing of fixed parameters, maximizes the aesthetic effect under the premise of controllable risk, and decisively disables the beautification when the risk is too high, ensuring the bottom line safety and robustness of the output result;
[0042] 3. By introducing the skin color non-physiological deviation index and calculating it based on the natural skin color space reference probability distribution, the method can effectively ensure that the beautified skin color is still within the natural range acceptable by human physiology; this mechanism actively avoids the unrealistic or strange color tone that may be caused by beautification, thereby eliminating the risk of potential cultural discomfort caused by the image; this is not only a precise technical control, but also improves the universality and user acceptance of the product at the application level;
[0043] 4. In specific implementation, the scheme uses a generative adversarial network for high-quality hair occlusion area repair to ensure the fine depiction of key structures such as hairline; at the same time, the lightweight skin beautification uses a lightweight student network based on knowledge distillation technology, which greatly reduces the computational load while achieving similar skin beautification effect as complex models; these optimizations significantly improve the efficiency, accuracy and deployment advantages on devices with limited computing resources of the entire preprocessing process. BRIEF DESCRIPTION OF DRAWINGS
[0044] The application will be further explained in conjunction with the accompanying drawings and embodiments:
[0045] Figure 1 is a flowchart of the method of the application.
[0046] Figure 2 is a structural diagram of the system of the application. DETAILED DESCRIPTION
[0047] To make the purpose, technical scheme and advantages of the application clearer, the application will be further described in detail below in conjunction with specific embodiments.
[0048] Embodiment 1:
[0049] Please refer to Figure 1 A face fusion preprocessing method based on hair repair and lightweight beautification, comprising:
[0050] Repairing the hair occlusion area of the original input image to obtain a repaired image;
[0051] Lightweight beautification is performed on the repaired image to obtain a beautified image;
[0052] Based on the beautified image and a preset reference variance, a skin area unnatural smoothness index is determined;
[0053] Based on the hairline boundary area of the repaired image and the beautified image, a hair repair boundary structure distortion index is determined;
[0054] Based on the color histogram of the beautified image and a preset natural skin color space reference probability distribution, a skin color non-physiological deviation index is determined;
[0055] Combining the skin area unnatural smoothness index, the hair repair boundary structure distortion index and the skin color non-physiological deviation index, a risk measure is calculated through a preset weighting model;
[0056] Based on the risk measure, a preset warning threshold and a preset critical threshold, a dynamic fusion coefficient is determined;
[0057] According to the dynamic fusion coefficient, the repaired image and the beautified image are weighted and fused to generate a final image.
[0058] The embodiment provides a face fusion preprocessing method based on hair repair and lightweight beautification; the method aims to automatically process the input original image before face fusion or virtual dressing and the like, repair the information loss caused by hair occlusion, and moderately improve the aesthetic quality of the image, and the core lies in introducing a risk assessment and dynamic intervention mechanism, so that the finally generated image is not only beautiful, but also has authenticity and cultural suitability, and unnatural or offensive artifacts are avoided;
[0059] In a specific implementation scenario, the method comprises the following steps:
[0060] The original input image as a processing object is subjected to hair occlusion area repair processing to obtain a repaired image ; the purpose of this step is to accurately reconstruct the face area, such as the forehead, which is occluded by hair, especially bangs; the original hair occlusion area in the repaired image is repaired to present a complete and natural face structure;
[0061] Based on the repaired image lightweight beautification is performed to obtain a beautified image ; the purpose of this step is to perform lightweight aesthetic enhancement on the repaired image, such as smoothing the skin and evening the skin tone, without introducing obvious distortion; the beautified image has better visual presentation while maintaining the identity features of the person;
[0062] The system enters the risk assessment stage, which quantifies the distortion risk that may be introduced by beautification processing by calculating three key indicators;
[0063] Based on the beautified image and the preset reference variance, a skin area unnatural smoothness indicator is determined ; the physical meaning of this indicator is to quantify whether the lightweight skin beautification has caused the skin texture details to be lost and appear false; its calculation is based on comparing the texture variance of the skin area in the beautified image with a baseline value;
[0064] Based on the hairline boundary area of the repaired image and the beautified image, a hair repair boundary structure distortion indicator is determined ; the physical meaning of this indicator is to evaluate whether the skin beautification process has caused secondary damage, such as blurring or distortion, to the previously repaired hairline area, in order to evaluate the coordination between the repair and beautification modules and prevent the superposition of artifacts;
[0065] Based on the color histogram of the beautified image and the preset natural skin color space reference probability distribution, a skin color non-physiological deviation indicator is determined ; the physical meaning of this indicator is to ensure that the skin color after beautification is still within the natural range acceptable to human physiology, avoiding unrealistic or strange color tones, in order to control the color authenticity of the generated image;
[0066] The specific calculation method of the hair repair boundary structure distortion indicator is as follows:
[0067] The hairline boundary area is determined; through a face key point detection algorithm, such as a key point detection network based on deep learning, the key point position of the hairline is located, and an area with a width of pixels is extracted upward along the hairline contour and an area with a width of pixels is extracted downward, and the corresponding boundary areas from the repaired image and the beautified image are extracted, denoted as and ;
[0068] Edge extraction is performed on the boundary areas; Canny edge detection algorithm or Sobel operator is used to extract the edges of and Edge detection is performed to obtain an edge map And In this embodiment, the Canny algorithm is used, with a low threshold of 50 and a high threshold of 150.
[0069] Calculate the edge structure similarity; use the structural similarity index SSIM to evaluate the structural consistency between And The calculation formula is:
[0070] ;
[0071] Wherein, And are the mean values of And , respectively, And are the standard deviations thereof, is the covariance, And are the stability constants, in this embodiment , ;
[0072] After that, calculate the distortion index; the hair repair boundary structure distortion index Is defined as:
[0073] ;
[0074] The value range of is [0, 1], wherein 0 represents that the boundary structure is completely consistent and has no distortion; 1 represents that the boundary structure is completely different and the distortion is serious; when Approaches 0, it indicates that the skin treatment does not damage the boundary structure of the hairline; when Increases, it indicates that the skin treatment causes secondary damage such as blurring and distortion to the hairline;
[0075] Specific calculation method of the skin color non-physiological deviation index Is as follows:
[0076] Color space conversion and histogram extraction; convert the beautified image From the RGB color space to the LAB color space, which is more consistent with human visual perception characteristics; extract the two-dimensional joint histogram of the a channel (red-green chroma) and the b channel (yellow-blue chroma) of the skin area, denoted as , and the resolution of the histogram is set to 32x32 bins;
[0077] Construct a natural skin color space reference probability distribution; the reference distribution Based on the statistics of 100,000 natural face images randomly selected from the CelebA public face dataset; Specifically, for each image, the skin region is segmented, the a and b channel values in the LAB space are extracted, the probability distribution in 32x32 bins is counted, and finally the distribution of 100,000 images is averaged and normalized to obtain The reference distribution covers a wide range of natural skin colors of different races, genders, and ages, and has wide representativeness.
[0078] Calculate the distribution deviation; use KL divergence to quantify the skin color distribution of the beautified image Relative to the natural reference distribution The calculation formula is:
[0079] ;
[0080] Among them, represents the KL divergence, which is used to measure the deviation of the skin color distribution of the beautified image from the natural skin color reference distribution ; and are the bins indexes of the a channel and the b channel respectively, and are the probability values corresponding to the bins; In order to avoid numerical instability in logarithmic operation, when or is 0, this term does not participate in summation.
[0081] Normalized as distortion index; In order to make the value range of convenient for threshold setting, the following normalization formula is used:
[0082] ;
[0083] Among them, is the normalization coefficient, whose value is the 99th percentile value of the KL divergence obtained by statistics in the verification dataset, and in this embodiment ; After normalization, the value range is about [0, 1], where 0 means that the skin color distribution is completely consistent with the natural distribution, and the larger the value, the more serious the deviation; When exceeds 1, it indicates that the skin color has deviated significantly from the human physiological acceptable range.
[0084] In order to form a unified risk judgment, combined with the skin region unnatural smoothness index , the hair repair boundary structure distortion index and the skin color non-physiological deviation index , through a pre-set weighting model, the risk measure ; risk metric is a non-dimensional scalar that comprehensively reflects the beautification image overall distortion risk, whose value range is [0, +∞), the greater the value, the higher the risk; this single scalar provides a clear decision basis for subsequent dynamic intervention;
[0085] The preset weighting model adopts a linear weighted sum form, and the specific calculation formula is:
[0086] ;
[0087] Among them, , and are the weight coefficients of the three indexes, which are non-negative real numbers, and their sources are obtained by multiple linear regression training on a training set containing 1000 artificially annotated images; during the training process, each image is subjectively scored by 5 expert reviewers from 0 to 10 for its distortion risk, and the average value is taken as the supervision label, and the optimal weight coefficient is obtained by least squares fitting; in this embodiment, the typical value of the weight coefficient determined by the above training process is , , ; the relative size of the weight coefficient reflects the degree of influence of different types of distortion on user perception, among which is larger, indicating that the hairline boundary structure distortion has the most significant impact on overall visual quality; in other embodiments, the values of the weight coefficients can be adjusted according to the preferences of specific application scenarios and target user groups, for example, for applications that pay more attention to skin color authenticity, the value of can be appropriately increased.
[0088] Based on the calculated risk metric , the preset warning threshold and the preset critical threshold , the dynamic fusion coefficient is determined; the purpose of this step is to generate an adjustment coefficient for controlling the intensity of beautification effect according to the comprehensive risk score calculated in the previous step; the dynamic fusion coefficient is a floating-point number whose value range is between 0 and a preset maximum value; the preset warning threshold and the preset critical threshold are two key risk dividing points to divide the safe zone, the warning zone and the danger zone; the determination method is as follows:
[0089] A calibration data set is constructed; 500 face images that have been repaired and beautified are selected, covering different risk metrics Level, ranging from 0.1 to 2.0; 50 test users were invited to rate the naturalness and acceptability of each image. The rating criteria were: 9-10 points indicate completely natural and acceptable (safe zone), 6-8 points indicate minor flaws but acceptable (warning zone), and 0-5 points indicate obvious distortion and unacceptable (danger zone).
[0090] Statistical analysis was used to determine the threshold; the average subjective score was calculated for each image, and the subjective scores were plotted against the risk metric. A scatter plot; analysis revealed that when At that time, 95% of the images scored above 9 points, placing them in the safe zone; when At that time, the image score gradually decreased and entered the warning zone; when At that time, 80% of the images scored below 6 points, placing them in the danger zone;
[0091] Based on the above statistical analysis, in this embodiment, a preset warning threshold is used. Set to 0.3, preset critical threshold The threshold is set to 0.7; in other embodiments, the threshold can be adjusted according to the quality requirements of the target application. For example, for application scenarios with more stringent requirements, the threshold can be... Reduce to 0.2, Reduce to 0.5; for scenarios where a higher level of enhancement can be tolerated, it can be... Increase to 0.4, Increased to 0.9.
[0092] To generate the final result, based on the dynamic fusion coefficient For image restoration and beautify images Weighted fusion is performed to generate the final image. The purpose of this step is to adaptively apply enhancement effects to the repaired image based on the results of the risk assessment; final image By safely restoring images And potentially risky beautified images The fusion ratio is obtained by linear interpolation and is determined by the dynamic fusion coefficient. Precise control; final image The specific calculation formula is as follows:
[0093] ;
[0094] in, This is the dynamic fusion coefficient, and its value range is [0, ...]. ];when hour, This means the output is a completely retouched image, without any enhancement effects. This carries the lowest risk but offers zero aesthetic improvement. When the beautification effect reaches the preset maximum intensity, the aesthetic improvement is maximum, but the controllable risk needs to be ensured. When the output image is between the repaired image and the beautified image, the balance between aesthetic improvement and risk control is achieved; the weighted fusion strategy ensures smooth transition of the final image at the pixel level, avoiding the effect mutation caused by directly using the beautified image or the repaired image.
[0095] The present application realizes a complete processing-evaluation-intervention closed loop by connecting hair repair, lightweight skin beautification and a risk assessment mechanism containing multi-dimensional indicators; compared with the prior art which processes repair and beautification as isolated steps, the present application can actively identify and quantify the risks such as unnatural smoothing, boundary structure distortion and skin color deviation possibly introduced in the beautification process, and dynamically adjust the intensity of the beautification effect based on the risk measurement; this not only solves the reconstruction problem of the hair-shielded area, realizes efficient aesthetic improvement, but more importantly, through the introduction of safety degradation logic, the authenticity and naturalness of the output image are ensured, the artifacts and potential cultural discomfort caused by over-processing are effectively avoided, and the robustness and user acceptance of the face fusion application are significantly improved.
[0096] Embodiment 2:
[0097] The dynamic fusion coefficient is determined, including: when the risk measurement is less than a preset warning threshold, the dynamic fusion coefficient is set to a preset maximum value;
[0098] When the risk measurement is greater than or equal to the preset warning threshold and less than or equal to a preset critical threshold, the dynamic fusion coefficient is linearly decreased from the preset maximum value to zero based on the risk measurement;
[0099] When the risk measurement is greater than the preset critical threshold, the dynamic fusion coefficient is set to zero.
[0100] This embodiment is a specific implementation of the step of determining the dynamic fusion coefficient in embodiment 1; the core purpose of this design is to establish a clear, smooth and rapid response risk-effect conversion function, i.e. safety degradation logic, to ensure that the system can make reasonable beautification intensity adjustment according to the size of the risk measurement; ;
[0101] This process is defined by a piecewise function:
[0102] When the risk measurement is less than the preset warning threshold , the system determines that the quality of the beautified image is very high and the distortion risk is extremely low, which is in the safety zone, at this time the dynamic fusion coefficient is set to the preset maximum value ; is a scalar hyper-parameter of a business preset, which defines the maximum beautification intensity allowed by the system, and its source is set according to the aesthetic requirements of the product and the preferences of the target user group, and its value range is constrained in [0, 1]; this is intended to provide the best visual enhancement effect under controllable risk;
[0103] When the risk metric is greater than or equal to the preset warning threshold and less than or equal to the preset critical threshold , the system determines that the beautified image begins to show perceptible but still acceptable defects, in the warning zone; for smooth transition and to avoid sudden changes in effect, the dynamic blending coefficient adopts a linear decay strategy, and its calculation formula is:
[0104] ;
[0105] In this formula, is the risk metric calculated by the previous step, which is a dimensionless scalar; is the preset warning threshold, which is a dimensionless scalar, and its source is determined by manually calibrating the risk level; is the preset critical threshold, which is a dimensionless scalar, and its source is also determined by manually calibrating the risk level; is the maximum value preset according to business needs, which is a dimensionless scalar;
[0106] The applicable range of this formula is ; within this range, increases from to , monotonically decreases from to 0; to ensure numerical stability, a boundary constraint can be added in actual implementation: , which ensures that is always in the range [0, ], avoiding abnormal values caused by numerical errors;
[0107] When the risk metric is greater than the preset critical threshold , the system determines that the beautified image has entered the danger zone and may cause user discomfort due to severe distortion or artifacts; at this time, the system takes the most conservative strategy and forces the dynamic blending coefficient to be set to 0;
[0108] This segmented dynamic coefficient determination method provides a refined and adaptive safety control mechanism for the system; it avoids the one-size-fits-all processing logic and realizes smooth and gradual intervention on the beautification effect by setting up a safety zone, a warning zone and a danger zone; when the image quality is excellent, the aesthetic effect can be maximized; when risks occur, the effect can be reduced in proportion rather than being directly turned off; when the risk is too high, the beautification can be disabled decisively to ensure the bottom-line safety of the output result; this graceful degradation capability greatly improves the stability and user experience of the system.
[0109] Embodiment 3:
[0110] The hair occlusion area repair processing is implemented by using a preset-based generative adversarial network;
[0111] Among them, the pixel loss and the edge loss are combined to form a composite loss function to optimize and constrain the generative adversarial network.
[0112] This embodiment is a specific implementation of the hair occlusion area repair processing step in Embodiment 1; the purpose is to use a deep learning model to generate the forehead and other areas occluded by hair with high fidelity, so that they seamlessly connect with other parts of the image in terms of texture, lighting and structure;
[0113] In this embodiment, the processing is implemented by using a preset-based generative adversarial network (GAN); the generative adversarial network in this embodiment is specifically a generator based on the U-Net architecture and a corresponding discriminator, which functions to generate highly realistic image content through adversarial training between the generator and the discriminator;
[0114] The specific structure of the generator is as follows: a standard U-Net architecture is adopted, including 4 layers of encoders and 4 layers of decoders; each layer of the encoder includes two convolutional layers (convolution kernel size of 3x3, step size of 1, padding of 1) and a maximum pooling layer (step size of 2), with channel numbers of 64, 128, 256 and 512 respectively; the decoder adopts a symmetrical structure, each layer including an up-sampling layer (step size of 2) and two convolutional layers, and the feature maps of the corresponding layers of the encoder and the decoder are spliced through a jump connection; the input of the generator is the original image containing occlusion and its corresponding binary occlusion mask, and the output is the repaired image;
[0115] The specific structure of the discriminator is as follows: a PatchGAN architecture is adopted, including 5 convolutional layers, with a convolution kernel size of 4x4 and step sizes alternately set to 2 and 1, and channel numbers of 64, 128, 256, 512 and 1 respectively; the discriminator discriminates the true and false of the local area of the image with a 70x70 receptive field, and finally outputs an NxN discrimination map, with each element representing the true and false probability of the corresponding area; compared with the global discriminator, this design can more effectively preserve the texture details of the repaired area;
[0116] Training strategy: the standard adversarial training method is adopted, and the generator and the discriminator are updated alternately; the learning rate is set to 0.0002, and the Adam optimizer is used, , the number of training iterations is 50000, and the batchsize is 16; during the training process, the model performance is evaluated on the validation set every 1000 iterations, and the model with the lowest validation loss is selected as the final model.
[0117] In order to effectively optimize and constrain the generated results of the generative adversarial network, not only the content is correct, but also the edge is clear, the embodiment combines pixel loss and edge loss to form a composite loss function; the innovation of this design lies in that it realizes that separate pixel-level matching is easy to cause blur, and structural constraints must be introduced to ensure the authenticity of details;
[0118] The calculation method of the composite loss function is as follows:
[0119] ;
[0120] In the formula, is the total loss of the repair network, which is used as a scalar to guide network training; is the pixel loss, which physically ensures the overall color and brightness consistency of the repair area, and its source is the commonly used L1 norm in image-to-image translation tasks, and the calculation method is where is the real reference image without occlusion; is the edge loss, which physically forces the network to generate a structurally real hairline edge, and its source is the structural similarity measure based on image gradient, and the calculation method is where is a preset Laplacian operator used to extract high-frequency edge information of the image; and are scalar weight coefficients, which are obtained by performing multiple experiments on a standard validation set to select the weight combination that best combines the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) indicators of the repaired image;
[0121] In the embodiment, in order to ensure that the pixel loss and the edge loss are comparable in numerical scale, the two loss terms are normalized respectively; specifically, the total number of pixels of the image is normalized, normalization; after normalization, the numerical ranges of the two loss terms are both about [0, 255], so they can be directly combined by weight coefficients and linear combination; the values of the weight coefficients reflect the relative importance of pixel consistency and edge sharpness, in this embodiment, , This configuration enables the network to moderately emphasize the structural sharpness of the hairline while ensuring overall color consistency.
[0122] By using a generative adversarial network and designing a composite loss function that combines pixel loss and edge loss, this embodiment can achieve high-quality hair occlusion repair; pixel loss ensures the macro-consistency of the generated content, while edge loss focuses on the fine depiction of key structures such as the hairline, and the combination of the two effectively overcomes the defects of traditional methods that are prone to blur and distortion artifacts, significantly improving the realism and sharpness of the repaired area.
[0123] Embodiment 4:
[0124] Lightweight skin beautification is achieved by using a pre-set lightweight student network.
[0125] Among them, the knowledge distillation technology is used to guide the training of the lightweight student network by using a pre-set teacher network.
[0126] The knowledge distillation technology uses soft label loss and feature map alignment loss for constraint.
[0127] This embodiment is a specific implementation of the lightweight skin beautification step in embodiment 1, and details the core training technology; the purpose of this design is to ensure strong skin beautification effect while significantly reducing the computational complexity of the model, so that it can efficiently run on devices with limited computing resources.
[0128] The lightweight skin beautification is achieved by using a pre-set lightweight student network; the lightweight student network is a neural network with relatively simple model structure, fewer parameters, and fast calculation speed, which serves as the model for performing skin beautification tasks in actual deployment.
[0129] To make this lightweight network have comparable performance to complex models, this embodiment uses knowledge distillation technology to guide the training of the lightweight student network using a pre-set teacher network; knowledge distillation technology is a model compression method, and its core idea is to let a powerful but complex teacher network transfer its learned knowledge to a lightweight student network; the teacher network is a pre-trained large neural network with top-notch skin beautification performance.
[0130] To achieve deeper and more effective knowledge transfer, the knowledge distillation technology of the embodiment utilizes soft label loss and feature map alignment loss for constraint; this dual constraint design not only requires the student to imitate the final answer of the teacher, but also imitates the feature extraction logic of the teacher;
[0131] The total distillation loss function of the technology For In the formula, is a total distillation loss scalar used to guide the training of the student network; is a scalar weight hyperparameter, which is obtained by external experimental tuning to balance the importance of final output alignment and intermediate feature alignment; is a soft label loss, which is used to make the student network learn the complete probability distribution output by the teacher network, rather than just the final prediction result, in the embodiment, it is obtained by fitting the KL divergence of the probability distribution output by the teacher network , wherein and are the final outputs of the student and teacher networks, respectively; is a feature map alignment loss, which is used to promote the student network to learn the feature expression ability of the corresponding layer of the teacher network in the intermediate layer, in the embodiment, it is obtained by using the L2 norm calculation , wherein and are the feature maps extracted by the student and teacher networks in the preset intermediate layer , respectively is a set containing multiple intermediate layer indexes, which is obtained by selecting the network layer containing rich texture and structure information after hierarchical analysis of the teacher network;
[0132] The present application realizes an efficient and high-quality lightweight skin beautifying method; by using the knowledge distillation technology, especially by the dual constraint of soft label loss and feature map alignment loss, the lightweight student network can deeply simulate and inherit the powerful ability of the complex teacher network; while obtaining similar skin beautifying effect as the top large model, the computational complexity and memory occupation are significantly reduced, so that the advanced skin beautifying function can be smoothly run on mobile devices and other equipment, which has very high practical value and deployment advantage.
[0133] Embodiment 5:
[0134] Determine the non-natural smoothness index of the skin region, including:
[0135] Calculate the texture variance of the skin region in the beautified image, and compare it with the preset reference variance.
[0136] The embodiment is a specific implementation of the step of determining the unnatural smoothness index of the skin region in Embodiment 1; the purpose is to provide an objective and quantitative method to detect whether light beautification is excessive, thereby causing the skin to lose its natural texture;
[0137] In this embodiment, the determination process of the index includes calculating the texture variance of the skin region in the beautified image and comparing it with the preset reference variance; the technical principle behind this design is that the natural skin surface has subtle texture changes, and such changes will be reflected as the variance of pixel values in the high-pass filtered image; and excessive light beautification will flatten these details, resulting in a significant reduction in variance;
[0138] The specific calculation formula is ; in the formula, the skin region unnatural smoothness index is a dimensionless scalar, and the larger the value, the more unnatural the smoothness; the preset reference variance has the physical meaning of providing a natural skin texture benchmark, and its source is the average variance value obtained by statistically calculating the variance of the skin region after the same high-pass filtering from a large and diverse natural human face skin texture database, which ensures the universality and objectivity of the benchmark; the variance calculation operation is used to calculate the pixel value variance of the input image region; the high-pass filter is, for example, a Laplacian operator, which is used to extract high-frequency texture details in the image; the skin region in the beautified image is automatically extracted from the beautified image by face key point positioning or semantic segmentation technology;
[0139] The method based on texture variance comparison used in this embodiment provides a no-reference and high-sensitivity technical means for evaluating skin smoothness; it innovatively compares the texture features of the processed image with a natural texture benchmark statistically obtained from large-scale real data, thereby accurately quantifying the degree of detail loss caused by excessive beautification; compared with traditional methods that rely on subjective judgment or reference images, this method has higher objectivity and automation level, and provides reliable data support for accurate decision-making of the risk control model.
[0140] Embodiment 6:
[0141] Please refer to Figure 2 , the hair occlusion region repair module is used for hair occlusion region repair processing of the original input image to obtain a repaired image;
[0142] The light beautification module is used for light beautification of the repaired image to obtain a beautified image;
[0143] a smoothness index determination module configured to determine a non-natural smoothness index of the skin region based on the beautified image and a preset reference variance;
[0144] a distortion index determination module configured to determine a hair repair boundary structure distortion index based on the hairline boundary region of the repair image and the beautified image;
[0145] a deviation index determination module configured to determine a non-physiological deviation index of the skin color based on a color histogram of the beautified image and a preset natural skin color space reference probability distribution;
[0146] a risk measurement calculation module configured to calculate a risk measurement by a preset weighting model in combination with the non-natural smoothness index of the skin region, the hair repair boundary structure distortion index, and the non-physiological deviation index of the skin color;
[0147] a coefficient generation module configured to determine a dynamic fusion coefficient based on the risk measurement, a preset warning threshold, and a preset critical threshold;
[0148] an image synthesis module configured to perform weighted fusion on the repair image and the beautified image according to the dynamic fusion coefficient to generate a final image.
[0149] The embodiment provides a face fusion preprocessing system based on hair repair and lightweight skin beautification, which is designed to execute the method in any of the foregoing embodiments; the system aims to provide an integrated and automated solution to realize high-quality and safe face image preprocessing;
[0150] The system comprises the following modules that cooperate with each other: a hair occlusion region repair module configured to perform hair occlusion region repair processing on an original input image to obtain a repair image; a lightweight skin beautification module configured to perform lightweight skin beautification on the repair image to obtain a beautified image; a smoothness index determination module configured to determine a non-natural smoothness index of the skin region based on the beautified image and a preset reference variance; a distortion index determination module configured to determine a hair repair boundary structure distortion index based on the hairline boundary region of the repair image and the beautified image; a deviation index determination module configured to determine a non-physiological deviation index of the skin color based on a color histogram of the beautified image and a preset natural skin color space reference probability distribution; a risk measurement calculation module configured to calculate a risk measurement by a preset weighting model in combination with the foregoing three indexes; a coefficient generation module configured to determine a dynamic fusion coefficient based on the risk measurement, a preset warning threshold, and a preset critical threshold; and an image synthesis module configured to perform weighted fusion on the repair image and the beautified image according to the dynamic fusion coefficient to generate a final image.
[0151] The system implements the complete process of the method in Example 1 through a clear modular design; each module has a clear responsibility and interface, forming an automated pipeline from image input, repair, beautification, multi-dimensional risk assessment to adaptive image synthesis; this system architecture not only ensures the feasibility of the technical solution, but also makes it easy to maintain, upgrade and expand, laying a solid system foundation for providing stable, efficient and reliable face image preprocessing services.
[0152] Example 7:
[0153] The coefficient generation module is configured as follows:
[0154] When the risk metric is less than the preset warning threshold, the dynamic fusion coefficient is set to the preset maximum value;
[0155] When the risk metric is greater than or equal to the preset warning threshold and less than or equal to the preset critical threshold, the dynamic fusion coefficient is linearly reduced from the preset maximum value to zero based on the risk metric.
[0156] When the risk metric exceeds the preset critical threshold, the dynamic fusion coefficient is set to zero.
[0157] The specific configuration method of the coefficient generation module in the system; the purpose of this configuration is to solidify the security degradation logic described in Example 2 into the internal working mechanism of the module, so that it can automatically and accurately execute risk response;
[0158] In this embodiment, the coefficient generation module is configured to operate according to the following logic: when the risk metric it receives is less than a preset warning threshold, the dynamic fusion coefficient is set to a preset maximum value; when the risk metric is greater than or equal to the preset warning threshold and less than or equal to a preset critical threshold, the dynamic fusion coefficient is linearly reduced from the preset maximum value to zero based on the risk metric; when the risk metric is greater than the preset critical threshold, the dynamic fusion coefficient is set to zero.
[0159] By configuring the coefficient generation module in such a clear three-stage logic, the entire system possesses intelligent and smooth risk avoidance capabilities. This configuration encapsulates complex decision-making logic within a single module, ensuring that risk assessment results can be accurately and flawlessly translated into refined control over the beautification effect. This not only enhances the system's automation level but also, through a predictable and reasonable response mechanism, guarantees that the system can output high-quality and reliable images at any risk level, thereby greatly enhancing the system's robustness and end-user trust.
[0160] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A face fusion preprocessing method based on hair repair and lightweight skin beautification, characterized in that, include: The original input image is processed to repair the hair-occluded areas, and the repaired image is obtained. Lightweight skin smoothing is applied to the repaired image to obtain a beautified image; Based on the variance between the beautified image and the preset reference, the unnatural smoothness index of the skin region is determined. Based on the hairline boundary region of the restored and beautified images, the hair restoration boundary structure distortion index is determined; Based on the color histogram of the beautified image and the preset spatial reference probability distribution of natural skin color, the non-physiological deviation index of skin color is determined. By combining the indicators of unnatural smoothness of skin regions, distortion of hair repair boundary structure, and non-physiological deviation of skin color, a risk measure is calculated through a pre-set weighted model. The dynamic fusion coefficient is determined based on risk measurement, preset warning threshold, and preset critical threshold. Based on the dynamic fusion coefficient, the repaired image and the beautified image are weighted and fused to generate the final image.
2. The face fusion preprocessing method based on hair repair and lightweight skin beautification according to claim 1, characterized in that, Determine the dynamic fusion coefficients, including: When the risk metric is less than the preset warning threshold, the dynamic fusion coefficient is set to the preset maximum value; When the risk metric is greater than or equal to the preset warning threshold and less than or equal to the preset critical threshold, the dynamic fusion coefficient is linearly reduced from the preset maximum value to zero based on the risk metric. When the risk metric exceeds the preset critical threshold, the dynamic fusion coefficient is set to zero.
3. The face fusion preprocessing method based on hair repair and lightweight skin beautification according to claim 1, characterized in that, The hair-covered area repair is achieved by using a pre-defined generative adversarial network; Among them, pixel loss and edge loss are combined to form a composite loss function, which is used to optimize and constrain the generative adversarial network.
4. The face fusion preprocessing method based on hair repair and lightweight skin beautification according to claim 1, characterized in that, Lightweight skincare is achieved through a pre-designed lightweight student network. Among them, knowledge distillation technology is used to guide the training of lightweight student networks using a pre-set teacher network.
5. The face fusion preprocessing method based on hair repair and lightweight skin beautification according to claim 4, characterized in that, Knowledge distillation techniques utilize soft label loss and feature map alignment loss for constraints.
6. The face fusion preprocessing method based on hair repair and lightweight skin beautification according to claim 1, characterized in that, Determine indicators of unnatural smoothness in skin areas, including: Calculate the texture variance of the skin region in the beautified image and compare it with a preset reference variance.
7. A face fusion preprocessing system based on hair repair and lightweight skin beautification, based on the face fusion preprocessing method based on hair repair and lightweight skin beautification as described in any one of claims 1-6, characterized in that, include: The hair-occluded area repair module is used to repair the hair-occluded areas of the original input image and obtain the repaired image. The lightweight skin-smoothing module is used to perform lightweight skin-smoothing on repaired images to obtain beautified images; The smoothness index determination module is used to determine the unnatural smoothness index of the skin area based on the beautified image and the preset reference variance. The distortion index determination module is used to determine the distortion index of the hair restoration boundary structure based on the hairline boundary region between the restored image and the beautified image. The deviation index determination module is used to determine non-physiological deviation indexes of skin color based on the color histogram of the beautified image and the preset spatial reference probability distribution of natural skin color. The risk measurement and calculation module is used to combine the skin region non-natural smoothness index, hair repair boundary structure distortion index, and skin color non-physiological deviation index, and calculate the risk measurement through a preset weighted model. The coefficient generation module is used to determine the dynamic fusion coefficient based on risk measurement, preset warning threshold, and preset critical threshold. The image synthesis module is used to perform weighted fusion of the repaired image and the beautified image based on dynamic fusion coefficients to generate the final image.
8. The face fusion preprocessing system based on hair repair and lightweight skin beautification according to claim 7, characterized in that, The coefficient generation module is configured as follows: When the risk metric is less than the preset warning threshold, the dynamic fusion coefficient is set to the preset maximum value; When the risk metric is greater than or equal to the preset warning threshold and less than or equal to the preset critical threshold, the dynamic fusion coefficient is linearly reduced from the preset maximum value to zero based on the risk metric. When the risk metric exceeds the preset critical threshold, the dynamic fusion coefficient is set to zero.
Citation Information
Patent Citations
Method and device for beautifying human face image
CN107369133A
Image processing method and device, storage medium and terminal
CN112784773A
Image processing method and device, electronic equipment and storage medium
CN113763285A
Image buffing processing method and device and storage medium
CN120278923A
Digital human image correction method, electronic equipment and storage medium
CN120355627A
Cited By
Character and background fusion method and system based on adaptive skin color protection
CN121213428A
Hairline detection method based on multi-scale feature fusion
CN121582985A