Image processing method and device, readable storage medium and program product
By acquiring the color card image and performing feature extraction, adjusting the color of the face image taken by the mobile terminal to show a stable skin condition under the preset lighting environment, solving the inaccuracy of skin detection caused by changing light, and achieving accurate skin condition representation in different environments.
Patent Information
- Application Number
- CN202510576261.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-29
AI Technical Summary
By taking face images through mobile terminals, the skin condition represented by face images is relatively accurate, especially affected by the variability of the light environment, resulting in unstable skin detection results.
Obtain preconfigured color card images, obtain skin color block features through color card features, and generate images based on the initial face area and color card features, adjust the color of the initial face image to make it show a stable skin condition under the preset lighting environment.
It improves the accuracy of the skin condition of facial images, ensures that the target face images generated under different lighting environments can accurately reflect the skin condition, and solves the problem of unstable skin detection caused by changing light.
Smart Images

Figure CN120564237A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, device, readable storage medium, and program product. Background Art
[0002] With the growing demand for skin condition testing, skin analyzers have emerged. These analyzers can include components such as a light box, a camera, and a light source. These components can provide a variety of fixed lighting environments, capture facial images of a person under these conditions, and analyze the facial images to perform skin testing. These analyzers can detect skin conditions such as skin age, skin color, wrinkles, and tear troughs. Skin analyzers can accurately detect skin conditions, but individual users typically do not purchase skin analyzers and need to use them at offline institutions such as skin management agencies. This makes it difficult to track the user's skin condition in a timely manner. To improve the timeliness of skin testing, a mobile terminal can often be used to capture a user's facial image as an alternative.
[0003] However, the method of taking facial images by mobile terminals has the problem that the accuracy of the skin condition represented by the facial images is low. Summary of the Invention
[0004] Based on this, the present application provides an image processing method, device, readable storage medium and program product, which can improve the accuracy of skin condition representation of facial images.
[0005] In one aspect, the present application provides an image processing method, comprising:
[0006] Acquire an initial face image; the initial face image is obtained by photographing with a mobile terminal, and the initial face image includes an initial face area;
[0007] Obtaining a preconfigured color card image, wherein the color card image includes a skin color block obtained by photographing the preset color card with a skin tester under a preset lighting environment;
[0008] Performing color card feature extraction on the color card image to obtain color card features, wherein the color card features include skin color block features of the skin color block;
[0009] An image is generated based on the initial facial region and the color card features, so as to adjust the color of the initial facial image based on the color card features and generate a target facial image under the preset lighting environment.
[0010] In one aspect, the present application further provides an image processing device, comprising:
[0011] An acquisition module is configured to acquire an initial facial image; the initial facial image is obtained by photographing the image with a mobile terminal, the initial facial image including an initial facial region; and acquire a pre-configured color card image, the color card image including a skin color block under a preset lighting environment obtained by photographing the preset color card with a skin tester.
[0012] An image generation module is configured to extract color card features from the color card image to obtain color card features, wherein the color card features include skin color block features of the skin color block; generate an image based on the initial facial area and the color card features, adjust the color of the initial facial image based on the color card features, and generate a target facial image under the preset lighting environment.
[0013] In one aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0014] Acquire an initial face image; the initial face image is obtained by photographing with a mobile terminal, and the initial face image includes an initial face area;
[0015] Obtaining a preconfigured color card image, wherein the color card image includes a skin color block obtained by photographing the preset color card with a skin tester under a preset lighting environment;
[0016] Performing color card feature extraction on the color card image to obtain color card features, wherein the color card features include skin color block features of the skin color block;
[0017] An image is generated based on the initial facial region and the color card features, so as to adjust the color of the initial facial image based on the color card features and generate a target facial image under the preset lighting environment.
[0018] In one aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:
[0019] Acquire an initial face image; the initial face image is obtained by photographing with a mobile terminal, and the initial face image includes an initial face area;
[0020] Obtaining a preconfigured color card image, wherein the color card image includes a skin color block obtained by photographing the preset color card with a skin tester under a preset lighting environment;
[0021] Performing color card feature extraction on the color card image to obtain color card features, wherein the color card features include skin color block features of the skin color block;
[0022] An image is generated based on the initial facial region and the color card features, so as to adjust the color of the initial facial image based on the color card features and generate a target facial image under the preset lighting environment.
[0023] In the above-mentioned image processing method, device, readable storage medium and program product, the initial face image is obtained by shooting with a mobile terminal. Due to the variable ambient light source when shooting with the mobile terminal, the initial face area in the initial face image has problems such as variable illumination, which makes the recognition of skin condition affected by illumination, resulting in low accuracy of skin condition represented by the initial face image. In order to solve the above-mentioned problem, a pre-configured color card image is obtained and color card feature extraction is performed on the color card image, and an image is generated based on the initial face area and the color card feature. Since the color card image contains the preset illumination obtained by the skin tester shooting the preset color card, The skin color block in the environment, the color card feature includes the skin color block feature of the skin color block, so that the skin color block feature in the color card feature can guide the adjustment of the color of the initial facial area to the color presented by the face in the preset lighting environment, and generate a target facial image under the preset lighting environment. In this way, no matter what lighting environment the initial facial image is in, the initial facial image can be aligned to the preset lighting environment of the skin tester, so as to provide a target facial image under a stable lighting environment. The target facial image can more accurately represent the skin condition, thereby improving the accuracy of the skin condition represented by the facial image. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0025] Figure 1 A diagram showing an application environment of an image processing method in one embodiment;
[0026] Figure 2 1 is a flow chart of an image processing method according to an embodiment;
[0027] Figure 3 A schematic diagram of an example of a color card in one embodiment;
[0028] Figure 4 Schematic diagram of the architecture of an image generation model in one embodiment;
[0029] Figure 5 Schematic diagram of the architecture of a graph generation model in one embodiment;
[0030] Figure 6A schematic diagram of a Scheduler step-by-step denoising workflow in one embodiment;
[0031] Figure 7 Schematic diagram of a simplified network structure of UNet in one embodiment;
[0032] Figure 8 Detailed network structure diagram of UNet in one embodiment;
[0033] Figure 9 This is a comparison diagram of examples of facial images taken by a skin tester under different types of light sources in one embodiment;
[0034] Figure 10 This is a comparison diagram of examples of facial images taken at different angles by a skin tester in one embodiment;
[0035] Figure 11 This is a comparison chart of examples of standard color card images taken by a skin tester under different types of light sources in one embodiment;
[0036] Figure 12 This is a comparison diagram of examples of first color card images captured by a skin tester under different types of light sources in one embodiment;
[0037] Figure 13 This is an example diagram of a pairing of a second color card image and a fourth face image in one embodiment;
[0038] Figure 14 This is an example diagram of a pairing of a fifth face image and an extracted third color card image in one embodiment;
[0039] Figure 15 A schematic diagram of the architecture of a generation model in the first stage in one embodiment;
[0040] Figure 16 A schematic diagram of the architecture of a generation model in the second stage in one embodiment;
[0041] Figure 17 An example of multiple color-adjusted images in one embodiment;
[0042] Figure 18 A comparison example diagram of a color adjustment image, a color card image, and a third generated image in one embodiment;
[0043] Figure 19 is a structural block diagram of an image processing device in one embodiment;
[0044] Figure 20 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and beneficial effects of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0046] The image processing method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the mobile terminal 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The mobile terminal 102 can capture an initial facial image, and the server 104 can obtain the initial facial image captured by the mobile terminal 102 and obtain a pre-configured color card image, perform color card feature extraction on the color card image to obtain color card features, and generate an image based on the initial facial area and the color card features, so as to adjust the color of the initial facial image based on the color card features and generate a target facial image under a preset lighting environment. The mobile terminal 102 can be a personal computer, a laptop, a smartphone, a tablet computer, and a portable wearable device. The portable wearable device can be a smart watch, a smart bracelet, etc. The server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. It is understandable that the image processing method provided in the embodiments of the present application can also be executed independently by the mobile terminal 102 or the server 104.
[0047] In an exemplary embodiment, Figure 2 As shown, an image processing method is provided, which is applied to Figure 1 The server 104 in FIG. 1 is used as an example to illustrate the method, which includes the following steps 202 to 208.
[0048] in:
[0049] Step 202 , obtaining an initial face image; the initial face image is obtained by photographing with a mobile terminal, and the initial face image includes an initial face area.
[0050] The initial facial image is a facial image obtained by photographing the face of the person to be subjected to skin detection via a mobile terminal. The initial facial region is the region representing the face in the initial facial image. Skin detection may include detection in multiple skin detection dimensions, such as skin age, skin color, skin color uniformity, skin texture, skin sensitivity, and skin blemishes. Skin blemishes may include wrinkles, tear troughs, eye bags, sagging eyelids, dark circles, pores, blackheads, acne, and spots. Wrinkles may include forehead wrinkles, glabellar wrinkles, periorbital wrinkles, crow's feet, nasolabial folds, and mouth wrinkles. The scores of each of the above dimensions can be quantified to generate a skin detection result.
[0051] In the skin detection scenario, facial images taken by a mobile terminal are used as the basis for skin detection, forming an online skin detection scenario based on mobile terminal shooting, which can improve the convenience of skin detection. Users do not need to go to offline institutions to use skin testers for skin detection. However, since users may be in various shooting environments when using mobile terminals for shooting, and the lighting environment is also diverse, even facial images taken by mobile terminals with a very short time interval may produce skin detection results with large differences. There is a problem of unstable skin detection results, and the skin condition identified thereby is inaccurate. Furthermore, in scenarios such as skin care or makeup, it is impossible to provide users with accurate skin care or makeup suggestions. Therefore, the initial facial image needs to be further processed, and the specific processing will be explained in subsequent steps.
[0052] For example, the server may receive an image processing request sent by a mobile terminal and obtain an initial facial image from the image processing request. The image processing request may be generated within a skin detection interface of the mobile terminal. For example, the skin detection interface may display a skin detection trigger control. When the skin detection trigger control is triggered after the initial facial image has been uploaded to the skin detection interface, an image processing request may be generated.
[0053] Step 204: Acquire a pre-configured color card image, where the color card image includes a skin color block obtained by photographing the preset color card with the skin tester under a preset lighting environment.
[0054] Among them, the color card image can be an image obtained by the skin tester when the preset color card is in a preset lighting environment. The preset color card is a pre-set color card. The color card is a color reference tool that brings together a variety of colors. Each color in the preset color card can be presented in blocks. The skin color block is the image area where each skin color in the preset color card is presented in the color card image. The color card image may include multiple different skin color blocks. It can be understood that the same color in the color card (referring to the color recorded by the color card) can present different colors (referring to the color recorded by the image) in different color card images taken under different lighting environments. In this way, the color card image can reflect the lighting environment in which it is located.
[0055] The preset color chart can be a standard color chart designed by a color research organization, such as a Pantone color chart, a RAL color chart, or others. The preset color chart can also be a redesigned skin color chart, for example, a skin color chart specifically designed for the online skin testing scenario described in this application. The colors in the skin color chart can be collected from different standard color charts to enrich the diversity of the preset color charts.
[0056] The colors in the skin color test card can be set according to the needs. For example, the skin color test card can include at least 20 skin colors of different shades, including white, gray of different shades, and other standard colors for auxiliary use. Figure 3 As shown in the schematic diagram of the color card example, the skin color measurement card may include 60 colors, among which the middle 6*6 area may include 36 skin colors, the lower row of the 6*6 area may include 6 different shades of gray, and the left and right columns of the 6*6 area and the upper row may include 18 auxiliary standard colors.
[0057] A skin analyzer is a skin testing instrument. It can have a standard light source system that produces a stable lighting environment. This ensures consistent skin tone across facial images captured under the same lighting environment. This means that even if the same face is captured within a short timeframe, the skin tone will remain consistent.
[0058] Skin analyzers can present skin detection results by establishing a skin quantification system and calculating quantitative scores for multiple skin detection dimensions. When performing skin detection using a skin analyzer, facial images can be captured using different light sources, such as white light, cross-polarized light, parallel polarized light, UV light, or other light sources. By combining facial images captured under different light sources, scores for multiple skin detection dimensions are calculated. Skin analyzers can be equipped with standard lighting, cameras, and other components, and combined with deep learning technology, they can achieve accurate and stable skin detection, resulting in highly consistent skin detection results.
[0059] A preset lighting environment is a pre-set lighting environment. It can be the lighting environment created when the skin analyzer is using a preset light source type. The preset type can be white light, cross-polarized light, parallel polarized light, UV light, or other. Using different light sources can create different lighting environments.
[0060] Color card images can be used to record the lighting environment created by a skin analyzer. Because different skin analyzers can create different lighting environments, such as different light source systems and light source layouts, text information alone can be difficult to fully capture. Color card images, however, provide a more complete record of the lighting environment created by a skin analyzer.
[0061] For example, in a skin analyzer, when white light hits a preset color card, each color block of the preset color card will show a different color, which can more comprehensively record the lighting environment formed by the white light source in the skin analyzer. It is difficult to describe the complete lighting environment with text information. Moreover, once the skin analyzer replaces or changes a light source (for example, changing the position, brightness, etc.), the entire text information needs to be re-corrected for the lighting environment formed by the light source. However, by recording the lighting environment through the color card image, you only need to re-photograph the preset color card in the skin analyzer to obtain a new color card image. The new color card image can fully describe the lighting environment formed by the skin analyzer under the light source, without the need to accurately record various information about the light source.
[0062] The color chart image can be obtained by photographing a preset color chart using a specific skin analyzer model in a preset lighting environment. In conjunction with steps 206 to 208, the initial facial image captured by the mobile terminal can be aligned with the preset lighting environment created by the skin analyzer to obtain the target facial image. If the requirements change, for example, if the initial facial image needs to be aligned with the lighting environment of another skin analyzer model, the preset color chart can be placed in the other skin analyzer model and photographed to obtain a color chart image corresponding to the other skin analyzer model. In conjunction with steps 206 to 208, the initial facial image captured by the mobile terminal can be aligned with the lighting environment created by the other skin analyzer model.
[0063] Exemplarily, multiple preset light sources of various types correspond one-to-one with multiple preconfigured color card images, and the preset lighting environment of each color card image is the preset lighting environment created by the light sources of the preset types corresponding to the color card images. In this embodiment, the server may obtain a color card image from the multiple preconfigured color card images.
[0064] The color card image obtained may be a color card image randomly obtained from a plurality of color card images, or may be a color card image corresponding to a preset type of light source indicated by the image processing request obtained from a plurality of color card images. It is understood that the user may specify the type of light source.
[0065] Step 206 , extracting color card features from the color card image to obtain color card features, where the color card features include skin color block features of the skin color blocks.
[0066] Color card feature extraction is the process of converting a color card image from image space to feature space. Specifically, the feature space can be a latent space. Color card features can be the data representation of the color card image in the feature space. Color card features can represent the color information of each color block in the color card image, such as the brightness, saturation, or other information displayed by each color block in the color card image. Skin color block features can represent the color information of skin color blocks. Color card features can be represented as vectors.
[0067] Exemplarily, the server can extract color card features from the color card image using a trained color card encoder to obtain color card features. The color card encoder can encode the color card image from a high-dimensional image space to a low-dimensional feature space. The color card encoder can be implemented using an encoder that processes images, for example, the image encoder in the CLIP model (Contrastive Language-Image Pre-training, a multimodal pre-training model) or the image encoder in the BLIP model (Bootstrapping Language-Image Pre-training, a multimodal framework that improves the joint understanding and generation capabilities of images and languages through self-guided Bootstrapping).
[0068] Step 208 : generating an image based on the initial face region and the color card features, adjusting the color of the initial face image based on the color card features, and generating a target face image under a preset lighting environment.
[0069] Image generation can be performed with the goal of adjusting the color in the initial facial region to match the color of the applicable skin color patch, and the overall color of the initial facial image can be adjusted accordingly. The applicable skin color patch can be a patch that matches the skin color of the face represented by the initial facial region when placed in a preset lighting environment. Image generation can be implemented based on a trained image information creator, which can understand the lighting environment of the facial image captured by the mobile terminal during training and learn how to reconstruct image information based on the color card features and information in the facial image. The reconstructed image information can be used to generate a facial image under the lighting environment corresponding to the color card features.
[0070] Exemplarily, the server may generate target image features based on the initial facial image and color card features through a trained image information creator and then perform decoding processing to adjust the color of the initial facial image based on the color card features to generate a target facial image under a preset lighting environment.
[0071] Among them, the image information creator can reconstruct image information through a diffusion process to generate image features. During the image information reconstruction process, the image information creator can focus on the information of the facial area in the facial image and the skin color block features in the color card features, and focus on reconstructing information according to the skin color block features in the color card features and the information of the facial area. The diffusion process of the image information creator can adopt the diffusion process of a diffusion model, such as DDIM (Denoising Diffusion Implicit Models), Stable Diffusion (Stable Diffusion Model) or others. The decoding process is the process of converting features into images. The decoding process can be implemented by a pre-trained decoder. The decoder can adopt the decoder in VAE (Variational Autoencoder), Consistency Decoder (Consistency Decoder) or others.
[0072] In the above-mentioned image processing method, the initial facial image is obtained by photographing with a mobile terminal. Due to the variable ambient light source during photographing with the mobile terminal, the initial facial region in the initial facial image is subject to variable lighting, resulting in the recognition of skin condition being affected by the lighting. This results in low accuracy of the skin condition represented by the initial facial image. To address the above-mentioned problem, a pre-configured color card image is obtained and color card features are extracted from the color card image. An image is generated based on the initial facial region and the color card features. Since the color card image includes skin color blocks obtained by a skin analyzer photographing the preset color card under a preset lighting environment, the color card features include skin color block features of the skin color blocks. Therefore, the skin color block features in the color card features can guide the adjustment of the color of the initial facial region to the color presented by the face under the preset lighting environment, thereby generating a target facial image under the preset lighting environment. In this way, regardless of the lighting environment in which the initial facial image is located, the initial facial image can be aligned to the preset lighting environment of the skin analyzer, thereby providing a target facial image under a stable lighting environment. The target facial image can more accurately represent the skin condition, thereby improving the accuracy of the skin condition represented by the facial image.
[0073] In an exemplary embodiment, color card feature extraction is implemented through a color card encoder in a configured image generation model. The image generation model also includes an image encoder, an image information creator, and an image decoder. Step 208 may include: based on the initial face image, performing feature extraction through the face image encoder and then performing noise processing to obtain initial face image features; through the image information creator, performing iterative denoising processing based on the initial face image features and the color card features to generate target image features, and then performing decoding processing through the image decoder to adjust the color of the initial face image based on the color card features to generate a target face image under a preset lighting environment.
[0074] Among them, the face image encoder can encode the face image from a high-dimensional image space to a low-dimensional feature space. The feature space can be a latent space. The face image encoder can adopt the encoder in VAE, TAESD (TinyAutoEncoder for Stable Diffusion, a micro autoencoder for a stable diffusion model) or others. Noise addition is the process of adding noise. The face image encoder can extract the initial face image features to obtain latent features, and by gradually adding noise to the latent features, the initial face image features can be obtained. The initial face image features may include noisy features corresponding to the initial face area. During the image information reconstruction process, the image information creator may pay attention to the noisy features corresponding to the initial face area and the skin color block features in the color card features, focusing on information reconstruction according to the skin color block features in the color card features and the noisy features corresponding to the initial face area.
[0075] Iterative denoising is a process of gradually denoising the initial facial image features with reference to the color card features. The image information creator may include a noise prediction network and a scheduler. Noise addition can be implemented by a scheduler, and iterative denoising can be implemented by a noise prediction network and a scheduler. The scheduler may include a calculation process for adding noise or removing noise, for example, it may include a mathematical formula or a process, and different schedulers have different calculation processes. The scheduler may adopt the Scheduler component in the Stable Diffusion model. The noise prediction network is used to predict noise, and the noise prediction network may be a UNet network (a network including an encoder and a decoder with jump connections between the encoder and the decoder).
[0076] In this embodiment, the initial facial image is encoded by a facial image encoder and then denoised to obtain initial facial image features for easy processing by an image information creator. The image information creator performs iterative denoising based on the initial facial image features and the color card features, and can gradually inject the color card features into the initial facial image, thereby generating target image features and then decoding them through an image decoder, thereby adjusting the color of the initial facial image based on the color card features and generating a target facial image under a preset lighting environment.
[0077] In one embodiment, the schematic diagram of the architecture of the image generation model can be as follows: Figure 4 shown. Figure 4The image generation model shown may include Card Encoder (which may represent a color card encoder), Image Encoder (which may represent a face image encoder), Image Information Creator (which may represent an image information creator) and Image Decoder (which may represent an image decoder), and the Image Information Creator may include a noise prediction network UNet and a scheduler Scheduler.
[0078] The image generation model can be constructed based on the graph generation model architecture. The architecture diagram of the graph generation model can be as follows: Figure 5 As shown, it includes a Text Encoder (which can represent a text encoder), an Image Encoder, an ImageInformation Creator, and an Image Decoder. This image-to-image model uses text to guide adjustments to the input facial image, hoping to adjust the lighting environment of the facial image according to the text. For example, the text can be "a young girl under the UV light source."
[0079] The image generation process of this graph-to-graph model based on text and face images can be achieved through the following steps: input text into the Text Encoder to obtain Token embeddings (which can represent text embedding features); input the face image into the Image Encoder to obtain the potential features in the latent space, and add noise to the potential features to obtain the Initial image information tensor (in the example Figure 5 The architecture shown can represent the initial image information tensor), input the Token embeddings and the Initial image information tensor into the Image Information Creator, and perform the Diffusion process (diffusion process) through the Image Information Creator to reconstruct the image information and generate a Processed image information tensor containing text information. The generated face image (Generated Image) is obtained by decoding it through the Image Decoder.
[0080] Combined with the actual application, it is found that the text is difficult to accurately and completely record the lighting environment formed by the skin analyzer, and the generated face image is difficult to meet expectations. Therefore, the image generation model is further improved, such as Figure 4 As shown in the figure, color card images are used to replace text, Card Encoder is used to replace Text Encoder, and Card embeddings (which can represent color card features) are used to replace Token embeddings. Based on this, the image generation model can generate the target face image based on the color card image and the initial face image through the following steps. The color card image is input into the Card Encoder to extract the color card features, and the Card embeddings can be obtained; the initial face image is input into the Image Encoder to obtain the potential features in the latent space, and the potential features are denoised to obtain the Initial image information tensor (as shown in the figure). Figure 4 The architecture shown in FIG can represent the initial face image features); input the Card embeddings and Initial image information tensor into the Image Information Creator, and perform the Diffusion process through the Image Information Creator to reconstruct the image information and generate a Processed image information tensor containing the reconstructed image information (as shown in FIG). Figure 4 The architecture shown can represent the target image features), and then input the processed image information tensor into the Image Decoder for decoding to generate the target face image.
[0081] The color card features can be represented as vectors with a size of 77*768. The initial facial image features can be represented as tensors. The target image features can be represented as tensors with a size of 64*64. For the feature maps in the image generation model (including the input image), more detailed size information can be described in the form of b*c*h*w, where b represents the batch size (i.e., the number of images processed by the image generation model in one generation process), c represents the number of channels (3 for an image), h represents the height of the feature map, and w represents the width of the feature map. The input size of the image generation model can be 2*3*512*512 (b=2, c=3, h=512, w=512, for example, two initial facial images with 3 channels and a width and height of 512 are input). This input is downsampled by 8 times by the Image Encoder to obtain a feature map of size 2*4*64*64. After noise addition, the initial facial image features are 2*4*64*64. The input size of the image generation model can also be 2*3*768*512 (b=2, c=3, h=768, w=512). The feature map size obtained by downsampling the input by 8 times is 2*3*96*64.
[0082] In one embodiment, during the iterative denoising process based on the initial face image features and color card features by the Image Information Creator, the iterative denoising process can be performed in the latent space by the noise prediction network UNet and the scheduler. Figure 6 The Scheduler step-by-step denoising workflow diagram shown and Figure 4 , iterative denoising processing can be implemented through the following steps: in the first denoising time step (Step 1), based on the color card features and the initial face image features, the UNet network is used to predict noise to obtain the predicted noise of the current denoising time step, and the Scheduler is used to remove the predicted noise of the current denoising time step from the initial face image features to obtain the Latent features (latent features) generated in the first denoising time step; from the second denoising time step (Step 2) to the Sth denoising time step (Step S), in each denoising time step, based on the color card features and the Latent features generated in the previous denoising time step, the UNet network is used to predict noise to obtain the predicted noise of the current denoising time step, and the Scheduler is used to remove the predicted noise of the current denoising time step from the Latent features generated in the previous denoising time step to obtain the Latent features generated in the current denoising time step.
[0083] Among them, S in the Sth denoising time step can represent the number of iterations or the number of sampling times, and S can be an integer between 50 and 100. The Latent features generated in the Sth denoising time step can be used as the processed image information tensor (which can represent the target image features) output by the Image Information Creator. The prediction noise can be called Prediction Noise or just Noise. The Latent features generated in each denoising time step have a reduced degree of noise compared to the input of the denoising time step. Through denoising processing of multiple denoising time steps, the information in the color card features can be injected into the Latent features, so that the target image features finally generated can be in the preset lighting environment corresponding to the color card features.
[0084] The UNet in the image information creator can adopt the UNet in Stable Diffusion, and the simplified network structure diagram of UNet can be as follows: Figure 7 As shown, the detailed network structure diagram of UNet can be shown as Figure 8 shown. Figure 7 It also includes the Scheduler denoising process, that is, the Scheduler removes the noise predicted by UNet from the Latent features input to UNet, and obtains the Latent features generated by the current denoising time step (that is, Figure 7 The Next Latent in can be understood as the Latent feature used to input into the next denoising time step).
[0085] UNet can include downsampling stage and upsampling stage, Figure 8 In the figure, T1 to T16 identify multiple Transformer2DModel modules, T1 to T6 are Transformer2DModel modules in the downsampling stage, and T8 to T16 are Transformer2DModel modules in the upsampling stage. At the same time, the figure identifies the implicit feature maps (featuremaps) output by different modules in the downsampling stage, such as OutDown64x64A, OutDown64x64B, OutDown32x32A, OutDown32x32B, as well as the implicit feature maps output by different modules in the upsampling stage, such as OutUp64x64A, OutUp64x64B, OutUp32x32A, OutUp32x32B. Taking OutDown64x64A and OutDown64x64B as examples, they represent an implicit feature map of size 64x64 obtained by downsampling. A and B can be used to distinguish implicit feature maps of the same size.
[0086] Combine Figure 8 It can be seen that in each denoising time step, the Latent feature generated by the previous denoising time step can be input into the Conv_in layer of UNet (in the first denoising time step, the initial face image feature is input into the Conv_in layer), the color card feature is input as prompts_embdding into each Transformer2DModel module (each Transformer2DModel module has prompts_embdding input), and the time step embedding feature time_embdding obtained by encoding the denoising time step is input into each ResnetBlock2D module (each ResnetBlock2D module has time_embdding input), so that UNet can be based on Figure 8 The network structure shown implements noise prediction.
[0087] In an exemplary embodiment, the training steps of training a color card encoder and an image information creator include: obtaining a graphic data set, the graphic data set including a first face data set and a second face data set; the first face data set includes each first face image taken by a skin tester and each paired description text, the second face data set includes each second face image taken by a mobile terminal and each paired description text, the description text is used to describe the lighting environment in which the paired images are located; based on the graphic data set, the pre-trained image information creator is trained to obtain a first-stage image information creator; obtaining a color card face data set, the color card face data set includes a first color card face data set and a second color card face data set; the first color card face data set includes multiple groups of paired first color card images and third face images taken by a skin tester; the second color card face data set includes multiple groups of paired second color card images and fourth face images taken by a mobile terminal; based on the color card face data set, the pre-trained color card encoder and the first-stage image information creator are trained to obtain a trained color card encoder and a trained image information creator.
[0088] Among them, pre-training refers to training in advance through a large data set. The large data set can be a general data set. The first face images in the first face data set may include first face images taken by different skin testers respectively; the first face images taken by each skin tester may include first face images taken by the skin tester under lighting environments formed by different types of light sources, and under each lighting environment, first face images at different angles can be taken, for example, faces at three angles: left face, front face, and right face. The number of first face images in the first face data set can be more than 1 million. In some scenarios, the first face data set can be called a skin tester face data set, which can be recorded as FaceStdSet1.
[0089] For example, in the first face dataset, the comparison chart of face images taken by the skin tester under different types of light sources can be shown as follows: Figure 9 As shown, Figure 9 (a) can represent a face image taken by the skin tester in a white light environment with a frontal angle. Figure 9 (b) can represent a face image taken by the skin tester at a frontal angle under a lighting environment formed by cross-polarized light, Figure 9 (c) can represent a face image taken by the skin tester at a frontal angle under a lighting environment formed by parallel polarized light. The comparison chart of face images taken by the skin tester at different angles can be shown as follows: Figure 10 As shown, Figure 10 (a) can represent a facial image of the right face angle taken by the skin tester under the lighting environment formed by white light, Figure 10 (b) can represent a face image taken by the skin tester in a white light environment with a frontal angle. Figure 10 (c) may represent a facial image of a person at a left face angle taken by the skin tester under a lighting environment formed by white light.
[0090] The second facial images in the second facial dataset may include second facial images captured using different models of mobile terminals (such as mobile phones). The second facial images captured by each model of mobile terminal may include second facial images captured by the mobile terminal under different lighting environments. Each lighting environment may include second facial images captured from different angles. The number of second facial images in the second facial dataset may be greater than one million. In some scenarios, the second facial dataset may be referred to as a mobile phone face dataset, which may be denoted as PhoneFaceSet1.
[0091] The descriptive text paired with the first facial image may include text describing the lighting environment formed by the skin tester in which the first facial image is located. Specifically, it may be represented by the light source information used by the skin tester to form the lighting environment, for example, it may include information such as light angle, light brightness, light color, and light category. The descriptive text paired with the second facial image may include text describing the lighting environment in which the second facial image is located. Specifically, it may be represented by the light source information of the mobile terminal, for example, it may include information such as whether the camera used is a front camera or a rear camera, whether the flash is enabled, whether it is backlit, whether it is backlit, whether it is indoors or outdoors, etc. It can be understood that each image in the graphic data set is paired with a descriptive text.
[0092] The image and text data set may also include a standard color card dataset. A standard color card dataset may include images of each standard color card and their corresponding descriptive text. A standard color card image contains a standard color patch captured by a skin tester. A standard color patch is the image representation of the standard colors in the standard color card. The standard color card may be a Pantone color card, a RAL color card, or other color cards, specifically a Pantone skin tone card with 138 colors. In some scenarios, the standard color card dataset may be denoted as ToneCardSet1.
[0093] The standard color card images in the standard color card dataset can include images taken using different skin analyzers. Each skin analyzer can capture images of the standard color card under different lighting conditions. Under each lighting condition, images of the standard color card can be captured from different angles, such as from the left, front, and right sides of the standard color card. Capturing images from different angles enriches the dataset and helps improve the model's generalization capabilities during subsequent training.
[0094] For example, in the standard color card dataset, the comparison chart of standard color card images taken by the skin analyzer under different types of light sources can be shown as follows: Figure 11 As shown, Figure 11 (a) can represent the standard color card image taken by the skin analyzer under the lighting environment formed by white light, Figure 11 (b) can represent the standard color card image taken by the skin analyzer under the lighting environment formed by cross-polarized light, Figure 11 (c) may represent an image of a standard color card taken by the skin analyzer under a lighting environment formed by parallel polarized light.
[0095] The image and text dataset can also include a standard object dataset. This dataset can include images of standard objects captured using a skin analyzer standard object, along with their corresponding descriptive text. Standard objects are less reflective than color charts and human faces. These objects can include silicone faces, silicone fruit, and other silicone objects. They can also include colored, low-reflectivity cardboard. In some scenarios, the standard object dataset can be denoted as StdObjectSet1.
[0096] The standard object images in the standard object data set may include standard object images taken of the standard objects using different skin testers; the standard object images taken by each skin tester may include standard object images taken of the standard objects by the skin tester under lighting environments formed by different types of light sources. Under each lighting environment, standard object images at different angles may be taken, for example, from the left side, front, and right side of the standard object.
[0097] The first color card face data set may include multiple color card face data sets. Each color card face data set may include a first color card image of a sample color card taken using a skin tester under a lighting environment formed by different types of light sources, and a third face image of at least one sample face taken using the skin tester under a lighting environment formed by different types of light sources after the first color card image of the current color card face data set is taken. The sample color card is removed from the skin tester and then a third face image is taken of the sample color card using the skin tester under a lighting environment formed by different types of light sources. The first color card face data set may include more than 300 color card face data sets. In some scenarios, the first color card face data set may be denoted as CardStdFacePairSet1.
[0098] In the first color card face data set, the position of the sample color card in the skin tester can vary between different color card face data groups, and the position of the sample color card in each group does not need to be absolutely correct, which can avoid the sample color cards being in the same position and improve the diversity of the training samples. The sample color card can be a standard color card, or a redesigned skin color test card, or a color card combination formed by a standard color card and a skin color test card. The color card combination can provide a wider range of color information and can more comprehensively reflect the lighting environment. The at least one sample face can be one sample face, or more than two sample faces, such as 5 to 10 sample faces. The first color card image and the third face image in the same color card face data group under the same lighting environment can be used as a set of paired first color card images and third face images.
[0099] In the first color card face dataset, the comparison chart of the first color card images taken by the skin tester under different types of light sources can be seen as follows Figure 12 As shown, here, the first color card image can be obtained by photographing a sample color card including a standard color card and a skin color measurement card. Figure 12 (a) may represent the first color card image captured by the skin analyzer under a lighting environment formed by white light, Figure 12 (b) may represent the first color card image captured by the skin analyzer under the illumination environment formed by cross-polarized light, Figure 12 (c) can represent the first color card image taken by the skin tester under the illumination environment formed by parallel polarized light. Figure 12 When the first color card image is in the image, the color card area in the first color card image can be extracted by the corner points, and the color card area similar to Figure 3 The results shown are shown, removing other redundant information.
[0100] The second color card face dataset may include multiple sets of paired second color card images and fourth face images, captured using different mobile terminal models and in different lighting environments. Each set of paired second color card images and fourth face images was captured sequentially using the same mobile terminal model under the same lighting environment. For example, a fourth face image may be captured first, followed by a color card image, from which the color card region is extracted to obtain the second color card image. Over 300 sets of paired second color card images and fourth face images may be collected. In some scenarios, the second color card face dataset may be denoted as CardPhoneFacePairSet1.
[0101] In the second color card face dataset, the second color card image and the fourth face image pairing example diagram can be seen as follows Figure 13 As shown, Figure 13 (a) can represent the fourth face image taken by the mobile phone, Figure 13 (b) It can indicate that the mobile phone is in Figure 13 (a) A color card image taken under the same lighting environment. The image contains the color card area. Figure 13 (c) can be expressed as Figure 13 (b) Extracted from Figure 13 (a) Paired second color card image.
[0102] The color card face dataset may also include a third color card face dataset that is different from the first and second color card face datasets. The third color card face dataset may include multiple paired third color card images and fifth facial images. The fifth facial image may be captured by a mobile terminal and include a sample facial region and a sample color card region. The third color card image is extracted from the sample color card region in the paired fifth facial image.
[0103] Each fifth face image in the third color card face dataset can be captured using different models of mobile terminals while a person is holding a skin color test card, allowing the face and skin color test card to be presented in a single image. When holding the skin color test card, the top of the skin color test card can be aligned with the area near the chin of the face. The sample face area can be the area representing the face in the fifth face image, and the sample color card area can be the area representing the skin color test block in the fifth face image. The sample color card area is extracted from the fifth face image to form a third color card image paired with the fifth face image. More than 300 sets of paired third color card images and fifth face images can be collected. In some scenarios, the third color card face dataset can be denoted as CardPhoneFaceOneSet1.
[0104] In the third color card face dataset, the fifth face image and the extracted third color card image pairing example can be seen as follows Figure 14 As shown, Figure 14 (a) may represent a fifth face image taken with a mobile phone while the person holds the skin color card near the chin. Figure 14 (b) can be expressed as Figure 14 (a) Extracted from Figure 14 (a) Paired third color card image.
[0105] In this embodiment, since the graphic data set includes facial images taken by a skin tester and a mobile terminal respectively, as well as descriptive texts of the facial images, and the descriptive texts are used to describe the lighting environment in which the paired images are located, the image information creator is trained based on the graphic data set, and the obtained first-stage image information creator is able to understand the lighting environment information of the facial images taken by the skin tester and the mobile terminal respectively; further, on this basis, since the color card facial data set includes paired color card images and facial images taken by a skin tester and a mobile terminal respectively, the color card encoder and the first-stage image information creator are trained based on the color card facial data set, the color card encoder is able to learn how to fully encode the lighting environment information recorded by the color card image into color card features, and the image information creator is able to learn how to fully inject the encoded color card features into the facial image for lighting adjustment based on the understanding of the lighting environment information, thereby obtaining an accurately encoded color card encoder and an image information creator that accurately reconstructs image information.
[0106] In an exemplary embodiment, the graphic data set also includes a standard color card data set and a standard object data set; the standard color card data set includes each standard color card image and each paired description text; the standard color card image contains a standard color block photographed by a skin tester on the standard color card; the standard object data set includes each standard object image photographed by a skin tester standard object and each paired description text; the reflectivity of the standard object is lower than that of the color card and the human face; the description text is used to describe the lighting environment of the paired image and the image category to which it belongs.
[0107] The image category can be a face, an object, or a color card. The description text for the standard color card image pairing and the description text for the standard object image pairing can include information about the lighting environment using the light source used by the skin analyzer to create the lighting environment. For example, this information may include lighting angle, brightness, color, and type.
[0108] For the second facial image, the description text paired with the second facial image can also be used to describe the model of the mobile terminal used to take the second facial image, such as the mobile phone model, so that through training, the image information creator can understand the lighting environment information of different models of mobile terminals.
[0109] In this embodiment, the graphic data set includes a first face data set, a second face data set, a standard color card data set, and a standard object data set, which further improves the richness of the training data. The standard color card data set can fully reflect the lighting environment of the skin analyzer through a variety of colors, and the reflectivity of the standard object is lower than that of the color card and the face. The standard object image thus formed can more accurately reflect the color information than the face image and the color card image; and the descriptive text is used to describe the lighting environment and image category of the paired image. When the image information creator is trained based on the graphic data set, the image information creator can fully understand the semantic category of the image and the lighting environment information, thereby improving the performance of the image information creator.
[0110] In an exemplary embodiment, a pre-trained image information creator is trained based on a graphic and text data set, and the step of obtaining a first-stage image information creator may include: obtaining a first-stage generation model, the first-stage generation model including a pre-trained text encoder, a pre-trained image information creator and a pre-trained image decoder; based on the graphic and text data set, the image information creator in the first-stage generation model is iteratively trained to obtain a first-stage image information creator; wherein, in each round of iteration, based on each descriptive text in the graphic and text data set, the first generated image corresponding to each descriptive text is generated through the first-stage generation model of the current round, and the image information creator in the first-stage generation model of the current round is updated according to the difference between each image in the graphic and text data set and the first generated image corresponding to each paired descriptive text.
[0111] The first-stage generative model uses a text-based graph architecture. For example, the first-stage generative model can use the text-based graph architecture of StableDiffusion. The text encoder can use the text encoder in the CLIP model or the BLIP model. The image information creator can be constructed using a noise prediction network and a scheduler. The noise prediction network can be a UNet model. The pre-trained image decoder can use the decoder in the VAE (Variational Autoencoder), the Consistency Decoder, or other models.
[0112] The architecture diagram of the first stage generation model can be shown as follows Figure 15 As shown, the model structure may include Text Encoder (which may represent a text encoder), Image Information Creator (which may represent an image information creator), and ImageDecoder (which may represent an image decoder). Figure 15 Based on each description text in the image-text dataset, the process of generating the first generated image corresponding to each description text through the first stage generation model of the current round can be as follows:
[0113] For each description text, the description text is input into the text encoder to obtain the text embedding features (Token embeddings) corresponding to the description text, obtain the random noise features (Random image information tensor), input the text embedding features and random noise features into the image information creator, perform the Diffusion process through the image information creator to reconstruct the image information, and obtain the text-integrated features (Processed image information tensor) that incorporate the information in the description text. The text-integrated features are then input into the image decoder for decoding to generate the first generated image corresponding to the description text.
[0114] The text embedding feature can be sized 77*768. The random noise feature can be a 64*64 image randomly generated by a LatentSeed (a random seed using Gaussian noise ~N(0,1)), representing a completely noisy image. The text-integrated feature can contain image information generated based on the semantic information of the descriptive text. The size of the first generated image can be 512*512.
[0115] Based on the text-image data set, fine-tune iterative training can be performed during the iterative training of the image information creator in the first-stage generative model. Specifically, the model parameters of both the text encoder and the image decoder in the first-stage generative model can be fixed, and the model parameters of the image information creator in the first-stage generative model can be opened. Specifically, only the model parameters of the noise prediction network in the image information creator can be opened. For example, the above-mentioned text-image data set can be used to fine-tune the model parameters in UNet on the existing Stable Diffusion text-image architecture. Fixed model parameters can mean that the model does not participate in training, and the model parameters remain unchanged during the training process. Open model parameters can mean that the model participates in training, and the model parameters can be updated during the training process.
[0116] During the training of the image information creator of the first-stage generative model, each image in the image-text dataset can be used as annotation data. The difference between each image in the image-text dataset and the first generated image corresponding to the corresponding paired descriptive text can be measured by calculating the loss value using a loss function. In each round of iteration, the backpropagation technology can be used to update the model parameters of the noise prediction network in the image information creator of the first-stage generative model of the current round to update the image information creator of the first-stage generative model of the current round. The current round refers to the round of the current iteration. In the first round of iteration, the first-stage generative model of the current round can refer to the acquired first-stage generative model; starting from the second round of iteration, in each round of iteration, the first-stage generative model of the current round can be the first-stage generative model after the image information creator is updated in the previous round.
[0117] In this embodiment, the image information creator in the first-stage generation model is iteratively trained through the graphic data set. Since in each round of iteration, the first generated image corresponding to each descriptive text is generated by the first-stage generation model of the current round, and then the image information creator in the first-stage generation model of the current round is updated according to the difference between each image in the graphic data set and the first generated image corresponding to each paired descriptive text, the image information creator can gradually and accurately understand the lighting environment of the images taken by the mobile terminal and the skin tester.
[0118] In an exemplary embodiment, the steps of training the pre-trained color card encoder and the first-stage image information creator based on the color card face dataset to obtain the trained color card encoder and the trained image information creator include:
[0119] A second-stage generative model is obtained, which includes a pre-trained color card encoder, a first-stage image information creator, and a pre-trained image decoder. Based on the color card face data set, the color card encoder and the image information creator in the second-stage generative model are iteratively trained to obtain the second-stage color card encoder and the second-stage image information creator. In each round of iteration, based on each color card image in the color card face data set, the second-stage generative model of the current round is used to generate a second generated image corresponding to each color card image. According to the difference between the face image paired with each color card image and the corresponding second generated image, the color card encoder and the image information creator in the second-stage generative model of the current round are updated. Data transformation is performed on each face image to adjust the color of each face image to obtain a color-adjusted image paired with each face image. Based on the color card face data set and the color-adjusted image paired with each face image, the second-stage color card encoder and the second-stage image information creator are trained to obtain a trained color card encoder and a trained image information creator.
[0120] The trained first-stage generative model includes the first-stage image information creator. The second-stage generative model can be formed by replacing the text encoder in the trained first-stage generative model with a pre-trained color card encoder. The pre-trained color card encoder can use the image encoder in the CLIP model or the BLIP model.
[0121] The architecture diagram of the second stage generation model can be shown as follows Figure 16 As shown, Figure 15 The difference between the architecture shown is that Figure 16 The architecture shown uses color card images instead of description text, and uses Card Encoder instead of Text Encoder. Card Encoder is used to encode the input color card images into color card features (Card embeddings), and Card embeddings replace Token embeddings. Based on this, based on each color card image in the color card face data set, the second stage generation model of the current round is used to generate the second generated image corresponding to each color card image. Combined with the above Figure 15 Based on each description text in the image and text data set, the process of generating the first generated image corresponding to each description text through the first stage generation model of the current round is similar logic and will not be repeated here.
[0122] Based on the color card face data set, fine-tune iterative training can be performed during the iterative training of the color card encoder and image information creator in the second-stage generative model. Specifically, the model parameters of the image decoder in the second-stage generative model can be fixed, and the model parameters of the color card encoder and image information creator in the second-stage generative model can be opened, among which the model parameters of the image information creator can only open the model parameters of the noise prediction network.
[0123] The noise prediction network, for example, can be a Unified Network (UNet) network, which can be fine-tuned using LoRA (Low-Rank Adaptation) technology. Specifically, the Transformer2DModel module in the UNet network can be trained. Modules other than the Transformer2DModel module in the UNet network are not trained, allowing the Transformer2DModel module to fully incorporate card embedding information during training. The color card encoder is also trained. The learning rate of the Transformer2DModel module can be 0.1 to 0.2 times the learning rate of the color card encoder.
[0124] LoRA technology is a parameter-efficient fine-tuning (PEFT) technique that trains only a subset of the original model's parameters during fine-tuning to accelerate the process. Compared to other PEFT methods, LoRA stands out due to several distinct advantages. Performance-wise, LoRA requires storing only a small number of fine-tuned parameters, rather than saving the entire new model. Furthermore, LoRA's new parameters can be merged with those of the original model, without increasing model execution time. Functionally, LoRA maintains the amount of change in the model during fine-tuning. By multiplying the change by a mixing ratio between 0 and 1, the degree of model modification can be controlled. Furthermore, multiple LoRAs trained independently from the same original model can be used simultaneously.
[0125] During the training of the color card encoder and image information generator of the second-stage generative model, each facial image in the color card face dataset can be used as annotation data. The difference between the facial image paired with each color card image in the color card face dataset and the corresponding second generated image of each color card image in the color card face dataset can be measured using a loss function to calculate a loss value. In each iteration, backpropagation technology can be used to update the model parameters of the color card encoder and the noise prediction network of the image information generator in the current second-stage generative model, thereby updating the color card encoder and image information generator in the current second-stage generative model.
[0126] Data transformation can be a process of adjusting the color of a facial image. For example, data transformation can be performed by adding shadows, adding highlights, adjusting color temperature, adjusting hue, reducing contrast, and other data gain processing. Data transformation can also use deep learning to learn lighting methods and relight to adjust the color of the facial image. Figure 17 Examples of multiple color-adjusted images are shown, Figure 17 The six color-adjusted images shown can be obtained by adjusting the color of the face image through deep learning relighting. Figure 17 (a) can be obtained by adding highlights to the face image; Figure 17 (b) can be obtained by adjusting the hue of the face image. Figure 17 The overall tone of (b) is reddish; Figure 17 (c) can be obtained by adding shadows to the face image; Figure 17 (d) can be obtained by reducing the contrast of the face image; Figure 17 (e) can be obtained by adjusting the hue of the face image, Figure 17 (e) has a greenish overall hue; Figure 17 (f) can be obtained by increasing the brightness of the face image.
[0127] In this embodiment, the color card encoder and the image information creator in the second-stage generation model are iteratively trained using the color card face data set. The color card encoder can gradually learn to fully encode the lighting environment information contained in the color card image, and the image information creator, based on the understanding of the lighting environment information through the first-stage training, gradually learns to fully inject the lighting environment information contained in the color card features output by the color card encoder into the image; moreover, by performing data transformation on each facial image in the color card face data set, the second-stage color card encoder and the second-stage image information creator are trained based on the color card face data set and the color adjustment images paired with each facial image, so that the color card encoder and the image information creator have the ability to adjust the lighting environment information of the facial image to the lighting environment information recorded by the color card image.
[0128] In an exemplary embodiment, the color card face data set further includes a third color card face data set, the third color card face data set including multiple sets of paired third color card images and fifth face images; the fifth face image is captured by a mobile terminal, and the fifth face image includes a sample face area and a sample color card area; the third color card image is extracted from the sample color card area in the paired fifth face image;
[0129] The above-mentioned steps of training the second-stage color card encoder and the second-stage image information creator based on the color card face data set and the color-adjusted images paired with each facial image to obtain the trained color card encoder and the trained image information creator include: obtaining an image generation model, the image generation model including a pre-trained facial image encoder, a second-stage color card encoder, a second-stage image information creator and a pre-trained image decoder; iteratively training the color card encoder and the image information creator in the image generation model based on the color card face data set and the color-adjusted images paired with the facial images in the color card face data set to obtain the trained color card encoder and the trained image information creator; wherein, in each round of iteration, based on the color card images in the color card face data set and the color-adjusted images paired with the facial images in the color card face data set, a third generated image corresponding to the facial image in the color card face data set is generated by the image generation model of the current round, and the color card encoder and image information creator in the image generation model of the current round are updated according to the difference between the facial image in the color card face data set and the corresponding third generated image.
[0130] Among them, the architectural diagram of the image generation model can be shown as follows: Figure 4 As shown. The third generated image corresponding to a face image in the color card face data set is generated based on the color card image paired with the face image in the color card face data set and the color adjustment image paired with the face image. The comparison example of the color adjustment image, the color card image and the third generated image can be shown as follows Figure 18 As shown, Figure 18 (a) can represent the color-adjusted image of a face image in the color card face image set after relighting, Figure 18 (b) may represent a color card image paired with the face image, Figure 18 (c) can be expressed based on Figure 18 (a) and Figure 18 (b), The third generated image generated by the image generation model.
[0131] The process of the image generation model outputting the third generated image during training corresponds to the process of the image generation model generating the target face image in the inference stage, wherein the color-adjusted image can correspond to the initial face image in the inference stage and be used to input into the face image encoder. The specific generation process of the third generated image will not be repeated here.
[0132] During the iterative training of the color card encoder and image information creator in the image generation model, the model parameters of the face image encoder and image decoder in the image generation model can be fixed, and the model parameters of the color card encoder and image information creator in the image generation model can be opened, among which the model parameters of the image information creator can only open the model parameters of the noise prediction network.
[0133] The noise prediction network, for example, can be a Unified Network (UNet) network, which can be fine-tuned using LoRA (Low-Rank Adaptation) technology. Specifically, the Transformer2DModel module in the UNet network can be trained. Modules other than the Transformer2DModel module in the UNet network are not trained, allowing the Transformer2DModel module to fully incorporate card embedding information during training. The color card encoder is also trained. The learning rate of the Transformer2DModel module can be 0.1 to 0.2 times the learning rate of the color card encoder.
[0134] In the process of training the color card encoder and image information creator of the image generation model, each face image in the color card face data set can be used as annotation data, and the color adjustment image corresponding to each face image is used as the input image. In this way, through the fine-tuning training at this stage, the image information creator can only adjust the color in the process of learning how to restore the color of the face image, avoid introducing transformations to non-color information in the face image, maintain the consistency of the face structure, and enable the generated face image to maintain the original structural information, defect information, etc. of the input face image, such as spots, wrinkles, and pores.
[0135] The difference between the facial images in the color card face dataset and their corresponding third generated images can be measured by calculating a loss value using a loss function. Since the fifth facial image in the third color card face dataset in the color card face dataset contains a sample facial region and a sample color card region, the third generated image corresponding to the fifth facial image also contains a facial region and a color card region. The supervised condition that the color of the color card region in the third generated image corresponding to the fifth facial image and the color of the sample color card region in the fifth facial image must remain unchanged can be used to further guide the color learning supervision of the image generation model.
[0136] For example, the average color value of each color block in the sample color card area of the fifth face image can be calculated and expressed as the value of the Lab channel (three channels of the Lab color model), which can be recorded as C k (L, a, b), where k = 1, 2, ..., N, which can represent the sequence number of the color block. There are N color blocks in total, such as 36, 60 or other. The average color prediction index of the color blocks in the third generated image generated by the image generation model is recorded as C x,k (L, a, b), the average color value C of the color block in the fifth face image corresponding to the third generated image y,k (L, a, b), the CIEDE2000 calculation formula (a color difference formula published by the International Commission on Illumination (CIE)) can be used to measure the error between two values. The calculation formula of the loss function is:
[0137] The Lab color model consists of three channels: one for lightness (L), and two for color (a and b). The L channel has a value range of (0, 100), and the a and b channels have a value range of (-128, +127). Colors in a range from dark green (low lightness) to gray (medium lightness) to bright pink (high lightness); colors in b range from bright blue (low lightness) to gray (medium lightness) to yellow (high lightness).
[0138] In each round of iteration, the back propagation technology is used to update the model parameters of the color card encoder in the current round of image generation model and the noise prediction network of the image information creator in the current round of image generation model, so as to update the color card encoder and image information creator in the current round of image generation model.
[0139] In this embodiment, the training process learns to restore the color-adjusted image to the corresponding facial image according to the color card image. In this way, the color card encoder can gradually learn how to fully encode the color card image in the scenario where the facial image is adjusted to the lighting environment of the color card image, and the image information creator can gradually learn how to fully inject the information of the color card image without changing the structural information of the face in the scenario where the facial image is adjusted to the lighting environment of the color card image, thereby further improving the model capability. Moreover, the color card face data set also includes a third color card face data set, and the fifth face image in the third color card face image set includes a sample face area and a sample color card area, so that more detailed image supervision can be introduced during the training process to improve the accuracy of model generation.
[0140] In a specific embodiment, the above image processing method may specifically include the following steps.
[0141] The server can obtain a graphic data set and a color card face data set. The graphic data set can include a first face data set FaceStdSet1, a second face data set PhoneFaceSet1, a standard color card data set ToneCardSet1, and a standard object data set StdObjectSet1. The color card face data set can include a first color card face data set CardStdFacePairSet1, a second color card face data set CardPhoneFacePairSet1, and a third color card face data set CardPhoneFaceOneSet1.
[0142] The server can train based on the image and text data set. Figure 15 The image information creator in the first stage generation model shown is obtained to obtain the image information creator of the first stage.
[0143] The server can obtain Figure 16 The second stage generation model shown, the image information creator in the obtained second stage generation model is the image information creator of the first stage.
[0144] The server can train the color card encoder and image information creator in the acquired second-stage generation model based on the first color card face data set and the second color card face data set in the color card face data set to obtain the second-stage color card encoder and the second-stage image information creator.
[0145] The server can obtain Figure 4 The image generation model shown in FIG, the color card encoder in the obtained image generation model is the color card encoder of the second stage, and the image information creator in the obtained image generation model is the image information creator of the second stage.
[0146] The server can train the acquired image generation model based on the color card face data set to obtain a trained image generation model, and the trained image generation model includes a trained color card encoder and a trained image information creator.
[0147] The server can obtain an initial facial image and a preconfigured color card image; perform color card feature extraction on the color card image through the color card encoder in the trained image generation model to obtain color card features; based on the initial facial image, perform feature extraction and then perform noise processing through the facial image encoder in the trained image generation model to obtain initial facial image features; perform iterative denoising processing based on the initial facial image features and color card features through the image information creator in the trained image generation model to generate target image features, and then perform decoding processing through the image decoder to adjust the color of the initial facial image based on the color card features to generate a target facial image under a preset lighting environment.
[0148] In order to solve the problems existing in online skin testing scenarios, traditional methods involve uniformly migrating the skin color of facial images captured by mobile terminals, or using a mobile terminal to capture a photo while the user is holding a standard color card. The face and color card are captured simultaneously under the same light source, and the color blocks in the color card are used to match the skin color of the face. However, the method of uniformly migrating the skin color of the initial facial image is not suitable for users of various skin colors, resulting in unstable color migration and the inability to truly reflect the skin color of different faces. Everyone's skin color becomes consistent, or the skin color is distorted, and different makeup looks cannot be recommended based on different skin colors. In the method where the user holds a standard color card, each user needs to be issued a standard color card and carry it with them for online skin testing, which is costly. In addition, under different light sources, the face has severe shadows or highlights, and the shadow removal or highlight removal algorithm is unstable, which will seriously affect the calculation of skin color, resulting in inconsistent matching of skin color and color blocks, resulting in deviation in the results. Through the above-mentioned image processing method, the user does not need to hold the color card. According to the color card image, the facial image captured by the user through the mobile terminal can be accurately aligned to the lighting environment of the skin tester.
[0149] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0150] Based on the same inventive concept, embodiments of the present application also provide an image processing device for implementing the aforementioned image processing method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more of the following image processing device embodiments can be found in the above-described limitations on the image processing method and will not be further elaborated here.
[0151] In an exemplary embodiment, Figure 19 As shown, an image processing device 1900 is provided, including: an acquisition module 1910 and an image generation module 1920, wherein:
[0152] Acquisition module 1910 is used to obtain an initial facial image; the initial facial image is obtained by photographing with a mobile terminal, and the initial facial image includes an initial facial area; a preconfigured color card image is obtained, and the color card image includes a skin color block obtained by photographing a preset color card with a skin tester under a preset lighting environment.
[0153] Image generation module 1920 is used to extract color card features from the color card image to obtain color card features, where the color card features include skin color block features of skin color blocks; generate an image based on the initial facial area and the color card features, adjust the color of the initial facial image based on the color card features, and generate a target facial image under a preset lighting environment.
[0154] In an exemplary embodiment, the image generation module 1920 may include an image generation model, which may include a color card encoder, a face image encoder, an image information creator, and an image decoder; the color card encoder is used to extract color card features from the color card image to obtain color card features; the face image encoder is used to perform noise processing after feature extraction based on the initial face image to obtain initial face image features; the image information creator is used to perform iterative denoising processing based on the initial face image features and the color card features to generate target image features; the image decoder is used to decode the target image features to adjust the color of the initial face image based on the color card features to generate a target face image under a preset lighting environment.
[0155] In an exemplary embodiment, the image processing device 1900 also includes a training module, which can be used to obtain a graphic data set, the graphic data set including a first face data set and a second face data set; the first face data set includes each first face image taken by a skin tester and each paired description text, the second face data set includes each second face image taken by a mobile terminal and each paired description text, the description text is used to describe the lighting environment in which the paired images are located; based on the graphic data set, the pre-trained image information creator is trained to obtain a first-stage image information creator; a color card face data set is obtained, the color card face data set includes a first color card face data set and a second color card face data set; the first color card face data set includes multiple groups of paired first color card images and third face images taken by a skin tester; the second color card face data set includes multiple groups of paired second color card images and fourth face images taken by a mobile terminal; based on the color card face data set, the pre-trained color card encoder and the first-stage image information creator are trained to obtain a trained color card encoder and a trained image information creator.
[0156] In an exemplary embodiment, the graphic data set also includes a standard color card data set and a standard object data set; the standard color card data set includes each standard color card image and each paired description text; the standard color card image contains a standard color block photographed by a skin tester on the standard color card; the standard object data set includes each standard object image photographed by a skin tester standard object and each paired description text; the reflectivity of the standard object is lower than that of the color card and the human face; the description text is used to describe the lighting environment of the paired image and the image category to which it belongs.
[0157] In an exemplary embodiment, the training module is also used to obtain a first-stage generative model, which includes a pre-trained text encoder, a pre-trained image information creator, and a pre-trained image decoder; based on the image-text data set, the image information creator in the first-stage generative model is iteratively trained to obtain the first-stage image information creator; wherein, in each round of iteration, based on each descriptive text in the image-text data set, the first generated image corresponding to each descriptive text is generated through the first-stage generative model of the current round, and the image information creator in the first-stage generative model of the current round is updated according to the difference between each image in the image-text data set and the first generated image corresponding to each paired descriptive text.
[0158] In an exemplary embodiment, the training module is further used to obtain a second-stage generative model, which includes a pre-trained color card encoder, a first-stage image information creator, and a pre-trained image decoder; based on the color card face data set, the color card encoder and image information creator in the second-stage generative model are iteratively trained to obtain the second-stage color card encoder and the second-stage image information creator, and in each round of iteration, based on each color card image in the color card face data set, the second-stage generative model of the current round is used to generate a second generated image corresponding to each color card image; according to the difference between the face image paired with each color card image and the corresponding second generated image, the color card encoder and image information creator in the second-stage generative model of the current round are updated; data transformation is performed on each face image to adjust the color of each face image to obtain a color-adjusted image paired with each face image; based on the color card face data set and the color-adjusted image paired with each face image, the second-stage color card encoder and the second-stage image information creator are trained to obtain a trained color card encoder and a trained image information creator.
[0159] In an exemplary embodiment, the color card face data set also includes a third color card face data set, and the third color card face data set includes multiple sets of paired third color card images and fifth face images; the fifth face image is obtained by taking a picture through a mobile terminal, and the fifth face image includes a sample face area and a sample color card area; the third color card image is extracted from the sample color card area in the paired fifth face image; the training module is also used to obtain an image generation model, the image generation model includes a pre-trained face image encoder, a second-stage color card encoder, a second-stage image information creator and a pre-trained image decoder; based on the color card face data set, and the color card face data The method comprises the following steps: using a color-adjusted image of a pair of face images in a color card face data set, iteratively training a color card encoder and an image information creator in an image generation model to obtain a trained color card encoder and a trained image information creator; wherein, in each round of iteration, based on the color card images in the color card face data set and the color-adjusted image of the pair of face images in the color card face data set, a third generated image corresponding to the face image in the color card face data set is generated through the image generation model of the current round, and according to the difference between the face image in the color card face data set and the respective corresponding third generated images, the color card encoder and the image information creator in the image generation model of the current round are updated.
[0160] Each module in the above-mentioned image processing device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0161] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 20 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data that needs to be stored when executing the above-mentioned image processing method. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an image processing method is implemented.
[0162] Those skilled in the art will understand that Figure 20 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0163] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0164] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0165] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0166] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0167] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile memory and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable logic unit (PLC), a data processing logic unit based on quantum computing, an artificial intelligence (AI) processor, and the like.
[0168] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0169] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire an initial face image; the initial face image is obtained by photographing with a mobile terminal, and the initial face image includes an initial face area; Obtaining a preconfigured color card image, wherein the color card image includes a skin color block obtained by photographing the preset color card with a skin tester under a preset lighting environment; Performing color card feature extraction on the color card image to obtain color card features, wherein the color card features include skin color block features of the skin color block; An image is generated based on the initial facial region and the color card features, so as to adjust the color of the initial facial image based on the color card features and generate a target facial image under the preset lighting environment.
2. The method according to claim 1, characterized in that The color card feature extraction is implemented by a color card encoder in a configured image generation model. The image generation model also includes a face image encoder, an image information creator, and an image decoder. The image generation is performed based on the initial face area and the color card features to adjust the color of the initial face image based on the color card features to generate a target face image under the preset lighting environment, including: Based on the initial facial image, extracting features through the facial image encoder and then performing noise processing to obtain initial facial image features; The image information creator performs iterative denoising based on the initial facial image features and the color card features to generate target image features. The image decoder then performs decoding processing to adjust the color of the initial facial image based on the color card features to generate the target facial image under the preset lighting environment.
3. The method according to claim 2, characterized in that The steps of training the color card encoder and the image information creator include: Obtaining a graphic data set, the graphic data set comprising a first face data set and a second face data set; the first face data set comprising first face images captured by a skin tester and their respective paired descriptive texts; the second face data set comprising second face images captured by a mobile terminal and their respective paired descriptive texts, the descriptive texts being used to describe the lighting environment in which the paired images exist; Based on the image and text data set, a pre-trained image information creator is trained to obtain a first-stage image information creator; Obtaining a color card face data set, the color card face data set comprising a first color card face data set and a second color card face data set; the first color card face data set comprising multiple sets of paired first color card images and third face images captured by a skin tester; and the second color card face data set comprising multiple sets of paired second color card images and fourth face images captured by a mobile terminal; Based on the color card face data set, the pre-trained color card encoder and the first-stage image information creator are trained to obtain a trained color card encoder and a trained image information creator.
4. The method according to claim 3, characterized in that The image and text data set also includes a standard color card data set and a standard object data set; The standard color card dataset includes each standard color card image and its paired description text; the standard color card image includes a standard color block photographed by a skin tester on the standard color card; The standard object dataset includes images of each standard object captured by a skin tester standard object and their corresponding description texts; the reflectivity of the standard object is lower than that of a color card and a human face; The description text is used to describe the lighting environment of the paired image and the image category to which it belongs.
5. The method according to claim 3, characterized in that The step of training a pre-trained image information creator based on the image and text data set to obtain a first-stage image information creator includes: Obtaining a first-stage generative model, the first-stage generative model comprising a pre-trained text encoder, a pre-trained image information creator, and a pre-trained image decoder; Based on the image-text data set, iteratively training the image information creator in the first-stage generation model to obtain the first-stage image information creator; In each round of iteration, based on each descriptive text in the graphic and text data set, the first generated image corresponding to each descriptive text is generated through the first stage generation model of the current round, and according to the difference between each image in the graphic and text data set and the first generated image corresponding to each paired descriptive text, the image information creator in the first stage generation model of the current round is updated.
6. The method according to claim 3, characterized in that The method of training the pre-trained color card encoder and the first-stage image information creator based on the color card face data set to obtain the trained color card encoder and the trained image information creator includes: Obtaining a second-stage generative model, the second-stage generative model comprising a pre-trained color card encoder, a first-stage image information creator, and a pre-trained image decoder; Iteratively training the color card encoder and the image information creator in the second-stage generative model based on the color card face dataset to obtain a second-stage color card encoder and a second-stage image information creator, and in each iteration, based on each color card image in the color card face dataset, generating a second generated image corresponding to each color card image through the second-stage generative model of the current round, and updating the color card encoder and the image information creator in the second-stage generative model of the current round based on the difference between the face image paired with each color card image and the second generated image corresponding to each color card image; Performing data transformation on each facial image to adjust the color of each facial image to obtain a color-adjusted image paired with each facial image; Based on the color card face data set and the color adjustment images paired with each face image, the second-stage color card encoder and the second-stage image information creator are trained to obtain a trained color card encoder and a trained image information creator.
7. The method according to claim 6, characterized in that The color card face data set also includes a third color card face data set, the third color card face data set including multiple sets of paired third color card images and fifth face images; the fifth face image is captured by a mobile terminal, and the fifth face image includes a sample face area and a sample color card area; the third color card image is extracted from the sample color card area in the paired fifth face image; The method of training the second-stage color card encoder and the second-stage image information creator based on the color card face data set and the color-adjusted images paired with the respective face images to obtain the trained color card encoder and the trained image information creator comprises: Acquire an image generation model, wherein the image generation model includes a pre-trained face image encoder, a color card encoder of the second stage, an image information creator of the second stage, and a pre-trained image decoder; Iteratively training a color card encoder and an image information creator in the image generation model based on the color card face dataset and color-adjusted images of paired face images in the color card face dataset to obtain a trained color card encoder and a trained image information creator; In each round of iteration, based on the color card images in the color card face data set and the color adjustment images paired with the face images in the color card face data set, the image generation model of the current round is used to generate a third generated image corresponding to the face image in the color card face data set. According to the difference between the face image in the color card face data set and the respective corresponding third generated images, the color card encoder and image information creator in the image generation model of the current round are updated.
8. An image processing device, characterized in that: The device comprises: An acquisition module is configured to acquire an initial facial image; the initial facial image is obtained by photographing the image with a mobile terminal, the initial facial image including an initial facial region; and acquire a pre-configured color card image, the color card image including a skin color block under a preset lighting environment obtained by photographing the preset color card with a skin tester. An image generation module is configured to extract color card features from the color card image to obtain color card features, wherein the color card features include skin color block features of the skin color block; generate an image based on the initial facial area and the color card features, adjust the color of the initial facial image based on the color card features, and generate a target facial image under the preset lighting environment.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.