Virtual staining method, device and electronic device
By obtaining the mask image of the area to be dyed in the video to be processed, the brightness is extracted and fused with the color to be dyed, and combining the mask image and the fusion image, the problem of not being able to adapt to different hair colors in the prior art is solved, and real-time and natural hair dyeing effect is achieved.
Patent Information
- Application Number
- CN202210730722.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-06-24
AI Technical Summary
The existing virtual hair dyeing methods cannot meet the needs of different original hair colors, especially yellow or silver-gray hair colors, and the gradient reconstruction method is time-consuming and difficult to achieve real-time and natural hair dyeing effects.
By obtaining the masked image of the area to be dyed in the video to be processed, the brightness is extracted and fused with the color to be dyed, combining the masked image and the fusion image, the post-stained image is determined, and the image segmentation model is trained using model distillation and timing consistency constraints to improve the stability and accuracy of the segmented area.
It achieves natural hair dyeing effects that adapt to different original hair colors, can handle hair dyeing needs in the video in real time, and improves the stability and accuracy of hair dyeing effects.
Smart Images

Figure CN114926617B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to a virtual dyeing method, apparatus, and electronic device. Background Art
[0002] In various video publishing, video live streaming and other platforms, virtual hair dyeing special effects are a commonly used function. Applying virtual hair dyeing special effects can dye the hair of the characters in the video into the desired color to meet the personalized needs of users.
[0003] Existing virtual hair dyeing methods usually generate color lookup tables before and after color replacement through the method of Look Up Table (LUT), and directly replace colors based on the color lookup table during use. This method cannot adapt to different original hair colors. For hair colors that are not dark, such as yellow or silver - gray, the expected coloring effect cannot be achieved. Another method is through gradient reconstruction. This method does not use the information of the hair area, making the entire hair dyeing effect unrealistic, and the gradient calculation and fusion are time - consuming, making it difficult to achieve real - time hair dyeing effects. Summary of the Invention
[0004] Embodiments of this application provide a virtual dyeing method, apparatus, and electronic device to meet the hair dyeing needs of users with natural effects and compatible with different original hair colors.
[0005] In a first aspect, embodiments of this application provide a virtual dyeing method, including:
[0006] Obtain a mask image of the area to be dyed in the image to be processed in the video to be processed;
[0007] Obtain the brightness of the area to be dyed, and perform a fusion process on the brightness and the color to be dyed to obtain a fused image of the area to be dyed;
[0008] Based on the image to be processed, the mask image, and the fused image, determine the dyed image as the image frame of the processed video.
[0009] In a second aspect, embodiments of this application provide a virtual dyeing apparatus, including:
[0010] An obtaining module, configured to obtain a mask image of the area to be dyed in the image to be processed in the video to be processed;
[0011] A fusion module, configured to obtain the brightness of the area to be dyed, and perform a fusion process on the brightness and the color to be dyed to obtain a fused image of the area to be dyed;
[0012] A determining module, configured to determine the dyed image as the image frame of the processed video based on the image to be processed, the mask image, and the fused image.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the method provided in any embodiment of the present application is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method provided in any embodiment of the present application is implemented.
[0015] Compared with the prior art, the present application has the following advantages:
[0016] The virtual coloring method, device, and electronic device provided in the embodiments of the present application obtain a mask image of a to-be-colored area of a to-be-processed image in a to-be-processed video; obtain the brightness of the to-be-colored area, fuse the brightness and the to-be-colored color to obtain a fused image of the to-be-colored area; and determine a colored image based on the to-be-processed image, the mask image, and the fused image as an image frame of the processed video. In the technical solution of the present application, after fusing the brightness of the to-be-colored area of the original image with the to-be-colored color, the colored image is obtained according to the original image, the mask image of the to-be-colored area of the original image, and the fused image, which can adapt to different original hair colors and can achieve a real-time and natural hair coloring effect.
[0017] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the present application will be readily apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments disclosed in the present application and should not be regarded as limiting the scope of the present application.
[0019] Figure 1 It is a flowchart of a virtual coloring method provided in an embodiment of the present application;
[0020] Figure 2 It is a schematic diagram of training an image segmentation model by a temporal consistency method provided in an embodiment of the present application;
[0021] Figure 3 It is a schematic diagram of a virtual coloring system provided in an embodiment of the present application;
[0022] Figure 4 It is a schematic diagram of a virtual coloring device provided in an embodiment of the present application;
[0023] Figure 5 It is a block diagram of an electronic device for implementing the embodiments of the present application. Detailed implementation manners
[0024] In the following, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present application. Therefore, the drawings and descriptions are considered to be exemplary in nature rather than restrictive.
[0025] To facilitate understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.
[0026] To more clearly show the virtual staining method provided in the embodiments of the present application, the application scenarios that can be used to implement this method are introduced first.
[0027] The technical solution of the present application can be applied to the scenario of performing staining processing on regions in image frames of a video. For example, when recording a video or video live broadcast, according to the staining instruction input by the user, the hair region of the person in the image frame of the video is stained. In addition, it can also be applied to the scenario of staining regions in images of non-video image frames.
[0028] The embodiments of the present application provide a virtual staining method. Figure 1 It is a flowchart of the virtual staining method of an embodiment of the present application. This method can be applied to a virtual staining device, and this device can be deployed in a user terminal, a server, or other processing devices. Among them, the user terminal can be a terminal device such as a user's personal computer, tablet computer, mobile phone, etc. In some possible implementation manners, this method can also be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 shown, this method includes:
[0029] Step S101, obtain a mask image of the region to be stained in the image to be processed in the video to be processed.
[0030] In this embodiment, taking the execution entity as the user terminal as an example for introduction, the processing can be performed by the graphics processing unit (GPU) of the user terminal. When the user is shooting a video or conducting a live broadcast, the area to be dyed and the color to be dyed can be determined according to the pre-configured dyeing special effects. For example, hair dyeing, red. When shooting a video or conducting a live broadcast, the user terminal segments the hair area in the video image frame through image segmentation. It should be noted that virtual hair dyeing can be performed during the video shooting process, or virtual hair dyeing can be performed after the video shooting. The embodiments of the present application do not limit this.
[0031] Among them, the pixel value of each pixel in the mask image of the area to be dyed can be 0 or 1. 0 indicates that the pixel is not a pixel in the area to be dyed, and 1 indicates that the pixel is a pixel in the area to be dyed. Optionally, the pixel value of each pixel in the mask image of the area to be dyed can also be the probability value that the area corresponding to the pixel is the area to be dyed.
[0032] Step S102, obtain the brightness of the area to be dyed, and perform a fusion process on the brightness and the color to be dyed to obtain a fused image of the area to be dyed.
[0033] Separate the color and brightness of the image to be processed, and extract the brightness information. Optionally, process the image to be processed into a grayscale image, and the grayscale value of the pixels in the area to be dyed is the brightness corresponding to each pixel. The color to be dyed can be a color determined according to the user's selection instruction. Using a fusion algorithm to fuse the brightness of the area to be dyed and the color to be dyed can obtain a fused image.
[0034] Step S103, based on the image to be processed, the mask image, and the fused image, determine the dyed image as the image frame of the processed video.
[0035] According to the pixel values of the pixels in the area to be dyed in the image to be processed, the mask image of the area to be dyed, and the fused image, the pixel values of the dyed image can be obtained, and the dyed image can be used as the image frame of the video.
[0036] The virtual dyeing method provided by the embodiments of the present application obtains the mask image of the area to be dyed of the image to be processed in the video to be processed; obtains the brightness of the area to be dyed, and performs a fusion process on the brightness and the color to be dyed to obtain a fused image of the area to be dyed; based on the image to be processed, the mask image, and the fused image, determine the dyed image as the image frame of the processed video. In the embodiments of the present application, after fusing the brightness of the area to be dyed in the original image with the color to be dyed, and then obtaining the dyed image according to the original image, the mask image of the area to be dyed in the original image, and the fused image, it can adapt to different original hair colors and can achieve a real-time and natural hair dyeing effect.
[0037] Among them, for step S101, the specific implementation method for obtaining the mask image of the area to be dyed is shown in the following embodiments:
[0038] In a possible implementation, obtaining the mask image of the area to be dyed in the image to be processed in the video to be processed includes:
[0039] Input the image to be processed into an image segmentation model, and the image segmentation model outputs the mask image of the area to be dyed; the image segmentation model is trained through model distillation and / or temporal consistency constraint; the model distillation is used to simplify the model structure; the temporal consistency constraint is used to improve the correlation of the areas to be dyed in adjacent frames of the images to be processed.
[0040] In practical applications, the image segmentation model can be a deep learning model such as a convolutional neural network model. When training the image segmentation model, it can be trained by at least one of model distillation or temporal consistency constraint. The model distillation is to supervise a deep learning model with a relatively simple structure through the training of a deep learning model with a more complex structure. Usually, the output result of the deep learning model with a more complex structure is more accurate, but the calculation time is longer. In contrast, the accuracy of the output result of the deep learning model with a relatively simple structure is lower, but the calculation time is shorter. The image segmentation model trained by model distillation has a relatively simple structure and short calculation time, can meet the real-time image segmentation on the user terminal, and has a relatively high accuracy of the output result.
[0041] In addition, when the image segmentation model segments the areas to be dyed in adjacent frame images, it is possible that due to the position change of the areas to be dyed, the stability of the segmented areas to be dyed is poor, or it is possible that due to the difference in noise and the difference in light transformation between adjacent frame images, the segmentation result is not accurate enough. When training the image segmentation model, through temporal consistency constraint, the correlation of the segmentation of the areas to be dyed in adjacent image frames in the video by the image segmentation model can be improved, thereby improving the stability and accuracy of the segmentation of the areas to be dyed in adjacent frame images.
[0042] Among them, the specific model distillation method is shown in the following embodiments:
[0043] In a possible implementation, training the image segmentation model through model distillation includes:
[0044] Training an initial first image segmentation model with a first sample image pair to obtain a first image segmentation model; the first sample image pair includes a first sample image and the mask image of the area to be dyed in the first sample image; obtaining an image segmentation model based on the first image segmentation model and the first sample image pair; the number of neurons in the first image segmentation model is greater than that of the image segmentation model.
[0045] Among them, compared with the image segmentation model, the first image segmentation model can be a model with a relatively complex structure. The number of neurons in the first image segmentation model is greater than that of the image segmentation model. For example, the first image segmentation model can be a high resolution network (HRNet), etc. The image segmentation model can be a model with a relatively simple structure. For example, it can be a Unet network using the mobilenet-v3 network, etc.
[0046] Use the first sample image to train the first image segmentation model to obtain a model with a relatively complex structure, and then use the first sample image as part of the training samples to train the image segmentation model, that is, a model with a relatively simple structure. Deploying the image segmentation model on the user terminal can meet the user terminal's demand for real-time image processing.
[0047] Among them, obtaining a model with a relatively simple structure based on a model with a relatively complex structure, the specific implementation method is shown in the following embodiments:
[0048] In a possible implementation manner, obtaining the image segmentation model based on the first image segmentation model and the first sample image pair includes:
[0049] Input the second sample image into the first image segmentation model to obtain a predicted mask image of the area to be stained corresponding to the second sample image; the second sample image and the predicted mask image of the area to be stained corresponding to the second sample image form a second sample image pair; use multiple first sample image pairs and multiple second sample image pairs to train the second image segmentation model to obtain the image segmentation model.
[0050] In practical applications, the first sample image pair includes a mask image of the area to be stained in the first sample image. It can be understood that the ground truth of the area to be stained is obtained through the annotation method. Inputting the first sample image pair into the second image segmentation model can obtain a predicted mask image of the area to be stained in the first sample image, that is, the predicted value of the area to be stained. The loss function corresponding to the first sample image can be calculated through the mask image of the area to be stained in the first sample image and the predicted mask image of the area to be stained in the first sample image. Among them, the second image segmentation model can be the initial model of the image segmentation model, which can be a model with a simpler structure than the first image segmentation model. After iterative training, the final image segmentation model is obtained.
[0051] Among them, the second sample image is a sample of a mask image without a corresponding area to be stained, that is, the ground truth of the area without the area to be stained. Then, through the first image segmentation model, the second image segmentation model is supervised. The second sample image is input into the first image segmentation model, and the predicted mask image of the area to be stained corresponding to the second sample image is obtained as the ground truth of the area to be stained corresponding to the second sample image. During the training process of the second image segmentation model, the second sample image is input into the second image segmentation model. Through the second image segmentation model, the predicted mask image of the area to be stained corresponding to the second sample image can be obtained. Using the predicted mask image of the area to be stained corresponding to the second sample image obtained through the first image segmentation model and the predicted mask image of the area to be stained corresponding to the second sample image obtained through the second image segmentation model, the loss function corresponding to the second sample image is calculated, and iterative training is performed to obtain the final image segmentation model.
[0052] In this embodiment, multiple pairs of first sample images and multiple pairs of second sample images are used to train the image segmentation model, making the samples more abundant. Moreover, there is also the first image segmentation model for supervision. In this way, the output result of the trained image segmentation model is more accurate when used. Moreover, the image segmentation model has a simpler structure and faster data processing than the first image segmentation model, and can meet the requirement of real-time image segmentation processing on the user terminal.
[0053] In addition, the image segmentation model can be further trained by a temporal consistency constraint method. See the following embodiments for details:
[0054] In a possible implementation, training the image segmentation model by a temporal consistency constraint method includes:
[0055] Training a third image segmentation model using multiple pairs of third sample images to obtain an image segmentation model; among them, the third sample image pair includes a third sample image and a processed image corresponding to the third sample image; the processed image is obtained by performing at least one of affine transformation of pixels, adding random noise, and adding random illumination to the third sample image; the loss function of the image segmentation model is determined using the predicted mask image of the area to be stained corresponding to the third sample image and the predicted mask image of the area to be stained corresponding to the processed image of the third sample image.
[0056] Among them, the third image segmentation model can be the second image segmentation model or a model trained based on the second image segmentation model. The third sample image segmentation model is used as the initial model of the image segmentation model and continues to be trained to obtain the final image segmentation model. The third sample image pair includes a third sample image and a processed image corresponding to the third sample image; the processed image is obtained by performing at least one of affine transformation of pixels, adding random noise, and adding random illumination on the third sample image. The affine transformation may include position transformation in the horizontal or vertical direction of the pixels, or size transformation such as stretching or compression, etc.
[0057] Optionally, the third sample image and the processed image obtained by performing affine transformation of pixels on the third sample image are used as a training sample pair to train the image segmentation model. The third sample image is input into the third image segmentation model to obtain a predicted mask image of the area to be stained corresponding to the third sample image, that is, the predicted value corresponding to the third sample image; the processed image is input into the third image segmentation model to obtain a predicted mask image of the area to be stained corresponding to the processed image, that is, the predicted value corresponding to the processed image. When calculating the loss function, the predicted value corresponding to the processed image is used as the true value corresponding to the third sample image, and the loss function corresponding to the third sample image is calculated using the predicted value corresponding to the third sample image and the predicted value corresponding to the processed image. The predicted value corresponding to the third sample image is used as the true value corresponding to the processed image, and the loss function corresponding to the processed image is calculated using the predicted value corresponding to the third sample image and the predicted value corresponding to the processed image. In this embodiment, the image segmentation model trained by the training method of temporal consistency can improve the stability and accuracy of the segmentation of the area to be stained in adjacent frame images when performing image segmentation.
[0058] Next, a specific embodiment is used to introduce the training of the image segmentation model by the temporal consistency constraint method. Figure 2 It is a schematic diagram of training an image segmentation model by the temporal consistency method in an embodiment of the present application. As Figure 2 shown, in this embodiment, the third sample image is a portrait image, which is input into the initial image segmentation network (Network shown in the figure) to obtain a predicted mask image of the portrait hair area; random illumination is added to the third sample image and affine transformation is performed to obtain a processed image; the inverse transformation of the affine transformation is performed on the predicted mask image of the portrait hair area to obtain the transformed predicted mask image of the portrait hair area; the processed image is input into the initial image segmentation network (Network shown in the figure) to obtain a predicted mask image of the portrait hair area of the processed image; the transformed predicted mask image of the portrait hair area and the predicted mask image of the portrait hair area of the processed image are used for temporal consistency constraint training (consistency loss constraint shown in the figure) to obtain the final image segmentation model.
[0059] Among them, after the image segmentation model segments the area to be dyed, it can be further optimized according to the specific application scenario. For specific examples, see the following embodiments:
[0060] In a possible implementation, the image to be processed is a portrait image, and the area to be dyed is the portrait hair area; the method further includes: obtaining a mask image of the skin area in the portrait image; fusing the mask image of the skin area and the mask image of the portrait hair area to obtain a processed mask image of the portrait hair area.
[0061] In practical applications, when the image to be processed is a portrait image and the area to be dyed is the portrait hair area, for the mask image of the portrait hair area output by the image segmentation model, it can be optimized by the mask image of the skin area. By obtaining the mask image of the skin area and removing the area detected as skin from the portrait hair area, the front forehead edge of the portrait hair area can be optimized, and an accurate portrait hair area can be obtained.
[0062] In an example, the skin area and the portrait hair area can be fused through the following formula:
[0063] mask = mask_seg * (1.0 - mask_skin) (1)
[0064] Among them, mask represents the pixel value of the fused image; mask_seg represents the pixel value of the mask image of the portrait hair area segmented by the image segmentation model; mask_skin represents the pixel value of the mask image of the skin area.
[0065] It should be noted that the skin area and the portrait hair area can also be fused in other ways, and the embodiments of the present application do not limit this.
[0066] Among them, the implementation method for obtaining the mask image of the skin area is shown in the following embodiments:
[0067] In a possible implementation, obtaining the mask image of the skin area in the portrait image includes:
[0068] Based on the skin detection lookup table, determine the mask image of the skin area; the skin detection lookup table is determined based on the pixel values of the pixels corresponding to the skin area and the pixel values of the pixels corresponding to the non-skin area.
[0069] In practical applications, a mask image of the skin area can be determined through a skin detection lookup table, which is obtained based on large-scale skin / non-skin pixel statistics in the RGB space. The skin detection lookup table includes multiple pixel values, and whether each pixel value is a pixel in the skin area, or the probability that the pixel value is a pixel in the skin area. According to the skin detection lookup table, the skin area in the image to be processed can be determined, and a mask image of the skin area can be obtained. In this embodiment, by using the skin detection lookup table to determine the mask image of the skin area, the calculation process is simple and the result is accurate.
[0070] After obtaining the mask image and the fused image of the area to be dyed, how to obtain the dyed image is specifically shown in the following embodiments:
[0071] In a possible implementation manner, based on the image to be processed, the mask image, and the fused image, determining the dyed image as the image frame of the processed video includes:
[0072] Fuse the image to be processed, the mask image, and the fused image to obtain the dyed image as the image frame of the processed video.
[0073] In practical applications, by fusing the image to be processed, the mask image, and the fused image, a dyed image can be obtained, and there are various fusion methods.
[0074] In one example, the following formula can be used for fusion processing:
[0075]
[0076] result = frame * (1 - mask) + mask * merge (3)
[0077] Where merge represents the pixel value of the fused image of the area to be dyed obtained by fusing the brightness and the color to be dyed; gray represents the gray value of the pixel of the image to be processed; new_color represents the pixel value of the color to be dyed; result represents the pixel value of the dyed image; frame represents the pixel value of the image to be processed; mask represents the pixel value of the mask image of the area to be dyed.
[0078] Optionally, before fusing the image to be processed, the mask image, and the fused image, perform Gaussian filtering on the extracted mask image of the area to be dyed to obtain a smoothed mask image of the area to be dyed, so that the edge of the dyed area is smoother after dyeing and the visual effect is better.
[0079] Figure 3 This is a schematic diagram of a virtual dyeing system according to an embodiment of the present application. As Figure 3As shown in the figure, in this embodiment, the image to be processed is a portrait image, and the area to be dyed is the hair area of the portrait. The virtual dyeing system includes a hair area extraction module and a hair dyeing fusion module. First, the mask image of the portrait hair area is extracted by the hair area extraction module. Specifically, for hair area segmentation, the portrait image is input into an image segmentation model to obtain the mask image of the portrait hair area; the portrait image is subjected to skin detection, and through a skin detection lookup table, the mask image of the skin area is obtained; the mask image of the portrait hair area and the mask image of the skin area are fused to obtain an optimized image of the mask image of the portrait hair area. Secondly, through the hair dyeing fusion module, fusion processing is performed to obtain the image after hair dyeing. Specifically, the brightness of the portrait hair area is extracted, and the new hair color to be dyed is determined. The brightness and the new hair color are fused to obtain a fused image, and then the portrait image, the optimized image, and the fused image are fused again to obtain the image after hair dyeing, and the image of the person with the new hair color is displayed in the video live broadcast.
[0080] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application further provides a virtual dyeing device. As Figure 4 shown, the virtual dyeing device may include:
[0081] An acquisition module 401, configured to acquire a mask image of an area to be dyed in an image to be processed in a video to be processed.
[0082] A fusion module 402, configured to acquire the brightness of the area to be dyed, and perform fusion processing on the brightness and the color to be dyed to obtain a fused image of the area to be dyed.
[0083] A determination module 403, configured to determine a dyed image based on the image to be processed, the mask image, and the fused image, and use it as an image frame of the processed video.
[0084] The virtual dyeing device provided in the embodiment of the present application acquires a mask image of an area to be dyed in an image to be processed in a video to be processed; acquires the brightness of the area to be dyed, and performs fusion processing on the brightness and the color to be dyed to obtain a fused image of the area to be dyed; based on the image to be processed, the mask image, and the fused image, determines a dyed image and uses it as an image frame of the processed video. In the embodiment of the present application, after fusing the brightness of the area to be dyed in the original image with the color to be dyed, and then obtaining the dyed image according to the original image, the mask image of the area to be dyed in the original image, and the fused image, it can adapt to different original hair colors and can achieve a real-time and natural hair dyeing effect.
[0085] In a possible implementation manner, the acquisition module 401 is configured to:
[0086] Input the image to be processed into an image segmentation model, and the image segmentation model outputs a mask image of the area to be dyed; the image segmentation model is trained by means of model distillation and / or temporal consistency constraint; the model distillation method is used to simplify the model structure; the temporal consistency constraint method is used to improve the correlation of the areas to be dyed in adjacent frames of the images to be processed.
[0087] In a possible implementation, the image to be processed is a portrait image, and the area to be dyed is the portrait hair area; the acquisition module 401 further includes an acquisition unit and a fusion unit:
[0088] The acquisition unit is used to acquire a mask image of the skin area in the portrait image;
[0089] The fusion unit is used to fuse the mask image of the skin area and the mask image of the portrait hair area to obtain a processed mask image of the portrait hair area.
[0090] In a possible implementation, the acquisition unit is specifically used for:
[0091] Based on the skin detection look-up table, determine the mask image of the skin area; the skin detection look-up table is determined based on the pixel values of the pixels corresponding to the skin area and the pixel values of the pixels corresponding to the non-skin area.
[0092] In a possible implementation, the determination module 403 is used for:
[0093] Fuse the image to be processed, the mask image and the fused image to obtain a dyed image as the image frame of the processed video.
[0094] In a possible implementation, training the image segmentation model by means of model distillation includes:
[0095] Use the first sample image pair to train the initial first image segmentation model to obtain the first image segmentation model; the first sample image pair includes the first sample image and the mask image of the area to be dyed in the first sample image;
[0096] Based on the first image segmentation model and the first sample image pair, obtain the image segmentation model; the number of neurons in the first image segmentation model is greater than that of the image segmentation model.
[0097] In a possible implementation, obtaining the image segmentation model based on the first image segmentation model and the first sample image pair includes:
[0098] Input the second sample image into the first image segmentation model to obtain a predicted mask image of the area to be dyed corresponding to the second sample image; the second sample image and the predicted mask image of the area to be dyed corresponding to the second sample image form a second sample image pair;
[0099] The second image segmentation model is trained using multiple first sample image pairs and multiple second sample image pairs to obtain an image segmentation model.
[0100] In a possible implementation, the image segmentation model is trained by a temporal consistency constraint method, including:
[0101] The third image segmentation model is trained using multiple third sample image pairs to obtain an image segmentation model;
[0102] Wherein, the third sample image pair includes a third sample image and a processed image corresponding to the third sample image; the processed image is obtained by performing at least one of affine transformation of pixels, adding random noise, and adding random illumination to the third sample image; the loss function of the image segmentation model is determined using the predicted mask image of the area to be stained corresponding to the third sample image and the predicted mask image of the area to be stained corresponding to the processed image of the third sample image.
[0103] For the functions of the modules in each device of the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, and they have the corresponding beneficial effects, which will not be elaborated here.
[0104] Figure 5 It is a block diagram of an electronic device for implementing the embodiments of the present application. As Figure 5 shown, the electronic device includes: a memory 510 and a processor 520. The memory 510 stores a computer program that can run on the processor 520. When the processor 520 executes the computer program, the method in the above embodiments is implemented. The number of the memory 510 and the processor 520 can be one or more.
[0105] The electronic device further includes:
[0106] A communication interface 530, configured to communicate with external devices and perform data interaction and transmission.
[0107] If the memory 510, the processor 520, and the communication interface 530 are independently implemented, the memory 510, the processor 520, and the communication interface 530 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 5It is represented only by a thick line, but it does not mean that there is only one bus or one type of bus.
[0108] Optionally, in a specific implementation, if the memory 510, the processor 520, and the communication interface 530 are integrated on a chip, the memory 510, the processor 520, and the communication interface 530 can communicate with each other through an internal interface.
[0109] The embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, the method provided in the embodiment of the present application is implemented.
[0110] The embodiment of the present application also provides a chip, which includes a processor for calling and running instructions stored in a memory, so that a communication device installed with the chip executes the method provided in the embodiment of the present application.
[0111] The embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiment of the application.
[0112] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the advanced reduced instruction set machine (ARM) architecture.
[0113] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may further include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0114] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0115] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0116] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means two or more, unless otherwise specifically defined.
[0117] Any process or method description represented in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed.
[0118] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices.
[0119] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by a program instructing relevant hardware, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiment.
[0120] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.
[0121] As mentioned above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions thereof, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A virtual staining method, characterized in that, The method includes: Obtaining a mask image of a to-be-stained area of a to-be-processed image in a to-be-processed video. When the to-be-processed image is a portrait image and the to-be-stained area is a portrait hair area, obtaining a mask image of a skin area in the portrait image, and performing a fusion process on the mask image of the skin area and the mask image of the portrait hair area to obtain the processed mask image of the portrait hair area. The mask image of the skin area is used to exclude the skin area from the portrait hair area; Obtaining the brightness of the to-be-stained area, and performing a fusion process on the brightness and a to-be-dyed color to obtain a fusion image of the to-be-stained area; Based on the to-be-processed image, the mask image, and the fusion image, determining a stained image as an image frame of the processed video.
2. The method according to claim 1, wherein The obtaining of the mask image of the to-be-stained area of the to-be-processed image in the to-be-processed video includes: Inputting the to-be-processed image into an image segmentation model, and the image segmentation model outputs the mask image of the to-be-stained area; the image segmentation model is trained by a model distillation method and / or a temporal consistency constraint method; the model distillation method is used to simplify the model structure; the temporal consistency constraint method is used to improve the correlation of the to-be-stained areas of adjacent frames of the to-be-processed images.
3. The method according to claim 1, wherein The obtaining of the mask image of the skin area in the portrait image includes: Based on a skin detection lookup table, determining the mask image of the skin area; the skin detection lookup table is determined based on the pixel values of the pixels corresponding to the skin area and the pixel values of the pixels corresponding to the non-skin area.
4. The method according to claim 1, characterized in that, The determining of the stained image as an image frame of the processed video based on the to-be-processed image, the mask image, and the fusion image includes: Performing a fusion process on the to-be-processed image, the mask image, and the fusion image to obtain the stained image as an image frame of the processed video.
5. The method according to claim 2, wherein Training the image segmentation model by the model distillation method includes: Training an initial first image segmentation model with a first sample image pair to obtain a first image segmentation model; the first sample image pair includes a first sample image and the mask image of the to-be-stained area in the first sample image; Based on the first image segmentation model and the first sample image pair, obtaining the image segmentation model; the number of neurons in the first image segmentation model is greater than that of the image segmentation model.
6. The method according to claim 5, wherein The obtaining of the image segmentation model based on the first image segmentation model and the first sample image pair includes: Inputting a second sample image into the first image segmentation model to obtain a predicted mask image of the to-be-stained area corresponding to the second sample image; the second sample image and the predicted mask image of the to-be-stained area corresponding to the second sample image form a second sample image pair; Training a second image segmentation model with multiple first sample image pairs and multiple second sample image pairs to obtain the image segmentation model.
7. The method according to claim 2, wherein Training the image segmentation model by the temporal consistency constraint method includes: Training a third image segmentation model using multiple third sample images to obtain the image segmentation model; Wherein, the third sample image pair includes a third sample image and a processed image corresponding to the third sample image; the processed image is obtained by performing at least one of affine transformation of pixels, adding random noise, and adding random illumination to the third sample image; the loss function of the image segmentation model is determined using the predicted mask image of the area to be stained corresponding to the third sample image and the predicted mask image of the area to be stained of the processed image corresponding to the third sample image.
8. A virtual dyeing device, characterized in that, The apparatus includes: An acquisition module, configured to acquire a mask image of an area to be stained in a to-be-processed image in a to-be-processed video. When the to-be-processed image is a portrait image and the area to be stained is a portrait hair area, acquire a mask image of the skin area in the portrait image, and perform a fusion process on the mask image of the skin area and the mask image of the portrait hair area to obtain a processed mask image of the portrait hair area. The mask image of the skin area is used to remove the skin area from the portrait hair area; A fusion module, configured to acquire the brightness of the area to be stained, and perform a fusion process on the brightness and the color to be stained to obtain a fused image of the area to be stained; A determination module, configured to determine a stained image based on the to-be-processed image, the mask image, and the fused image, and use it as an image frame of the processed video.
9. An electronic device, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
A virtual hair dyeing method based on image semantic segmentation
CN109903257A
Virtual hair dyeing method and device, electronic equipment and storage medium
CN113450431A
Skin color segmentation method, device, electronic equipment and storage medium
CN113888543A