Method for generating images with metallic texture and method for training models

By designing a metal texture image generation model, the problem of high computational complexity of the gold rendering method based on normal phase estimation was solved, the real-time performance of gold portrait rendering was improved, and the video rendering lag was reduced.

CN114241387BActive Publication Date: 2025-10-03FACE CUTE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111580389.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-10-03
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

The existing gold rendering method based on normal phase estimation has a large amount of calculation, resulting in poor real-time performance of gold rendering of portraits.

Method used

A metal texture image generation model is designed. The first video is input into a pre-trained metal texture image generation model through a training model to generate a second video with metal texture. This avoids the normal phase estimation step, reduces the amount of calculation, and improves real-time performance.

Benefits of technology

The real-time performance of golden portrait rendering has been improved, the lag in video rendering has been reduced, and processing efficiency has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114241387B_ABST
    Figure CN114241387B_ABST
Patent Text Reader

Abstract

This application provides a method for generating an image with a metallic texture and a model training method. The method comprises: obtaining a first video; inputting the first video into a pre-trained metallic texture image generation model to obtain a second video, wherein each frame image in the second video is an image with a metallic texture; wherein the metallic texture image generation model is trained based on multiple first sample images and second sample images with metallic texture corresponding to each first sample image. The method for generating an image with a metallic texture and the model training method provided in this application are used to improve the real-time performance of obtaining the second video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method for generating an image with a metallic texture and a method for training a model. Background Art

[0002] Metallization of portraits refers to giving a metallic texture to the portrait, making it look like a metal sculpture.

[0003] In related technologies, in order to realize the metallic materialization of human portraits, a gold rendering method based on normal estimation is usually adopted. The gold rendering method based on normal estimation first needs to perform normal estimation on the human portrait, and then perform gold rendering on the human portrait based on the estimated normal phase.

[0004] In the above-mentioned related technologies, the algorithm of the gold rendering method based on normal phase estimation has a large amount of calculation, resulting in poor real-time performance of gold rendering of portraits. Summary of the Invention

[0005] The present application provides a method for generating an image with a metallic texture and a method for training a model to solve the problem of poor real-time performance in rendering gold portraits.

[0006] In a first aspect, the present application provides a method for generating an image with a metallic texture, comprising:

[0007] Get the first video;

[0008] Inputting the first video into a pre-trained metal texture image generation model to obtain a second video, wherein each frame image in the second video is an image with a metal texture;

[0009] The metal texture image generation model is trained based on a plurality of first sample images and second sample images with metal texture corresponding to each first sample image.

[0010] Optionally, the metal texture image generation model is trained based on multiple sample image pairs;

[0011] A plurality of sample image pairs is determined based on a plurality of second sample images and a plurality of target sample images;

[0012] A plurality of second sample images are determined based on the first sample image;

[0013] The first sample image is determined based on the target sample image.

[0014] Optionally, the second sample image is obtained by inputting a grayscale image corresponding to the target sample image into the second network model;

[0015] The second network model is obtained by training the first network model based on the grayscale image corresponding to the third sample image and the target sample image;

[0016] The third sample image is determined based on the first sample image.

[0017] Optionally, the third sample image is obtained by aligning the second facial key points in the first sample image with the first facial key points in the target sample image.

[0018] Optionally, the sample image pair is obtained based on the fourth sample image and the target sample image;

[0019] The fourth sample image is obtained by fusing the second sample image and the third sample image.

[0020] Optionally, the fourth sample image is obtained by fusing the second sample image and the corresponding third sample image based on the positions of facial features and the first mask image;

[0021] The positions of facial features and mask images are obtained based on face segmentation of the target sample image.

[0022] Optionally, the fourth sample image is obtained by fusing the second sample image and the corresponding third sample image based on the positions of facial features and the second mask image;

[0023] The second mask image is obtained based on performing morphological operation processing and / or Gaussian blur processing on the first mask image.

[0024] Optionally, the sample image pair is obtained based on the second sample image and the processed image;

[0025] The processed image is obtained based on attribute information adjustment processing performed on the target sample image.

[0026] Optionally, the second network model includes a plurality of third network models obtained during the training of the first network model;

[0027] The second sample image is obtained by performing transparency blending processing on a plurality of fifth sample images;

[0028] The plurality of fifth sample images are obtained by processing the grayscale images based on the plurality of third network models respectively.

[0029] In a second aspect, the present application provides a method for training a metal texture image generation model, comprising:

[0030] Determining a first sample image corresponding to each target sample image, where the first sample image is an image with a metallic texture;

[0031] Determining, based on the first sample images corresponding to the target sample images, a second sample image corresponding to each target sample image, wherein a difference between a position of facial features in the second sample image and a position of facial features in the corresponding target sample image is less than a preset value;

[0032] A plurality of sample image pairs are determined according to each second sample image and each target sample image, and model parameters of an initial metal texture image generation model are updated according to the plurality of sample image pairs to obtain a metal texture image generation model.

[0033] Optionally, determining a second sample image corresponding to the target sample image according to the first sample image includes:

[0034] Determining, based on each of the first sample images, a third sample image corresponding to each of the first sample images, wherein a difference between positions of facial features in the third sample image and positions of the facial features in the target sample image is smaller than a difference between positions of the facial features in the first sample image and positions of the facial features in the target sample image;

[0035] Training the first network model according to the grayscale images corresponding to each third sample image and each target sample image to obtain a second network model;

[0036] The grayscale image corresponding to the target sample image is processed by the second network model to obtain a second sample image.

[0037] Optionally, determining, according to the first sample image, a third sample image corresponding to the first sample image includes:

[0038] extracting a plurality of first facial key points in the target sample image and a plurality of second facial key points in the first sample image respectively;

[0039] The second facial key point in the first sample image is aligned with the first facial key point in the target sample image to obtain a third sample image.

[0040] Optionally, determining a sample image pair according to the second sample image and the target sample image includes:

[0041] fusing the second sample image and the third sample image to obtain a fourth sample image;

[0042] A sample image pair is determined according to the fourth sample image and the target sample image.

[0043] Optionally, fusing the second sample image and the third sample image to obtain a fourth sample image includes:

[0044] Perform face segmentation on the target sample image to obtain the positions of facial features and the first mask image;

[0045] According to the positions of the facial features and the first mask image, the second sample image and the corresponding third sample image are fused to obtain a fourth sample image.

[0046] Optionally, according to the positions of the facial features and the first mask image, the second sample image and the corresponding third sample image are fused to obtain a fourth sample image, including:

[0047] Performing morphological operation processing and / or Gaussian blur processing on the first mask image to obtain a second mask image;

[0048] According to the positions of the facial features and the second mask image, the second sample image and the corresponding third sample image are fused to obtain a fourth sample image.

[0049] Optionally, determining a sample image pair according to the second sample image and the target sample image includes:

[0050] Performing attribute information adjustment processing on the target sample image to obtain a processed image;

[0051] The second sample image and the processed image are determined as a sample image pair.

[0052] Optionally, the second network model includes a plurality of third network models obtained during the training of the first network model;

[0053] Processing the grayscale image corresponding to the target sample image through the second network model to obtain a second sample image includes:

[0054] Processing the grayscale image respectively through multiple third network models to obtain multiple fifth sample images;

[0055] Perform transparency blending processing on the multiple fifth sample images to obtain a second sample image.

[0056] In a third aspect, the present application provides a device for generating an image with a metallic texture, comprising: a processing module; the processing module is configured to:

[0057] Get the first video;

[0058] Inputting the first video into a pre-trained metal texture image generation model to obtain a second video, wherein each frame image in the second video is an image with a metal texture;

[0059] The metal texture image generation model is trained based on a plurality of first sample images and second sample images with metal texture corresponding to each first sample image.

[0060] Optionally, the metal texture image generation model is trained based on multiple sample image pairs;

[0061] A plurality of sample image pairs is determined based on a plurality of second sample images and a plurality of target sample images;

[0062] A plurality of second sample images are determined based on the first sample image;

[0063] The first sample image is determined based on the target sample image.

[0064] Optionally, the second sample image is obtained by inputting a grayscale image corresponding to the target sample image into the second network model;

[0065] The second network model is obtained by training the first network model based on the grayscale image corresponding to the third sample image and the target sample image;

[0066] The third sample image is determined based on the first sample image.

[0067] Optionally, the third sample image is obtained by aligning the second facial key points in the first sample image with the first facial key points in the target sample image.

[0068] Optionally, the sample image pair is obtained based on the fourth sample image and the target sample image;

[0069] The fourth sample image is obtained by fusing the second sample image and the third sample image.

[0070] Optionally, the fourth sample image is obtained by fusing the second sample image and the corresponding third sample image based on the positions of facial features and the first mask image;

[0071] The positions of facial features and mask images are obtained based on face segmentation of the target sample image.

[0072] Optionally, the fourth sample image is obtained by fusing the second sample image and the corresponding third sample image based on the positions of facial features and the second mask image;

[0073] The second mask image is obtained based on performing morphological operation processing and / or Gaussian blur processing on the first mask image.

[0074] Optionally, the sample image pair is obtained based on the second sample image and the processed image;

[0075] The processed image is obtained based on attribute information adjustment processing performed on the target sample image.

[0076] Optionally, the second network model includes a plurality of third network models obtained during the training of the first network model;

[0077] The second sample image is obtained by performing transparency blending processing on a plurality of fifth sample images;

[0078] The plurality of fifth sample images are obtained by processing the grayscale images based on the plurality of third network models respectively.

[0079] In a fourth aspect, the present application provides a training device for a metal texture image generation model, comprising: a processing module; the processing module is configured to:

[0080] Determining a first sample image corresponding to each target sample image, where the first sample image is an image with a metallic texture;

[0081] Determining, based on the first sample images corresponding to the target sample images, a second sample image corresponding to each target sample image, wherein a difference between a position of facial features in the second sample image and a position of facial features in the corresponding target sample image is less than a preset value;

[0082] A plurality of sample image pairs are determined according to each second sample image and each target sample image, and model parameters of an initial metal texture image generation model are updated according to the plurality of sample image pairs to obtain a metal texture image generation model.

[0083] Optionally, the processing module is specifically configured to:

[0084] Determining, based on each of the first sample images, a third sample image corresponding to each of the first sample images, wherein a difference between positions of facial features in the third sample image and positions of the facial features in the target sample image is smaller than a difference between positions of the facial features in the first sample image and positions of the facial features in the target sample image;

[0085] Training the first network model according to the grayscale images corresponding to each third sample image and each target sample image to obtain a second network model;

[0086] The grayscale image corresponding to the target sample image is processed by the second network model to obtain a second sample image.

[0087] Optionally, the processing module is specifically configured to:

[0088] extracting a plurality of first facial key points in the target sample image and a plurality of second facial key points in the first sample image respectively;

[0089] The second facial key point in the first sample image is aligned with the first facial key point in the target sample image to obtain a third sample image.

[0090] Optionally, the processing module is specifically configured to:

[0091] fusing the second sample image and the third sample image to obtain a fourth sample image;

[0092] A sample image pair is determined according to the fourth sample image and the target sample image.

[0093] Optionally, the processing module is specifically configured to:

[0094] Perform face segmentation on the target sample image to obtain the positions of facial features and the first mask image;

[0095] According to the positions of the facial features and the first mask image, the second sample image and the corresponding third sample image are fused to obtain a fourth sample image.

[0096] Optionally, the processing module is specifically configured to:

[0097] Performing morphological operation processing and / or Gaussian blur processing on the first mask image to obtain a second mask image;

[0098] According to the positions of the facial features and the second mask image, the second sample image and the corresponding third sample image are fused to obtain a fourth sample image.

[0099] Optionally, the processing module is specifically configured to:

[0100] Performing attribute information adjustment processing on the target sample image to obtain a processed image;

[0101] The second sample image and the processed image are determined as a sample image pair.

[0102] Optionally, the second network model includes multiple third network models obtained during the training of the first network model; the processing module is specifically configured to:

[0103] Processing the grayscale image respectively through multiple third network models to obtain multiple fifth sample images;

[0104] Perform transparency blending processing on the multiple fifth sample images to obtain a second sample image.

[0105] In a fifth aspect, the present application provides a device for generating an image with a metallic texture, comprising: a processor, and a memory communicatively connected to the processor;

[0106] Memory stores computer-executable instructions;

[0107] The processor executes the computer-executable instructions stored in the memory to implement any method of the first aspect.

[0108] In a sixth aspect, the present application provides a training device for a metal texture image generation model, comprising: a processor, and a memory communicatively connected to the processor;

[0109] Memory stores computer-executable instructions;

[0110] The processor executes the computer-executable instructions stored in the memory to implement any method of the second aspect.

[0111] In a seventh aspect, the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer execution instructions, which are used to implement any method as in the first aspect when executed by a processor.

[0112] In an eighth aspect, the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer execution instructions, which are used to implement any method as in the second aspect when executed by a processor.

[0113] In a ninth aspect, the present application provides a computer program product, characterized in that it includes a computer program, which implements any method of the first aspect when executed by a processor.

[0114] In a tenth aspect, the present application provides a computer program product, characterized in that it includes a computer program, which implements any method of the second aspect when executed by a processor.

[0115] The present application provides a method for generating an image with a metallic texture and a method for training a model, the method comprising: obtaining a first video; inputting the first video into a pre-trained metallic texture image generation model to obtain a second video, wherein each frame image in the second video is an image with a metallic texture; wherein the metallic texture image generation model is trained based on a plurality of first sample images and second sample images with a metallic texture corresponding to each first sample image. In the above method, the metallic texture image generation model is trained based on a plurality of first sample images and second sample images with a metallic texture corresponding to each first sample image, thereby avoiding the need for the normal phase estimation-based gold rendering method to first perform normal phase estimation on the portrait, thereby reducing the amount of data calculation. Therefore, the first video is input into the pre-trained metallic texture image generation model to obtain the second video in real time, thereby improving the real-time performance of the second video. BRIEF DESCRIPTION OF THE DRAWINGS

[0116] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0117] Figure 1 Schematic diagram of the application scenario provided for this application;

[0118] Figure 2 A flowchart of a method for generating an image with a metallic texture provided in this application;

[0119] Figure 3Flowchart of the training method for the metal texture image generation model provided in this application;

[0120] Figure 4 A flow chart of the method for obtaining a second sample image provided in this application;

[0121] Figure 5 A flow chart of the method for obtaining a sample image pair provided in this application;

[0122] Figure 6 A schematic diagram of the high-precision metal texture human model provided for this application;

[0123] Figure 7 A schematic diagram of the structure of a device for generating an image with a metallic texture provided in this application;

[0124] Figure 8 A schematic diagram of the structure of the training device for the metal texture image generation model provided in this application;

[0125] Figure 9 A hardware diagram of a device for generating metallic texture images provided in this application;

[0126] Figure 10 Hardware diagram of the training device for the metal texture image generation model provided in this application.

[0127] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0128] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0129] Related technologies use a gold rendering method based on normal phase estimation. This method first estimates the normal phase of a person's image and then applies the gold rendering to the person based on the estimated normal phase. Due to the computationally intensive nature of this method, the real-time performance of gold rendering on a person's image is poor. For example, when a person in a video is rendered gold, the video may experience significant lag after the rendering.

[0130] In this application, in order to improve the real-time performance of gold rendering of portraits, the inventors thought of designing a metal texture image generation model with low computational complexity. After inputting the first video into the metal texture image generation model, the second video can be quickly output. The first sample image of each frame in the second video is an image with a metallic texture, thereby improving the real-time performance of gold rendering of portraits.

[0131] The following combination Figure 1 The application scenarios of the method for generating an image with a metallic texture provided in this application are described.

[0132] Figure 1 This is a schematic diagram of the application scenario provided by this application. Figure 1 As shown, it includes: a first video, a second video and a metal texture image generation model.

[0133] The first video includes multiple frames of images, for example, the multiple frames of images include images 11 , 12 , and 13 .

[0134] The second video includes multiple frames of images, for example, the multiple frames of images include images 21 , 22 , and 23 .

[0135] The metallic texture image generation model is used to process images 11, 12, and 13 in sequence and output images 21, 22, and 23. Image 21 is a sample image with metallic texture corresponding to image 11, image 22 is a sample image with metallic texture corresponding to image 12, and image 23 is a sample image with metallic texture corresponding to image 13.

[0136] Next, the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0137] Figure 2 This is a flow chart of the method for generating a metallic texture image provided by this application. Figure 2 As shown, the method includes:

[0138] S201, obtaining a first video.

[0139] Optionally, the execution subject of the method for generating a metallic texture image can be a metallic texture image generating device, or a metallic texture image generating device set in a metallic texture image generating device. The metallic texture image generating device can be implemented by a combination of software and / or hardware.

[0140] Optionally, the first video may be a video captured in real time by the electronic device, or may be a video pre-stored in the electronic device. The first video includes X frames of images, where X is an integer greater than or equal to 2.

[0141] S202, input the first video into a pre-trained metal texture image generation model to obtain a second video; each frame image in the second video is an image with a metal texture, and the metal texture image generation model is trained based on multiple first sample images and second sample images with a metal texture corresponding to each first sample image.

[0142] The number of images included in the second video is equal to the number of images included in the first video.

[0143] exist Figure 2 In the method for generating an image with a metallic texture provided in the embodiment, the metallic texture image generation model is trained based on multiple first sample images and second sample images with a metallic texture corresponding to each first sample image. Therefore, it is possible to avoid the need for the gold rendering method based on normal estimation to first perform normal estimation on the portrait, thereby reducing the amount of data calculation. Therefore, by inputting the first video into the pre-trained metallic texture image generation model, the second video can be obtained in real time, thereby improving the real-time performance of the second video.

[0144] Based on the above embodiments, the training method of the metal texture image generation model is described below in conjunction with specific embodiments.

[0145] Figure 3 This is a flow chart of the training method for the metal texture image generation model provided in this application. Figure 3 As shown, the method includes:

[0146] S301 : Determine a first sample image corresponding to each target sample image, where the first sample image is an image with a metallic texture.

[0147] Optionally, the execution entity of the training method of the metal texture image generation model can be the training device of the metal texture image generation model, or the training device of the metal texture image generation model set in the training device, and the training device can be implemented by a combination of software and / or hardware.

[0148] The training device may be, for example, a server, or the above-mentioned generating device.

[0149] The target sample image is a color image. For example, the target sample image can be an RGB image.

[0150] In some embodiments, for each target sample image, the first sample image corresponding to the target sample image can be determined by the following method: adjusting the lighting, color, and posture angle of N pre-rendered three-dimensional (3D) high-precision metal texture human models respectively to generate multiple two-dimensional (2D) metal texture images; using multiple two-dimensional (2D) metal texture images, training the initial metal texture addition model to obtain the metal texture addition model; using the metal texture addition model, processing the target sample image to obtain the first sample image corresponding to the target sample image.

[0151] N is an integer greater than or equal to 2. For example, N may be 6, 7, 8, 10, etc.

[0152] The initial metal texture addition can be used for the Agile Generative adversarialNets (Agile-GAN) model.

[0153] For the first sample image corresponding to the target sample image, there is a large difference in the shapes of facial features between the first sample image and the target sample image, but the first sample image has a better metallic texture.

[0154] S302, determining a second sample image corresponding to each target sample image based on the first sample image corresponding to each target sample image, wherein the difference between the position of the facial features in the second sample image and the position of the facial features in the target sample image is less than a preset value.

[0155] Each target sample image has a corresponding second sample image.

[0156] In some embodiments, for each target sample image, a second sample image corresponding to the target sample image can be obtained through the following methods 01 and 02.

[0157] Method 01 can use the first sample image corresponding to each target sample image and the grayscale image corresponding to each target sample image to train the first network model to obtain the second network model; and use the second network model to process the target sample image to obtain the second sample image corresponding to the target sample image.

[0158] Method 02 can use the first sample image corresponding to each target sample image and the grayscale image corresponding to each target sample image to train the first network model to obtain the second network model; adjust the attribute information of the target sample image to obtain a processed image; and process the processed image through the second network model to obtain a second sample image corresponding to the target sample image.

[0159] The first network model is the Cycle Generative adversarial Nets (Cycle-GAN) model.

[0160] Optionally, the attribute information may include any at least one of brightness, contrast, hue, resolution, blurring, etc.

[0161] S303 , determining a plurality of sample image pairs according to each second sample image and each target sample image, and updating model parameters of an initial metal texture image generation model according to the plurality of sample image pairs to obtain a metal texture image generation model.

[0162] In some embodiments, for each sample image pair, a target sample image and a second sample image corresponding to the target sample image are determined as the sample image pair.

[0163] In some embodiments, for each sample image pair, attribute information adjustment processing is performed on the target sample image to obtain a processed image; and the second sample image and the processed image are determined as the sample image pair.

[0164] In the present application, attribute information adjustment processing is performed on the target sample image to obtain a processed image; based on the processed image, a sample image pair is obtained, which can improve the model adaptability and robustness of the metal texture image generation model.

[0165] The initial metal texture image generation model is the Pix2pix model.

[0166] exist Figure 3 In the training method of the metal texture image generation model provided in the embodiment, each target sample image is processed separately through the metal texture addition model to obtain the first sample image corresponding to each target sample image, so as to determine multiple first sample images, so that the first sample image can have a better metal texture. In addition, based on the first sample image, multiple second sample images corresponding to the target sample image are determined, and the difference between the position of the facial features in each second sample image and the position of the facial features in the target sample image is less than a preset value, so that the facial features of the second sample image and the target sample image are more aligned. Furthermore, by determining multiple sample image pairs based on each second sample image and each target sample image, and updating the model parameters of the initial metal texture image generation model based on the multiple sample image pairs to obtain the metal texture image generation model, the accuracy of the metal texture image generation model can be improved.

[0167] Based on the above embodiments, Figure 4A method of obtaining the second sample image is described.

[0168] Figure 4 This is a flow chart of the method for obtaining the second sample image provided by this application. Figure 4 As shown, the method includes:

[0169] S401 : Determine, based on each first sample image, a third sample image (denoted as warp_agile_image) corresponding to each first sample image.

[0170] The difference between the positions of the facial features in the third sample image and the positions of the facial features in the target sample image is smaller than the difference between the positions of the facial features in the first sample image and the positions of the facial features in the target sample image.

[0171] In some embodiments, the third sample image can be obtained by the following method: extracting multiple first facial key points in the target sample image and multiple second facial key points in the first sample image respectively; aligning the second facial key points in the first sample image with the first facial key points in the target sample image to obtain the third sample image.

[0172] Optionally, multiple first facial key points in the target sample image may be extracted in the following three ways.

[0173] Method 11: extract key points from the target sample image using a facial key point detection algorithm model to obtain multiple first facial key points in the target sample image.

[0174] Method 12, through the facial key point detection algorithm model, key point extraction is performed on the target sample image to obtain multiple facial key points; through the pupil key point detection algorithm model, key point extraction is performed on the target sample image to obtain multiple pupil key points; through the facial outer contour key point detection algorithm model, key point extraction is performed on the target sample image to obtain multiple facial outer contour key points; based on the multiple facial key points, multiple pupil key points and multiple facial outer contour key points, multiple target facial key points are determined.

[0175] Optionally, key points corresponding to the four parts of the nose, mouth, eyes and eyebrows among the multiple facial key points, as well as multiple pupil key points and multiple facial contour key points can be determined as multiple target facial key points.

[0176] Optionally, the key points corresponding to the five parts of the nose, mouth, eyes, eyebrows and facial contour (contour of the lower half of the face) among multiple facial key points, multiple pupil key points, and the key points corresponding to the contour of the upper half of the face among multiple facial contour key points can also be determined as multiple target facial key points.

[0177] Method 13, through the face key point detection algorithm model, key point extraction is performed on the target sample image to obtain multiple facial key points; through the pupil key point detection algorithm model, key point extraction is performed on the target sample image to obtain multiple pupil key points; through the mouth dense key point detection algorithm model, key point extraction is performed on the target sample image to obtain multiple mouth key points; through the face outer contour key point detection algorithm model, key point extraction is performed on the target sample image to obtain multiple face outer contour key points; based on the multiple facial key points, multiple pupil key points, multiple mouth key points and multiple face outer contour key points, multiple target face key points are determined.

[0178] Optionally, key points corresponding to the three parts of the nose, eyes and eyebrows among the multiple facial key points, as well as multiple pupil key points, multiple mouth key points and multiple facial contour key points can be determined as multiple target facial key points.

[0179] Optionally, the key points corresponding to the four parts of the nose, eyes, eyebrows and facial contour (contour of the lower half of the face) among multiple facial key points, multiple pupil key points, multiple mouth key points, and key points corresponding to the contour of the upper half of the face among multiple facial contour key points can also be determined as multiple target facial key points.

[0180] Optionally, a method similar to method 11, 12, or 13 may be used to extract multiple second facial key points in the first sample image, which will not be described in detail here.

[0181] In some embodiments, the second facial key points in the first sample image and the first facial key points in the target sample image may be aligned by affine transformation (thin plate spline warping) to obtain a third sample image.

[0182] It should be noted that the positions of the facial features in the third sample image are more aligned with the positions of the facial features in the target sample image, and the degree of fit is higher.

[0183] S402: Training the first network model according to the grayscale images corresponding to each third sample image and each target sample image to obtain a second network model.

[0184] S403: Process the grayscale image corresponding to the target sample image through the second network model to obtain a second sample image.

[0185] exist Figure 4In the second sample image obtained by the embodiment, the first network model is trained according to the grayscale images corresponding to each third sample image and each target sample image to obtain the second network model, and the grayscale image corresponding to the target sample image is processed to obtain the second sample image, so that the second sample image can better retain the facial features in the target sample image.

[0186] In some embodiments, the second network model includes multiple third network models obtained during the training of the first network model; the above S403 specifically includes: processing the grayscale image through multiple third network models respectively to obtain multiple fifth sample images; performing alpha-blending processing on the multiple fifth sample images to obtain the second sample image.

[0187] It should be noted that the facial features of the second sample image and the target sample image are relatively aligned, but the metal texture is usually poor. Therefore, based on the fact that the facial features of the second sample image and the target sample image are relatively aligned, a sample image pair is obtained through a third sample image with better metal texture.

[0188] Based on the above embodiments, Figure 5 The method of obtaining sample image pairs is described. For details, see Figure 5 Example.

[0189] Figure 5 The method flow for obtaining sample image pairs provided in this application Figure 1 .like Figure 5 As shown, the method includes:

[0190] S501 : Fusing the second sample image and the third sample image corresponding to the target sample image to obtain a fourth sample image (denoted as final_target_image).

[0191] In some embodiments, S501 specifically includes: performing face segmentation on the target sample image to obtain the positions of facial features and a first mask image; based on the positions of facial features and the first mask image, fusing the second sample image (denoted as cycle_image) and the corresponding third sample image to obtain a fourth sample image.

[0192] The first mask image is a mask image corresponding to the facial features.

[0193] In the present application, since the facial features in the second sample image are better and the metal texture of the third sample image is better, the facial features image can be obtained from the second sample image based on the first mask image and the position of the facial features, and other images except the above-mentioned facial features image can be obtained from the third sample image corresponding to the second sample image, and the facial features image and other images can be combined to obtain the fourth sample image.

[0194] In some embodiments, the second sample image and the corresponding third sample image are fused according to the positions of the facial features and the first mask image to obtain a fourth sample image, including: performing morphological operations and / or Gaussian blur processing on the first mask image to obtain the second mask image; and fusing the second sample image and the corresponding third sample image according to the positions of the facial features and the second mask image to obtain the fourth sample image.

[0195] In this application, the first mask image is subjected to morphological operation processing and / or Gaussian blur processing to obtain a second mask image; according to the position of the facial features and the second mask image, a fourth sample image is obtained.

[0196] In some embodiments, the morphological operation process includes dilation and / or erosion.

[0197] S502: Determine a sample image pair according to the fourth sample image and the target sample image.

[0198] In some embodiments, the fourth sample image and the target sample image may be determined as a sample image pair.

[0199] In some embodiments, attribute information adjustment processing is performed on the target sample image to obtain a processed image;

[0200] The fourth sample image and the processed image are determined as a sample image pair.

[0201] exist Figure 5 In an embodiment, the second sample image and the third sample image corresponding to the target sample image are fused to obtain a fourth sample image, which can improve the similarity of facial features between the fourth sample image and the target sample image, as well as the metallic texture of the facial image in the fourth sample image.

[0202] It should be noted that in the process of adjusting the attribute information of the image involved in this application, the hair color in the image can also be changed according to the hair mask image of the target sample image corresponding to the image, so that the initial metal texture image generation model can learn the hair color under variable conditions and is no longer limited to generating images in a single black color, thereby improving the adaptability and robustness of the metal texture image generation model.

[0203] Optionally, after or before adjusting the attribute information of the image, image noise can be added to the image to allow the initial metal texture image generation model to learn sample image pairs under extreme conditions to further improve the adaptability and robustness of the metal texture image generation model.

[0204] Figure 6 This is a schematic diagram of the high-precision metal texture human model provided in this application. Figure 6 As shown, it includes: three-dimensional high-precision metal texture human models 61, 62, and 63.

[0205] Figure 7 This is a schematic diagram of the structure of the device for generating a metallic texture image provided by this application. Figure 7 As shown, the device 10 for generating a metallic texture image includes: a processing module 101; the processing module 101 is used to:

[0206] Get the first video;

[0207] Inputting the first video into a pre-trained metal texture image generation model to obtain a second video, wherein each frame image in the second video is an image with a metal texture;

[0208] The metal texture image generation model is trained based on a plurality of first sample images and second sample images with metal texture corresponding to each first sample image.

[0209] The device 10 for generating an image with a metallic texture provided in the present application can execute the above-mentioned method for generating an image with a metallic texture. Its implementation principles and beneficial effects are similar and will not be described in detail here.

[0210] Optionally, the metal texture image generation model is trained based on multiple sample image pairs;

[0211] A plurality of sample image pairs is determined based on a plurality of second sample images and a plurality of target sample images;

[0212] A plurality of second sample images are determined based on the first sample image;

[0213] The first sample image is determined based on the target sample image.

[0214] Optionally, the second sample image is obtained by inputting a grayscale image corresponding to the target sample image into the second network model;

[0215] The second network model is obtained by training the first network model based on the grayscale image corresponding to the third sample image and the target sample image;

[0216] The third sample image is determined based on the first sample image.

[0217] Optionally, the third sample image is obtained by aligning the second facial key points in the first sample image with the first facial key points in the target sample image.

[0218] Optionally, the sample image pair is obtained based on the fourth sample image and the target sample image;

[0219] The fourth sample image is obtained by fusing the second sample image and the third sample image.

[0220] Optionally, the fourth sample image is obtained by fusing the second sample image and the corresponding third sample image based on the positions of facial features and the first mask image;

[0221] The positions of facial features and mask images are obtained based on face segmentation of the target sample image.

[0222] Optionally, the fourth sample image is obtained by fusing the second sample image and the corresponding third sample image based on the positions of facial features and the second mask image;

[0223] The second mask image is obtained based on performing morphological operation processing and / or Gaussian blur processing on the first mask image.

[0224] Optionally, the sample image pair is obtained based on the second sample image and the processed image;

[0225] The processed image is obtained based on attribute information adjustment processing performed on the target sample image.

[0226] Optionally, the second network model includes a plurality of third network models obtained during the training of the first network model;

[0227] The second sample image is obtained by performing transparency blending processing on a plurality of fifth sample images;

[0228] The plurality of fifth sample images are obtained by processing the grayscale images based on the plurality of third network models respectively.

[0229] The device 10 for generating an image with a metallic texture provided in the present application can execute the above-mentioned method for generating an image with a metallic texture. Its implementation principles and beneficial effects are similar and will not be described in detail here.

[0230] Figure 8 This is a schematic diagram of the structure of the training device for the metal texture image generation model provided in this application. Figure 8 As shown, the training device 20 for the metal texture image generation model includes: a processing module 201; the processing module 201 is used to:

[0231] Determining a first sample image corresponding to each target sample image, where the first sample image is an image with a metallic texture;

[0232] Determining, based on the first sample images corresponding to the target sample images, a second sample image corresponding to each target sample image, wherein a difference between a position of facial features in the second sample image and a position of facial features in the corresponding target sample image is less than a preset value;

[0233] A plurality of sample image pairs are determined according to each second sample image and each target sample image, and model parameters of an initial metal texture image generation model are updated according to the plurality of sample image pairs to obtain a metal texture image generation model.

[0234] The training device 20 for the metal texture image generation model provided in this application can execute the above-mentioned training method for the metal texture image generation model. Its implementation principle and beneficial effects are similar and will not be repeated here.

[0235] Optionally, the processing module is specifically configured to:

[0236] Determining, based on each of the first sample images, a third sample image corresponding to each of the first sample images, wherein a difference between positions of facial features in the third sample image and positions of the facial features in the target sample image is smaller than a difference between positions of the facial features in the first sample image and positions of the facial features in the target sample image;

[0237] Training the first network model according to the grayscale images corresponding to each third sample image and each target sample image to obtain a second network model;

[0238] The grayscale image corresponding to the target sample image is processed by the second network model to obtain a second sample image.

[0239] Optionally, the processing module is specifically configured to:

[0240] extracting a plurality of first facial key points in the target sample image and a plurality of second facial key points in the first sample image respectively;

[0241] The second facial key point in the first sample image is aligned with the first facial key point in the target sample image to obtain a third sample image.

[0242] Optionally, the processing module is specifically configured to:

[0243] fusing the second sample image and the third sample image to obtain a fourth sample image;

[0244] A sample image pair is determined according to the fourth sample image and the target sample image.

[0245] Optionally, the processing module is specifically configured to:

[0246] Perform face segmentation on the target sample image to obtain the positions of facial features and the first mask image;

[0247] According to the positions of the facial features and the first mask image, the second sample image and the corresponding third sample image are fused to obtain a fourth sample image.

[0248] Optionally, the processing module is specifically configured to:

[0249] Performing morphological operation processing and / or Gaussian blur processing on the first mask image to obtain a second mask image;

[0250] According to the positions of the facial features and the second mask image, the second sample image and the corresponding third sample image are fused to obtain a fourth sample image.

[0251] Optionally, the processing module is specifically configured to:

[0252] Performing attribute information adjustment processing on the target sample image to obtain a processed image;

[0253] The second sample image and the processed image are determined as a sample image pair.

[0254] Optionally, the second network model includes multiple third network models obtained during the training of the first network model; the processing module is specifically configured to:

[0255] Processing the grayscale image respectively through multiple third network models to obtain multiple fifth sample images;

[0256] Perform transparency blending processing on the multiple fifth sample images to obtain a second sample image.

[0257] The training device 20 for the metal texture image generation model provided in this application can execute the above-mentioned training method for the metal texture image generation model. Its implementation principle and beneficial effects are similar and will not be repeated here.

[0258] Figure 9 This is a hardware diagram of the device for generating a metallic texture image provided by this application. Figure 9 As shown, the device 30 for generating an image with a metallic texture may include: a transceiver 301 , a memory 302 and a processor 303 .

[0259] The transceiver 301 may include a transmitter and / or a receiver. The transmitter may also be referred to as a transmitter, a transmitter, a transmitting port, a transmitting interface, or similar descriptions. The receiver may also be referred to as a receiver, a receiver, a receiving port, a receiving interface, or similar descriptions.

[0260] Exemplarily, the transceiver 301 , the memory 302 , and the processor 303 are interconnected via a bus.

[0261] The memory 302 is used to store computer-executable instructions.

[0262] The processor 303 is configured to execute the computer-executable instructions stored in the memory 302 , so that the processor 303 executes the above-mentioned method for generating an image with a metallic texture.

[0263] Figure 10 This is a hardware diagram of the training device for the metal texture image generation model provided in this application. Figure 10 As shown, the training device 40 for the metal texture image generation model may include: a transceiver 401 , a memory 402 and a processor 403 .

[0264] The transceiver 401 may include a transmitter and / or a receiver. The transmitter may also be referred to as a transmitter, a transmitter, a transmitting port, a transmitting interface, or similar descriptions. The receiver may also be referred to as a receiver, a receiver, a receiving port, a receiving interface, or similar descriptions.

[0265] Exemplarily, the transceiver 401 , the memory 402 , and the processor 403 are interconnected via a bus.

[0266] The memory 402 is used to store computer-executable instructions.

[0267] The processor 403 is configured to execute the computer-executable instructions stored in the memory 402 , so that the processor 403 executes the above-mentioned training method for the metal texture image generation model.

[0268] The present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the above-mentioned method for generating an image with a metallic texture is implemented.

[0269] The present application provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the training method of the above-mentioned metal texture image generation model is implemented.

[0270] The present application also provides a computer program product, including a computer program, which, when executed by a processor, can implement the above-mentioned method for generating an image with a metallic texture.

[0271] The present application also provides a computer program product, including a computer program, which, when executed by a processor, can implement the above-mentioned training method for the metal texture image generation model.

[0272] All or part of the steps of the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, it performs the steps of the above-mentioned method embodiments; and the aforementioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disc, and any combination thereof.

[0273] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processing unit of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0274] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0275] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0276] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

[0277] In this application, the term "include" and its variations may refer to non-restrictive inclusion; the term "or" and its variations may refer to "and / or". The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. In this application, "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship.

[0278] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0279] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for generating an image with a metallic texture, characterized in that: include: Get the first video; Inputting the first video into a pre-trained metal texture image generation model to obtain a second video, wherein each frame image in the second video is an image with a metal texture; The metal texture image generation model is trained based on a plurality of first sample images and second sample images with metal texture corresponding to each first sample image; The metal texture image generation model is obtained by training based on multiple sample image pairs; The plurality of sample image pairs are determined based on a plurality of second sample images and a plurality of target sample images, wherein the target sample images are color images; and the first sample images are images with a metallic texture; The plurality of second sample images are determined based on the first sample image, and the difference between the positions of facial features in the second sample images and the positions of facial features in the corresponding target sample images is less than a preset value; The first sample image is determined based on the target sample image.

2. The method according to claim 1, characterized in that The second sample image is obtained by inputting a grayscale image corresponding to the target sample image into the second network model; The second network model is obtained by training the first network model based on a third sample image and a grayscale image corresponding to the target sample image; the difference between the positions of the facial features in the third sample image and the positions of the facial features in the target sample image is smaller than the difference between the positions of the facial features in the first sample image and the positions of the facial features in the target sample image; The third sample image is determined based on the first sample image.

3. The method according to claim 2, characterized in that The third sample image is obtained by aligning the second facial key points in the first sample image and the first facial key points in the target sample image.

4. The method according to claim 2 or 3, characterized in that The sample image pair is obtained based on a fourth sample image and the target sample image; The fourth sample image is obtained by fusing the second sample image and the third sample image.

5. The method according to claim 4, characterized in that The fourth sample image is obtained by fusing the second sample image and the corresponding third sample image based on the positions of facial features and the first mask image; The facial feature positions and the mask image are obtained based on face segmentation performed on the target sample image.

6. The method according to claim 5, characterized in that The fourth sample image is obtained by fusing the second sample image and the corresponding third sample image based on the positions of the facial features and the second mask image; The second mask image is obtained by performing morphological operation processing and / or Gaussian blur processing on the first mask image.

7. The method according to claim 1 or 2, characterized in that The sample image pair is obtained based on the second sample image and the processed image; The processed image is obtained based on attribute information adjustment processing performed on the target sample image.

8. The method according to claim 2 or 3, characterized in that The second network model includes a plurality of third network models obtained during the training process of the first network model; The second sample image is obtained by performing transparency blending processing on a plurality of fifth sample images; The multiple fifth sample images are obtained by processing the grayscale images respectively based on the multiple third network models.

9. A method for training a metal texture image generation model, characterized in that: include: Determining a first sample image corresponding to each target sample image, where the first sample image is an image with a metallic texture; The target sample image is a color image; Determining, based on the first sample images corresponding to the target sample images, a second sample image corresponding to each target sample image, wherein a difference between a position of facial features in the second sample image and a position of facial features in the corresponding target sample image is less than a preset value; A plurality of sample image pairs are determined according to each second sample image and each target sample image, and model parameters of an initial metal texture image generation model are updated according to the plurality of sample image pairs to obtain a metal texture image generation model.

10. The method according to claim 9, characterized in that Determining, according to the first sample image, a second sample image corresponding to the target sample image, comprising: Determining, based on each of the first sample images, a third sample image corresponding to each of the first sample images, wherein a difference between positions of facial features in the third sample image and positions of the facial features in the target sample image is smaller than a difference between positions of facial features in the first sample image and positions of the facial features in the target sample image; Training the first network model according to the grayscale images corresponding to each third sample image and each target sample image to obtain a second network model; The grayscale image corresponding to the target sample image is processed by the second network model to obtain the second sample image.

11. The method according to claim 10, characterized in that Determining, according to the first sample image, a third sample image corresponding to the first sample image includes: extracting a plurality of first facial key points in the target sample image and a plurality of second facial key points in the first sample image respectively; The second facial key points in the first sample image and the first facial key points in the target sample image are aligned to obtain the third sample image.

12. The method according to claim 10 or 11, characterized in that Determining a sample image pair according to the second sample image and the target sample image includes: fusing the second sample image and the third sample image to obtain a fourth sample image; A sample image pair is determined according to the fourth sample image and the target sample image.

13. The method according to claim 12, characterized in that The fusing the second sample image and the third sample image to obtain a fourth sample image includes: Performing face segmentation on the target sample image to obtain facial features and a first mask image; The second sample image and the corresponding third sample image are fused according to the positions of the facial features and the first mask image to obtain the fourth sample image.

14. The method according to claim 13, wherein: The fusing the second sample image and the corresponding third sample image according to the positions of the facial features and the first mask image to obtain the fourth sample image includes: performing morphological operation processing and / or Gaussian blur processing on the first mask image to obtain a second mask image; According to the positions of the facial features and the second mask image, the second sample image and the corresponding third sample image are fused to obtain the fourth sample image.

15. The method according to claim 9 or 10, characterized in that Determining a sample image pair according to the second sample image and the target sample image includes: Performing attribute information adjustment processing on the target sample image to obtain a processed image; The second sample image and the processed image are determined as a sample image pair.

16. The method according to claim 10 or 11, characterized in that The second network model includes a plurality of third network models obtained during the training process of the first network model; Processing the grayscale image corresponding to the target sample image by the second network model to obtain the second sample image includes: Processing the grayscale images respectively through the multiple third network models to obtain multiple fifth sample images; Perform transparency blending processing on the multiple fifth sample images to obtain the second sample image.

17. A device for generating an image with a metallic texture, characterized in that: include: Processing module; the processing module is used to: Get the first video; Inputting the first video into a pre-trained metal texture image generation model to obtain a second video, wherein each frame image in the second video is an image with a metal texture; The metal texture image generation model is trained based on a plurality of first sample images and second sample images with metal texture corresponding to each first sample image; The metal texture image generation model is obtained by training based on multiple sample image pairs; The plurality of sample image pairs are determined based on a plurality of second sample images and a plurality of target sample images, wherein the target sample images are color images; and the first sample images are images with a metallic texture; The plurality of second sample images are determined based on the first sample image, and the difference between the positions of facial features in the second sample images and the positions of facial features in the corresponding target sample images is less than a preset value; The first sample image is determined based on the target sample image.

18. A training device for a metal texture image generation model, characterized in that: include: Processing module; the processing module is used to: Determining a first sample image corresponding to each target sample image, where the first sample image is an image with a metallic texture; and the target sample image is a color image; Determining, based on the first sample images corresponding to the target sample images, a second sample image corresponding to each target sample image, wherein a difference between a position of facial features in the second sample image and a position of facial features in the corresponding target sample image is less than a preset value; A plurality of sample image pairs are determined according to each second sample image and each target sample image, and model parameters of an initial metal texture image generation model are updated according to the plurality of sample image pairs to obtain a metal texture image generation model.

19. A device for generating an image having a metallic texture, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.

20. A training device for a metal texture image generation model, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 9 to 16.

21. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 8 when executed by a processor.

22. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 9 to 16 when executed by a processor.

23. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 8 when executed by a processor.

24. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 9 to 16 when executed by a processor.

Citation Information

Patent Citations

  • Style image generation method and device, model training method and device, equipment and medium

    CN112989904A