Face image processing method and electronic device
By employing a phased training approach, combining unsupervised and supervised learning, an algorithm model is generated and optimized. This addresses the issues of high cost and computational complexity in existing technologies, enabling low-cost, real-time facial stylization processing suitable for mobile devices.
Patent Information
- Application Number
- CN202210860811.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2042-07-21
AI Technical Summary
Existing technologies require significant manual and time costs to train facial stylization algorithm models, and struggle to keep up with the rapid changes on the internet. Furthermore, existing models are computationally intensive and difficult to run in real-time on mobile devices.
A combination of unsupervised and supervised learning is used to train the algorithm model in stages: first, a stylized face image dataset without pairing relationships is generated, and the first algorithm model is trained through unsupervised learning; then, the second algorithm model is trained using the dataset with pairing relationships, which is adapted for mobile devices.
It achieves low-cost, real-time facial stylization, reduces computational load, enables models to run efficiently on mobile devices, reduces reliance on designers, and supports multiple stylization methods.
Smart Images

Figure CN115393177B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a face image processing method and an electronic device. BACKGROUND
[0002] In the fields of live broadcast, short video, text and picture scene, digital human, etc., the facial form always plays an important role. Facial stylization can not only realize scene functions such as atmosphere creation and improved visual perception, but also has the functions of facial privacy protection, added fun, and IP style promotion. Facial stylization refers to converting a real facial image collected in reality into a facial image of a certain style, such as a Disney cartoon style facial image, and meanwhile, the converted facial image can also retain the landmark attribute features of the real facial image.
[0003] In the process of realizing facial stylization, an algorithm model needs to be pre-trained so that the algorithm model learns the image processing method in the process of mapping from a real facial image to a facial image of a certain style, so that after inputting a real facial image into the algorithm model, the algorithm model can output a facial image of the style.
[0004] In the prior art, when training an algorithm model for facial stylization, paired training samples are needed, i.e., a real facial image and a stylized facial image corresponding to the real facial image. However, such a stylized facial image corresponding to a real facial image does not exist in reality, so a large number of professional designers are often needed to design the style, draw the stylized facial image, and repeatedly modify it for a long period of time for a real facial image, and then use the obtained stylized facial image corresponding to the real facial image to train the algorithm model. This mode not only requires a large amount of manual and time cost for the production of each stylization algorithm model, but also is difficult to keep up with the rapidly changing Internet world. SUMMARY
[0005] The present application provides a face image processing method and an electronic device, which can realize real-time facial stylization processing on a mobile terminal at a low cost.
[0006] The present application provides the following solutions:
[0007] A face image processing method, comprising:
[0008] obtaining a first data set composed of a plurality of stylized facial images with a target style and a second data set composed of a plurality of real facial images, wherein there is no pairing relationship between the stylized facial images in the first data set and the real facial images in the second data set;
[0009] The first data set and the second data set are taken as training samples, and the first algorithm model is trained by an unsupervised learning mode, so that a third data set composed of real face images and stylized face images with a pairing relationship is obtained by the first algorithm model;
[0010] The third data set is taken as a training sample, and the second algorithm model is trained by a supervised learning mode, so that the second algorithm model is distributed to a terminal device where a client is located, and the client is used to convert a real face image collected by the terminal device into a stylized face image of the target style by the second algorithm model.
[0011] The first data set composed of a plurality of stylized face images with a target style is obtained, including:
[0012] A first number of stylized face image raw materials related to the target style are collected;
[0013] The third algorithm model is trained by using the raw materials;
[0014] A second number of stylized face images are obtained according to the trained third algorithm model to form the first data set;
[0015] The first number is less than the second number.
[0016] Further comprising:
[0017] The third algorithm model is pre-trained by using a plurality of real face images;
[0018] The third algorithm model is trained by using the raw materials, including:
[0019] The raw materials are used to perform secondary training on the third algorithm model after the pre-training is completed.
[0020] Further comprising:
[0021] Before the secondary training is performed, the values of part of the parameters obtained by the pre-training in the third algorithm model are fixed, and the part of the parameters are parameters related to common features between real face images and stylized face images.
[0022] Further comprising:
[0023] The parameter values of the pre-training result and the secondary training result are fused to correct errors of the parameter values obtained by the secondary training;
[0024] The second number of stylized face images are obtained according to the trained third algorithm model, including:
[0025] obtain a second number of stylized face images according to the third algorithm model after fusing the parameter values.
[0026] Further comprising:
[0027] training a fourth algorithm model for generating random vectors based on the third algorithm model after pre-training is completed;
[0028] obtaining a second number of stylized face images according to the third algorithm model after secondary training, generating a random vector through the fourth algorithm model and taking the random vector as an input of the third algorithm model to control the distribution of stylized face images output by the third algorithm model.
[0029] Further comprising:
[0030] collecting stylized face image raw materials about at least two styles respectively;
[0031] training the third algorithm model by using the raw materials corresponding to the at least two styles respectively to obtain at least two groups of parameter values corresponding to the at least two styles respectively;
[0032] obtaining fused parameter values by fusing the at least two groups of parameter values;
[0033] obtaining a plurality of stylized face images with the target style according to the third algorithm model and the fused parameter values.
[0034] Further comprising:
[0035] Further comprising:
[0036] adding an average error L1 loss term about a background area in the generation network part, ignoring an adversarial loss term of the background area, and removing a branch related to discriminating a whole image in the discriminant network part, so as to avoid the image of the background area being stylized in the process of converting a real face image into a stylized face image.
[0037] Further comprising:
[0038] Further comprising:
[0039] performing edge recognition on the stylized face images in the third data set and performing blur processing on the edge part;
[0040] The third data set is taken as a training sample, and the second algorithm model is trained in a supervised learning manner, including:
[0041] The third data set and the stylized face image with the blurred edge part are taken as training samples, and the second algorithm model is trained in a supervised learning manner, and an edge enhancement adversarial loss is provided in the second algorithm model, so that the stylized face image generated by the second algorithm model obtains edge enhancement.
[0042] The second algorithm model includes a discriminator network, and the discriminator network has global discrimination ability, local discrimination ability and attention mechanism.
[0043] Further comprising:
[0044] The face image in the third data set is subjected to data enhancement processing, and the data enhancement processing includes random cropping, random scaling or random optical distortion processing.
[0045] The client includes a client provided by a commodity information service system;
[0046] The client is used to:
[0047] After receiving the request of the user to live broadcast or shoot short video / photograph of the target commodity, the stylization processing option is provided;
[0048] In response to the request initiated through the stylization processing option, the target style is determined, and the real face image is intercepted from the original image collected by the terminal device;
[0049] The real face image is converted into a stylized face image of the target style through the second algorithm model corresponding to the target style, and the stylized face image is pasted back into the original image.
[0050] A face image processing device, comprising:
[0051] A data generation unit is configured to obtain a first data set composed of a plurality of stylized face images with a target style and a second data set composed of a plurality of real face images, wherein there is no pairing relationship between the stylized face images in the first data set and the real face images in the second data set;
[0052] An unsupervised learning unit is configured to take the first data set and the second data set as training samples, and train a first algorithm model in an unsupervised learning manner, so as to obtain a third data set composed of real face images and stylized face images with a pairing relationship through the first algorithm model;
[0053] A supervised learning unit is configured to train the second algorithm model by using the third data set as training samples in a supervised learning manner, and distribute the second algorithm model to a terminal device where a client is located, and the client is configured to convert a real face image collected by the terminal device into a stylized face image of the target style by using the second algorithm model.
[0054] A computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the steps of the method of any preceding method.
[0055] An electronic device comprising:
[0056] One or more processors; and
[0057] A memory associated with the one or more processors, the memory configured to store program instructions that, when executed by the one or more processors, perform the steps of the method of any preceding method.
[0058] According to the specific embodiments provided in the present application, the present application discloses the following technical effects:
[0059] By the embodiments of the present application, the generation process of the stylization processing model can be divided into multiple stages. In the first stage, a first data set composed of multiple stylized face images with a target style and a second data set composed of multiple real face images can be obtained. The stylized face images in the first data set and the real face images in the second data set can not necessarily have a pairing relationship. Then, in the second stage, the first data set and the second data set can be used as training samples to train the first algorithm model by an unsupervised learning manner, so as to obtain a third data set composed of real face images and stylized face images with a pairing relationship by the first algorithm model. In the third stage, the third data set can be used as training samples to train the second algorithm model by a supervised learning manner, and then the second algorithm model can be distributed to a terminal device where a client is located, so that the client converts a collected real face image into a stylized face image with the target style by using the second algorithm model. In this way, the decoupling of the three tasks of stylized data production, paired image data making, and mobile terminal image translation can be realized. Moreover, the designer, expert, or the like does not need to design the stylized face image paired with the real face image, and the training of the second algorithm model can be completed at a lower cost, so that the second algorithm model can output a stylized face image with a certain style corresponding to the input of a real face image. Moreover, since the training of the second algorithm model can be realized by a supervised manner, the control of the operation amount can be realized, so that the second algorithm model can be run on the client to realize real-time stylization processing.
[0060] In optional embodiments, the algorithms in each stage, the algorithms or data between stages can also be optimized and improved. For example, in the first stage, the pre-training of the algorithm model by using real face data, the fixing of part of the parameters, and the like can be used to reduce the demand for the number of stylized face image raw materials, so that a third model can be trained by using a small amount of stylized face image raw materials to generate a large amount of stylized face data. Style innovation can also be realized by parameter fusion of different styles and the like. In the second stage, the average error L1 loss term about the background area can be added to the generation network part, the background area loss term can be ignored, and the branch related to the discrimination of the whole image can be removed from the discrimination network part, so as to realize the fixation of the background image in the stylization process and avoid the background blur caused by the stylization processing of the image of the background area. In the third stage, the loss and perception function can be increased, the stylized face image after the edge blur processing can be added to the training data, so that the stylized face image generated by the second algorithm model has an improved edge; the robustness of the algorithm model can be improved by the data enhancement processing of the face image in the third data set, and the like.
[0061] Of course, implementing any of the products of the present application does not necessarily require all of the advantages described above to be achieved simultaneously. BRIEF DESCRIPTION OF DRAWINGS
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0063] Figure 1 is a schematic diagram of the system architecture provided by the embodiments of the present application;
[0064] Figure 2 is a flowchart of the method provided by the embodiments of the present application;
[0065] Figure 3 is a schematic diagram of the apparatus provided by the embodiments of the present application;
[0066] Figure 4 is a schematic diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0067] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art belong to the scope of protection of the present application.
[0068] In order to facilitate understanding of the specific implementation schemes provided by the embodiments of the present application, it is first necessary to explain that in actual application scenarios, a user usually initiates a request for face stylization processing in the process of live streaming or shooting short videos through a specific application client. At this time, the real face image actually collected needs to be recognized, extracted, and stylized, and the stylized face image obtained is pasted back to the image collection interface. In this process, in order to better support real-time performance, the specific processing process usually needs to be completed by the client. That is, the server needs to distribute the specific stylization algorithm model to the client, and the client runs the specific algorithm model locally on the terminal device and generates a stylized face image, etc. The specific terminal device is usually a mobile terminal such as a mobile phone, and therefore the calculation amount in the algorithm model running process cannot be too large, otherwise the performance of the mobile terminal may not be able to support it.
[0069] Therefore, the improvement target of the embodiments of the present application is not only to reduce the production cost of the stylization algorithm model, but also to meet the demand of real-time face stylization processing on the mobile end. To this end, in the scheme provided by the embodiments of the present application, the training process of the stylization algorithm model is divided into three stages, wherein:
[0070] The first stage is used to produce a plurality of stylized face images and collect a plurality of real face images. Of course, in the embodiments of the present application, the specific stylized face images produced do not need to correspond to specific real face images, that is, although the image acquisition of the two domains of stylized face images and real face images is involved, the two do not need to have a pairing relationship. It is also because the stylized face images do not need to have a pairing relationship with the real face images that the production of the stylized face images does not need to be designed by a designer, but can be generated by an algorithm model.
[0071] The second stage can train the first algorithm model according to the stylized face images and the real face images that do not have a pairing relationship, so that the first algorithm model learns from the data the first mapping relationship of converting the real face image to the stylized face image of the target style, so that the stylized face image data having a pairing relationship with the real face image can be output under the condition that a real face image is input. Of course, since the stylized face images and the real face images as training data do not have a pairing relationship, the learning process of the first algorithm model belongs to unsupervised learning, and accordingly, the operator, depth, width, resolution, etc. of the algorithm have relatively high requirements, therefore, although the first algorithm model can output the stylized face image having the target style according to the input real face image, such a first algorithm model is usually difficult to run on the mobile end due to the high computational complexity. To this end, the embodiments of the present application also provide a third stage.
[0072] The third stage can convert a plurality of real face images to stylized face images by using the first algorithm model trained in the second stage, so that a data set composed of real face images and stylized face images having a pairing relationship is obtained, and the paired real face images and stylized face images in the data set are used to train the second algorithm model, so that the second algorithm model can also convert the real face image to the stylized face image. At this time, the second algorithm model can be a supervised model, therefore, compared with the first algorithm model, the computational complexity when running will be relatively low, and it is more suitable for running on the mobile end.
[0073] In each specific stage, the existing algorithm model can be trained using training data suitable for the application scenario of the embodiments of the present application. In addition, the existing algorithm model can be optimized or improved according to the characteristics of the specific application scenario to improve the effect of the image output by the algorithm. These specific optimization or improvement points will be described in detail below.
[0074] From the perspective of system architecture, referring to Figure 1 The embodiments of the present application can involve the client and server of the related application, wherein the server can be mainly used to generate a specific stylized algorithm model, that is, the above three stages can be completed on the server, and after obtaining the above second algorithm model, it can be distributed to the client, so that the client can complete the process of converting the real face image into the stylized face image according to the specific algorithm model on the terminal device locally.
[0075] The specific implementation scheme provided by the embodiments of the present application will be described in detail below.
[0076] First, the embodiment provides a face image processing method from the perspective of the aforementioned server, referring to Figure 2 The method can specifically include:
[0077] S201: Obtain a first data set composed of a plurality of stylized face images with a target style and a second data set composed of a plurality of real face images, wherein there is no pairing relationship between the stylized face images in the first data set and the real face images in the second data set.
[0078] This step S201 is the aforementioned first stage. Specifically, the specific target style required can be determined first, for example, a series of animation face style, etc., and then a first data set composed of a plurality of stylized face images with a target style and a second data set composed of a plurality of real face images can be obtained. Since the stylized face images in the first data set and the real face images in the second data set do not need to have a pairing relationship, that is, the specific stylized face data does not need to have the characteristics of a certain real face image. Therefore, such stylized face data can be obtained by collecting from some related picture library, official website, etc.
[0079] Of course, the collected stylized face images are used to train a specific model, so the number of stylized face images usually required is relatively large (for example, usually at least thousands), but the number of stylized face images of the same style that can actually be collected may be relatively small, for example, may be only dozens, etc. Therefore, in an optional implementation, for the case where the number of collected stylized face images is relatively small, a specific algorithm model (in order to distinguish from the subsequent algorithm model, it can be referred to as a third algorithm model) can also be used to generate such stylized face images.
[0080] Specifically, the first number of collected stylized face images can be used as training samples to train the third algorithm model. Specifically, when collecting such stylized face images, the data distribution can be optimized and adjusted, for example, the specific training samples preferably include various different angles (frontal, profile, overhead, etc.), whether wearing glasses, and various different facial features. For example, by collecting from an official website, etc. 100 stylized face images of a certain target style are obtained, then the 100 original materials can be optimized and adjusted in data distribution to make the images of various facial features evenly distributed, etc. After training the third algorithm model using such original materials, the second number of stylized face images can be obtained according to the third algorithm model trained to form the first data set. Wherein, the first number is less than the second number. That is, the number of original materials can be relatively small, but after the third algorithm model is trained, more stylized face images of a certain style can be obtained through the third algorithm model.
[0081] Specifically, the third algorithm model can be implemented by using a generative discriminative adversarial network model (for example, Style GAN, etc.). Such a generative discriminative adversarial network model includes a generative network model part and a discriminative network model part. The generative network model part can take any data (for example, a random vector of length N, etc.) as input, and the target is to output a stylized face image. Of course, at the beginning of training, the output of the generative network model cannot really have the required style, and the discriminative network model part is used to compare the image output by the generative network model part with the real stylized face image in the sample. If they are not close enough, the generative network model part will be notified to continue learning and modify the parameter value. After multiple iterations, when the image output by the generative network model part is close enough to the real stylized face image in the sample, the training can be stopped. After that, the generative network model part and the corresponding parameter value can be used to generate stylized face images. That is, a random vector is input into the generative network model part, and a stylized face image with the target style can be output.
[0082] Since the amount of raw material of the stylized face data collected under the same style can be less, in order to ensure the training effect of the third algorithm model under the condition of less training sample, in the optional implementation, after the third algorithm model is trained by using the raw material of the stylized face data, the third algorithm model can be pre-trained by using a plurality of real face images, and then the raw material of the stylized face data can be used to perform secondary training on the third algorithm model after the pre-training is completed.
[0083] The target of the pre-training is to obtain a set of parameter values for the third algorithm model, so that the third algorithm model can output the real face image according to the input random vector under the condition of the set of parameter values. Since the amount of real face images can be collected, all of which can be used as pre-training samples, an ideal pre-training effect can be obtained. In addition, since the set of parameter values has been obtained, the raw material of the stylized face data can be used as the training sample to perform secondary training on the third algorithm model based on the set of parameter values, which can reduce the iteration number in the training process and reduce the demand for the amount of stylized training sample.
[0084] Since the stylized face image and the real face image also have some common features, for example, the face contour, in the optional manner, after the pre-training is completed, before the secondary training is performed, the numerical value of part of the parameters obtained by the pre-training in the third algorithm model can be fixed, wherein the part of the parameters can be the parameters related to the common features between the real face image and the stylized face image. That is, by analyzing and comparing, it can be found which parameters affect the features such as the face contour, and the parameter values of these parameters can be fixed, so that only the parameter values of other parameters need to be learned in the secondary training process, and therefore the iteration number and the demand for the amount of training sample can also be reduced.
[0085] In addition, since the original material of the collected stylized face image is usually drawn and does not need to correspond to a real face image, some features may not be obvious enough, for example, most stylized face images may not include teeth, so when pre-training is performed using real face images and then secondary training is performed using stylized face images as training data, the model may not be able to correctly learn the processing method of the tooth feature when it involves a tooth-related feature. Therefore, in the preferred manner, the parameter values obtained by secondary training can also be error-corrected by fusing the pre-training results and the secondary training results. Then, the second number of stylized face images can be obtained according to the third algorithm model after parameter value fusion. When performing parameter fusion, a set of parameter values obtained after pre-training is fused with a set of parameter values obtained after secondary training, for example, the same parameter values in the two sets of parameter values are averaged or weighted averaged, etc. In the case of weighted average, the specific weight selection can be determined by repeated testing and the like.
[0086] Furthermore, since the input of the generation network part in the generative adversarial network model is a random vector, if the vector input is completely random, the distribution of the output face image may not be controlled. Therefore, in the optional implementation manner, the distribution control of the stylized face data generation can be realized by using GAN inversion and the like. Specifically, an encoding network can be trained according to the E4E (Encoder 4 Editing) algorithm or the like. The encoding network can be a fourth algorithm model for generating a random vector, or can be obtained by training. For example, when implemented, the fourth algorithm model for generating a random vector can be trained on the trained third algorithm model. That is, inputting a real face image into the fourth algorithm model (i.e., the encoding network) can obtain a latent space code, and inputting the code into the pre-trained Style GAN generator can obtain an output image basically consistent with the input real face image. In this way, the training of the fourth algorithm model can be completed based on the pre-trained third algorithm model. Furthermore, when the second number of stylized face images is obtained according to the trained third algorithm model, a random vector can be generated by the fourth algorithm model and used as the input of the third algorithm model to control the distribution of the stylized face images output by the third algorithm model.
[0087] In the case of generating the stylized face image through the third algorithm model, the style innovation can also be performed by fusing multiple different styles, so as to support the generation of more styles of stylized face images. For example, 20 styles of face image original materials can be collected in reality, the third algorithm model is trained through the 20 styles of original materials respectively, and 20 groups of parameter values are obtained. By fusing the 20 groups of parameter values two by two or fusing multiple groups of parameter values together, more combined styles can be obtained, that is, the third algorithm model after parameter fusion can be used to produce stylized face images of such combined styles, and then the conversion from real face images to stylized face images of such combined styles can be realized. For example, style A is Disney style, and style B is a child-related style. After fusing the parameter values corresponding to styles A and B, the third algorithm model can produce stylized face images of the "child Disney" style. In the subsequent face stylization process, not only "Disney style" and "child style" can be provided, but also "child Disney style" can be provided.
[0088] That is, in the above manner, at least two styles of stylized face image original materials can be collected respectively, and then the third algorithm model is trained using the original materials corresponding to the at least two styles to obtain at least two groups of parameter values corresponding to the at least two styles respectively. By fusing the at least two groups of parameter values, the fused parameter values can be obtained. Then, according to the third algorithm model and the fused parameter values, a plurality of stylized face images with the target style can be obtained.
[0089] It should be noted that in the process of style innovation by fusing two or more different styles, the specific parameter fusion method can also include averaging, weighted averaging and other fusion methods of parameter values on corresponding parameters. The weight used in the weighted averaging process can also be determined by repeated testing and other methods, which will not be described here.
[0090] S202: The first data set and the second data set are used as training samples to train the first algorithm model through unsupervised learning, so as to obtain a third data set of real face images and stylized face images with a paired relationship through the first algorithm model.
[0091] After obtaining the first data set and the second data set, a second stage can be entered, that is, the first data set and the second data set are taken as training samples to train the first algorithm model. Since there is no pairing relationship between the stylized face images in the first data set and the real face images in the second data set, the first algorithm model can be trained in an unsupervised manner. The training goal is to enable the first algorithm model to learn the mapping relationship from the real face image to the stylized face image of the target style under unsupervised manner, so that in the case of inputting a certain real face image into the first algorithm model, a stylized face image with the target style can be output. Of course, at this time, there will be a pairing relationship between the output stylized face image and the input real face image, that is, the output stylized face image will have part of the features in the input real face image, such as face shape, etc. (even for the user himself corresponding to the real face image, or for the user who is familiar with the person, the output stylized face image not only has the target style, but also can roughly judge who this person is).
[0092] The first algorithm model can be U-GAT-IT (unpaired image translation) and the like. However, in the embodiments of the present application, considering that there may be a requirement for fixed background image in some scenarios, the structure of the existing algorithm model can be improved to meet this requirement.
[0093] Specifically, since the input image and the output image of the first algorithm model are usually rectangular images of the same size, the actual face image is only part of the input image, and the other part is the background image (the face area cannot be a rectangle, and the input image must include a complete face image, so part of the background image will be inevitably included in the rectangular input image). If all the pixels in the rectangular input image are directly stylized, the background image part will also be "stylized", but in fact, the background image does not need to be "stylized", and even if it is "stylized", it will cause the background image to become blurred and distorted. In some scenarios (for example, real-time stylization during short video shooting, etc.), the stylized face image generated by the algorithm model also needs to be pasted back to the original image, at this time, the stylized face image may not be able to better connect with the original image due to the blurred background image part, etc., affecting the final visual effect, so in the above scenarios, there is a requirement for fixing the background image during stylization.
[0094] To meet the requirement, in the embodiments of the present application, the existing generative adversarial network model can be improved. For example, since the first algorithm model includes a generative network part and a discriminative network part, an L1 (mean error) loss term about the background region can be added to the generative network part, the adversarial loss of the background region is ignored, and the image background generated by the generative network is constrained to be the same as the original image. At the same time, since the image generated by the generative network is used as the input of the discriminator for discrimination, there is a discriminative loss term in the discrimination process, and usually, this loss term is based on the whole image for discrimination, which makes the discriminator punish the generated background part if it finds that the generated background image is different from other stylized patterns, so that the background part of the image generated by the generative network is also more like stylization. Therefore, while adding the L1 (mean error) loss term about the background region in the generative network, the branch involving the discrimination of the whole image can also be removed from the discriminative network part, so that the discriminator does not need to discriminate the background image of the image output by the generative network. In this way, the image of the background region can be avoided from being stylized in the process of converting the real face image into a stylized face image.
[0095] In summary, after the first algorithm model is trained by using the first data set and the second data set, a real face image can be input into the first algorithm model to output a stylized face image with a certain specific style, and the input image and the output image of the first algorithm model have a pairing relationship.
[0096] S203: The third data set is used as a training sample to train the second algorithm model in a supervised learning manner, so as to distribute the second algorithm model to a terminal device where a client is located, and the client is used to convert a real face image collected by the terminal device into a stylized face image of the target style by using the second algorithm model.
[0097] After the third data set composed of real face images and stylized face images with a pairing relationship is obtained by using the first algorithm model, the second algorithm model can be trained by using the third data set as a training sample. The training target is to learn the mapping relationship from a real face image to a stylized face image of the target style, so that the second algorithm model can know how to process the input real face image to output a stylized face image of the target style after the training is completed. That is, after the training is completed, a real face image input into the second algorithm model can output a stylized face image of the target style.
[0098] Wherein, since the images between the two domains in the training sample have a pairing relationship, supervised learning of the second algorithm model can be realized. And the second algorithm model based on supervised learning can be different from the first algorithm model in terms of operators, modules, depth, width, resolution, etc., and can be greatly simplified in terms of operation amount. Specifically, the running efficiency of operators, modules, depth, width, resolution, etc. on the central processor, graphics processor, etc. of the terminal device can be analyzed to determine suitable operators, modules, depth, width, resolution to obtain a lightweight second algorithm model and make it suitable for running on a mobile terminal device.
[0099] It should be noted here that different mobile terminal devices will have different hardware resource configurations (including the aforementioned central processor, graphics processor, etc.), and the ability to run algorithm models will also vary. For the second algorithm model, although a smaller model structure than the first algorithm model can be selected, if the second algorithm model is as large as possible within the capacity of the mobile terminal, it is also beneficial to obtain better processing results. Therefore, in specific implementation, the models of multiple mobile terminal devices can also be collected, the hardware conditions of various models are analyzed respectively, and the appropriate model size for each model is determined, then multiple second algorithm models of different sizes can be trained respectively, and then the trained second algorithm models can be distributed to the corresponding terminal devices for running. In this way, while achieving lightweight end-side models, the effect of face stylization processing can be improved as much as possible. Of course, the training data and training targets used in the training of second algorithm models of different sizes can be consistent overall, but the number of training data, the number of iterations, the accuracy of the final output face stylization result, etc. may be slightly different.
[0100] In summary, after the design of the second algorithm model is completed, the stylized face images and real face images in the aforementioned third data set with pairing relationship can be used to train it. And since the fixed background image has been realized when training the first algorithm model, that is, the background part of the image will not be stylized, therefore, when using such data as training samples, the processing result of the second algorithm model can also realize a fixed background image.
[0101] Specifically, a model such as U-Net can be used as the second algorithm model. However, if the existing U-Net model is directly used, since the L1 loss is generally performed on the entire input and output image, that is, each pixel point is considered to be equally important and all pixel points are required to approach the original image, the entire image will be smoother (or blurred). However, the human eye perceives different contents in the image differently. For example, the human eye perceives edges more obviously, and accordingly, for non-edge regions, even if the image is not very clear, the human eye may not perceive it obviously. Similarly, the human eye perceives the foreground more obviously, and even if the background is blurred, the human eye may not perceive it obviously or may not care, and the like. Therefore, from the perspective of human eye perception, this problem exists. Therefore, a discrimination loss and a perception loss can be added to the loss function in the existing model to improve the clarity of the edges and the foreground of the image, thereby obtaining a better experience.
[0102] In addition, in order to improve the edges (including face edges, glasses edges, nose edges, edges of hair accessories, and the like) of the generated stylized face image, edge recognition can be performed on the stylized face images in the third data set, and the edge part can be blurred to obtain an edge-blurred stylized face image. In this way, when the second algorithm model is trained, the training data can include three types, which are the original stylized face images in the third data set, the paired real face images, and the above-mentioned edge-blurred stylized face images. By inputting the edge-clear stylized face image and the edge-blurred stylized face image to the discriminator for training, the discriminator is punished for the edge-blurred image. In this way, the trained second algorithm model can improve the edges of the output stylized face image.
[0103] In addition, the second algorithm model can further include a discrimination network, which can have global discrimination ability, local discrimination ability, and attention mechanism. The global discrimination ability refers to processing the entire image and outputting a value representing whether the image is a real stylized image or a fake stylized face image, and then using the value to perform loss to make the generated entire image more inclined to a real stylized image. The local discrimination ability refers to that after the generated image is input to the discriminator, the discriminator can generate a matrix grid, each value of which corresponds to a pixel block of the original image, such as a pixel block in a region near the nose, and judges whether the generated image is a real stylized image or a fake stylized image in the pixel block dimension, and the like. This discrimination ability is referred to as local discrimination ability. The attention mechanism refers to enabling the algorithm model to selectively focus on part of the information while ignoring other information, so as to more reasonably utilize limited computing resources.
[0104] Further, in order to improve the robustness of the second algorithm model, data enhancement processing can also be performed on the face images in the third data set. For example, specific data enhancement processing can include random cropping, random scaling, or random optical distortion processing, etc. Random cropping refers to performing consistent cropping with the same offset direction and offset pixel quantity on a pair of paired face images. For example, ten pixel points on the left side are cropped in the x-axis direction to obtain a new image. For different face image pairs, the specific offset direction and offset quantity can be random. Random scaling refers to performing scaling processing on the face images in the third data set. The same scaling processing is performed on the same image pair, and for different image pairs, the scaling ratio can be random. Optical distortion processing is mainly performed on the real face images in the third data set. For example, by adding noise or changing the brightness and darkness, the quality of the real face images can be damaged, etc. After the above processing is performed on the data in the third data set, the second algorithm model is trained again. In this way, the second algorithm model can have higher robustness. For example, in the process of using the second algorithm model to perform real-time stylization processing on the actually collected real face images, even if the quality of the actually collected real face images is poor, including the presence of noise, or the image is collected in a dark environment, so that the collected real face image is dark, etc., but the stylized image generated by the second algorithm model can still have high quality.
[0105] After the second algorithm model is trained, the server can distribute the second algorithm model to the terminal device where the corresponding client is installed. Of course, in specific implementation, the server can train the second algorithm model for multiple target styles to obtain multiple sets of different parameter values. Therefore, the server can distribute the multiple sets of different parameter values to the terminal device where the client is installed. In this way, when a user uses the client to perform live streaming or short video shooting, photo shooting, etc., the stylization processing option can be provided in the shooting interface. After the user selects the stylization processing function, the user can be provided with multiple selectable styles. After the user selects one of the styles, the second algorithm model can be locally run on the terminal device where the client is installed. Correspondingly, the real face image can be extracted from the image collected by the terminal device, and the second algorithm model and the parameter value corresponding to the currently selected style can be used to generate a stylized face image with the corresponding style. Then, the stylized face image can be pasted back to the image collected by the terminal device, so that the user who shoots can view the stylized face image in the terminal device screen. In addition, the stylized face image can be used to replace the real face image in each frame of the generated short video or the video recorded in the live streaming.
[0106] The above style processing can be applied in various practical scenarios. For example, in a commodity information service system, a certain service can be provided, which can be directed to buyers, sellers and various users. Among them, when facing the buyer user, the user can publish "buyer show" based on the service, that is, shoot short videos, photos, etc. of the clothes bought from the system and publish them for other buyer users to view. Among them, when shooting short videos or photos, in order to protect user privacy, enhance fun, etc., a style processing option can be provided after the user triggers shooting, at which time the algorithm model provided by the embodiments of the present application can be used on the client to perform style processing on the real face entering the shooting picture. In addition, when facing the seller user, the seller user can shoot "seller show", publish videos or photos of the goods they sell, etc. for the buyer user to browse, helping the user make a purchase decision. Among them, the seller user may need a real person to appear in the process of shooting their own goods, at which time the real face is processed by style processing, and since the face image after style processing is usually more beautiful, the seller user can not need to find a very beautiful model, and can use the style processing function to beautify and improve the visual effect. In addition, for sellers selling goods for a certain specific group of people, during the shooting of short videos or live broadcasts, the face of the host can be converted into a face image with the style of the corresponding group of people through style processing, thereby playing a role in creating a sales atmosphere. For example, a certain seller user mainly sells children's clothes, but when recording an explanation video or live broadcast, an adult host is needed to explain the goods, at which time the face image of the host is converted into a child-style face image, at which time the face image can be more consistent with the atmosphere of the current sales scene, thereby playing a role in creating a sales atmosphere, etc.
[0107] In summary, by the embodiments of the present application, the generation process of the stylized processing model can be divided into multiple stages. In the first stage, a first data set composed of multiple stylized face images with a target style and a second data set composed of multiple real face images can be obtained. The stylized face images in the first data set and the real face images in the second data set can not necessarily have a pairing relationship. Then, in the second stage, the first data set and the second data set can be used as training samples to train the first algorithm model by an unsupervised learning manner, so as to obtain a third data set composed of real face images and stylized face images with a pairing relationship by the first algorithm model. In the third stage, the third data set can be used as training samples to train the second algorithm model by a supervised learning manner, and then the second algorithm model can be distributed to a terminal device where a client is located, so that the client converts a collected real face image into a stylized face image with the target style by using the second algorithm model. In this way, the decoupling of the stylized data production, the paired image data production, and the mobile terminal image translation three tasks can be realized. Moreover, the designer, the expert, or the like does not need to design the stylized face image paired with the real face image, and the training of the second algorithm model can be completed at a lower cost, so that the second algorithm model can output a stylized face image with a certain style corresponding to the input of a real face image. Moreover, since the training of the second algorithm model can be realized by a supervised manner, the control of the operation amount can be realized, so that the second algorithm model can be run on the client to realize the real-time stylized processing on the client.
[0108] In optional embodiments, the algorithms in each stage, the algorithms or data between stages can also be optimized and improved. For example, in the first stage, the pre-training of the algorithm model by using the real face data, the fixing of part of the parameters, and the like can reduce the demand for the number of stylized face image raw materials, and a third model can be trained by using a small amount of stylized face image raw materials to generate a large amount of stylized face data. Style innovation can also be realized by parameter fusion of different styles and the like. In the second stage, the average error L1 loss term about the background area can be added to the generation network part, the background area loss term can be ignored, and the branch related to the discrimination of the whole image can be removed from the discrimination network part, so as to realize the fixation of the background image in the stylized process and avoid the background blur caused by the stylized processing of the image of the background area. In the third stage, the loss and the perception function can be increased, the stylized face image after the edge blur processing can be added to the training data, so that the stylized face image generated by the second algorithm model has an improved edge; the robustness of the algorithm model can be improved by the data enhancement processing of the face image in the third data set, and the like.
[0109] It should be noted that the embodiments of the present application can involve the use of user data. In actual applications, user-specific personal data can be used in the schemes described herein within the scope permitted by applicable laws and regulations, provided that the applicable laws and regulations are met (for example, the user has given explicit consent, the user has been effectively notified, etc.).
[0110] Corresponding to the foregoing method embodiments, the embodiments of the present application also provide a face image processing apparatus, see Figure 3 The apparatus can include:
[0111] A data generation unit 301 configured to obtain a first data set composed of a plurality of stylized face images having a target style and a second data set composed of a plurality of real face images, wherein there is no pairing relationship between the stylized face images in the first data set and the real face images in the second data set;
[0112] An unsupervised learning unit 302 configured to train a first algorithm model by an unsupervised learning manner with the first data set and the second data set as training samples, so as to obtain a third data set composed of real face images and stylized face images having a pairing relationship through the first algorithm model;
[0113] A supervised learning unit 303 configured to train a second algorithm model by a supervised learning manner with the third data set as training samples, so as to distribute the second algorithm model to a terminal device where a client is located, and the client is configured to convert a real face image collected by the terminal device into a stylized face image of the target style through the second algorithm model.
[0114] The data generation unit can be specifically configured to:
[0115] An original material collection sub-unit configured to collect a first number of stylized face image original materials about the target style;
[0116] A generation model training sub-unit configured to train a third algorithm model by using the original materials;
[0117] A stylized face image generation sub-unit configured to obtain a second number of stylized face images according to the trained third algorithm model, so as to compose the first data set;
[0118] The first number is less than the second number.
[0119] In order to reduce the demand of the third algorithm model for the amount of training data, the apparatus can further include:
[0120] The pre-training unit is configured to pre-train the third algorithm model by using a plurality of real face images.
[0121] The generation model training subunit can be specifically configured to:
[0122] The second training is performed on the basis of the third algorithm model after the pre-training is completed by using the original material.
[0123] In addition, before the second training is performed, the numerical values of part of the parameters obtained by the pre-training in the third algorithm model can be fixed, and the part of the parameters are related to common features between real face images and stylized face images.
[0124] Further, the device can further include:
[0125] The first parameter fusion unit is configured to fuse the pre-training result and the second training result in terms of parameter values, so as to correct errors of the parameter values obtained by the second training.
[0126] The stylized face image generation subunit can be specifically configured to:
[0127] The second number of stylized face images are obtained according to the third algorithm model after the parameter value fusion.
[0128] In addition, the device can further include:
[0129] The fourth algorithm model training unit is configured to train a fourth algorithm model for generating a random vector on the basis of the third algorithm model after the pre-training is completed.
[0130] The random vector generation unit is configured to generate a random vector by the fourth algorithm model as an input of the third algorithm model to control the distribution of the stylized face images output by the third algorithm model when the second number of stylized face images are obtained according to the third algorithm model after the second training.
[0131] Alternatively, in another mode, the data generation unit can specifically include:
[0132] The multi-style original material collection subunit is configured to collect stylized face image original materials about at least two styles respectively.
[0133] The multi-style training subunit is configured to train a third algorithm model by using original materials corresponding to the at least two styles respectively, to obtain at least two groups of parameter values corresponding to the at least two styles respectively.
[0134] The fusion subunit is configured to obtain fused parameter values by fusing the at least two groups of parameter values.
[0135] A generating sub-unit is configured to obtain a plurality of stylized face images with the target style according to the third algorithm model and the fused parameter value.
[0136] In addition, in the second stage, the first algorithm model can include a generating network part and a discriminative network part; wherein, an average error L1 loss term about the background region can be added to the generating network part, an adversarial loss term for the background region can be ignored, and a branch related to the discrimination of the whole image can be removed from the discriminative network part, so as to avoid the image of the background region being stylized in the process of converting the real face image into the stylized face image.
[0137] In addition, the pixel loss function of the second algorithm model can further include an adversarial loss and a perception function.
[0138] Further, the device can further include:
[0139] A blur processing unit is configured to perform edge recognition on the stylized face images in the third data set and perform blur processing on the edge part.
[0140] The supervised learning unit can be specifically configured to:
[0141] The third data set and the stylized face images with the blurred edge part are used as training samples, and the second algorithm model is trained by a supervised learning manner, and an edge enhancement adversarial loss is provided in the second algorithm model, so that the stylized face images generated by the second algorithm model have edge enhancement.
[0142] The second algorithm model includes a discriminative network, and the discriminative network has global discrimination ability, local discrimination ability and attention mechanism.
[0143] Further, the device can further include:
[0144] A data enhancement processing unit is configured to perform data enhancement processing on the face images in the third data set, and the data enhancement processing includes random cropping, random scaling or random optical distortion processing.
[0145] The client includes a client provided by a commodity information service system;
[0146] The client is configured to:
[0147] After receiving a request of the user to shoot a short video / photo of a target commodity, a stylization processing option is provided.
[0148] In response to a request initiated through the stylization processing option, a target style is determined, and a real face image is cropped from an original image collected by a terminal device;
[0149] The real face image is converted into a stylized face image of the target style through a second algorithm model corresponding to the target style, and the stylized face image is pasted back into the original image.
[0150] In addition, the embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the steps of the method in any one of the preceding method embodiments.
[0151] And an electronic device, comprising:
[0152] one or more processors; and
[0153] a memory associated with the one or more processors, the memory being configured to store program instructions, which, when executed by the one or more processors, perform the steps of the method in any one of the preceding method embodiments.
[0154] wherein, Figure 4 An exemplary architecture of the electronic device is shown, which can specifically include a processor 410, a video display adapter 411, a disk drive 412, an input / output interface 413, a network interface 414, and a memory 420. The processor 410, the video display adapter 411, the disk drive 412, the input / output interface 413, the network interface 414, and the memory 420 can be communicatively connected through a communication bus 430.
[0155] The processor 410 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided by the present application.
[0156] The memory 420 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 420 can store an operating system 421 for controlling the operation of the electronic device 400, a basic input / output system (BIOS) for controlling the low-level operation of the electronic device 400. In addition, a web browser 423, a data storage management system 424, and a face image processing system 425, etc. can also be stored. The face image processing system 425 described above can be an application program for implementing the operations of the above steps in the embodiments of the present application. In summary, when the technical solutions provided by the present application are implemented by software or firmware, the relevant program codes are stored in the memory 420 and executed by the processor 410.
[0157] The input / output interface 413 is configured to connect an input / output module to realize information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0158] The network interface 414 is configured to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0159] The bus 430 includes a path for transmitting information between various components (such as the processor 410, the video display adapter 411, the disk drive 412, the input / output interface 413, the network interface 414, and the memory 420) of the device.
[0160] It should be noted that although the above device only shows the processor 410, the video display adapter 411, the disk drive 412, the input / output interface 413, the network interface 414, the memory 420, the bus 430, etc., in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only contain the components necessary to implement the solutions of the present application, and does not have to contain all the components shown in the figure.
[0161] Those skilled in the art can clearly understand the application by the description of the above embodiments that the application can be implemented by means of software and the necessary universal hardware platforms. Based on such an understanding, the technical solutions of the application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, and the like) execute the methods described in each embodiment or some parts of the embodiments of the application.
[0162] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the system or the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts can be referred to the part of the method embodiment. The above-described system and system embodiment are merely illustrative, and the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to actual needs, some or all of the modules can be selected to achieve the purpose of the embodiment. Those skilled in the art can understand and implement it without creative labor.
[0163] The face image processing method and the electronic device provided by the application are described in detail above, and the principle and implementation manner of the application are described by applying specific examples. The above embodiment is only used to help understand the method of the application and its core idea; meanwhile, for those skilled in the art, according to the idea of the application, the specific implementation manner and application range can be changed. In conclusion, the content of the specification should not be understood as a limitation of the application.
Claims
1. A facial image processing method, characterized in that, include: Obtain a first dataset consisting of multiple stylized face images with a target style, and a second dataset consisting of multiple real face images, wherein there is no pairing relationship between the stylized face images in the first dataset and the real face images in the second dataset. The step of obtaining a first dataset consisting of multiple stylized face images with a target style includes: The third algorithm model is pre-trained using multiple real face images to obtain a set of parameter values, so that the third algorithm model outputs a real face image based on the input random vector under the condition of the set of parameter values. A first number of original stylized face images of the target style are collected, and these original images are used as training samples to perform secondary training based on the parameter value set of the pre-trained third algorithm model. This allows the third algorithm model to output stylized face images with the target style based on the input random vector. Before the secondary training, the values of some parameters obtained through the pre-training in the third algorithm model are fixed. These parameters are parameters related to the common features between the real face image and the stylized face image. A second number of stylized face images are obtained based on the third algorithm model completed after secondary training, to form the first dataset; The first dataset and the second dataset are used as training samples. The first algorithm model is trained through unsupervised learning so that a third dataset consisting of real face images and stylized face images with paired relationships can be obtained through the first algorithm model. The third dataset is used as training samples to train the second algorithm model through supervised learning, so that the second algorithm model can be distributed to the terminal device where the client is located. The client is used to convert real face images collected by the terminal device into stylized face images of the target style through the second algorithm model.
2. The method according to claim 1, characterized in that, Also includes: By fusing the parameter values of the pre-training results and the secondary training results, error correction can be performed on the parameter values obtained from the secondary training. The step of obtaining a second number of stylized face images based on the trained third algorithm model includes: The second number of stylized face images are obtained based on the third algorithm model after parameter fusion.
3. The method according to claim 1, characterized in that, Also includes: Based on the pre-trained third algorithm model, the fourth algorithm model for generating random vectors is trained. When obtaining a second number of stylized face images based on the third algorithm model obtained from secondary training, a random vector is generated by the fourth algorithm model and used as the input of the third algorithm model to control the distribution of the stylized face images output by the third algorithm model.
4. The method according to claim 1, characterized in that, The first algorithm model includes a generator network and a discriminator network. The method further includes: In the generative network part, an average error L1 loss term for the background region is added, while the adversarial loss term for the background region is ignored. In the discriminative network part, the branch involved in discriminating the entire image is removed, so as to avoid the background region image being stylized during the process of converting the real face image into a stylized face image.
5. The method according to claim 1, characterized in that, The pixel loss function of the second algorithm model includes adversarial loss and a perceptual function.
6. The method according to claim 1, characterized in that, Also includes: Edge recognition is performed on the stylized face images in the third dataset, and the edge parts are blurred. The step of using the third dataset as training samples to train the second algorithm model through supervised learning includes: The third dataset and the stylized face images with blurred edges are used as training samples. The second algorithm model is trained using supervised learning. An edge enhancement adversarial loss is provided in the second algorithm model so that the stylized face images generated by the second algorithm model can achieve edge enhancement.
7. The method according to claim 1, characterized in that, The second algorithm model includes a discriminant network, which has global discrimination ability, local discrimination ability, and attention mechanism.
8. The method according to any one of claims 1 to 7, characterized in that, The client includes the client provided by the commodity information service system; The client is used for: Upon receiving a user's request to livestream or shoot a short video / photo of a target product, provide stylization options; In response to a request initiated via the stylization processing option, a target style is determined, and a real face image is extracted from the original image captured by the terminal device. The real face image is converted into a stylized face image of the target style using a second algorithm model corresponding to the target style, and the stylized face image is then pasted back into the original image.
9. A facial image processing method, characterized in that, include: Obtain a first dataset consisting of multiple stylized face images with a target style, and a second dataset consisting of multiple real face images, wherein there is no pairing relationship between the stylized face images in the first dataset and the real face images in the second dataset. The target style includes a style obtained by fusing multiple different styles; the acquisition of a first dataset consisting of multiple stylized face images with the target style includes: Collect original source materials of stylized facial images in at least two different styles; The third algorithm model is trained using the original materials corresponding to the at least two styles respectively, to obtain at least two sets of parameter values corresponding to the at least two styles respectively; By fusing the at least two sets of parameter values, a fused parameter value is obtained, so that the third algorithm model can generate multiple stylized face images with the target style based on the input random vector under the condition of the fused parameter value. The first dataset and the second dataset are used as training samples. The first algorithm model is trained through unsupervised learning so that a third dataset consisting of real face images and stylized face images with paired relationships can be obtained through the first algorithm model. The third dataset is used as training samples to train the second algorithm model through supervised learning, so that the second algorithm model can be distributed to the terminal device where the client is located. The client is used to convert real face images collected by the terminal device into stylized face images of the target style through the second algorithm model.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1 to 9.
11. An electronic device, characterized in that, include: One or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Character image gender conversion model training method and device, image generation method and device
CN114078082A