Model training method and apparatus

By progressively training multiple models and leveraging the ability of generative adversarial networks and loss functions to optimize each model, the problem of unstable beautification effects in existing technologies has been solved, achieving high-quality facial image processing in different scenarios.

CN117196007BActive Publication Date: 2026-05-29VIVO MOBILE COMM CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2023-09-21
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to train facial image processing models that produce good and stable beautification effects, especially in scenarios with poor shooting quality or diverse skin textures, where issues such as over-beautification or incomplete removal of blemishes and imperfections arise.

Method used

By progressively training multiple models, the first model is trained using randomized model parameters, the second model inherits its parameters, the third model is trained, and so on up to the sixth model, gradually improving the beautification effect. Generative adversarial networks and loss functions are used to optimize the capabilities of each model.

Benefits of technology

It reduces the difficulty of model training, improves the stability and quality of beautification effects, and ensures high-quality facial image processing in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117196007B_ABST
    Figure CN117196007B_ABST
Patent Text Reader

Abstract

The application discloses a model training method and device, and belongs to the field of image processing. The method comprises the following steps: training a first model based on random model parameters to obtain a second model, wherein the first model is used for performing first processing on a face image; training a third model based on first model parameters of the second model to obtain a fourth model, wherein the third model is used for performing second processing on the face image, and the second processing is different from the first processing; and training a fifth model based on second model parameters of the fourth model to obtain a sixth model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing, and specifically relates to a model training method and apparatus. Background Technology

[0002] As electronic devices become increasingly feature-rich, users are paying particular attention to the beautification effects on portraits when taking photos. Currently, convolutional neural network models are typically used to learn a single-task end-to-end beautification mapping image to achieve beautification of portrait images.

[0003] However, directly using single-task end-to-end learning of beautification mapping images makes training beautification models more difficult and prevents the trained beautification models from achieving high-quality facial retouching and beautification effects. Especially in input scenarios with poor shooting quality or in scenarios where users have diverse skin textures, there is a high probability of over-beautification or incomplete removal of blemishes and imperfections; thus, it is impossible to train a facial image processing model with good and stable beautification effects. Summary of the Invention

[0004] The purpose of this application is to provide a model training method and apparatus that can solve the problem of being unable to train a stable face image processing model with good beautification effects.

[0005] In a first aspect, embodiments of this application provide a model training method, which includes: training a first model based on random model parameters to obtain a second model, the first model being used to perform a first processing on a face image; training a third model based on the first model parameters of the second model to obtain a fourth model, the third model being used to perform a second processing on the face image, the second processing being different from the first processing; and training a fifth model based on the second model parameters of the fourth model to obtain a sixth model.

[0006] Secondly, embodiments of this application provide a model training apparatus, which includes a training module; the training module is used to train a first model based on random model parameters to obtain a second model, the first model being used to perform a first processing on a face image; the training module is also used to train a third model based on the first model parameters of the second model to obtain a fourth model, the third model being used to perform a second processing on the face image, the second processing being different from the first processing; the training module is also used to train a fifth model based on the second model parameters of the fourth model to obtain a sixth model.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0010] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0011] In this embodiment, a first model can be trained based on random model parameters to obtain a second model, which is used to perform a first processing on a face image. A third model is then trained based on the first model parameters of the second model to obtain a fourth model, which is used to perform a second processing on the face image, different from the first processing. Finally, a fifth model is trained based on the second model parameters of the fourth model to obtain a sixth model. This approach allows for the training of the fifth model based on the model parameters of the second model obtained from the first model used for the first processing of the face image, and the model parameters of the fourth model obtained from the third model used for the second processing of the face image. This enables the training of progressively more complex face image processing tasks from easy to difficult, reducing the training difficulty of the fifth model and resulting in a stable and effective face image processing model with good beautification effects. Attached Figure Description

[0012] Figure 1 This is one of the flowcharts of the model training method provided in the embodiments of this application;

[0013] Figure 2 This is the second flowchart of the model training method provided in the embodiments of this application;

[0014] Figure 3 This is the third flowchart of the model training method provided in the embodiments of this application;

[0015] Figure 4 This is a schematic diagram of a method for training a first model using a seventh model in the model training method provided in the embodiments of this application;

[0016] Figure 5 This is a schematic diagram of a method for training a third model using a seventh model in the model training method provided in the embodiments of this application;

[0017] Figure 6This is a schematic diagram of the method for training the fifth model using the seventh model in the model training method provided in the embodiments of this application;

[0018] Figure 7 This is a schematic diagram of the model training device provided in the embodiments of this application;

[0019] Figure 8 This is a schematic diagram of the electronic device provided in the embodiments of this application;

[0020] Figure 9 This is a hardware schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0022] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0023] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0024] The model training method and apparatus provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0025] As smartphones become increasingly feature-rich, users pay particular attention to their camera capabilities when choosing a smartphone, with portrait image quality and aesthetics being key considerations, especially beautification effects. Almost all smartphones include a portrait mode, which can be used to capture beautified portraits. Currently, portrait beautification primarily achieves its effects through traditional computer vision (CV) algorithms or by using convolutional neural networks to learn a single-task, end-to-end beautification mapping image.

[0026] However, most portrait beautification methods implemented using the aforementioned CV algorithms suffer from problems such as over-beautification and incomplete removal of large blemishes and imperfections. Directly utilizing the single-task, end-to-end learning of the beautification mapping map is too difficult and cannot effectively train a beautification model with good and stable beautification results. Especially in scenarios with poor image quality or where users have diverse skin textures, simply using this CV algorithm or a single-task, single-mapping image face retouching training architecture cannot achieve high-quality face retouching and beautification effects, and is likely to result in over-beautification or incomplete removal of blemishes and imperfections. Therefore, when implementing portrait beautification, the relevant solutions cannot train a facial image processing model with good and stable beautification results.

[0027] To address the aforementioned problems, embodiments of this application provide a model training method and apparatus. The model training method provided in this application can be applied to the training scenarios of beautification models.

[0028] In the model training method provided in this application embodiment, a first model can be trained based on random model parameters to obtain a second model. The first model is used to perform a first processing on a face image. Based on the first model parameters of the second model, a third model is trained to obtain a fourth model. The third model is used to perform a second processing on the face image, which is different from the first processing. Based on the second model parameters of the fourth model, a fifth model is trained to obtain a sixth model. Through this scheme, since the training of the fifth model can be based on the model parameters of the second model obtained from the first model used for the first processing of the face image, and the model parameters of the fourth model obtained from the third model used for the second processing of the face image, the progressively more difficult face image processing tasks can be trained from easy to difficult, reducing the training difficulty of the fifth model and thus training a face image processing model with good beautification effects and stability.

[0029] It should be noted that the model training method provided in this application can be executed by a model training device, an electronic device, or a functional module within an electronic device. Some embodiments of this application use an electronic device executing the model training method as an example to illustrate the model training method provided in this application.

[0030] Figure 1 A flowchart of the model training method provided in an embodiment of this application is shown. Figure 1 As shown, the model training method provided in this application embodiment may include the following steps 101 to 103.

[0031] Step 101: The electronic device trains the first model based on the random model parameters to obtain the second model.

[0032] The first model mentioned above is used to perform a first process on the face image.

[0033] Optionally, in the embodiments of this application, the first processing described above can be used to remove larger blemishes such as pimples, moles, or scars from a face image.

[0034] Optionally, in this embodiment of the application, based on the image editing principles and processes of image editing software, larger blemishes and scars in the face image are first removed, so that the electronic device can design and train the above-mentioned first model.

[0035] For example, taking the first processing described above for removing blemishes from a facial image as an example, the specific principle of designing the first model is as follows:

[0036] 1. Prepare a training dataset for acne removal; for example, by batch operation, only enable the acne removal layer of the photo editing software on the electronic device to obtain the training dataset Set-despot for acne removal. The training dataset includes at least one training data pair, and each training data pair includes the image data of an input face image and the image data of an acne removal result image corresponding to the input face image.

[0037] 2. Design the first model MulPSGAN-A based on generative adversarial networks. MulPSGAN-A includes a generator network and a discriminator network D_despot. The generator network consists of a generator G and a branch network A. The generator network is used to generate image data of the acne removal result image, and the discriminator network is used to determine whether the generated image data is real.

[0038] 3. Design the loss function loss_all_despot for the first model learning above, as shown in formula (1) below:

[0039] Loss_all_despot=a1*Loss_pixel_despot+b1*Loss_adv_despot+c1*Loss_rec_despot; (1)

[0040] Wherein, Loss_pixel_despot is the pixel loss for acne removal, which directly uses the general L1 loss; Loss_adv_despot is the adversarial loss of the first model, calculated using the softplus function; Loss_rec_despot is the image reconstruction loss for acne removal, which is the average L1 loss from the first conv1 to the fifth conv5 of the Visual Geometry Group (VGG) 19 classification network; a1, b1, and c1 are all weight coefficients, where a1 can be set to 1, b1 can be set to 0.1, and c1 can be set to 1.

[0041] Optionally, in the embodiments of this application, combined with Figure 1 ,like Figure 2 As shown, step 101 above can be implemented through step 101a below.

[0042] Step 101a: The electronic device uses the random model parameters as the initial model parameters of the first model to train the first model and obtain the second model.

[0043] In this embodiment of the application, the aforementioned initial model parameters are the initial training parameters after model initialization.

[0044] Optionally, in this embodiment of the application, the above-mentioned random model parameters are model parameters randomly generated by the electronic device for training the above-mentioned first model.

[0045] Optionally, in this embodiment of the application, the electronic device can use the general generative adversarial network training method adopted by the above-mentioned random model parameter initialization model parameters, alternately backpropagate to update the generator network and discriminator network of the first model, and continuously train and update the parameters of the generator network and the discriminator network until the training converges, so that the first model can perform the above-mentioned first processing on the face image.

[0046] In this embodiment of the application, since the electronic device can use the random model parameters as the initial model parameters of the first model when training the first model, it is convenient to train the first model and simplify the training process of the first model.

[0047] Step 102: The electronic device trains the third model based on the first model parameters of the second model to obtain the fourth model.

[0048] The third model is used to perform a second processing on the face image, which is different from the first processing.

[0049] Optionally, in this embodiment, the second processing described above can be used to remove smaller blemishes such as color blocks, gray blocks, or fat particles in a facial image, and to soften skin texture in a facial image (i.e., skin smoothing). Optionally, in this embodiment, based on the image editing principles and processes of image editing software, after removing larger blemishes and scars from a facial image, a skin smoothing layer is created to remove color blocks, gray blocks, fat particles, and to soften skin texture, etc., thus allowing the electronic device to design and train the third model based on the second model that already possesses the capabilities of the first processing described above.

[0050] For example, taking the second process described above for smoothing skin on a face image as an example, the specific principle of designing the third model described above is as follows:

[0051] 1. Prepare the skin smoothing training dataset; for example, by batch operation, only enable the skin smoothing layer of the photo editing software on the electronic device to obtain the skin smoothing training dataset Set-soften;

[0052] 2. Design the third model MulPSGAN-B based on generative adversarial networks. MulPSGAN-B includes a generator network and a discriminator network D_soften. The generator network consists of the generator G and the branch network B from the second model mentioned above.

[0053] 3. Design the loss function loss_all_soften for the third model learning described above, as shown in formula (2) below:

[0054] Loss_all_soften=a2*Loss_pixel_soften+b2*Loss_adv_soften+c2*Loss_rec_soften; (2)

[0055] Wherein, Loss_pixel_soften is the pixel loss for skin smoothing, which directly uses the general L1 loss; Loss_adv_soften is the adversarial loss of the second model, which is calculated using the softplus function; Loss_rec_soften is the image reconstruction loss for skin smoothing, which is the average L1 loss from the first conv1 to the fifth conv5 of the VGG19 classification network; a2, b2, and c2 are all weight coefficients, where a2 can be set to 1, b2 can be set to 0.1, and c2 can be set to 1.

[0056] It should be noted that since the first model parameters are the model parameters of the second model obtained after the first model has converged, training the third model based on these first model parameters can enable the fourth model to simultaneously possess the capabilities of the first and second processing methods.

[0057] Step 103: The electronic device trains the fifth model based on the second model parameters of the fourth model to obtain the sixth model.

[0058] In this embodiment of the application, the fifth model is used to perform third processing on the face image.

[0059] Optionally, in the embodiments of this application, the third processing described above can be used to uniformly fine-tune the image parameters of the face image, and to blur or soften the details in the face image, etc.

[0060] Optionally, in this embodiment of the application, based on the image editing principles and processes of image editing software, after removing large blemishes and scars from the face image, as well as removing color blocks and gray blocks and other fine-tuning processes, the face image will be uniformly fine-tuned and the overall details will be processed to achieve the final beautification effect. Thus, the electronic device can design and train the fifth model based on the first model that already has the first processing capability and the second model that already has the second processing capability.

[0061] For example, the specific principles behind designing the fifth model described above are as follows:

[0062] 1. Prepare a beauty training dataset; for example, by enabling all the blemish removal and skin smoothing layers in the photo editing software on the electronic device, and batch exporting the already edited beauty result images, you can obtain the beauty processing training dataset Set_beauty.

[0063] 2. Design the fifth model MulPSGAN based on generative adversarial networks. MulPSGAN includes a generator network and a discriminator network D_beauty. The generator network consists of the generator G and branch network C from the second and fourth models mentioned above.

[0064] 3. Design the loss function loss_all_beauty for the fifth model above, as shown in formula (3):

[0065] Loss_all_beauty=a3*Loss_pixel_beauty+b3*Loss_adv_beauty+c3*Loss_rec_beauty; (3)

[0066] Wherein, Loss_pixel_beauty is the pixel loss for beautification, which directly uses the general L1 loss; Loss_adv_beauty is the adversarial loss of the third model, which is calculated using the softplus function; Loss_rec_beauty is the image reconstruction loss for beautification, which is the average L1 loss from the first conv1 to the fifth conv5 of the VGG19 classification network; a3, b3 and c3 are all weight coefficients, where a3 can be set to 1, b3 can be set to 0.1 and c3 can be set to 1.

[0067] It should be noted that since the second model parameters are the model parameters of the fourth model obtained after the third model has converged, training the fifth model based on these second model parameters can enable the sixth model to simultaneously possess the capabilities of the first, second, and third processing.

[0068] Optionally, in the embodiments of this application, combined with Figure 1 ,like Figure 3 As shown, step 102 can be implemented by step 102a, and step 103 can be implemented by step 103a.

[0069] Step 102a: The electronic device uses the parameters of the first model as the initial parameters of the third model to train the third model and obtain the fourth model.

[0070] In this embodiment of the application, the generator G in the third model is the same as the generator G in the second model. Thus, the parameters of the first model can be directly inherited into the third model through the generator G, so that the trained fourth model embeds the underlying feature capabilities of the second model (i.e., the capabilities of the first processing).

[0071] Optionally, in this embodiment of the application, the electronic device can use the first model parameters and the random parameters of the branch network B to alternately backpropagate and update the generator network and discriminator network of the third model based on the general training method of generative adversarial networks. By continuously training and updating the parameters of the generator network and the discriminator network, the training eventually converges to obtain the fourth model.

[0072] It should be noted that after obtaining the fourth model, the generator G shared by the second model and the fourth model already has the capabilities of the first and second processing described above.

[0073] Step 103a: The electronic device uses the second model parameters as the initial model parameters of the fifth model to train the fifth model and obtain the sixth model.

[0074] In this embodiment of the application, the generator G in the fifth model is the same as the generator G in the second model and the fourth model. Thus, the parameters of the second model can be directly inherited into the fifth model through the generator G, so that the trained sixth model embeds the underlying feature capabilities of the second model (i.e., the capabilities of the first processing) and the underlying feature capabilities of the fourth model (i.e., the capabilities of the second processing).

[0075] Optionally, in this embodiment of the application, the electronic device can use the second model parameters and the random parameters of the branch network C to alternately backpropagate and update the generator network and discriminator network of the fifth model based on the general training method of generative adversarial networks. By continuously training and updating the parameters of the generator network and the discriminator network, the training eventually converges to obtain the sixth model.

[0076] It should be noted that after obtaining the sixth model, the generator G shared by the second model, the fourth model, and the sixth model already has the capabilities of the first processing, the second processing, and the third processing.

[0077] In this embodiment, since the electronic device can use the first model parameters as the initial model parameters of the third model to train the third model, and use the second model parameters of the fourth model obtained by training the third model as the initial model parameters of the fifth model to train the fifth model, the learning and training can be carried out in a progressive manner, thereby reducing the difficulty of model training and improving the training effect of the final sixth model.

[0078] Optionally, in the embodiments of this application, each of the first model, the third model, and the fifth model includes a generator network; the generator network included in the first model, the generator network included in the third model, and the generator network included in the fifth model are the same.

[0079] Optionally, in the embodiments of this application, the generator network included in each of the above models can be the generator G mentioned above.

[0080] In this embodiment of the application, since the generator networks included in the first model, the third model and the fifth model are the same, the inheritance of model parameters can be facilitated during model training, thereby improving the training speed of the model.

[0081] For the specific training methods of the first, third, and fifth models mentioned above, please refer to the specific descriptions of training methods for generative adversarial network models in related technologies. To avoid repetition, they will not be repeated here.

[0082] Optionally, in this embodiment, the first model, the third model, and the fifth model can be designed within a large beautification model. In the architecture of this beautification model, the first model can be trained first to learn the capabilities of the first processing; and after the training of the first model converges, the third model is trained based on the weight parameters of the common generator G to learn the capabilities of the second processing; and after the training of the third model converges, the fifth model is trained based on the updated weight parameters of the generator G to learn the capabilities of the third processing; thus, the beautification model that is finally trained can simultaneously possess the capabilities of the first processing, the second processing, and the third processing.

[0083] Thus, the aforementioned progressive network architecture design and training method can greatly reduce the difficulty of directly training the model to learn the beautification function, allowing the model to be trained step by step from easy to difficult based on the model parameters of the previous stage's effect function, greatly improving the training effect and stability.

[0084] Optionally, in this embodiment of the application, after obtaining the sixth model, in the actual test inference stage, only the ability of the first processing plus the second processing is needed, so that the branch network A and branch network B can be directly deleted, so as to improve the model training effect without adding any time overhead.

[0085] In the model training method provided in this application embodiment, since the model parameters of the second model obtained by training the first model for the first processing of face images and the model parameters of the fourth model obtained by training the third model for the second processing of face images can be used when training the fifth model, the face image processing tasks can be trained from easy to difficult, reducing the training difficulty of the fifth model, thereby training a face image processing model with good beautification effect and stability.

[0086] Optionally, in the embodiments of this application, the model training method provided in the embodiments of this application may further include the following step 104.

[0087] Step 104: During the training of the target model, the electronic device uses the seventh model to segment the facial feature image region and the non-facial feature image region in the training image of the target model.

[0088] The seventh model is used to segment facial feature regions and non-facial feature regions in a face image, and the target model includes at least one of the first model, the third model, and the fifth model.

[0089] Optionally, in the embodiments of this application, the aforementioned facial features image region refers to the region in the face image where the eyes, eyebrows, mouth, and other facial features are located; the aforementioned non-facial features image region refers to the region in the face image other than the facial features image region.

[0090] It should be noted that when beautifying facial images, the processing level and detail differ between the aforementioned facial feature areas and the aforementioned non-facial feature areas. For example, the eyebrow area needs to be rendered clearly, with individual hairs removed; the eye area needs to emphasize contrast and depth; the mouth area needs to highlight contours and lip lines; while the non-facial feature areas are processed with a uniform degree of detail. Therefore, the seventh model described above allows the target model to consciously learn the directional information regarding the different processing requirements for facial feature areas and non-facial feature areas.

[0091] Optionally, in this embodiment of the application, the seventh model may include a facial skin segmentation model and a facial region multi-task discriminator, with the specific network architecture as follows:

[0092] 1. Since the facial features image region and the non-facial features image region need to learn different processing effects and learning directions, and if the above target model is directly and uniformly allowed to learn the overall effect of the face image, it will increase the training difficulty and make the training fluctuate, affecting the training effect. Therefore, the above facial features skin segmentation model MulPSGAN-D is designed. MulPSGAN-D consists of the above generator G and branch network D.

[0093] First, a training dataset Set-segment for facial feature and skin segmentation can be prepared. This training dataset includes at least one training data pair, each pair consisting of image data of a face input image and image data of the corresponding facial feature and skin segmentation image. Then, a simple L2 loss calculation for pixels is designed, and the MulPSGAN-D network is trained using a general Artificial Intelligence (AI) network training method, continuously backpropagating to update the gradient. After the training converges, MulPSGAN-D will have the ability to segment facial features and skin. Thus, the generator G mentioned above also has the ability to distinguish regions of facial features and skin during the training process of MulPSGAN-D.

[0094] 2. To further enhance the learning ability of facial feature image regions for detail effects, a multi-task discriminator for facial feature regions can be designed. This discriminator uses multiple facial feature region discriminator networks to discriminate and learn the detail effects of facial feature regions. In this way, the global discriminator can focus on the overall face image effect, while the facial feature region discriminator can selectively discriminate and learn the effects of local image regions. Through joint discrimination training of the global discriminator and the facial feature discriminator, the overall face effect can be guaranteed, while the learning and improvement of facial feature region effects can be highlighted, thus further improving the overall performance of the algorithm.

[0095] The aforementioned facial discriminator can include a left eyebrow discriminator, a right eyebrow discriminator, a left eye discriminator, a right eye discriminator, and a mouth discriminator. During training, in addition to using the aforementioned global discriminator to discriminate the overall facial effect, a local discriminator can also be used to discriminate the facial features separately. This is mainly reflected in the adversarial loss of the beautification function. The global adversarial loss without the addition of a local discriminator is expressed as follows in formula (4):

[0096] Loss_adv_beauty =-softplus(D(y)); (4)

[0097] The overall adversarial loss of the network-generated face image, after adding the facial feature discriminator, is expressed as follows by formula (5):

[0098] Loss_adv_beauty=-softplus(D(y))+(-softplus(D(y_lefteye))+(-softplus(D(y_righteye)))+

[0099] (-softplus(D(y_lefteyebrow)))+(-softplus(D(y_righteyebrow)))(-softplus(D(y_mouth)));(5)

[0100] It can be seen that after adding the discrimination loss for local facial features, the training loss of the facial feature area effect will be backpropagated and updated, thereby enabling the network to learn the ability to focus on the local facial feature area to learn details and improve.

[0101] For example, taking the target model as including the first model, the third model, and the fifth model, as an example, Figure 4As shown, the electronic device first trains the first model according to the channel rules of the image editing software. At the same time, during the training of the first model, the seventh model is used to segment the facial feature image region and the non-facial feature image region in the training image, so as to guide the first model to learn different levels of attention to the processing effect of the facial feature image region and the non-facial feature image region of the face image.

[0102] Then, as Figure 5 As shown, after the first model is trained and converged, the second model obtained directly feeds the parameters of the first model in the generator G to the third model, so that the third model does not start training from the random model parameters, but starts training from the first model parameters. At the same time, during the training of the third model, the seventh model is used to segment the facial feature image region and the non-facial feature image region in the training image, so as to guide the third model to learn different levels of attention to the processing effect of the facial feature image region and the non-facial feature image region of the face image.

[0103] Finally, as Figure 6 As shown, the fourth model obtained after the third model converges will directly feed the second model parameters from the generator G to the fifth model. This means that the fifth model will not start training from the random model parameters, but from the second model parameters. At the same time, during the training of the fifth model, the seventh model is used to segment the facial feature regions and non-facial feature regions in the training image. This guides the fifth model to learn different levels of attention to the processing effects of facial feature regions and non-facial feature regions in the face image. This can improve the training effect and stability of the entire model training process.

[0104] In this embodiment of the application, since the electronic device can use the seventh model to segment the facial feature image region and non-facial feature image region in the training image of the target model during the training process of the target model, the target model can have an attention mechanism with different processing degrees for facial feature images and skin images, thereby achieving better model performance and algorithm stability.

[0105] The above-described method embodiments, or various possible implementations of the method embodiments, can be executed individually, or, provided there are no contradictions, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0106] The model training method provided in this application can be executed by a model training device. This application uses the example of a model training device executing the model training method to illustrate the model training device provided in this application.

[0107] like Figure 7 As shown in the figure, this application embodiment provides a model training device 70, which may include a training module 71.

[0108] The training module 71 can be used to train a first model based on random model parameters to obtain a second model, which is used to perform a first processing on the face image. The training module 71 can also be used to train a third model based on the first model parameters of the second model to obtain a fourth model, which is used to perform a second processing on the face image. The training module 71 can also be used to train a fifth model based on the second model parameters of the fourth model to obtain a sixth model.

[0109] In one possible implementation, the training module 71 can be used to train the first model by using the random model parameters as the initial model parameters of the first model, thereby obtaining the second model.

[0110] In one possible implementation, the training module 71 can be specifically used to use the first model parameters as the initial model parameters of the third model to train the third model and obtain the fourth model; and to use the second model parameters as the initial model parameters of the fifth model to train the fifth model and obtain the sixth model.

[0111] In one possible implementation, the model training device 70 may further include a processing module. The processing module may be used to segment the facial feature region and non-facial feature region in the training image of the target model using a seventh model during the training process of the target model by the training module 71. The seventh model is used to segment the facial feature region and non-facial feature region in the face image, and the target model includes at least one of the first model, the third model, and the fifth model described above.

[0112] In one possible implementation, each of the first, third, and fifth models mentioned above includes a generator network. The generator network included in the first model, the third model, and the fifth model is the same.

[0113] In the model training apparatus provided in this application embodiment, since the training of the fifth model can be based on the model parameters of the second model obtained by training the first model for the first processing of face images, and the model parameters of the fourth model obtained by training the third model for the second processing of face images, the face image processing tasks can be trained from easy to difficult, reducing the training difficulty of the fifth model, thereby training a face image processing model with good beautification effect and stability.

[0114] The model training device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0115] The model training device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0116] The model training apparatus provided in this application embodiment can implement all the processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0117] like Figure 8 As shown, this application embodiment also provides an electronic device 800, including a processor 801 and a memory 802. The memory 802 stores a program or instructions that can run on the processor 801. When the program or instructions are executed by the processor 801, they implement the various steps of the above-described model training method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0118] It should be noted that the electronic devices in the embodiments of this application include mobile electronic devices and non-mobile electronic devices.

[0119] Figure 9 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0120] like Figure 9As shown, the electronic device 1000 includes, but is not limited to, components such as: radio frequency unit 1001, network module 1002, audio output unit 1003, input unit 1004, sensor 1005, display unit 1006, user input unit 1007, interface unit 1008, memory 1009, and processor 1010.

[0121] Those skilled in the art will understand that the electronic device 1000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 9 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0122] The processor 1010 can be used to train a first model based on random model parameters to obtain a second model, which is used to perform a first processing on a face image; and to train a third model based on the first model parameters of the second model to obtain a fourth model, which is used to perform a second processing on the face image; and to train a fifth model based on the second model parameters of the fourth model to obtain a sixth model.

[0123] In one possible implementation, the processor 1010 can be used to train the first model by using the aforementioned random model parameters as the initial model parameters of the first model, thereby obtaining the aforementioned second model.

[0124] In one possible implementation, the processor 1010 can be used to train the third model by using the first model parameters as the initial model parameters of the third model to obtain the fourth model; and to train the fifth model by using the second model parameters as the initial model parameters of the fifth model to obtain the sixth model.

[0125] In one possible implementation, the processor 1010 can also be used to segment the facial feature image region and the non-facial feature image region in the training image of the target model during the training process of the target model using a seventh model. The seventh model is used to segment the facial feature image region and the non-facial feature image region in the face image, and the target model includes at least one of the first model, the third model, and the fifth model described above.

[0126] In one possible implementation, each of the first, third, and fifth models mentioned above includes a generator network. The generator network included in the first model, the third model, and the fifth model is the same.

[0127] In the electronic device provided in this application embodiment, since the model parameters of the second model obtained by training the first model for the first processing of face images and the model parameters of the fourth model obtained by training the third model for the second processing of face images can be used when training the fifth model, the face image processing tasks can be trained from easy to difficult, reducing the training difficulty of the fifth model, thereby training a face image processing model with good beautification effect and stability.

[0128] It should be understood that, in this embodiment, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0129] The memory 1009 can be used to store software programs and various data. The memory 1009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1009 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0130] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor 1010.

[0131] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described model training method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0132] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0133] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described model training method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0134] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0135] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the model training method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0136] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0137] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0138] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A model training method, characterized in that, The method includes: Based on random model parameters, a first model is trained to obtain a second model. The first model is used to perform a first process on face images. The first model includes a generator G and a branch network A. Based on the first model parameters of the second model and the random parameters of the branch network B, a third model is trained to obtain a fourth model. The third model is used to perform a second processing on the face image. The second processing is different from the first processing. The third model includes the generator G and the branch network B. The first model parameters are inherited into the third model through the generator G. Based on the second model parameters of the fourth model and the random parameters of the branch network C, the fifth model is trained to obtain the sixth model. The fifth model includes the generator G and the branch network C, and the second model parameters are inherited into the fifth model through the generator G.

2. The method according to claim 1, characterized in that, The process of training the first model based on random model parameters to obtain the second model includes: The random model parameters are used as the initial model parameters of the first model to train the first model and obtain the second model.

3. The method according to claim 1 or 2, characterized in that, The method further includes: During the training of the target model, the seventh model is used to segment the facial feature image region and the non-facial feature image region in the training image of the target model. The seventh model is used to segment facial feature regions and non-facial feature regions in a face image, and the target model includes at least one of the first model, the third model, and the fifth model.

4. A model training device, characterized in that, The device includes a training module; The training module is used to train a first model based on random model parameters to obtain a second model. The first model is used to perform a first process on a face image. The first model includes a generator G and a branch network A. The training module is also used to train a third model based on the first model parameters of the second model and the random parameters of the branch network B to obtain a fourth model. The third model is used to perform a second processing on the face image. The second processing is different from the first processing. The third model includes the generator G and the branch network B. The first model parameters are inherited into the third model through the generator G. The training module is also used to train a fifth model based on the second model parameters of the fourth model and the random parameters of the branch network C to obtain a sixth model. The fifth model includes the generator G and the branch network C, and the second model parameters are inherited into the fifth model through the generator G.

5. The apparatus according to claim 4, characterized in that, The training module is specifically used to train the first model by using the random model parameters as the initial model parameters of the first model, thereby obtaining the second model.

6. The apparatus according to claim 4 or 5, characterized in that, The device also includes a processing module; The processing module is used to segment the facial feature image region and non-facial feature image region in the training image of the target model using the seventh model during the training process of the target model in the training module. The seventh model is used to segment facial feature regions and non-facial feature regions in a face image, and the target model includes at least one of the first model, the third model, and the fifth model.