Method, device, electronic device and storage medium for generating stylized images
The construction of a stylized image generation model through transfer learning and a small number of training samples solves the problems of high costs and difficult sample acquisition in the prior art, and realizes efficient generation of multiple stylized images.
Patent Information
- Application Number
- CN202210067042.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-20
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2042-01-20
AI Technical Summary
The prior art requires a large number of training samples when generating stylized images, resulting in high cost and difficulty in building a model of a specific style type.
The parameters of the facial image generation model are obtained through transfer learning, the first and second samples to be trained are constructed, and a small number of training samples are used for training, the target sample generation model of the two style types is fused to generate the fused stylized image.
Efficiently constructing target style data generation models without large amounts of training samples reduces model construction costs and generates multiple stylized images.
Smart Images

Figure CN114429418B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of data processing technology, and in particular to a method, device, electronic device, and storage medium for generating a stylized image. Background Art
[0002] With the continuous development of image processing technology, users can use a variety of applications to process images so that the processed images can present the style they expect.
[0003] In the existing technology, the relevant algorithms used for image processing often need to use a large amount of data to train the model before providing corresponding services to users. However, this method consumes a lot of costs. At the same time, when relevant images of a certain style type cannot be obtained, it is impossible to build an effective algorithm model for this style type. Summary of the Invention
[0004] The embodiments of the present disclosure provide a method, apparatus, electronic device, and storage medium for generating stylized images, which can efficiently construct a target style data generation model without requiring a large number of training samples with two style types to be fused, thereby reducing the cost consumed in the model construction process.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for generating a stylized image, the method comprising:
[0006] Obtaining model parameters to be transferred of the facial image generation model, so as to construct a first sample generation model to be trained and a second sample generation model to be trained based on the model parameters to be transferred;
[0007] Training the first to-be-trained sample generation model based on training samples of the first style type to obtain a first target sample generation model;
[0008] Training the second to-be-trained sample generation model based on the training samples of the second style type to obtain a second target sample generation model;
[0009] Based on the to-be-fitted model parameters of the first target sample generation model and the second target sample generation model, a target style data generation model is determined to generate a stylized image that fuses the first style type and the second style type based on the target style generation model.
[0010] In a second aspect, an embodiment of the present disclosure further provides a device for generating a stylized image, the device comprising:
[0011] A module for obtaining model parameters to be transferred, used to obtain model parameters to be transferred of the facial image generation model, so as to construct a first sample generation model to be trained and a second sample generation model to be trained based on the model parameters to be transferred;
[0012] A first to-be-trained sample generation model training module, configured to train the first to-be-trained sample generation model based on training samples of a first style type to obtain a first target sample generation model;
[0013] A second to-be-trained sample generation model training module, configured to train the second to-be-trained sample generation model based on training samples of a second style type to obtain a second target sample generation model;
[0014] a target style data generation model determination module, configured to determine a target style data generation model based on the to-be-fitted model parameters of the first target sample generation model and the second target sample generation model, so as to generate a stylized image that fuses the first style type and the second style type based on the target style generation model.
[0015] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0016] one or more processors;
[0017] a storage device for storing one or more programs,
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating a stylized image as described in any one of the embodiments of the present disclosure.
[0019] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the method for generating a stylized image as described in any one of the embodiments of the present disclosure.
[0020] The technical solution of the embodiment of the present disclosure first obtains the model parameters to be transferred of the facial image generation model, and then constructs a first sample generation model to be trained and a second sample generation model to be trained based on these parameters. Furthermore, the corresponding sample generation models to be trained are trained based on the training samples of two style types. After the model training is completed, the model parameters to be fitted of the two target sample generation models are obtained, and then the target style data generation model is determined based on the model parameters to be fitted, so as to generate a stylized image that fuses the two style types based on the target style data generation model. Without the need to fuse a large number of training samples with two style types, the target style data generation model can be efficiently constructed, which not only allows users to use the model to generate images of the target style type, but also reduces the cost consumed in the model construction process. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0022] Figure 1 A schematic flow chart of a method for generating a stylized image provided in the first embodiment of the present disclosure;
[0023] Figure 2 A schematic diagram of constructing a first to-be-trained sample generation model and a second to-be-trained sample generation model based on a facial image generation model provided in the first embodiment of the present disclosure;
[0024] Figure 3 A schematic diagram of constructing a target style data generation model based on the first target sample generation model and the second target sample generation model provided in the first embodiment of the present disclosure;
[0025] Figure 4 A schematic flow chart of a method for generating a stylized image provided in the second embodiment of the present disclosure;
[0026] Figure 5 This is a structural block diagram of a device for generating a stylized image provided in the third embodiment of the present disclosure;
[0027] Figure 6 This is a structural diagram of an electronic device provided in Example 4 of the present disclosure. DETAILED DESCRIPTION
[0028] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0029] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0030] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0031] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0032] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0033] Example 1
[0034] Figure 1 This is a flowchart of a method for generating a stylized image provided in the first embodiment of the present disclosure. This embodiment is applicable to scenarios where a specific style data generation model is constructed. The constructed model is used to generate a stylized image that combines two style types. The method can be executed by a device for generating stylized images. The device can be implemented in the form of software and / or hardware. The hardware can be an electronic device such as a mobile terminal, PC, or server. Any image display scenario is usually implemented by the cooperation of a client and a server. The method provided in this embodiment can be executed by the server, the client, or the cooperation of the client and the server.
[0035] like Figure 1 , the method of this embodiment includes:
[0036] S101. Obtain model parameters to be transferred of a facial image generation model, and construct a first sample generation model to be trained and a second sample generation model to be trained based on the model parameters to be transferred.
[0037] In this embodiment, the facial image generation model can be a neural network model for generating a user's facial image. It can be understood that after the user's facial features are input into the facial image generation model, a facial image consistent with the user's facial features can be obtained after model processing.
[0038] In actual application, the facial image generation model can be a stylegan model based on a generative adversarial network (GAN). The generative adversarial network consists of a generator network and a discriminator network. The generator network randomly samples from the latent space as input, and its output needs to imitate the real samples in the training set as much as possible. The input of the discriminator network is the real sample and the output of the generator network. Based on this, it can be understood that the stylegan model in this embodiment also includes a generator and a discriminator. Specifically, the generator can be used to process the Gaussian noise corresponding to the user's facial image to regenerate a user's facial image; the discriminator can be used to adjust the relevant parameters in the generator. The advantage of using a discriminator containing a discriminator network is that the user's facial image regenerated by the stylegan model after parameter correction can be almost completely consistent with the user's facial image corresponding to the Gaussian noise as input. It should be noted that in the field of high-definition image generation, the stylegan model has very excellent expression capabilities and can generate at least high-definition images with a resolution of up to 1024*1024.
[0039] by Figure 2 Taking the schematic diagram of constructing the first to-be-trained sample generation model and the second to-be-trained sample generation model based on the facial image generation model as an example, G1 is the facial image generation model, and a clear facial schematic diagram can be obtained after inputting Gaussian noise. In this embodiment, in order to make the output of the facial image generation model almost completely consistent with the facial image corresponding to the input Gaussian vector, it is also necessary to train the facial image generation model. Optionally, multiple basic training samples are obtained; Gaussian noise is processed based on the to-be-trained image generator to generate an image to be discriminated; the discriminator discriminates between the image to be discriminated and the collected real facial image to determine a baseline loss value; the model parameters in the to-be-trained image generation model are corrected based on the baseline loss value; the convergence of the loss function in the to-be-trained image generator is used as the training goal to obtain the facial image generation model.
[0040] The basic training samples are the data used to train the facial image generation model. Each basic training sample consists of Gaussian noise corresponding to the target subject's facial information. The target subject's facial information is an image containing the user's facial information, such as a user's ID photo or daily photos. The Gaussian noise can be understood as a high-dimensional vector corresponding to the target subject's facial information. It should be noted that in actual applications, a large number of basic training samples can be obtained based on the large public dataset FFHQ (a facial feature dataset).
[0041] Furthermore, according to the above description, when the facial image generation model to be trained is a StyleGAN model, the model consists of a generator for training images and a discriminator. Therefore, after obtaining multiple basic training samples, the generator for training images can be used to process a large amount of Gaussian noise to generate a discriminant image, i.e., an image that may differ from the real facial image input by the user. Furthermore, after this determination, a baseline loss value between the image to be discriminated and the real facial image can be determined based on the discriminator. When using the baseline loss value to correct the model parameters in the training image generation model, the training error of the loss function in the training image generator, i.e., the loss parameter, can be used as a criterion for detecting whether the loss function has reached convergence, such as whether the training error is less than a preset error, whether the error trend is stable, or whether the current number of iterations is equal to a preset number. If convergence conditions are met, such as if the training error of the loss function is less than a preset error, or if the error trend is stable, the training of the training image generation model is complete, and iterative training can be terminated. If it is detected that the convergence condition has not been reached, other basic training samples can be further obtained to continue training the model until the training error of the loss function is within the preset range. It can be understood that when the training error of the loss function reaches convergence, the trained facial image generation model can be obtained. At this time, after inputting the Gaussian vector corresponding to the user's facial image into the model, an image that is almost completely consistent with the user's facial image can be obtained. Figure 2For example, after training, the image output by G1 is almost identical to the image corresponding to the input Gaussian noise. Generally speaking, training a facial image generation model using large amounts of data is difficult and requires significant computing resources. Furthermore, if one wishes to train a model for generating images of a specific style, a large number of images of that style must be obtained as training samples. However, samples of a certain style are almost non-existent or difficult to obtain. Consequently, in practical applications, it is impossible to train a model for that style, and consequently, it is impossible to convert captured images into images of that style. Therefore, in this embodiment, after the parameters of the facial image generation model are trained, transfer learning can be used to obtain a model for generating images of a specific style. In the field of artificial intelligence, transfer learning is the process of applying knowledge or patterns learned in a specific domain or task to a different but related domain or problem. This involves transferring annotated data or knowledge structures from a related domain to achieve or improve learning outcomes in the target domain or task. In this embodiment, the advantage of using transfer learning is that a model for generating a specific style can be trained using only a small number of samples.
[0042] Specifically, in order to obtain a model for generating images of a specific style type, the trained parameters in the facial image generation model can be used as the model parameters to be transferred, and the first sample generation model to be trained and the second sample generation model to be trained can be constructed based on the parameters.
[0043] It can be understood that the benefit of constructing the first to-be-trained sample generation model and the second to-be-trained sample generation model through transfer learning is that the model parameters that have been trained can be used to efficiently construct a model for generating images of a specific style type, which not only avoids the tedious process of obtaining a large number of style model images as training data, that is, eliminating the problem of difficulty in sample acquisition, but also reduces the consumption of computing resources.
[0044] Continue with Figure 2 For example, when the facial image generation model is determined to be G1, the parameters of the model to be transferred of G1 can be obtained, and the first generation model of the sample to be trained G2 and the second generation model of the sample to be trained G3 can be generated based on transfer learning. Figure 2It can be seen that after processing the Gaussian noise input G2 corresponding to the user's facial image, the image output by the model retains the user's unique facial features while presenting the style of dressing in a specific region. For example, the image under the first style type output by G2 can be an image with local regional characteristics such as clothing, hairstyle, hair accessories and makeup added on the basis of the user's original facial features; after processing the Gaussian noise input G3 corresponding to the user's facial image, the image output by the model retains the user's unique facial features while presenting the style of ancient materials. For example, the image under the second style type output by G3 can be an image with the features of characters in ancient paintings added on the basis of the user's original facial features. It can be understood that the user's realistic facial image presents the visual effect of ancient character paintings.
[0045] S102: Train a first to-be-trained sample generation model based on training samples of the first style type to obtain a first target sample generation model.
[0046] In this embodiment, after obtaining the first training sample generation model, the training sample of the first style type can be obtained to train the model. The first style type is a regional style image, for example, a facial image of a user dressed in a unique style, and this style of dress corresponds to a certain region. It can be understood that the first style type is a style type that presents the characteristics of clothing, hairstyle, hair accessories, and makeup of users in a certain region. Each training sample includes the first facial image of the first style type. The first facial image can be processed based on the trained target compilation model to generate Gaussian noise corresponding to the first facial image. Figure 2 For example, when the first to-be-trained sample generation model is a model for generating images of a specific regional style, the corresponding training samples are multiple images of the dressing style of users in the region, and these images are the first facial images.
[0047] The process of training the first to-be-trained sample generation model is as follows: obtaining multiple training samples under a first style type; inputting Gaussian noise into the first to-be-trained sample generation model to obtain a first actual output image; performing discriminant processing on the first actual output image and the corresponding first facial image based on a discriminator to determine a loss value, and correcting the model parameters in the first to-be-trained sample generation model based on the loss value; taking the convergence of the loss function in the first to-be-trained image generation model as a training target to obtain a first target sample generation model.
[0048] Specifically, after obtaining multiple training samples of the first style type, the image generator in the model can be used to process multiple Gaussian noises to generate a first actual output image to be discriminated, i.e., an image that differs from the first facial image. Furthermore, after determining the first actual output image and the corresponding first facial image, multiple corresponding loss values can be determined based on the discriminator. When using the multiple loss values to correct model parameters in the model generating the first training samples, the training error of the loss function in the model, i.e., the loss parameter, can be used as a criterion for detecting whether the loss function has reached convergence, such as whether the training error is less than a preset error, whether the error trend is stable, or whether the current number of iterations is equal to a preset number. If convergence is detected, such as if the training error of the loss function is less than a preset error, or if the error trend is stable, training of the model generating the first training samples is complete, and iterative training can be terminated. If convergence is not detected, additional training samples of the first style type can be obtained to continue training the model until the training error of the loss function is within a preset range. It can be understood that when the training error of the loss function reaches convergence, the first target sample generation model that has been trained can be obtained. At this time, after inputting the user's facial image into the model, a user facial image that retains the user's unique facial features while presenting the first style type can be obtained.
[0049] It should be noted that since the first sample generation model to be trained is constructed based on the trained facial image generation model, it is only necessary to use a small number of training samples of the first style type to train the model to obtain the first target sample generation model. In actual application, the training samples can be about 200 images of the first style type (i.e., the first facial images). At the same time, these images should have a similar structure to the facial images input by the user. For example, the images must have features such as the user's facial features and hair.
[0050] In this way, not only the convenience of model training is improved, but also the corresponding target sample generation model can be trained when there are fewer images of a specific style type, which greatly reduces the demand for training samples for the model to be trained.
[0051] S103: Train a second to-be-trained sample generation model based on the training samples of the second style type to obtain a second target sample generation model.
[0052] In this embodiment, after obtaining the second training sample generation model, the second style type training sample can be obtained to train the model. The second style type is an ancient material image, for example, an image in the style of ancient figure painting. It can be understood that the second style type is a style type that presents the characteristics of ancient fine brushwork, oil painting, etc. Each training sample includes a second facial image of the second style type. After processing the second facial image, Gaussian noise reflecting the corresponding facial features can also be obtained. Figure 2 For example, when the second to-be-trained sample generation model is a model for generating images of ancient material style, the corresponding training samples are multiple images of ancient material style, and these images are the second facial images.
[0053] The process of training the second to-be-trained sample generation model is as follows: obtaining multiple training samples under a second style type; inputting Gaussian noise into the second to-be-trained sample generation model to obtain a second actual output image; performing discriminant processing on the second actual output image and the corresponding second facial image based on a discriminator to determine a loss value, and correcting the model parameters in the second to-be-trained sample generation model based on the loss value; taking the convergence of the loss function in the second to-be-trained image generation model as a training target to obtain a second target sample generation model.
[0054] Those skilled in the art should understand that the process of training the second to-be-trained sample generation model based on multiple training samples under the second style type is similar to the process of training the first to-be-trained sample generation model based on multiple training samples under the first style type, and the embodiments of the present disclosure will not be repeated here. At the same time, in the actual application process, training the second to-be-trained sample generation model to obtain the second target sample generation model also requires only a small amount of second style type training data, for example, about 200 second style type images (i.e., second facial images). At the same time, these images also have a similar structure to the facial image input by the user, for example, the images must have features such as the user's facial features and hair. It can be understood that this model training method, which is similar to the first to-be-trained sample generation model, is also convenient and reduces the demand for second style type images. The embodiments of the present disclosure will not be repeated here.
[0055] S104 : Determine a target style data generation model based on the to-be-fitted model parameters of the first target sample generation model and the second target sample generation model, so as to generate a stylized image that fuses the first style type and the second style type based on the target style generation model.
[0056] In this embodiment, after training the first target sample generation model and the second target sample generation model, the parameters of these two models are obtained, and the target style data generation model is obtained through model fusion. Model fusion is the process of training multiple models and then integrating them according to a specific method. Furthermore, after the integrated target style generation model processes a user's input facial image, the output image not only retains the user's unique facial features, but also exhibits both the first and second style types. These images exhibiting multiple style types are known as stylized images.
[0057] Specifically, when constructing a target style data generation model, it is first necessary to obtain pre-set fitting parameters; based on the fitting parameters, the parameters of the models to be fitted in the first target sample generation model and the second target sample generation model are fitted to obtain target model parameters; based on the target model parameters, the target style data generation model is determined. Among them, the fitting parameters can be coefficients that characterize the degree of fusion of two style types. In the output stylized image, the fitting parameters are at least used to adjust the weights of different style types, which can be understood as being used to control which of the two style types the style type presented by the output stylized image is more inclined to. In actual application, developers can pre-edit or modify the fitting parameters based on corresponding controls or programs, and the embodiments of the present disclosure will not be described in detail here.
[0058] Furthermore, based on the preset fitting parameters, the model parameters of the first target sample generation model and the second target sample generation model can be linearly combined to obtain the target model parameters, that is, the parameters required for constructing the target style data generation model. Therefore, based on these parameters, the target style data generation model can be obtained.
[0059] by Figure 3 Taking the schematic diagram of constructing the target style data generation model based on the first target sample generation model and the second target sample generation model as an example, when the model parameters of G2 and G3 are linearly combined based on the preset fitting parameters, the target style data generation model G4 can be constructed. Figure 3It can be seen that since G2 can obtain an image of a specific regional style based on user input, and G3 can obtain an image of an ancient material style based on user input, after processing the user input using the constructed G4, the obtained image not only retains the user's unique facial features, but also presents a specific regional style and an ancient material style. For example, when the image of the first style type is an image with added features such as local regional clothing, hairstyle, hair accessories, and makeup, and the image of the second style type is an image with added features of characters in ancient paintings, the stylized image output by G4 that combines the first style type and the second style type can present the user's original facial features while presenting local regional clothing, hairstyle, hair accessories, and makeup, and make the image present the visual effect of ancient character paintings.
[0060] The technical solution of this embodiment first obtains the model parameters to be transferred of the facial image generation model, and then constructs a first sample generation model to be trained and a second sample generation model to be trained based on these parameters. Furthermore, the corresponding sample generation models to be trained are trained based on the training samples of the two style types. After the model training is completed, the model parameters to be fitted of the two target sample generation models are obtained, and then the target style data generation model is determined based on the model parameters to be fitted, so as to generate a stylized image that fuses the two style types based on the target style data generation model. Without the need to fuse a large number of training samples of the two style types, the target style data generation model can be efficiently constructed, which not only allows users to use the model to generate images of the target style type, but also reduces the cost consumed in the model construction process.
[0061] Based on the above solution, once the target style data generation model is obtained, the user's input facial image can be processed to produce an image with multiple styles. However, since the model is based on a weighted average of the parameters in the first target sample generation model and the parameters in the second target sample generation model, the output image may be poor. To address this issue, the target style data generation model can be further optimized using the following method.
[0062] Specifically, Gaussian noise is input into the target style data generation model to obtain a stylized image to be corrected that combines the first style type and the second style type; the target style image is determined by correcting the stylized image to be corrected, and the target style image is used as a target training sample to correct the model parameters in the target style data generation model based on the target training sample to obtain an updated target style data generation model.
[0063] by Figure 3For example, after obtaining the Gaussian noise z corresponding to the user's facial image, it can be input into the target style data generation model G4. Correspondingly, the image output by G4 is the stylized image to be corrected. It can be understood that although the stylized image to be corrected retains the user's unique facial features, it may not achieve a high degree of accuracy when embodying the first style type and the second style type, or the fusion of the two style types is relatively stiff. At this time, the stylized image to be corrected can be corrected based on relevant applications. For example, based on a pre-written script or related drawing software, the image parameters such as saturation, contrast, blur, and texture are adjusted to obtain a target style image that better meets the user's expectations. Those skilled in the art should understand that the corrected target style image can be used as training data to train the target style data generation model in the subsequent process.
[0064] In this embodiment, the model parameters may be corrected to achieve model updating by inputting Gaussian noise into the target style data generation model and outputting a stylized image to be corrected; processing the stylized image to be corrected and the target style image based on a discriminator to determine a loss value; and correcting the model parameters in the target style data generation model based on the loss value to obtain an updated target style data generation model.
[0065] In this embodiment, after obtaining Gaussian noise corresponding to the user's facial features, the target style data generation model can be used to process multiple Gaussian noises to generate a stylized image to be corrected, i.e., an image that does not fully exhibit the target style type. Furthermore, after determining the stylized image to be corrected and the target style image, the discriminator can determine multiple corresponding loss values. When using the multiple loss values to correct the model parameters in the target style data generation model, the training error of the loss function in the model, i.e., the loss parameter, can be used as a criterion to determine whether the loss function has reached convergence. For example, whether the training error is less than a preset error, whether the error trend is stable, or whether the current number of iterations is equal to a preset number. If convergence conditions are met, such as whether the training error of the loss function is less than a preset error, or whether the error trend is stable, training of the target style data generation model is complete, and iterative training can be terminated. If convergence conditions are not met, additional Gaussian noise can be processed to generate new stylized images to be corrected, thereby continuing model training until the training error of the loss function is within a preset range. It can be understood that when the training error of the loss function reaches convergence, the trained target style data generation model can be obtained. At this time, after inputting the user's facial image into the model, a user facial image can be obtained that retains the user's unique facial features while presenting the first style type and the second style type.
[0066] It should also be noted that, in this technical solution, the target stylized image corresponds to the target special effect image mentioned in this technical solution.
[0067] It should be noted that, in actual application, the constructed target style data generation model can be deployed in related application software. It can be understood that when it is detected that the user triggers the special effects control related to the target style data generation model, the special effects-related program can be run. Furthermore, if the user's facial image is received based on the user import operation (such as the user uploads a photo through the relevant button), or the user's facial image is captured by the camera device based on the mobile terminal (such as the user is performing real-time video), these images can be converted to display a stylized image that combines two style types.
[0068] Example 2
[0069] Figure 4 This is a flow chart of a method for generating stylized images provided in the second embodiment of the present disclosure. Based on the aforementioned embodiment, after obtaining the target style data generation model, the trained target compilation model can be combined with the target style data generation model to obtain a complete special effects image generation model. Furthermore, the special effects image generation model is deployed on a mobile terminal, thereby providing users with a service for generating special effects images in various styles based on input images. For specific implementation methods, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the aforementioned embodiments are not repeated here.
[0070] like Figure 4 As shown, the method specifically includes the following steps:
[0071] S201: Obtain model parameters to be transferred of a facial image generation model, and construct a first sample generation model to be trained and a second sample generation model to be trained based on the model parameters to be transferred.
[0072] S202: Train a first to-be-trained sample generation model based on training samples of the first style type to obtain a first target sample generation model.
[0073] S203: Train a second to-be-trained sample generation model based on the training samples of the second style type to obtain a second target sample generation model.
[0074] S204 : Determine a target style data generation model based on the to-be-fitted model parameters of the first target sample generation model and the second target sample generation model, so as to generate a stylized image that fuses the first style type and the second style type based on the target style generation model.
[0075] S205: Determine a special effects image generation model.
[0076] In this embodiment, after obtaining the target style data generation model, in order to provide corresponding services to users, that is, to enable users to use the model to make the input facial image present corresponding special effects, it is also necessary to construct a corresponding special effects image generation model based on the target style data generation model.
[0077] Typically, after obtaining a special effects image generation model, it needs to be deployed on a terminal device. Since terminal devices generally have the ability to capture user facial images, the trained target style data generation model can only process Gaussian noise corresponding to the user's facial image. Therefore, in order for the special effects image generation model to run effectively on the terminal device, a model that can generate the corresponding Gaussian noise based on the user's facial image needs to be determined, namely the target compilation model.
[0078] Specifically, based on the facial image generation model and each facial image, the training compilation model is trained to obtain a target compilation model; based on the target compilation model and the target style data generation model, the special effects image generation model is determined, and the acquired facial image to be processed is stylized based on the special effects image generation model to obtain a target special effects image that integrates the first style type and the second style type.
[0079] The facial image is an image containing facial features input by the user, such as a user's ID photo or daily photo, and the compiled model to be trained can be an encoder coding model. Those skilled in the art should understand that the encoder-decoder framework is a deep learning model framework, and the embodiments of the present disclosure will not be described in detail here. After inputting multiple facial images into the encoder coding model and processing the Gaussian noise output by the encoder coding model based on the facial image generation model, corresponding images can be obtained that can be used as training data for the compiled model to be trained.
[0080] The specific training process of the compilation model to be trained is to obtain multiple first training images; for each first training image to be trained, input the current first training image into the compilation model to be trained to obtain the Gaussian noise to be used corresponding to the current first training image; input the Gaussian noise to be used into the facial image generation model to obtain a third actual output image; based on the third actual output image and the current first training image, determine the image loss value; based on the image loss value, correct the model parameters in the compilation model to be trained, and take the convergence of the loss function in the compilation model to be trained as the training target to obtain the target compilation model, so as to determine the special effect image generation model based on the target compilation model and the target style data generation model.
[0081] In this embodiment, after obtaining a first training image containing a user's facial features, the trained compiled model can be used to process multiple of these images to generate corresponding Gaussian noise to be used. This Gaussian noise is actually a high-dimensional vector that does not accurately and completely reflect the user's facial features. Furthermore, the facial image generation model is used to process these Gaussian noises to generate a third actual output image that is not completely consistent with the first training image. After determining the third actual output image and the current first training image, a discriminator can be used to determine multiple corresponding loss values. When using these multiple loss values to correct model parameters in the trained compiled model, the training error of the model's loss function, i.e., the loss parameter, can be used as a criterion for detecting whether the loss function has reached convergence. For example, whether the training error is less than a preset error, whether the error trend is stable, or whether the current number of iterations is equal to a preset number. If convergence conditions are met, such as if the training error of the loss function is less than a preset error, or if the error trend is stable, the training of the trained compiled model is complete, and iterative training can be terminated. If it is detected that the convergence condition has not been met, the other first training images can be further processed, and a third actual output image corresponding to the obtained Gaussian vector can be generated based on the facial image generation model. The model training continues until the training error of the loss function is within a preset range. When the training error of the loss function reaches convergence, the trained target compilation model can be obtained. It can be understood that the target compilation model is used to process the input facial image into corresponding Gaussian noise. After the user's facial image is input into the target compilation model, the facial image generation model can output an image that is almost completely consistent with the user's facial image based on the Gaussian noise output by the target compilation model.
[0082] In this embodiment, after obtaining the target compilation model, the target compilation model and the target style data generation model are combined to obtain the special effect image generation model. Figure 3 For example, after obtaining the target compilation model (i.e. Figure 3 After obtaining the model corresponding to the identifier E shown in the figure, the model can be combined with G4 to obtain a special effects image generation model. After the user inputs the facial image into the special effects image generation model, the target compilation model in the model can process the image and input the processed Gaussian noise z into G4. After processing by G4, the user's unique facial features can be retained while presenting a specific regional style and ancient material style.
[0083] S206: deploying a special effect image generation model in the mobile terminal, so that when a special effect display control is detected, the collected image to be processed is processed into a target special effect image that fuses the first style type and the second style type.
[0084] In this embodiment, after obtaining the special effects image generation model, in order to use the model to provide corresponding services to users, the model can be deployed in a mobile terminal. For example, based on a specific program algorithm, the special effects image generation model can be integrated into an application (Application, APP) developed for a mobile platform.
[0085] Specifically, a corresponding control can be developed in the app for the special effect image. For example, a button named "Multi-style Special Effects" can be developed in the app interface. At the same time, the button is associated with a function that generates images with multiple styles based on the special effect image generation model. Based on this, when it is detected that the user has triggered the button, the image input by the user in real time based on the mobile terminal can be called, or the image pre-stored in the mobile terminal can be called. It can be understood that the called image must at least contain the user's facial information. These images are the images to be processed.
[0086] Furthermore, the image to be processed can be processed based on the program code corresponding to the special effect image generation model, so as to obtain a target special effect image that not only retains the user's unique facial features but also integrates the first style type and the second style type, that is, Figure 3 Special effects image output by G4.
[0087] The technical solution of this embodiment can also combine the trained target compilation model with the target style data generation model after obtaining the target style data generation model to obtain a complete special effects image generation model; further, the special effects image generation model is deployed on a mobile terminal to provide users with a service for generating special effects images of various styles based on the input image.
[0088] Example 3
[0089] Figure 5 This is a structural block diagram of a device for generating a stylized image provided by the third embodiment of the present disclosure, which can execute the method for generating a stylized image provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method. Figure 5 As shown, the device specifically includes: a module for obtaining model parameters to be transferred 301, a module for training a first model to be trained sample generation 302, a module for training a second model to be trained sample generation 303 and a module for determining a target style data generation model 304.
[0090] The module 301 for obtaining model parameters to be transferred is used to obtain model parameters to be transferred of the facial image generation model, so as to construct a first generation model of samples to be trained and a second generation model of samples to be trained based on the model parameters to be transferred.
[0091] The first to-be-trained sample generation model training module 302 is configured to train the first to-be-trained sample generation model based on training samples of a first style type to obtain a first target sample generation model.
[0092] The second to-be-trained sample generation model training module 303 is configured to train the second to-be-trained sample generation model based on training samples of the second style type to obtain a second target sample generation model.
[0093] The target style data generation model determination module 304 is configured to determine a target style data generation model based on the model parameters to be fitted of the first target sample generation model and the second target sample generation model, so as to generate a stylized image that fuses the first style type and the second style type based on the target style generation model.
[0094] Based on the above technical solutions, the device for generating a stylized image further includes a facial image generation model determination module.
[0095] A facial image generation model determination module is used to obtain multiple basic training samples; each basic training sample is Gaussian noise corresponding to the target subject's facial information; the Gaussian noise is processed based on the image generator to be trained to generate an image to be distinguished; the image to be distinguished and the collected real facial image are distinguished and processed by the discriminator to determine a baseline loss value; the model parameters in the image generation model to be trained are corrected based on the baseline loss value; the convergence of the loss function in the image generator to be trained is used as the training goal to obtain the facial image generation model.
[0096] Based on the above technical solutions, the first to-be-trained sample generation model training module 302 includes a first style type training sample acquisition unit, a first actual output image determination unit, a first correction unit, and a first target sample generation model determination unit.
[0097] The first style type training sample acquisition unit is configured to acquire a plurality of training samples of the first style type; wherein each training sample includes a first facial image of the first style type.
[0098] The first actual output image determining unit is configured to input Gaussian noise corresponding to the first facial image into the first to-be-trained sample generation model to obtain a first actual output image.
[0099] The first correction unit is used to perform discriminant processing on the first actual output image and the corresponding first facial image based on the discriminator, determine a loss value, and correct the model parameters in the first to-be-trained sample generation model based on the loss value.
[0100] The first target sample generation model determination unit is configured to take the convergence of the loss function in the first image generation model to be trained as a training target to obtain the first target sample generation model.
[0101] Based on the above technical solutions, the second to-be-trained sample generation model training module 303 includes a second style type training sample acquisition unit, a second actual output image determination unit, a second correction unit and a second target sample generation model determination unit.
[0102] The second style type training sample acquisition unit is used to acquire multiple training samples of the second style type; wherein each training sample includes a second facial image of the second style type.
[0103] The second actual output image determining unit is configured to input Gaussian noise corresponding to the second facial image into the second to-be-trained sample generation model to obtain a second actual output image.
[0104] The second correction unit is used to perform discriminant processing on the second actual output image and the corresponding second facial image based on the discriminator, determine a loss value, and correct the model parameters in the second to-be-trained sample generation model based on the loss value.
[0105] The second target sample generation model determination unit is used to take the convergence of the loss function in the second image generation model to be trained as a training target to obtain the second target sample generation model.
[0106] On the basis of the above technical solutions, the target style data generation model determination module 304 includes a fitting parameter acquisition unit, a target model parameter determination unit and a target style data generation model determination unit.
[0107] The fitting parameter acquisition unit is used to obtain preset fitting parameters.
[0108] The target model parameter determination unit is used to perform fitting processing on the model parameters to be fitted in the first target sample generation model and the second target sample generation model based on the fitting parameters to obtain target model parameters.
[0109] The target style data generation model determination unit is configured to determine a target style data generation model based on the target model parameters.
[0110] Based on the above technical solutions, the device for generating a stylized image further includes a target style data generation model updating module.
[0111] A target style data generation model update module is configured to input Gaussian noise into the target style data generation model to obtain a stylized image to be corrected that fuses the first style type and the second style type; determine a target style image by correcting the stylized image to be corrected, and use the target style image as a target training sample to correct model parameters in the target style data generation model based on the target training sample to obtain an updated target style data generation model.
[0112] Based on the above technical solutions, the device for generating a stylized image further includes a model parameter correction module.
[0113] The model parameter correction module is configured to input Gaussian noise into the target style data generation model and output a stylized image to be corrected; process the stylized image to be corrected and the target style image based on a discriminator to determine a loss value; and correct the model parameters in the target style data generation model based on the loss value to obtain an updated target style data generation model.
[0114] Based on the above technical solutions, the device for generating a stylized image further includes a stylization processing module.
[0115] A stylization processing module is configured to train a compilation model to be trained based on the facial image generation model and each facial image to obtain a target compilation model; wherein the target compilation model is configured to process the input facial image into corresponding Gaussian noise; and determine a special effects image generation model based on the target compilation model and the target style data generation model, so as to stylize the acquired facial image to be processed based on the special effects image generation model to obtain a target special effects image that fuses the first style type and the second style type.
[0116] Based on the above technical solutions, the device for generating a stylized image further includes a target compilation model determination module.
[0117] A target compilation model determination module is used to obtain multiple first training images; for each first image to be trained, input the current first training image into the compilation model to be trained to obtain the Gaussian noise to be used corresponding to the current first training image; input the Gaussian noise to be used into the facial image generation model to obtain a third actual output image; determine an image loss value based on the third actual output image and the current first training image; correct the model parameters in the compilation model to be trained based on the image loss value, and use the convergence of the loss function in the compilation model to be trained as a training target to obtain a target compilation model, so as to determine a special effect image generation model based on the target compilation model and the target style data generation model.
[0118] Based on the above technical solutions, the device for generating stylized images also includes a model deployment module.
[0119] The model deployment module is used to deploy the special effect image generation model in the mobile terminal so that when a special effect display control is detected, the collected image to be processed is processed into a target special effect image that integrates the first style type and the second style type.
[0120] On the basis of the above technical solutions, the first style type is a regional style image, and the second style type is an ancient style material image.
[0121] The technical solution provided in this embodiment first obtains the model parameters to be transferred of the facial image generation model, and then constructs a first sample generation model to be trained and a second sample generation model to be trained based on these parameters. Furthermore, the corresponding sample generation models to be trained are trained based on the training samples of the two style types. After the model training is completed, the model parameters to be fitted of the two target sample generation models are obtained, and then the target style data generation model is determined based on the model parameters to be fitted, so as to generate a stylized image that fuses the two style types based on the target style data generation model. Without the need to fuse a large number of training samples of the two style types, the target style data generation model can be efficiently constructed, which not only allows users to use the model to generate images of the target style type, but also reduces the cost consumed in the model construction process.
[0122] The apparatus for generating a stylized image provided by the embodiments of the present disclosure can execute the method for generating a stylized image provided by any embodiment of the present disclosure, and has the functional modules and beneficial effects corresponding to the execution method.
[0123] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
[0124] Example 4
[0125] Figure 6 This is a structural diagram of an electronic device provided by the fourth embodiment of the present disclosure. Figure 6 , which shows an electronic device (eg Figure 6The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0126] like Figure 6 As shown, the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 406 into a random access memory (RAM) 403. Various programs and data required for the operation of the electronic device 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An edit / output (I / O) interface 405 is also connected to the bus 404.
[0127] Typically, the following devices may be connected to the I / O interface 405: an editing device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 400 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0128] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 406, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0129] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0130] The electronic device provided in the embodiment of the present disclosure and the method for generating stylized images provided in the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0131] Example 5
[0132] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the method for generating a stylized image provided in the above embodiment is implemented.
[0133] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0134] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0135] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0136] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0137] In response to a special effect triggering operation, determining at least one special effect image to be displayed, and displaying the at least one special effect image to be displayed according to a preset image display mode;
[0138] If it is detected during the display process that the page turning condition is met, the target page turning special effect is performed on the currently displayed special effect image to be displayed, and the remaining special effect images to be displayed are displayed according to the image display method until a stop special effect display operation is received.
[0139] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0141] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0142] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0143] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0144] According to one or more embodiments of the present disclosure, [Example 1] provides a method for generating a stylized image, the method comprising:
[0145] Obtaining model parameters to be transferred of the facial image generation model, so as to construct a first sample generation model to be trained and a second sample generation model to be trained based on the model parameters to be transferred;
[0146] Training the first to-be-trained sample generation model based on training samples of the first style type to obtain a first target sample generation model;
[0147] Training the second to-be-trained sample generation model based on the training samples of the second style type to obtain a second target sample generation model;
[0148] Based on the to-be-fitted model parameters of the first target sample generation model and the second target sample generation model, a target style data generation model is determined to generate a stylized image that fuses the first style type and the second style type based on the target style generation model.
[0149] According to one or more embodiments of the present disclosure, [Example 2] provides a method for generating a stylized image, further comprising:
[0150] Optionally, a plurality of basic training samples are obtained; wherein each basic training sample is Gaussian noise corresponding to facial information of the target subject;
[0151] Processing the Gaussian noise based on the image generator to be trained to generate an image to be distinguished;
[0152] Determine a benchmark loss value based on a discriminator performing discriminative processing on the image to be discriminated and the collected real facial image;
[0153] Modifying model parameters in the image generation model to be trained based on the benchmark loss value;
[0154] The convergence of the loss function in the image generator to be trained is used as a training goal to obtain the facial image generation model.
[0155] According to one or more embodiments of the present disclosure, [Example 3] provides a method for generating a stylized image, further comprising:
[0156] Optionally, a plurality of training samples of a first style type are obtained; wherein each training sample includes a first facial image of the first style type;
[0157] Inputting Gaussian noise corresponding to the first facial image into the first to-be-trained sample generation model to obtain a first actual output image;
[0158] performing discriminative processing on the first actual output image and the corresponding first facial image based on a discriminator to determine a loss value, and modifying model parameters in the first to-be-trained sample generation model based on the loss value;
[0159] The convergence of the loss function in the first to-be-trained image generation model is used as a training goal to obtain the first target sample generation model.
[0160] According to one or more embodiments of the present disclosure, [Example 4] provides a method for generating a stylized image, further comprising:
[0161] Optionally, a plurality of training samples of a second style type are obtained; wherein each training sample includes a second facial image of the second style type;
[0162] Inputting Gaussian noise corresponding to the second facial image into the second to-be-trained sample generation model to obtain a second actual output image;
[0163] performing discriminative processing on the second actual output image and the corresponding second facial image based on the discriminator to determine a loss value, and modifying model parameters in the second to-be-trained sample generation model based on the loss value;
[0164] The convergence of the loss function in the second to-be-trained image generation model is used as a training goal to obtain the second target sample generation model.
[0165] According to one or more embodiments of the present disclosure, [Example 5] provides a method for generating a stylized image, further comprising:
[0166] Optionally, obtain pre-set fitting parameters;
[0167] Performing fitting processing on the model parameters to be fitted in the first target sample generation model and the second target sample generation model based on the fitting parameters to obtain target model parameters;
[0168] Based on the target model parameters, a target style data generation model is determined.
[0169] According to one or more embodiments of the present disclosure, [Example 6] provides a method for generating a stylized image, further comprising:
[0170] Optionally, Gaussian noise is input into the target style data generation model to obtain a stylized image to be corrected that combines the first style type and the second style type;
[0171] By correcting the stylized image to be corrected, a target style image is determined, and the target style image is used as a target training sample, so as to correct the model parameters in the target style data generation model based on the target training sample to obtain the updated target style data generation model.
[0172] According to one or more embodiments of the present disclosure, [Example 7] provides a method for generating a stylized image, further comprising:
[0173] Optionally, Gaussian noise is input into the target style data generation model, and a stylized image to be corrected is output;
[0174] Processing the stylized image to be corrected and the target style image based on the discriminator to determine a loss value;
[0175] Model parameters in the target style data generation model are modified based on the loss value to obtain an updated target style data generation model.
[0176] According to one or more embodiments of the present disclosure, [Example 8] provides a method for generating a stylized image, further comprising:
[0177] Optionally, based on the facial image generation model and each facial image, a compilation model to be trained is trained to obtain a target compilation model; wherein the target compilation model is used to process the input facial image into corresponding Gaussian noise;
[0178] Based on the target compilation model and the target style data generation model, a special effects image generation model is determined, so as to perform stylized processing on the acquired facial image to be processed based on the special effects image generation model to obtain a target special effects image that integrates the first style type and the second style type.
[0179] According to one or more embodiments of the present disclosure, [Example 9] provides a method for generating a stylized image, further comprising:
[0180] Optionally, obtaining a plurality of first training images;
[0181] For each first image to be trained, input the current first training image into the compiled model to be trained to obtain the Gaussian noise to be used corresponding to the current first training image;
[0182] Inputting the Gaussian noise to be used into the facial image generation model to obtain a third actual output image;
[0183] Determining an image loss value based on the third actual output image and the current first training image;
[0184] Based on the image loss value, the model parameters in the to-be-trained compilation model are corrected, and the convergence of the loss function in the to-be-trained compilation model is used as a training target to obtain a target compilation model, so as to determine a special effects image generation model based on the target compilation model and the target style data generation model.
[0185] According to one or more embodiments of the present disclosure, [Example 10] provides a method for generating a stylized image, further comprising:
[0186] Optionally, the special effect image generation model is deployed in a mobile terminal so that when a special effect display control is detected, the collected image to be processed is processed into a target special effect image that fuses the first style type and the second style type.
[0187] According to one or more embodiments of the present disclosure, [Example 11] provides a method for generating a stylized image, further comprising:
[0188] Optionally, the first style type is a regional style image, and the second style type is an ancient style material image.
[0189] According to one or more embodiments of the present disclosure, [Example 12] provides a device for generating a stylized image, including:
[0190] A module for obtaining model parameters to be transferred, used to obtain model parameters to be transferred of the facial image generation model, so as to construct a first sample generation model to be trained and a second sample generation model to be trained based on the model parameters to be transferred;
[0191] A first to-be-trained sample generation model training module, configured to train the first to-be-trained sample generation model based on training samples of a first style type to obtain a first target sample generation model;
[0192] A second to-be-trained sample generation model training module, configured to train the second to-be-trained sample generation model based on training samples of a second style type to obtain a second target sample generation model;
[0193] a target style data generation model determination module, configured to determine a target style data generation model based on the to-be-fitted model parameters of the first target sample generation model and the second target sample generation model, so as to generate a stylized image that fuses the first style type and the second style type based on the target style generation model.
[0194] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0195] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0196] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for generating a stylized image, characterized in that: include: Obtaining model parameters to be transferred of a facial image generation model, and constructing a first sample generation model to be trained and a second sample generation model to be trained based on the model parameters to be transferred; wherein the facial image generation model is a neural network model for generating a user's facial image; Training the first to-be-trained sample generation model based on training samples of the first style type to obtain a first target sample generation model; Training the second to-be-trained sample generation model based on the training samples of the second style type to obtain a second target sample generation model; Determining a target style data generation model based on the to-be-fitted model parameters of the first target sample generation model and the second target sample generation model, so as to generate a stylized image that combines the first style type and the second style type based on the target style generation model; The determining of the target style data generation model based on the to-be-fitted model parameters of the first target sample generation model and the second target sample generation model includes: Obtaining a preset fitting parameter; wherein the fitting parameter is a coefficient of the degree of fusion of two style types, and is used to adjust the weight of the style types; Performing fitting processing on the model parameters to be fitted in the first target sample generation model and the second target sample generation model based on the fitting parameters to obtain target model parameters; Based on the target model parameters, a target style data generation model is determined.
2. The method according to claim 1, characterized in that Before obtaining the model parameters to be transferred of the facial image generation model, the method further includes: Acquire multiple basic training samples; wherein each basic training sample is Gaussian noise corresponding to the target subject's facial information; Processing the Gaussian noise based on the image generator to be trained to generate an image to be distinguished; Determine a benchmark loss value based on a discriminator performing discriminative processing on the image to be discriminated and the collected real facial image; Modifying model parameters in the image generation model to be trained based on the benchmark loss value; The convergence of the loss function in the image generator to be trained is used as a training goal to obtain the facial image generation model.
3. The method according to claim 1, characterized in that The step of training the first to-be-trained sample generation model based on the training sample of the first style type to obtain a first target sample generation model includes: Acquire multiple training samples of a first style type, wherein each training sample includes a first facial image of the first style type; Inputting Gaussian noise corresponding to the first facial image into the first to-be-trained sample generation model to obtain a first actual output image; performing discriminative processing on the first actual output image and the corresponding first facial image based on a discriminator to determine a loss value, and modifying model parameters in the first to-be-trained sample generation model based on the loss value; The convergence of the loss function in the first to-be-trained image generation model is used as a training goal to obtain the first target sample generation model.
4. The method according to claim 1, wherein The step of training the second to-be-trained sample generation model based on the training sample of the second style type to obtain the second target sample generation model includes: Acquire a plurality of training samples of a second style type, wherein each training sample includes a second facial image of the second style type; Inputting Gaussian noise corresponding to the second facial image into the second to-be-trained sample generation model to obtain a second actual output image; performing discriminative processing on the second actual output image and the corresponding second facial image based on the discriminator to determine a loss value, and modifying model parameters in the second to-be-trained sample generation model based on the loss value; The convergence of the loss function in the second to-be-trained image generation model is used as a training goal to obtain the second target sample generation model.
5. The method according to claim 1, wherein After obtaining the target style data generation model, the method further includes: Inputting Gaussian noise into the target style data generation model to obtain a stylized image to be corrected that combines the first style type and the second style type; By correcting the stylized image to be corrected, a target style image is determined, and the target style image is used as a target training sample, so as to correct the model parameters in the target style data generation model based on the target training sample to obtain the updated target style data generation model.
6. The method according to claim 5, characterized in that After obtaining the target training sample, the method further includes: Inputting Gaussian noise into the target style data generation model and outputting a stylized image to be corrected; Processing the stylized image to be corrected and the target style image based on the discriminator to determine a loss value; Model parameters in the target style data generation model are modified based on the loss value to obtain an updated target style data generation model.
7. The method according to claim 2, characterized in that Also includes: Based on the facial image generation model and each facial image, the to-be-trained compilation model is trained to obtain a target compilation model; wherein the target compilation model is used to process the input facial image into corresponding Gaussian noise; Based on the target compilation model and the target style data generation model, a special effects image generation model is determined, so as to perform stylized processing on the acquired facial image to be processed based on the special effects image generation model to obtain a target special effects image that integrates the first style type and the second style type.
8. The method according to claim 7, characterized in that After obtaining the facial image generation model, the method further includes: acquiring a plurality of first training images; For each first image to be trained, input the current first training image into the compiled model to be trained to obtain the Gaussian noise to be used corresponding to the current first training image; Inputting the Gaussian noise to be used into the facial image generation model to obtain a third actual output image; Determining an image loss value based on the third actual output image and the current first training image; Based on the image loss value, the model parameters in the to-be-trained compilation model are corrected, and the convergence of the loss function in the to-be-trained compilation model is used as a training target to obtain a target compilation model, so as to determine a special effects image generation model based on the target compilation model and the target style data generation model.
9. The method according to claim 7, characterized in that Also includes: The special effect image generation model is deployed in a mobile terminal so that when a special effect display control is detected, the collected image to be processed is processed into a target special effect image that fuses the first style type and the second style type.
10. The method according to any one of claims 1 to 9, characterized in that: The first style type is a regional style image, and the second style type is an ancient style material image.
11. A device for generating a stylized image, characterized in that: include: a module for acquiring model parameters to be transferred, configured to acquire model parameters to be transferred of a facial image generation model, so as to construct a first training sample generation model and a second training sample generation model based on the model parameters to be transferred; wherein the facial image generation model is a neural network model for generating a user's facial image; A first to-be-trained sample generation model training module, configured to train the first to-be-trained sample generation model based on training samples of a first style type to obtain a first target sample generation model; A second to-be-trained sample generation model training module, configured to train the second to-be-trained sample generation model based on training samples of a second style type to obtain a second target sample generation model; a target style data generation model determination module, configured to determine a target style data generation model based on the to-be-fitted model parameters of the first target sample generation model and the second target sample generation model, so as to generate a stylized image that fuses the first style type and the second style type based on the target style generation model; The target style data generation model determination module includes: A fitting parameter acquisition unit, configured to acquire a preset fitting parameter; wherein the fitting parameter is a coefficient of the degree of fusion of two style types, and is used to adjust the weight of the style types; a target model parameter determination unit, configured to perform a fitting process on the model parameters to be fitted in the first target sample generation model and the second target sample generation model based on the fitting parameters to obtain target model parameters; The target style data generation model determination unit is configured to determine a target style data generation model based on the target model parameters.
12. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating a stylized image according to any one of claims 1 to 10.
13. A storage medium comprising computer executable instructions, wherein when the computer executable instructions are executed by a computer processor, the computer executable instructions are used to perform the method for generating a stylized image according to any one of claims 1 to 10.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN110516201A
Training method and device of image style conversion model, and image style conversion method and device
CN113850712A