Training method and device of object stylization model, and style transfer method and device

CN115908107BActive Publication Date: 2026-09-08BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211241241.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-11
Publication Date
2026-09-08
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

[0004]上述技术方案中,对于每种风格,都需要基于数百张真实的样本风格对象图像来对对象风格化模型进行训练,而真实的样本风格对象图像的获取难度较高,这就导致对象风格化模型所能转换的风格有限,使得模型训练的效率不高

Benefits of technology

[0043] This disclosure provides a training method for an object stylization model. By fusing sample object encoding features and sample style encoding features, a style fusion encoding is obtained that can represent the object features of a certain type of object and effectively incorporate style information. This style fusion encoding enables the generation of a pair of high-quality object images. By using the generated pair of images as sample images for training the object stylization model, model training can be performed without any real sample images, reducing the difficulty of sample acquisition, increasing the diversity of styles that the object stylization model can convert, and improving the training efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908107B_ABST
    Figure CN115908107B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and device for training an object stylization model, and a method and device for style transfer, belonging to the technical field of computers. The method comprises: obtaining sample object encoding features and sample style encoding features; fusing the sample object encoding features and the sample style encoding features to obtain sample fusion features; generating a sample natural object image and a sample style object image based on the sample fusion features; and training a preset model based on the sample natural object image, the sample style object image, and the sample style encoding features to obtain an object stylization model. The above scheme can train the model without any real sample image, thereby reducing the difficulty of sample acquisition, improving the diversity of styles that can be converted by the object stylization model, and improving the training efficiency of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and in particular to a training method, style transfer method, apparatus and device for an object stylization model. Background Technology

[0002] Stylization techniques can transform a given natural object image into a styled object image with a target style, giving the styled object image an artistic effect similar to the target style. For example, it can transform a natural object image into a styled object image with the artistic style of anime, oil painting, or pencil drawing. Because stylization techniques are highly engaging and fun, they have promising business prospects.

[0003] In related technologies, a pre-trained object stylization model is typically used to convert natural object images into stylized object images. This object stylization model is built upon GAN (Generative Adversarial Networks) image translation technology. This model effectively maintains the feature consistency between the stylized object image and the natural object image, preventing abnormal distortions.

[0004] In the above technical solutions, for each style, it is necessary to train the object stylization model based on hundreds of real sample style object images. However, it is difficult to obtain real sample style object images, which results in a limited range of styles that the object stylization model can convert, making the model training inefficient. Summary of the Invention

[0005] This disclosure provides a training method, style transfer method, apparatus, and device for an object stylization model. It enables model training without any real sample images, reducing the difficulty of sample acquisition, increasing the diversity of styles that the object stylization model can transfer, and improving the training efficiency of the model. The technical solution of this disclosure is as follows:

[0006] According to one aspect of the embodiments of this disclosure, a method for training an object stylization model is provided, the method comprising:

[0007] Obtain sample object encoding features and sample style encoding features, wherein the sample object encoding features are used to represent sample objects and the sample style encoding features are used to represent sample styles;

[0008] The sample object encoding features and the sample style encoding features are fused to obtain sample fusion features;

[0009] Based on the sample fusion features, a sample natural object image and a sample style object image are generated. The sample natural object image is used to represent the sample object in the natural environment, and the sample style object image is used to represent the sample object in the environment belonging to the sample style. Both the sample natural object image and the sample style object image include the sample object.

[0010] Based on the sample natural object image, the sample style object image, and the sample style encoding features, a preset model is trained to obtain an object stylization model, which is used to convert natural object images into style object images.

[0011] According to another aspect of the embodiments of this disclosure, a style transfer method is provided, the method comprising:

[0012] Obtain a target natural object image and target style encoding features, wherein the target natural object image is used to represent the target object in the natural environment, and the target style encoding features are used to represent the target style;

[0013] The target natural object image and the target style encoding features are input into the object stylization model for style transfer to obtain the style object image;

[0014] The object stylization model is obtained based on the training method of the above-mentioned object stylization model.

[0015] According to another aspect of the embodiments of this disclosure, a training apparatus for an object stylization model is provided, the apparatus comprising:

[0016] The acquisition unit is configured to acquire sample object encoding features and sample style encoding features, wherein the sample object encoding features are used to represent sample objects and the sample style encoding features are used to represent sample styles.

[0017] The fusion unit is configured to perform fusion of the sample object coding features and the sample style coding features to obtain sample fusion features;

[0018] The generation unit is configured to generate a sample natural object image and a sample style object image based on the sample fusion features. The sample natural object image is used to represent the sample object in a natural environment, and the sample style object image is used to represent the sample object in an environment belonging to the sample style. Both the sample natural object image and the sample style object image include the sample object.

[0019] The training unit is configured to train a preset model based on the sample natural object image, the sample style object image, and the sample style encoding features to obtain an object stylization model, which is used to convert the natural object image into a style object image.

[0020] In some embodiments, the acquisition unit is configured to perform the following operations: randomly acquire a plurality of first random numbers based on a first preset distribution; generate the sample object encoding feature based on the plurality of first random numbers, wherein the number of elements in the sample object encoding feature is equal to the number of the plurality of first random numbers; randomly acquire a plurality of second random numbers based on a second preset distribution; and generate the sample style encoding feature based on the plurality of second random numbers, wherein the number of elements in the sample style encoding feature is equal to the number of the plurality of second random numbers.

[0021] In some embodiments, the acquisition unit is configured to perform the following: randomly acquire a plurality of third random numbers based on a third preset distribution; generate the sample object encoding features based on the plurality of third random numbers, wherein the number of elements in the sample object encoding features is equal to the number of the plurality of third random numbers; and extract features from a style reference image to obtain the sample style encoding features, wherein the style reference image is used to provide style information.

[0022] In some embodiments, the acquisition unit is configured to perform feature extraction on an object reference image to obtain the sample object encoding features, wherein the object reference image is used to provide object information; randomly acquire a plurality of fourth random numbers based on a fourth preset distribution; and generate the sample style encoding features based on the plurality of fourth random numbers, wherein the number of elements in the sample style encoding features is equal to the number of the plurality of fourth random numbers.

[0023] In some embodiments, the fusion unit is configured to perform fusion parameter determination based on fusion data and fusion weights in an image generation model, wherein the fusion data is used to indicate whether sample style coding features in each network layer of the image generation model participate in fusion, the fusion weights are associated with the sample style of the sample style coding features, and the fusion parameter is used to represent the degree of fusion between the sample object coding features and the sample style coding features; and based on the fusion parameter, the sample object coding features and the sample style coding features are fused to obtain the sample fusion feature.

[0024] In some embodiments, the training unit includes:

[0025] The transfer subunit is configured to perform style transfer on the sample natural object image based on the sample style encoding features through the generator in the preset model to obtain the target style object image;

[0026] The training subunit is configured to execute a discriminator in the preset model to train the preset model based on the difference between the sample style object image and the target style object image, thereby obtaining the object stylization model.

[0027] In some embodiments, the migration subunit includes:

[0028] The encoding subunit is configured to encode the sample natural object image through the generator in the preset model to obtain a first encoded feature;

[0029] The transfer subunit is configured to perform style transfer on the first coding feature based on the sample style coding feature to obtain a second coding feature, wherein the second coding feature includes style information in the sample style coding feature;

[0030] The decoding subunit is configured to perform decoding on the second encoded feature to obtain the target style object image.

[0031] In some embodiments, the transfer subunit is configured to perform destylation on the first coding feature based on the mean and standard deviation of multiple elements in the first coding feature to obtain intermediate coding features; and to perform style transfer on the intermediate coding features based on the mean and standard deviation of multiple elements in the sample style coding features to obtain the second coding feature.

[0032] In some embodiments, the training subunit is configured to perform the following operations: determine a first loss, which is a mean squared error loss, based on the difference between pixels in the sample style object image and pixels in the target style object image; determine a second loss, which is a perceptual loss, based on the difference between image features in the sample style object image and image features in the target style object image; discriminate between the target style object image and the sample style object image based on the discriminator to obtain a third loss, which is a discrimination loss; and train the preset model based on the first loss, the second loss, and the third loss to obtain the object stylization model.

[0033] According to another aspect of the present disclosure, a style transfer apparatus is provided, the method comprising:

[0034] The acquisition unit is configured to acquire a target natural object image and target style encoding features, wherein the target natural object image represents a target object in a natural environment and the target style encoding features represent a target style;

[0035] The style transfer unit is configured to perform style transfer on the target natural object image and the target style encoding features, inputting them into an object stylization model to obtain a style object image;

[0036] The object stylization model is obtained based on the training method of the above-mentioned object stylization model.

[0037] According to another aspect of the embodiments of this disclosure, an electronic device is provided, the electronic device comprising:

[0038] One or more processors;

[0039] Memory used to store the executable program code of the processor;

[0040] The processor is configured to execute the program code to implement the above-mentioned object stylization model training method, or to implement the above-mentioned style transfer method.

[0041] According to another aspect of the present disclosure, a computer-readable storage medium is provided, which, when the program code in the computer-readable storage medium is executed by a processor of an electronic device, enables the electronic device to perform the above-described object stylization model training method, or to implement the above-described style transfer method.

[0042] According to another aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described object stylization model training method, or implements the above-described style transfer method.

[0043] This disclosure provides a training method for an object stylization model. By fusing sample object encoding features and sample style encoding features, a style fusion encoding is obtained that can represent the object features of a certain type of object and effectively incorporate style information. This style fusion encoding enables the generation of a pair of high-quality object images. By using the generated pair of images as sample images for training the object stylization model, model training can be performed without any real sample images, reducing the difficulty of sample acquisition, increasing the diversity of styles that the object stylization model can convert, and improving the training efficiency of the model.

[0044] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0045] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0046] Figure 1 This is a schematic diagram illustrating the implementation environment of a training method for an object stylization model according to an exemplary embodiment.

[0047] Figure 2 This is a flowchart illustrating a training method for an object stylization model according to an exemplary embodiment.

[0048] Figure 3 This is a flowchart illustrating another method for training an object stylization model according to an exemplary embodiment.

[0049] Figure 4 This is a schematic diagram illustrating model training according to an exemplary embodiment.

[0050] Figure 5 This is a flowchart illustrating a style transfer method according to an exemplary embodiment.

[0051] Figure 6 This is an illustration of the effect of generating a styled object image using an object stylization model according to an exemplary embodiment.

[0052] Figure 7 This is a structural block diagram illustrating a training apparatus for an object stylization model according to an exemplary embodiment.

[0053] Figure 8 This is a structural block diagram of a training apparatus for another object stylization model, illustrated according to an exemplary embodiment.

[0054] Figure 9 This is a structural block diagram of a style transfer device according to an exemplary embodiment.

[0055] Figure 10 This is a block diagram illustrating a server according to an exemplary embodiment. Detailed Implementation

[0056] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0057] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0058] It should be noted that the information (including but not limited to the target object's device information, the target object's personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are all authorized by the target object or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the style reference images and object reference images involved in this disclosure were obtained with full authorization.

[0059] Figure 1 This is a schematic diagram illustrating the implementation environment of a training method for an object stylization model according to an exemplary embodiment. See also Figure 1 The implementation environment specifically includes: terminal 101 and server 102.

[0060] Terminal 101 is at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, MP3 (Moving Picture Experts Group Audio Layer III) player, MP4 (Moving Picture Experts Group Audio Layer IV) player, and laptop computer. An application is installed and runs on terminal 101. This application can be a multimedia application, a communication application, or a conferencing application, etc., and this embodiment of the disclosure is not limited to this. The target object can log in to the application through terminal 101 to obtain the services provided by the application. Taking a multimedia application as an example, the target object can use the multimedia application to convert a natural object image containing the target object into an artistic object image in the style of animation, oil painting, pencil drawing, etc. Terminal 101 can connect to server 102 via a wireless network or wired network. Optionally, terminal 101 can receive an object stylization model trained by server 102, and use this object stylization model to convert a natural object image containing the target object into a stylized object image.

[0061] Terminal 101 generally refers to one of a plurality of terminals; this embodiment uses terminal 101 as an example. Those skilled in the art will understand that the number of terminals can be more or less. For example, there may be several terminals, or dozens or hundreds of terminals, or even more. This disclosure does not limit the number of terminals or the type of device.

[0062] Server 102 can be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 102 can connect to terminal 101 and other terminals via a wireless or wired network. Server 102 can receive natural object images containing target objects sent by terminal 101, and convert these natural object images into stylized object images using a trained object stylization model. The stylized object images are then returned to terminal 101 for display. Alternatively, server 101 can also send a trained object stylization model to terminal 101, allowing terminal 101 to convert the natural object images into stylized object images locally. In some embodiments, the number of servers can be more or less, and this disclosure does not limit this. Of course, server 102 may also include other functional servers to provide more comprehensive and diversified services.

[0063] The implementation environment for the style transfer method is similar in principle to that for the training method of the object stylization model, and will not be elaborated further here.

[0064] Figure 2 This is a flowchart illustrating a training method for an object stylization model according to an exemplary embodiment. See also... Figure 2 The training method for this object stylization model is performed by an electronic device and includes the following steps:

[0065] In step S201, the electronic device acquires sample object encoding features and sample style encoding features. The sample object encoding features are used to represent sample objects, and the sample style encoding features are used to represent sample styles.

[0066] In this embodiment of the disclosure, the sample object encoding feature can represent the sample object to be style transferred. The sample object can be a person, animal, or building, etc., and this embodiment of the disclosure is not limited thereto. The sample style encoding feature can represent the art style to be converted. The art style can be comics, oil paintings, or pencil drawings, etc., and this embodiment of the disclosure is not limited thereto. For any iteration in the training process of the object stylization model, the electronic device can acquire a sample object encoding feature and a sample style encoding feature as a pair of input features.

[0067] In step S202, the electronic device fuses the sample object coding features and the sample style coding features to obtain sample fusion features.

[0068] In this embodiment of the disclosure, the electronic device can add the sample style represented by the sample style encoding feature to the sample object encoding feature. That is, the electronic device fuses the sample object encoding feature and the sample style encoding feature. The fused sample feature can represent both the sample object and the sample style.

[0069] In step S203, the electronic device generates a sample natural object image and a sample style object image based on sample fusion features. The sample natural object image is used to represent sample objects in a natural environment, and the sample style object image is used to represent sample objects in an environment belonging to a sample style. Both the sample natural object image and the sample style object image include sample objects.

[0070] In this embodiment of the disclosure, the electronic device can decode the sample fusion features to obtain a pair of images: a sample natural object image and a sample style object image. The sample natural object image represents how the sample object appears in its natural environment. The sample style object image represents how the sample object appears in an environment belonging to the sample style. The sample style object image is equivalent to adding a filter of the sample style to the sample natural object image.

[0071] In step S204, the electronic device trains a preset model based on sample natural object images, sample style object images, and sample style coding features to obtain an object stylization model, which is used to convert natural object images into style object images.

[0072] In this embodiment of the disclosure, the electronic device can process a sample natural object image and sample style encoding features using a preset model to obtain a target style object image. That is, the electronic device can convert a sample natural object image into a target style object image belonging to the sample style using a preset model. Then, the electronic device can train the preset model based on the differences between the sample style object image and the target style object image to obtain an object stylization model.

[0073] The solution provided in this disclosure can fuse sample object encoding features and sample style encoding features to obtain a style fusion encoding that can both represent the object features of a certain type of sample object and effectively incorporate style information of the sample style. This style fusion encoding can generate a pair of high-quality object images. By using the generated pair of images as sample images for training the object stylization model, model training can be performed without any real sample images, reducing the difficulty of sample acquisition. This not only improves the diversity of styles that the object stylization model can convert, but also improves the training efficiency of the model.

[0074] In some embodiments, obtaining sample object encoding features and sample style encoding features includes:

[0075] Based on a first preset distribution, multiple first random numbers are randomly obtained;

[0076] Based on multiple first random numbers, a sample object encoding feature is generated, wherein the number of elements in the sample object encoding feature is equal to the number of the multiple first random numbers;

[0077] Based on the second preset distribution, multiple second random numbers are randomly obtained;

[0078] Based on multiple second random numbers, a sample style coding feature is generated, in which the number of elements is equal to the number of the multiple second random numbers.

[0079] The solution provided in this disclosure, during model training, randomly acquires a sample object encoding feature in each iteration through a first preset distribution, enabling the model to obtain a large number of different sample object encoding features, thus enriching the variety of sample objects. Similarly, randomly acquires sample style encoding features in each iteration through a second preset distribution, enabling the model to obtain a large number of different sample style encoding features, thus enriching the variety of sample styles. By randomly acquiring sample object encoding features and sample style encoding features, a rich variety of sample natural object images and sample style object images can be generated. This eliminates the need to prepare real sample images during the training of the object stylization model, reducing the difficulty of sample acquisition, increasing the diversity of styles that the object stylization model can convert, and improving the model's training efficiency.

[0080] In some embodiments, obtaining sample object encoding features and sample style encoding features includes:

[0081] Based on a third preset distribution, multiple third random numbers are randomly obtained;

[0082] Based on multiple third random numbers, a sample object encoding feature is generated, wherein the number of elements in the sample object encoding feature is equal to the number of the multiple third random numbers;

[0083] Feature extraction is performed on the style reference image to obtain the sample style encoding features, which are used to provide style information.

[0084] The solution provided in this disclosure allows for the random acquisition of a sample object's encoding feature during model training through a third preset distribution in each iteration. This enables the model to obtain a large number of different sample object encoding features, enriching the variety of sample objects. By extracting features from style reference images, sample style encoding features are obtained, allowing for the subsequent generation of various sample natural object images and sample style object images belonging to that sample style. By randomly acquiring sample object encoding features, it is unnecessary to prepare real sample images during the training of the object stylization model. This not only reduces the difficulty of sample acquisition but also allows for training on a specific style according to specific needs, improving the model's training efficiency.

[0085] In some embodiments, obtaining sample object encoding features and sample style encoding features includes:

[0086] Feature extraction is performed on the object reference image to obtain the sample object encoding features. This object reference image is used to provide object information.

[0087] Based on the fourth preset distribution, multiple fourth random numbers are randomly obtained;

[0088] Based on multiple fourth random numbers, a sample style coding feature is generated, in which the number of elements is equal to the number of the multiple fourth random numbers.

[0089] The solution provided in this disclosure extracts features from the object reference image during model training to obtain sample object encoding features, enabling the generation of sample natural object images and sample style object images of the sample object. By randomly acquiring a sample style encoding feature in each iteration through a fourth preset distribution, a large number of different sample style encoding features can be obtained during model training, enriching the variety of sample styles. By randomly acquiring sample style encoding features, it is unnecessary to prepare real sample images during the training of the object stylization model. This not only reduces the difficulty of sample acquisition but also allows training on a specific object according to specific needs, increasing the diversity of styles that the object stylization model can convert and improving the model's training efficiency.

[0090] In some embodiments, the sample object coding features and sample style coding features are fused to obtain sample fused features, including:

[0091] Based on the fusion data and fusion weights in the image generation model, the fusion parameters are determined. The fusion data is used to indicate whether the sample style coding features in each network layer of the image generation model participate in the fusion. The fusion weights are associated with the sample style. The fusion parameters are used to represent the degree of fusion between the sample object coding features and the sample style coding features.

[0092] Based on the fusion parameters, the sample object coding features and sample style coding features are fused to obtain the sample fusion features.

[0093] The solution provided in this disclosure determines the fusion parameters between sample object encoding features and sample style encoding features by using fusion data and fusion weights in an image generation model. This allows the fusion parameters to determine whether sample style encoding features in each network layer of the image generation model participate in the fusion. Since different sample styles correspond to different fusion weights, fusion is performed based on the fusion weights between the sample object encoding features and sample style encoding features. This ensures that the fused sample features can represent both the object features of the sample object and effectively incorporate the style information of the sample style. Consequently, based on this style fusion encoding, a pair of high-quality object images can be generated, providing high-quality samples for training the object stylization model.

[0094] In some embodiments, a preset model is trained based on sample natural object images, sample style object images, and sample style encoding features to obtain an object stylization model, including:

[0095] By using the generator in the preset model, style transfer is performed on the sample natural object image based on the sample style encoding features to obtain the target style object image;

[0096] By using the discriminator in the preset model, the preset model is trained based on the differences between the sample style object image and the target style object image to obtain the object stylization model.

[0097] The solution provided in this disclosure uses generated sample natural object images and sample style object images as sample images for training an object stylization model. This enables the generator in the preset model to generate a target style object image based on the sample natural object images and sample style encoding features. The discriminator in the preset model can train the preset model based on the differences between the sample style object images and the target style object images, thereby obtaining an object stylization model. This achieves model training without any real sample images, reducing the difficulty of sample acquisition. It not only increases the diversity of styles that the object stylization model can convert but also improves the model's training efficiency.

[0098] In some embodiments, a generator in a preset model performs style transfer on a sample natural object image based on sample style encoding features to obtain a target style object image, including:

[0099] The sample natural object image is encoded by the generator in the preset model to obtain the first encoded feature;

[0100] Based on the sample style coding features, style transfer is performed on the first coding features to obtain the second coding features, which include style information of the sample style.

[0101] The second encoded feature is decoded to obtain the target style object image.

[0102] The solution provided in this embodiment uses a generator in a preset model to perform style transfer on the encoding features of a sample natural object image based on the sample style encoding features, generating a target style object image with the sample style. This achieves the effect of transferring the style information represented in the sample style encoding features to the sample natural object image, ensuring the structural consistency between the stylized object and background and the sample natural object image, and improving the fidelity of the object stylization result.

[0103] In some embodiments, style transfer is performed on the first coding features based on the sample style coding features to obtain the second coding features, including:

[0104] Based on the mean and standard deviation of multiple elements in the first coding feature, the first coding feature is destylated to obtain the intermediate coding feature;

[0105] Based on the mean and standard deviation of multiple elements in the sample style coding features, style transfer is performed on the intermediate coding features to obtain the second coding features.

[0106] The scheme provided in this disclosure first destylates the first encoded features by using the mean and standard deviation of multiple elements in the first encoded features. This means removing style information from the sample natural object image, resulting in intermediate encoded features that contain no style information. Then, style transfer is performed on the intermediate encoded features using the mean and standard deviation of multiple elements in the sample style encoded features. This transfers the style information represented by the sample style encoded features to the encoded features in the sample natural object image, thereby generating a target style object image with the sample style, providing data for model training.

[0107] In some embodiments, the preset model is trained based on the difference between the sample style object image and the target style object image using a discriminator in the preset model to obtain an object stylization model, including:

[0108] Based on the difference between the pixels of the sample style object image and the pixels of the target style object image, a first loss is determined, which is the mean squared error loss.

[0109] Based on the difference between the image features of the sample style object image and the image features of the target style object image, a second loss is determined, which is the perceptual loss.

[0110] The discriminator distinguishes between the target style object image and the sample style object image, resulting in a third loss, which is the discrimination loss.

[0111] Based on the first loss, second loss, and third loss, the preset model is trained to obtain the object stylization model.

[0112] The solution provided in this disclosure uses mean squared error loss, perceptual loss, and discriminative loss to train a preset model. This allows the preset model to be trained from three perspectives: pixels, image features, and overall differences between the sample style object image and the target style object image. This results in an object stylization model, which improves the structural consistency between the stylized object and background and the sample natural object image, thereby enhancing the fidelity of the object stylization result.

[0113] The above Figure 2 The diagram illustrates the basic process of this disclosure. The following section will further elaborate on the solution provided in this disclosure based on one implementation method. Figure 3 This is a flowchart illustrating another method for training an object stylization model according to an exemplary embodiment. Taking an electronic device provided as a server as an example, see [link to example]. Figure 3 The method includes:

[0114] In step S301, the server obtains the sample object encoding feature and the sample style encoding feature. The sample object encoding feature is used to represent the sample object, and the sample style encoding feature is used to represent the sample style.

[0115] In this embodiment of the disclosure, the sample object encoding feature is encoding information that can represent the object features of a certain object. This embodiment of the disclosure does not limit the size and number of dimensions of the sample object encoding feature. The sample style encoding feature is encoding information that can represent the style features of a certain style. This embodiment of the disclosure does not limit the size and number of dimensions of the sample style encoding feature. The server can randomly obtain sample object encoding features and sample style encoding features; it can also obtain sample style encoding features of a certain style and randomly obtain sample object encoding features; it can also obtain sample object encoding features of a certain object and randomly obtain sample style encoding features. This embodiment of the disclosure does not limit these aspects.

[0116] In some embodiments, the server randomly obtains sample object encoding features and sample style encoding features. Accordingly, the server randomly obtains multiple first random numbers based on a first preset distribution. Then, the server generates sample object encoding features based on the multiple first random numbers, where the number of elements in the sample object encoding features is equal to the number of the multiple first random numbers. The server randomly obtains multiple second random numbers based on a second preset distribution. Then, the server generates sample style encoding features based on the multiple second random numbers, where the number of elements in the sample style encoding features is equal to the number of the multiple second random numbers. The preset distribution can be a Gaussian distribution, an exponential distribution, or a Laplace distribution, etc., and this embodiment does not limit this. The first preset distribution and the second preset distribution can be the same or different, and this embodiment does not limit this. The elements in the sample object encoding features are the aforementioned multiple first random numbers, and the elements in the sample style encoding features are the aforementioned multiple second random numbers. The solution provided by this embodiment allows for the random acquisition of a sample object encoding feature in each iteration through the first preset distribution during model training, enabling the acquisition of a large number of different sample object encoding features during model training, thus enriching the types of sample objects. By randomly acquiring sample style encoding features in each iteration through the second preset distribution, the model training process can obtain a large number of different sample style encoding features, enriching the variety of sample styles. By randomly acquiring sample object encoding features and sample style encoding features, a rich variety of sample natural object images and sample style object images can be generated. This eliminates the need to prepare real sample images during the training of the object stylization model, reducing the difficulty of sample acquisition, increasing the diversity of styles that the object stylization model can convert, and improving the model's training efficiency.

[0117] In some embodiments, the server obtains sample style encoding features of a certain style and randomly obtains sample object encoding features. Accordingly, the server randomly obtains multiple third random numbers based on a third preset distribution. Then, the server generates sample object encoding features based on the multiple third random numbers, where the number of elements in the sample object encoding features is equal to the number of the multiple third random numbers. The server performs feature extraction on a style reference image to obtain sample style encoding features, which are used to provide style information. The third preset distribution may be the same as or different from the first preset distribution described above; this embodiment does not limit this. Optionally, the server can use a pre-trained style encoder to extract features from the style reference image to obtain sample style encoding features; this embodiment does not limit this. The elements in the sample object encoding features are the multiple third random numbers mentioned above. The solution provided by this embodiment allows for the random acquisition of a sample object encoding feature in each iteration of the third preset distribution during model training, enabling the model to obtain a large number of different sample object encoding features during training, thus enriching the variety of sample objects. By extracting features from style reference images, sample style encoding features are obtained, enabling the generation of various sample natural object images and sample style object images belonging to that sample style. By randomly acquiring sample object encoding features, it is unnecessary to prepare real sample images during the training of the object stylization model. This not only reduces the difficulty of sample acquisition but also allows for training on a specific style according to specific needs, improving the model's training efficiency.

[0118] In some embodiments, the server obtains sample object encoding features of a certain object and randomly obtains sample style encoding features. Correspondingly, the server performs feature extraction on the object reference image to obtain sample object encoding features, which are used to provide object information. The server randomly obtains multiple fourth random numbers based on a fourth preset distribution. Then, the server generates sample style encoding features based on the multiple fourth random numbers. The number of elements in the sample style encoding features is equal to the number of the multiple fourth random numbers. Optionally, the fourth preset distribution may be the same as or different from the aforementioned second preset distribution; this embodiment does not limit this. The server can extract features from the object reference image using a pre-trained object encoder to obtain sample object encoding features; this embodiment does not limit this. The elements in the sample style encoding features are the aforementioned multiple fourth random numbers. The scheme provided by this embodiment obtains sample object encoding features by extracting features from the object reference image during model training, enabling the subsequent generation of sample natural object images and sample style object images of the sample object. By randomly obtaining a sample style encoding feature in each iteration through the fourth preset distribution, a large number of different sample style encoding features can be obtained during model training, enriching the variety of sample styles. By randomly acquiring sample style encoding features, it is no longer necessary to prepare real sample images during the training of the object stylization model. This not only reduces the difficulty of sample acquisition, but also allows training to be performed on a specific object according to one's own needs, thereby increasing the diversity of styles that the object stylization model can convert and improving the training efficiency of the model.

[0119] It should be noted that the steps of obtaining the encoding features of the sample object and obtaining the encoding features of the sample style can be performed at the same or different times, and this embodiment does not limit this.

[0120] In step S302, the server fuses the sample object encoding features and sample style encoding features using an image generation model to obtain sample fusion features.

[0121] In this embodiment, the server can transform the sample object encoding features and sample style encoding features into the feature space of the image generation model. For example, the server can use a multilayer perceptron to transform the sample object encoding features and sample style encoding features into a specified feature space; this embodiment does not limit this method. Then, the server uses the image generation model to fuse the sample object encoding features and sample style encoding features to generate fused sample features. This image generation model is pre-trained, and the server can fuse the sample object encoding features and sample style encoding features based on the image generation model. The image generation model can be a BlendGAN (Blend Generative Adversarial Networks) model; this embodiment does not limit this model.

[0122] In some embodiments, the server fuses the two encoded features using fusion data and fusion weights in the image generation model. The fusion data indicates whether the sample style encoding features in each network layer of the image generation model participate in the fusion. The fusion weights are associated with the sample style of the sample style encoding features. Accordingly, the server determines fusion parameters based on the fusion data and fusion weights in the image generation model. Then, the server fuses the sample object encoding features and the sample style encoding features based on the fusion parameters to obtain sample fused features. The fusion parameters represent the degree of fusion between the sample object encoding features and the sample style encoding features. The scheme provided by this disclosure determines the fusion parameters between sample object encoding features and sample style encoding features using fusion data and fusion weights in the image generation model, enabling the determination of whether the sample style encoding features in each network layer of the image generation model participate in the fusion based on the fusion parameters. Since different sample styles correspond to different fusion weights, fusion is performed based on the fusion weight between the sample object encoding features and the sample style encoding features. This allows the fused sample fusion features to represent both the object features of the sample object and effectively incorporate the style information of the sample style. As a result, a pair of high-quality object images can be generated based on the style fusion encoding, providing high-quality samples for the training of the object stylization model.

[0123] In step S303, the server generates a sample natural object image and a sample style object image based on the sample fusion features using an image generation model. The sample natural object image is used to represent sample objects in a natural environment, and the sample style object image is used to represent sample objects in an environment belonging to a sample style. Both the sample natural object image and the sample style object image include sample objects.

[0124] In this embodiment, the image generation model is used to generate sample image pairs of natural object images and style object images. When the fusion data indicates that the sample style encoding features in each network layer of the image generation model do not participate in the fusion, the sample fusion feature is the sample object encoding feature. The server decodes the sample fusion feature using the image generation model to obtain the sample natural object image. When the fusion data indicates that the sample style encoding features in each network layer of the image generation model participate in the fusion, the server decodes the sample fusion feature using the image generation model to obtain the sample style object image.

[0125] In step S304, the server uses the generator in the preset model to perform style transfer on the sample natural object image based on the sample style encoding features to obtain the target style object image.

[0126] In this embodiment, the preset model includes a generator and a discriminator. The generator generates style object images. The discriminator determines whether the input style object image was generated by the generator. The server inputs a sample natural object image and sample style encoding features into the generator, and outputs a target style object image through the generator, thereby achieving style transfer on the sample natural object image. The target style object image contains the sample style represented by the sample style encoding features.

[0127] In some embodiments, the generator includes an encoder and a decoder. The encoder extracts coded features from the sample natural object image. The decoder adds style information from the sample style coded features to the coded features of the sample natural object image to generate a target style object image. Accordingly, the server encodes the sample natural object image using the generator in the preset model to obtain a first coded feature. Then, the server performs style transfer on the first coded feature based on the sample style coded feature to obtain a second coded feature. Then, the server decodes the second coded feature to obtain the target style object image. The second coded feature includes style information from the sample style. The solution provided by this disclosure, through the generator in the preset model, performs style transfer on the coded features of the sample natural object image based on the sample style coded features to generate a target style object image with the sample style. This achieves the effect of transferring the style information represented in the sample style coded features to the sample natural object image, ensuring the structural consistency between the stylized object and the background and the sample natural object image, and improving the fidelity of the object stylization result.

[0128] The process of style transfer of the first encoded feature by the server through the decoder is as follows: the server destylates the first encoded feature based on the mean and standard deviation of multiple elements in the first encoded feature to obtain intermediate encoded features. Then, the server performs style transfer on the intermediate encoded feature based on the mean and standard deviation of multiple elements in the sample style encoded feature to obtain the second encoded feature. The scheme provided in this embodiment first destylates the first encoded feature using the mean and standard deviation of multiple elements in the first encoded feature, that is, it first removes style information from the sample natural object image, resulting in intermediate encoded features that do not contain any style information. Then, style transfer is performed on the intermediate encoded feature using the mean and standard deviation of multiple elements in the sample style encoded feature, achieving the effect of transferring the style information represented in the sample style encoded feature to the encoded features in the sample natural object image, thereby generating a target style object image with the sample style, providing data for model training.

[0129] For example, the decoder includes a network layer using the AdaIN (Adaptive Instance Normalization) algorithm. The server can use this AdaIN algorithm to transfer style information represented in the style encoding features of samples to the natural object image of the samples. This disclosure does not limit the style transfer algorithm used.

[0130] In some embodiments, the server can generate a target style object image using the following formula 1.

[0131] Formula 1:

[0132]

[0133] in, G represents the target style object image; G represents the image generation function used by the generator; x represents the sample natural object image; z s Used to represent sample style coding features.

[0134] In step S305, the server trains the preset model using the discriminator in the preset model based on the difference between the sample style object image and the target style object image to obtain an object stylization model.

[0135] In this embodiment of the disclosure, the server inputs the target style object image output by the generator and the sample style object image output by the image generation model into the discriminator. The discriminator can use the sample style object image as a reference to identify the differences between the target style object image and the sample style object image, thereby training a preset model to obtain an object stylization model.

[0136] In some embodiments, the server can determine the loss of the object stylization model based on the difference between the target style object image and the sample style object image, thereby enabling training. Accordingly, the server determines a first loss based on the difference between the pixels of the sample style object image and the pixels of the target style object image. The server determines a second loss based on the difference between the image features of the sample style object image and the image features of the target style object image. The server discriminates between the target style object image and the sample style object image using a discriminator to obtain a third loss. Then, the server trains a preset model based on the first, second, and third losses to obtain the object stylization model. The first loss is the Mean-Squared Error (MSE) loss. The second loss is the Learned Perceptual Image Patch Similarity (LPIPS) loss. The third loss is the GAN loss. The solution provided in this disclosure uses mean squared error loss, perceptual loss, and discriminative loss to train a preset model. This allows the model to be trained from three perspectives: the pixels and image features of the sample style object image and the target style object image, as well as the discriminator's ability to distinguish between the two. This results in an object stylization model, which improves the structural consistency between the stylized object and background and the sample natural object image, thereby enhancing the fidelity of the object stylization result.

[0137] In some embodiments, the server can determine the first loss using the following formula 2.

[0138] Formula 2:

[0139]

[0140] in, Used to represent the first loss; y is used to represent the sample style object image; This is used to represent images of the target style object. During model training, the first loss gradually decreases as the number of iterations increases.

[0141] In some embodiments, the server can determine the second loss using the following Formula 3.

[0142] Formula 3:

[0143]

[0144] in, Used to represent the second loss; y is used to represent the sample style object image; The second loss is used to represent the target style object image; φ(·) represents a pre-trained feature extractor. During model training, the second loss gradually decreases as the number of iterations increases.

[0145] In some embodiments, the server can determine the third loss using the following formula four.

[0146] Formula 4:

[0147]

[0148] in, Used to indicate third-party loss; D(·) is used to represent the target style object image; D(·) is used to represent the discriminator. This represents the expected loss achieved during training. During model training, as the number of iterations gradually increases, the probability that the discriminator identifies the target style object image as an image input from the generator gradually approaches 0.5 and oscillates around 0.5, reaching a dynamic equilibrium. In this case, the discriminator has difficulty distinguishing whether the input image is a target style object image or a sample style object image. At this point, the third loss gradually stabilizes.

[0149] In some embodiments, the server can determine the target loss by weighting the first loss, the second loss, and the third loss, thereby training a preset model to obtain an object stylization model. Accordingly, the server can determine the target loss using the following formula (Formula 5).

[0150] Formula 5:

[0151]

[0152] in, Used to represent target loss; Used to indicate the first loss; Used to indicate the second loss; Used to represent the third loss; λ1 and λ2 are the weights of the first loss and the second loss, respectively. λ1 and λ2 may be equal or unequal, and this disclosure does not limit this.

[0153] To more clearly describe the model training process, the following section, in conjunction with accompanying diagrams, further elaborates on the training process. Figure 4 This is a schematic diagram illustrating model training according to an exemplary embodiment. See also... Figure 4 The training process mainly consists of two parts: sample generation and model training. During sample generation, the server obtains the encoded features z of the sample objects by randomly sampling from a pre-defined distribution. f and sample style coding features z sThen, the server transforms the sample object encoding features and sample style encoding features into a specified feature space to obtain the sample object encoding features w. f and sample style coding features w s Then, the server encodes the features w of the sample objects. f and sample style coding features w s The input is fed into an image generation model, which encodes the features w of the sample object. f and sample style coding features w s The process generates a sample natural object image x and a sample style object image y. Then, the server inputs the sample natural object image x into the generator within a pre-defined model. The encoder in this generator extracts features from the sample natural object image x, obtaining its encoded features. The server then combines the encoded features of the sample natural object image x with the sample style encoded features z. s The input is fed into the decoder in the generator, which generates an image of the target style object. Then, the server trains the preset model using a discriminator based on the three loss functions mentioned above, thereby obtaining the object stylization model.

[0154] This disclosure provides a training method for an object stylization model. Since the image generation model can fuse the sample object encoding features and the sample style encoding features to obtain a style fusion encoding that can both represent the object features of a certain type of sample object and effectively incorporate the style information of the sample style, a pair of high-quality object images can be generated based on the style fusion encoding. By using the pair of images generated by the image generation model as sample images for training the object stylization model, model training can be performed without any real sample images, reducing the difficulty of sample acquisition. This not only improves the styles that the object stylization model can convert, but also improves the training efficiency of the model.

[0155] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.

[0156] Figure 5 This is a flowchart illustrating a style transfer method according to an exemplary embodiment, see [link to flowchart]. Figure 5 This style transfer method is performed by an electronic device and includes the following steps:

[0157] In step S501, the electronic device acquires a target natural object image and a target style encoding feature. The target natural object image is used to represent a target object in a natural environment, and the target style encoding feature is used to represent a target style.

[0158] In this embodiment of the disclosure, the target object can be a person, animal, or building, etc., and this embodiment of the disclosure is not limited thereto. The target style encoding feature can represent the target style to be converted. The target style can be comics, oil paintings, or pencil drawings, etc., and this embodiment of the disclosure is not limited thereto. The principle of obtaining the target style encoding feature in step S501 is similar to the principle of obtaining the sample style feature encoding in step S301, and will not be repeated here.

[0159] In step S502, the electronic device inputs the target natural object image and the target style encoding features into the object style model for style transfer to obtain the style object image.

[0160] In this embodiment of the disclosure, the object stylization model is based on the above. Figure 3 The corresponding embodiment provides a training method for the object stylization model. During style transfer, the electronic device inputs the target natural object image and target style encoding features into the object stylization model. The object stylization model then performs style transfer on the target natural object image to obtain a style object image. This style object image contains the target style represented by the target style encoding features. The object stylization model is trained based on the methods described in steps S301 to S305, which will not be elaborated further here. The principle of style transfer of the target natural object image by the electronic device using the object stylization model is similar to the principle of style transfer of the sample natural object image in step S304, and will not be elaborated further here. By using this trained object stylization model to perform style transfer on the natural object image to obtain the style object image, the electronic device can achieve good stylization results.

[0161] For example, Figure 6 This is an illustration of the effect of generating a styled object image using an object stylization model according to an exemplary embodiment. See also... Figure 6 , Figure 6 The example demonstrates the effects of converting four natural object images into style object images of their respective styles. Taking the three images in the lower left corner as examples, the first image is the natural object image input to the object stylization model. The second image is a style reference image, providing style information. The third image is the style object image generated by the object stylization model. By comparing the natural face image and the style object image, it is clear that the object and background in both images are roughly the same. This means the object stylization model effectively maintains structural consistency between the style object image and the natural object image, improving the fidelity of the object stylization result.

[0162] This disclosure provides a style transfer method that uses a trained object stylization model with generated sample images to perform style transfer on a target natural object image, generating a style object image with the target style. This achieves the effect of transferring the style information represented in the target style encoding features to the target natural object image, ensuring the structural consistency between the stylized object and background and the sample natural object image, and improving the fidelity of the object stylization result.

[0163] Figure 7 This is a structural block diagram illustrating a training apparatus for an object stylization model according to an exemplary embodiment. See also... Figure 7 The device includes:

[0164] The acquisition unit 701 is configured to acquire sample object encoding features and sample style encoding features, wherein the sample object encoding features are used to represent sample objects and the sample style encoding features are used to represent sample styles.

[0165] Fusion unit 702 is configured to perform fusion of sample object coding features and sample style coding features to obtain sample fusion features;

[0166] The generation unit 703 is configured to perform sample fusion feature generation to generate a sample natural object image and a sample style object image, wherein the sample natural object image is used to represent sample objects in a natural environment and the sample style object image is used to represent sample objects in an environment belonging to a sample style, and both the sample natural object image and the sample style object image include sample objects.

[0167] Training unit 704 is configured to train a preset model based on sample natural object images, sample style object images, and sample style encoding features to obtain an object stylization model, which is used to convert natural object images into style object images.

[0168] This disclosure provides a training device for an object stylization model. Because it can fuse the sample object encoding features and the sample style encoding features, it obtains a style fusion encoding that can both represent the object features of a certain type of sample object and effectively incorporate the style information of the sample style. This style fusion encoding enables the generation of a pair of high-quality object images. By using the generated pair of images as sample images for training the object stylization model, model training can be performed without any real sample images, reducing the difficulty of sample acquisition. This not only improves the diversity of styles that the object stylization model can convert, but also improves the training efficiency of the model.

[0169] In some embodiments, Figure 8This is a structural block diagram of a training apparatus for another object stylization model, illustrated according to an exemplary embodiment. See also Figure 8 The acquisition unit 701 is configured to perform the following operations: randomly acquire multiple first random numbers based on a first preset distribution; generate a sample object encoding feature based on the multiple first random numbers, wherein the number of elements in the sample object encoding feature is equal to the number of the multiple first random numbers; randomly acquire multiple second random numbers based on a second preset distribution; and generate a sample style encoding feature based on the multiple second random numbers, wherein the number of elements in the sample style encoding feature is equal to the number of the multiple second random numbers.

[0170] In some embodiments, see continue to see Figure 8 The acquisition unit 701 is configured to perform the following operations: randomly acquire multiple third random numbers based on a third preset distribution; generate sample object encoding features based on the multiple third random numbers, wherein the number of elements in the sample object encoding features is equal to the number of the multiple third random numbers; and extract features from the style reference image to obtain sample style encoding features, wherein the style reference image is used to provide style information.

[0171] In some embodiments, see continue to see Figure 8 The acquisition unit 701 is configured to perform feature extraction on the object reference image to obtain sample object encoding features, the object reference image being used to provide object information; based on a fourth preset distribution, a plurality of fourth random numbers are randomly acquired; based on the plurality of fourth random numbers, sample style encoding features are generated, the number of elements in the sample style encoding features being equal to the number of the plurality of fourth random numbers.

[0172] In some embodiments, see continue to see Figure 8 The fusion unit 702 is configured to perform fusion based on fusion data and fusion weights in the image generation model to determine fusion parameters. The fusion data is used to indicate whether the sample style coding features in each network layer of the image generation model participate in the fusion. The fusion weights are associated with the sample style of the sample style coding features. The fusion parameters are used to represent the degree of fusion between the sample object coding features and the sample style coding features. Based on the fusion parameters, the sample object coding features and the sample style coding features are fused to obtain the sample fusion features.

[0173] In some embodiments, see continue to see Figure 8 Training unit 704 includes:

[0174] The transfer subunit 801 is configured to perform style transfer on the sample natural object image based on the sample style encoding features through the generator in the preset model to obtain the target style object image;

[0175] Training subunit 802 is configured to execute the training of the preset model by the discriminator in the preset model based on the difference between the sample style object image and the target style object image, to obtain the object stylization model.

[0176] In some embodiments, see continue to see Figure 8 Migration subunit 801 includes:

[0177] Encoding subunit 8011 is configured to encode the sample natural object image through a generator in a preset model to obtain the first encoded feature;

[0178] The transfer subunit 8012 is configured to perform style transfer on the first coding feature based on the sample style coding feature to obtain the second coding feature, the second coding feature including the style information in the sample style coding feature;

[0179] The decoding subunit 8013 is configured to perform decoding on the second encoded feature to obtain the target style object image.

[0180] In some embodiments, see continue to see Figure 8 The transfer subunit 801 is configured to perform destylation on the first coding feature based on the mean and standard deviation of multiple elements in the first coding feature to obtain intermediate coding features; and to perform style transfer on the intermediate coding features based on the mean and standard deviation of multiple elements in the sample style coding features to obtain second coding features.

[0181] In some embodiments, see continue to see Figure 8 Training subunit 802 is configured to perform the following operations: determine a first loss based on the difference between pixels in the sample style object image and pixels in the target style object image, the first loss being the mean squared error loss; determine a second loss based on the difference between image features in the sample style object image and image features in the target style object image, the second loss being the perceptual loss; discriminate between the target style object image and the sample style object image based on a discriminator to obtain a third loss, the third loss being the discriminative loss; and train a preset model based on the first loss, the second loss, and the third loss to obtain an object stylization model.

[0182] It should be noted that the object stylization model training device provided in the above embodiments uses the division of the above functional units as an example when training the object stylization model. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the object stylization model training device and the object stylization model training method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0183] Regarding the apparatus in the above embodiments, the specific manner in which each unit performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0184] Figure 9 This is a structural block diagram illustrating a style transfer device according to an exemplary embodiment. See also... Figure 9 The device includes:

[0185] The acquisition unit 901 is configured to acquire a target natural object image and a target style encoding feature, wherein the target natural object image represents a target object in a natural environment and the target style encoding feature represents a target style;

[0186] Style transfer unit 902 is configured to perform style transfer by inputting the target natural object image and the target style encoding features into the object stylization model to obtain a style object image;

[0187] The object stylization model is obtained based on the training method of the above-mentioned object stylization model.

[0188] The style transfer apparatus provided in this embodiment uses the generated sample image to train the object stylization model, performs style transfer on the target natural object image, and generates a style object image with the target style. This achieves the effect of transferring the style information represented in the target style encoding features to the target natural object image, ensuring the structural consistency between the stylized object and the background and the sample natural object image, and improving the fidelity of the object stylization result.

[0189] It should be noted that the style transfer apparatus provided in the above embodiments is illustrated using the division of the above functional units when performing style transfer on natural object images. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the style transfer apparatus and style transfer method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0190] Regarding the apparatus in the above embodiments, the specific manner in which each unit performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0191] When electronic devices are provided as servers, Figure 10 This is a block diagram illustrating a server 1000 according to an exemplary embodiment. The server 1000 can vary significantly due to differences in configuration or performance. It may include one or more Central Processing Units (CPUs) 1001 and one or more memories 1002. The memories 1002 store at least one line of program code, which is loaded and executed by the processor 1001 to implement the object stylization model training method provided in the above-described method embodiments, or to implement the style transfer method provided in the above-described method embodiments. Of course, the server 1000 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1000 may also include other components for implementing device functions, which will not be elaborated here.

[0192] In an exemplary embodiment, a computer-readable storage medium including program code is also provided, such as a memory 1002 including program code. This program code can be executed by the processor 1001 of the server 1000 to complete the above-described object stylization model training method, or to complete the style transfer method provided in the various method embodiments described above. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.

[0193] A computer program product includes a computer program that, when executed by a processor, implements the training method for the object stylization model described above, or implements the style transfer method provided in the various method embodiments described above.

[0194] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0195] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for training an object stylization model, characterized in that, The method includes: Obtain sample object encoding features and sample style encoding features, wherein the sample object encoding features are used to represent sample objects and the sample style encoding features are used to represent sample styles; By using an image generation model, the encoded features of the sample object and the encoded features of the sample style are fused to obtain the sample fusion features; Using the image generation model, and based on the sample fusion features, a natural object image and a style object image are generated. The natural object image represents the sample object in a natural environment, and the style object image represents the sample object in an environment belonging to the sample style. Both the natural object image and the style object image include the sample object. The sample natural object image, the sample style object image, and the sample style encoding features are used as training data for a preset model. These images are then input into the preset model. The generator in the preset model encodes the sample natural object image to obtain a first encoding feature. Based on the sample style encoding feature, style transfer is performed on the first encoding feature to obtain a second encoding feature, which includes style information of the sample style. The second encoding feature is then decoded to obtain the target style object image. Based on the difference between the pixels of the sample style object image and the pixels of the target style object image, a first loss is determined, wherein the first loss is the mean squared error loss; Based on the difference between the image features of the sample style object image and the image features of the target style object image, a second loss is determined, which is a perceptual loss. The discriminator distinguishes between the target style object image and the sample style object image to obtain a third loss, which is the discrimination loss. Based on the first loss, the second loss, and the third loss, the preset model is trained to obtain an object stylization model. The object stylization model is used to convert natural object images into stylized object images. The training of the preset model does not depend on real sample images.

2. The training method for the object stylization model according to claim 1, characterized in that, The acquisition of sample object encoding features and sample style encoding features includes: Based on a first preset distribution, multiple first random numbers are randomly obtained; Based on the plurality of first random numbers, the sample object encoding feature is generated, wherein the number of elements in the sample object encoding feature is equal to the number of the plurality of first random numbers; Based on the second preset distribution, multiple second random numbers are randomly obtained; Based on the plurality of second random numbers, the sample style coding features are generated, wherein the number of elements in the sample style coding features is equal to the number of the plurality of second random numbers.

3. The training method for the object stylization model according to claim 1, characterized in that, The acquisition of sample object encoding features and sample style encoding features includes: Based on a third preset distribution, multiple third random numbers are randomly obtained; Based on the plurality of third random numbers, the sample object encoding feature is generated, wherein the number of elements in the sample object encoding feature is equal to the number of the plurality of third random numbers; Feature extraction is performed on the style reference image to obtain the sample style encoding features, and the style reference image is used to provide style information.

4. The training method for the object stylization model according to claim 1, characterized in that, The acquisition of sample object encoding features and sample style encoding features includes: Feature extraction is performed on the object reference image to obtain the sample object encoding features, and the object reference image is used to provide object information; Based on the fourth preset distribution, multiple fourth random numbers are randomly obtained; Based on the plurality of fourth random numbers, the sample style coding features are generated, wherein the number of elements in the sample style coding features is equal to the number of the plurality of fourth random numbers.

5. The training method for the object stylization model according to claim 1, characterized in that, The process of fusing the sample object encoding features and the sample style encoding features using an image generation model to obtain sample fusion features includes: Based on the fusion data and fusion weights in the image generation model, fusion parameters are determined. The fusion data is used to indicate whether the sample style coding features in each network layer of the image generation model participate in the fusion. The fusion weights are associated with the sample style. The fusion parameters are used to represent the degree of fusion between the sample object coding features and the sample style coding features. Based on the fusion parameters, the sample object encoding features and the sample style encoding features are fused to obtain the sample fusion features.

6. The training method for the object stylization model according to claim 1, characterized in that, The step of performing style transfer on the first coding feature based on the sample style coding feature to obtain the second coding feature includes: Based on the mean and standard deviation of multiple elements in the first coding feature, the first coding feature is destylated to obtain intermediate coding features; Based on the mean and standard deviation of multiple elements in the sample style coding features, style transfer is performed on the intermediate coding features to obtain the second coding features.

7. A style transfer method, characterized in that, The method includes: Obtain a target natural object image and target style encoding features, wherein the target natural object image is used to represent the target object in the natural environment, and the target style encoding features are used to represent the target style; The target natural object image and the target style encoding features are input into the object stylization model for style transfer to obtain the style object image; The object stylization model is obtained based on the training method of the object stylization model according to any one of claims 1 to 6.

8. A training device for an object stylization model, characterized in that, The device includes: The acquisition unit is configured to acquire sample object encoding features and sample style encoding features, wherein the sample object encoding features are used to represent sample objects and the sample style encoding features are used to represent sample styles. The fusion unit is configured to perform a process using an image generation model to fuse the sample object encoding features and the sample style encoding features to obtain sample fusion features; The generation unit is configured to perform the image generation model to generate a sample natural object image and a sample style object image based on the sample fusion features. The sample natural object image is used to represent the sample object in a natural environment, and the sample style object image is used to represent the sample object in an environment belonging to the sample style. Both the sample natural object image and the sample style object image include the sample object. The training unit is configured to take the sample natural object image, the sample style object image, and the sample style encoding features as training data for a preset model; input the sample natural object image, the sample style object image, and the sample style encoding features into the preset model; encode the sample natural object image using a generator in the preset model to obtain a first encoding feature; perform style transfer on the first encoding feature based on the sample style encoding features to obtain a second encoding feature, the second encoding feature including style information of the sample style; decode the second encoding feature to obtain a target style object image; and perform style transfer based on the pixel sum of the sample style object image. The differences between pixels in the target style object image are used to determine a first loss, which is a mean squared error loss. A second loss, a perceptual loss, is determined based on the differences between the image features of the sample style object image and the image features of the target style object image. A third loss, a discriminant loss, is obtained by discriminating between the target style object image and the sample style object image using a discriminator. The preset model is trained based on the first loss, the second loss, and the third loss to obtain an object stylization model. This object stylization model is used to convert natural object images into style object images, and the training of the preset model does not depend on real sample images.

9. A style transfer device, characterized in that, The method includes: The acquisition unit is configured to acquire a target natural object image and target style encoding features, wherein the target natural object image represents a target object in a natural environment and the target style encoding features represent a target style; The style transfer unit is configured to perform style transfer on the target natural object image and the target style encoding features, inputting them into an object stylization model to obtain a style object image; The object stylization model is obtained based on the training method of the object stylization model according to any one of claims 1 to 6.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the training method for the object stylization model as described in any one of claims 1 to 6, or to implement the style transfer method as described in claim 7.

11. A computer-readable storage medium, characterized in that, When the program code in the computer-readable storage medium is executed by the processor of the electronic device, the electronic device is able to perform the training method of the object stylization model as described in any one of claims 1 to 6, or implement the style transfer method as described in claim 7.

Citation Information

Patent Citations

  • Image processing method and device, computer equipment and storage medium

    CN111127378A