Style image generation method and device, equipment, medium and program product
By introducing multi-image segmentation models into the diffusion model, encoding identification and style features, fusion and denoising features, and generating images that maintain consistency of identification information and stylized, the problem of low-level image generation accuracy caused by insufficient text concreteness is solved, and training costs are reduced.
Patent Information
- Application Number
- CN202510020586.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
AI Technical Summary
现有技术中,由于文本的具象性不足,导致风格图像生成的精度不高,且训练成本高。
Multi-image segmentation model (Multi-UNet) combination is used to encode identification features and style features, and insert them into the attention network of the image segmentation model (UNet) in the diffusion model. By acquiring identification images, stylistic reference images and Gaussian noise, fuse and denoising features, we generate images that maintain consistency in identification information and stylized.
It realizes that under zero-sample learning conditions, images that maintain the consistency and stylized are generated and stylized, reducing training costs and improving the accuracy and concreteness of style image generation.
Smart Images

Figure CN119941493A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence, and in particular to a style image generation method, apparatus, device, medium and program product. Background Art
[0002] A style image refers to an image with a certain style, such as anime style, fairy style, line style, and dark style. The stylization model is used to convert a real person image into a character style image. For example, a real person image is converted into a fairy style character style image.
[0003] In the related art, a diffusion model can perform image rendering based on text or image, and a stylized model is obtained by fine-tuning the diffusion model. The stylized model is used to generate a character style image with a certain style based on the text, taking the text describing a certain style and character as a condition.
[0004] However, due to the lack of concreteness of text, the accuracy of style image generation is not high. Summary of the invention
[0005] The present application provides a method, device, equipment, medium and program product for generating a style image. The technical solution is as follows:
[0006] In one aspect, a method for generating a style image is provided, the method comprising:
[0007] Acquire a logo image, at least one style reference image, and Gaussian noise, wherein the logo image is an image including designated logo information and not having a designated style, the at least one style reference image is an image having the designated style and not including the designated logo information, and the Gaussian noise is used to simulate the characteristic distribution of the logo style image after noise addition, and the logo style image is an image including the designated logo information and having the designated style;
[0008] encoding identification features of the identification image; encoding style features of the at least one style reference image; and encoding noise features of the Gaussian noise;
[0009] fusing the noise feature, the identification feature and the style feature to obtain a fused feature; and performing denoising on the fused feature to obtain a denoised feature;
[0010] The denoising features are decoded to generate the logo style image.
[0011] In some embodiments, the fusing the sample identification feature, the sample identification style feature and the sample style feature to obtain a sample fusion feature includes:
[0012] The sample identification feature and the sample style feature are transferred to the sample identification style feature to obtain a sample splicing feature;
[0013] Mapping the sample identification style feature and the sample splicing feature based on the attention mechanism to obtain a sample attention feature;
[0014] A convolution operation is performed based on the sample attention feature to obtain the sample fusion feature.
[0015] In some embodiments, the transferring of the sample identification feature and the sample style feature to the sample identification style feature to obtain the sample splicing feature includes:
[0016] The sample identification style feature, the sample identification feature and the sample features located at the same position in the sample style feature are spliced together to obtain the sample splicing feature.
[0017] In some embodiments, the sample splicing feature includes a sample splicing K feature and a sample splicing V feature, the sample identification style feature includes a sample identification style K matrix and a sample identification style V matrix, the sample identification feature includes a sample identification K matrix and a sample identification V matrix, and the sample style feature includes a sample style K matrix and a sample style V matrix;
[0018] The step of splicing the sample identification style feature, the sample identification feature and the sample features located at the same position in the sample style feature to obtain the sample splicing feature includes:
[0019] The sample identification style K matrix, the sample identification K matrix and the sample style K matrix are concatenated to obtain the sample concatenation K feature; and the sample identification style V matrix, the sample identification V matrix and the sample style V matrix are concatenated to obtain the sample concatenation V matrix.
[0020] In some embodiments, the sample identification style feature further includes a sample identification style Q matrix, and the sample splicing feature includes a sample splicing K feature and a sample splicing V feature;
[0021] The step of mapping the sample identification style feature and the sample splicing feature based on the attention mechanism to obtain the sample attention feature includes:
[0022] The sample identification style Q matrix, the sample splicing K matrix and the sample splicing V matrix are mapped based on the attention mechanism to obtain the sample attention feature.
[0023] In some embodiments, performing denoising on the sample fusion feature to obtain the sample denoising feature includes:
[0024] Perform noise prediction based on the sample fusion features to obtain predicted noise;
[0025] De-noising the sample fusion feature according to the predicted noise to obtain the sample denoised feature.
[0026] In some embodiments, encoding the sample identification feature of the sample identification image includes:
[0027] Mapping the sample identification image to obtain a sample identification mapping feature;
[0028] Performing noise addition on the sample identification mapping feature to obtain a noisy sample identification mapping feature;
[0029] The noisy sample identification mapping feature is encoded to obtain the sample identification feature.
[0030] In some embodiments, encoding the sample style feature of the sample style reference image includes:
[0031] Mapping the sample style reference image to obtain a sample style mapping feature;
[0032] Adding noise to the sample style mapping feature to obtain a noisy sample style mapping feature;
[0033] The noisy sample style mapping feature is encoded to obtain the sample style feature.
[0034] In some embodiments, encoding the sample identification style features of the sample identification style image includes:
[0035] Mapping the sample identification style image to obtain a sample identification style mapping feature;
[0036] Performing noise addition on the sample identification style mapping feature to obtain a noisy sample identification style mapping feature;
[0037] The noisy sample identification style mapping feature is encoded to obtain the sample identification style feature.
[0038] In some embodiments, the method further comprises:
[0039] Screening the sample image pairs to determine sample image pairs that meet a second screening condition;
[0040] The second screening condition includes: a third similarity between the sample identification style image and the sample identification image in the sample image pair is greater than a third threshold.
[0041] In another aspect, a style image generation device is provided, the device comprising:
[0042] an acquisition module, configured to acquire an identification image, at least one style reference image, and Gaussian noise, wherein the identification image is an image including designated identification information and not having a designated style, the at least one style reference image is an image having the designated style and not including the designated identification information, and the Gaussian noise is used to simulate a feature distribution of the identification style image after noise addition, and the identification style image is an image including the designated identification information and having the designated style;
[0043] An encoding module, used for encoding the identification feature of the identification image; encoding the style feature of the at least one style reference image; and encoding the noise feature of the Gaussian noise;
[0044] A processing module, configured to fuse the noise feature, the identification feature and the style feature to obtain a fused feature; and to perform denoising on the fused feature to obtain a denoised feature;
[0045] A decoding module is used to decode the denoising feature and generate the logo style image.
[0046] On the other hand, a computer device is provided, comprising: a processor and a memory, wherein the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the style image generation method as described above.
[0047] On the other hand, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the style image generation method as described above.
[0048] On the other hand, a computer program product is provided, the computer program product comprising computer instructions, the computer instructions being stored in a computer-readable storage medium, a processor obtaining the computer instructions from the computer-readable storage medium, so that the processor loads and executes the computer instructions to implement the style image generation method as described above.
[0049] The beneficial effects brought by the technical solution provided by the embodiment of the present application include at least:
[0050] A style image generation method is provided, which is a method capable of generating a stylized image with consistent identification information. By obtaining an identification image with specified identification information, encoding its identification features, and obtaining at least one style reference image with a specified style, encoding its style features, and fusing the identification features, style features, and noise features of Gaussian noise, a fusion feature is obtained. On the one hand, the identification features and style features can be used as reference features to guide the denoising process of the fusion features, thereby restoring an identification style image including the specified identification information and having a specified style, so that the identification style image maintains consistency with the identification information of the identification image and maintains consistency with the style of the style reference image. On the other hand, compared with the text of the related art, the identification image and the style reference image have stronger concreteness, and realize precise control of the identification information and style, thereby improving the accuracy of the identification style image. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0052] Figure 1 is a structural block diagram of a computer system provided by an exemplary embodiment of the present application;
[0053] Figure 2 is a processing schematic diagram of a potential diffusion model provided by an exemplary embodiment of the present application;
[0054] Figure 3 is a schematic diagram of a style image generation method provided by an exemplary embodiment of the present application;
[0055] Figure 4 is a schematic block diagram of a style image generation model provided by an exemplary embodiment of the present application;
[0056] Figure 5 is a schematic block diagram of a style image generation model provided by an exemplary embodiment of the present application;
[0057] Figure 6 is a flowchart of a style image generation method provided by an exemplary embodiment of the present application;
[0058] Figure 7 is a schematic diagram of feature splicing provided by an exemplary embodiment of the present application;
[0059] Figure 8 is a schematic diagram of a style image generation method provided by an exemplary embodiment of the present application;
[0060] Fig. 9 is a schematic diagram of training data provided by an exemplary embodiment of the present application;
[0061] Fig.10 is a schematic diagram of a style image generation method provided by an exemplary embodiment of the present application;
[0062] Fig.11 is a schematic diagram of a sample style reference image provided by an exemplary embodiment of the present application;
[0063] Fig.12 is a schematic diagram of screening sample style reference images provided by an exemplary embodiment of the present application;
[0064] Fig.13 is a schematic diagram of a style image generation method provided by an exemplary embodiment of the present application;
[0065] Fig.14 is a schematic diagram of the structure of a style model provided by an exemplary embodiment of the present application;
[0066] Fig.15 is a block diagram of a style image generating device provided by an exemplary embodiment of the present application;
[0067] Fig.16 It is a structural block diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0068] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0069] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0070] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in this application refers to and includes any or all possible combinations of one or more associated listed items.
[0071] It should be understood that, although the terms first, second, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first parameter may also be referred to as the second parameter, and similarly, the second parameter may also be referred to as the first parameter. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0072] It should be noted that before collecting relevant data of users and characters (for example, identification information, identification images, style reference images, sample identification information, sample identification images, sample style reference images, and sample identification style images) and during the process of collecting relevant data of users, this application can display a prompt interface, pop-up window, or output voice prompt information. The prompt interface, pop-up window, or voice prompt information is used to prompt the user that its relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining user-related data after obtaining the user's confirmation operation on the prompt interface or pop-up window. Otherwise, that is, when the user's confirmation operation on the prompt interface or pop-up window is not obtained, the relevant steps of obtaining user-related data are terminated, that is, the user's relevant data is not obtained. In other words, all user data collected by this application are collected with the user's consent and authorization, and the collection, use, and processing of relevant user data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0073] First, briefly introduce the terms involved in the embodiments of this application:
[0074] Identification image: an image that includes specified identification (ID) information. The specified identification information may be identification information of a living being, where the living being may be a person, an animal, etc. In the embodiment of the present application, a person is taken as an example, then the identification image is also referred to as a person image, and the specified identification information may be at least one of a person's name, nickname, gender, age, facial features, and film and television dramas in which the person has participated. Optionally, facial features include at least one of the distribution, size, shape, face shape, hairstyle, and muscle distribution of facial features. The identification image may be used as a reference for the identification information in the identification style image, that is, the identification style image maintains consistency in identification information with the identification image.
[0075] In the embodiment of the present application, the logo image is an image that includes the specified logo information and does not have a specified style. The specified style here is the style in the logo style image to be generated, for example, the specified style is the fairy style. Then, the logo image is an image that includes the specified logo information and has no style, or the logo image is an image that includes the specified logo information and has other styles, and the other styles are different from the specified style, for example, the specified style is the fairy style, and the other style is the line style.
[0076] Style reference image: an image that does not include specified identification information and has a specified style. The style reference image can be used as a reference for the style in the identification style image, that is, the identification style image maintains consistency in style with the style reference image.
[0077] Logo style image: an image that includes specified logo information and has a specified style. The logo style image maintains the consistency of logo information with the logo image and the consistency of style with the style reference image. It is also called a stylized image with consistent logo information.
[0078] For example, taking a person as an example, the identification image is a non-styled image of the person, and the style reference image is an image of an animal in a fairy-tale style. Then the identification style image is an image of the person in a fairy-tale style.
[0079] Style image generation model: a model for executing the style image generation method provided in the embodiment of the present application, which is used to generate a logo style image based on the logo image and the style reference image, wherein the logo style image maintains the consistency of logo information with the logo image and the consistency of style with the style reference image. The style image generation model is implemented based on the diffusion model.
[0080] Concreteness: refers to the real existence form and state of things or phenomena, as well as their specific manifestations and characteristics. Things or phenomena with strong concreteness are easier to be directly perceived and understood.
[0081] Zero-shot learning: Using a large amount of training data in the training set to train the neural network model, so that the generalization of the neural network model can be used to transfer the neural network model capabilities to the test set without further training. There is no intersection between the training set categories and the test set categories. Since there is no need to continue training, the training cost can be significantly reduced.
[0082] Figure 1 1 is a block diagram of a computer system provided by an exemplary embodiment of the present application. The computer system 100 can be implemented as a system architecture of a style image generation method. The computer system 100 includes: a terminal 120 and a server 140.
[0083] The terminal 120 may be an electronic device such as a mobile phone, a tablet computer, a vehicle terminal (vehicle computer), a wearable device, a PC (Personal Computer), an unmanned reservation terminal, a smart home appliance, an intelligent voice interaction device, an unmanned vending terminal, etc. The terminal 120 may run a client of a target application. Optionally, the target application may be an application that provides a style image generation function, for example, at least one of a game application, a chat application, a life service application, a video application, an entertainment application, a travel application, a social application, a health application, and an application download application. Among them, the game may be any one of a battle royale shooting game, a virtual reality (VR) client, an augmented reality (AR) program, a three-dimensional map program, a virtual reality game, an augmented reality game, a first-person shooting game (FPS), a third-person shooting game (TPS), a multiplayer online tactical competitive game (MOBA), a strategy game (SLG), and a party game. Alternatively, the terminal 120 stores the style image generation model, and the target application may also be an application that provides training and / or reasoning functions for the style image generation model, which is not limited in the embodiments of the present application. In addition, the embodiments of the present application do not limit the form of the target application, including but not limited to App (Application, application) installed in the terminal 120, applets, etc., and may also be in the form of a web page.
[0084] Server 140 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and cloud servers for basic cloud computing services such as big data and artificial intelligence platforms. Server 140 can be the background server of the above-mentioned target application, used to provide background services for the client of the target application. In some optional embodiments, server 140 can also be implemented as a node in a blockchain system.
[0085] The terminal 120 and the server 140 may communicate with each other via a network, such as a wired or wireless network.
[0086] The method for generating a style image provided in the embodiment of the present application can be performed by a computer device in each step. A computer device refers to an electronic device with data calculation, processing and storage capabilities. Figure 1 Taking the implementation environment of the scheme shown as an example, the style image generation method can be executed by the terminal 120, for example, the target application running in the terminal 120 executes the style image generation method, or the style image generation method can be executed by the server 140, or the terminal 120 and the server 140 interact and cooperate to execute it, which is not limited to the embodiments of the present application.
[0087] Those skilled in the art will appreciate that the number of terminals 120 may be more or less. For example, there may be only one terminal 120, or there may be dozens or hundreds of terminals 120, or more. The embodiment of the present application does not limit the number and device type of the terminals 120.
[0088] Diffusion Models (DM) are models that perform image rendering based on input text or images. The core idea of the diffusion model includes two main processes: forward diffusion process and reverse diffusion process. In the forward diffusion process, Gaussian noise is gradually added to the original image until the image becomes pure noise. The whole process is usually a series of Markov chains. In the reverse diffusion process, a series of Markov chains are used to gradually remove the prediction noise from the noise to restore the original image from the noise. Due to the ability of diffusion models to preserve the semantic structure of data, they have been proven to be immune to mode collapse.
[0089] The Latent Diffusion Models (LDM) model is a type of diffusion model. The latent diffusion model is used to diffuse in latent space rather than pixel space, which can save memory. At the same time, it combines the text semantic feedback from the Transformer to produce diverse and highly detailed images while retaining the semantic structure of the data. The training data used for fine-tuning the latent diffusion model is text and image pairs. During the training process, the latent diffusion model uses text as a condition and a noisy image as input, and performs denoising on the noisy image to achieve the purpose of restoring the image from the noise.
[0090] Figure 2It is a processing diagram of a potential diffusion model provided by an exemplary embodiment of the present application. The potential diffusion model mainly includes three parts: a text encoder (Text Encoder), an image encoder (VAE Encoder) and an image segmentation model (UNet). During the pre-training process, the image encoder adds noise to the original image, and the image segmentation model continuously removes noise under the condition of the text (Text) encoded by the text encoder to restore the original image. The pre-trained potential diffusion model can generate images with corresponding meanings based on the text.
[0091] Based on the pre-trained latent diffusion model, objects unknown to the model can be embedded into the model through a variety of fine-tuning schemes. In related technologies, fine-tuning schemes include: text embedding (Text_Inversion) fine-tuning, low-rank adaptation (Low-Rank Adaptation, LORA) fine-tuning and personalized fine-tuning (Dreambooth).
[0092] The difference between these three fine-tuning schemes lies in the different trainable parameters. Specifically, text embedding fine-tuning changes a word vector in the text encoder. Low-rank adaptive fine-tuning introduces an additional bypass matrix in the linear layer of the attention network of the image segmentation model. Personalized fine-tuning changes more parameters, involving all parameters of the text encoder and image segmentation model. Text embedding fine-tuning changes the least number of parameters, so the model capability is limited. When applied to new concept training, it is difficult to generate high-quality new concept maps. Personalized fine-tuning modifies the parameters of the entire model structure and has strong fitting ability, but due to the modification of the parameters of the original model, it leads to a large number of parameters and poor transferability. Low-rank adaptive fine-tuning only introduces an additional bypass matrix in the linear layer of the attention network of the image segmentation model, with fewer parameters and stronger fitting ability.
[0093] Taking people as an example, using the fine-tuning solution of related technologies, under the training data of real person images, style texts, and style images of the same person, a stylized person style image that retains the person identification information can be generated according to the text. However, the related technology has at least the following shortcomings:
[0094] 1. Poor scalability: Due to the lack of concreteness of text, it is difficult for text to accurately describe detailed information about the style, such as complex shapes and colors, and it is impossible to expand to styles that are difficult to describe in text, which leads to low accuracy in style image generation. 2. High training cost: When fine-tuning the pre-trained latent diffusion model, a large amount of training data (hundreds of thousands) is required to effectively achieve zero-shot learning capabilities.
[0095] Based on this, the embodiment of the present application provides a model based on an innovative multi-image segmentation model (Multi-UNet) that maintains the consistency of identification information and is stylized, called a style image generation model, which can achieve zero-sample learning capabilities and generate images that maintain the consistency of identification information and are stylized, called identification style images. Accordingly, the improvements of the embodiment of the present application include at least the following aspects:
[0096] 1. Introduce image features to solve the problem of insufficient control and concreteness of text and improve scalability. Images are more concrete and controllable than text. On the basis of the diffusion model, use images to replace text to achieve precise style control. 2. Use a combination of multi-image segmentation models (Multi-UNet) to achieve zero-shot learning capabilities and reduce training costs. Taking people as an example, take the identification information and style of the people as prior conditions, add a model with the same structure as the image segmentation model (UNet), encode identification features and style features, and insert the identification features and style features into the attention network (Attention) of the image segmentation model (UNet) in the diffusion model. The functions of each model of the multi-image segmentation model are independent and have the same semantic space. Only a small amount of training data (tens of thousands) is required, and it can converge quickly to generate stylized images that maintain consistent identification information.
[0097] Figure 3 2 is a schematic diagram of a style image generation method provided by an exemplary embodiment of the present application. The method is executed by a computer device, and the computer device stores a style image generation model 200. The computer device may be Figure 1 The terminal 120 and / or server 140 shown. The steps of the method are briefly described as follows:
[0098] Acquire an identification image 210, where the identification image 210 is an image that includes specified identification information and does not have a specified style, for example, the identification image 210 is a person image of a real person; acquire at least one style reference image 220, where at least one style reference image 220 is an image that has a specified style and does not include specified identification information, for example, the specified style is anime style, and at least one style reference image is an anime-style image of an anime character with anime style; acquire Gaussian noise, where the Gaussian noise is randomly sampled from a Gaussian distribution from 0 to 1, and the Gaussian noise is used to simulate the feature distribution of the identification style image 230 after noise addition, where the identification style image 230 is an image that includes specified identification information and has a specified style, for example, the identification style image is an anime-style image of the real person with anime style.
[0099] The logo image 210, at least one style reference image 220 and Gaussian noise are input to the style image generation model 200, and the logo style image 230 is output based on the style image generation model 200; wherein, in the style image generation model 200, the logo feature of the logo image 210 is encoded; the style feature of at least one style reference image 220 is encoded; the noise feature of the Gaussian noise is encoded; the noise feature, the logo feature and the style feature are fused to obtain a fused feature, the fused feature is denoised to obtain a denoised feature; the denoised feature is decoded to generate the logo style image 230. Please refer to the following for specific embodiments.
[0100] (I) Structure of the style image generation model
[0101] Figure 4 The figure is a schematic block diagram of a style image generation model provided by an exemplary embodiment of the present application. The style image generation model for executing the style image generation method comprises: an encoder 10, a feature network 20 and a decoder 30 cascaded in sequence; wherein:
[0102] In the inference stage, the encoder 10 and the feature network 20 are used to encode the identification features of the identification image, encode the style features of at least one style reference image, and encode the noise features of Gaussian noise; the feature network 20 is also used to fuse the noise features, the identification features and the style features to obtain the fused features, and perform denoising on the fused features to obtain the denoised features; the decoder 30 is used to decode the denoised features to generate the identification style image.
[0103] During the training phase, the encoder 10 and the feature network 20 are used to encode the sample identification features of the sample identification image, the sample identification style features of the sample identification style image, and the sample style features of the sample style reference image; the feature network 20 is also used to fuse the sample identification features, the sample identification style features, and the sample style features to obtain sample fusion features, and to perform denoising on the sample fusion features to obtain sample denoising features; the decoder 30 is used to decode the sample denoising features to obtain a predicted identification style image.
[0104] Figure 5 It is a schematic block diagram of a style image generation model provided by an exemplary embodiment of the present application. The encoder 10 includes a first encoder 11, a second encoder 12 and a third encoder 13 connected in parallel, the feature network 20 includes a first feature network 21, a second feature network 22 and a third feature network 23 connected in parallel, the first encoder 11 is also connected to the first feature network 21, the second encoder 12 is also connected to the second feature network 22, the third encoder 13 is also connected to the third feature network 23, the first feature network 21 and the third feature network 23 are also connected to the second feature network 22 respectively, and the second feature network 22 is also connected to the decoder 30; wherein:
[0105] In the inference stage, the first encoder 11 is used to map the identification image to obtain the identification mapping features, and the first feature network 21 is used to encode the noisy identification mapping features to obtain the identification features; the third encoder 13 is used to map at least one style reference image to obtain the style mapping features, and the third feature network 23 is used to encode the noisy style mapping features to obtain the style features; the second feature network 22 is used to encode the noise features of Gaussian noise; the noise features, identification features and style features are fused to obtain the fused features, and the fused features are denoised to obtain the denoised features; the decoder 30 is used to decode the denoised features to generate the identification style image.
[0106] In the training stage, the first encoder 11 is used to map the sample identification image to obtain the sample identification mapping feature, and the first feature network 21 is used to encode the noisy sample identification mapping feature to obtain the sample identification feature; the third encoder 13 is used to map the sample style reference image to obtain the sample style mapping feature, and the third feature network 23 is used to encode the noisy sample style mapping feature to obtain the sample style feature; the second encoder 12 is used to map the sample identification style image to obtain the sample identification style mapping feature, and the second feature network 22 is used to encode the noisy sample identification style mapping feature to obtain the sample identification style feature, and the sample identification feature, the sample identification style feature and the sample style feature are fused to obtain the sample fusion feature, and the sample fusion feature is denoised to obtain the sample denoised feature; the decoder 30 is used to decode the sample denoised feature to generate a predicted identification style image.
[0107] In some embodiments, the first encoder 11, the second encoder 12 and the third encoder 13 can be completely identical, and the first feature network 21, the second feature network 22 and the third feature network 23 are models of the same structure, which can be completely identical or keep the same size. The first encoder 11 and the first feature network 21, the second encoder 12 and the second feature network 22, and the third encoder 13 and the third feature network 23 respectively have the function of feature encoding, and the second feature network 22 also has the function of feature fusion and denoising. In one example, the first feature network 21, the second feature network 22 and the third feature network 23 can respectively include an attention network (Attention), the identification feature can also be called an identification attention feature (Person AttentionFeature), the noise feature can also be called a noise attention feature (UNet Attention Feature), and the style feature can also be called a style attention feature (Style Attention Feature).
[0108] In a specific implementation, the first encoder 11, the second encoder 12 and the third encoder 13 are image encoders (Encoders), for example, image encoders (VAE Encoders) in a diffusion model, also known as variational encoders. The decoder 30 is an image decoder, for example, an image encoder (VAE Decoder) in a diffusion model, also known as a variational encoder. The second feature network 22 is implemented as an image segmentation model (UNet) in a diffusion model, also known as a denoising model. The first feature network 21, the third feature network 23 and the second feature network 22 are models with the same structure.
[0109] Implementation method 1, the first feature network 21 and the third feature network 23 can be exactly the same as the second feature network 22, the first feature network 21 and the third feature network 23 are also image segmentation models (UNet), the first feature network 21 is also called the identity image segmentation model (PersonNet), and the third feature network 23 is also called the style image segmentation model (StyleNet).
[0110] For example, the image segmentation model (UNet) in the diffusion model can generate images of different styles. The input of the image segmentation model (UNet) is the noisy feature X t , whose function is to predict noise and generate denoising feature X after removing the prediction noise t-1 After t steps of iterative processing, the noisy feature X t Converted to denoised feature X0. The smaller t is, the less predicted noise is. When t=1, the image segmentation model (UNet) can obtain nearly noise-free data. After this denoising step, the denoising result is denoised feature X0. After the denoised feature X0 is decoded by the decoder, a meaningful image can be generated. Corresponding to this embodiment, a logo style image is generated.
[0111] The image segmentation model (UNet) mainly includes two structures: Residual Network (ResNet) and Transformer. ResNet is mainly composed of convolutional networks, which are used to encode image features. Transformer is mainly composed of attention networks (Attention), which are used to encode image features and introduce conditional features. The conditional features can be at least one of the following: text features, image features, audio features. The mapping formula of the attention network is:
[0112]
[0113] The attention network in Transformer mainly includes two forms: self-attention and cross-attention. When the Q matrix, K matrix and V matrix are the same, it is self-attention, and the attention network is used to calculate the relationship between image features. When the Q matrix is the image feature, and the K matrix and V matrix are the conditional features, it is cross-attention, and the attention network is used to calculate the relationship between image features and conditional features. It should be noted that although the number of parameters of the attention network accounts for a small proportion in the entire style image generation model, its role is very important.
[0114] When the first feature network 21, the second feature network 22 and the third feature network 23 are all image segmentation models (UNet), the structure of the entire feature network 20 is: a multi-image segmentation model (Multi-UNet). The overall function of the multi-image segmentation model (Multi-UNet) is to encode the identification features of the identification image and the style features of at least one style reference image, and to help the image segmentation model (UNet) restore the identification style image by passing the identification features and style features to the attention network of the image segmentation model (UNet) of the second feature network 22 in the form of vector splicing.
[0115] In implementation mode 2, the first feature network 21 and the third feature network 23 may also be convolutional neural networks (CNN), for example, residual networks (ResNet). In this case, the first feature network 21 is used to encode the identification features of the identification image, and the third feature network 23 is used to encode the style features of the style reference image. The first feature network 21 and the third feature network 23 need to maintain the same size as the second feature network 22.
[0116] In this embodiment, the structure of the multi-image segmentation model (Multi-UNet) in the style image generation model can achieve zero-sample learning capability and reduce training costs. Among them, the identification features and style features are inserted into the attention network (Attention) of the image segmentation model (UNet) in the diffusion model. The functions of each model of the multi-image segmentation model (Multi-UNet) are independent, with the same semantic space, and only a small amount of training data (tens of thousands) is required. It can converge quickly and generate stylized images that maintain consistency in identification information.
[0117] (II) Reasoning stage of style image generation model
[0118] Figure 6is a flowchart of a style image generation method provided by an exemplary embodiment of the present application. The method is executed by a computer device, and the computer device stores a style image generation model, and the style image generation model is used to perform steps related to encoding, fusion, denoising, and decoding. The computer device can be Figure 1 The terminal 120 and / or server 140 shown. The method includes at least some of steps 320, 340, 360, and 380:
[0119] Step 320, obtaining a logo image, at least one style reference image and Gaussian noise, wherein the logo image is an image that includes specified logo information and does not have a specified style, and at least one style reference image is an image that has a specified style and does not include the specified logo information. Gaussian noise is used to simulate the feature distribution of the logo style image after noise addition, and the logo style image is an image that includes specified logo information and has a specified style.
[0120] The designated identification (ID) information is predetermined identification information and is also the identification information in the identification style image to be generated. The identification image is an image that includes the designated identification information. The designated identification information may be identification information of a creature, where the creature may be a person, an animal, etc. Taking a person as an example, the identification image is called a person image, and the designated identification information may be at least one of a person's name, nickname, gender, age, facial features, and a film or TV series in which the person has participated. Optionally, the facial features include at least one of the distribution, size, shape, face shape, hairstyle, and muscle distribution of the facial features.
[0121] The logo image is an image that includes the specified logo information and does not have a specified style. The specified style here is the style of the logo style image to be generated, for example, the specified style is the fairy style. Then, the logo image is an image that includes the specified logo information and has no style, or the logo image is an image that includes the specified logo information and has other styles, and the other styles are different from the specified style, for example, the specified style is the fairy style, and the other styles are line styles. The logo image can be one of the avatars, images, and photos uploaded by the user through an application or application platform, or it can be obtained from a public dataset.
[0122] Style is the visual effect presented by an image. An image usually has a style. The style can be at least one of the following: Anime style, Xianxia style, oil painting style, watercolor style, Oneline-Drawing style, Black Magic style, Ink style, Ice Scupture style. Anime style also includes at least one of the following: 3D Real Cartoon style, 3D Cartoon style, 2D Anime style, 2D Comic style.
[0123] A style can be characterized by both structure and texture. The structure is characterized by at least one of the following parameters: two-dimensional (2D), three-dimensional (3D), structure, shape and distribution of facial features. The texture is characterized by at least one of the following parameters: color, material, light.
[0124] For example, take characters and 3D real-life cartoon style as an example, refer to Fig.11 As shown in (1), the structure is three-dimensional, the shape, structure and distribution of the facial features are close to real people, and the special materials, colors and light are used. Fig.11 As shown in (2), the structure is three-dimensional, the shape, structure and distribution of the facial features are close to real people, and the special materials, colors and light are used. Fig.11 As shown in (3), the structure is three-dimensional, the facial features are exaggerated, and the material, color and light are special. Fig.11 As shown in (4), the structure is three-dimensional, the shape, structure and distribution of the facial features are close to real people, and the special materials, colors and light are used. Fig.11 As shown in (5), the structure is two-dimensional, the facial features are exaggerated, and special materials, colors and light are used. Taking the figure and line style as an example, refer to Fig.11 As shown in (6), the structure is two-dimensional, the facial features are exaggerated, and special materials, colors and light are used. Fig.11 As shown in (7), the structure is two-dimensional, the shape, structure and distribution of the facial features are close to real people, special materials, colors and light. Taking the character and dark style as an example, refer to Fig.11 As shown in (8), the structure is three-dimensional, the facial features are exaggerated, and the materials, colors and light are special.
[0125] The specified style is a pre-specified style and is also the style in the identification style image to be generated. The specified style can be one of the aforementioned multiple styles, or the specified style can be a style in the training data used in the training phase, or, since the style image generation model also has the ability of zero-sample learning, the specified style can also be a style that does not appear in the training data used in the training phase, and is a completely new style. The style reference image is an image that does not include the specified identification information and has the specified style. There is at least one style reference image, and when there are multiple style reference images, the multiple style reference images are images that do not include the specified identification information and have the specified style.
[0126] The logo style image is an image that includes specified logo information and has a specified style. The logo style image maintains the consistency of logo information with the logo image and the consistency of style with the style reference image, and is also called a stylized image with consistent logo information. For example, taking a person as an example, the logo image is a non-styled image of the person, and the style reference image is an image of an animal in the style of a fairy hero, then the logo style image is an image of the person in the style of a fairy hero.
[0127] Referring to the description of the diffusion model in the above embodiment, the core idea of the diffusion model includes two main processes: a forward diffusion process and a reverse diffusion process. In the forward diffusion process, Gaussian noise is gradually added to the original image until the image becomes pure noise. The whole process is usually a series of Markov chains. In the reverse diffusion process, a series of Markov chains are used to gradually remove the prediction noise from the pure noise to restore the original image from the pure noise.
[0128] Corresponding to this embodiment, the logo style image is regarded as the original image to be restored, and the obtained Gaussian noise is regarded as pure noise after adding Gaussian noise to the original image. Therefore, the Gaussian noise obtained in this embodiment is used to simulate the characteristic distribution of the logo style image after noise addition.
[0129] The Gaussian noise is obtained by random sampling from a Gaussian distribution ranging from 0 to 1. Ideally, the feature distribution of the logo style image after adding noise obeys a Gaussian distribution ranging from 0 to 1. It can be understood that in the inference stage, the logo style image actually does not exist, or is understood to be unknown to the style image generation model. In this embodiment, the feature distribution of the logo style image after adding noise is simulated by Gaussian noise.
[0130] Exemplarily, a logo image, at least one style reference image and Gaussian noise are obtained. It should also be noted that the above-mentioned obtaining steps can be performed in parallel or in sequence, which is not limited in this embodiment.
[0131] Step 340 , encoding identification features of the identification image; encoding style features of at least one style reference image; and encoding noise features of Gaussian noise.
[0132] The image features of the logo image are referred to as logo features. Exemplarily, the logo image is obtained, and the logo features of the logo image are encoded by the first encoder and the first feature network of the style image generation model. In some embodiments, the logo image can be used as a reference for the logo information in the logo style image, that is, the logo style image maintains the consistency of the logo information with the logo image.
[0133] The image features of the style reference image are referred to as style features. Exemplarily, at least one style reference image is obtained, and the style features of the at least one style reference image are encoded by a third encoder and a third feature network of the style image generation model. In some embodiments, the style reference image can be used as a reference for identifying the style in the style image, that is, the style image is identified to maintain consistency in style with the style reference image.
[0134] Exemplarily, Gaussian noise is obtained, and the noise features of the Gaussian noise are encoded through the second feature network of the style image generation model to facilitate subsequent feature fusion. It should also be noted that the aforementioned encoding steps can be performed in parallel or in sequence, and this embodiment does not limit this.
[0135] Step 360: Fuse the noise feature, the identification feature and the style feature to obtain a fused feature; and perform denoising on the fused feature to obtain a denoised feature.
[0136] Feature fusion is used to optimize the combination of noise features, identification features and style features. The feature fusion methods include at least one of the following: splicing, multiplication, addition, and weighting. Denoising refers to the process of reducing or eliminating noise, and the original image is restored through continuous denoising.
[0137] Exemplarily, through the second feature network of the style image generation model, the noise feature, the identification feature and the style feature are fused to obtain a fused feature, which is a noisy feature, and denoising is performed on the fused feature to obtain a denoised feature. In the denoising process, the identification feature and the style feature are regarded as reference features to guide the denoising direction. It should also be noted that the second encoder is not used in the inference stage, but is used in the training stage, which will be explained in the corresponding steps of the following embodiments.
[0138] Step 380: Decode the denoising features to generate a logo style image.
[0139] The image restored by continuous denoising is the logo style image. Exemplarily, the denoising features are decoded by the decoder of the style image generation model to generate the logo style image. The logo style image maintains the consistency of logo information with the logo image and the consistency of style with the style reference image.
[0140] In summary, the style image generation method provided by the embodiment of the present application is that a computer device obtains a logo image, at least one style reference image and Gaussian noise, the logo image is an image including specified logo information and not having a specified style, at least one style reference image is an image having a specified style and not including specified logo information, Gaussian noise is used to simulate the feature distribution of the logo style image after noise addition, the logo style image is an image including specified logo information and having a specified style; encodes the logo feature of the logo image; encodes the style feature of at least one style reference image; and, encodes the noise feature of Gaussian noise; fuses the noise feature, logo feature and style feature to obtain a fused feature; and, denoises the fused feature to obtain a denoised feature; decodes the denoised feature to generate a logo style image. Accordingly, a style image generation method is provided, which is a method that can generate a stylized image that maintains consistency in logo information. By obtaining a logo image with specified logo information, encoding its logo features, and obtaining at least one style reference image with a specified style, encoding its style features, and fusing the logo features, style features and noise features of Gaussian noise, a fusion feature is obtained. On the one hand, the logo features and style features can be used as reference features to guide the denoising process of the fusion features, thereby restoring a logo style image including the specified logo information and having the specified style, so that the logo style image maintains consistency with the logo information of the logo image and maintains consistency with the style of the style reference image. On the other hand, compared with the text of the related technology, the logo image and the style reference image have stronger concreteness, realize the precise control of the logo information and style, thereby improving the accuracy of the logo style image.
[0141] Feature fusion
[0142] In some embodiments, the second feature network of the style image generation model is an image segmentation model (UNet). The image segmentation model (UNet) mainly includes two structures: ResNet and Transformer. ResNet is mainly composed of a convolutional network for encoding image features. Transformer is mainly composed of an attention network (Attention), which is used to encode image features and introduce conditional features. In this embodiment, the conditional features are image features. That is, the second feature network of the style image generation model includes an attention network (Attention). Exemplarily, step 360 fuses noise features, identification features, and style features to obtain fused features, which are implemented as steps 361, 362, and 363:
[0143] Step 361, transferring the identification feature and the style feature to the noise feature to obtain a splicing feature;
[0144] Step 362, mapping the noise feature and the splicing feature based on the attention mechanism to obtain the attention feature;
[0145] Step 363, perform a convolution operation based on the attention feature to obtain a fusion feature.
[0146] In the denoising process, in order to regard the identification features and style features as reference features and guide the denoising direction, the identification features and style features are transferred to the noise features by vector splicing, so as to transfer the identification features and style features to the attention network in the second feature network, and the splicing features can be obtained after splicing. The attention network continues to fuse these features, that is, the noise features and splicing features are mapped based on the attention mechanism to obtain the attention features; the convolution operation is performed on the attention features through the convolution network in the second feature network to obtain the fused features. After obtaining the fused features, the noise prediction and denoising process can be performed based on the fused features.
[0147] In this embodiment, the logo feature and the style feature can be transferred to the noise feature, and the fusion of the three is achieved, which is conducive to generating a stylized logo style image that maintains the consistency of the logo information.
[0148] Specifically, step 361 is implemented as step 3610:
[0149] Step 3610, concatenate the features located at the same position among the noise feature, the identification feature, and the style feature to obtain a concatenated feature.
[0150] During splicing, the identification features and the style features are spliced to the same position in the attention network in the second feature network. Specifically, the features at the same position among the noise features, the identification features, and the style features are spliced to obtain the spliced features.
[0151] In this embodiment, a vector concatenation method is used to transfer the identification feature and the style feature to the noise feature, thereby achieving a fusion of the three.
[0152] In some embodiments, the first feature network and the third feature network of the style image generation model are also image segmentation models (UNet). That is, the first feature network and the third feature network of the style image generation model also include an attention network (Attention). The noise feature is also called the noise attention feature, the identification feature is also called the identification attention feature, and the style feature is also called the style attention feature. At this time, the noise feature includes the noise K matrix (K unet ), noise V matrix (V unet ), the identification features include the identification K matrix (K person ), identify the V matrix (V person ), the style features include the style K matrix (K style ), style V matrix (V style ). Then the splicing features include splicing K matrix and splicing V matrix. Specifically, step 3610 is implemented as step 3611:
[0153] Step 3611, concatenate the noise K matrix, the identification K matrix and the style K matrix to obtain the concatenated K features; and concatenate the noise V matrix, the identification V matrix and the style V matrix to obtain the concatenated V matrix.
[0154] The concatenation can be performed in the form of rows or columns, and is implemented using the concatenation function (concat, cat).
[0155] Exemplarily, a concatenation function is used to concatenate the noise K matrix, the identification K matrix and the style K matrix to obtain a concatenated K feature; and a noise V matrix, an identification V matrix and a style V matrix are concatenated to obtain a concatenated V matrix. The concatenated K matrix is expressed as: cat(K unet , K style , K person ). The concatenated V matrix is expressed as: cat(V unet , V style , V person ).
[0156] This embodiment provides a specific splicing method in the attention network, which can not only realize the splicing with the logo features and style features, but also does not change the original feature dimensions of the noise features, which is conducive to generating a stylized logo style image that maintains consistency in the logo information.
[0157] In some embodiments, the noise signature also includes a noise Q matrix (Q unet ), the splicing features include splicing K matrix and splicing V matrix. Specifically, step 362 is implemented as step 3620:
[0158] Step 3620, mapping the noise Q matrix, the splicing K matrix and the splicing V matrix based on the attention mechanism to obtain the attention feature.
[0159] After splicing, based on the attention mechanism of the attention network in the second feature network, the noise Q matrix, the splicing K matrix and the splicing V matrix are mapped to obtain the attention feature y, which is expressed as:
[0160] y=Attention(Q unet , cat(K unet , K style , K person ), cat(V unet , Vs tyle , V person ))
[0161] Among them, Q unet , K unet and V unet represents the attention network input in the second feature network, K person and V person is the attention network input in the concatenated first feature network, K style and V style It is used as the input of the attention network in the spliced third feature network. Next, the convolution operation is performed through the convolution network in the second feature network to output the fused features.
[0162] This embodiment provides a method for processing splicing features in an attention network, thereby achieving feature fusion.
[0163] As an example, Figure 7 : is a schematic diagram of feature splicing provided by an exemplary embodiment of the present application. The noise feature is also called the noise attention feature (UNet Attention Feature), the identification feature is also called the identification attention feature (Person Attention Feature), and the style feature is also called the style attention feature (Style Attention Feature). In the denoising process, the identification attention feature and the style attention feature are regarded as reference attention features (Reference Attention Feature). When splicing, the splicing noise K matrix (K unet ), identify the K matrix (K person ) and style K matrix (K style ), get the spliced K features; and the spliced noise V matrix (V unet ), identify the V matrix (V person ) and style V matrix (V style), and obtain the spliced V matrix. Based on the attention mechanism of the attention network in the second feature network, the noise Q matrix, the spliced K matrix and the spliced V matrix are mapped to obtain the attention feature y, which is expressed as:
[0164] y=Attention(Q unet , cat(K unet , K style , K person ), cat(V unet , V style , V person ))
[0165] Next, the convolution network in the second feature network performs a convolution operation based on the attention feature, outputs a fused feature, and then performs denoising on the fused feature to obtain a denoised feature. By decoding the denoised feature, the logo style image can be obtained.
[0166] In some embodiments, the second feature network of the style image generation model is an image segmentation model (UNet). The input of the image segmentation model (UNet) is the noisy feature X t , whose function is to predict noise and generate denoising feature X after removing the prediction noise t-1 After t steps of iterative processing, the noisy feature X t Convert to denoised feature X0. The smaller t is, the less predicted noise is. When t=1, the image segmentation model (UNet) can obtain nearly noise-free data. After this denoising step, the denoising result is denoised feature X0. After the denoised feature X0 is decoded by the decoder, a meaningful image can be generated. Corresponding to this embodiment, a logo style image is generated. Specifically, step 360 performs denoising on the fused feature to obtain the denoised feature, which is implemented as steps 364 and 365:
[0167] Step 364, performing noise prediction based on the fused features to obtain predicted noise;
[0168] Step 365, denoising the fused features according to the predicted noise to obtain denoised features.
[0169] Exemplarily, through the second feature network in the style image generation model, based on the fused feature, noise prediction is performed according to the time step (time step, t) to obtain the predicted noise, and denoising is performed on the fused feature according to the predicted noise, specifically, the predicted noise can be subtracted to obtain the denoised feature X t-1 , after t steps of iterative processing, the denoised feature X0 is obtained.
[0170] This embodiment provides a denoising method. By performing multiple iterations of denoising, denoising features can be obtained. By decoding the denoising features, a stylized logo style image that maintains consistency of logo information can be obtained.
[0171] Feature encoding
[0172] In some embodiments, the first encoder and the first feature network of the style image generation model have the function of feature encoding. Specifically, step 340 encodes the identification features of the identification image, which is implemented as steps 341, 342 and 343:
[0173] Step 341, mapping the identification image to obtain identification mapping features;
[0174] Step 342, performing noise addition on the identification mapping feature to obtain a noisy identification mapping feature;
[0175] Step 343, encode the noise-added identification mapping feature to obtain the identification feature.
[0176] Exemplarily, through the first encoder in the style image generation model, the identification image is mapped to obtain the identification mapping feature. Noise is added to the identification mapping feature according to the time step (time step, t) to obtain the noisy identification mapping feature. Through the first feature network in the style image generation model, the noisy identification mapping feature is encoded to obtain the identification feature. Among them, the identification feature may also be called the identification attention feature. In some embodiments, the amount of noise added to each time step may be the same. Noise addition can be achieved by a noise addition algorithm or by another noise addition encoder.
[0177] This embodiment provides a method for encoding a logo feature of a logo image. The logo feature is an image feature that is more concrete than text and is conducive to the subsequent generation of a logo style image.
[0178] In some embodiments, the third encoder and the third feature network of the style image generation model have the function of feature encoding. Specifically, step 340 encodes the style features of at least one style reference image, which is implemented as steps 344, 345 and 346:
[0179] Step 344, mapping at least one style reference image to obtain a style mapping feature;
[0180] Step 345, performing noise addition on the style mapping feature to obtain a noisy style mapping feature;
[0181] Step 346: Encode the noise-added style mapping feature to obtain a style feature.
[0182] Exemplarily, at least one style reference image is mapped through a third encoder in the style image generation model to obtain a style mapping feature. Noise is added to the style mapping feature according to a time step (time step, t) to obtain a noisy style mapping feature. The noisy style mapping feature is encoded through a third feature network in the style image generation model to obtain a style feature. Among them, the style feature may also be referred to as a style attention feature. In some embodiments, the amount of noise added to each time step may be the same. Noise addition may be implemented by a noise addition algorithm or by another noise addition encoder.
[0183] This embodiment provides a method for encoding style features of a style reference image. The style feature is a type of image feature that is more concrete than text and is conducive to the subsequent generation of a logo style image.
[0184] As an example, Figure 8 Schematic diagram of a style image generation method provided by an exemplary embodiment of the present application. Taking a person as an example, the first encoder 11 in the style image generation model 200 is specifically an image encoder (VAE Encoder) 14, the third encoder 13 is specifically an image encoder (VAE Encoder) 16, the first feature network 21 is specifically an identification image segmentation model (PersonNet) 24, the second feature network 22 is specifically an image segmentation model (UNet) 25, the third feature network 23 is specifically a style image segmentation model (StyleNet) 26, and the decoder 30 is specifically an image encoder (VAE Decoder) 31.
[0185] Obtain a logo image 210, map the logo image 210 based on the image encoder 14, and obtain a logo mapping feature; perform noise addition on the logo mapping feature to obtain a noisy logo mapping feature 211; and encode the noisy logo mapping feature 211 based on a logo image segmentation model (PersonNet) 24 to obtain a logo feature.
[0186] A style reference image 220 is obtained, and the style reference image 220 is mapped based on the image encoder 16 to obtain a style mapping feature; the style mapping feature is denoised to obtain a noisy style mapping feature 221; and the noisy style mapping feature 221 is encoded based on a style image segmentation model (StyleNet) 26 to obtain a style feature.
[0187] Gaussian noise 230 is obtained, and the noise features of Gaussian noise 230 are encoded based on the image segmentation model (UNet) 25; the identification features and style features are transferred to the noise features based on the image segmentation model (UNet) 25 to obtain the splicing features, specifically: splicing the noise K matrix, the identification K matrix and the style K matrix to obtain the splicing K features, splicing the noise V matrix, the identification V matrix and the style V matrix to obtain the splicing V matrix, and the splicing effect is referenced Figure 7 As shown; mapping noise features and splicing features based on the attention mechanism, specifically: mapping the noise Q matrix, splicing K matrix and splicing V matrix based on the attention mechanism to obtain the attention feature, performing a convolution operation on the attention feature to obtain a fused feature, which is a noisy feature, and performing denoising on the fused feature to obtain a denoised feature.
[0188] The denoising features are decoded based on the image encoder (VAE Decoder) 31 to generate a logo style image 240 .
[0189] (III) Training phase of style image generation model
[0190] In some embodiments, the training phase of the style image generation model is performed by a computer device, which may be Figure 1 The terminal 120 and / or server 140 shown. Optionally, the computer devices corresponding to the training phase and the inference phase of the style image generation model can be the same or different, which can be determined according to the performance of the computer device or actual technical needs. The method also includes at least some of the steps 410, 420, 430, and 440 (not shown in the figure):
[0191] Step 410, obtaining training data, each training data includes a sample identification image, a sample style reference image and a sample identification style image, the sample identification image is an image including sample identification information and not having the sample style, the sample style reference image is an image having the sample style and not having the sample identification information, and the sample identification style image is an image having the sample identification information and having the sample style;
[0192] The training data is the data used in the training phase of the style image generation model. The training data includes multiple data, each of which consists of three sample images, namely: a sample identification image, a sample style reference image, and a sample identification style image. The construction method of the training data is described in the following embodiment.
[0193] The sample identification image is an image including sample identification information and not having a sample style, the sample style reference image is an image having a sample style and not having sample identification information, and the sample identification style image is an image having sample identification information and having a sample style. The sample identification style image maintains consistency with the sample identification information in the sample identification image, and the sample identification style image maintains consistency with the sample style of the sample style reference image.
[0194] For example, refer to Fig. 9 , training data 1 includes: Fig. 9 The sample identification image shown in (1-1), Fig. 9 The sample logo style image shown in (1-2), Fig. 9 The sample style reference images shown in (1-3) are as follows. Training data 2 includes: Fig. 9 The sample identification image shown in (2-1), Fig. 9 The sample logo style image shown in (2-2), Fig. 9 The sample style reference image shown in (2-3).
[0195] In some embodiments, the sample identification image is an image including sample identification information and no style, or the sample identification image is an image including sample identification information and having other styles, and the other styles are different from the sample style, for example, the sample style is a fairy-tale style, and the other styles are line styles. Whether the sample identification image in the training phase has a style needs to be consistent with the identification image in the reasoning phase. The sample identification information can be the sample identification information of the sample organism, where the sample organism can be a person, an animal, etc. Taking a person as an example, the sample identification image is called a sample person image, and the sample identification information can be at least one of a sample person's name, nickname, gender, age, facial features, and a film or television drama in which he has participated. Optionally, the facial features include at least one of the distribution, size, shape, face shape, hairstyle, and muscle distribution of facial features.
[0196] The sample style is the visual effect presented by a sample image. A sample image usually has a sample style. The sample style can be at least one of the following: Anime style, Xianxia style, Oil painting style, Watercolor style, Oneline-Drawing style, Black Magic style, Ink style, Ice Scupture style. Anime style also includes at least one of the following: 3D Real Cartoon style, 3D Cartoon style, 2D Anime style, 2D Comic style.
[0197] A sample style can be characterized by both structure and texture. The structure is characterized by at least one of the following parameters: two-dimensional (2D), three-dimensional (3D), structure, shape and distribution of facial features. The texture is characterized by at least one of the following parameters: color, material, light.
[0198] For example, take characters and 3D real-life cartoon style as an example, refer to Fig.11 As shown in (1), the structure is three-dimensional, the shape, structure and distribution of the facial features are close to real people, and the special materials, colors and light are used. Fig.11 As shown in (2), the structure is three-dimensional, the shape, structure and distribution of the facial features are close to real people, and the special materials, colors and light are used. Fig.11 As shown in (3), the structure is three-dimensional, the facial features are exaggerated, and the material, color and light are special. Fig.11 As shown in (4), the structure is three-dimensional, the shape, structure and distribution of the facial features are close to real people, and the special materials, colors and light are used. Fig.11 As shown in (5), the structure is two-dimensional, the facial features are exaggerated, and special materials, colors and light are used. Taking the figure and line style as an example, refer to Fig.11 As shown in (6), the structure is two-dimensional, the facial features are exaggerated, and special materials, colors and light are used. Fig.11 As shown in (7), the structure is two-dimensional, the shape, structure and distribution of the facial features are close to real people, special materials, colors and light. Taking the character and dark style as an example, refer to Fig.11 As shown in (8), the structure is three-dimensional, the facial features are exaggerated, and the materials, colors and light are special.
[0199] Step 420: Encode the sample identification feature of the sample identification image; encode the sample identification style feature of the sample identification style image; and encode the sample style feature of the sample style reference image.
[0200] The image features of the sample identification image are referred to as sample identification features. The image features of the sample style reference image are referred to as sample style features. The image features of the sample identification style image are referred to as sample identification style features. Exemplarily, the sample identification features of the sample identification image are encoded by a first encoder and a first feature network of the style image generation model, the sample identification style features of the sample identification style image are encoded by a second encoder and a second feature network of the style image generation model, and the sample style features of the sample style reference image are encoded by a third encoder and a third feature network of the style image generation model.
[0201] Step 430: Fuse the sample identification feature, the sample identification style feature and the sample style feature to obtain a sample fusion feature; and perform denoising on the sample fusion feature to obtain a sample denoising feature.
[0202] Feature fusion is used to optimize the combination of sample identification features, sample identification style features and sample style features. The feature fusion methods include at least one of the following: splicing, multiplication, addition, and weighting. Denoising refers to the process of reducing or eliminating noise, and the original image is restored by continuous denoising. Among them, in the training stage, the original image is the sample identification style image in the training data.
[0203] Exemplarily, the sample identification feature, the sample identification style feature and the sample style feature are fused through the second feature network of the style image generation model to obtain a sample fusion feature, which is a noisy feature, and the sample fusion feature is denoised to obtain a sample denoised feature. In the denoising process, the sample identification feature and the sample style feature are regarded as sample reference features to guide the denoising direction.
[0204] Step 440 , decode the sample denoising features to obtain a predicted identification style image; and adjust the model parameters of the style image generation model by taking reducing the difference between the predicted identification style image and the sample identification style image as a training goal.
[0205] The image restored by continuous denoising is called the predicted signature style image. When the model accuracy of the style image generation model is high enough, the predicted signature style image is very close to or even completely identical to the sample signature style image. Exemplarily, the sample denoising features are decoded by the decoder of the style image generation model to obtain the predicted signature style image. The model parameters of the style image generation model are adjusted with the goal of reducing the difference between the predicted signature style image and the sample signature style image.
[0206] In some embodiments, based on the predicted identification style image and the sample identification style image, a loss function is determined, and the model parameters of the style image generation model are adjusted with minimizing the loss function as the training goal. In some possible implementations, the model parameters in the second feature network of the style image generation model are adjusted. In a specific implementation, the parameters of the attention network in the second feature network are adjusted. Optionally, the loss function can be at least one of the following: mean square error (MSE), cross entropy loss function (Cross-Entropy Loss).
[0207] This embodiment provides a training method for a style image generation model, which enables the style image generation model to fully utilize the sample identification features of the sample identification image and the sample style features of the sample style reference image to achieve feature fusion and generate a predicted identification style image, thereby generating an identification style image in the inference stage.
[0208] Feature fusion
[0209] In some embodiments, the second feature network of the style image generation model is an image segmentation model (UNet). The image segmentation model (UNet) mainly includes two structures: ResNet and Transformer. ResNet is mainly composed of a convolutional network for encoding image features. Transformer is mainly composed of an attention network (Attention), which is used to encode image features and introduce conditional features. In this embodiment, the conditional features are image features. That is, the second feature network of the style image generation model includes an attention network (Attention). Exemplarily, step 430 fuses the sample identification feature, the sample identification style feature and the sample style feature to obtain a sample fusion feature, which is implemented as steps 431, 432 and 433:
[0210] Step 431, transferring the sample identification feature and the sample style feature to the sample identification style feature to obtain the sample splicing feature;
[0211] Step 432, mapping the sample identification style feature and the sample splicing feature based on the attention mechanism to obtain the sample attention feature;
[0212] Step 433, performing a convolution operation based on the sample attention feature to obtain a sample fusion feature.
[0213] In the denoising process, in order to regard the sample identification features and sample style features as sample reference features and guide the denoising direction, the sample identification features and sample style features are transferred to the sample identification style features by means of vector splicing, so as to transfer the sample identification features and sample style features to the attention network in the second feature network, and the sample splicing features can be obtained after splicing. The attention network continues to fuse these features, that is, the sample identification style features and the sample splicing features are mapped based on the attention mechanism to obtain the sample attention features; the convolution operation is performed on the sample attention features through the convolution network in the second feature network to obtain the sample fusion features. After obtaining the sample fusion features, the noise prediction and denoising process can be performed based on the sample fusion features.
[0214] In this embodiment, the sample identification feature and the sample style feature can be transferred to the sample identification style feature, and the fusion of the three is achieved, which is conducive to generating a stylized predicted identification style image that maintains the consistency of identification information.
[0215] Specifically, step 431 is implemented as step 4310:
[0216] Step 4310, concatenate the sample identification style feature, the sample identification feature, and the sample features at the same position in the sample style feature to obtain a sample concatenated feature.
[0217] During splicing, the sample identification feature and the sample style feature are spliced to the same position in the attention network in the second feature network. Specifically, the sample identification style feature, the sample identification feature and the sample features at the same position in the sample style feature are spliced to obtain the sample splicing feature.
[0218] In this embodiment, a vector concatenation method is used to transfer the sample identification feature and the sample style feature to the sample identification style feature, thereby achieving a fusion of the three.
[0219] In some embodiments, the first feature network and the third feature network of the style image generation model are also image segmentation models (UNet). That is, the first feature network and the third feature network of the style image generation model also include attention networks (Attention). At this time, the sample identification style feature includes the sample identification style K matrix (K unet ) and the sample identification style V matrix (V unet ), the sample identification features include the sample identification K matrix (K person ), sample identification matrix V (V person ), the sample style features include the sample style K matrix (K style ), sample style V matrix (V style ), the sample splicing features include sample splicing K features and sample splicing V features. Specifically, step 4310 is implemented as step 4311:
[0220] Step 4311, concatenate the sample identification style K matrix, the sample identification K matrix and the sample style K matrix to obtain the sample concatenation K feature; and concatenate the sample identification style V matrix, the sample identification V matrix and the sample style V matrix to obtain the sample concatenation V matrix.
[0221] The concatenation can be performed in the form of rows or columns, and is implemented using the concatenation function (concat, cat).
[0222] Exemplarily, a concatenation function is used to concatenate the sample identification style K matrix, the sample identification K matrix and the sample style K matrix to obtain the sample concatenation K feature; and concatenate the sample identification style V matrix, the sample identification V matrix and the sample style V matrix to obtain the sample concatenation V matrix. The sample concatenation K matrix is expressed as: cat(K unet , K style , K person ). The sample concatenation V matrix is expressed as: cat(V unet , V style , V person ).
[0223] This embodiment provides a specific splicing method in the attention network, which can realize the splicing with the sample identification features and the sample style features without changing the original feature dimensions of the sample identification style features, which is conducive to generating a stylized predicted identification style image that maintains consistency in identification information and improves the training speed.
[0224] In some embodiments, the sample identification style feature further includes a sample identification style Q matrix (Q unet ), the sample splicing features include sample splicing K features and sample splicing V features. Specifically, step 432 is implemented as step 4320:
[0225] Step 4320, based on the attention mechanism, maps the sample identification style Q matrix, the sample splicing K matrix and the sample splicing V matrix to obtain the sample attention features.
[0226] After splicing, based on the attention mechanism of the attention network in the second feature network, the sample identification style Q matrix, the sample splicing K matrix and the sample splicing V matrix are mapped to obtain the sample attention feature y, which is expressed as:
[0227] y=Attention(Q unet , cat(K unet , K style , K person ), cat(V unet , V style , V person ))
[0228] Among them, Q unet , K unet and V unet represents the attention network input in the second feature network, K person and V person is the attention network input in the concatenated first feature network, K style and V styleIt is used as the input of the attention network in the spliced third feature network. Next, the convolution operation is performed through the convolution network in the second feature network to output the sample fusion feature.
[0229] This embodiment provides a method for processing sample splicing features in an attention network, thereby achieving feature fusion.
[0230] In some embodiments, the second feature network of the style image generation model is an image segmentation model (UNet). The input of the image segmentation model (UNet) is the noisy feature X t , whose function is to predict noise and generate denoising feature X after removing the prediction noise t-1 After t steps of iterative processing, the noisy feature X t Converted to denoised feature X0. The smaller t is, the less predicted noise is. When t=1, the image segmentation model (UNet) can obtain nearly noise-free data. After this denoising step, the denoising result is denoised feature X0. After the denoised feature X0 is decoded by the decoder, a meaningful image can be generated. Corresponding to this embodiment, a predicted signature style image is generated. Specifically, step 430 performs denoising on the sample fusion feature to obtain a sample denoising feature, which is implemented as steps 434 and 435:
[0231] Step 434, performing noise prediction based on the sample fusion feature to obtain predicted noise;
[0232] Step 435 , denoising the sample fusion feature according to the predicted noise to obtain the sample denoising feature.
[0233] Exemplarily, through the second feature network in the style image generation model, based on the sample fusion feature, noise prediction is performed according to the time step (time step, t) to obtain the predicted noise, and the sample fusion feature is denoised according to the predicted noise, specifically, the predicted noise can be subtracted to obtain the sample denoised feature X t-1 , after t steps of iterative processing, the sample denoising feature X0 is obtained.
[0234] This embodiment provides a denoising method. By performing multiple iterations of denoising, sample denoising features can be obtained. By decoding the sample denoising features, a predicted logo style image that maintains consistency in logo information and is stylized can be obtained, which is beneficial to improving the training speed.
[0235] Feature encoding
[0236] In some embodiments, the first encoder and the first feature network of the style image generation model have the function of feature encoding. Step 420 encodes the sample identification feature of the sample identification image, which is implemented as steps 421, 422 and 423:
[0237] Step 421, mapping the sample identification image to obtain sample identification mapping features;
[0238] Step 422, performing noise addition on the sample identification mapping feature to obtain a noisy sample identification mapping feature;
[0239] Step 423, encode the noise-added sample identification mapping feature to obtain the sample identification feature.
[0240] Exemplarily, the sample identification image is mapped by the first encoder in the style image generation model to obtain the sample identification mapping feature. The sample identification mapping feature is denoised according to the time step (time step, t) to obtain the noisy sample identification mapping feature. The noisy sample identification mapping feature is encoded by the first feature network in the style image generation model to obtain the sample identification feature. In some embodiments, the amount of noise added to each time step can be the same. Noise addition can be implemented by a noise addition algorithm or by another noise addition encoder.
[0241] This embodiment provides a method for encoding sample identification features of a sample identification image. The sample identification feature is an image feature that is more concrete than text, and is conducive to the subsequent generation of a predicted identification style image and improves model accuracy.
[0242] In some embodiments, the third encoder and the third feature network of the style image generation model have the function of feature encoding. Specifically, step 420 encodes the sample style features of the sample style reference image, which is implemented as steps 424, 425 and 426:
[0243] Step 424, mapping the sample style reference image to obtain sample style mapping features;
[0244] Step 425, performing noise addition on the sample style mapping feature to obtain a noisy sample style mapping feature;
[0245] Step 426: Encode the noisy sample style mapping feature to obtain the sample style feature.
[0246] Exemplarily, the sample style reference image is mapped through the third encoder in the style image generation model to obtain the sample style mapping feature. Noise is added to the sample style mapping feature according to the time step (time step, t) to obtain the noisy sample style mapping feature. The noisy sample style mapping feature is encoded through the third feature network in the style image generation model to obtain the sample style feature. In some embodiments, the amount of noise added to each time step can be the same. Noise addition can be implemented by a noise addition algorithm or by another noise addition encoder.
[0247] This embodiment provides a method for encoding sample style features of a sample style reference image. The sample style feature is an image feature that is more concrete than text, and is conducive to the subsequent generation of a predicted identification style image, thereby improving model accuracy.
[0248] In some embodiments, the second encoder and the second feature network of the style image generation model have the function of feature encoding. Specifically, step 420 encodes the sample identification style features of the sample identification style image, which is implemented as steps 427, 428 and 429:
[0249] Step 427, mapping the sample identification style image to obtain sample identification style mapping features;
[0250] Step 428, performing noise addition on the sample identification style mapping feature to obtain a noisy sample identification style mapping feature;
[0251] Step 429: Encode the noisy sample identification style mapping feature to obtain the sample identification style feature.
[0252] Exemplarily, the sample identification style image is mapped by the second encoder in the style image generation model to obtain the sample identification style mapping feature. The sample identification style mapping feature is denoised according to the time step (time step, t) to obtain the noisy sample identification style mapping feature. The noisy sample identification style mapping feature is used to characterize the feature distribution of the sample identification style image after noisy, and the noisy sample identification style mapping feature corresponds to the Gaussian noise in the inference stage. The noisy sample identification style mapping feature is encoded by the second feature network in the style image generation model to obtain the sample identification style feature. In some embodiments, the amount of noise added for each time step can be the same. Noising can be achieved by a noisy algorithm or by another noisy encoder.
[0253] This embodiment provides a method for encoding the logo style features of a sample logo style image, which is beneficial for the subsequent generation of a predicted logo style image.
[0254] As an example, Fig.10Schematic diagram of a style image generation method provided by an exemplary embodiment of the present application. Taking a person as an example, the first encoder 11 in the style image generation model 200 is specifically an image encoder (VAE Encoder) 14, the second encoder 12 is specifically an image encoder (VAE Encoder) 15, the third encoder 13 is specifically an image encoder (VAE Encoder) 16, the first feature network 21 is specifically an identification image segmentation model (PersonNet) 24, the second feature network 22 is specifically an image segmentation model (UNet) 25, the third feature network 23 is specifically a style image segmentation model (StyleNet) 26, and the decoder 30 is specifically an image encoder (VAE Decoder) 31.
[0255] Obtain a sample identification image 250, map the sample identification image 250 based on the image encoder 14, and obtain identification mapping features; perform noise addition on the sample identification mapping features to obtain noisy sample identification mapping features 251; based on the identification image segmentation model (PersonNet) 24, encode the noisy sample identification mapping features 251 to obtain sample identification features.
[0256] A sample style reference image 260 is obtained, and the sample style reference image 260 is mapped based on the image encoder 16 to obtain a sample style mapping feature; the sample style mapping feature is denoised to obtain a noisy sample style mapping feature 261; and the noisy sample style mapping feature 261 is encoded based on a style image segmentation model (StyleNet) 26 to obtain a sample style feature.
[0257] A sample identification style image 270 is obtained, and the sample identification style image 270 is mapped based on the image encoder 15 to obtain a sample identification style mapping feature; the sample identification style mapping feature is denoised to obtain a noisy sample identification style mapping feature 271; and the noisy sample identification style mapping feature 271 is encoded based on the image segmentation model (UNet) 25 to obtain a sample identification style feature.
[0258] Based on the image segmentation model (UNet) 25, the sample identification features and sample style features are transferred to the sample identification style features to obtain the sample splicing features. Specifically, the sample identification style K matrix, the sample identification K matrix and the sample style K matrix are spliced to obtain the sample splicing K features, and the sample identification style V matrix, the sample identification V matrix and the sample style V matrix are spliced to obtain the sample splicing V matrix. The splicing effect can be referenced Figure 7As shown; based on the attention mechanism, the sample identification style features and the sample splicing features are mapped, specifically: based on the attention mechanism, the sample identification style Q matrix, the sample splicing K matrix and the sample splicing V matrix are mapped to obtain the sample attention features, and a convolution operation is performed on the sample attention features to obtain the sample fusion features, which are noisy features. De-noising is performed on the sample fusion features to obtain the sample denoising features.
[0259] Based on the sample denoising features decoded by the image encoder (VAE Decoder) 31, a predicted identification style image 280 is generated. With the training goal of reducing the difference between the sample identification style image 270 and the predicted identification style image 280, the model parameters of the style image generation model 200 are adjusted, and specifically, the parameters of the attention network of the image segmentation model (UNet) 25 are adjusted.
[0260] Training data
[0261] In some embodiments, before step 410, training data needs to be constructed. Then the method further includes steps 520, 540 and 560:
[0262] Step 520, constructing a sample style reference image set, each sample style reference image set including a plurality of sample style reference images having the same sample style;
[0263] Step 540, constructing sample image pairs, each sample image pair including a sample identification style image and a sample identification image corresponding to the sample identification style image;
[0264] Step 560: construct training data based on the sample image pairs and the sample style reference image set.
[0265] The training data includes multiple, each training data consists of three sample images, namely: a sample identification image, a sample style reference image and a sample identification style image. For a training data, the sample style reference image is obtained from the sample style reference image set, and the sample style reference image and its corresponding sample identification image are obtained from a pre-constructed sample image pair.
[0266] Exemplarily, a sample style reference image set is first constructed, each sample style reference image set includes several sample style reference images with the same sample style, and then sample image pairs are constructed, each sample image pair includes a sample identification style image, and a sample identification image corresponding to the sample identification style image, and finally, training data is constructed based on the sample image pairs and the sample style reference image sets.
[0267] This embodiment provides a method for constructing training data, each training data includes three images, so that the style image generation model can learn separate identification information and style at the same time in each training, and also learn fused identification information and style, which is conducive to generating a predicted identification style image that maintains consistency in identification information and is stylized.
[0268] In some embodiments, a large number of sample style reference images are first obtained. A sample style reference image has a sample style, and it is necessary to cover as many sample styles as possible. Then, based on these sample style reference images, sample style reference images of the same sample style are clustered in a clustering manner to construct a set of sample style reference images. The method of obtaining a large number of sample style reference images includes:
[0269] Method 1: Use crawler technology to obtain sample style reference images of various sample styles from public websites, large model platforms, and text image generation model sharing platforms. Among them, the text image generation model is a text image model, such as the Flux1dev model, which can generate images based on text.
[0270] For example, the text image generation model sharing platform provides an application programming interface (API), where users can upload and / or download sample style reference images of different sample styles. It also provides a variety of self-trained style models, which can significantly improve style coverage. For example, if you select images generated by the Flux1 dev model, you will get a total of 700+ style models and 50,000+ images.
[0271] Method 2: Obtain image-text pairs from a public dataset, evaluate the similarity between the image and text in each image-text pair, and use the images in the image-text pairs whose similarity is greater than a set threshold as sample style reference images.
[0272] Exemplarily, the public dataset is a high-quality image-text pair dataset (LAION-5B), which contains multi-style images. An open source multimodal pre-training model (Contrastive Language-Image Pre-training, CLIP) is used to evaluate the similarity between the image and text in each image-text pair in the dataset, and image-text pairs with a similarity lower than a set threshold are deleted. Among them, when the text is in English, the English threshold is set to 0.28, and the remaining thresholds are set to 0.26. It should also be noted that the multimodal pre-training model can map images and texts into a shared vector space, so that the semantic relationship between images and texts can be understood. This shared vector space enables the multimodal pre-training model to achieve unsupervised joint learning between images and texts, which can be used for various visual tasks and language tasks.
[0273] Method 3: Use the text image generation model's text image capability to obtain several sample style reference images. Specifically, step 520 is implemented as steps 522, 524, and 526:
[0274] Step 522, constructing a plurality of style texts based on the first text template;
[0275] Step 524, inputting a plurality of style texts into a text image generation model, and generating a plurality of sample style reference images corresponding to each of the plurality of style texts;
[0276] Step 526, constructing a set of sample style reference images based on several sample style reference images corresponding to each style text of several style texts; wherein the first text template includes: a first identifier placeholder, a second identifier placeholder, a third style placeholder and an image quality descriptor, the first identifier placeholder is used to fill in the first identifier information, the second identifier placeholder is used to fill in the second identifier information, the third style placeholder is used to fill in the style, and the image quality descriptor is used to fill in the image quality description.
[0277] The first text template is a template for constructing a style text, and the style text is used to describe a style and at least one identification information. The first text template includes: a first identification placeholder (Token1), a second identification placeholder (Token2), a third style placeholder (Token3), and an image quality descriptor. The first identification placeholder is used to fill in the first identification information, the second identification placeholder is used to fill in the second identification information, the third style placeholder is used to fill in the style, and the image quality descriptor is used to fill in the image quality description.
[0278] For example, the first identification information is gender, which is at least one of the following: {"male / boy", "female / girl"}, the second identification information is age (years old), which is at least one of the following: {5, 10, 20, 30, 40, 50, 60, 70}, and the style (style) can be a style word, which is at least one of the following: 3D Real Cartoon, 3D Cartoon, 2D Anime, 2D Comic, Oneline-Drawing, Black Magic, Ink, Ice Scupture. The image quality description can be a description word or a description text, and the image quality description includes at least one of the following: masterpiece, best quality, crop top, perfectly symmetrical face, detailed skin, amazing photograph, 8K, high quality, photorealistic, and realism.
[0279] Exemplarily, based on the first text template, several style texts are constructed. The several style texts are input into the text image generation model, and the noise random number of the initialization input of the text image generation model is changed to generate several sample style reference images corresponding to each style text of the several style texts, thereby generating sample style reference images of different genders, multiple age groups, and different styles. Then, based on the several sample style reference images corresponding to each style text of the several style texts, a sample style reference image set is constructed.
[0280] This embodiment provides a method for constructing a set of sample style reference images under each style, which is beneficial to the subsequent optimization of the style model and the construction of training data.
[0281] Specifically, step 526 is implemented as step 5262, step 5264 and step 5266:
[0282] Step 5262, determining a sample style reference image that meets a first screening condition from a plurality of sample style reference images corresponding to each style text of a plurality of style texts;
[0283] Step 5264, clustering the sample style reference images that meet the first screening condition to obtain a plurality of clusters;
[0284] Step 5266, constructing a set of sample style reference images based on several clusters; wherein, the first screening condition includes: a first similarity between the style text and the corresponding sample style reference image is greater than a first threshold, or a second similarity between different sample style reference images corresponding to the style text is greater than a second threshold; each of the several clusters corresponds to a set of sample style reference images.
[0285] Since the quality of sample style reference images varies, filtering is required. The filtering method can be:
[0286] Method 1: Manual screening: Manually select sample style reference images with consistent style and high quality.
[0287] Method 2, screening based on image recognition model: The image recognition model is used to perform face detection and face alignment. For example, taking people as an example, the image recognition model is the face detection model (RetinaFace) model, which is used to realize face recognition. Its principle is based on a one-stage target detector, which uses multi-scale feature fusion and multi-task learning methods to accurately detect and align faces at different scales and directions. Reference Fig.12 , input each of the sample style reference images of several sample style reference images into the image recognition model, and obtain the coordinates (x, y, w, h) of the rectangle containing the face, where (x, y) is the coordinate of the upper left corner of the rectangle, and w and h are the width and height of the rectangle respectively. The sample style reference image whose size of the contained rectangle is greater than the size threshold is determined as the sample style reference image that meets the first screening condition.
[0288] Method 3, based on the open source multimodal pre-training model (Contrastive Language-Image Pre-training, CLIP), evaluates the first similarity between each style text and its corresponding sample style reference image, that is, the similarity between the text and the image, and evaluates the second similarity between different sample style reference images corresponding to each style text, that is, the similarity between the images. The sample style reference image with the first similarity greater than the first threshold, or the second similarity greater than the second threshold, is regarded as the sample style reference image that meets the first screening condition.
[0289] In order to obtain a data set of the same style, that is, to construct a sample style reference image set, a clustering algorithm is used to cluster the sample style reference images that meet the first screening condition to obtain several cluster clusters; based on the several cluster clusters, a sample style reference image set is constructed, and each of the several cluster clusters corresponds to a sample style reference image set.
[0290] In some embodiments, a K-means clustering algorithm is used to cluster sample style reference images that meet the first screening condition to obtain a number of cluster clusters. K-means clustering is a distance-based clustering algorithm that uses distance as a similarity evaluation index. It is believed that the closer the distance between two sample style reference images x and y, the greater the similarity, which is expressed as:
[0291]
[0292] The basic idea of K-means clustering is that, when the number of clusters K is specified, K is greater than 0, and K sample style reference images are randomly selected from the sample style reference images that meet the first screening condition as the starting cluster center points, and the Euclidean distance between the points represented by other sample style reference images and the initial cluster center points is calculated, and the sample style reference images are classified into the class closest to the cluster center. After all sample style reference images that meet the first screening condition are divided into categories, K clusters are formed, and the mean of the sample style reference images in each cluster is recalculated, and the mean is used as the new cluster center. The cluster center is in a state of change, and the process is repeated until the cluster center point no longer changes.
[0293] Among them, the selection of K value significantly affects the accuracy of dividing the sample style reference image set. In this embodiment, K=10 is initially set, and the K-means clustering algorithm is executed. The style consistency in the clustering results is manually sampled and checked. If the style consistency is met, the cluster is deleted and the K value is reduced by 1. If the style consistency is not met, the K value is increased by 1. The above scheme is repeated until all sample style reference images that meet the first screening condition are successfully classified.
[0294] In this embodiment, a method of constructing a sample style reference image set by clustering is provided, which is conducive to improving data processing efficiency.
[0295] In some embodiments, each sample image pair includes a sample identification style image and a sample identification image corresponding to the sample identification style image. The sample image pair is constructed by:
[0296] Method 1: Use the existing generation capability of the Stable Diffusion XL (SDXL) model to input the logo text into the stable diffusion model to generate a sample logo style image corresponding to the logo text. The logo text includes logo information, which includes at least one of the nickname, name, nickname and name of the character in the film and television drama in which the character has participated. Whether the logo text includes style is not limited. The sample logo style image is an image with the logo information in the logo text and a certain style.
[0297] Method 2: First, based on a set of sample style reference images, train an optimized style model, use the optimized style model to generate a sample identification style image, and then use the de-stylization model to remove the sample style in the sample identification style image to obtain a sample identification image that only retains the sample identification information. In some embodiments, the style model is optimized by low-rank adaptation (LORA) fine-tuning, and the optimized style model is also called a LORA style model. The de-stylization model can be at least one of a diffusion model, a background removal model, and an image processing model. Specifically, step 540 is implemented as steps 542, 544, 546, and 548:
[0298] Step 542, constructing a plurality of logo style texts based on the second text template;
[0299] Step 544, inputting a plurality of logo style texts into the optimized style model, and generating a plurality of sample logo style images corresponding to each of the plurality of logo style texts;
[0300] Step 546, removing the sample style in each sample identification style image of the plurality of sample identification style images, to obtain a sample identification image corresponding to each sample identification style image;
[0301] Step 548, constructing a sample image pair based on each sample identification style image and the sample identification image corresponding to each sample identification style image; wherein the second text template includes: a first identification placeholder, a second style placeholder and an image quality descriptor, the first identification placeholder is used to fill in the first identification information, the second style placeholder is used to fill in the style, and the image quality descriptor is used to fill in the image quality description.
[0302] The second text template is a template for constructing a logo style text, and the logo style text is used to describe a style and a logo information. The second text template includes: a first logo placeholder (Token1), a second style placeholder (Token2), and an image quality descriptor. The first logo placeholder is used to fill in the first logo information, the second logo placeholder is used to fill in the style, and the image quality descriptor is used to fill in the image quality description.
[0303] For example, the first identification information is gender, which is at least one of the following: {"male / boy", "female / girl"}, and the style can be a style word, which is at least one of the following: 3D RealCartoon, 3D Cartoon, 2D Anime, 2D Comic, Oneline-Drawing, Black Magic, Ink, Ice Scupture. The image quality description can be a description word or a description text, and the image quality description includes at least one of the following: masterpiece, best quality, crop top, perfectly symmetrical face, detailed skin, amazing photograph, 8K, high quality, photorealistic, and realism.
[0304] Exemplarily, based on the second text template, several logo style texts are constructed. Several logo style texts are input into the optimized style model, and the noise random number of the optimized style model initialization input is changed to generate several sample logo style images corresponding to each logo style text of the several logo style texts, thereby realizing the generation of sample logo style images with multiple different logo information. Based on the de-stylization model, the sample style in each sample logo style image of several sample logo style images is removed to obtain the sample logo image corresponding to each sample logo style image; based on each sample logo style image and the sample logo image corresponding to each sample logo style image, a sample image pair is constructed.
[0305] Each sample image pair includes a sample identification image and a sample identification style image corresponding to the sample identification image. Different sample image pairs may have intersections in their respective sample identification images, and may also have intersections in the sample styles indicated by their respective sample identification style images. For example, sample image pair 1 includes: an image of person 1 and an image of person 1 in a fairy-tale style; sample image pair 2 includes: an image of person 2 and an image of person 2 in a fairy-tale style; sample image pair 3 includes: an image of person 1 and an image of person 1 in a dark style.
[0306] In this embodiment, based on the style image capability of the optimized style model, a sample identification style image can be constructed, which is conducive to quickly constructing training data and improving the efficiency of constructing training data.
[0307] In some embodiments, in order to ensure the quality of the sample image pair and improve the consistency of the sample identification information between the sample identification style image in the sample image pair and the corresponding sample identification image, the constructed sample image pair is further screened to obtain a screened sample image pair. Subsequently, training data is constructed based on the screened sample image pair. Specifically, after step 540, step 550 is also included:
[0308] Step 550 , screening the sample image pairs to determine sample image pairs that meet a second screening condition; wherein the second screening condition includes: a third similarity between the sample identification style image and the sample identification image in the sample image pair is greater than a third threshold.
[0309] Exemplarily, screening is performed based on an image recognition model: the image recognition model is used to calculate facial similarity. For example, taking a person as an example, the image recognition model is a face recognition model (ArcFace) model, which is used to calculate the third similarity between the sample identification style image and the sample identification image in the sample image pair.
[0310] Specifically, the face recognition model proposes a face recognition loss function of angular cosine distance, which improves the performance of the face recognition model by enhancing the distinguishability of facial features in the feature space. The core idea is to project facial features onto the unit hypersphere, making the feature vector more compact in angle. Using the image recognition model, the sample identification style image X and the sample identification image Y in the sample image pair are encoded into 512-dimensional vectors respectively, and the third similarity is calculated using cosine similarity, which is expressed as:
[0311]
[0312] The cosine value range is [-1, 1]. After linear transformation, the value range is [0, 1]. The larger the cosine value, the greater the probability that the sample identification style image and the sample identification image in the sample image pair correspond to the same sample identification information, and the sample image pair whose third similarity is greater than the third threshold is regarded as the sample image pair that meets the second screening condition. For example, the third threshold can be set to: 0.8.
[0313] The method of this embodiment can ensure the consistency of identification information of sample image pairs used in constructing training data, thereby improving the data accuracy and quality of the training data.
[0314] In some embodiments, the style model is optimized by low-rank adaptation (LORA) fine-tuning, and the optimized style model is also called a LORA style model.
[0315] Low-rank adaptive fine-tuning achieves the purpose of introducing new concepts by making small modifications to the original model. Specifically, the main modification of low-rank adaptive fine-tuning is the attention network of the image segmentation model (UNet) in the diffusion model. The calculation formula of the attention network is:
[0316]
[0317] Among them, the K, Q, and V matrices are used in the calculation process of the attention network to calculate the mutual dependence between different tokens. Q and K can be used to calculate the similarity between the current token and other tokens. The similarity is used as a weight to perform weighted summation on V, and the result can be used as the token of the next layer. In the attention network used by the image segmentation model (UNet), K and V come from text, and Q comes from the image. The attention network can calculate the encoding result of the image under the text condition through attention operation. Accordingly, this embodiment adds trainable parameters to the original style model. After optimization, the style model can stably generate sample identification style images corresponding to the identification style text according to the identification style text.
[0318] Specifically, after step 520 and before step 540, the method further includes step 530:
[0319] Step 530, based on the sample style reference image in the sample style reference image set, optimize the trainable parameters in the style model to obtain an optimized style model; wherein the optimized style model is used to generate the sample identification style image, and the trainable parameters include adding at least two bypass matrices in the attention network of the style model.
[0320] The style model can adopt one of the diffusion model and the stable diffusion model. Based on the low-rank adaptive fine-tuning method, trainable parameters are added to the style model. Specifically, trainable parameters are added to the linear projection part of the QKV of the attention network of the style model and to the linear part of the feedforward neural network (FFN). These incremental trainable parameters can be transformed into fewer trainable parameters through matrix decomposition, greatly reducing the number of parameters required to optimize the fine-tuning style model.
[0321] After adding trainable parameters to the style model, the calculation formula is:
[0322] h=W0x+ΔWx=W0x+BAx
[0323] Among them, x represents input, h represents output, W0 represents the original parameters of the style model, and ΔW represents the trainable parameters of the inserted layer. In order to minimize the number of parameters of the inserted layer and improve the optimization efficiency, ΔW is decomposed into the product of at least two matrices A and B, called the bypass matrix. Fig.13 As shown in the figure. When ΔW is a d×d dimensional matrix, the size of the A and B matrices can be reduced to d×r, where r can be much smaller than d. In theory, the smaller the rank r of the multiplication between the B and A matrices, the smaller the number of parameters in the inserted layer.
[0324] The trainable parameters are decomposed into at least two bypass matrices, and the trainable parameters include at least two bypass matrices added to the attention network of the style model. Next, based on the sample style reference image in the sample style reference image set, the trainable parameters in the style model are optimized to obtain an optimized style model. The optimized style model is used to generate a sample identification style image.
[0325] As an example, Fig.14 It is a schematic diagram of a style image generation method provided by an exemplary embodiment of the present application. Taking the style model as a diffusion model, specifically an image segmentation model (UNet) in the diffusion model as an example, a LORA module is mounted on the style model, and the LORA module includes trainable parameters. First, based on the sample style reference image in the sample style reference image set, the trainable parameters in the style model are optimized to obtain an optimized style model; a number of identification style texts (Text) are constructed, and a number of identification style texts are input into the optimized style model to generate a number of sample identification style images corresponding to each identification style text.
[0326] This embodiment provides a LORA-based fine-tuning method to optimize the style model to obtain an optimized style model, which can stably generate sample identification style images, thereby facilitating the construction of training data.
[0327] Next, the style image generation method provided in the embodiment of the present application is generally described.
[0328] In the following embodiments, taking a person as an example, the sample identification image is described as a real person image, the sample identification information is described as person identification information, the sample style reference image is described as a style image, the sample identification style image is described as a person style image, the optimized style model is described as a LORA style model, and the sample image pair is described as a stylized image pair.
[0329] Application scenarios
[0330] The style image generation model provided in the embodiment of the present application is a model based on a multi-image segmentation model (Multi-UNet) that can maintain the consistency of character identification information and is stylized. It can achieve zero-shot learning capability and generate images that maintain the consistency of character identification information and are stylized.
[0331] The embodiment of the present application can be applied in applications such as chat applications, game applications, video applications, entertainment applications, travel applications, social applications, health applications, application download applications, etc. The application provides at least one function of character avatar conversion, character image generation, and image fusion of character images and style images. The style image generation model is a model for converting real character images to character style images with zero-sample learning capabilities, which can form stylized images with consistent character identification information, improve generation accuracy and generation efficiency, and also improve user experience.
[0332] Technical implementation
[0333] The style image generation model provided in the embodiment of the present application is based on the diffusion model, which improves the training data and model structure. The overall architecture can be referred to Figure 8 and Fig.10 And method embodiments thereof. Among them, for training data, it includes: collecting multi-style images; training LORA style model; generating character style images including specified character identification information and having different styles; constructing training data. For the model structure, on the basis of the diffusion model, multiple models with the same structure as the image segmentation model (UNet) in the diffusion model are added, which are respectively used to encode real character images, extract character identification features, and encode style images and extract style features; the UNet model is used to fuse character identification features and style features, and finally generate a stylized image that maintains consistency in character identification information. After optimizing the training data and model structure, a style image generation model is obtained, which can achieve zero-sample learning capability, that is, a stylized image that maintains consistency in character identification information is generated without further training.
[0334] 1. Training Data
[0335] This embodiment mainly includes: collecting multi-style images; training the LORA style model; generating character style images including specified character identification information and having different styles; and constructing training data.
[0336] (1) Collecting images in multiple styles
[0337] In order to cover as many styles as possible, we first collect a large number of style images, and then use clustering to cluster style images of the same style. Among them, there are several ways to collect a large number of style images: Method 1, use the text image capability of the Flux1 dev model to generate style images of a specified style; Method 2, use crawler technology to obtain images from public websites, which can be a text image model sharing platform, where users can upload and download images of different styles; Method 3, use an open source data set to obtain style images in the data set.
[0338] For method 1, refer to Fig.11 , 8 different styles are selected: 3D RealCartoon, 3D Cartoon, 2D Anime, 2D Comic, Oneline-Drawing, Black Magic, Ink, Ice Scupture. Use the Flux1 dev model's text image capability to generate style images with a resolution of 1024*1024 containing characters. The input text template is: "a Token1, Token2 years old, Token3 style, masterpiece, best quality, (crop top), perfectly symmetrical face, detailed skin, amazing photograph, masterpiece, best quality, 8K, high quality, photorealistic, realism". Among them, the placeholder Token1 can select gender {"male / boy", "female / girl"}, the placeholder Token2 can select age {5,10,20,30,40,50,60,70}, and the placeholder Token3 can select 8 different styles. Based on this text template, the noise random number of the initialization input of the Flux1 dev model is changed to generate images of multiple age groups, different genders, and different styles. After generating the style images, the style images with consistent style and high quality are manually selected, and only 80 style images (combination: 2 genders, 8 age groups, 5 style images) are retained for each style.
[0339] For method 2, the Wensheng Image Model Sharing Platform provides relevant APIs, allowing users to upload and download images of different styles. The Wensheng Image Model Sharing Platform contains a variety of self-trained style models, which can significantly improve style coverage. Select the style images generated by the Flux1 dev model, and obtain a total of 700+ models and 50,000+ style images.
[0340] For method 3, the open source LAION-5B dataset is used, which contains multi-style image-text pairs. The CLIP model is used to evaluate the similarity of the image and text in each image-text pair, and the image-text pairs with similarity below the set threshold are deleted, where the English threshold is set to 0.28 and the remaining thresholds are set to 0.26. The dataset finally retains 5.85 billion image-text pairs, including 2.32 billion in English, 2.26 billion in 100+ languages, and 1.27 billion in unknown languages. Post-processing can also be performed on the 2.32 billion English data.
[0341] Based on the above three methods, we collect multi-style images to significantly improve the style coverage of the training data. Since the quality of these style images varies, there may be non-person style images, or there may be style images with face size, face distortion, and blur, so screening is also required.
[0342] Method 1, face recognition. Keep style images with a face area larger than 256*256. Use an image recognition model, such as the RetinaFace model, to achieve face recognition. It is a neural network model for face detection and face alignment. It is based on a one-stage target detector and uses multi-scale feature fusion and multi-task learning methods to accurately detect and align faces at different scales and directions. The RetinaFace model has good robustness in multi-scale, multi-directional and occlusion situations, and is suitable for real-time face detection and related applications. Reference to the RetinaFace model face recognition process Fig.12 As shown, a style image containing a person is input to the RetinaFace model, and the coordinates (x, y, w, h) of the rectangle containing the face are returned, where (x, y) is the coordinate of the upper left corner of the rectangle, w and h are the width and height of the rectangle respectively, and only style images with a rectangle size greater than the size threshold are retained.
[0343] Method 2: Use the CLIP model to vectorize images. The CLIP model can map images and text into a shared vector space and understand the semantic relationship between images and text. This shared vector space enables the CLIP model to achieve unsupervised joint learning between images and text, and can be used for various visual and language tasks. Based on the principle of the CLIP model, text or images with the same semantics have a high degree of similarity. The image-text pair is input into the CLIP model, and only the style images in the image-text pair whose similarity to the text is greater than the threshold are retained.
[0344] In order to obtain the same style data set, the K-Means clustering algorithm is used to perform clustering on the filtered style images. K-means clustering is a typical distance-based clustering algorithm, which uses distance as the evaluation index of similarity. It believes that the closer the distance between two objects x and y, the greater the similarity, which is expressed as:
[0345]
[0346] The basic idea of K-means clustering is to randomly select K style images from the data set as the starting cluster center points when the number of clusters K is specified, K is greater than 0, and the Euclidean distance between the points represented by other style images and the initial cluster center points is calculated. The formula is as above. The style images are classified into the class closest to the cluster center. After all style images are divided into categories, K clusters are formed. The mean of the style images in each cluster is recalculated and the mean is used as the new cluster center. The cluster center is in a state of change, and the process is repeated until the cluster center point no longer changes.
[0347] The choice of K value significantly affects the accuracy of dividing the style data set. In this embodiment, K=10 is initially set, and the K-means clustering algorithm is executed. The style consistency in the clustering results is manually sampled and checked. If the style consistency is met, the cluster is deleted and the K value is reduced by 1. If the style consistency is not met, the K value is increased by 1. The above scheme is repeated until all style images are successfully classified.
[0348] (2) Training LOAR style model
[0349] Using the above style image, the diffusion model or the stable diffusion model (Stable Diffusion XL, SDXL) is used as the style model. Based on the LORA fine-tuning method, trainable parameters are added to the style model to train the LORA style model. Among them, LORA fine-tuning achieves the purpose of introducing new concepts by making small modifications to the original model. Specifically, the modified location is the attention network in the image segmentation model (UNet) in the diffusion model. The calculation formula of the attention network is:
[0350]
[0351] Among them, the K, Q, and V matrices are used in the calculation process of the attention network to calculate the mutual dependencies between different tokens. Q and K can be used to calculate the similarity between the current token and other tokens. The similarity is used as a weight to perform a weighted sum on V, and the result can be used as the token of the next layer. In the attention network used in the image segmentation model (UNet), K and V come from text, and Q comes from the image. The attention network can calculate the encoding results of the image under the condition of text. The attention network can calculate the encoding results of the image under the condition of text through attention operations.
[0352] More specifically, LORA adds trainable parameters to the attention layer of the attention network, specifically the linear projection part of the QKV of the Cross-Attention layer, and the linear part of the subsequent Feedforward Networks (FFN). These incremental trainable parameters can be transformed into fewer trainable parameters through matrix decomposition, greatly reducing the number of parameters that need to be optimized for fine-tuning the style model. Fig.13 , the corresponding calculation formula is as follows:
[0353] h=W0x+ΔWx=W0x+BAx
[0354] Among them, x represents input, h represents output, and W o Represents the original parameters of the style model, and ΔW represents the trainable parameters of the inserted layer. In order to minimize the number of parameters of the inserted layer and improve the optimization efficiency, ΔW is decomposed into the product of two matrices A and B. When ΔW is a d×d dimensional matrix, the size of the A and B matrices can be reduced to d×r, where r can be much smaller than d. In theory, the smaller the rank r of the multiplication between the B and A matrices, the smaller the number of parameters of the inserted layer. After training the LORA style model, character style images of the corresponding style can be stably generated.
[0355] (3) Generate a character style image with specified character identification information and different styles
[0356] This embodiment mainly includes: constructing stylized image pairs; screening the stylized image pairs, and only retaining stylized image pairs whose face similarity between the real person image and the person style image is greater than a threshold.
[0357] There are several ways to construct stylized image pairs: Method 1, use the existing generation capabilities of the SDXL model to generate styled character images with specified character identification information; Method 2, in order to further increase the number of styled character images and real character images containing different character identification information, first randomly generate character style images, and then use the de-stylization model to generate real character images with the same character identification information.
[0358] For method 1, the training data of the SDXL model is a public text-image pair, which already contains a large number of public person identification information, and can be directly used to generate a large number of stylized person style images that include the specified person identification information. The input text template is: "a Token1,{Token2 style,}masterpiece,best quality,(croptop),perfectly symmetrical face,detailed skin,amazing photograph,masterpiece,best quality,8K,high quality,photorealistic,realism". Among them, the placeholder Token1 is a required option, which is filled with the nickname, name, nickname and name of the film and television drama in which the character has participated. The placeholder Token2 is an optional option, and multiple different styles can be selected. When the specified style does not exist, the real person image of the character is generated. Based on the text template, the noise random number of the SDXL model initialization input is changed to generate person style images with different styles containing the specified person identification information, and construct stylized image pairs.
[0359] For method 2, in order to further increase the number of character style images and real character images containing different character identification information, character style images are first randomly generated, and then the de-stylization model is used to generate real character images with the same character identification information. Combined with the LORA style model, the input text template is "a Token1, Token2 style, masterpiece, best quality, (crop top), perfectly symmetrical face, detailed skin, amazing photograph, masterpiece, best quality, 8K, high quality, photorealistic, realism". Among them, the placeholder Token1 is a required option, and the gender is randomly filled, and the placeholder Token2 can choose a variety of different styles. Based on the text template, the noise random number initialized by the LORA style model is changed to generate character style images with a variety of different character identification information. Based on the de-stylization model, the style in the character style image is removed, and the real character image with the same character identification information is generated to construct a stylized image pair.
[0360] In order to further improve the consistency of the person identification information in the stylized image pairs, an image recognition model, such as the ArcFace model, is used to calculate the facial similarity between the person style image in the stylized image pair and the real person image, and only the stylized image pairs with facial similarity higher than the threshold are retained. For example, the threshold is set to 0.8, and the similarity is expressed as:
[0361]
[0362] Among them, ArcFace is a face recognition model based on a deep convolutional neural network architecture. It proposes a face recognition loss function of angular cosine distance, which improves the model performance by enhancing the distinguishability of facial features in the feature space. The core idea is to project facial features onto the unit hypersphere, making the feature vector more compact in angle. The ArcFace model is used to encode the character style image and the real character image in the stylized image pair into a 512-dimensional vector. The cosine similarity is used to calculate the face similarity. The cosine value range is [-1,1]. After linear transformation, the cosine value range is [0,1]. The larger the cosine value, the greater the probability that the two images represent the same character identification information.
[0363] (4) Constructing training data
[0364] Each training data consists of three images: a real person image and two person style images of the same style corresponding to different persons. Fig. 9As shown. One of the character style images contains the same character identification information as the real character image, and the other character style image contains different character identification information from the real character image. Optionally, a random sampling strategy can be used on the same style data set to randomly select these two character style images to construct training data.
[0365] 2. Model Structure
[0366] The style image generation model is based on the diffusion model. The image segmentation model (UNet) in the diffusion model is responsible for image generation, and the input is the noisy feature X t , whose function is to predict noise and generate denoising feature X after removing the prediction noise t-1 After t steps of iterative processing, the noisy feature X t Converted to denoised feature X0. The smaller t is, the less predicted noise is. When t=1, the UNet model can obtain nearly noise-free data. After this denoising step, the denoising result is denoised feature X0. After the denoised feature X0 is decoded by the decoder, a meaningful image can be generated. Based on the above principle, when t=1, the UNet model can be used as an image encoder to encode image features.
[0367] Based on this, in order to extract character identification features and style features and generate stylized character style images that maintain consistency of character identification information, it is proposed to use multiple models (t=1) with the same structure as the image segmentation model (UNet) to encode the character identification features of real character images and the style features of character style images. Among them, the attention network of the UNet model is responsible for fusing image-related features. Therefore, the character identification features and style features are spliced to the same position of the UNet model. In order to minimize the impact of the original model's capabilities, only the parameters of the attention network of the UNet model are optimized.
[0368] Multiple UNet models have the same semantic space, which can achieve fast convergence with a small amount of training data, and generate character style images that fuse styles and maintain character identification information. Model structure reference Figure 8 , Fig.10 . In order to introduce character identification features and style features, a multi-image segmentation model (Multi-UNet) is proposed to encode real character images and style images. In order to reduce the impact of the newly introduced features on the original UNet model and maintain the same semantic space, the structure and parameters of Multi-UNet are consistent with UNet. Specifically, Multi-UNet includes: an image segmentation model (UNet) for generating character style images, a character image segmentation model (PersonNet) for processing real character images, and a style image segmentation model (StyleNet) for processing style images.
[0369] UNet mainly includes two structures: ResNet and Transformer. ResNet is composed of convolutional networks and is responsible for encoding image features. Transformer is composed of attention networks, and the calculation formula is as follows:
[0370]
[0371] Among them, the attention network is responsible for encoding image features and introducing conditional features, such as text features, image features, and audio features. The attention network in Transformer mainly includes two forms, including self-attention and cross-attention. When Q, K, and V are the same, it is specified as Self-Attention, which is responsible for calculating the relationship between image features; when Q is an image feature and K and V are other conditional features, it is specified as Cross-Attention, which is responsible for calculating the relationship between image features and other features. Studies have shown that the amount of Attention parameters accounts for a small proportion of the entire model, but plays an important role. Taking LORA as an example, only the linear projection part of QKV in the Cross-Attention layer and the linear part of the subsequent Feedforward Networks (FFN) are used. The amount of fine-tuning parameters is much smaller than all the model parameters of the UNet model, but the fine-tuning effect is similar to the effect of fine-tuning the model parameters of the entire UNet model.
[0372] Based on this, this embodiment transfers the Self-Attention features of the attention network of PersonNet and StyleNet to UNet. Since the role of Multi-UNet is to encode the character identification features and style features, helping UNet to restore the character and style, only image features are involved, and the Self-Attention features encoded by PersonNet and StyleNet are transferred to the Self-Attention in UNet in the form of vector splicing. For splicing effect, refer to Figure 7 As shown, the calculation formula is as follows:
[0373] y=Attention(Q unet , cat(K unet , K style , K person ), cat(V unet , V style , V person ))
[0374] Among them, Q unet , K unet and Vunet Represents the Self-Attetnion input of UNet, K person and V person For the Self-Attetnion input of the spliced PersonNet, K style and V style It is the Self-Attemion input of the spliced StyleNet. The UNet model is used to fuse the character identification features encoded by PersonNet and the style features encoded by StyleNet, and finally generate a high-quality character style image that maintains the consistency of the character identification information and is stylized.
[0375] 3. Effect description
[0376] Based on the diffusion model, the embodiment of the present application provides a model based on Multi-UNet that maintains the consistency of character identification information and is stylized, which can achieve zero-sample learning ability and generate character style images that maintain the consistency of character identification information and are stylized. First, collect style images and real character images to train the LORA style model. Under the action of the LORA style model, generate stylized image pairs and construct training data. Based on the Multi-UNet structure and training data, a style image generation model is obtained to achieve the ability to integrate character identification information and style. Figure 4 ,In the inference stage, any real person image can be converted into a stylized person style image that maintains the consistency of the person identification information.
[0377] Fig.15 800 is a block diagram of a style image generation device provided by an exemplary embodiment of the present application. The style image generation device 800 includes: at least part of the acquisition module 820, the encoding module 840, the processing module 860, and the decoding module 880.
[0378] The acquisition module 820 is used to acquire a logo image, at least one style reference image, and Gaussian noise, wherein the logo image is an image including designated logo information and not having a designated style, the at least one style reference image is an image having the designated style and not including the designated logo information, and the Gaussian noise is used to simulate the characteristic distribution of the logo style image after noise addition, and the logo style image is an image including the designated logo information and having the designated style;
[0379] The encoding module 840 is used to encode the identification feature of the identification image; encode the style feature of the at least one style reference image; and encode the noise feature of the Gaussian noise;
[0380] The processing module 860 is used to fuse the noise feature, the identification feature and the style feature to obtain a fused feature; and perform denoising on the fused feature to obtain a denoised feature;
[0381] The decoding module 880 is used to decode the denoising feature and generate the logo style image.
[0382] In some embodiments, the processing module 860 is configured to:
[0383] Transferring the identification feature and the style feature to the noise feature to obtain a splicing feature;
[0384] Mapping the noise feature and the splicing feature based on an attention mechanism to obtain an attention feature;
[0385] A convolution operation is performed based on the attention feature to obtain the fusion feature.
[0386] In some embodiments, the processing module 860 is configured to:
[0387] The noise feature, the identification feature and the features located at the same position in the style feature are spliced to obtain the spliced feature.
[0388] In some embodiments, the splicing feature includes a splicing K matrix and a splicing V matrix, the noise feature includes a noise K matrix and a noise V matrix, the identification feature includes an identification K matrix and an identification V matrix, and the style feature includes a style K matrix and a style V matrix; the processing module 860 is used to:
[0389] The noise K matrix, the identification K matrix and the style K matrix are concatenated to obtain the concatenated K features; and the noise V matrix, the identification V matrix and the style V matrix are concatenated to obtain the concatenated V matrix.
[0390] In some embodiments, the noise feature further includes a noise Q matrix, and the splicing feature includes a splicing K matrix and a splicing V matrix; the processing module 860 is used to:
[0391] The noise Q matrix, the splicing K matrix and the splicing V matrix are mapped based on the attention mechanism to obtain the attention feature.
[0392] In some embodiments, the processing module 860 is configured to:
[0393] Perform noise prediction based on the fusion feature to obtain predicted noise;
[0394] De-noising the fused feature according to the predicted noise to obtain the de-noised feature.
[0395] In some embodiments, the encoding module 840 is configured to:
[0396] Mapping the identification image to obtain identification mapping features;
[0397] Performing noise addition on the identification mapping feature to obtain a noisy identification mapping feature;
[0398] The noise-added identification mapping feature is encoded to obtain the identification feature.
[0399] In some embodiments, the encoding module 840 is configured to:
[0400] Mapping the at least one style reference image to obtain a style mapping feature;
[0401] Adding noise to the style map feature to obtain a noisy style map feature;
[0402] The noise-added style map feature is encoded to obtain the style feature.
[0403] In some embodiments, a style image generation model for executing the style image generation method includes: an encoder, a feature network, and a decoder cascaded in sequence;
[0404] The encoder and the feature network are used to encode the identification features of the identification image, the style features of the at least one style reference image, and the noise features of the Gaussian noise; the feature network is also used to fuse the noise features, the identification features and the style features to obtain fused features, and perform denoising on the fused features to obtain denoised features; the decoder is used to decode the denoised features to generate the identification style image.
[0405] In some embodiments, the encoder includes a first encoder, a second encoder, and a third encoder connected in parallel, the feature network includes a first feature network, a second feature network, and a third feature network connected in parallel, the first encoder is also connected to the first feature network, the second encoder is also connected to the second feature network, the third encoder is also connected to the third feature network, the first feature network and the third feature network are also respectively connected to the second feature network, and the second feature network is also connected to the decoder;
[0406] Among them, the first encoder is used to map the identification image to obtain the identification mapping feature, and the first feature network is used to encode the noisy identification mapping feature to obtain the identification feature; the third encoder is used to map the at least one style reference image to obtain the style mapping feature, and the third feature network is used to encode the noisy style mapping feature to obtain the style feature; the second feature network is used to encode the noise feature of the Gaussian noise, fuse the noise feature, the identification feature and the style feature to obtain a fused feature, and perform denoising on the fused feature to obtain a denoised feature.
[0407] In some embodiments, the invention further includes: an acquisition module; the acquisition module is used to:
[0408] Acquire training data, each of the training data includes a sample identification image, a sample style reference image, and a sample identification style image, wherein the sample identification image is an image including sample identification information and not having the sample style, the sample style reference image is an image having the sample style and not having the sample identification information, and the sample identification style image is an image having the sample identification information and having the sample style;
[0409] In some embodiments, the encoding module 840 is configured to:
[0410] Encoding the sample identification features of the sample identification image; encoding the sample identification style features of the sample identification style image; and encoding the sample style features of the sample style reference image;
[0411] In some embodiments, the processing module 860 is configured to:
[0412] Fusion of the sample identification feature, the sample identification style feature and the sample style feature to obtain a sample fusion feature; and denoising of the sample fusion feature to obtain a sample denoising feature;
[0413] In some embodiments, the decoding module 880 is configured to:
[0414] Decoding the sample denoising features to obtain a predicted signature style image;
[0415] In some embodiments, the invention further comprises: a training module; a training module for:
[0416] The model parameters of the style image generation model are adjusted by taking reducing the difference between the predicted identification style image and the sample identification style image as a training goal.
[0417] In some embodiments, the processing module 860 is configured to:
[0418] The sample identification feature and the sample style feature are transferred to the sample identification style feature to obtain a sample splicing feature;
[0419] Mapping the sample identification style feature and the sample splicing feature based on the attention mechanism to obtain a sample attention feature;
[0420] A convolution operation is performed based on the sample attention feature to obtain the sample fusion feature.
[0421] In some embodiments, the processing module 860 is configured to:
[0422] The sample identification style feature, the sample identification feature and the sample features located at the same position in the sample style feature are spliced together to obtain the sample splicing feature.
[0423] In some embodiments, the sample splicing feature includes a sample splicing K feature and a sample splicing V feature, the sample identification style feature includes a sample identification style K matrix and a sample identification style V matrix, the sample identification feature includes a sample identification K matrix and a sample identification V matrix, and the sample style feature includes a sample style K matrix and a sample style V matrix; the processing module 860 is used to:
[0424] The sample identification style K matrix, the sample identification K matrix and the sample style K matrix are concatenated to obtain the sample concatenation K feature; and the sample identification style V matrix, the sample identification V matrix and the sample style V matrix are concatenated to obtain the sample concatenation V matrix.
[0425] In some embodiments, the sample identification style feature further includes a sample identification style Q matrix, and the sample splicing feature includes a sample splicing K feature and a sample splicing V feature; the processing module 860 is used to:
[0426] The sample identification style Q matrix, the sample splicing K matrix and the sample splicing V matrix are mapped based on the attention mechanism to obtain the sample attention feature.
[0427] In some embodiments, the processing module 860 is configured to:
[0428] Perform noise prediction based on the sample fusion features to obtain predicted noise;
[0429] De-noising the sample fusion feature according to the predicted noise to obtain the sample denoised feature.
[0430] In some embodiments, the encoding module 840 is configured to:
[0431] Mapping the sample identification image to obtain a sample identification mapping feature;
[0432] Performing noise addition on the sample identification mapping feature to obtain a noisy sample identification mapping feature;
[0433] The noisy sample identification mapping feature is encoded to obtain the sample identification feature.
[0434] In some embodiments, the encoding module 840 is configured to:
[0435] Mapping the sample style reference image to obtain a sample style mapping feature;
[0436] Adding noise to the sample style mapping feature to obtain a noisy sample style mapping feature;
[0437] The noisy sample style mapping feature is encoded to obtain the sample style feature.
[0438] In some embodiments, the encoding module 840 is configured to:
[0439] Mapping the sample identification style image to obtain a sample identification style mapping feature;
[0440] Performing noise addition on the sample identification style mapping feature to obtain a noisy sample identification style mapping feature;
[0441] The noisy sample identification style mapping feature is encoded to obtain the sample identification style feature.
[0442] In some embodiments, the invention further comprises: a building block; a building block for:
[0443] Constructing a sample style reference image set, each of the sample style reference image sets including a plurality of sample style reference images having the same sample style;
[0444] Constructing sample image pairs, each of the sample image pairs comprising a sample identification style image and a sample identification image corresponding to the sample identification style image;
[0445] The training data is constructed according to the sample image pairs and the sample style reference image set.
[0446] In some embodiments, a building block is provided for:
[0447] Based on the first text template, construct a plurality of style texts;
[0448] Inputting the plurality of style texts into a text image generation model to generate a plurality of sample style reference images corresponding to each of the plurality of style texts;
[0449] Constructing the sample style reference image set based on the sample style reference images corresponding to each style text of the plurality of style texts;
[0450] Among them, the first text template includes: a first identifier placeholder, a second identifier placeholder, a third style placeholder and an image quality descriptor, the first identifier placeholder is used to fill in the first identifier information, the second identifier placeholder is used to fill in the second identifier information, the third style placeholder is used to fill in the style, and the image quality descriptor is used to fill in the image quality description.
[0451] In some embodiments, a building block is provided for:
[0452] Determining a sample style reference image that meets a first screening condition from a plurality of sample style reference images corresponding to each of the plurality of style texts;
[0453] Clustering the sample style reference images that meet the first screening condition to obtain a plurality of clusters;
[0454] Based on the plurality of clusters, constructing the sample style reference image set;
[0455] Among them, the first screening condition includes: the first similarity between the style text and the corresponding sample style reference image is greater than a first threshold, or the second similarity between different sample style reference images corresponding to the style text is greater than a second threshold; each of the plurality of clusters corresponds to a set of sample style reference images.
[0456] In some embodiments, a building block is provided for:
[0457] Based on the second text template, construct a plurality of logo style texts;
[0458] Inputting the plurality of logo style texts into the optimized style model to generate a plurality of sample logo style images corresponding to each of the plurality of logo style texts;
[0459] Removing the sample style from each sample identification style image of the plurality of sample identification style images to obtain the sample identification image corresponding to each sample identification style image;
[0460] Constructing the sample image pair based on each sample identification style image and the sample identification image corresponding to each sample identification style image;
[0461] The second text template includes: a first identification placeholder, a second style placeholder and an image quality descriptor, the first identification placeholder is used to fill in the first identification information, the second style placeholder is used to fill in the style, and the image quality descriptor is used to fill in the image quality description.
[0462] In some embodiments, a building block is provided for:
[0463] Screening the sample image pairs to determine sample image pairs that meet a second screening condition;
[0464] The second screening condition includes: a third similarity between the sample identification style image and the sample identification image in the sample image pair is greater than a third threshold.
[0465] In some embodiments, a building block is provided for:
[0466] Based on the sample style reference images in the sample style reference image set, optimizing the trainable parameters in the style model to obtain the optimized style model;
[0467] The optimized style model is used to generate the sample identification style image, and the trainable parameters include adding at least two bypass matrices in the attention network of the style model.
[0468] It should be noted that the specific limitations in the embodiments of the one or more style image generation devices 800 provided above can refer to the limitations of the style image generation method above, and will not be repeated here. Each module of the above device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor of the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each module.
[0469] An embodiment of the present application further provides a computer device, which includes: a processor and a memory, wherein a computer program is stored in the memory; and the processor is configured to execute the computer program in the memory to implement the style image generation method provided by the above-mentioned method embodiments.
[0470] For example, Fig.16 is a block diagram of a computer device provided by an exemplary embodiment of the present application. Optionally, the computer device is Figure 1 The server or terminal, in this embodiment, is described by taking the computer device being the server 1000 as an example.
[0471] Typically, the server 1000 includes: a processor 1001 and a memory 1002 .
[0472] The processor 1001 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1001 may be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1001 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.
[0473] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1002 is used to store at least one instruction, which is used to be executed by the processor 1001 to implement the style image generation method provided by the above-mentioned method embodiments.
[0474] In some embodiments, the server 1000 may also optionally include: an input interface 1003 and an output interface 1004. The processor 1001, the memory 1002, the input interface 1003, and the output interface 1004 may be connected via a bus or a signal line. Each peripheral device may be connected to the input interface 1003 and the output interface 1004 via a bus, a signal line, or a circuit board. The input interface 1003 and the output interface 1004 may be used to connect at least one peripheral device related to input / output (I / O) to the processor 1001 and the memory 1002. In some embodiments, the processor 1001, the memory 1002, the input interface 1003, and the output interface 1004 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1001, the memory 1002, the input interface 1003, and the output interface 1004 may be implemented on a separate chip or circuit board, which is not limited in the embodiments of the present application.
[0475] Those skilled in the art will understand that Fig.16 The structure shown in the figure does not constitute a limitation on the computer device, and may include more or less components than shown in the figure, or combine some components, or adopt a different component arrangement.
[0476] In an exemplary embodiment, the embodiment of the present application further provides a chip, which includes a programmable logic circuit and / or computer instructions. When the chip runs on a computer device, it is used to implement the style image generation method provided by the above-mentioned method embodiments.
[0477] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the style image generation method provided by the above-mentioned method embodiments.
[0478] The present application embodiment also provides a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the processor of the computer device loads and executes to implement the style image generation method provided by the above-mentioned method embodiments.
[0479] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0480] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned computer-readable storage medium may be a read-only memory, a disk or an optical disk, etc.
[0481] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented with hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein the communication media include any media that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium that a general or special-purpose computer can access.
[0482] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A style image generation method, characterized in that: The method comprises: Acquire a logo image, at least one style reference image, and Gaussian noise, wherein the logo image is an image including designated logo information and not having a designated style, the at least one style reference image is an image having the designated style and not including the designated logo information, and the Gaussian noise is used to simulate the characteristic distribution of the logo style image after noise addition, and the logo style image is an image including the designated logo information and having the designated style; encoding identification features of the identification image; encoding style features of the at least one style reference image; and encoding noise features of the Gaussian noise; fusing the noise feature, the identification feature and the style feature to obtain a fused feature; and performing denoising on the fused feature to obtain a denoised feature; The denoising features are decoded to generate the logo style image.
2. The method according to claim 1, characterized in that: The fusing the noise feature, the identification feature and the style feature to obtain a fused feature includes: Transferring the identification feature and the style feature to the noise feature to obtain a splicing feature; Mapping the noise feature and the splicing feature based on an attention mechanism to obtain an attention feature; A convolution operation is performed based on the attention feature to obtain the fusion feature.
3. The method according to claim 2, characterized in that The step of transferring the identification feature and the style feature to the noise feature to obtain a splicing feature includes: The noise feature, the identification feature and the features located at the same position in the style feature are spliced to obtain the spliced feature.
4. The method according to claim 3, characterized in that The splicing features include a splicing K matrix and a splicing V matrix, the noise features include a noise K matrix and a noise V matrix, the identification features include an identification K matrix and an identification V matrix, and the style features include a style K matrix and a style V matrix; The step of splicing the noise feature, the identification feature, and the features located at the same position in the style feature to obtain the spliced feature includes: Concatenating the noise K matrix, the identification K matrix, and the style K matrix to obtain the concatenated K features; And, concatenating the noise V matrix, the identification V matrix and the style V matrix to obtain the concatenated V matrix.
5. The method according to any one of claims 2 to 4, characterized in that: The noise feature also includes a noise Q matrix, and the splicing feature includes a splicing K matrix and a splicing V matrix; The step of mapping the noise feature and the splicing feature based on the attention mechanism to obtain the attention feature includes: The noise Q matrix, the splicing K matrix and the splicing V matrix are mapped based on the attention mechanism to obtain the attention feature.
6. The method according to any one of claims 1 to 5, characterized in that: The performing denoising on the fused features to obtain denoised features includes: Perform noise prediction based on the fusion feature to obtain predicted noise; De-noising the fused feature according to the predicted noise to obtain the de-noised feature.
7. The method according to any one of claims 1 to 6, characterized in that: The encoding of the identification features of the identification image includes: Mapping the identification image to obtain identification mapping features; Performing noise addition on the identification mapping feature to obtain a noisy identification mapping feature; The noise-added identification mapping feature is encoded to obtain the identification feature.
8. The method according to any one of claims 1 to 6, characterized in that: The encoding of the style feature of the at least one style reference image comprises: Mapping the at least one style reference image to obtain a style mapping feature; Adding noise to the style map feature to obtain a noisy style map feature; The noise-added style map feature is encoded to obtain the style feature.
9. The method according to any one of claims 1 to 8, characterized in that: A style image generation model for executing the style image generation method includes: The encoder, feature network and decoder are cascaded in sequence; The encoder and the feature network are used to encode the identification features of the identification image, the style features of the at least one style reference image, and the noise features of the Gaussian noise; the feature network is also used to fuse the noise features, the identification features and the style features to obtain fused features, and perform denoising on the fused features to obtain denoised features; the decoder is used to decode the denoised features to generate the identification style image.
10. The method according to claim 9, characterized in that The encoder includes a first encoder, a second encoder, and a third encoder connected in parallel, the feature network includes a first feature network, a second feature network, and a third feature network connected in parallel, the first encoder is also connected to the first feature network, the second encoder is also connected to the second feature network, the third encoder is also connected to the third feature network, the first feature network and the third feature network are also respectively connected to the second feature network, and the second feature network is also connected to the decoder; Among them, the first encoder is used to map the identification image to obtain the identification mapping feature, and the first feature network is used to encode the noisy identification mapping feature to obtain the identification feature; the third encoder is used to map the at least one style reference image to obtain the style mapping feature, and the third feature network is used to encode the noisy style mapping feature to obtain the style feature; the second feature network is used to encode the noise feature of the Gaussian noise, fuse the noise feature, the identification feature and the style feature to obtain a fused feature, and perform denoising on the fused feature to obtain a denoised feature.
11. The method according to any one of claims 1 to 10, characterized in that: The method further comprises: Acquire training data, each of the training data includes a sample identification image, a sample style reference image, and a sample identification style image, wherein the sample identification image is an image including sample identification information and not having the sample style, the sample style reference image is an image having the sample style and not having the sample identification information, and the sample identification style image is an image having the sample identification information and having the sample style; Encoding the sample identification features of the sample identification image; encoding the sample identification style features of the sample identification style image; and encoding the sample style features of the sample style reference image; Fusion of the sample identification feature, the sample identification style feature and the sample style feature to obtain a sample fusion feature; and denoising of the sample fusion feature to obtain a sample denoising feature; The sample denoising feature is decoded to obtain a predicted identification style image; and the model parameters of the style image generation model are adjusted by reducing the difference between the predicted identification style image and the sample identification style image as a training goal.
12. The method according to claim 11, characterized in that The method further comprises: Constructing a sample style reference image set, each of the sample style reference image sets including a plurality of sample style reference images having the same sample style; Constructing sample image pairs, each of the sample image pairs comprising a sample identification style image and a sample identification image corresponding to the sample identification style image; The training data is constructed according to the sample image pairs and the sample style reference image set.
13. The method according to claim 12, characterized in that The step of constructing a sample style reference image set includes: Based on the first text template, construct a plurality of style texts; Inputting the plurality of style texts into a text image generation model to generate a plurality of sample style reference images corresponding to each of the plurality of style texts; Constructing the sample style reference image set based on the sample style reference images corresponding to each style text of the plurality of style texts; Among them, the first text template includes: a first identifier placeholder, a second identifier placeholder, a third style placeholder and an image quality descriptor, the first identifier placeholder is used to fill in the first identifier information, the second identifier placeholder is used to fill in the second identifier information, the third style placeholder is used to fill in the style, and the image quality descriptor is used to fill in the image quality description.
14. The method according to claim 13, characterized in that The constructing the sample style reference image set based on the plurality of sample style reference images corresponding to each style text of the plurality of style texts includes: Determining a sample style reference image that meets a first screening condition from a plurality of sample style reference images corresponding to each of the plurality of style texts; Clustering the sample style reference images that meet the first screening condition to obtain a plurality of clusters; Based on the plurality of clusters, constructing the sample style reference image set; Among them, the first screening condition includes: the first similarity between the style text and the corresponding sample style reference image is greater than a first threshold, or the second similarity between different sample style reference images corresponding to the style text is greater than a second threshold; each of the plurality of clusters corresponds to a set of sample style reference images.
15. The method according to claim 12, characterized in that The constructing of the sample image pair comprises: Based on the second text template, construct a plurality of logo style texts; Inputting the plurality of logo style texts into the optimized style model to generate a plurality of sample logo style images corresponding to each of the plurality of logo style texts; Removing the sample style from each sample identification style image of the plurality of sample identification style images to obtain the sample identification image corresponding to each sample identification style image; Constructing the sample image pair based on each sample identification style image and the sample identification image corresponding to each sample identification style image; The second text template includes: a first identification placeholder, a second style placeholder and an image quality descriptor, the first identification placeholder is used to fill in the first identification information, the second style placeholder is used to fill in the style, and the image quality descriptor is used to fill in the image quality description.
16. The method according to claim 15, characterized in that The method further comprises: Based on the sample style reference images in the sample style reference image set, optimizing the trainable parameters in the style model to obtain the optimized style model; The optimized style model is used to generate the sample identification style image, and the trainable parameters include adding at least two bypass matrices in the attention network of the style model.
17. A style image generating device, characterized in that: The device comprises: an acquisition module, configured to acquire an identification image, at least one style reference image, and Gaussian noise, wherein the identification image is an image including designated identification information and not having a designated style, the at least one style reference image is an image having the designated style and not including the designated identification information, and the Gaussian noise is used to simulate a feature distribution of the identification style image after noise addition, and the identification style image is an image including the designated identification information and having the designated style; An encoding module, used for encoding the identification feature of the identification image; encoding the style feature of the at least one style reference image; and encoding the noise feature of the Gaussian noise; A processing module, configured to fuse the noise feature, the identification feature and the style feature to obtain a fused feature; and to perform denoising on the fused feature to obtain a denoised feature; A decoding module is used to decode the denoising feature and generate the logo style image.
18. A computer device, characterized in that: The computer device includes: a processor and a memory, the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the style image generation method according to any one of claims 1 to 16.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the style image generation method according to any one of claims 1 to 16.
20. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. The processor obtains the computer instructions from the computer-readable storage medium, so that the processor loads and executes to implement the style image generation method according to any one of claims 1 to 16.