Image style transfer method and device, computer device, storage medium and product
By training and generating a preset style transfer network, adjusting the fusion ratio of content images and style images, and using an adaptive instance normalization network, the problem of poor performance in traditional style transfer is solved, achieving a more efficient style transfer effect.
Patent Information
- Application Number
- CN202211625387.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-16
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2042-12-16
AI Technical Summary
Traditional image style transfer techniques struggle to balance the fusion ratio between content and style images, resulting in poor style transfer performance.
The initial style transfer network is trained based on a preset set of weight parameters, sample content images, and sample style images to generate a preset style transfer network. The fusion ratio between the content images and style images is adjusted, and an adaptive instance normalization network is used to improve the style transfer effect.
It improves the accuracy and robustness of style transfer networks, achieves better style transfer results, and avoids the problems of semantic information loss or unclear stylization caused by improper fusion ratio of content images and style images.
Smart Images

Figure CN115861041B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to an image style transfer method and device, computer equipment, storage medium and product. BACKGROUND
[0002] At present, style transfer technology has been widely used in image processing, computer picture synthesis and computer vision and other aspects. The style transfer technology refers to a kind of image processing technology that image is transferred from original style to another style, and at the same time, the image content does not change. For example, when the real face image is based on the animation image and the style is transferred, the style of the animation image can be transferred to the real face image, so that the original style of the real face image is transferred to the animation style, and the animation face image containing the real face image is generated.
[0003] Conventionally, when the image is processed by style transfer, the style transfer effect is poor. SUMMARY
[0004] Therefore, it is necessary to provide an image style transfer method and device, computer equipment, computer readable storage medium and computer program product, which can improve the style transfer effect of the image.
[0005] In a first aspect, the present application provides an image style transfer method. The method comprises:
[0006] obtaining a content image to be processed;
[0007] inputting the content image into a preset style transfer network for style transfer processing to obtain a target style transfer image; the preset style transfer network is used to convert the original style of the content image into a preset style; wherein the preset style transfer network is obtained by training an initial style transfer network based on a preset weight parameter set, a sample content image and a sample style image;
[0008] outputting the target style transfer image.
[0009] In this embodiment, by acquiring a content image to be processed, the content image is input into a preset style transfer network for style transfer processing to obtain a target style transfer image, and then the target style transfer image is output; wherein the preset style transfer network is used to convert the original style of the content image into a preset style; the preset style transfer network is obtained by training an initial style transfer network based on a preset weight parameter set, a sample content image and a sample style image; that is, in this embodiment, the fusion ratio between the sample content image and the sample style image in the training process of the initial style transfer network is adjusted based on different preset weight parameters, each preset weight parameter in the preset weight parameter set is traversed and iterated to obtain a preset weight parameter with better style transfer effect, which is used as the fusion ratio between the content image and the style image to achieve better style transfer effect; compared with the problem that the fusion ratio between the content image and the style image cannot be balanced in the traditional technology, resulting in poor style transfer effect, the image style transfer method proposed in this application can improve the accuracy of the style transfer network and improve the image effect of style transfer.
[0010] In one embodiment, the training process of the preset style transfer network includes:
[0011] obtaining a preset weight parameter set, a sample content image and a sample style image;
[0012] training the initial style transfer network according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network to generate a candidate style transfer network;
[0013] determining the preset style transfer network from the candidate style transfer network according to the sample style image
[0014] In this embodiment, a preset weight parameter set, a sample content image and a sample style image are obtained; the initial style transfer network is trained according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network, to generate a candidate style transfer network; and the preset style transfer network is determined from the candidate style transfer network according to the sample style image. In this embodiment, the preset weight parameter set is determined according to the value range of the fusion ratio, and the initial style transfer network is iteratively trained by each preset weight parameter in the preset weight parameter set, to obtain the candidate style transfer network corresponding to each preset weight parameter set, and the final preset style transfer network is selected from the candidate style transfer networks with better style transfer effect; that is, the preset weight parameter with the best style transfer effect is determined by the traversal training of different preset weight parameters, and the best image fusion ratio is obtained, which can greatly improve the accuracy of the trained preset style transfer network and improve the processing effect of image style transfer.
[0015] In one of the embodiments, the initial style transfer network is trained according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network, to generate a candidate style transfer network, including:
[0016] The sample content image is subjected to style transfer processing and the initial style transfer network is trained according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network, to generate a candidate style transfer image corresponding to the sample content image and a candidate style transfer network corresponding to the candidate style transfer image.
[0017] In this embodiment, the sample content image is subjected to style transfer processing and the initial style transfer network is trained according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network, to generate a candidate style transfer image corresponding to the sample content image and a candidate style transfer network corresponding to the candidate style transfer image; that is, for each different preset weight parameter, the candidate style transfer image and the candidate style transfer network corresponding to each preset weight parameter are trained, so as to subsequently determine the fusion weight parameter with the best style transfer effect based on each candidate style transfer image, and finally determine the preset style transfer network from each candidate style transfer network; the style transfer effect and the robustness of the preset style transfer network are improved.
[0018] In one of the embodiments, the style transfer processing on the sample content image and the training on the initial style transfer network are performed according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network, to generate a candidate style transfer image corresponding to the sample content image and a candidate style transfer network corresponding to the candidate style transfer image, including:
[0019] The preset weight parameter in the preset weight parameter set, the sample content image and the sample style image are input to the initial style transfer network to perform the style transfer processing on the sample content image, to generate a candidate style transfer image corresponding to the sample content image;
[0020] The initial style transfer network is trained according to the candidate style transfer image and the sample style image, to obtain a candidate style transfer network
[0021] The next preset weight parameter of the preset weight parameter is taken as a new preset weight parameter, and a new candidate style transfer image corresponding to the sample content image and a new candidate style transfer network are iteratively generated until the last preset weight parameter in the preset weight parameter set is reached. The candidate style transfer image and the new candidate style transfer image are taken as the candidate style transfer image, and the candidate style transfer image and the new candidate style transfer network are taken as the candidate style transfer network.
[0022] In the embodiment, each preset weight parameter in the preset weight parameter set is taken in turn as the network parameter of the style transfer network, and the network iterative training is performed according to the sample content image and the sample style image, to obtain the candidate style transfer network and the candidate style transfer image corresponding to the preset weight parameter under different fusion ratios, so as to select the image fusion ratio with the best style transfer effect according to each candidate style transfer image and the sample style image in the subsequent process, to obtain the preset style transfer network. The training process of the preset style transfer network in the embodiment can obtain the candidate style transfer network under different fusion ratios, to balance the fusion ratio between the content image and the style image, and to obtain the preset style transfer network with the best style transfer effect. Not only the network training can be efficiently realized, but also the robustness and accuracy of the style transfer network can be improved.
[0023] In one of the embodiments, the initial style transfer network includes an initial feature extraction network and an initial adaptive instance normalization network; the preset weight parameter includes a first type of preset weight parameter; the preset weight parameter in the preset weight parameter set, the sample content image and the sample style image are input to the initial style transfer network to perform the style transfer processing on the sample content image, to generate a candidate style transfer image corresponding to the sample content image, including:
[0024] input the sample content image and the sample style image into the initial feature extraction network for feature extraction to obtain a first feature image of the sample content image and a second feature image of the sample style image;
[0025] update the network parameters of the adaptive instance normalization network to the first preset weight parameters to generate an initial adaptive instance normalization network;
[0026] input the first feature image and the second feature image into the initial adaptive instance normalization network for style transfer to generate a candidate style transfer image corresponding to the sample content image.
[0027] In this embodiment, the initial style transfer network includes the initial feature extraction network and the initial adaptive instance normalization network; the preset weight parameters include the first preset weight parameters, which are the network parameters in the adaptive instance normalization network and are used to represent the fusion proportion between the content image and the style image; when the server trains the initial style transfer network to generate the candidate style transfer image corresponding to the sample content image, the sample content image and the sample style image are input into the initial feature extraction network for feature extraction to obtain the first feature image of the sample content image and the second feature image of the sample style image; then, the network parameters of the adaptive instance normalization network are updated to the first preset weight parameters to generate the initial adaptive instance normalization network; and the first feature image and the second feature image are input into the initial adaptive instance normalization network for style transfer to generate the candidate style transfer image corresponding to the sample content image; that is, in this embodiment, by improving the traditional adaptive instance normalization network and increasing the weight parameters to balance the fusion proportion between the content image and the style image, the image processing effect of style fusion can be greatly improved.
[0028] In one of the embodiments, inputting the first feature image and the second feature image into the initial adaptive instance normalization network for style transfer to generate the first candidate style transfer image corresponding to the sample content image includes:
[0029] calculating a weighted sum of the mean of the first feature image and the mean of the second feature image based on the first preset weight parameters;
[0030] calculating a weighted sum of the variance of the first feature image and the variance of the second feature image based on the first preset weight parameters;
[0031] inputting the weighted sum of the mean, the weighted sum of the variance, the mean and the variance of the first feature image, and the sample content image into the initial adaptive instance normalization network for style transfer to generate the first candidate style transfer image corresponding to the sample content image.
[0032] In this embodiment, the preset weight parameter is used to calculate the fusion mean and fusion variance between the sample content image and the sample style image, so as to balance the sample content image and the sample style image, improve the balance of image fusion, and further improve the fusion effect of the stylized image.
[0033] In one of the embodiments, the preset weight parameter further includes a second type of preset weight parameter; the initial feature extraction network includes a preset convolutional neural network, a first encoder network and a second encoder network; the sample content image and the sample style image are input into the initial feature extraction network for feature extraction to obtain a first feature image of the sample content image and a second feature image of the sample style image, including:
[0034] The sample content image is input into the preset convolutional neural network for feature extraction to obtain a first feature of the sample content image, and the sample content image is input into the first encoder network for feature extraction to obtain a second feature of the sample content image;
[0035] According to the second type of preset weight parameter, the first feature of the sample content image is fused with the second feature of the sample content image to generate the first feature image of the sample content image;
[0036] The sample style image is input into the second encoder network for feature extraction to obtain a second feature image of the sample style image.
[0037] In this embodiment, the preset weight parameter further includes a second type of preset weight parameter; the initial feature extraction network includes a preset convolutional neural network, a first encoder network and a second encoder network; wherein the preset convolutional neural network and the first encoder network are used for feature extraction of the content image, and the second encoder network is used for feature extraction of the style image, and the second type of preset weight parameter is used to balance the output proportion of the two networks of the preset convolutional neural network and the first encoder network; when the server extracts the feature image of the sample content image and the sample style image through the initial feature extraction network, the sample content image can be input into the preset convolutional neural network for feature extraction to obtain a first feature of the sample content image, and the sample content image can be input into the first encoder network for feature extraction to obtain a second feature of the sample content image; and according to the second type of preset weight parameter, the first feature of the sample content image is fused with the second feature of the sample content image to generate the first feature image of the sample content image; in addition, the sample style image is input into the second encoder network for feature extraction to obtain a second feature image of the sample style image. In this embodiment, the features of the content image are mainly extracted through two different networks, especially for the face image, the problem of loss of face semantic information can be solved, and the accuracy and integrity of the face feature extraction are improved.
[0038] In one embodiment, the preset convolutional neural network is trained based on a preset image database corresponding to the sample content image.
[0039] In this embodiment, the convolutional neural network for extracting facial features is pre-trained based on the preset image database. More semantic information of the facial image can be extracted through the convolutional neural network, the completeness of subsequent facial feature extraction is improved, and the problem of semantic information loss is avoided.
[0040] In a second aspect, the present application also provides an image style transfer device. The device comprises:
[0041] The acquisition module is configured to acquire a content image to be processed.
[0042] The processing module is configured to input the content image into a preset style transfer network for style transfer processing to obtain a target style transfer image. The preset style transfer network is configured to convert the original style of the content image into a preset style. The preset style transfer network is obtained by training an initial style transfer network based on a preset weight parameter set, a sample content image, and a sample style image.
[0043] The output module is configured to output the target style transfer image.
[0044] In a third aspect, the present application also provides a computer device. The computer device comprises a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the image style transfer method in the first aspect are implemented.
[0045] In a fourth aspect, the present application also provides a computer readable storage medium. The computer readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the image style transfer method in the first aspect are implemented.
[0046] In a fifth aspect, the present application also provides a computer program product. The computer program product comprises a computer program. When the computer program is executed by a processor, the steps of the image style transfer method in the first aspect are implemented.
[0047] The aforementioned image style transfer method, apparatus, computer device, storage medium, and computer program product acquire a content image to be processed, input the content image into a preset style transfer network for style transfer processing, obtain a target style transfer image, and then output the target style transfer image. The preset style transfer network is used to convert the original style of the content image into a preset style. The preset style transfer network is obtained by training an initial style transfer network based on a preset set of weight parameters, sample content images, and sample style images. In other words, in this embodiment, the fusion ratio between sample content images and sample style images during the initial style transfer network training process is adjusted based on different preset weight parameters. By iterating through each preset weight parameter in the preset weight parameter set, preset weight parameters with better style transfer effect are trained and used as the fusion ratio between the content image and the style image, achieving a better style transfer effect. Compared to the problem of poor style transfer effect caused by the difficulty in balancing the fusion ratio between the content image and the style image in traditional technologies, the image style transfer method proposed in this application can improve the accuracy of the style transfer network and improve the image effect of style transfer. Attached Figure Description
[0048] Figure 1 This is a diagram illustrating the application environment of the image style transfer method in one embodiment;
[0049] Figure 2 This is a flowchart illustrating an image style transfer method in one embodiment;
[0050] Figure 3 This is a flowchart illustrating the image style transfer method in another embodiment;
[0051] Figure 4 This is a flowchart illustrating the image style transfer method in another embodiment;
[0052] Figure 5 This is a flowchart illustrating the image style transfer method in another embodiment;
[0053] Figure 6 This is a flowchart illustrating the image style transfer method in another embodiment;
[0054] Figure 7 This is a schematic diagram of the overall network structure of a preset style transfer network in one embodiment;
[0055] Figure 8 This is a schematic diagram of the network structure of a face feature extraction network in one embodiment;
[0056] Figure 9 This is a structural block diagram of an image style transfer device in one embodiment;
[0057] Figure 10 is a structural block diagram of an image style transfer device in another embodiment.
[0058] Figure 11 is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION
[0059] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0060] The generative adversarial network belongs to a kind of generative model, and its principle is to iteratively train a model, so that the trained model can generate data according to the data rules given during training. The generative adversarial network can generate samples in parallel, uses very few function restrictions, and has better performance than previous generative models. In the training process of the neural network, the weight values of each layer of the entire network will change constantly, so the layers located behind the network need to constantly adapt to the distribution change of the data, and at the same time, the problem of gradient disappearance is also prone to occur, which reduces the convergence learning speed of the model, therefore, batch normalization (BN) is used to integrate the data of each batch, so that the input features satisfy the distribution of mean value 0 and variance 1. In the initial style transfer network, most of them use batch normalization, and after research, it is found that instance normalization (IN) is more suitable for application in style transfer. The difference lies in that instance normalization is standardized for each sample, and the convergence speed and loss function of instance normalization are better than batch normalization as a whole. After that, on the basis of instance normalization, adaptive instance normalization (AdaIN) appeared, which is a method of fusing the styles of two images to realize style transfer. Its advantages are:
[0061] (1) AdaIN is based on a forward neural network, and the generation speed is very fast.
[0062] (2) It supports the transfer of any style instead of being limited to a specific style. Its input is a content image and a style image, and it can realize the transfer of any style.
[0063] Based on the features x of the given content image and the features y of the style image, the adaptive instance normalization can migrate the style of y to x, and the model mainly consists of a down-sampling encoder (VGG Encoder), an adaptive instance normalization (AdaIN) and an up-sampling decoder (Decoder). The down-sampling encoder Encoder adopts a pre-trained VGG model, and the parameters of the encoder Encoder are fixed during the training process. The first 8 layers (such as the ReLU4-1 layer) of the VGG are used. The feature maps of the content image and the style image are obtained by using the down-sampling encoder, and are input into the adaptive instance normalization residual layer together, so that the features of the style image are migrated to the content image to obtain a new feature image. The new feature image is input into the up-sampling decoder Decoder to obtain the picture after style migration.
[0064] The traditional style migration network is a recurrent generative adversarial network and some deformations thereof, which are used to convert two styles into each other. However, the particularity of the human face feature makes the stylization of the human face not suitable for using the recurrent generative adversarial network to realize, and a deep neural network with stronger performance is needed to extract the human face feature.
[0065] The traditional adaptive normalization algorithm adopts a channel-level method to calculate the pixels, and the specific method is as follows: the feature map region pixels of the input content image are subtracted from the mean of the content image pixels and divided by the variance thereof, then multiplied by the variance of the style image pixels and added to the mean of the style image pixels, and the obtained is the stylized image. Such a method has a disadvantage that it is difficult to balance the proportion of the content image and the style image features. When the proportion of the style image is too large, the generated image is prone to lose the original semantic information of the foreground or the background, affecting the visual effect, and when the proportion of the content image is too large, the generated image is not obviously stylized, and the effect is not outstanding.
[0066] Therefore, the present application proposes a new weight-based adaptive instance normalization method, so that the generated style image is more in line with people's visual perception.
[0067] The technical solutions involved in the embodiments of the present application will be introduced in combination with the scene to which the embodiments of the present application are applied.
[0068] The image style migration method provided by the embodiments of the present application can be applied to, for example Figure 1The application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data required by the server 104 to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things device can be a smart speaker, a smart TV, a smart air conditioner, a smart vehicle device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented by a stand-alone server or a server cluster composed of multiple servers.
[0069] Exemplarily, the terminal 102 can respond to the content image to be processed input by the user, and send the content image to be processed to the server 104; the server 104, after receiving the content image to be processed sent by the terminal 102, performs style transfer processing on the content image to be processed based on a preset style transfer network, to obtain a target image after style transfer processing; the target image is an image including a preset style and content of the content image after converting the original style of the content image into the preset style; then, the server 104 sends the stylized target image to the terminal 102, so that the terminal 102 can output and display the target image.
[0070] In one embodiment, as Figure 2 shown, an image style transfer method is provided, which is applied to Figure 1 the server in the above-mentioned application environment, including the following steps:
[0071] Step 220, obtaining a content image to be processed.
[0072] Optionally, the content image can be any type of image that needs to be converted in style, such as images taken by mobile phones, cameras, etc., images generated by image processing software scanning, drawing, processing, and synthesizing, etc.; in addition, the content image can be any content image such as landscape images, building images, human images, animal images, and object images.
[0073] Optionally, the server can receive the content image to be processed sent by the terminal, or can obtain the content image to be processed from a preset storage location according to a preset path, or can obtain the content image to be processed from a preset database, etc.; the application embodiment does not make specific limitation on the acquisition method of the server to obtain the content image to be processed.
[0074] Step 240, inputting the content image into a preset style transfer network for style transfer processing to obtain a target style transfer image.
[0075] The preset style transfer network is used to convert the original style of the content image into a preset style. The preset style transfer network is obtained by training an initial style transfer network based on a preset weight parameter set, a sample content image, and a sample style image. The sample style image is a sample image corresponding to the preset style.
[0076] Optionally, the preset weight parameter set includes at least one preset weight parameter, which can be used to represent the fusion ratio of the content image and the style image during image fusion. Through the preset weight parameter, the content image and the style image can be balanced to avoid the problem that the semantic information of the stylized image is lost or the stylization is not obvious due to the high proportion of one of the content image and the style image. Since the preset weight parameter represents the fusion ratio of the content image and the style image, the value range of the preset weight parameter should be between 0 and 1, and 0 and 1 should not be included. In order to avoid the high proportion of one of them, the value range of the preset weight parameter can be set to 0.3-0.7, or 0.4-0.6, etc. The value range of the preset weight parameter is not limited in the embodiments of the present application. In actual application, it can be adaptively set according to the stylization requirement.
[0077] Further, after determining the value range of the preset weight parameter, a plurality of preset weight parameters can be determined from the value range according to a preset granularity to generate the preset weight parameter set. For example, the preset granularity can be 1, 0.5, 0.1, etc. The finer the granularity, the higher the accuracy of the preset style transfer network obtained by training. Based on this, when the initial style transfer network is trained based on the sample content image and the sample style image, each preset weight parameter in the preset weight parameter set can be used for iterative training, and the preset weight parameter with good style conversion effect can be determined according to the training result, that is, the fusion ratio between the content image and the style image with good style conversion effect is determined. It should be noted that the fusion ratio corresponding to different content images and style images can be different.
[0078] Optionally, after obtaining the preset style transfer network by training the initial style transfer network based on the preset weight parameter set, the sample content image, and the sample style image, the preset style transfer network can be used to perform style conversion processing on the content image to be processed to obtain a target style transfer image.
[0079] Step 260: output the target style transfer image.
[0080] Optionally, after the server performs style transfer processing on the content image to be processed to obtain a target style transfer image, the server can send the target style transfer image to the terminal, so that the terminal can output and display the target style transfer image to the user.
[0081] In the image style transfer method, the server obtains a content image to be processed, inputs the content image into a preset style transfer network for style transfer processing to obtain a target style transfer image, and then outputs the target style transfer image. The preset style transfer network is used to convert the original style of the content image into a preset style. The preset style transfer network is obtained by training an initial style transfer network based on a preset weight parameter set, a sample content image, and a sample style image. That is, in the embodiment of the application, the fusion ratio between the sample content image and the sample style image in the training process of the initial style transfer network is adjusted based on different preset weight parameters. By traversing each preset weight parameter in the preset weight parameter set, a preset weight parameter with better style transfer effect is trained as the fusion ratio between the content image and the style image to achieve better style transfer effect. Compared with the problem that the traditional technology is difficult to balance the fusion ratio between the content image and the style image, resulting in poor style transfer effect, the image style transfer method proposed in the application can improve the accuracy of the style transfer network and improve the image effect of style transfer.
[0082] Figure 3 The flowchart of the image style transfer method in another embodiment is shown. The embodiment relates to an optional training process of the preset style transfer network. Based on the above-mentioned embodiment, as shown in Figure 3 The method further includes the following steps:
[0083] In step 320, a preset weight parameter set, a sample content image, and a sample style image are obtained.
[0084] Taking the face style transfer processing as an example, the preset weight parameter can be set to a range of 0.3-0.7 according to the experience of controlling the proportion of face images and animation images in the past style fusion. If the weight is divided by 0.5 granularity, the preset weight parameter set can be [0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7]. The sample content image can be a real face image, and the sample style image can be an animation face image.
[0085] In step 340, the initial style transfer network is trained based on each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image, and the initial style transfer network to generate a candidate style transfer network.
[0086] The preset weight parameter is used as a network parameter in the initial style transfer network. When the preset weight parameter is different, that is, the content image and the style image are fused with different fusion ratios, the style transfer image generated by the initial style transfer network is different.
[0087] Optionally, each preset weight parameter in the preset weight parameter set can be used as a network parameter in the initial style transfer network in turn, and the initial style transfer network can be trained by using the sample content image and the sample style image. In the process of iterative training, other network parameters in the initial style transfer network, such as neuron parameters in the network, are optimized. After the training meets a certain preset iterative stop condition, a candidate style transfer network corresponding to each preset weight parameter can be obtained.
[0088] Optionally, the first preset weight parameter, such as 0.3, in the preset weight parameter set can be used as a network parameter of the initial style transfer network, and the initial style transfer network can be iteratively trained by using the sample content image and the sample style image. After the preset iterative stop condition is met, a candidate style transfer network corresponding to the first preset weight parameter is obtained. Then, the second preset weight parameter, such as 0.35, in the preset weight parameter set is used as a network parameter of the candidate style transfer network, and the candidate style transfer network is iteratively trained by using the sample content image and the sample style image. After the preset iterative stop condition is met, a candidate style transfer network corresponding to the second preset weight parameter is obtained. In this way, the iteration is repeated until a candidate style transfer network corresponding to the last preset weight parameter in the preset weight parameter set is obtained.
[0089] In step 360, the preset style transfer network is determined from the candidate style transfer networks according to the sample style image.
[0090] Optionally, after the candidate style transfer network corresponding to each preset weight parameter in the preset weight parameter set is obtained, one candidate style transfer network with better style transfer effect can be further determined from each candidate style transfer network according to the sample style image, and used as the preset style transfer network.
[0091] Exemplarily, one candidate style transfer image with better style transfer effect can be determined according to the similarity between the candidate style transfer image corresponding to the sample content image and output by each candidate style transfer network and the sample style image, and then the candidate style transfer network corresponding to the candidate style transfer image can be determined as the preset style transfer network; that is, the preset weight parameter corresponding to the candidate style transfer network can be determined as the network parameter of the preset style transfer network. It should be noted that after the candidate style transfer network corresponding to the target weight parameter with better style transfer is determined, the preset style transfer network can also be determined based on the candidate style transfer network, and the network parameter corresponding to the image transfer in the preset style transfer network, that is, the network parameter used for image fusion, is the target weight parameter corresponding to the candidate style transfer network; the other network parameters in the preset style transfer network except the target weight parameter can be the same as or different from the network parameters in the candidate style transfer network, or can be determined in other ways, which is not limited in the embodiments of the present application.
[0092] In the embodiment, the preset weight parameter set, the sample content image and the sample style image are obtained; the initial style transfer network is trained according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network, to generate a candidate style transfer network; and the preset style transfer network is determined from the candidate style transfer network according to the sample style image. In the embodiment, the preset weight parameter set is determined according to the value range of the fusion ratio, and the initial style transfer network is iteratively trained by each preset weight parameter in the preset weight parameter set, to obtain the candidate style transfer network corresponding to each preset weight parameter set, and the final preset style transfer network is determined from the candidate style transfer network with better style transfer effect. That is, the preset weight parameter with the best style transfer effect is determined by the iterative training of different preset weight parameters, the best image fusion ratio is obtained, and the accuracy of the preset style transfer network trained can be greatly improved, and the processing effect of image style transfer is improved.
[0093] In an optional embodiment of the present application, the step 302 of training the initial style transfer network according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network to generate a candidate style transfer network can include: performing style transfer processing on the sample content image and training the initial style transfer network according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network to generate a candidate style transfer image corresponding to the sample content image and a candidate style transfer network corresponding to the candidate style transfer image. That is, in the process of iterative training, the candidate style transfer network corresponding to each preset weight parameter and the candidate style transfer image processed by the candidate style transfer network are saved; based on this, when the step 303 of determining the preset style transfer network from the candidate style transfer network according to the sample style image is performed, the candidate style transfer network with the best style transfer effect can be determined from the multiple candidate style transfer networks according to the similarity between the candidate style transfer image corresponding to each preset weight parameter and the sample style image, that is, the best fusion ratio is determined, that is, the preset weight parameter with the best fusion effect is determined from the preset weight parameter set; and then, the preset style transfer network is determined based on the preset weight parameter with the best fusion effect.
[0094] Optionally, as shown in Figure 4 The process of performing style transfer processing on the sample content image and training the initial style transfer network according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network to generate a candidate style transfer image corresponding to the sample content image and a candidate style transfer network corresponding to the candidate style transfer image can include:
[0095] In step 420, the preset weight parameter in the preset weight parameter set, the sample content image and the sample style image are input into the initial style transfer network to perform style transfer processing on the sample content image to generate a candidate style transfer image corresponding to the sample content image.
[0096] In step 440, the initial style transfer network is trained according to the candidate style transfer image and the sample style image to obtain a candidate style transfer network.
[0097] Step 460, taking the next preset weight parameter of the preset weight parameter set as a new preset weight parameter, iteratively generating a new candidate style transfer image corresponding to the sample content image and a new candidate style transfer network until the iteration reaches the last preset weight parameter in the preset weight parameter set, taking the candidate style transfer image and the new candidate style transfer image as the candidate style transfer image, and taking the candidate style transfer image and the new candidate style transfer network as the candidate style transfer network.
[0098] As described in the content of step 302, that is, each preset weight parameter in the preset weight parameter set is iteratively trained as the network parameter of the style transfer network in turn, and the candidate style transfer network and the candidate style transfer image corresponding to each preset weight parameter are obtained. For example, the first preset weight parameter can be taken as the network parameter of the initial style transfer network for a preset number of iterations. After the preset number of iterations is reached, the candidate style transfer network and the candidate style transfer image corresponding to the first preset weight parameter are obtained. At this time, the second preset weight parameter is taken as the network parameter of the candidate style transfer network for a preset number of iterations. After the preset number of iterations is reached, the candidate style transfer network and the candidate style transfer image corresponding to the second preset weight parameter are obtained. Then, the third preset weight parameter is taken as the network parameter of the candidate style transfer network for a preset number of iterations. The iteration is repeated in turn until the last preset weight parameter is taken as the network parameter of the candidate style transfer network for a preset number of iterations. The candidate style transfer network and the candidate style transfer image corresponding to the last preset weight parameter are obtained.
[0099] In this embodiment, by taking each preset weight parameter in the preset weight parameter set as the network parameter of the style transfer network in turn, the network is iteratively trained according to the sample content image and the sample style image, thereby obtaining the candidate style transfer network and the candidate style transfer image corresponding to the preset weight parameter under different fusion ratios. In order to subsequently select the image fusion ratio with the best style transfer effect according to each candidate style transfer image and the sample style image, thereby obtaining the preset style transfer network. The training process of the preset style transfer network in this embodiment can obtain candidate style transfer networks under different fusion ratios, thereby balancing the fusion ratio between the content image and the style image, and obtaining the preset style transfer network with the best style transfer effect. Not only can the network training be efficiently realized, but also the robustness and accuracy of the style transfer network can be improved.
[0100] In an optional embodiment of the present application, the initial style transfer network described above can include an initial feature extraction network and an initial adaptive instance normalization network; wherein the initial feature extraction network is configured to perform feature extraction on the sample content image and the sample style image to obtain a feature map of the sample content image and a feature map of the sample style image; the initial adaptive instance normalization network is configured to perform fusion processing on the feature map of the sample content image and the feature map of the sample style image, i.e. to realize style transfer, to obtain a style-fused image; the preset weight parameters include first preset weight parameters, which are network parameters in the initial adaptive instance normalization network, such as a fusion ratio for image fusion of the feature map of the sample content image and the feature map of the sample style image.
[0101] Based on this, as shown in Figure 5 Step 401 inputs the preset weight parameters in the preset weight parameter set, the sample content image and the sample style image into the initial style transfer network to perform style transfer processing on the sample content image to generate a candidate style transfer image corresponding to the sample content image, including:
[0102] Step 520 inputs the sample content image and the sample style image into the initial feature extraction network to perform feature extraction to obtain a first feature image of the sample content image and a second feature image of the sample style image.
[0103] Optionally, the initial feature extraction network can be a convolutional neural network-based feature extraction network, including but not limited to a convolutional neural network (CNN), a deep neural network (DNN), a generative adversarial network (GAN), etc. Optionally, the feature extraction network can also be an encoder network, a visual geometry group network (VGG), etc., which is not limited in the embodiments of the present application. It should be noted that for the sample content image and the sample style image, one feature extraction network can be used to extract the features of the sample content image and the features of the sample style image, or two feature extraction networks can be used to extract the features of the sample content image and the features of the sample style image. In the case of using two feature extraction networks to extract the features of the sample content image and the features of the sample style image, the two feature extraction networks can be the same or different. In addition, for the feature extraction network for extracting the features of the sample content image or for the feature extraction network for extracting the features of the sample style image, it can be one network or a combination of multiple networks, that is, multiple different types of networks can be used for feature extraction, and the extracted features are fused, and the fused feature map is used as the final feature map. The network structure and type are not limited in the embodiments of the present application.
[0104] For the initial feature extraction network with different network structures, the sample content image and the sample style image are input into the initial feature extraction network to obtain a first feature image corresponding to the sample content image and a second feature image corresponding to the sample style image.
[0105] In step 540, the network parameters of the adaptive instance normalization network are updated to the first preset weight parameters to generate an initial adaptive instance normalization network.
[0106] Conventionally, the adaptive instance normalization network can be implemented by an adaptive instance normalization algorithm, which can be represented by the following formula (1).
[0107]
[0108] wherein x is an input content image (when network training is performed, it can refer to a sample content image), y is an input style image (when network training is performed, it can refer to a sample style image), μ(x) is a variance of the content image, σ(x) is a mean value of the content image, μ(y) is a variance of the style image, σ(y) is a mean value of the style image, and AdaIN(x, y) is a style transfer image.
[0109] Optionally, in the embodiment, the improved adaptive instance normalization algorithm can be represented by the following formula (2).
[0110]
[0111] wherein σ(x, y) is the mean value combined with the content image and the style image, μ(x, y) is the variance combined with the content image and the style image, and W-AdaIN(x, y) is the style transfer image.
[0112] Optionally, σ(x, y) can be represented as:
[0113] σ(x, y) = σ(x) x λ + σ(y) x (1 - λ) (3)
[0114] μ(x, y) can be represented as:
[0115] μ(x, y) = μ(x) x λ + μ(y) x (1 - λ) (4)
[0116] wherein λ is a preset weight parameter, representing the fusion ratio between the content image and the style image; λ is also the network parameter of the adaptive instance normalization network, that is, the first type of preset weight parameter.
[0117] When performing iterative training, the network parameter of the adaptive instance normalization network can be updated to the first type of preset weight parameter to generate an initial adaptive instance normalization network; for example, based on the preset weight parameter set in the above example, the first preset weight parameter in the preset weight parameter set, such as 0.3, can be taken as the network parameter of the adaptive instance normalization network to obtain the initial adaptive instance normalization network, that is, the initial adaptive instance normalization network can be represented as:
[0118]
[0119] Step 560, input the first feature image and the second feature image into the initial adaptive instance normalization network to perform style transfer to generate a candidate style transfer image corresponding to the sample content image.
[0120] Optionally, the first feature image of the sample content image can be taken as x in the above formula (5), and the second feature image of the sample style image can be taken as y in the above formula (5) to be input into the above formula (5) to perform style transfer calculation to obtain the candidate style transfer image corresponding to the sample content image.
[0121] That is, a weighted sum of the mean of the first feature image and the mean of the second feature image can be calculated based on the first preset weight parameter λ, i.e., σ(x, y) = 0.3σ(x) + 0.7σ(y); and a weighted sum of the variance of the first feature image and the variance of the second feature image can be calculated based on the first preset weight parameter λ, i.e., μ(x, y) = 0.3μ(x) + 0.7μ(y); then, the weighted sum of the mean σ(x, y), the weighted sum of the variance μ(x, y), the mean σ(x) and the variance μ(x) of the first feature image, and the sample content image x are input into the initial adaptive instance normalization network for style transfer to generate the first candidate style transfer image corresponding to the sample content image.
[0122] In the embodiment, the initial style transfer network includes an initial feature extraction network and an initial adaptive instance normalization network; the preset weight parameter includes a first preset weight parameter, which is a network parameter in the adaptive instance normalization network and is used to represent the fusion ratio between the content image and the style image; when the server trains the initial style transfer network to generate the candidate style transfer image corresponding to the sample content image, the sample content image and the sample style image are input into the initial feature extraction network for feature extraction to obtain the first feature image of the sample content image and the second feature image of the sample style image; then, the network parameter of the adaptive instance normalization network is updated to the first preset weight parameter to generate the initial adaptive instance normalization network; and the first feature image and the second feature image are input into the initial adaptive instance normalization network for style transfer to generate the candidate style transfer image corresponding to the sample content image; that is, in the embodiment, by improving the traditional adaptive instance normalization network and adding a weight parameter to balance the fusion ratio between the content image and the style image, the image processing effect of style fusion can be greatly improved.
[0123] In an optional embodiment of the present application, the above-mentioned preset weight parameter further includes a second preset weight parameter; the initial feature extraction network can include a preset convolutional neural network, a first encoder network and a second encoder network; wherein the preset convolutional neural network and the first encoder network are used for feature extraction of the content image, and the second encoder network is used for feature extraction of the style image; the second preset weight parameter is used as the fusion ratio of the preset convolutional neural network and the first encoder network, i.e., the feature map of the content image output by the preset convolutional neural network and the feature map of the content image output by the first encoder network are fused to obtain the final feature map of the content image.
[0124] Based on this, as Figure 6As shown, the step 501 inputs the sample content image and the sample style image into the initial feature extraction network for feature extraction to obtain a first feature image of the sample content image and a second feature image of the sample style image, including:
[0125] The step 620 inputs the sample content image into a preset convolutional neural network for feature extraction to obtain a first feature of the sample content image, and inputs the sample content image into a first encoder network for feature extraction to obtain a second feature of the sample content image.
[0126] Optionally, the network structure of the preset convolutional neural network can be a network structure of a traditional convolutional neural network, or an improved network structure based on the traditional convolutional neural network, or a combined network structure based on the traditional convolutional neural network and other neural networks, and the embodiments of the present application do not make specific limitations on the network structure of the preset convolutional neural network. The first encoder network can also be a network structure based on any neural network, and the embodiments of the present application do not make specific limitations on this.
[0127] Optionally, the preset convolutional neural network can be a network obtained after being trained in advance based on a preset image database corresponding to the sample content image. For example, the preset convolutional neural network can be a front 8-layer network of a VGG19 network pre-trained based on a face database. It should be noted that the front 8-layer network of the pre-trained VGG19 network listed in the embodiment is only an example of the preset convolutional neural network, and is not used to limit the preset convolutional neural network.
[0128] Optionally, the network parameters in the first encoder network can be initialized network parameters, which are continuously optimized and updated in the training process to obtain a finally trained first encoder network. Optionally, the network parameters in the first encoder network can include neuron parameters in the network.
[0129] Optionally, the server can input the sample content image into a plurality of different feature extraction networks to obtain a plurality of features corresponding to the sample image, such as inputting the sample content image into a preset convolutional neural network for feature extraction to obtain a first feature of the sample content image, and inputting the sample content image into a first encoder network for feature extraction to obtain a second feature of the sample content image.
[0130] The step 640 performs feature fusion on the first feature of the sample content image and the second feature of the sample content image according to the second type of preset weight parameters to generate a first feature image of the sample content image.
[0131] Optionally, the second type of preset weight parameter can be used as a fusion weight or a fusion ratio for feature fusion between the first feature of the sample content image and the second feature of the sample content image. For example, the fusion process of different features of the content image can be represented by the following formula (6).
[0132] F(X) = ωV(X) + (1-ω)E(X) (6)
[0133] where ω is the second type of preset weight parameter, V(X) is the first feature of the sample content image, E(X) is the second feature of the sample content image, and F(X) is the first feature image of the sample content image.
[0134] Optionally, the value range of the second type of preset weight parameter ω can be 0-1. For example, the initialization of ω can be 0, which gradually increases during the network training process as the network parameters are updated. Alternatively, the initialization of ω can be 1, which gradually decreases during the network training process as the network parameters are updated. Of course, the value range of the second type of preset weight parameter ω can also be 0-0.7, 0.1-0.9, 0.2-0.8, 0.1-0.7, etc., which is not limited in the embodiments of the present application.
[0135] After network training, the second type of preset weight parameter ω is continuously updated to obtain the weight parameter ω with the best fusion effect.
[0136] Step 660, inputting the sample style image into the second encoder network for feature extraction to obtain the second feature image of the sample style image.
[0137] Optionally, the network parameters in the second encoder network can be initialized network parameters, which are continuously updated and optimized in the training process to obtain the finally trained second encoder network. Optionally, the network parameters in the second encoder network can include neuron parameters in the network.
[0138] Optionally, the server can input the sample style image into the second encoder network for feature extraction to obtain the second feature image of the sample style image.
[0139] In the embodiment, the preset weight parameters further include second type preset weight parameters; the initial feature extraction network includes a preset convolutional neural network, a first encoder network and a second encoder network; wherein the preset convolutional neural network and the first encoder network are used for feature extraction of the content image, and the second encoder network is used for feature extraction of the style image; the second type preset weight parameters are used for balancing the output proportion of the preset convolutional neural network and the first encoder network; when the server extracts the feature image of the sample content image and the sample style image through the initial feature extraction network, the first feature of the sample content image can be obtained by inputting the sample content image into the preset convolutional neural network for feature extraction, and the second feature of the sample content image can be obtained by inputting the sample content image into the first encoder network for feature extraction; and according to the second type preset weight parameters, the first feature of the sample content image and the second feature of the sample content image are fused to generate the first feature image of the sample content image; in addition, the second feature image of the sample style image is obtained by inputting the sample style image into the second encoder network for feature extraction. In the embodiment, the features of the content image are mainly extracted through two different networks, especially for the face image, the problem of face semantic information loss can be solved, and the accuracy and integrity of face feature extraction are improved.
[0140] The technical scheme of the present application is described as a whole below. Exemplarily, the overall network structure of the preset style transfer network of the present application can refer to Figure 7As shown, the input image of the network can include a 512x512 long and wide RGB real face image with 3 channels and a 512x512 long and wide RGB cartoon face image with 3 channels; optionally, if the size of the input image exceeds 512x512 pixels, it can be scaled to 512x512 size to adapt to the input of the network, and the dataset images used in the present application can be manually adjusted to a suitable size to adapt to the model training requirements. The entire network is divided into three parts, the first part is a face feature extraction network, including a VGG network and a first encoder network, responsible for extracting face features and providing them to the second network, i.e., a W-AdaIN adaptive instance normalization network, for fusion upsampling; the second part is a cartoon style generation network, i.e., a second encoder network, responsible for extracting cartoon face features and inputting them into the W-AdaIN adaptive instance normalization network for normalization and combination with the face features, and finally outputting the final cartoon face stylized image through the decoder; the third part is the discriminator part in the network model of the present application, responsible for improving the performance of the first encoder network, the second encoder network and the decoder network, providing the generated cartoon face stylized image and the cartoon face image input into the network to the discriminator, scoring the generated cartoon face image, and improving the neuron parameters of the first encoder network, the second encoder network and the decoder network through the score and the adjustment of the loss function.
[0141] Optionally, the network structure diagram of the face feature extraction network based on the fusion of the VGG network and the first encoder network used in the present application can be as shown in Figure 8 As shown, the input image is a real face image with 512x512 long and wide and 3 pixel channels, denoted as X, wherein the first face feature extraction network intercepts the first 8 layers of the pre-trained VGG19 network, denoted as V, and the input real face image is subjected to convolution downsampling and other operations through the V network to extract a feature image with obvious face feature contours, denoted as V(X); the second face feature extraction network is an encoder structure with parameters that can be learned through back propagation, denoted as E, and the input real face image is subjected to convolution downsampling and other operations through the E network to extract another ordinary face feature image, denoted as E(X), which will obtain stronger feature extraction capability after several iterations of learning, and the two feature images are subjected to channel-level weighted summation by assigning a weight parameter ω (i.e., the second type of preset weight parameter described above) to each of them, to obtain a new face feature image, denoted as F, which is the final face style feature image output by the network. The ω initialization can be 0, which gradually increases with the parameter update of the network, and the value range can be 0 to 0.7, and the final weight value is obtained after training.
[0142] The face feature extraction model provided in the application can ensure that the semantic information is not lost in the style transfer image generation process. In addition, the application also provides a new weight-based adaptive instance normalization method, so that the generated style image is more in line with people's perception.
[0143] The formula of the weight-based adaptive instance normalization animation face image generation algorithm can be shown in formula (2) above, wherein the fusion mean value can be shown in formula (3) above, the fusion variance can be shown in formula (4) above, and the fusion ratio between the content image and the style image is balanced by the weight parameter λ.
[0144] Exemplarily, according to the experience of controlling the proportion of face images and animation images in the past style fusion, the change range thereof can be set to 0.3-0.7, and the weight parameter value most suitable for the preset style transfer network of the application is finally determined through the update iteration of the network parameter; optionally, a preset weight parameter set can be set, which can be a preset weight parameter array, for example: [0.3, 0.35, 0.4, 0.45, 0.5, 0.5, 0.55, 0.6, 0.65, 0.7], the length of the array is 10, which corresponds to the values in 100 iterations of epoch, in the training process, the array is traversed, and a weight parameter value is taken out each time as an initialization weight value λ in every 10 iterations of epoch, and finally the best checkpoint is selected through the generated image quality and evaluation index, and the corresponding λ value is taken as the final weight parameter value.
[0145] Optionally, the weight parameter ω can be updated once every 100 iterations of epoch; exemplarily, if the parameter range is 0-1 and the update granularity is 0.05, then after every 100 iterations of epoch, the new weight parameter ω is obtained by increasing 0.05, and the process is repeated.
[0146] The application proposes a new face image feature extraction network based on VGG and encoder fusion to solve the problem of face semantic information loss in the style transfer task, which consists of two feature extraction networks. One is a VGG network trained by a high-definition face dataset. The network is first trained using a face dataset, and the first eight layers of the network are extracted for face feature extraction in the application after the parameters are fixed. The other network is an encoder network, which sets gradually increasing parameters to balance the output proportion of the two networks. It should be noted that the first eight layers of the VGG network selected in the application are only used for illustration and do not limit the VGG network.
[0147] The original adaptive instance normalization algorithm directly aligns the variance and mean of the real image to the variance and mean of the animation style image, resulting in the generated image having the problems of discontinuity and color confusion; therefore, the application proposes a new weight-based adaptive instance normalization animation image generation method, which introduces a weight parameter to adjust the proportion of the variance and mean of the content image and the animation image, can better perform style transfer and improve the image effect of style transfer.
[0148] The application proposes a face image feature extraction network based on VGG network and encoder fusion, and an improved weight-based adaptive instance normalization animation face style image generation algorithm, which can exhibit excellent performance and effect in the field of face animation image generation. According to the generated animation image results and evaluation indicators, it has been shown that the method proposed in the application can surpass most style transfer networks in performance and image quality.
[0149] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps.
[0150] Based on the same inventive concept, the embodiments of the application also provide an image style transfer device for implementing the above-mentioned image style transfer method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more image style transfer device embodiments provided below can refer to the limitations of the image style transfer method in the above text, which will not be repeated here.
[0151] In one embodiment, as shown in Figure 9 An image style transfer device is provided, comprising: an acquisition module 920, a processing module 940 and an output module 960, wherein:
[0152] The acquisition module 920 is configured to acquire a content image to be processed;
[0153] The processing module 940 is used to input the content image into a preset style transfer network for style transfer processing to obtain a target style transfer image; the preset style transfer network is used to convert the original style of the content image into a preset style; wherein, the preset style transfer network is trained on the initial style transfer network based on a preset set of weight parameters, sample content images and sample style images;
[0154] Output module 960 is used to output the target style transfer image.
[0155] In one embodiment, such as Figure 10 As shown, the device also includes a network training module 980, which is used to train a preset style transfer network. The network training module 980 includes an acquisition unit, a training unit, and a determination unit. The acquisition unit is used to acquire a preset set of weight parameters, sample content images, and sample style images. The training unit is used to train the initial style transfer network based on each preset weight parameter in the preset set of weight parameters, the sample content images, the sample style images, and the initial style transfer network to generate candidate style transfer networks. The determination unit is used to determine the preset style transfer network from the candidate style transfer networks based on the sample style images.
[0156] In one embodiment, the training unit is configured to perform style transfer processing on the sample content image and train the initial style transfer network based on each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image, and the initial style transfer network, to generate a candidate style transfer image corresponding to the sample content image and a candidate style transfer network corresponding to the candidate style transfer image.
[0157] In one embodiment, the training unit is configured to input preset weight parameters, sample content images, and sample style images from a preset set of weight parameters into an initial style transfer network to perform style transfer processing on the sample content images, generating candidate style transfer images corresponding to the sample content images; train the initial style transfer network based on the candidate style transfer images and the sample style images to obtain a candidate style transfer network; use the next preset weight parameter of the preset weight parameters as a new preset weight parameter, iteratively generate new candidate style transfer images and new candidate style transfer networks corresponding to the sample content images until iterates to the last preset weight parameter in the preset set of weight parameters, and use the candidate style transfer images and the new candidate style transfer images as candidate style transfer images, and use the candidate style transfer images and the new candidate style transfer networks as candidate style transfer networks.
[0158] In one of the embodiments, the initial style transfer network comprises an initial feature extraction network and an initial adaptive instance normalization network; the preset weight parameters comprise first preset weight parameters; the training unit is configured to input the sample content image and the sample style image into the initial feature extraction network to perform feature extraction, to obtain a first feature image of the sample content image and a second feature image of the sample style image; update the network parameters of the adaptive instance normalization network to the first preset weight parameters to generate the initial adaptive instance normalization network; and input the first feature image and the second feature image into the initial adaptive instance normalization network to perform style transfer, to generate a candidate style transfer image corresponding to the sample content image.
[0159] In one of the embodiments, the training unit is configured to calculate a weighted sum of the mean of the first feature image and the mean of the second feature image based on the first preset weight parameters; calculate a weighted sum of the variance of the first feature image and the variance of the second feature image based on the first preset weight parameters; and input the weighted sum of the mean, the weighted sum of the variance, the mean and the variance of the first feature image, and the sample content image into the initial adaptive instance normalization network to perform style transfer, to generate a first candidate style transfer image corresponding to the sample content image.
[0160] In one of the embodiments, the preset weight parameters further comprise second preset weight parameters; the initial feature extraction network comprises a preset convolutional neural network, a first encoder network and a second encoder network; the training unit is configured to input the sample content image into the preset convolutional neural network to perform feature extraction, to obtain a first feature of the sample content image, and input the sample content image into the first encoder network to perform feature extraction, to obtain a second feature of the sample content image; perform feature fusion on the first feature of the sample content image and the second feature of the sample content image according to the second preset weight parameters, to generate a first feature image of the sample content image; and input the sample style image into the second encoder network to perform feature extraction, to obtain a second feature image of the sample style image.
[0161] In one of the embodiments, the preset convolutional neural network is obtained based on a preset image database corresponding to the sample content image.
[0162] The above various modules of the image style transfer device can be realized by software, hardware and combinations thereof in whole or in part. The above various modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above various modules.
[0163] In one of the embodiments, a computer device is provided, which can be a server, and the internal structure diagram thereof can be as shown in Figure 11As shown in the figure. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the network data of the preset style transfer network. The network interface of the computer device is used to communicate with the external terminal through the network connection. The computer program is executed by the processor to implement an image style transfer method.
[0164] Those skilled in the art can understand that, Figure 11 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0165] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the image style transfer method in each of the above embodiments.
[0166] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by the processor to implement the steps of the image style transfer method in each of the above embodiments.
[0167] In one embodiment, a computer program product is provided, including a computer program, and the computer program is executed by the processor to implement the steps of the image style transfer method in each of the above embodiments.
[0168] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0169] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0170] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. An image style transfer method, characterized in that, The method comprises: acquiring a content image to be processed; inputting the content image into a preset style transfer network for style transfer processing to obtain a target style transfer image; the preset style transfer network is used to convert an original style of the content image into a preset style; wherein the preset style transfer network is obtained by training an initial style transfer network based on a preset weight parameter set, a sample content image and a sample style image; outputting the target style transfer image; wherein the initial style transfer network comprises an initial feature extraction network and an initial adaptive instance normalization network; the initial feature extraction network comprises a preset convolutional neural network, a first encoder network and a second encoder network; the initial feature extraction network is used to perform feature extraction on the sample content image and the sample style image to obtain a feature map of the sample content image and a feature map of the sample style image; the initial adaptive instance normalization network is used to perform fusion processing on the feature map of the sample content image and the feature map of the sample style image; the preset weight parameters comprise first preset weight parameters and second preset weight parameters; the first preset weight parameters are used to adjust the fusion ratio of content images and style images; and the second preset weight parameters are used to adjust the output proportion of the preset convolutional neural network and the first encoder network.
2. The method of claim 1, wherein, The training process of the preset style transfer network comprises: acquiring the preset weight parameter set, the sample content image and the sample style image; training the initial style transfer network according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network to generate a candidate style transfer network; determining the preset style transfer network from the candidate style transfer network according to the sample style image.
3. The method of claim 2, wherein, The training of the initial style transfer network according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network to generate a candidate style transfer network comprises: performing style transfer processing on the sample content image and training the initial style transfer network according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network to generate a candidate style transfer image corresponding to the sample content image and a candidate style transfer network corresponding to the candidate style transfer image.
4. The method of claim 3, wherein, The training of the initial style transfer network according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network to generate a candidate style transfer network comprises: performing style transfer processing on the sample content image and training the initial style transfer network according to each preset weight parameter in the preset weight parameter set, the sample content image, the sample style image and the initial style transfer network to generate a candidate style transfer image corresponding to the sample content image and a candidate style transfer network corresponding to the candidate style transfer image. inputting the preset weight parameter in the preset weight parameter set, the sample content image and the sample style image into the initial style transfer network to perform style transfer processing on the sample content image to generate a candidate style transfer image corresponding to the sample content image; training the initial style transfer network according to the candidate style transfer image and the sample style image to obtain a candidate style transfer network; taking a next preset weight parameter of the preset weight parameter as a new preset weight parameter, iteratively generating a new candidate style transfer image corresponding to the sample content image and a new candidate style transfer network, until the last preset weight parameter in the preset weight parameter set is iterated to, taking the candidate style transfer image and the new candidate style transfer image as a candidate style transfer image, and taking the candidate style transfer image and the new candidate style transfer network as a candidate style transfer network.
5. The method of claim 4, wherein, The inputting the preset weight parameter in the preset weight parameter set, the sample content image and the sample style image into the initial style transfer network to perform style transfer processing on the sample content image to generate a candidate style transfer image corresponding to the sample content image, comprises: inputting the sample content image and the sample style image into an initial feature extraction network to perform feature extraction to obtain a first feature image of the sample content image and a second feature image of the sample style image; updating network parameters of the initial adaptive instance normalization network to the first preset weight parameter to generate the initial adaptive instance normalization network; inputting the first feature image and the second feature image into the initial adaptive instance normalization network to perform style transfer to generate a candidate style transfer image corresponding to the sample content image.
6. The method of claim 5, wherein, The inputting the first feature image and the second feature image into the initial adaptive instance normalization network to perform style transfer to generate a first candidate style transfer image corresponding to the sample content image, comprises: calculating a weighted sum of a mean value of the first feature image and a mean value of the second feature image based on the first preset weight parameter; calculating a weighted sum of a variance of the first feature image and a variance of the second feature image based on the first preset weight parameter; inputting the weighted sum of the mean value, the weighted sum of the variance, the mean value and the variance of the first feature image, and the sample content image into the initial adaptive instance normalization network to perform style transfer to generate the first candidate style transfer image corresponding to the sample content image.
7. The method according to claim 5 or 6, characterized in that, The inputting the sample content image and the sample style image into an initial feature extraction network to perform feature extraction to obtain a first feature image of the sample content image and a second feature image of the sample style image, comprises: inputting the sample content image into the preset convolutional neural network to perform feature extraction, to obtain a first feature of the sample content image, and inputting the sample content image into the first encoder network to perform feature extraction, to obtain a second feature of the sample content image; performing feature fusion on the first feature of the sample content image and the second feature of the sample content image according to the second preset weight parameter, to generate a first feature image of the sample content image; inputting the sample style image into the second encoder network to perform feature extraction, to obtain a second feature image of the sample style image.
8. The method of claim 7, wherein, The preset convolutional neural network is obtained based on a preset image database corresponding to the sample content image.
9. An image processing apparatus characterized by comprising: The device comprises: an acquisition module configured to acquire a content image to be processed; a processing module configured to input the content image into a preset style transfer network to perform style transfer processing, to obtain a target style transfer image; the preset style transfer network is configured to convert an original style of the content image into a preset style; wherein the preset style transfer network is obtained based on an initial style transfer network trained by using a preset weight parameter set, a sample content image, and a sample style image; an output module configured to output the target style transfer image; The initial feature extraction network comprises a preset convolutional neural network, a first encoder network, and a second encoder network; the initial feature extraction network is configured to perform feature extraction on the sample content image and the sample style image, to obtain a feature map of the sample content image and a feature map of the sample style image; and the initial adaptive instance normalization network is configured to perform fusion processing on the feature map of the sample content image and the feature map of the sample style image. The preset weight parameter comprises a first preset weight parameter and a second preset weight parameter; the first preset weight parameter is configured to adjust a fusion ratio of a content image and a style image; and the second preset weight parameter is configured to adjust an output proportion of the preset convolutional neural network and the first encoder network. 10.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-9. The processor executes the computer program to implement the steps of the method in any one of claims 1 to 8.
11. A computer readable storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 8.
12. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 8. The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 8.