Style transfer model training method, video processing method and related devices
By constructing a generative adversarial network model with multiple encoder-decoder structures, style feature transfer maps of different resolutions are generated, solving the problem of poor performance in the transfer of complex targets and complex styles in existing technologies, and achieving high stability and high quality video style transfer.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-20
- Publication Date
- 2026-03-27
AI Technical Summary
Existing image and video style transfer techniques have poor performance and stability on complex targets and styles, and cannot effectively identify and transfer styles in complex regions at multiple scales.
A generative adversarial network (GAN) model is adopted. By constructing a generator with multiple encoder-decoder structures, style feature transfer maps of different resolutions are generated. The GAN model is trained using a training sample set. The generator generates multiple style feature transfer maps and judges them through a discriminator. The loss function is optimized to improve the model's style transfer capability.
It improves the style transfer effect and stability for complex targets, and can achieve rich and stable style transfer of video streams, which is suitable for model training and application on terminal devices and servers.
Smart Images

Figure CN115171023B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video processing, in particular to a style transfer model training method, a video processing method and related devices. BACKGROUND
[0002] Image style transfer is a technology that migrates the picture style in a reference image to an original image, which maintains the subject content structure of the original image and makes it have the corresponding picture style in the reference image. Video style transfer is a picture style transfer on the video level, which requires higher stability and accuracy compared to image style transfer.
[0003] Existing image style transfer technologies usually use deep learning and other methods to extract several feature layers of the image to be transferred, and distinguish content features and style features, and finally mix the content features and style features of different images to achieve style transfer. In order to ensure the quality of the style and content of the transferred image, multiple iterative optimization training is required, or a single model is trained for individual styles. However, these methods are generally simple migration of texture or pixels and other single styles, and the effect and stability of complex targets and complex styles are very poor. SUMMARY
[0004] One of the purposes of the present application is to provide a style transfer model training method, a video processing method and related devices to improve the migration effect and stability of complex targets and complex styles.
[0005] In a first aspect, the present application provides a style transfer model training method, which comprises: obtaining a training sample set; wherein the training sample set comprises at least one content image and at least one reference image, and the reference image has target style features; constructing an initial generative adversarial model; wherein the generative adversarial model comprises a generator and a discriminator; the generator is used to generate multiple style feature transfer images corresponding to each training sample; the resolutions of the multiple style feature transfer images corresponding to each other are different; using the training sample set, the generative adversarial model is trained to obtain a style transfer model corresponding to the target style features; the style transfer model is used to process a video stream to be processed, so that each frame of the video stream to be processed has the target style features.
[0006] In a second aspect, the present application provides a video processing method, comprising: obtaining a video stream to be processed and a target style; inputting each frame image of the video stream to be processed into a style transfer model corresponding to the target style to obtain a target image corresponding to each frame image; wherein the target image has the target style, and the style transfer model is obtained by the style transfer model training method of the first aspect; and obtaining a processed video stream based on all the target images.
[0007] In a third aspect, the present application provides a style transfer model training device, comprising: an obtaining module configured to obtain a training sample set; wherein the training sample set comprises at least one content image and at least one reference image, and the reference image has a target style feature; a constructing module configured to construct an initial generative adversarial model; wherein the generative adversarial model comprises a generator and a discriminator; the generator is configured to generate a plurality of style feature transfer images corresponding to each training sample; the plurality of style feature transfer images each have a different resolution; and a training module configured to train the generative adversarial model by using the training sample set to obtain a style transfer model corresponding to the target style feature; and the style transfer model is configured to process a video stream to be processed so that each frame image of the video stream to be processed has the target style feature.
[0008] In a fourth aspect, the present application provides a video processing device, comprising: an obtaining module configured to obtain a video stream to be processed and a target style; a transfer module configured to input each frame image of the video stream to be processed into a style transfer model corresponding to the target style to obtain a target image corresponding to each frame image; wherein the target image has the target style, and the style transfer model is obtained by the style transfer model training method of the first aspect; and a processing module configured to obtain a processed video stream based on all the target images.
[0009] In a fifth aspect, the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores a computer program capable of being executed by the processor, and the processor is capable of executing the computer program to implement the method of the first aspect or the second aspect.
[0010] In a sixth aspect, the present application provides a readable storage medium, wherein the readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect or the second aspect.
[0011] The application provides a style transfer model training method, a video processing method and related devices, and the method comprises the following steps: acquiring a content image and a reference image for training, then constructing a generative adversarial model, training the generative adversarial model by using a training sample set, and obtaining a style transfer model corresponding to a target style feature. The application takes the generative adversarial model as an initial training model, and uses a generator contained in the generative adversarial model to generate multiple style feature transfer images with different resolutions for each training sample. In the training process, the model can be trained based on the style feature transfer images with different resolutions. In this way, the model can learn image features under different resolutions, thereby enhancing the ability of the model to detect and extract features of a complex target. Finally, the style transfer model obtained through training can accurately realize style transfer of the complex target, thereby improving the accuracy of style transfer. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0013] Figure 1 An application scenario diagram of the style transfer model training method provided by the embodiments of the application is shown.
[0014] Figure 2 A flowchart of the style transfer model training method provided by the embodiments of the application is shown.
[0015] Figure 3 A structure diagram of the generator provided by the embodiments of the application is shown.
[0016] Figure 4 Another structure diagram of the generator provided by the embodiments of the application is shown.
[0017] Figure 5 A flowchart of step S203 provided by the embodiments of the application is shown.
[0018] Figure 6 A flowchart of the video processing method provided by the embodiments of the application is shown.
[0019] Figure 7 A user interface diagram provided by the embodiments of the application is shown.
[0020] Figure 8 A functional module diagram of the style transfer model training device provided by the embodiments of the application is shown.
[0021] Figure 9This is a functional block diagram of a video processing apparatus provided in an embodiment of the present invention;
[0022] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0024] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0025] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0026] In the description of this invention, it should be noted that if terms such as "upper," "lower," "inner," or "outer" are used to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of this invention is usually placed, they are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0027] Furthermore, the terms "first" and "second" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0028] It should be noted that, where there is no conflict, the features in the embodiments of the present invention can be combined with each other.
[0029] Style transfer refers to the transformation of images in two different domains. Specifically, it involves providing a style image and transforming any image into that style while preserving as much of the original image's content as possible. For example, transforming a real photo into a cartoon-style photo, or a real photo into an oil painting-style photo, or a real photo into a hand-drawn-style photo, etc.
[0030] Existing image style transfer techniques typically employ deep learning and other methods to extract several feature layers from the image to be transferred, distinguishing between content features and style features, and finally blending the content and style features of different images to achieve style transfer. To ensure the quality of the style and content of the transferred image, multiple iterations of optimization training are required, or a single model may be trained for individual styles. However, these methods generally only perform simple transfers of texture and pixel styles, and their performance is poor for transferring complex targets and styles. Furthermore, when extended to video style transfer, they suffer from poor style stability.
[0031] For example, related technologies provide a method of replacing the pixel groups of frame images in the video to be transferred with the pixels of the style-transferred cluster centers by extracting the pixel points of the frame images in the video stream and performing pixel clustering. One or more frames are extracted by random sampling or equal interval sampling and processed to obtain the style-transferred video stream.
[0032] However, the above methods are only applicable to style transfer at the pixel level. They do not distinguish the high-level semantics of the transfer target and cannot identify complex targets for detailed style transfer. Furthermore, their style transfer scale is a cluster of pixels with similar conditions, which is a simple, non-fixed region at a single scale. It can only perform simple style transfer on similar regions such as skin, but cannot perform simple segmentation and transfer on complex regions at multiple scales such as faces, eyes, and pupils through clustering.
[0033] To address the aforementioned issues, this invention proposes a style transfer model and a corresponding training method, which enables style transfer for complex targets and effectively improves the quality and stability of style transfer, thus achieving style transfer for detailed, stable, and high-quality live video streams.
[0034] The style transfer model and its training method provided in the embodiments of the present invention will be described in detail below.
[0035] The style transfer model training method provided in this application can be applied to devices with model training capabilities, such as terminal devices and servers. Specifically, the terminal device can be a smartphone, computer, personal digital assistant (PDA), tablet computer, etc.; the server can be an application server or a web server. In actual application deployment, the server can be a standalone server or a cluster server.
[0036] In practical applications, terminal devices and servers can train style transfer models independently or interactively. When training style transfer models interactively, the terminal device can obtain a training sample set from the server and then use the training sample set to train the model to obtain the style transfer model. Alternatively, the server can obtain a training sample set from the terminal and then use the training sample set to train the model to obtain the style transfer model.
[0037] It should be understood that after a terminal device or server executes the training method provided in the embodiments of this application to train a style transfer model, it can send the style transfer model to other terminal devices to run the style transfer model on these terminal devices and realize the corresponding functions; or it can send the style transfer model to other servers to run the style transfer model on other servers and realize the corresponding functions through these servers.
[0038] To facilitate understanding of the technical solutions provided in the embodiments of this application, the training method provided in the embodiments of this application will be introduced below using server-based style transfer model training as an example, combined with actual application scenarios.
[0039] See Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of the style transfer model training method provided in this application embodiment. The scenario includes a terminal device 101 and a server 102 for model training, which are connected via a network. The terminal device 101 can provide the server with content images and reference images. The content images can be images containing arbitrary image content, such as images of people, animals, or scenery. The reference images can be style feature transfer images obtained based on the content images and a target style. The reference images appear in pairs with the content images, exhibiting similar content but different styles. For example, the target style can be, but is not limited to, oil painting style, bright style, colored pencil style, etc., and is not limited here.
[0040] After obtaining the content image and reference image from the terminal device 101 via the network, the server 102 combines the content image and reference image into a training sample set. Next, the server can construct an initial generative adversarial model and use the training sample set to execute the training method provided in this embodiment of the invention on the constructed generative adversarial model, and finally obtain the style transfer model corresponding to the target style. The structure and training method of the generative adversarial model constructed in this embodiment of the invention will be described in detail in the following content.
[0041] After the server 102 generates the style transfer model, it can further send the style transfer model to the terminal device 101 so that the style transfer model can be run on the terminal device 101 and the corresponding functions can be implemented using these style transfer models.
[0042] It is understood that the embodiments of the present invention can pre-train a multi-style transfer model, each transfer model corresponds to an image style, and each transfer model corresponds to a different image style, that is, each transfer model can transfer an image style to an image.
[0043] In some implementations, multiple style transfer models can be stored locally on terminal 101 or server 102. In scenarios where style transfer models are needed, the file of the style transfer model corresponding to the target style can be read directly from the local storage.
[0044] It should be noted that the above Figure 1 The application scenario shown is only one example. In practical applications, the style transfer model training method provided in this application embodiment can also be applied to other application scenarios. No limitation is made here on the application scenario of this style transfer model training method.
[0045] Please see Figure 2 , Figure 2 This is a flowchart illustrating a style transfer model training method provided in an embodiment of this application. For ease of description, the following embodiments use a server as the execution subject. It should be understood that the execution subject of this style transfer model training method is not limited to a server, but can also be applied to devices with model training capabilities, such as terminal devices. Figure 2 As shown, the training method for this style transfer model includes the following steps:
[0046] S201, Obtain the training sample set; wherein, the training sample set includes at least one content image and at least one reference image, and the reference image has target style features;
[0047] S202, Construct the initial generative adversarial model; wherein, the generative adversarial model includes a generator and a discriminator; the generator is used to generate multiple style feature transfer maps corresponding to each training sample; the resolutions of the multiple style feature transfer maps are different for each other;
[0048] S203. Using the training sample set, the generative adversarial model is trained to obtain the style transfer model corresponding to the target style features. The style transfer model is used to process the video stream to be processed so that each frame of the video stream to be processed has the target style features.
[0049] According to the style transfer model training method provided in the embodiments of the present invention, firstly, content images and reference images for training are acquired, then a generative adversarial model is constructed, and then the generative adversarial model is trained using a training sample set to obtain a style transfer model corresponding to the target style features. Since the generator in the generative adversarial model constructed in the embodiments of the present invention can generate multiple style feature transfer maps of different resolutions corresponding to each training sample, the model is trained based on style feature transfer maps of different resolutions, which enables the style transfer model to accurately detect complex targets and extract complex style features of complex targets, and finally obtains a model that can accurately achieve style transfer of complex targets.
[0050] The steps S201 to S203 provided in the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0051] In step S201, the training sample set is obtained.
[0052] In this embodiment of the application, the training sample set includes content images and reference images. The reference image is a style transfer map obtained based on the content image, which is similar to the content image in content but has a different style. The reference image contains target style features. That is, when it is necessary to transfer any style feature to the content image to be processed, the reference image can be obtained based on any style feature and the content image.
[0053] In some implementations, the number of content images and the number of reference images are not limited. The number of content images can be one or more, and the number of reference images is the same as the number of content images. When training using multiple content images and reference images, the model can be trained simultaneously using both the content images and the style image to improve the accuracy of the style transfer model obtained after training.
[0054] To ensure the efficiency and effectiveness of model training, after obtaining the training sample set, the content images and reference images can be normalized first, and the normalized content images and reference images can be used in subsequent training processes.
[0055] In step S202, an initial generative adversarial model is constructed.
[0056] The generative adversarial model in this embodiment of the invention may include a generator and a discriminator, wherein the generator is used to generate multiple style feature transfer maps corresponding to each training sample, and the resolutions of the multiple style feature transfer maps are different from each other.
[0057] To enable the generator to output multiple style feature transfer maps with different resolutions, in one alternative implementation, please refer to [link to implementation details]. Figure 3 ,Figure 3 This is a schematic diagram of the generator provided in an embodiment of the present invention. The generator includes M sub-generator models. The M sub-generator models are connected in series through pooling layers. Each sub-generator model corresponds to a discriminator. The resolution of the style feature transfer map generated by the m-th sub-generator model is greater than the resolution of the style feature transfer map generated by the (m+1)-th sub-generator model. M is greater than or equal to 2. This can be understood as follows: starting with the first sub-generator model that receives the input training sample, all sub-generator models are numbered in ascending order. From the 1st sub-generator model to the Mth sub-generator model, the resolution of the output style feature transfer map decreases sequentially by a preset multiple.
[0058] To achieve the effect of progressively decreasing resolution of the style feature transfer map according to a preset multiplier, please refer to [link to relevant documentation]. Figure 4 , Figure 4 Another schematic diagram of the generator provided in this embodiment of the invention:
[0059] Each sub-generator model can be composed of an encoder and a decoder; each decoder has a first output branch and a second output branch; the input of the encoder of the m-th sub-generator model is the output of the encoder of the (m-1)-th sub-generator model; the input of the decoder of the m-th sub-generator model is the output of the encoder of the m-th sub-generator model and the output of the first output branch of the decoder of the (m+1)-th sub-generator model; the second output branch of each decoder is used to output a first style feature transfer map or a second style feature transfer map.
[0060] Understandably, the encoder-decoder model is a common convolutional neural network model. After training, it can detect targets in the encoder stage and restore image content in the decoder stage. In the process of using the encoder-decoder model, existing technologies extract features by inputting feature maps into the encoder to obtain feature maps with reduced resolution. Then, the feature maps are passed to the decoder for decoding. The intermediate features of each level of the encoder are passed to the decoder through skip connections to help with decoding. Finally, the decoder outputs a set of decoded maps.
[0061] As can be seen from the above, the existing encoder-decoder model only has a single output, which leads to an imbalance in the model's encoding and decoding capabilities for multi-scale target features. Complex targets are typically composed of features at multiple scales, and this imbalance in encoding and decoding capabilities is detrimental to style transfer for complex targets. Furthermore, since the decoder only has one output branch, only the highest resolution information is explicitly supervised during training. The lack of constraints on the intermediate layers makes them prone to output perturbations, resulting in unstable style transfer performance in the video stream.
[0062] Therefore, embodiments of the present invention construct such a model based on multiple encoder-decoder models.Figure 4 The generator shown can be configured such that each encoder-decoder pair can serve as a sub-generator model. To achieve the effect of each sub-generator model outputting style feature transfer maps at different resolutions, this embodiment of the invention employs a special design for the encoder-decoder pair within each sub-generator model, as detailed below:
[0063] Step 1: For the encoder within each sub-generator model, determine the numbers N of the convolutional and pooling layers based on the number of the sub-generator model and the total number of sub-generator models. Concatenate the N groups of convolutional and pooling layers to form the encoding. For example, for the m-th (1≤m≤M) sub-generator model, N = M-m+1.
[0064] As can be seen, multiple sub-generator models can be numbered sequentially according to the data flow between encoders to obtain the corresponding number for each sub-generator model.
[0065] Second: For the decoder within each sub-generator model, it consists of concatenated convolutional layers and upsampling layers with the same number N as the encoder. The decoder has a first output branch and a second output branch. The first output branch is used to output feature maps, and the second output branch is used to output generated style feature transfer maps. The second output branch can be composed of convolutional layers with 3 channels.
[0066] Third: The decoder of the m-th (1≤m≤M) sub-generator model has two inputs, namely the output of the m-th encoder and the output of the (m+1)-th decoder. For the m-th decoder, its output feature map is encoded by m-1 encoders and decoded by m-1 decoders, merged, and then subjected to M-m+1 encoding and decoding cycles within the sub-generator model.
[0067] In other words, the feature map output from the first output branch of each decoder will undergo a constant number of encoding and decoding operations (m-1 + M-m+1 = M), thus enabling encoding and decoding of cardinality from 1 to 2. M-1 The target features in the receptive field have considerable encoding and decoding capabilities, which alleviates the problem of imbalance in the encoding and decoding capabilities of multi-scale target features. This is beneficial for the model to implicitly learn to detect targets at different scales and achieve style transfer for complex targets at multiple scales.
[0068] Furthermore, since each decoder also has a second output branch, meaning the entire generator has M output branches, it can output images with a resolution reduced by 1 to 2 times compared to the original image. M-1 The style feature transfer map is multiplied by 10, while the low-resolution style feature transfer map is less affected by brightness changes, noise and motion. This method of gradually adding transfer details from coarse to fine can significantly improve the stability of the style feature transfer map.
[0069] In step S203, the generative adversarial model is trained using the training sample set to obtain a style transfer model corresponding to the target style features. The style transfer model is used to process the video stream to be processed so that each frame of the video stream to be processed has the target style features.
[0070] In conjunction with the above appendix Figure 3 and attached Figure 4 The above step S203 can be understood as follows: For the content image, the generator obtains M first style feature transfer maps y' with the resolution reduced by a factor. m Where m = 1, 2, ..., M, y' m The first discrimination information is obtained by inputting it into the m-th discriminator. The reference image is then downsampled to obtain M second style feature transfer maps y with the resolution reduced by a factor. m , m = 1, 2, ... M, y i The input is fed into the m-th discriminator to obtain the second discriminant information. Then, based on the first and second discriminant information obtained from all discriminators, the loss values of multiple preset loss functions are calculated until the preset conditions are met, and the training ends.
[0071] It is understood that, in the process of training the style transfer model, this embodiment of the invention aims to obtain style transfer maps at different scales by downsampling the reference image, using these maps as supervisory information. The generator then downsamples and encodes the content image. During this process, the generator can detect the various targets to be transferred in the content image. Then, in the decoding stage, the decoded features are transformed into style feature transfer maps at the corresponding resolution. When style transfer maps at different scales exist as supervisory information, the features corresponding to the decoding stage can be explicitly trained as multi-scale features. In other words, each scale of the reference image requires features of the same scale to generate. This ensures that each stage of the decoder focuses on generating features at its corresponding scale, thereby obtaining a refined style transfer map at the final decoding output.
[0072] In this embodiment of the invention, the discriminator is used to determine whether the input image is a style feature transfer map generated by the generator or a reference image. For example, suppose the feature map output by the first output branch of the decoder of the m-th sub-generator model is s'. m The second output branch of the decoder outputs the style feature transfer map y'. m Then y' is determined by the m-th discriminator. m Whether to output discriminative information is a style feature transfer map generated by the generator or a reference image.
[0073] This uses a common generative adversarial training scenario: assuming the style feature transfer map y' generated by the generator is false and the reference image y is true, the generator's goal is to classify y' as true, and the discriminator's goal is to classify y as true and y' as false. Specifically, the input (which could be y or y') is passed through several convolutional pooling layers to obtain one or more loss values. Then, a GAN loss (LSGAN is used here) is applied to derive the losses corresponding to the generator and discriminator, thus performing generative adversarial training.
[0074] The significance of the above training method lies in enabling the style transfer map generated by the generator during training to be consistent with the reference style transfer map. Figure 1 Therefore, they are indistinguishable. This allows the style transfer map during actual inference to more closely resemble the style of the reference image.
[0075] Therefore, in one alternative implementation, please refer to Figure 5 , Figure 5 A schematic flowchart of step S203 provided in an embodiment of the present invention:
[0076] S203-1, the training sample set is input into the generative adversarial model, and each sub-generative model is used to generate each first style feature transfer map corresponding to each content image and each second style feature transfer map corresponding to each reference image by downsampling.
[0077] In other words, during training, the content image is input into the generator to generate first style feature transfer maps at different resolution scales. The reference image is input into the model and is directly obtained at different resolution scales through downsampling, which is the second style feature transfer map. Both the first and second style feature transfer maps are input into the discriminator for discrimination. The discriminator gives different discrimination information by judging whether the input image is the generated image or the reference image.
[0078] S203-2, Input the first style feature transfer map and the second style feature transfer map corresponding to each sub-generation model into the discriminator corresponding to the sub-generation model respectively, to obtain the first discrimination information of each first style feature transfer map and the first discrimination information of each second style feature transfer map.
[0079] It is understood that, since the content image and the reference image in the embodiments of this application appear in pairs, for each content image, a generator generates multiple first style feature transfer maps corresponding to it, and then a second style feature transfer map of the corresponding reference image is obtained by downsampling. Then, the first style feature transfer map and the second style feature transfer map are input into the discriminator for discrimination.
[0080] The first discriminant information mentioned above refers to the discriminator's determination of whether the first style feature transfer map of the input is generated by the generator or a reference image; the second discriminant information mentioned above refers to the discriminator's determination of whether the second style feature transfer map of the input is generated by the generator or a reference image.
[0081] S203-3, Based on the first and second discriminant information obtained from all discriminators, determine the loss values of the multiple loss functions corresponding to the generative adversarial model;
[0082] S203-4, the loss values of multiple loss functions are backpropagated to each layer of the generative adversarial model to iteratively update the model parameters until the preset conditions are met. The trained generative adversarial model is then used as the style transfer model corresponding to the target style features.
[0083] In optional implementations, the multiple loss functions in the embodiments of the present invention can be respectively: least squares loss function. Learning the perceptual image patch similarity loss function Feature matching loss function Style feature loss function The expressions for the above loss functions are as follows:
[0084]
[0085] Where x is the normalized content image, y is the normalized reference image, Flpips and FVGG are the LPIPS and VGG inference models respectively, j is the output feature of a certain layer of the model, and w j The weights are the weights of the loss function at layer j.
[0086] In an alternative implementation, to optimize the computational graph of the trained model to accelerate inference, the following steps may also be performed:
[0087] Remove the second output branch from the remaining sub-generator models (excluding the initial sub-generator model) in the style transfer model, and use the removed style transfer model as the style transfer model corresponding to the target style feature.
[0088] Understandably, in conjunction with the appendix Figure 4 It can be seen that the initial sub-generator model (which can also be understood as the first sub-generator model) has the highest output resolution. Therefore, in practical applications, it is only necessary to retain the output of the initial sub-generator model.
[0089] In practical applications, the style transfer model trained above can be deployed on the target device via an inference engine (including but not limited to TensorRT, MNN, etc.), and run multiple times on the target device to achieve the effect of inference warm-up, so that the model can be compatible with the target device.
[0090] After obtaining the style transfer model corresponding to each style feature, style transfer processing can be performed on the video stream or image. Therefore, this embodiment of the invention also provides a video processing method, please refer to [link to relevant documentation]. Figure 6 , Figure 6 This is a schematic flowchart illustrating a video processing method provided in an embodiment of the present invention. The execution entity of this video processing method may be... Figure 1 In the terminal 101 or server 102, the method may include:
[0091] S301, acquire the video stream to be processed and the target style.
[0092] In this embodiment of the invention, the real-time video stream can be a video stream originating from the network or a video stream stored locally on the device.
[0093] In an alternative implementation, the method for obtaining the aforementioned video stream and target style can be found in [reference needed]. Figure 7 , Figure 7 This is a schematic diagram of a user interface provided in an embodiment of the present invention. Based on this user interface, the above step S301 can be implemented as follows:
[0094] a1 displays the user interface; the user interface includes a video acquisition area and a style selection area.
[0095] a2 responds to user actions in the video acquisition area and obtains the video stream;
[0096] a3, select the desired style from the style selection area.
[0097] It is understandable that the above user interface is merely an example and not a limitation on the user interface. In actual scenarios, the video acquisition area and style selection area can also be displayed through two sub-interfaces. That is, the sub-interface corresponding to the video acquisition area can be displayed first, and the sub-interface of the style selection area can be triggered to display after the video stream is obtained.
[0098] It is also understandable that the sub-interface for the style selection area can not only display the corresponding icons for various styles, but also display the preview interface for each style. In other words, the preview interface can show the user the effect of each style so that the user can choose the target style that meets their needs.
[0099] S302, each frame of the video stream to be processed is input into the style transfer model corresponding to the target style to obtain the target image corresponding to each frame.
[0100] The target image possesses the target style. The style transfer model described above is obtained through the style transfer model training method provided in this embodiment of the invention.
[0101] S303, the processed video stream is obtained based on all target images.
[0102] According to the video processing method provided in the embodiments of this application, after obtaining the video stream and the target style, the style transfer model corresponding to the target style is used to perform style transfer processing on each frame of the video stream to be processed, so as to obtain the target image with the target style. Then, the style-transferred video stream is obtained based on the target image. The whole process uses a pre-trained style transfer model, which can realize style transfer for complex targets and improve the accuracy and stability of the transfer.
[0103] In optional implementations, the video stream in this application is typically stored as a YUV video frame byte string. Therefore, after obtaining the video stream, the frame data needs to be converted into an original image in RGB format. The format of each frame image is an RGB image. Therefore, before inputting each frame image into the style transfer model, the video frame byte string needs to be converted into a frame image. Therefore, this embodiment of the invention provides a conversion method, namely:
[0104] b1 reads the frame data corresponding to the video stream and preprocesses the frame data to obtain the YUV component data corresponding to each frame image.
[0105] Understandably, the frame data is in YUV420p / I420 format. To construct the data structures needed for the model and transmit them to the graphics card, the expression for the frame data is as follows:
[0106] {I Y ,I U ,I V} = DecodeToDevice(ByteStream) in )
[0107] The above relational expression means that: byte data is read from memory, transmitted to the processing device (GPU or NPU, no additional transmission is required if it is CPU), and divided into three YUV component data according to the total data volume in a 4:1:1 ratio, namely the Y component subgraph, the U component subgraph and the V component subgraph.
[0108] b2, based on the YUV component data corresponding to each frame of the image and the preset color space conversion matrix, obtains the RGB format data corresponding to each frame of the image.
[0109] The frame data obtained in step b1 is processed to separate YUV sub-images. Nearest-neighbor interpolation upsampling is then performed on the U and V component sub-images respectively. After recombinizing the sub-images, the result is multiplied by the YUV-to-RGB conversion matrix. Finally, the RGB format data corresponding to each frame image is obtained, as shown in the following expression:
[0110]
[0111] RGB=MY UV2RGB *I YUV
[0112] b3 converts the RGB format data corresponding to each frame of the image into the data format corresponding to the style transfer model.
[0113] Since style transfer models are based on data structures suitable for the model itself, in order for the model to process video stream data, it is also necessary to convert the RGB format data corresponding to each frame of the image into a data structure that the model can process. The specific conversion method is as follows:
[0114]
[0115] y = G(x)
[0116] Where x represents each frame of the transformed image, G represents the style transfer model, and y represents the output of the style transfer model;
[0117] In an optional implementation, in order to display the model output data on the terminal, the data format of the model output data needs to be converted into RGB format first, and then the RGB format needs to be converted into YUV format. Therefore, this embodiment of the invention provides an optional implementation, namely, the implementation based on the processed video stream of all target images may include the following steps:
[0118] c1, based on the color space transformation matrix, converts each target image into YUV format data;
[0119] The YUV format data here refers to frame data in YUV420p / I420 format.
[0120] c1, based on YUV format data, obtains the YUV component data corresponding to the target image.
[0121] The YUV component data here refers to the sub-image in YUV format.
[0122] After obtaining the data output by the model, it can first be converted into RGB format data according to the following formula:
[0123] RGB = [Resize(y) + 1] * 127.5
[0124] Then, the obtained RGB format data is multiplied by the RGB to YUV format conversion matrix, and the U component sub-image and V component sub-image are downsampled. Finally, after byte encoding, the output YUV420p video frame is obtained and transmitted to the terminal device. The specific conversion process is shown in the table below:
[0125] YUV=M RGB2YUV *O RGB
[0126]
[0127] yteStream out =EncodeToHost(O Y O U O V )
[0128] This involves acquiring three data components from the processing device, concatenating them together, and converting them into byte data.
[0129] In the above formula, O Y O U O V For the three components of the output stylized YUV video stream frame data, O YUV This is the output image in YUV444 format. MRGB2YUV is the color space conversion matrix from RGB to YUV.
[0130] Through the above implementation method, a style-transferred video stream can be obtained, and users can intuitively experience the style-transferred video effect on their terminal devices.
[0131] The style transfer model training method provided in this application can be executed in a hardware device or as a software module. When the style transfer model training method is implemented as a software module, this application also provides a style transfer model training method apparatus. Please refer to [link to relevant documentation]. Figure 8 , Figure 8 This is a functional block diagram of the style transfer model training device provided in the embodiments of this application. The style transfer model training device 400 may include:
[0132] The first acquisition module 410 is used to acquire a training sample set; wherein the training sample set includes at least one content image and at least one reference image, and the reference image has target style features;
[0133] Module 420 is used to build the initial generative adversarial model; wherein, the generative adversarial model includes a generator and a discriminator; the generator is used to generate multiple style feature transfer maps corresponding to each training sample; the resolutions of the multiple style feature transfer maps are different from each other;
[0134] Training module 430 is used to train the generative adversarial model using the training sample set to obtain the style transfer model corresponding to the target style features; the style transfer model is used to process the video stream to be processed so that each frame of the video stream to be processed has the target style features.
[0135] It is understandable that the first acquisition module 410, the construction module 420, and the training module 430 can be executed collaboratively. Figure 2 Each step in the process is used to achieve the corresponding technical effect.
[0136] In an optional implementation, the aforementioned building module 420 is used to build such as Figure 3 or Figure 4 The generator shown.
[0137] In an optional implementation, the training module 430 is specifically used to perform Figure 6 Each step in the process is used to achieve the corresponding technical effect.
[0138] In an optional embodiment, the above-mentioned apparatus may further include a processing module, which is used to remove the second output branch from the remaining sub-generator models other than the initial group sub-generator model in the style transfer model, and use the removed style transfer model as the style transfer model corresponding to the target style feature.
[0139] The video processing method provided in this application can be executed in a hardware device or as a software module. When the video processing method is implemented as a software module, this application also provides a style transfer model training method apparatus. Please refer to [link to relevant documentation]. Figure 9 , Figure 9 This is a functional block diagram of a video processing apparatus provided in an embodiment of this application. The video processing apparatus 500 may include:
[0140] The second acquisition module 510 is used to acquire the video stream to be processed and the target style;
[0141] The transfer module 520 is used to input each frame of the video stream to be processed into the style transfer model corresponding to the target style to obtain the target image corresponding to each frame; wherein, the target image has the target style, and the style transfer model is obtained by the style transfer model training method provided in the embodiment of the present invention;
[0142] Processing module 530 is used to obtain the processed video stream based on all target images.
[0143] It is understandable that the second acquisition module 510, the migration module 520, and the processing module 530 can execute collaboratively. Figure 7 Each step in the process is used to achieve the corresponding technical effect.
[0144] In an optional implementation, the second acquisition module 510 is specifically used to: display a user interface; the user interface has a video acquisition area and a style selection area; respond to user operations on the video acquisition area to acquire a video stream; and respond to selection operations on the style selection area to acquire a target style.
[0145] In an optional implementation, the processing module 530 is further configured to read the frame data corresponding to the video stream, and preprocess the frame data to obtain the YUV component data corresponding to each frame image; based on the YUV component data corresponding to each frame image and the preset color space conversion matrix, obtain the RGB format corresponding to each frame image; and convert the RGB format data corresponding to each frame image into the data format corresponding to the style transfer model.
[0146] In an optional implementation, the processing module 530 is specifically used to convert each target image into YUV format data based on a color space conversion matrix; obtain the YUV component data corresponding to the target image based on the YUV format data; and generate a processed video stream based on the YUV component data corresponding to all target images.
[0147] This invention also provides an electronic device, please refer to [link to relevant documentation]. Figure 10 , Figure 10 This is a structural block diagram of an electronic device provided in an embodiment of the present invention.
[0148] like Figure 10 As shown, the electronic device 600 includes a memory 601, a processor 602, and a communication interface 603. The memory 601, processor 602, and communication interface 603 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0149] The memory 601 can be used to store software programs and modules, such as the instructions / modules of the style transfer model training device 400 or the video processing device 500 provided in this embodiment of the invention. These can be stored in the memory 601 in the form of software or firmware, or embedded in the operating system (OS) of the electronic device 600. The processor 602 executes various functional applications and data processing by executing the software programs and modules stored in the memory 601. The communication interface 603 can be used to communicate with other node devices for signaling or data.
[0150] The memory 601 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0151] Processor 602 can be an integrated circuit chip with signal processing capabilities. Processor 602 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0152] Understandable. Figure 10 The structure shown is for illustrative purposes only; the electronic device 600 may also include more than [other components]. Figure 10 The more or fewer components shown, or having the same Figure 10 The different configurations shown. Figure 10 The components shown can be implemented using hardware, software, or a combination thereof.
[0153] This application also provides a readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the style transfer model training method or video processing method as described in any of the foregoing embodiments. The computer-readable storage medium can be, but is not limited to, various media capable of storing program code, such as a USB flash drive, external hard drive, ROM, RAM, PROM, EPROM, EEPROM, magnetic disk, or optical disk.
[0154] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A style transfer model training method, characterized in that, The method comprises: obtaining a training sample set; wherein the training sample set comprises at least one content image and at least one reference image, and the reference image has a target style feature; constructing an initial generative adversarial model; wherein the generative adversarial model comprises a generator and a discriminator; the generator is used to generate a plurality of style feature transfer images corresponding to each training sample; the plurality of style feature transfer images each correspond to a resolution different from each other; the generator comprises M groups of sub-generative models; the resolution of the style feature transfer image generated by the mth group of sub-generative models is greater than the resolution of the style feature transfer image generated by the m+1th group of sub-generative models; wherein M is greater than or equal to 2, and m is a positive integer; training the generative adversarial model using the training sample set to obtain a style transfer model corresponding to the target style feature, comprising: inputting the training sample set into the generative adversarial model, using each sub-generative model to generate a first style feature transfer image corresponding to each content image, and obtaining a second style feature transfer image corresponding to each reference image through downsampling; inputting the first style feature transfer image and the second style feature transfer image corresponding to each sub-generative model into the discriminator corresponding to the sub-generative model respectively to obtain first discrimination information of each first style feature transfer image and second discrimination information of each second style feature transfer image, determining a loss value of each loss function corresponding to the generative adversarial model based on the first discrimination information and the second discrimination information obtained by all the discriminators; and propagating the loss value of each loss function to each layer of the generative adversarial model in reverse to perform model parameter iterative update until a preset condition is reached, and using the trained generative adversarial model as the style transfer model corresponding to the target style feature.
2. The style transfer model training method of claim 1, wherein, The M groups of sub-generative models are connected in series through a pooling layer; each group of sub-generative models corresponds to a discriminator.
3. The style transfer model training method of claim 2, wherein each group of sub-generative models is combined by an encoder and a decoder; each decoder has a first output branch and a second output branch; the input of the encoder of the mth group of sub-generative models is the output of the encoder of the m-1th group of sub-generative models; the input of the decoder of the mth group of sub-generative models is the output of the encoder of the mth group of sub-generative models and the output of the first output branch of the decoder of the m+1th group of sub-generative models; and the second output branch of each decoder is used to output the first style feature transfer image or the second style feature transfer image.
4. The style transfer model training method of claim 3, characterized in that, The plurality of loss functions are respectively: a least square loss function, a learning perceptual image block similarity loss function, a feature matching loss function, and a style feature loss function.
5. The style transfer model training method of claim 2, wherein, After training the generative adversarial model using the training sample set to obtain the style transfer model corresponding to the target style feature, the method further comprises: The second output branch in the remaining sub-generation models in the style transfer model except the initial group of sub-generation models is pruned, and the pruned style transfer model is taken as the style transfer model corresponding to the target style feature.
6. A method for video processing, comprising: The method comprises: obtaining a video stream to be processed and a target style; inputting each frame image of the video stream to be processed into the style transfer model corresponding to the target style to obtain a target image corresponding to each frame image; wherein the target image has the target style, and the style transfer model is obtained by the style transfer model training method in any one of claims 1-5; obtaining a processed video stream based on all the target images.
7. The video processing method of claim 6, wherein, Obtaining a video stream to be processed and a target style comprises: displaying a user interface; the user interface has a video acquisition area and a style selection area; obtaining the video stream in response to user operation on the video acquisition area; obtaining the target style in response to selection operation on the style selection area.
8. The video processing method of claim 6, wherein, Before inputting each frame image of the video stream to be processed into the style transfer model corresponding to the target style to obtain a target image corresponding to each frame image, the method further comprises: reading frame data corresponding to the video stream and pre-processing the frame data to obtain YUV component data corresponding to each frame image; obtaining an RGB format corresponding to each frame image based on the YUV component data corresponding to each frame image and a preset color space conversion matrix; converting the RGB format data corresponding to each frame image into a data format corresponding to the style transfer model.
9. The video processing method of claim 6, wherein, Obtaining a processed video stream based on all the target images comprises: converting each target image into YUV format data based on a color space conversion matrix; obtaining YUV component data corresponding to the target image based on the YUV format data, and generating the processed video stream based on the YUV component data corresponding to all the target images.
10. A style transfer model training apparatus, comprising: comprises: an acquisition module configured to acquire a training sample set; wherein the training sample set comprises at least one content image and at least one reference image, and the reference image has target style features; a construction module configured to construct an initial generative adversarial model; wherein the generative adversarial model comprises a generator and a discriminator; the generator is configured to generate a plurality of style feature transfer images corresponding to each training sample; the plurality of style feature transfer images each correspond to a resolution different from each other; the generator comprises M groups of sub-generation models; a style feature transfer image generated by an mth group of sub-generation models has a resolution greater than a style feature transfer image generated by an (m+1)th group of sub-generation models; wherein M is greater than or equal to 2, and m is a positive integer; The training module is configured to train the generative adversarial model by using the training sample set to obtain a style transfer model corresponding to the target style feature, including: inputting the training sample set into the generative adversarial model, generating each first style feature transfer image corresponding to each content image by using each sub generative model, and obtaining each second style feature transfer image corresponding to each reference image by down-sampling; inputting the first style feature transfer image and the second style feature transfer image corresponding to each sub generative model into the discriminator corresponding to the sub generative model respectively to obtain first discrimination information of each first style feature transfer image and second discrimination information of each second style feature transfer image, determining a loss value of each loss function corresponding to the generative adversarial model based on the first discrimination information and the second discrimination information obtained by all the discriminators; and performing backward propagation of the loss value of each loss function to each layer of the generative adversarial model to perform iterative update of model parameters until a preset condition is reached, and taking the trained generative adversarial model as the style transfer model corresponding to the target style feature.
11. A video processing apparatus, comprising: The method comprises: The acquisition module is configured to acquire a video stream to be processed and a target style; The transfer module is configured to input each frame image of the video stream to be processed into the style transfer model corresponding to the target style to obtain a target image corresponding to each frame image; The target image has the target style, and the style transfer model is obtained by the style transfer model training method according to any one of claims 1 to 5. The processing module is configured to obtain a processed video stream based on all the target images.
12. An electronic device, comprising: The computer program is executed by the processor to implement the method according to any one of claims 1 to 5 or the method according to any one of claims 6 to 9.
13. A readable storage medium, having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image processing method and apparatus, electronic device and storage medium
CN110598781A
Chinese character style migration method and system based on multi-task adversarial learning network
CN111553246A
Identity migration model construction method and device, electronic equipment and readable storage medium
CN113592982A