Special effect processing method and device, electronic equipment, storage medium and program product
By deploying a lightweight generative model on mobile devices and utilizing server-side training data, the problems of slow response and poor effects in mobile special effects processing requests were solved, enabling fast and high-quality special effects processing and improving the user experience.
Patent Information
- Application Number
- CN202410976851.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-20
AI Technical Summary
Mobile devices have slower response times for special effects processing requests and poorer effects, which negatively impacts the user experience.
A lightweight first generative model is deployed on the mobile device. The first generative adversarial network is trained to perform special effects processing. The second generative model on the server provides training data to improve the accuracy and response efficiency of the model's special effects processing.
It enables rapid processing of complex special effects on mobile devices, improves the response efficiency of special effects requests and the quality of special effects images, and enhances the user experience.
Smart Images

Figure CN121364804A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to computer application technology, and in particular, to a special effect processing method and device, electronic equipment, storage medium and program product. BACKGROUND
[0002] In the scene of image processing or video production, the application of special effects is favored by users. Users can process images or videos with selected special effects to make the resulting special effect images present special effect effects corresponding to the special effects.
[0003] In related technologies, when requesting special effect processing on a mobile terminal, due to the computing power limitation of the mobile terminal, the response rate of the special effect processing request is slow, or the special effect effect of the processed special effect image is poor, thereby affecting the use experience of the special effect. SUMMARY
[0004] The present disclosure provides a special effect processing method, device, electronic equipment, storage medium and program product to improve the response efficiency of special effect requests and the presented special effect effect.
[0005] In a first aspect, the embodiments of the present disclosure provide a special effect processing method, which comprises:
[0006] In response to a special effect processing request input on a mobile terminal, obtaining a to-be-processed image;
[0007] Processing the to-be-processed image through a first generation model deployed on the mobile terminal to obtain a target special effect image, and displaying the target special effect image;
[0008] The first generation model is obtained by training a first generative adversarial network, at least part of the training data of the first generative adversarial network is generated by a second generation model deployed on a server, and the second generation model is obtained by training a second generative adversarial network.
[0009] In a second aspect, the embodiments of the present disclosure also provide a special effect processing device, which comprises:
[0010] A special effect request module configured to obtain a to-be-processed image in response to a special effect processing request input on a mobile terminal;
[0011] A special effect processing module configured to process the to-be-processed image through a first generation model deployed on the mobile terminal to obtain a target special effect image, and display the target special effect image; wherein the first generation model is obtained by training a first generative adversarial network, at least part of the training data of the first generative adversarial network is generated by a second generation model deployed on a server, and the second generation model is obtained by training a second generative adversarial network.
[0012] In a third aspect, the embodiments of the present disclosure further provide an electronic device, comprising:
[0013] one or more processors;
[0014] a storage device configured to store one or more programs,
[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the special effect processing method according to any of the embodiments of the present disclosure.
[0016] In a fourth aspect, the embodiments of the present disclosure further provide a storage medium containing computer executable instructions for performing the special effect processing method according to any of the embodiments of the present disclosure when executed by a computer processor.
[0017] In a fifth aspect, the embodiments of the present disclosure further provide a computer program product, which comprises a computer program for implementing the special effect processing method according to any of the embodiments of the present disclosure when executed by a processor.
[0018] The technical scheme of the embodiments of the present disclosure, by responding to the special effect processing request input in the mobile terminal, obtaining the to-be-processed image, obtaining the to-be-processed image in the mobile terminal, improving the acquisition efficiency of the to-be-processed image, providing a data basis for subsequent special effect processing, and by initiating the special effect processing request through the mobile terminal, the application object of the special effect can be wider. Further, by deploying the first generation model in the mobile terminal to process the to-be-processed image to obtain the target special effect image, and displaying the target special effect image, since the first generation model is obtained by training the first generative adversarial network, at least part of the training data of the first generative adversarial network is generated by deploying the second generation model in the server, and the second generation model is obtained by training the second generative adversarial network, compared with the client, the server can support more complex special effect processing, and support larger data volume training of the second generation model. The training data of the first generative adversarial network is generated by deploying the second generation model in the server, which guarantees the quality of the training data, thereby improving the special effect processing precision of the first generation model, solving the problems of slow response rate of the mobile terminal to the special effect processing request and poor special effect of the obtained special effect image in the related art, realizing the effect that the first generation model deployed in the mobile terminal can realize complex special effect, and improving the response efficiency of the special effect request, which can complete the special effect processing more quickly and present the special effect image. BRIEF DESCRIPTION OF DRAWINGS
[0019] The above-described and other features, advantages, and aspects of the present disclosure will become more apparent as various embodiments of the present disclosure are described in greater detail. It should be understood that the drawings are not to scale, and in some instances, the drawings have been simplified for the sake of clarity. Like reference numbers in different drawings represent the same or similar elements unless otherwise indicated.
[0020] Figure 1 A flowchart of a special effect processing method provided by an embodiment of the present disclosure;
[0021] Figure 2 A model structure diagram of a second generation model for a special effect processing method provided by an embodiment of the present disclosure;
[0022] Figure 3 A flowchart of another special effect processing method provided by an embodiment of the present disclosure;
[0023] Figure 4 A structure diagram of a special effect processing device provided by an embodiment of the present disclosure;
[0024] Figure 5 A structure diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] Embodiments of the present disclosure will be described in greater detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein, but rather, the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure are for illustrative purposes only and are not intended to limit the scope of the present disclosure.
[0026] It should be understood that the various steps of the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this regard.
[0027] The term "comprising" and variations thereof as used herein are open-ended, and mean "including but not limited to". The term "based on" means "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related terms are defined as follows.
[0028] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the sequence or interdependence of the functions performed by these devices, modules or units.
[0029] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that "one" or "multiple" should be understood as "one or more" unless otherwise explicitly indicated in the context.
[0030] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0031] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.
[0032] For example, in response to receiving the active request of the user, the user is sent prompt information to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the electronic device, application program, server or storage medium, etc. software or hardware performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0033] As an optional but non-limiting implementation, in response to receiving the active request of the user, the user can be sent prompt information in the form of a pop-up window, for example, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0034] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation of the present disclosure. Other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0035] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0036] Figure 1A flowchart of a special effect processing method provided by an embodiment of the present disclosure is shown. The embodiment of the present disclosure is applicable to a case of performing special effect processing on an acquired image to be processed on a mobile terminal. The method can be executed by a special effect processing apparatus. The apparatus can be implemented in the form of software and / or hardware. Optionally, the apparatus is implemented by an electronic device, which can be a mobile terminal, a PC terminal, or a server, etc.
[0037] As shown in Figure 1 , the method of the embodiment can specifically include the following steps.
[0038] S110, in response to a special effect processing request input on a mobile terminal, acquiring an image to be processed.
[0039] The special effect processing request can be understood as an instruction for requesting special effect processing on an image. The special effect processing request can include various input modes. Optionally, in a case where a triggering operation input by a user to a preset special effect processing control of the mobile terminal is detected, it can be determined that the special effect processing request input on the mobile terminal is received. Alternatively, in a case where it is detected that audio information received by the mobile terminal includes a triggering word associated with special effect processing, it can be determined that the special effect processing request input on the mobile terminal is received. Alternatively, in a case where a special effect processing instruction input by the user on the mobile terminal is detected, it can be determined that the special effect processing request input on the mobile terminal is received. The image to be processed can be an image to be processed. Optionally, the image to be processed can be a default template image, or an image acquired based on a terminal device, or an image acquired from a target storage space (such as an image library of an application software, or a local terminal album, etc.) in response to a triggering operation of the user, or an image received from an external device, etc. The terminal device can refer to a camera, a smart phone, a tablet computer, and other electronic devices with an image capturing function.
[0040] It should be noted that the mobile terminal acquiring the image to be processed can be a terminal device supporting special effect processing on an image, for example, a mobile terminal registered in an application software with a special effect processing function, or a mobile terminal registered in an application software with a special effect prop making function. The embodiment of the present disclosure does not make a specific limitation in this regard.
[0041] In the embodiment of the present disclosure, in a case where the special effect processing request input on the mobile terminal is received on the mobile terminal, the special effect processing request can be responded to. Further, the image to be processed to be processed can be acquired based on the special effect processing request.
[0042] As an optional implementation of the embodiment one of the present disclosure, in a case that a trigger operation input by a user to a preset special effect processing control of a mobile terminal is detected, it is determined that a special effect processing request input on the mobile terminal is received. Then, the special effect processing request is responded to, and a shooting interface is entered. At this time, a picture displayed in the shooting interface is a picture in a field of view area of a shooting device. Further, in a case that an image shooting operation is detected, the shooting device can be used to collect an image displayed in the current shooting interface, and the collected image is taken as a to-be-processed image.
[0043] In S120, a first generation model deployed on the mobile terminal is used to perform special effect processing on the to-be-processed image to obtain a target special effect image, and the target special effect image is displayed.
[0044] In the embodiment of the present disclosure, the first generation model can be understood as a neural network model that takes an image as an input object, performs stylization special effect processing on the image, and outputs an image with a specific stylization special effect. The first generation model can be a neural network model with any model structure. For example, the first generation model can be a generative adversarial network (GAN).
[0045] In actual application, the calculation amount required for performing special effect processing on an image is large, and the to-be-processed image is usually processed based on a special effect image generation model deployed on a server. Then, the obtained special effect image is fed back to the mobile terminal. Such a processing manner can result in a low special effect processing efficiency. In addition, in a case that the training data used to train the special effect image generation model has a low quality, the generated special effect can be unstable and poor in a case that the special effect processing is performed based on the special effect image generation model.
[0046] In view of the above, in the embodiment of the present disclosure, the first generation model can be deployed on the mobile terminal. Then, in a case that the to-be-processed image is obtained on the mobile terminal, the to-be-processed image can be directly processed based on the deployed first generation model to obtain a target special effect image. It should be noted that, in order to adapt to the computing power of the mobile terminal, the first generation model can be a lightweight neural network model. In addition, in order to improve the stability and image quality of the special effect image output by the first generation model, at least part of the training data used to train the first generation model can be generated by a second generation model deployed on a server.
[0047] In actual application, different mobile terminals can correspond to different computing powers. If the first generation model with the same model structure is deployed on different mobile terminals, the computing power of some mobile terminals can not meet the operation requirement of the first generation model, and thus the mobile terminal can be stuck in the special effect processing process.
[0048] To address the above issues, in the embodiments of the present disclosure, in order to make the mobile terminal more suitable for the deployed first generation model, the model structure of the first generation model can be associated with the target parameter of the mobile terminal. The target parameter can be used to represent the computing power of the mobile terminal. The advantage of this setting is that the computing power of the mobile terminal can meet the running requirements of the first generation model with the corresponding model structure, thereby realizing the effect of improving the image processing efficiency on the basis of ensuring the special effect image generation effect, and improving the special effect processing experience of the mobile terminal user. The target parameter can include various parameters that can represent the computing power of the mobile terminal. Optionally, the target parameter at least includes the image resolution corresponding to the mobile terminal and / or the number of floating-point operations performed per second. The model structure of the first generation model deployed by the mobile terminal with different image resolutions can be different. Generally, the image resolution corresponding to the mobile terminal can include multiple levels, and the mobile terminal with the same image resolution level can deploy the first generation model with the same model structure.
[0049] For example, it is assumed that the image resolution level corresponding to the mobile terminal can include 512x512 pixels, 384x384 pixels and 256x256 pixels. The mobile terminal with an image resolution level of 512x512 pixels can deploy the first generation model with the same model structure; the mobile terminal with an image resolution level of 384x384 pixels can deploy the first generation model with the same model structure; and the mobile terminal with an image resolution level of 256x256 pixels can deploy the first generation model with the same model structure. It can be understood that in the case where there are other factors representing the computing power of the mobile terminal (such as the number of floating-point operations performed per second), the mobile terminal with the same image resolution level can also deploy the first generation model with different model structures, and the embodiments of the present disclosure do not make specific limitations thereto.
[0050] The number of floating-point operations performed per second (FLOPS) is a measure of how many floating-point operations a computer or processor can complete per second. The model structure of the first generation model deployed by the mobile terminal with different FLOPS can be different. Generally, the FLOPS corresponding to the mobile terminal can include multiple levels, and the mobile terminal with the same FLOPS can deploy the first generation model with the same model structure. Alternatively, the mobile terminal with the same FLOPS can also deploy the first generation model with different model structures.
[0051] In the embodiments of the present disclosure, the first generation model is obtained by training a first generative adversarial network, and at least part of the training data of the first generative adversarial network is generated by a second generation model deployed on a server.
[0052] As an optional implementation of the embodiment one of the present disclosure, before the first generation model deployed on the mobile terminal is used to perform special effect processing on the to-be-processed image, the first generative adversarial network can be trained in advance. The training process of the first generation model can be: obtaining a second sample image; performing special effect processing on the second sample image based on the second generation model to obtain a second special effect image, and constructing training data for training the first generative adversarial network based on the second sample image and the corresponding second special effect image; training the first generative adversarial network based on the training data to obtain the first generation model. Wherein, training the first generative adversarial network based on the training data to obtain the first generation model, comprising: inputting the second sample image into the first generative adversarial network to obtain a second actual output image; determining a loss value based on the second actual output image and the second special effect image corresponding to the second sample image; modifying the model parameters in the first generative adversarial network based on the loss value, and converging the loss function in the first generative adversarial network as a training target to obtain the first generation model.
[0053] It should be noted that, for the deployment of the first generation model on the mobile terminal, the first generation model can be trained on the server, and then the obtained first generation model is deployed on the corresponding mobile terminal according to the target parameters of the mobile terminal. Alternatively, for at least one mobile terminal to be deployed, the model structure of the first generative adversarial network to be trained can be determined according to the target parameters corresponding to the mobile terminal. Further, the first generative adversarial network to be trained can be deployed on the mobile terminal. Then, the training data can be obtained based on the mobile terminal, and the first generative adversarial network is trained based on the training data on the mobile terminal to obtain the first generation model.
[0054] In the embodiment of the present disclosure, in the case of obtaining the first generation model, the to-be-processed image can be input into the first generation model. Further, the to-be-processed image can be processed by the first generation model, and the image obtained after the special effect processing can be used as a target special effect image, and the target special effect image can be displayed. Wherein, the target special effect image can be an image that meets the special effect processing requirements and conforms to the expected special effect. It should be noted that the first generation model and the special effect added in the target special effect image are one-to-one corresponding, that is, the first generation model corresponding to any special effect can only output the target special effect image corresponding to the special effect.
[0055] In the embodiment of the present disclosure, the second generation model can be understood as a neural network model deployed on the server, which takes an image as an input object, performs stylized special effect processing on the image, and outputs an image with a specific stylized special effect. The second generation model can be a neural network model with any model structure. Optionally, the second generation model can include a convolution module and a generator in a third generative adversarial network based on style connected with the convolution module.
[0056] As Figure 2 shown, the convolution module 21 can be a neural network model constructed by at least one convolution layer. The convolution module can serve as an encoder in the second generation model to extract features of the input image. It can be understood that the convolution module 22 can be used to extract local image features of the image, which have the characteristics of rich quantity of information contained in the image, small correlation between features, and no influence on detection and matching of other features due to disappearance of part of the features in the case of occlusion, etc. Moreover, in the second generation model, the convolution module 22 is set as an encoder, which has the advantages of: the input image can be extracted to eliminate redundant image features, thereby improving the processing efficiency of the second generation model and enhancing the image effect of the special effect image.
[0057] Continuing to use Figure 2 , the generator 22 of the style-based third generative adversarial network can include a mapping module and a synthesis module. The mapping module can be specifically used to map the input random noise vector and the image feature vector output by the convolution module to a new latent space through a multi-layer fully connected layer (i.e., a multi-layer perceptron), thereby generating a new latent feature vector. By adding the image features output by the convolution module, the content features in the synthesized image in the synthesis module can be better controlled by the latent feature vector. The mapping module can realize feature decoupling of the random noise vector, so that each dimension or each subspace can control different features of the image. The synthesis module includes multiple resolution levels, starting from very low resolution and gradually increasing to high resolution. Each resolution level is usually composed of multiple style blocks, and each style block includes a convolution layer, a normalization layer and an activation function. The synthesis module is used to gradually generate high-resolution images according to the latent feature vector output by the mapping module. By applying the normalization layer at each resolution level, the latent feature vector is injected into the generation process, thereby controlling the style and features of the image. The style-based third generative adversarial network can enable the generator 22 to control and separate content and style, thereby realizing more fine and diversified image generation. The content vector determines the theme or basic information of the generated image, such as the outline of the object, facial features, etc.; the style vector can control the texture and style of the details. In the embodiments of the present disclosure, the generator 22 of the style-based third generative adversarial network is used as a decoder in the second generation model, which can realize control of the special effect style, so that the second generation model can generate special effect images in accordance with the expected special effect style.
[0058] As an optional implementation of the first embodiment of the present disclosure, in the case of obtaining the second sample image, the second sample image can be input to the second generation model. Further, the second sample image is encoded based on the convolution module in the second generation model to obtain the second sample feature. Further, the second sample feature and the second style vector corresponding to the second sample image can be input to the generator of the third style-based generative adversarial network, so as to generate the special effect image based on the generator according to the second sample feature and the second style vector, to obtain the second special effect image corresponding to the second sample image.
[0059] In the embodiments of the present disclosure, in order to adapt the second generation model to the needs of multiple special effects, that is, to enable the second generation model to generate special effect images of multiple different special effects, the output channels of the second generation model can be processed. Optionally, the second generation model includes multiple output channels. The number of output channels included in the second generation model can be associated with the model input. Optionally, in the case of image as the model input of the second generation model, the second generation model can include six output channels, which are image texture channels (including R channel, G channel and B channel), image mask layer channel and pixel displacement channel (including pixel horizontal displacement channel and pixel vertical displacement channel). The output data of the output channel at least includes texture data. Wherein, the texture data (RGB data) can be data representing image texture details. The output data of the output channel also includes image mask data and / or pixel displacement data. The image mask data can be used to indicate the special effect area in the output image. The image mask data can be a matrix or array with the same size as the original image, and the element value therein can be used to identify whether the pixel at the corresponding position needs to be processed. The pixel displacement data can be used to represent the displacement of each pixel point included in the image, that is, the displacement of each pixel point after model processing. It should be noted that the second generation model adopts multiple output channels, and the output data of the output channel includes texture data, image mask data and / or pixel displacement data. The advantage is that by adopting multiple output channels, whether full-image special effect processing or local special effect processing is based on the second generation model, the output data of the model output is adapted to the effect of full-image special effect or local special effect.
[0060] In the embodiments of the present disclosure, the second generation model is obtained by training a second generative adversarial network. Before the second generation model is applied to generate the training data for training the first generation model, the pre-established second generative adversarial network can be trained to obtain the second generation model. The specific training process can be: obtaining a first sample image, and processing the first sample image into a first special effect image; training the second generative adversarial network deployed on the server based on the first sample image and the first special effect image corresponding to the first sample image to obtain the second generation model. Wherein, training the second generative adversarial network deployed on the server based on the first sample image and the first special effect image corresponding to the first sample image to obtain the second generation model comprises: inputting the first sample image into the second generative adversarial network to process the first sample image based on the convolution module and the generator in the second generative adversarial network to obtain a first actual output image; determining a loss value based on the first actual output image and the first special effect image corresponding to the first sample image; modifying the model parameters in the second generative adversarial network based on the loss value, and converging the loss function in the second generative adversarial network as a training target to obtain the second generation model.
[0061] In the embodiments, the model loss of the second generation model is determined based on a generation loss of the second generation model and a discrimination loss of the discriminator in the first generative adversarial network; and the generation loss of the second generation model includes a perceptual loss and / or a semantic loss between the output image of the second generation model and the input image. The discrimination loss can be used to measure the prediction accuracy of the discriminator, and can also be used to measure the learning ability of the second generation model for the special effect image. The discrimination loss can be any loss function. Optionally, the discrimination loss can be a cross-entropy loss. The visual perceptual loss can be used to evaluate the perceptual similarity between two images. Generally, the perceptual loss can be used to evaluate the perceptual difference between two images through a deep learning model (such as a discriminant model), which is more capable of reflecting the perception of the visual system on the image quality than the traditional pixel-level loss. The perceptual loss extracts high-level features of the image through a pre-trained deep network, which can capture high-level attributes such as texture and shape of the image, so as to more accurately evaluate the perceptual similarity between images. Then, the perceptual loss calculates the distance between these features to quantify the perceptual difference between images.
[0062] There are various ways to determine the semantic loss between the output image and the input image. As an optional implementation manner of the embodiments of the present disclosure, the input image and the output image can be first segmented by a pre-trained semantic segmentation model, and then the semantic loss between the input image and the output image is determined according to the semantic segmentation results of the input image and the output image.
[0063] The technical scheme of the embodiments of the present disclosure is that, in response to a special effect processing request input on a mobile terminal, a to-be-processed image is acquired, the to-be-processed image is acquired on the mobile terminal, the acquisition efficiency of the to-be-processed image is improved, a data basis is provided for subsequent special effect processing, and the applicable object of the special effect can be made wider by initiating the special effect processing request on the mobile terminal. Further, the to-be-processed image is processed by a first generation model deployed on the mobile terminal to obtain a target special effect image, and the target special effect image is displayed. Since the first generation model is obtained by training a first generative adversarial network, at least part of the training data of the first generative adversarial network is generated by a second generation model deployed on a server, the second generation model is obtained by training a second generative adversarial network, compared with the client, the server can support more complex special effect processing, and support training of the second generation model with a larger amount of data. The training data of the first generative adversarial network is generated by the second generation model deployed on the server, the quality of the training data is ensured, so that the special effect processing precision of the first generation model is improved, the problems of slow response rate of the mobile terminal to the special effect processing request and poor special effect of the obtained special effect image in the related art are solved, the first generation model deployed on the mobile terminal can realize complex special effect, the response efficiency of the special effect request is improved, and the special effect processing can be completed more quickly, and the special effect image is presented.
[0064] Figure 3 The flowchart of another special effect processing method provided by the embodiments of the present disclosure is shown. The technical scheme of the present embodiment is based on the above-mentioned embodiments, and before the to-be-processed image is processed by the first generation model deployed on the mobile terminal, the second sample image is acquired, the second sample image is processed by the second generation model to obtain the second special effect image, and the first generative adversarial network is trained based on the second sample image and the second special effect image to obtain the first generation model. The specific implementation can be referred to the description of the present embodiment. The same or similar technical features as the foregoing embodiments are not described herein.
[0065] As shown in Figure 3 , the method of the present embodiment can specifically include:
[0066] S310, in response to a special effect processing request input on a mobile terminal, a to-be-processed image is acquired.
[0067] S320, a second sample image is acquired, a second generation model deployed on a server is used to process the second sample image to obtain a second special effect image.
[0068] The second sample image can be understood as an image captured by a camera, or an image reconstructed by an image reconstruction model, or an image pre-stored in a storage space. The second special effect image can be a special effect image generated by performing a special effect on the second sample image based on a second generation model.
[0069] In the embodiments of the present disclosure, the second generation model is obtained by training a second generative adversarial network. Before performing a special effect on the second sample image by using the second generation model, the method further includes: generating a first sample image by a third generation model deployed on a server, and processing the first sample image into a first special effect image; and training the second generative adversarial network deployed on the server by using the first sample image and the corresponding first special effect image, to obtain the second generation model.
[0070] The third generation model can be understood as a neural network model that takes an image as an input object, performs a stylized special effect on the image, and outputs an image with a specific stylized special effect. The third generation model can be a neural network model with any model structure. For example, the first generation model can be a Stable Diffusion model. Stable Diffusion is a diffusion process-based image generation model that can generate high-quality, high-resolution images, and is a relatively new diffusion model. The core idea of Stable Diffusion is to gradually approach a real image by continuously adjusting the implicit representation of the image. In the embodiments of the present disclosure, the third generation model is obtained by training a diffusion model.
[0071] In the embodiments of the present disclosure, generating a first sample image by a third generation model deployed on a server can include various implementation manners. Optionally, a text feature vector and a noise vector can be input into the third generation model, to process the text feature vector and the noise vector based on the third generation model, and output an image of the model as the first sample image. Alternatively, an image feature vector and a noise vector can be input into the third generation model, to process the image feature vector and the noise vector based on the third generation model, and output an image of the model as the first sample image. Alternatively, an audio feature vector and a noise vector can be input into the third generation model, to process the audio feature vector and the noise vector based on the third generation model, and output an image of the model as the first sample image.
[0072] In the embodiments of the present disclosure, in the case where the first sample image is obtained, the first sample image can be processed according to a preset special effect processing manner, and a first special effect image is obtained.
[0073] In actual application, different first sample images can include different image contents, and the same special effect processing manner is used to perform special effect processing on different first sample images, which can cause the special effect probability distribution of the obtained special effect images to be inconsistent, and further, using such training data to train the second generative adversarial network can cause the model processing effect of the trained second generative model to be unstable and difficult to converge.
[0074] In view of the above, in the embodiments of the present disclosure, for first sample images of different image contents, different special effect processing manners can be used to perform special effect processing on the corresponding first sample images to obtain first special effect images.
[0075] Optionally, processing the first sample image into the first special effect image includes: determining an image category corresponding to the first sample image; determining a special effect processing manner according to the image category, and processing the first sample image into the first special effect image using the special effect processing manner.
[0076] The image category is associated with attribute data of the image content in the first sample image to be processed. The attribute data can be understood as data used to represent specific features of the image content. For example, assuming that the image content to be processed is the face of an object in the first sample image, the corresponding attribute data can be data representing the features of the face in different attribute dimensions; assuming that the image content to be processed is all image contents in the first sample image, the corresponding attribute data can be data representing the first sample image in different attribute dimensions. Different image categories can correspond to different special effect processing manners. The special effect processing manner can include multiple ways of performing special effect processing on images. Optionally, the special effect processing manner can include a special effect processing algorithm and / or special effect processing based on a neural network model.
[0077] As an optional implementation manner of the embodiments of the present disclosure, a plurality of image categories and a special effect processing manner corresponding to each image category can be determined in advance, and a mapping relationship between the image categories and the corresponding special effect processing manners can be established. Further, when the first sample image is obtained, the image content to be processed in the first sample image can be determined. Further, the image category corresponding to the first sample image can be determined according to the attribute data of the image content. Further, the special effect processing manner corresponding to the image category can be determined according to the mapping relationship established in advance. Further, the first sample image can be processed using the special effect processing manner to obtain the first special effect image.
[0078] In the embodiments of the present disclosure, when the first sample image and the corresponding first special effect image are obtained, the first sample image and the corresponding first special effect image can be used to train the second generative adversarial network deployed on the server.
[0079] It should be noted that in order to enable sufficient training data for subsequent model training of the first generation model, the third generation model usually generates a large number of first sample images, and then a large number of first special effect images can be obtained. Among these large numbers of first sample images and first special effect images, there may be generated first sample images with low image quality and / or obtained first special effect images with low image quality. In addition, the data volume requirement of the training data of the second generation model is reduced. Therefore, in order to improve the model processing effect and model processing stability of the second generation model, the first sample images and / or the first special effect images can be screened to exclude images with low image quality. Then, the second generative adversarial network is trained based on the screened first sample images and first special effect images to obtain the second generation model.
[0080] Optionally, the first sample image and the first special effect image corresponding to the first sample image are used to train the second generative adversarial network deployed on the server, including: screening the first sample image and / or the first special effect image, determining a plurality of groups of paired data according to the screening result, and taking the plurality of groups of paired data as training data to train the second generative adversarial network deployed on the server.
[0081] As an optional implementation manner of the embodiment one of the present disclosure, the first sample image and / or the first special effect image is screened, and a plurality of groups of paired data are determined according to the screening result, including: screening the first sample image according to a preset first screening manner, and screening the first special effect image corresponding to the screened first sample image; and constructing a plurality of groups of paired data according to the remaining first sample image and the first special effect image corresponding thereto.
[0082] Optionally, the first screening manner can include manual screening and / or automatic screening based on at least one preset first image quality indicator. The first image quality indicator can be used to evaluate the image quality of the first sample image. Optionally, the first image quality indicator can include image clarity, image content fit degree, and / or image resolution. For example, the first sample image can be screened according to at least one first image quality indicator, and the first sample image that does not meet the standard value corresponding to the first image quality indicator is screened out.
[0083] As another optional implementation manner of the embodiment of the present disclosure, the first sample image and / or the first special effect image is screened, and a plurality of groups of paired data are determined according to the screening result, including: screening the first special effect image according to a preset second screening manner, and screening the first sample image corresponding to the screened first special effect image; and constructing a plurality of groups of paired data according to the remaining first special effect image and the first sample image corresponding thereto.
[0084] Optionally, the second screening manner can include manual screening and / or automatic screening based on at least one preset second image quality index. The second image quality index can be used to evaluate the image quality of the first special effect image. Optionally, the second image quality index can include image definition, image resolution, peak signal-to-noise ratio, mean square error, and the like. For example, the first special effect image can be screened according to the at least one second image quality index, and the first special effect image that does not meet the standard value corresponding to the second image quality index is excluded.
[0085] As another optional implementation of the embodiments of the present disclosure, the first sample image and / or the first special effect image are screened, and a plurality of sets of pairing data are determined according to the screening result, including: screening the first sample image according to a preset first screening manner, and excluding the first special effect image corresponding to the screened first sample image; screening the first special effect image after screening according to a preset second screening manner, and excluding the first sample image corresponding to the screened first special effect image; and constructing a plurality of sets of pairing data according to the first special effect image and the first sample image corresponding thereto after screening. The pairing data is training data containing the first sample image and the first special effect image corresponding to the first sample image.
[0086] In the embodiments of the present disclosure, after obtaining the plurality of sets of pairing data, the plurality of sets of pairing data can be used as training data, and the second generative adversarial network deployed on the server is trained based on the training data to obtain a second generation model. Further, the second sample image can be obtained, and the second sample image is input into the second generation model to perform special effect processing on the second sample image based on the second generation model, and output a second special effect image.
[0087] S330, training the first generative adversarial network based on the second sample image and the second special effect image to obtain a first generation model.
[0088] In the embodiments of the present disclosure, the second sample image can be input into the first generative adversarial network to perform special effect processing on the second sample image based on the first generative adversarial network, and obtain a second actual output image. Further, the loss value can be obtained based on the second actual output image and the second special effect image corresponding to the second sample image; the model parameters in the first generative adversarial network are corrected based on the loss value, and the loss function in the first generative adversarial network is converged as a training target to obtain a first generation model. It should be noted that the training of the first generative adversarial network can be performed on the server or on the client, and the specific manner can be set according to the requirements, which is not limited here.
[0089] S340, performing special effect processing on the to-be-processed image by the first generation model deployed on the mobile terminal to obtain a target special effect image, and displaying the target special effect image.
[0090] The technical scheme of the embodiment of the present disclosure obtains a second sample image, performs special effect processing on the second sample image by using a second generation model to obtain a second special effect image, and further trains a first generative adversarial network based on the second sample image and the second special effect image to obtain the first generation model, thereby achieving the effect of generating training data for training the first adversarial network based on the second generation model deployed on the server, improving the generation efficiency of the training data, and improving the data distribution consistency of the training data, and further improving the model performance of the first generation model.
[0091] Figure 4 A structural schematic diagram of a special effect processing device provided by the embodiment of the present disclosure is shown in Figure 4 The device includes a special effect request module 410 and a special effect processing module 420.
[0092] The special effect request module 410 is configured to obtain a to-be-processed image in response to a special effect processing request input on a mobile terminal. The special effect processing module 420 is configured to perform special effect processing on the to-be-processed image by a first generation model deployed on the mobile terminal to obtain a target special effect image, and display the target special effect image. The first generation model is obtained by training a first generative adversarial network. At least part of the training data of the first generative adversarial network is generated by a second generation model deployed on a server. The second generation model is obtained by training a second generative adversarial network.
[0093] On the basis of the above-mentioned optional technical solutions, the device further includes a sample image acquisition module and a first generation model determination module. The sample image acquisition module is configured to obtain a second sample image before performing special effect processing on the to-be-processed image by the first generation model deployed on the mobile terminal, and perform special effect processing on the second sample image by using the second generation model to obtain a second special effect image. The first generation model determination module is configured to train a first generative adversarial network based on the second sample image and the second special effect image to obtain the first generation model.
[0094] On the basis of each of the optional technical solutions described above, the device further comprises an image generation module and a second generation model determination module. The image generation module is configured to generate a first sample image by using a third generation model deployed on a server before the second sample image is processed by using the second generation model, and process the first sample image into a first special effect image; the third generation model is obtained by training a diffusion model; and the second generation model determination module is configured to train a second generative adversarial network deployed on the server by using the first sample image and the corresponding first special effect image, so as to obtain the second generation model.
[0095] On the basis of each of the optional technical solutions described above, the image generation module comprises an image category determination unit and an image processing unit. The image category determination unit is configured to determine an image category corresponding to the first sample image; the image category is associated with attribute data of image content to be processed by using a special effect in the first sample image; and the image processing unit is configured to determine a special effect processing mode according to the image category, and process the first sample image into the first special effect image by using the special effect processing mode.
[0096] On the basis of each of the optional technical solutions described above, the second generation model determination module is specifically configured to filter the first sample image and / or the first special effect image, determine a plurality of sets of paired data according to a filtering result, use the plurality of sets of paired data as training data, and train the second generative adversarial network deployed on the server; the paired data comprises the first sample image and the first special effect image corresponding to the first sample image.
[0097] On the basis of each of the optional technical solutions described above, the second generation model comprises a convolution module and a generator in a style-based third generative adversarial network connected to the convolution module.
[0098] On the basis of each of the optional technical solutions described above, the second generation model comprises a plurality of output channels; output data of the output channels at least comprises texture data; and the output data of the output channels further comprises image mask data and / or pixel displacement data.
[0099] On the basis of each of the optional technical solutions described above, a model loss of the second generation model is determined based on a generation loss of the second generation model and a discrimination loss of a discriminator in the first generative adversarial network; and the generation loss of the second generation model comprises a perceptual loss and / or a semantic loss between an output image of the second generation model and an input image.
[0100] On the basis of each optional technical scheme above, optionally, a model structure of the first generation model is associated with a target parameter of the mobile terminal; the target parameter is used to represent the computing power of the mobile terminal; and the target parameter at least includes an image resolution corresponding to the mobile terminal and / or a number of floating point operations performed per second.
[0101] The technical scheme of the embodiment of the disclosure, through the special effect request module 410, responds to the special effect processing request input in the mobile terminal, acquires the to-be-processed image, acquires the to-be-processed image in the mobile terminal, improves the acquisition efficiency of the to-be-processed image, provides a data basis for subsequent special effect processing, and through the mobile terminal initiating the special effect processing request, can also make the application object of the special effect wider; further, based on the special effect processing module 420, the first generation model deployed in the mobile terminal is used to perform special effect processing on the to-be-processed image to obtain a target special effect image, and the target special effect image is displayed. Since the first generation model is obtained by training the first generative adversarial network, at least part of the training data of the first generative adversarial network is generated by the second generation model deployed in the server, and the second generation model is obtained by training the second generative adversarial network. Compared with the client, the server can support more complex special effect processing, and support training of a larger amount of data for the second generation model. The training data of the first generative adversarial network is generated by the second generation model deployed in the server, which guarantees the quality of the training data, thereby improving the special effect processing precision of the first generation model, solving the problems of slow response rate of the mobile terminal to the special effect processing request and poor special effect of the obtained special effect image in the related art, realizing the effect that the first generation model deployed in the mobile terminal can realize complex special effect, and improving the response efficiency of the special effect request, which can complete the special effect processing more quickly and present the special effect image.
[0102] The special effect processing apparatus provided in the embodiments of the disclosure can execute the special effect processing method provided in any of the embodiments of the disclosure, and has the function modules and beneficial effects corresponding to the execution method.
[0103] It should be noted that each unit and module included in the apparatus is only divided according to the function logic, but is not limited to the division described above, as long as the corresponding function can be implemented; in addition, the specific names of each functional unit are only for convenient distinction, and do not limit the protection scope of the embodiments of the disclosure.
[0104] Figure 5 A structural schematic diagram of an electronic device provided in the embodiments of the disclosure. The following refers to Figure 5 which shows an electronic device (for example Figure 5FIG. 1 is a structural diagram of an electronic device 500 according to an embodiment of the disclosure. The electronic device 500 can include a terminal device or a server. The electronic device 500 can include a communication unit 510, a processor 520, a memory 530, a user interface 540, and a display unit 550. The electronic device 500 can include at least one of the communication unit 510, the processor 520, the memory 530, the user interface 540, and the display unit 550. The electronic device 500 can include a terminal device or a server. The terminal device can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (for example, a car navigation terminal), and the like, and a stationary terminal such as a digital TV, a desktop computer, and the like. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functions and the range of use of the embodiments of the disclosure.
[0105] As shown in FIG. 1, the electronic device 500 can include a communication unit 510, a processor 520, a memory 530, a user interface 540, and a display unit 550. The electronic device 500 can include at least one of the communication unit 510, the processor 520, the memory 530, the user interface 540, and the display unit 550. Figure 5 The processor 520 can include a central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), a programmable logic unit (PLU), a field programmable gate array (FPGA), a microprocessor, a microcomputer, or the like. The processor 520 can perform data processing and / or control operation of the electronic device 500.
[0106] Generally, the following devices can be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 508 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or via a wire to exchange data. Although not shown, the electronic device 500 can further include a power supply device (for example, a battery, a solar cell, and the like) to supply power to the electronic device 500. Figure 5 The electronic device 500 having various devices is shown, but it is understood that all of the shown devices are not required to be implemented or provided. More or less devices can be alternatively implemented or provided.
[0107] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the disclosure. For example, the embodiments of the disclosure include a computer program product including a computer program carried on a non-transitory computer-readable medium, the computer program containing program codes for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processor 520, the above-described functions defined in the methods of the embodiments of the disclosure are performed.
[0108] Names of messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.
[0109] The electronic device provided by the embodiments of the present disclosure and the special effect processing method provided by the above embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiment can be referred to the above embodiments, and the present embodiment has the same beneficial effects as the above embodiments.
[0110] The embodiments of the present disclosure provide a computer storage medium, which stores a computer program, and the program is executed by a processor to implement the special effect processing method provided by the above embodiments.
[0111] It should be noted that the computer readable medium of the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or component. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.
[0112] In some embodiments, the mobile terminal and the server can communicate using any currently known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication (e.g., a communication network) of any form or medium (e.g., the Internet). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.
[0113] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled into the electronic device.
[0114] The computer readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to: in response to a special effect processing request input on the mobile terminal, acquire a to-be-processed image; perform special effect processing on the to-be-processed image by a first generative model deployed on the mobile terminal, to obtain a target special effect image, and display the target special effect image; wherein the first generative model is obtained by training a first generative adversarial network, at least part of training data of the first generative adversarial network is generated by a second generative model deployed on a server, and the second generative model is obtained by training a second generative adversarial network.
[0115] Computer program code for carrying out operations of the present disclosure can be written in any one or more programming languages, including object oriented programming languages such as Java, Smalltalk, C++ or the like, conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0116] The computer program product of the first aspect can include one or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform the operations of the first aspect. The one or more non-transitory computer-readable media can include, for example, magnetic media such as one or more magnetic disks, magnetic tapes or cassettes; optical media such as one or more compact discs, optical discs or Blu-ray discs; magneto-optical media such as one or more floptical discs; solid state media such as one or more solid state drives or other flash memory arrays; or any suitable combination of these. The one or more non-transitory computer-readable media can be encoded with instructions that, when executed, cause one or more processors to perform the operations of the first aspect.
[0117] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware, or by a combination of software and hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses.
[0118] The functions described in this document can be implemented in part or in whole using one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0119] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0120] According to one or more embodiments of the present disclosure, Example One provides a special effect processing method, comprising: in response to a special effect processing request input on a mobile terminal, obtaining a to-be-processed image; performing special effect processing on the to-be-processed image by a first generation model deployed on the mobile terminal to obtain a target special effect image, and displaying the target special effect image; wherein the first generation model is obtained by training a first generative adversarial network, at least part of training data of the first generative adversarial network is generated by a second generation model deployed on a server, and the second generation model is obtained by training a second generative adversarial network.
[0121] According to one or more embodiments of the present disclosure, Example Two provides the method of Example One, further comprising:
[0122] Optionally, before the special effect processing on the to-be-processed image by the first generation model deployed on the mobile terminal, further comprising: obtaining a second sample image, performing special effect processing on the second sample image by the second generation model to obtain a second special effect image; and training a first generative adversarial network based on the second sample image and the second special effect image to obtain the first generation model.
[0123] According to one or more embodiments of the present disclosure, Example Three provides the method of Example Two, further comprising:
[0124] Optionally, before the special effect processing on the second sample image by the second generation model, further comprising: generating a first sample image by a third generation model deployed on the server, and processing the first sample image into a first special effect image; wherein the third generation model is obtained by training a diffusion model; and training a second generative adversarial network deployed on the server by using the first sample image and the corresponding first special effect image to obtain the second generation model.
[0125] According to one or more embodiments of the present disclosure, Example Four provides the method of Example Three, further comprising:
[0126] Optionally, the processing of the first sample image into a first special effect image comprises: determining an image category corresponding to the first sample image; wherein the image category is associated with attribute data of image content to be processed in the first sample image; determining a special effect processing mode according to the image category, and processing the first sample image into a first special effect image by using the special effect processing mode.
[0127] According to one or more embodiments of the present disclosure, Example Five provides the method of Example Three, further comprising:
[0128] Optionally, the training of the second generative adversarial network deployed on the server side by using the first sample image and the first special effect image corresponding to the first sample image comprises: screening the first sample image and / or the first special effect image, determining a plurality of sets of paired data according to a screening result, taking the plurality of sets of paired data as training data, and training the second generative adversarial network deployed on the server side; wherein the paired data is training data containing the first sample image and the first special effect image corresponding to the first sample image.
[0129] According to one or more embodiments of the present disclosure, Example Six provides the method of Example One, further comprising:
[0130] Optionally, the second generation model comprises a convolution module and a generator in a style-based third generative adversarial network connected with the convolution module.
[0131] According to one or more embodiments of the present disclosure, Example Seven provides the method of Example One, further comprising:
[0132] Optionally, the second generation model comprises a plurality of output channels; the output data of the output channels at least comprises texture data; and the output data of the output channels further comprises image mask data and / or pixel displacement data.
[0133] According to one or more embodiments of the present disclosure, Example Eight provides the method of Example One, further comprising:
[0134] Optionally, the model loss of the second generation model is determined based on a generation loss of the second generation model and a discrimination loss of a discriminator in the first generative adversarial network; and the generation loss of the second generation model comprises a perceptual loss and / or a semantic loss between an output image of the second generation model and an input image.
[0135] According to one or more embodiments of the present disclosure, Example Nine provides the method of Example One, further comprising:
[0136] Optionally, the model structure of the first generation model is associated with a target parameter of the mobile terminal; the target parameter is used to represent the computing power of the mobile terminal; and the target parameter at least comprises an image resolution corresponding to the mobile terminal and / or a number of floating point operations performed per second.
[0137] According to one or more embodiments of the present disclosure, Example Ten provides a special effect processing apparatus, comprising: a special effect request module configured to acquire a to-be-processed image in response to a special effect processing request input on a mobile terminal; a special effect processing module configured to perform special effect processing on the to-be-processed image by a first generative model deployed on the mobile terminal to obtain a target special effect image, and display the target special effect image; wherein the first generative model is obtained by training a first generative adversarial network, and at least part of training data of the first generative adversarial network is generated by a second generative model deployed on a server, and the second generative model is obtained by training a second generative adversarial network.
[0138] The above description is merely that of preferred embodiments of the present disclosure and a description of the principles of the technology used. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0139] In addition, although each operation is depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.
[0140] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A special effects processing method, characterized in that, include: In response to a special effects processing request input on the mobile device, obtain the image to be processed; The image to be processed is processed with special effects by a first generation model deployed on the mobile device to obtain a target special effects image, and the target special effects image is displayed. The first generative model is obtained by training a first generative adversarial network, and at least a portion of the training data of the first generative adversarial network is generated by a second generative model deployed on a server, which is obtained by training a second generative adversarial network.
2. The special effects processing method according to claim 1, characterized in that, Before applying special effects to the image to be processed using the first generation model deployed on the mobile device, the method further includes: A second sample image is obtained, and the second generation model is used to apply special effects to the second sample image to obtain a second special effects image. The first generative adversarial network is trained based on the second sample image and the second special effects image to obtain the first generative model.
3. The special effects processing method according to claim 2, characterized in that, Before applying the second generative model to the second sample image for special effects processing, the method further includes: A first sample image is generated by a third generation model deployed on the server, and the first sample image is processed into a first special effects image; wherein, the third generation model is obtained by training a diffusion model; The second generative adversarial network deployed on the server is trained using the first sample image and its corresponding first special effect image to obtain the second generative model.
4. The special effects processing method according to claim 3, characterized in that, The step of processing the first sample image into a first special effects image includes: Determine the image category corresponding to the first sample image; wherein, the image category is associated with the attribute data of the image content to be processed by special effects in the first sample image; The special effects processing method is determined according to the image category, and the first sample image is processed into a first special effects image using the special effects processing method.
5. The special effects processing method according to claim 3, characterized in that, The step of training a second generative adversarial network deployed on a server using the first sample image and the corresponding first special effects image includes: The first sample image and / or the first special effects image are filtered, and multiple sets of paired data are determined based on the filtering results. The multiple sets of paired data are used as training data to train the second generative adversarial network deployed on the server. The paired data is training data containing the first sample image and the first special effects image corresponding to the first sample image.
6. The special effects processing method according to claim 1, characterized in that, The second generative model includes a convolutional module and a generator in a style-based third generative adversarial network connected to the convolutional module.
7. The special effects processing method according to claim 1, characterized in that, The second generation model includes multiple output channels; the output data of the output channels includes at least texture data; the output data of the output channels also includes image mask data and / or pixel displacement data.
8. The special effects processing method according to claim 1, characterized in that, The model loss of the second generative model is determined based on the generation loss of the second generative model and the discriminant loss of the discriminator in the first generative adversarial network; the generation loss of the second generative model includes the perceptual loss and / or semantic loss between the output image and the input image of the second generative model.
9. The special effects processing method according to claim 1, characterized in that, The model structure of the first generative model is associated with the target parameters of the mobile device; the target parameters are used to characterize the computing power of the mobile device; the target parameters include at least the image resolution corresponding to the mobile device and / or the number of floating-point operations performed per second.
10. A special effects processing device, characterized in that, include: The special effects request module is used to respond to special effects processing requests entered on the mobile device and obtain the image to be processed; The special effects processing module is used to perform special effects processing on the image to be processed by a first generative model deployed on the mobile terminal to obtain a target special effects image and display the target special effects image; wherein, the first generative model is obtained by training a first generative adversarial network, at least part of the training data of the first generative adversarial network is generated by a second generative model deployed on the server, and the second generative model is obtained by training a second generative adversarial network.
11. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the special effects processing method as described in any one of claims 1-9.
12. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the special effects processing method as described in any one of claims 1-9.
13. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the special effects processing method according to any one of claims 1-9.